ai safety
- AdvPrefix: An Objective for Nuanced LLM Jailbreaks
- From Noise to Narrative: Tracing the Origins of Hallucinations in Transformers
- LLMs Encode Harmfulness and Refusal Separately
- Language Models Are Capable of Metacognitive Monitoring and Control of Their Internal Activations
- Position: Require Frontier AI Labs To Release Small "Analog" Models
- Who Speaks for the Trigger? Dynamic Expert Routing in Backdoored Mixture-of-Experts Transformers