faithfulness
- FlexAC: Towards Flexible Control of Associative Reasoning in Multimodal Large Language Models
- Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models
- OVERT: A Benchmark for Over-Refusal Evaluation on Text-to-Image Models
- One SPACE to Rule Them All: Jointly Mitigating Factuality and Faithfulness Hallucinations in LLMs
- OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and Metrics
- Rethinking Circuit Completeness in Language Models: AND, OR, and ADDER Gates
- Transformer Key-Value Memories Are Nearly as Interpretable as Sparse Autoencoders