needle-in-a-haystack
- MIR-Bench: Can Your LLM Recognize Complicated Patterns via Many-Shot In-Context Reasoning?
- NeedleInATable: Exploring Long-Context Capability of Large Language Models towards Long-Structured Tables
- The Fragile Truth of Saliency: Improving LLM Input Attribution via Attention Bias Optimization
- VideoGameQA-Bench: Evaluating Vision-Language Models for Video Game Quality Assurance