benchmark construction
- Can Agent Fix Agent Issues?
- Diagnosing and Addressing Pitfalls in KG-RAG Datasets: Toward More Reliable Benchmarking
- SECODEPLT: A Unified Benchmark for Evaluating the Security Risks and Capabilities of Code GenAI
- Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
- Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding