model capability
- CTRL-ALT-DECEIT Sabotage Evaluations for Automated AI R&D
- Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
- Seeking and Updating with Live Visual Knowledge
- To Think or Not To Think: A Study of Thinking in Rule-Based Visual Reinforcement Fine-Tuning
- Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding