state-of-the-art models
These refer to the most advanced models in AI that achieve the best performance on benchmark tasks, representing the current pinnacle of research and development in specific domains.
- A Token is Worth over 1,000 Tokens: Efficient Knowledge Distillation through Low-Rank Clone
- AR-RAG: Autoregressive Retrieval Augmentation for Image Generation
- Accelerating 3D Molecule Generative Models with Trajectory Diagnosis
- AudSemThinker: Enhancing Audio-Language Models Through Reasoning over Semantics of Sound
- Automated Detection of Visual Attribute Reliance with a Self-Reflective Agent
- Autoregressive Motion Generation with Gaussian Mixture-Guided Latent Sampling
- C3Po: Cross-View Cross-Modality Correspondence by Pointmap Prediction
- CAMILA: Context-Aware Masking for Image Editing with Language Alignment
- CoralVQA: A Large-Scale Visual Question Answering Dataset for Coral Reef Image Understanding
- DBLoss: Decomposition-based Loss Function for Time Series Forecasting
- DecompNet: Enhancing Time Series Forecasting Models with Implicit Decomposition
- DisMo: Disentangled Motion Representations for Open-World Motion Transfer
- EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
- Evaluating Robustness of Monocular Depth Estimation with Procedural Scene Perturbations
- FlexVAR: Flexible Visual Autoregressive Modeling without Residual Prediction
- IndEgo: A Dataset of Industrial Scenarios and Collaborative Work for Egocentric Assistants
- LexiCon: a Benchmark for Planning under Temporal Constraints in Natural Language
- MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models
- MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
- Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models
- MuSLR: Multimodal Symbolic Logical Reasoning
- Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving
- Multi-View Oriented GPLVM: Expressiveness and Efficiency
- OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-time Emotional Speech Synthesis
- RayFusion: Ray Fusion Enhanced Collaborative Visual Perception
- SWE-bench Goes Live!
- Scaling Physical Reasoning with the PHYSICS Dataset
- Selective Learning for Deep Time Series Forecasting
- Stochastic Process Learning via Operator Flow Matching
- SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
- THUNDER: Tile-level Histopathology image UNDERstanding benchmark
- WolBanking77: Wolof Banking Speech Intent Classification Dataset