benchmark tasks
Standardized tasks and datasets used to evaluate and compare the performance of different AI algorithms and models. They facilitate consistent assessments of advancements in the field.
- Acceleration via silver step-size on Riemannian manifolds with applications to Wasserstein space
- DisCO: Reinforcing Large Reasoning Models with Discriminative Constrained Optimization
- EDBench: Large-Scale Electron Density Data for Molecular Modeling
- FNOPE: Simulation-based inference on function spaces with Fourier Neural Operators
- Generative Trajectory Stitching through Diffusion Composition
- Local-Global Associative Frames for Symmetry-Preserving Crystal Structure Modeling
- SYMPHONY: Synergistic Multi-agent Planning with Heterogeneous Language Model Assembly
- SpecEM: Training-Free LLM Ensembling via Iterative Drafting, Verification, and Online Feedback
- Structure-Aware Fusion with Progressive Injection for Multimodal Molecular Representation Learning
- Think or Not? Exploring Thinking Efficiency in Large Reasoning Models via an Information-Theoretic Lens