large-scale dataset
A large-scale dataset refers to a collection of data encompassing a vast number of examples, often in the millions or billions, that is typically used to train artificial intelligence models. In the context of deep learning, these datasets allow models to learn complex patterns and achieve high performance across various tasks by providing comprehensive diversity and representation.
- Alligat0R: Pre-Training through Covisibility Segmentation for Relative Camera Pose Regression
- Aux-Think: Exploring Reasoning Strategies for Data-Efficient Vision-Language Navigation
- BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset
- EDBench: Large-Scale Electron Density Data for Molecular Modeling
- From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D
- GenColor: Generative and Expressive Color Enhancement with Pixel-Perfect Texture Preservation
- Hi3DEval: Advancing 3D Generation Evaluation with Hierarchical Validity
- MeshCoder: LLM-Powered Structured Mesh Code Generation from Point Clouds
- Multimodal 3D Genome Pre-training
- OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation
- Open-Insect: Benchmarking Open-Set Recognition of Novel Species in Biodiversity Monitoring
- OpenCUA: Open Foundations for Computer-Use Agents
- Optimize Any Topology: A Foundation Model for Shape- and Resolution-Free Structural Topology Optimization
- OrthoLoC: UAV 6-DoF Localization and Calibration Using Orthographic Geodata
- OrthoLoC: UAV 6-DoF Localization and Calibration Using Orthographic Geodata
- SSIMBaD: Sigma Scaling with SSIM-Guided Balanced Diffusion for AnimeFace Colorization
- Seg2Any: Open-set Segmentation-Mask-to-Image Generation with Precise Shape and Semantic Control
- UAV-Flow Colosseo: A Real-World Benchmark for Flying-on-a-Word UAV Imitation Learning
- Walking the Tightrope: Autonomous Disentangling Beneficial and Detrimental Drifts in Non-Stationary Custom-Tuning