training datasets
Training datasets are collections of data used to train machine learning models. They are crucial for teaching a model to recognize patterns and make predictions. The quality, size, and diversity of training datasets directly impact a model's performance and its ability to generalize to unseen data.
- Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning Abilities
- EvoLM: In Search of Lost Language Model Training Dynamics
- Harnessing the Universal Geometry of Embeddings
- Salient Concept-Aware Generative Data Augmentation
- Señorita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists
- Stitch and Tell: A Structured Data Augmentation Method for Spatial Understanding
- V2V: Scaling Event-Based Vision through Efficient Video-to-Voxel Simulation
- Vid-SME: Membership Inference Attacks against Large Video Understanding Models