pretraining
The phase in which a model is trained on a large dataset to learn general representations before being fine-tuned on a specific task-related dataset.
- AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
- AdaLRS: Loss-Guided Adaptive Learning Rate Search for Efficient Foundation Model Pretraining
- Axial Neural Networks for Dimension-Free Foundation Models
- Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
- Ditch the Denoiser: Emergence of Noise Robustness in Self-Supervised Learning from Data Curriculum
- Do You Really Need Public Data? Surrogate Public Data for Differential Privacy on Tabular Data
- Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation
- Flattening Hierarchies with Policy Bootstrapping
- From Faults to Features: Pretraining to Learn Robust Representations against Sensor Failures
- KnowMol: Advancing Molecular Large Language Models with Multi-Level Chemical Knowledge
- LLM Layers Immediately Correct Each Other
- MMaDA: Multimodal Large Diffusion Language Models
- Mixtures of Subspaces for Bandwidth Efficient Context Parallel Training
- MuRating: A High Quality Data Selecting Approach to Multilingual Large Language Model Pretraining
- Object-centric binding in Contrastive Language-Image Pretraining
- Positional Fragility in LLMs: How Offset Effects Reshape Our Understanding of Memorization Risks
- Principled Data Augmentation for Learning to Solve Quadratic Programming Problems
- Prior Forgetting and In-Context Overfitting
- Recursive Inference Scaling: A Winning Path to Scalable Inference in Language and Multimodal Systems
- ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code
- Scaling Laws for Optimal Data Mixtures
- Self-Improving Embodied Foundation Models
- Time Series Generation Under Data Scarcity: A Unified Generative Modeling Approach
- Who Speaks for the Trigger? Dynamic Expert Routing in Backdoored Mixture-of-Experts Transformers