data augmentation
Techniques used to artificially expand the size and diversity of a training dataset by applying transformations (e.g., rotations, scaling, cropping) to existing examples, helping to improve model robustness.
- A Few Moments Please: Scalable Graphon Learning via Moment Matching
- A Statistical Theory of Contrastive Learning via Approximate Sufficient Statistics
- A Unified Framework for Fair Graph Generation: Theoretical Guarantees and Empirical Advances
- AMBER: Adaptive Mesh Generation by Iterative Mesh Resolution Prediction
- An Analysis of Causal Effect Estimation using Outcome Invariant Data Augmentation
- AugGen: Synthetic Augmentation using Diffusion Models Can Improve Recognition
- AutoSciDACT: Automated Scientific Discovery through Contrastive Embedding and Hypothesis Testing
- Behavior Injection: Preparing Language Models for Reinforcement Learning
- Boosting Adversarial Transferability with Spatial Adversarial Alignment
- Coloring Learning for Heterophilic Graph Representation
- DexFlyWheel: A Scalable and Self-improving Data Generation Framework for Dexterous Manipulation
- Diffusion-Classifier Synergy: Reward-Aligned Learning via Mutual Boosting Loop for FSCIL
- Doctor Approved: Generating Medically Accurate Skin Disease Images through AI-Expert Feedback
- Domain-RAG: Retrieval-Guided Compositional Image Generation for Cross-Domain Few-Shot Object Detection
- Enhancing Visual Prompting through Expanded Transformation Space and Overfitting Mitigation
- Explicitly Modeling Subcortical Vision with a Neuro-Inspired Front-End Improves CNN Robustness
- From Style to Facts: Mapping the Boundaries of Knowledge Injection with Finetuning
- GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
- Hybrid-Collaborative Augmentation and Contrastive Sample Adaptive-Differential Awareness for Robust Attributed Graph Clustering
- HyperMixup: Hypergraph-Augmented with Higher-order Information Mixup
- Is Artificial Intelligence Generated Image Detection a Solved Problem?
- Joint‑Embedding vs Reconstruction: Provable Benefits of Latent Space Prediction for Self‑Supervised Learning
- Limited Preference Data? Learning Better Reward Model with Latent Space Synthesis
- Long-tailed Recognition with Model Rebalancing
- Lorentz Local Canonicalization: How to make any Network Lorentz-Equivariant
- MASTER: Enhancing Large Language Model via Multi-Agent Simulated Teaching
- MIDAS: Misalignment-based Data Augmentation Strategy for Imbalanced Multimodal Learning
- NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation
- Non-Asymptotic Analysis Of Data Augmentation For Precision Matrix Estimation
- Principled Data Augmentation for Learning to Solve Quadratic Programming Problems
- Puzzles: Unbounded Video-Depth Augmentation for Scalable End-to-End 3D Reconstruction
- Rao-Blackwell Gradient Estimators for Equivariant Denoising Diffusion
- ReCon: Region-Controllable Data Augmentation with Rectification and Alignment for Object Detection
- Self-Adapting Language Models
- Synthetic-powered predictive inference
- Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation
- Transfer Learning on Edge Connecting Probability Estimation Under Graphon Model
- Transforming Generic Coder LLMs to Effective Binary Code Embedding Models for Similarity Detection
- UtilGen: Utility-Centric Generative Data Augmentation with Dual-Level Task Adaptation
- What Do Latent Action Models Actually Learn?
- Zero-Shot Context Generalization in Reinforcement Learning from Few Training Contexts
- seq-JEPA: Autoregressive Predictive Learning of Invariant-Equivariant World Models