dataset curation
Dataset curation involves the collection, organization, and maintenance of datasets to ensure their quality and suitability for specific tasks. It includes processes like cleaning, labeling, and structuring data to improve model training.
- CGS-GAN: 3D Consistent Gaussian Splatting GANs for High Resolution Human Head Synthesis
- CPO: Condition Preference Optimization for Controllable Image Generation
- CheMixHub: Datasets and Benchmarks for Chemical Mixture Property Prediction
- DATE-LM: Benchmarking Data Attribution Evaluation for Large Language Models
- DataRater: Meta-Learned Dataset Curation
- Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
- Escaping the SpuriVerse: Can Large Vision-Language Models Generalize Beyond Seen Spurious Correlations?
- Frame In-N-Out: Unbounded Controllable Image-to-Video Generation
- Instance-Level Composed Image Retrieval
- LISAt: Language-Instructed Segmentation Assistant for Satellite Imagery
- Less is More: Improving LLM Alignment via Preference Data Selection
- STARC-9: A Large-scale Dataset for Multi-Class Tissue Classification for CRC Histopathology
- SWE-smith: Scaling Data for Software Engineering Agents
- SensorLM: Learning the Language of Wearable Sensors
- The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
- VTON-VLLM: Aligning Virtual Try-On Models with Human Preferences
- VisDiff: SDF-Guided Polygon Generation for Visibility Reconstruction, Characterization and Recognition
- Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLM