foundation models
Foundation models are large-scale pre-trained models that serve as a basis for downstream tasks. They are designed to be general-purpose, capturing broad knowledge from vast datasets to be fine-tuned or adapted for specific applications.
- A Generalist Intracortical Motor Decoder
- A Unified Reasoning Framework for Holistic Zero-Shot Video Anomaly Analysis
- AION-1: Omnimodal Foundation Model for Astronomical Sciences
- Activated LoRA: Fine-tuned LLMs for Intrinsics
- Aeolus: A Multi-structural Flight Delay Dataset
- An Investigation of Memorization Risk in Healthcare Foundation Models
- Analog Foundation Models
- Axial Neural Networks for Dimension-Free Foundation Models
- Blindfolded Experts Generalize Better: Insights from Robotic Manipulation and Videogames
- CPathAgent: An Agent-based Foundation Model for Interpretable High-Resolution Pathology Image Analysis Mimicking Pathologists' Diagnostic Logic
- CSBrain: A Cross-scale Spatiotemporal Brain Foundation Model for EEG Decoding
- Certifying Deep Network Risks and Individual Predictions with PAC-Bayes Loss via Localized Priors
- ChemPile: A 250 GB Diverse and Curated Dataset for Chemical Foundation Models
- ChromFound: Towards A Universal Foundation Model for Single-Cell Chromatin Accessibiltiy Data
- Continuous Subspace Optimization for Continual Learning
- CrossSpectra: Exploiting Cross-Layer Smoothness for Parameter-Efficient Fine-Tuning
- CrypticBio: A Large Multimodal Dataset for Visually Confusing Species
- Data-Juicer 2.0: Cloud-Scale Adaptive Data Processing for and with Foundation Models
- DataRater: Meta-Learned Dataset Curation
- Diffusion Transformers as Open-World Spatiotemporal Foundation Models
- Doctor Approved: Generating Medically Accurate Skin Disease Images through AI-Expert Feedback
- DynaNav: Dynamic Feature and Layer Selection for Efficient Visual Navigation
- Dynamic Risk Assessments for Offensive Cybersecurity Agents
- DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling
- EMLoC: Emulator-based Memory-efficient Fine-tuning with LoRA Correction
- Efficient Randomized Experiments Using Foundation Models
- EgoDTM: Towards 3D-Aware Egocentric Video-Language Pretraining
- Exploiting the Asymmetric Uncertainty Structure of Pre-trained VLMs on the Unit Hypersphere
- FaCT: Faithful Concept Traces for Explaining Neural Network Decisions
- Fantastic Features and Where to Find Them: A Probing Method to combine Features from Multiple Foundation Models
- Foundation Models for Scientific Discovery: From Paradigm Enhancement to Paradigm Transition
- From Pretraining to Pathology: How Noise Leads to Catastrophic Inheritance in Medical Models
- GeoLink: Empowering Remote Sensing Foundation Model with OpenStreetMap Data
- GraphMaster: Automated Graph Synthesis via LLM Agents in Data-Limited Environments
- HermesFlow: Seamlessly Closing the Gap in Multimodal Understanding and Generation
- How Many Tokens Do 3D Point Cloud Transformer Architectures Really Need?
- Image Super-Resolution with Guarantees via Conformalized Generative Models
- Is the acquisition worth the cost? Surrogate losses for Consistent Two-stage Classifiers
- Iterative Foundation Model Fine-Tuning on Multiple Rewards
- KGGen: Extracting Knowledge Graphs from Plain Text with Language Models
- LISAt: Language-Instructed Segmentation Assistant for Satellite Imagery
- LT-Soups: Bridging Head and Tail Classes via Subsampled Model Soups
- Leader360V: A Large-scale, Real-world 360 Video Dataset for Multi-task Learning in Diverse Environment
- Long-tailed Recognition with Model Rebalancing
- MMTU: A Massive Multi-Task Table Understanding and Reasoning Benchmark
- Mars-Bench: A Benchmark for Evaluating Foundation Models for Mars Science Tasks
- MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants
- MoORE: SVD-based Model MoE-ization for Conflict- and Oblivion-Resistant Multi-Task Adaptation
- Modeling the Economic Impacts of AI Openness Regulation
- Motion4D: Learning 3D-Consistent Motion and Semantics for 4D Scene Understanding
- Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables
- Navigating the MIL Trade-Off: Flexible Pooling for Whole Slide Image Classification
- NeurIPT: Foundation Model for Neural Interfaces
- OPTFM: A Scalable Multi-View Graph Transformer for Hierarchical Pre-Training in Combinatorial Optimization
- On the creation of narrow AI: hierarchy and nonlocality of neural network skills
- PRIMT: Preference-based Reinforcement Learning with Multimodal Feedback and Trajectory Synthesis from Foundation Models
- PRIMT: Preference-based Reinforcement Learning with Multimodal Feedback and Trajectory Synthesis from Foundation Models
- Parameter Efficient Fine-tuning via Explained Variance Adaptation
- PathVQ: Reforming Computational Pathology Foundation Model for Whole Slide Image Analysis via Vector Quantization
- PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts
- Position: Biology is the Challenge Physics-Informed ML Needs to Evolve
- Provable Meta-Learning with Low-Rank Adaptations
- QSCA: Quantization with Self-Compensating Auxiliary for Monocular Depth Estimation
- REOBench: Benchmarking Robustness of Earth Observation Foundation Models
- REVE: A Foundation Model for EEG - Adapting to Any Setup with Large-Scale Pretraining on 25,000 Subjects
- RoFt-Mol: Benchmarking Robust Fine-tuning with Molecular Graph Foundation Models
- SEMPO: Lightweight Foundation Models for Time Series Forecasting
- STEER-ME: Assessing the Microeconomic Reasoning of Large Language Models
- Scaling Laws for Optimal Data Mixtures
- Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets
- Scaling Unlocks Broader Generation and Deeper Functional Understanding of Proteins
- SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
- Self-Guided Hierarchical Exploration for Generalist Foundation Model Web Agents
- Self-Improving Embodied Foundation Models
- Succeed or Learn Slowly: Sample Efficient Off-Policy Reinforcement Learning for Mobile App Control
- Synthesizing Photorealistic and Dynamic Urban Environments for Multimodal Robot Navigation and Collaboration
- Synthetic Series-Symbol Data Generation for Time Series Foundation Models
- THUNDER: Tile-level Histopathology image UNDERstanding benchmark
- TS-RAG: Retrieval-Augmented Generation based Time Series Foundation Models are Stronger Zero-Shot Forecaster
- TabArena: A Living Benchmark for Machine Learning on Tabular Data
- Tabula: A Tabular Self-Supervised Foundation Model for Single-Cell Transcriptomics
- Towards Physics-informed Spatial Intelligence with Human Priors: An Autonomous Driving Pilot Study
- VESSA: Video-based objEct-centric Self-Supervised Adaptation for Visual Foundation Models
- VIPAMIN: Visual Prompt Initialization via Embedding Selection and Subspace Expansion
- VMDT: Decoding the Trustworthiness of Video Foundation Models
- scGeneScope: A Treatment-Matched Single Cell Imaging and Transcriptomics Dataset and Benchmark for Treatment Response Modeling