pre-trained models
Pre-trained models are those that have been initially trained on large datasets and can be further refined or fine-tuned for specific applications. They save time and compute resources, leveraging previously learned knowledge for new tasks.
- AdaSTaR: Adaptive Data Sampling for Training Self-Taught Reasoners
- AnaCP: Toward Upper-Bound Continual Learning via Analytic Contrastive Projection
- Building 3D Representations and Generating Motions From a Single Image via Video-Generation
- ChatVLA-2: Vision-Language-Action Model with Open-World Reasoning
- CogVLA: Cognition-Aligned Vision-Language-Action Models via Instruction-Driven Routing & Sparsification
- Continuous Subspace Optimization for Continual Learning
- Covariances for Free: Exploiting Mean Distributions for Training-free Federated Learning
- DP²O-SR: Direct Perceptual Preference Optimization for Real-World Image Super-Resolution
- Dataset Distillation for Pre-Trained Self-Supervised Vision Models
- Epistemic Uncertainty for Generated Image Detection
- Exploiting the Asymmetric Uncertainty Structure of Pre-trained VLMs on the Unit Hypersphere
- FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities
- FedRACE: A Hierarchical and Statistical Framework for Robust Federated Learning
- Federated Continual Learning via Orchestrating Multi-Scale Expertise
- FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation
- FlowerTune: A Cross-Domain Benchmark for Federated Fine-Tuning of Large Language Models
- FoGE: Fock Space inspired encoding for graph prompting
- HiFlow: Training-free High-Resolution Image Generation with Flow-Aligned Guidance
- Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale Problems
- Implicit Modeling for Transferability Estimation of Vision Foundation Models
- Knowledge Graph Enhanced Generative Multi-modal Models for Class-Incremental Learning
- LOMIA: Label-Only Membership Inference Attacks against Pre-trained Large Vision-Language Models
- Learning Diffusion Models with Flexible Representation Guidance
- Learning Multi-Source and Robust Representations for Continual Learning
- Mesh-RFT: Enhancing Mesh Generation via Fine-grained Reinforcement Fine-Tuning
- Mixture of Noise for Pre-Trained Model-Based Class-Incremental Learning
- Model Inversion with Layer-Specific Modeling and Alignment for Data-Free Continual Learning
- Mysteries of the Deep: Role of Intermediate Representations in Out of Distribution Detection
- NormFit: A Lightweight Solution for Few-Shot Federated Learning with Non-IID Data
- On the Loss of Context Awareness in General Instruction Fine-tuning
- PROFIT: A Specialized Optimizer for Deep Fine Tuning
- Parameter Dynamics of Online Machine Learning and Test-time Adaptation
- Probabilistic Token Alignment for Large Language Model Fusion
- RobustMerge: Parameter-Efficient Model Merging for MLLMs with Direction Robustness
- SSTAG: Structure-Aware Self-Supervised Learning Method for Text-Attributed Graphs
- Single GPU Task Adaptation of Pathology Foundation Models for Whole Slide Image Analysis
- State Space Prompting via Gathering and Spreading Spatio-Temporal Information for Video Understanding
- TTRL: Test-Time Reinforcement Learning
- Towards Robust Parameter-Efficient Fine-Tuning for Federated Learning
- True Zero-Shot Inference of Dynamical Systems Preserving Long-Term Statistics
- Understanding Bias Terms in Neural Representations
- Unified Transferability Metrics for Time Series Foundation Models
- Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive Activations
- VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
- ZeroPatcher: Training-free Sampler for Video Inpainting and Editing
- ZeroSep: Separate Anything in Audio with Zero Training