knowledge distillation
Knowledge distillation is a technique for transferring knowledge from a large, complex model (teacher) to a smaller, simpler model (student). The smaller model aims to mimic the behavior of the larger model to achieve comparable performance with reduced complexity.
- A Token is Worth over 1,000 Tokens: Efficient Knowledge Distillation through Low-Rank Clone
- ATLAS: Autoformalizing Theorems through Lifting, Augmentation, and Synthesis of Data
- AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders
- Advancing Machine-Generated Text Detection from an Easy to Hard Supervision Perspective
- Better Estimation of the Kullback--Leibler Divergence Between Language Models
- Centralized Reward Agent for Knowledge Sharing and Transfer in Multi-Task Reinforcement Learning
- Continuous Concepts Removal in Text-to-image Diffusion Models
- DKDR: Dynamic Knowledge Distillation for Reliability in Federated Learning
- DeepKD: A Deeply Decoupled and Denoised Knowledge Distillation Trainer
- Deliberation on Priors: Trustworthy Reasoning of Large Language Models on Knowledge Graphs
- Detecting Data Deviations in Electronic Health Records
- Efficient Hybrid Language Model Compression through Group-Aware SSM Pruning
- Efficient Verified Unlearning For Distillation
- Enhanced Expert Merging for Mixture-of-Experts in Graph Foundation Models
- Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
- HPSERec: A Hierarchical Partitioning and Stepwise Enhancement Framework for Long-tailed Sequential Recommendation
- Hawaii: Hierarchical Visual Knowledge Transfer for Efficient Vision-Language Models
- Improving Task-Specific Multimodal Sentiment Analysis with General MLLMs via Prompting
- InfiFPO: Implicit Model Fusion via Preference Optimization in Large Language Models
- Interaction-Centric Knowledge Infusion and Transfer for Open Vocabulary Scene Graph Generation
- Knowledge Distillation of Uncertainty using Deep Latent Factor Model
- MOSDT: Self-Distillation-Based Decision Transformer for Multi-Agent Offline Safe Reinforcement Learning
- MTL-KD: Multi-Task Learning Via Knowledge Distillation for Generalizable Neural Vehicle Routing Solver
- MURKA: Multi-Reward Reinforcement Learning with Knowledge Alignment for Optimization Tasks
- Mitigating Intra- and Inter-modal Forgetting in Continual Learning of Unified Multimodal Models
- MonoLift: Learning 3D Manipulation Policies from Monocular RGB via Distillation
- Multi-order Orchestrated Curriculum Distillation for Model-Heterogeneous Federated Graph Learning
- NaturalReasoning: Reasoning in the Wild with 2.8M Challenging Questions
- NegoCollab: A Common Representation Negotiation Approach for Heterogeneous Collaborative Perception
- PLD: A Choice-Theoretic List-Wise Knowledge Distillation
- SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning
- SSTAG: Structure-Aware Self-Supervised Learning Method for Text-Attributed Graphs
- Segment Anything Model Meets Semi-supervised Medical Image Segmentation: A Novel Perspective
- SemCoT: Accelerating Chain-of-Thought Reasoning through Semantically-Aligned Implicit Tokens
- Single-Teacher View Augmentation: Boosting Knowledge Distillation via Angular Diversity
- SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
- Synergy Between the Strong and the Weak: Spiking Neural Networks are Inherently Self-Distillers
- Taught Well Learned Ill: Towards Distillation-conditional Backdoor Attack
- Token-Level Self-Play with Importance-Aware Guidance for Large Language Models
- Unleashing the Power of One-Step Diffusion based Image Super-Resolution via a Large-Scale Diffusion Discriminator
- Unlocking SLM Potential for Data Analysis Code Generation via Non-Parametric Knowledge Distillation
- Why Knowledge Distillation Works in Generative Models: A Minimal Working Explanation