mixture-of-experts
A model architecture that consists of multiple expert sub-models, where a gating mechanism determines which experts to activate for a given input, enhancing efficiency and performance on diverse tasks.
- Advancing Expert Specialization for Better MoE
- Advancing Expert Specialization for Better MoE
- CALM: Culturally Self-Aware Language Models
- CryptoMoE: Privacy-Preserving and Scalable Mixture of Experts Inference via Balanced Expert Routing
- DiEP: Adaptive Mixture-of-Experts Compression through Differentiable Expert Pruning
- Discovering Important Experts for Mixture-of-Experts Models Pruning Through a Theoretical Perspective
- Domain-Specific Pruning of Large Mixture-of-Experts Models with Few-shot Demonstrations
- Enhanced Expert Merging for Mixture-of-Experts in Graph Foundation Models
- FAME: Adaptive Functional Attention with Expert Routing for Function-on-Function Regression
- FlashMoE: Fast Distributed MoE in a Single Kernel
- FlexOLMo: Open Language Models for Flexible Data Use
- FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts Training
- FlyLoRA: Boosting Task Decoupling and Parameter Efficiency via Implicit Rank-Wise Mixture-of-Experts
- Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
- Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
- JanusDNA: A Powerful Bi-directional Hybrid DNA Foundation Model
- Learning and Planning Multi-Agent Tasks via an MoE-based World Model
- Learning to Specialize: Joint Gating-Expert Training for Adaptive MoEs in Decentralized Settings
- MJ-Video: Benchmarking and Rewarding Video Generation with Fine-Grained Video Preference
- Mixture-of-Experts Meets In-Context Reinforcement Learning
- Mixture-of-Experts Operator Transformer for Large-Scale PDE Pre-Training
- MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
- MoEMeta: Mixture-of-Experts Meta Learning for Few-Shot Relational Learning
- Model Merging in Pre-training of Large Language Models
- Mozart: Modularized and Efficient MoE Training on 3.5D Wafer-Scale Chiplet Architectures
- OVS Meets Continual Learning: Towards Sustainable Open-Vocabulary Segmentation
- On Linear Mode Connectivity of Mixture-of-Experts Architectures
- On Linear Mode Connectivity of Mixture-of-Experts Architectures
- On the Expressive Power of Mixture-of-Experts for Structured Complex Tasks
- PT-MoE: An Efficient Finetuning Framework for Integrating Mixture-of-Experts into Prompt Tuning
- Robust Ego-Exo Correspondence with Long-Term Memory
- S'MoRE: Structural Mixture of Residual Experts for Parameter-Efficient LLM Fine-tuning
- SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert Identification
- SimulMEGA: MoE Routers are Advanced Policy Makers for Simultaneous Speech Translation
- The Omni-Expert: A Computationally Efficient Approach to Achieve a Mixture of Experts in a Single Expert Model
- Toward Efficient Inference Attacks: Shadow Model Sharing via Mixture-of-Experts
- Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training
- Unifying and Enhancing Graph Transformers via a Hierarchical Mask Framework
- Where Graph Meets Heterogeneity: Multi-View Collaborative Graph Experts
- Who Speaks for the Trigger? Dynamic Expert Routing in Backdoored Mixture-of-Experts Transformers