附录A:MoE相关论文推荐与阅读指南 按主题分类的论文列表 基础理论与原始模型 Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer (Shazeer et al., ICLR 2017) 意义:MoE领域的开山之作,提出了稀疏门控MoE层 核心贡献:首次在LSTM中引入MoE,证明了稀疏激活的可行性 阅读建议:必读,理解MoE的基本思想 Sparsely-Gated Mixture-of-Experts Layer (Shazeer et al.
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer (Shazeer et al., ICLR 2017)
Sparsely-Gated Mixture-of-Experts Layer (Shazeer et al., 2017)
GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding (Lepikhin et al., ICLR 2021)
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity (Fedus et al., JMLR 2022)
ST-MoE: Design and Stability of Multi-Expert Transformer Models (Zoph et al., TMLR 2022)
Base Layers: Simplifying Training of Large, Local Models (Lewis et al., ICML 2021)
Expert Choice Routing (Zhou et al., NeurIPS 2022)
HashMoE: Hashtable-Based Routing for Mixture-of-Experts (Levine et al., NeurIPS 2023)
Softmax Attention is Not All You Need: Achieving Linear Complexity with Attention-free Routing (Poe et al., 2023)
Mixtral of Experts (Jiang et al., arXiv 2024)
Qwen2.5 MoE Technical Report (Qwen Team, 2024)
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model (DeepSeek, 2024)
MoEfication: Converting Dense Models to Mixture-of-Experts (Kim et al., ACL 2023)
Speed Is All You Need: On-Device Acceleration of Large Language Models with Speculative Decoding and Adaptive Sparse MoE (Miao et al., 2024)
Mixture-of-Experts with Expert Choice Routing for Language Modeling (Puigcerver et al., NeurIPS 2024)
Mixture-of-Experts with Fine-grained Expert Expansion for Parameter-Efficient Multi-task Learning (Cheng et al., 2023)
Vision MoE: Scaling Vision Transformers with Sparse Mixture of Experts (Riquelme et al., 2024)
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models (Lin et al., 2024)