Machine Learning News
Read post

Alibaba Releases Qwen1.5-MoE-A2.7B: A Small MoE Model with only 2.7B Activated Parameters yet Matching the Performance of State-of-the-Art 7B models like Mistral 7B

Qwen1.5-MoE-A2.7B is an improved version of Qwen, a Large Language Model (LLM) series developed by the Qwen team at Alibaba Cloud. It performs on par with heavyweight 7B models like Mistral 7B and Qwen1.5-7B, despite having only 2.7 billion activated parameters. The architecture of Qwen1.5-MoE-A2.7B utilizes fine-grained experts and a generalized MoE routing paradigm.

    #ai#alibaba#machine-learning#mixture-of-experts
Mar 29, 2024•4m read time•From marktechpost.com
Post cover image
18 Impressions

Would you recommend this post?

Copy link
WhatsApp
Facebook
X
New Squad
  • © 2026 Daily Dev Ltd.
  • Guidelines
  • Explore
  • Tags
  • Sources
  • Squads
  • Leaderboard