> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

---
title: "Open-Weight Models"
url: https://daily.dev/agentic-ai-hub/open-weight-models/
description: "The models you can download, self-host, and fine-tune. As of July 2026. Two things to know: (1) \"open weight\" ≠ open source."
lastUpdated: "2026-07-22"
---

> The models you can download, self-host, and fine-tune. **As of July 2026.** Two things to know: (1) *"open weight" ≠ open source*. Some ship under permissive licenses (Apache-2.0 / MIT), others under custom, use-restricted ones, so the **License** column matters. Check it before you build on a model. (2) The open frontier is now **led by Chinese labs** (DeepSeek, Qwen, Kimi, GLM), while Meta has pivoted toward closed models.

| Model (latest open) | Released | Params (total / active) | License | Context | Best at | Run with |
| :---- | :---- | :---- | :---- | :---- | :---- | :---- |
| **DeepSeek V4** (Pro / Flash) | Apr 2026 | Pro 1.6T/49B · Flash 284B/13B | MIT ✅ | 1M | Reasoning, coding, cheap inference | vLLM, SGLang |
| **Qwen3.6** (35B-A3B) | Apr 2026 | 35B / 3B active | Apache-2.0 ✅ | 262K (~1M YaRN) | Agentic coding, repo-level reasoning | Ollama, vLLM, SGLang |
| **Kimi K3** (Moonshot) | Jul 2026 | 2.8T / ~50B active | Modified MIT ✅ | 1M | Coding/agentic, near-Opus-4.8 | vLLM, SGLang |
| **GLM-5.2** (Zhipu / Z.ai) | 2026 | ~744B / 40B active | MIT ✅ | ~1M | Coding/agentic/reasoning | vLLM, SGLang |
| **Llama 4 Scout** (Meta) | Apr 2025 | 109B / 17B active | Llama 4 Community ⚠️ | 10M | Very long context, single-GPU | Ollama, vLLM |
| **Llama 4 Maverick** (Meta) | Apr 2025 | 400B / 17B active | Llama 4 Community ⚠️ | 1M | Multimodal, general chat | vLLM |
| **Mistral Large 3** | Dec 2025 | 675B / 41B active | Apache-2.0 ✅ | Large | Multimodal, multilingual | vLLM |
| **Ministral 3** (3/8/14B) | Dec 2025 | dense | Apache-2.0 ✅ | n/a | Best small cost/perf, edge | Ollama, llama.cpp |
| **Gemma 4** (Google) | Apr 2026 | 12B (+ MoE 26B-A4B) | Apache-2.0 ✅ | 256K | On-device → server, multimodal | Ollama, llama.cpp, MLX |
| **Qwen3** (0.6B-235B) | 2025 | dense + MoE | Apache-2.0 ✅ | 32K-128K+ | Widest size range, 119 languages | Ollama, vLLM |
| **Phi-4** family (Microsoft) | 2024-26 | 3.8B-~15B | MIT ✅ | 16K-128K | Reasoning-per-parameter, on-device | Ollama, llama.cpp |

✅ permissive · ⚠️ source-available with use restrictions (check the model card).  
Rough VRAM rule of thumb: a dense model needs ~(params × bytes-per-param) of VRAM, e.g. a 27B model at 4-bit quantization ≈ 16 GB. MoE models load all experts into memory but only compute the active set, so they're fast but still memory-hungry.

Sources: [Qwen](https://huggingface.co/Qwen) · [DeepSeek](https://github.com/deepseek-ai/deepseek-v3) · [Mistral](https://mistral.ai/news/mistral-3/) · [Gemma](https://ai.google.dev/gemma/docs/releases) · [Kimi](https://huggingface.co/moonshotai) · [GLM](https://huggingface.co/zai-org/GLM-5) · [Llama](https://ai.meta.com/blog/).

**Track open models live:** [Hugging Face trending](https://huggingface.co/models) (the canonical release feed; the original Open LLM Leaderboard was archived in 2025), [LMArena](https://lmarena.ai/) (includes open models), [llm-stats.com](https://llm-stats.com/llm-updates).