Local LLM Model Database
Every local model worth running in 2026 in one sortable table, real VRAM at Q4, context window, license, use-case tags and a directional quality tier. Search it, filter by use case, and pick your GPU to grey out anything that won't fit. No cloud-only models, no fabricated benchmarks.
Most "best local LLM" lists are frozen the day they're published, or they rank cloud APIs you can't actually download. This one is different: 64 genuinely open-weight, runs-on-your-hardware models, Qwen3, Llama, Gemma 3, Mistral, DeepSeek, gpt-oss, GLM, Command-R and the uncensored/roleplay finetunes, with the one number that decides everything, how much VRAM the weights take at Q4_K_M. Pair it with the VRAM calculator for an exact fit check on your card.
Reading the table: VRAM @ Q4 is the weight footprint at Q4_K_M (the quality sweet spot); Min VRAM is a tighter Q3/low estimate for squeezing a model onto a smaller card. Mixture-of-Experts rows show total · active params, the total drives VRAM (every expert lives in memory), the active count drives speed. The Quality column is an editorial tier, explained below.
| total / active | Q4_K_M weights, GB | Q3 / low, GB | Type | editorial tier | Notes | |||
|---|---|---|---|---|---|---|---|---|
| DeepSeek V3 DeepSeek V3 · MoE | 671B · 37B active | 403 GB | 294 GB | 128K | DeepSeek | chatreasoningcoding | S | 671B MoE, server-class, listed for reference. |
| DeepSeek V3.2 DeepSeek V3 · MoE | 671B · 37B active | 403 GB | 294 GB | 128K | DeepSeek | chatreasoningcoding | S | Sparse-attention refresh of V3; still server-class. |
| Qwen3-Coder 480B-A35B Qwen3-Coder · MoE | 480B · 35B active | 288 GB | 210 GB | 256K | Apache-2.0 | codingreasoning | S | Agentic-coding SOTA, cloud-scale VRAM, not consumer-local. |
| Qwen3 235B-A22B Qwen3 · MoE | 235B · 22B active | 141 GB | 103 GB | 256K | Apache-2.0 | chatreasoningmultilingual | S | Flagship MoE for multi-GPU / server rig territory. |
| gpt-oss 120B gpt-oss (MoE) · MoE | 116.8B · 5.1B active | 63 GB | 59 GB | 128K | Apache-2.0 | reasoningchat | S | MXFP4 native; fits ~64-80GB (or one 96GB card). |
| GLM-4.5-Air GLM · MoE | 106B · 12B active | 64 GB | 46 GB | 128K | MIT | chatreasoningcoding | S | Agentic MoE; ~64GB at Q4, only ~12B active per token. |
| DeepSeek-R1-Distill Llama 70B DeepSeek-R1 (distill) | 70.6B | 42 GB | 31 GB | 128K | MIT | reasoning | S | R1 reasoning distilled into a Llama 70B body. |
| Llama 3.3 70B Llama | 70.6B | 42 GB | 31 GB | 128K | Llama-3.3 | chatreasoningmultilingual | S | 70B quality; needs ~48GB or a dual-GPU / offload setup. |
| Qwen3 32B Qwen3 | 32.8B | 20 GB | 14 GB | 128K | Apache-2.0 | chatreasoningmultilingual | S | Best all-round dense model for a 24GB card. |
| Qwen2.5-Coder 32B Qwen2.5-Coder | 32.5B | 20 GB | 14 GB | 128K | Apache-2.0 | coding | S | Top open coding model that fits a single 24GB GPU. |
| GLM-5.2 GLM · MoE | 743B · 39B active | 446 GB | 325 GB | 1024K | MIT | chatreasoningcoding | A | 743B MoE frontier model, open weights on HF, cloud-only convenience on Ollama; server-scale VRAM. |
| DeepSeek-V4-Flash DeepSeek V4 · MoE | 284B · 13B active | 170 GB | 124 GB | 1000K | MIT | chatreasoningcoding | A | 284B MoE, ~13B active, 1M context, server-class VRAM, listed for reference. |
| Mixtral 8x22B Mixtral (MoE) · MoE | 141B · 39B active | 85 GB | 62 GB | 64K | Apache-2.0 | chatreasoning | A | 141B in VRAM, fast per token, but server-class memory. |
| Mistral Medium 3.5 128B Mistral | 128B | 77 GB | 56 GB | 256K | Modified MIT | chatcodingreasoningvisionmultilingual | A | 128B dense flagship, multimodal, needs a multi-GPU / big-rig setup. |
| Command R+ 104B Cohere Command-R | 104B | 62 GB | 46 GB | 128K | CC-BY-NC-4.0 | chatreasoning | A | RAG / tool-use specialist; non-commercial license. |
| Llama 3.1 70B Llama | 70.6B | 42 GB | 31 GB | 128K | Llama-3.1 | chatmultilingual | A | The workhorse 70B before 3.3; still very capable. |
| Qwen3.6 35B-A3B Qwen3.6 · MoE | 35B · 3B active | 21 GB | 15 GB | 256K | Apache-2.0 | chatcodingreasoningvisionmultilingual | A | Agentic-coding MoE, 24GB-friendly, only ~3B active per token. |
| Laguna XS.2 Laguna (Poolside) · MoE | 33B · 3B active | 20 GB | 14 GB | 256K | Apache-2.0 | codingreasoning | A | Poolside 33B-A3B coding MoE (68% SWE-bench Verified); ~36GB local. |
| DeepSeek-R1-Distill Qwen 32B DeepSeek-R1 (distill) | 32.8B | 20 GB | 14 GB | 128K | MIT | reasoning | A | Best reasoning model that fits a 24GB card. |
| GLM-4 32B GLM | 32B | 19 GB | 14 GB | 32K | MIT | chatcodingreasoning | A | Dense 32B strong at code & function-calling. |
| Gemma 4 31B Gemma 4 | 31B | 19 GB | 14 GB | 256K | Gemma | chatvisionmultilingual | A | Dense multimodal flagship for a 24GB card; 256K context. |
| Qwen3 30B-A3B Qwen3 · MoE | 30.5B · 3.3B active | 18 GB | 13 GB | 256K | Apache-2.0 | chatreasoningmultilingual | A | Fast MoE, 24GB-friendly, only ~3B active per token. |
| Qwen3-Coder 30B-A3B Qwen3-Coder · MoE | 30.5B · 3.3B active | 18 GB | 13 GB | 256K | Apache-2.0 | codingreasoning | A | Local agentic coder, 24GB-friendly MoE. |
| North Mini Code 1.0 Cohere North · MoE | 30B · 3B active | 18 GB | 13 GB | 256K | Apache-2.0 | codingreasoning | A | Cohere's 30B-A3B agentic-coding MoE; runs on one 24-48GB GPU. |
| Gemma 3 27B Gemma 3 | 27.4B | 16 GB | 12 GB | 128K | Gemma | chatvisionmultilingual | A | Strong multimodal all-rounder for a 24GB card. |
| Qwen3.6 27B Qwen3.6 | 27B | 16 GB | 12 GB | 256K | Apache-2.0 | chatcodingreasoningmultilingual | A | Dense 27B all-rounder for a 24GB card; thinking + coding. |
| Gemma 4 26B-A4B Gemma 4 · MoE | 26B · 4B active | 16 GB | 11 GB | 256K | Gemma | chatvisionmultilingual | A | MoE with all experts in VRAM (~16GB), only ~4B active per token. |
| Cydonia 24B Uncensored / roleplay | 23.6B | 14 GB | 10 GB | 32K | Apache-2.0 | roleplayuncensoredchat | A | Uncensored roleplay finetune of Mistral Small. |
| Dolphin Mistral 24B Venice Uncensored / roleplay | 23.6B | 14 GB | 10 GB | 32K | Apache-2.0 | chatuncensoredroleplay | A | Uncensored assistant; Venice edition of Dolphin. |
| Mistral Small 3.2 24B Mistral | 23.6B | 14 GB | 10 GB | 128K | Apache-2.0 | chatvision | A | 24GB sweet spot with vision and a permissive license. |
| gpt-oss 20B gpt-oss (MoE) · MoE | 20.9B · 3.6B active | 12 GB | 10 GB | 128K | Apache-2.0 | reasoningchat | A | Ships MXFP4 4-bit natively at a ~12GB footprint, runs on 16GB. |
| Qwen3 14B Qwen3 | 14.8B | 9 GB | 6.5 GB | 128K | Apache-2.0 | chatreasoningmultilingual | A | Strong 14B that fits comfortably in 12-16GB. |
| Phi-4 14B Phi | 14.7B | 9 GB | 6.4 GB | 16K | MIT | reasoning | A | Reasoning-dense 14B; note the shorter 16K context. |
| Qwen2.5-Coder 14B Qwen2.5-Coder | 14.7B | 9 GB | 6.4 GB | 128K | Apache-2.0 | coding | A | Best coder for 12-16GB. |
| Mixtral 8x7B Mixtral (MoE) · MoE | 46.7B · 12.9B active | 28 GB | 20 GB | 32K | Apache-2.0 | chat | B | All 8 experts sit in VRAM; ~13B active per token. |
| Command R 35B Cohere Command-R | 35B | 21 GB | 15 GB | 128K | CC-BY-NC-4.0 | chat | B | Grounded RAG on a 24GB card; non-commercial. |
| Ornith 1.0 35B-A3B Ornith · MoE | 35B · 3B active | 21 GB | 15 GB | 256K | MIT | codingreasoning | B | Self-scaffolding agentic-coding MoE (Qwen3.5 base); ~3B active. |
| Yi-1.5 34B Yi | 34.4B | 21 GB | 15 GB | 32K | Apache-2.0 | chatmultilingual | B | Solid bilingual 34B; shorter native context. |
| Granite 4.1 30B Granite (IBM) | 30B | 18 GB | 13 GB | 128K | Apache-2.0 | chatcoding | B | Enterprise dense 30B tuned for RAG, tool-use and function-calling. |
| Nemotron 3 Nano Omni 30B-A3B Nemotron 3 (NVIDIA) · MoE | 30B · 3B active | 18 GB | 13 GB | 128K | NVIDIA Open Model | chatvisionreasoning | B | Omni-modal (text / image / audio / video) MoE; ~3B active per token. |
| DeepSeek-R1-Distill Qwen 14B DeepSeek-R1 (distill) | 14.8B | 9 GB | 6.5 GB | 128K | MIT | reasoning | B | Chain-of-thought reasoning for 12-16GB. |
| Gemma 3 12B Gemma 3 | 12.2B | 7.3 GB | 5.3 GB | 128K | Gemma | chatvisionmultilingual | B | Multimodal 12B for 12-16GB. |
| Mistral NeMo 12B Mistral | 12.2B | 7.3 GB | 5.3 GB | 128K | Apache-2.0 | chatmultilingual | B | Popular 12B base for finetunes; 128K context. |
| Rocinante 12B Uncensored / roleplay | 12.2B | 7.3 GB | 5.3 GB | 128K | Apache-2.0 | roleplayuncensoredchat | B | Popular 12B roleplay finetune (Mistral NeMo). |
| Gemma 4 12B Gemma 4 | 12B | 7.3 GB | 5.3 GB | 256K | Gemma | chatvisionmultilingual | B | Multimodal 12B for 12-16GB; 256K context. |
| Ornith 1.0 9B Ornith | 9B | 5.4 GB | 3.9 GB | 256K | MIT | coding | B | DeepReinforce self-scaffolding coder; edge-friendly 9B dense. |
| LFM2.5-8B-A1B LiquidAI LFM · MoE | 8.3B · 1.5B active | 5 GB | 3.6 GB | 128K | LFM Open License | chatcoding | B | On-device MoE (~1.5B active) built for fast, reliable tool-calling. |
| Qwen3 8B Qwen3 | 8.2B | 5 GB | 3.6 GB | 128K | Apache-2.0 | chatmultilingual | B | Great daily driver for 8GB cards. |
| Qwen3 8B Abliterated Uncensored / roleplay | 8.2B | 5 GB | 3.6 GB | 128K | Apache-2.0 | chatuncensored | B | Refusal-removed Qwen3 8B for an 8GB card. |
| DeepSeek-R1-Distill Llama 8B DeepSeek-R1 (distill) | 8B | 4.8 GB | 3.5 GB | 128K | MIT | reasoning | B | Reasoning traces on an 8GB budget. |
| Dolphin 2.9 Llama-3 8B Uncensored / roleplay | 8B | 4.8 GB | 3.5 GB | 8K | Llama-3 | chatuncensored | B | Classic uncensored 8B; short 8K context. |
| Granite 4.1 8B Granite (IBM) | 8B | 4.8 GB | 3.5 GB | 128K | Apache-2.0 | chatcoding | B | Dense 8B that matches IBM's prior 32B MoE flagship; RAG / tools. |
| Hermes 3 Llama 3.1 8B Uncensored / roleplay | 8B | 4.8 GB | 3.5 GB | 128K | Llama-3.1 | chatroleplay | B | Steerable, low-refusal generalist / roleplay 8B. |
| Llama 3.1 8B Llama | 8B | 4.8 GB | 3.5 GB | 128K | Llama-3.1 | chatmultilingual | B | Ubiquitous 8B baseline; huge finetune ecosystem. |
| Llama 3.1 8B Abliterated Uncensored / roleplay | 8B | 4.8 GB | 3.5 GB | 128K | Llama-3.1 | chatuncensored | B | Abliterated Llama 8B; broad tooling support. |
| MiniCPM-V 4.5 MiniCPM-V | 8B | 4.8 GB | 3.5 GB | 32K | Apache-2.0 | chatvision | B | Pocket-sized 8B multimodal (image + high-FPS video) on a Qwen3-8B base. |
| Ministral 8B Mistral | 8B | 4.8 GB | 3.5 GB | 128K | Mistral Research License | chatmultilingual | B | Edge model on a research license (non-commercial). |
| Qwen2.5-Coder 7B Qwen2.5-Coder | 7.6B | 4.6 GB | 3.3 GB | 128K | Apache-2.0 | coding | B | Fast autocomplete / small-repo coder for 8GB. |
| Mistral 7B v0.3 Mistral | 7.2B | 4.3 GB | 3.2 GB | 32K | Apache-2.0 | chat | B | The classic 7B; light and fast. |
| Gemma 3 4B Gemma 3 | 4.3B | 2.6 GB | 1.9 GB | 128K | Gemma | chatvision | B | Small multimodal model with image input. |
| Qwen3 4B Qwen3 | 4B | 2.4 GB | 1.8 GB | 128K | Apache-2.0 | chatmultilingual | B | Punches above its weight; runs on almost anything. |
| Phi-4-mini 3.8B Phi | 3.8B | 2.3 GB | 1.7 GB | 128K | MIT | reasoningchat | B | Compact reasoning model with 128K context. |
| Granite 4.1 3B Granite (IBM) | 3B | 1.8 GB | 1.3 GB | 128K | Apache-2.0 | chat | C | Edge / enterprise 3B for cheap on-device RAG and tools. |
| Gemma 3 1B Gemma 3 | 1B | 0.7 GB | 0.5 GB | 32K | Gemma | chat | C | Edge / draft model (text only). |
| No models match these filters. | ||||||||
Quality tiers are editorial, not benchmarks. S flagship · A excellent · B solid · C niche/edge. They reflect our honest judgment for local use as of 2026, not measured scores, always verify against your own workload. VRAM figures are weight-only footprints; add KV-cache and overhead with the VRAM calculator.
Skip the model hunt, Ember ships her own brain
Skip the model hunt. Ember ships her own tuned uncensored brain as a ~6.6 GB download that installs itself. No GGUF wrangling, no quant guesswork, no VRAM math.
- Runs on your own GPU · no cloud, no telemetry, no account
- $29 paid once in crypto, yours forever · 30-day refund
- Uncensored 18+ · voice, memory, she starts conversations
How to pick a model for your machine
Start from VRAM, not vibes. Find your card's memory, then read down the VRAM @ Q4 column for models that land under it with a little headroom for context. On 8GB you're in 7-8B territory (Qwen3 8B, Llama 3.1 8B); on 12-16GB the 12-14B class opens up (Qwen3 14B, Gemma 3 12B, gpt-oss 20B); on 24GB the 24-32B sweet spot is yours (Qwen3 32B, Qwen2.5-Coder 32B, Mistral Small 3.2 24B); and at 48GB+ the 70B dense models and larger MoEs become practical.
Then narrow by job. For code, see the best local coding model by VRAM. For characters and long scenes, the best local roleplay models. To understand the license labels and MoE-vs-dense trade-offs, read the open-weight model families of 2026. Confused by the quant column? The GGUF quantization cheat sheet explains why Q4_K_M is the default here. And for the single most-searched capacity, the best local LLMs for 24GB goes deep on the 24GB tier.
Local model database FAQ
What is the best local LLM in 2026?
It depends on your VRAM. On a 24GB card, Qwen3 32B and Qwen2.5-Coder 32B are the strongest all-rounders, with DeepSeek-R1-Distill Qwen 32B for reasoning. On 12-16GB, Qwen3 14B, Gemma 3 12B and gpt-oss 20B lead. Under 8GB, Qwen3 8B and Llama 3.1 8B are the safe picks. Use the GPU filter above to see exactly which of these fit your card at Q4.
How much VRAM does each model need?
As a rule of thumb the weights take parameters × 4.8 ÷ 8 GB at Q4_K_M (the local sweet spot), plus a little KV-cache and overhead. So an 8B model needs ~5GB, a 32B needs ~20GB, and a 70B needs ~42GB. Mixture-of-Experts models cost their TOTAL parameter count in VRAM even though only the active experts run per token. The VRAM @ Q4 column above is that weight figure; use the VRAM calculator for an exact fit check.
What does the quality tier mean, is it a benchmark score?
No. The S/A/B/C tier is a directional, editorial ranking for LOCAL use as of mid-2026, our honest judgment of how useful each model is on your own hardware, not a measured benchmark. Public leaderboards shift weekly and are often contaminated, so we deliberately do not print fabricated benchmark numbers. Treat the tier as a starting point and verify against your own workload.
Which local models are uncensored or good for roleplay?
Filter by the 'Uncensored' or 'Roleplay' chips above. The strongest current picks are Cydonia 24B and the Dolphin Mistral 24B Venice edition (both 24B, fit a 24GB card), with Rocinante 12B for 12GB and abliterated Qwen3/Llama 8B variants for 8GB. Abliterated models have had their refusal behaviour removed; roleplay finetunes are tuned for character consistency and longer scenes.
Can I use these models commercially?
Most can. Apache-2.0 and MIT models (Qwen, Mistral, Phi, gpt-oss, GLM, Yi, the DeepSeek-R1 distills) allow commercial use. Llama and Gemma models are permissive under their community licenses. The exceptions flagged as non-commercial are Cohere's Command-R / Command-R+ (CC-BY-NC) and Ministral 8B (research license). Filter by license class above and always read the actual license before shipping.
VRAM figures are weight-only footprints computed from published parameter counts; quality tiers are editorial and directional. Specs web-verified 2026-07, verify against the model card before you build. We link to tools we make. No ads, no trackers.