Database

Local LLM Model Database

Every local model worth running in 2026 in one sortable table, real VRAM at Q4, context window, license, use-case tags and a directional quality tier. Search it, filter by use case, and pick your GPU to grey out anything that won't fit. No cloud-only models, no fabricated benchmarks.

Most "best local LLM" lists are frozen the day they're published, or they rank cloud APIs you can't actually download. This one is different: 64 genuinely open-weight, runs-on-your-hardware models, Qwen3, Llama, Gemma 3, Mistral, DeepSeek, gpt-oss, GLM, Command-R and the uncensored/roleplay finetunes, with the one number that decides everything, how much VRAM the weights take at Q4_K_M. Pair it with the VRAM calculator for an exact fit check on your card.

Reading the table: VRAM @ Q4 is the weight footprint at Q4_K_M (the quality sweet spot); Min VRAM is a tighter Q3/low estimate for squeezing a model onto a smaller card. Mixture-of-Experts rows show total · active params, the total drives VRAM (every expert lives in memory), the active count drives speed. The Quality column is an editorial tier, explained below.

Use case
Size

64 of 64 models

Local LLMs by VRAM, context, license and directional quality tier. Sortable and filterable.
total / active Q4_K_M weights, GB Q3 / low, GB Type editorial tier Notes
DeepSeek V3 DeepSeek V3 · MoE 671B · 37B active 403 GB 294 GB 128K DeepSeek chatreasoningcoding S 671B MoE, server-class, listed for reference.
DeepSeek V3.2 DeepSeek V3 · MoE 671B · 37B active 403 GB 294 GB 128K DeepSeek chatreasoningcoding S Sparse-attention refresh of V3; still server-class.
Qwen3-Coder 480B-A35B Qwen3-Coder · MoE 480B · 35B active 288 GB 210 GB 256K Apache-2.0 codingreasoning S Agentic-coding SOTA, cloud-scale VRAM, not consumer-local.
Qwen3 235B-A22B Qwen3 · MoE 235B · 22B active 141 GB 103 GB 256K Apache-2.0 chatreasoningmultilingual S Flagship MoE for multi-GPU / server rig territory.
gpt-oss 120B gpt-oss (MoE) · MoE 116.8B · 5.1B active 63 GB 59 GB 128K Apache-2.0 reasoningchat S MXFP4 native; fits ~64-80GB (or one 96GB card).
GLM-4.5-Air GLM · MoE 106B · 12B active 64 GB 46 GB 128K MIT chatreasoningcoding S Agentic MoE; ~64GB at Q4, only ~12B active per token.
DeepSeek-R1-Distill Llama 70B DeepSeek-R1 (distill) 70.6B 42 GB 31 GB 128K MIT reasoning S R1 reasoning distilled into a Llama 70B body.
Llama 3.3 70B Llama 70.6B 42 GB 31 GB 128K Llama-3.3 chatreasoningmultilingual S 70B quality; needs ~48GB or a dual-GPU / offload setup.
Qwen3 32B Qwen3 32.8B 20 GB 14 GB 128K Apache-2.0 chatreasoningmultilingual S Best all-round dense model for a 24GB card.
Qwen2.5-Coder 32B Qwen2.5-Coder 32.5B 20 GB 14 GB 128K Apache-2.0 coding S Top open coding model that fits a single 24GB GPU.
GLM-5.2 GLM · MoE 743B · 39B active 446 GB 325 GB 1024K MIT chatreasoningcoding A 743B MoE frontier model, open weights on HF, cloud-only convenience on Ollama; server-scale VRAM.
DeepSeek-V4-Flash DeepSeek V4 · MoE 284B · 13B active 170 GB 124 GB 1000K MIT chatreasoningcoding A 284B MoE, ~13B active, 1M context, server-class VRAM, listed for reference.
Mixtral 8x22B Mixtral (MoE) · MoE 141B · 39B active 85 GB 62 GB 64K Apache-2.0 chatreasoning A 141B in VRAM, fast per token, but server-class memory.
Mistral Medium 3.5 128B Mistral 128B 77 GB 56 GB 256K Modified MIT chatcodingreasoningvisionmultilingual A 128B dense flagship, multimodal, needs a multi-GPU / big-rig setup.
Command R+ 104B Cohere Command-R 104B 62 GB 46 GB 128K CC-BY-NC-4.0 chatreasoning A RAG / tool-use specialist; non-commercial license.
Llama 3.1 70B Llama 70.6B 42 GB 31 GB 128K Llama-3.1 chatmultilingual A The workhorse 70B before 3.3; still very capable.
Qwen3.6 35B-A3B Qwen3.6 · MoE 35B · 3B active 21 GB 15 GB 256K Apache-2.0 chatcodingreasoningvisionmultilingual A Agentic-coding MoE, 24GB-friendly, only ~3B active per token.
Laguna XS.2 Laguna (Poolside) · MoE 33B · 3B active 20 GB 14 GB 256K Apache-2.0 codingreasoning A Poolside 33B-A3B coding MoE (68% SWE-bench Verified); ~36GB local.
DeepSeek-R1-Distill Qwen 32B DeepSeek-R1 (distill) 32.8B 20 GB 14 GB 128K MIT reasoning A Best reasoning model that fits a 24GB card.
GLM-4 32B GLM 32B 19 GB 14 GB 32K MIT chatcodingreasoning A Dense 32B strong at code & function-calling.
Gemma 4 31B Gemma 4 31B 19 GB 14 GB 256K Gemma chatvisionmultilingual A Dense multimodal flagship for a 24GB card; 256K context.
Qwen3 30B-A3B Qwen3 · MoE 30.5B · 3.3B active 18 GB 13 GB 256K Apache-2.0 chatreasoningmultilingual A Fast MoE, 24GB-friendly, only ~3B active per token.
Qwen3-Coder 30B-A3B Qwen3-Coder · MoE 30.5B · 3.3B active 18 GB 13 GB 256K Apache-2.0 codingreasoning A Local agentic coder, 24GB-friendly MoE.
North Mini Code 1.0 Cohere North · MoE 30B · 3B active 18 GB 13 GB 256K Apache-2.0 codingreasoning A Cohere's 30B-A3B agentic-coding MoE; runs on one 24-48GB GPU.
Gemma 3 27B Gemma 3 27.4B 16 GB 12 GB 128K Gemma chatvisionmultilingual A Strong multimodal all-rounder for a 24GB card.
Qwen3.6 27B Qwen3.6 27B 16 GB 12 GB 256K Apache-2.0 chatcodingreasoningmultilingual A Dense 27B all-rounder for a 24GB card; thinking + coding.
Gemma 4 26B-A4B Gemma 4 · MoE 26B · 4B active 16 GB 11 GB 256K Gemma chatvisionmultilingual A MoE with all experts in VRAM (~16GB), only ~4B active per token.
Cydonia 24B Uncensored / roleplay 23.6B 14 GB 10 GB 32K Apache-2.0 roleplayuncensoredchat A Uncensored roleplay finetune of Mistral Small.
Dolphin Mistral 24B Venice Uncensored / roleplay 23.6B 14 GB 10 GB 32K Apache-2.0 chatuncensoredroleplay A Uncensored assistant; Venice edition of Dolphin.
Mistral Small 3.2 24B Mistral 23.6B 14 GB 10 GB 128K Apache-2.0 chatvision A 24GB sweet spot with vision and a permissive license.
gpt-oss 20B gpt-oss (MoE) · MoE 20.9B · 3.6B active 12 GB 10 GB 128K Apache-2.0 reasoningchat A Ships MXFP4 4-bit natively at a ~12GB footprint, runs on 16GB.
Qwen3 14B Qwen3 14.8B 9 GB 6.5 GB 128K Apache-2.0 chatreasoningmultilingual A Strong 14B that fits comfortably in 12-16GB.
Phi-4 14B Phi 14.7B 9 GB 6.4 GB 16K MIT reasoning A Reasoning-dense 14B; note the shorter 16K context.
Qwen2.5-Coder 14B Qwen2.5-Coder 14.7B 9 GB 6.4 GB 128K Apache-2.0 coding A Best coder for 12-16GB.
Mixtral 8x7B Mixtral (MoE) · MoE 46.7B · 12.9B active 28 GB 20 GB 32K Apache-2.0 chat B All 8 experts sit in VRAM; ~13B active per token.
Command R 35B Cohere Command-R 35B 21 GB 15 GB 128K CC-BY-NC-4.0 chat B Grounded RAG on a 24GB card; non-commercial.
Ornith 1.0 35B-A3B Ornith · MoE 35B · 3B active 21 GB 15 GB 256K MIT codingreasoning B Self-scaffolding agentic-coding MoE (Qwen3.5 base); ~3B active.
Yi-1.5 34B Yi 34.4B 21 GB 15 GB 32K Apache-2.0 chatmultilingual B Solid bilingual 34B; shorter native context.
Granite 4.1 30B Granite (IBM) 30B 18 GB 13 GB 128K Apache-2.0 chatcoding B Enterprise dense 30B tuned for RAG, tool-use and function-calling.
Nemotron 3 Nano Omni 30B-A3B Nemotron 3 (NVIDIA) · MoE 30B · 3B active 18 GB 13 GB 128K NVIDIA Open Model chatvisionreasoning B Omni-modal (text / image / audio / video) MoE; ~3B active per token.
DeepSeek-R1-Distill Qwen 14B DeepSeek-R1 (distill) 14.8B 9 GB 6.5 GB 128K MIT reasoning B Chain-of-thought reasoning for 12-16GB.
Gemma 3 12B Gemma 3 12.2B 7.3 GB 5.3 GB 128K Gemma chatvisionmultilingual B Multimodal 12B for 12-16GB.
Mistral NeMo 12B Mistral 12.2B 7.3 GB 5.3 GB 128K Apache-2.0 chatmultilingual B Popular 12B base for finetunes; 128K context.
Rocinante 12B Uncensored / roleplay 12.2B 7.3 GB 5.3 GB 128K Apache-2.0 roleplayuncensoredchat B Popular 12B roleplay finetune (Mistral NeMo).
Gemma 4 12B Gemma 4 12B 7.3 GB 5.3 GB 256K Gemma chatvisionmultilingual B Multimodal 12B for 12-16GB; 256K context.
Ornith 1.0 9B Ornith 9B 5.4 GB 3.9 GB 256K MIT coding B DeepReinforce self-scaffolding coder; edge-friendly 9B dense.
LFM2.5-8B-A1B LiquidAI LFM · MoE 8.3B · 1.5B active 5 GB 3.6 GB 128K LFM Open License chatcoding B On-device MoE (~1.5B active) built for fast, reliable tool-calling.
Qwen3 8B Qwen3 8.2B 5 GB 3.6 GB 128K Apache-2.0 chatmultilingual B Great daily driver for 8GB cards.
Qwen3 8B Abliterated Uncensored / roleplay 8.2B 5 GB 3.6 GB 128K Apache-2.0 chatuncensored B Refusal-removed Qwen3 8B for an 8GB card.
DeepSeek-R1-Distill Llama 8B DeepSeek-R1 (distill) 8B 4.8 GB 3.5 GB 128K MIT reasoning B Reasoning traces on an 8GB budget.
Dolphin 2.9 Llama-3 8B Uncensored / roleplay 8B 4.8 GB 3.5 GB 8K Llama-3 chatuncensored B Classic uncensored 8B; short 8K context.
Granite 4.1 8B Granite (IBM) 8B 4.8 GB 3.5 GB 128K Apache-2.0 chatcoding B Dense 8B that matches IBM's prior 32B MoE flagship; RAG / tools.
Hermes 3 Llama 3.1 8B Uncensored / roleplay 8B 4.8 GB 3.5 GB 128K Llama-3.1 chatroleplay B Steerable, low-refusal generalist / roleplay 8B.
Llama 3.1 8B Llama 8B 4.8 GB 3.5 GB 128K Llama-3.1 chatmultilingual B Ubiquitous 8B baseline; huge finetune ecosystem.
Llama 3.1 8B Abliterated Uncensored / roleplay 8B 4.8 GB 3.5 GB 128K Llama-3.1 chatuncensored B Abliterated Llama 8B; broad tooling support.
MiniCPM-V 4.5 MiniCPM-V 8B 4.8 GB 3.5 GB 32K Apache-2.0 chatvision B Pocket-sized 8B multimodal (image + high-FPS video) on a Qwen3-8B base.
Ministral 8B Mistral 8B 4.8 GB 3.5 GB 128K Mistral Research License chatmultilingual B Edge model on a research license (non-commercial).
Qwen2.5-Coder 7B Qwen2.5-Coder 7.6B 4.6 GB 3.3 GB 128K Apache-2.0 coding B Fast autocomplete / small-repo coder for 8GB.
Mistral 7B v0.3 Mistral 7.2B 4.3 GB 3.2 GB 32K Apache-2.0 chat B The classic 7B; light and fast.
Gemma 3 4B Gemma 3 4.3B 2.6 GB 1.9 GB 128K Gemma chatvision B Small multimodal model with image input.
Qwen3 4B Qwen3 4B 2.4 GB 1.8 GB 128K Apache-2.0 chatmultilingual B Punches above its weight; runs on almost anything.
Phi-4-mini 3.8B Phi 3.8B 2.3 GB 1.7 GB 128K MIT reasoningchat B Compact reasoning model with 128K context.
Granite 4.1 3B Granite (IBM) 3B 1.8 GB 1.3 GB 128K Apache-2.0 chat C Edge / enterprise 3B for cheap on-device RAG and tools.
Gemma 3 1B Gemma 3 1B 0.7 GB 0.5 GB 32K Gemma chat C Edge / draft model (text only).

Quality tiers are editorial, not benchmarks. S flagship · A excellent · B solid · C niche/edge. They reflect our honest judgment for local use as of 2026, not measured scores, always verify against your own workload. VRAM figures are weight-only footprints; add KV-cache and overhead with the VRAM calculator.

Runs on your GPU · paid once · 18+

Skip the model hunt, Ember ships her own brain

Skip the model hunt. Ember ships her own tuned uncensored brain as a ~6.6 GB download that installs itself. No GGUF wrangling, no quant guesswork, no VRAM math.

  • Runs on your own GPU · no cloud, no telemetry, no account
  • $29 paid once in crypto, yours forever · 30-day refund
  • Uncensored 18+ · voice, memory, she starts conversations
See Ember →

How to pick a model for your machine

Start from VRAM, not vibes. Find your card's memory, then read down the VRAM @ Q4 column for models that land under it with a little headroom for context. On 8GB you're in 7-8B territory (Qwen3 8B, Llama 3.1 8B); on 12-16GB the 12-14B class opens up (Qwen3 14B, Gemma 3 12B, gpt-oss 20B); on 24GB the 24-32B sweet spot is yours (Qwen3 32B, Qwen2.5-Coder 32B, Mistral Small 3.2 24B); and at 48GB+ the 70B dense models and larger MoEs become practical.

Then narrow by job. For code, see the best local coding model by VRAM. For characters and long scenes, the best local roleplay models. To understand the license labels and MoE-vs-dense trade-offs, read the open-weight model families of 2026. Confused by the quant column? The GGUF quantization cheat sheet explains why Q4_K_M is the default here. And for the single most-searched capacity, the best local LLMs for 24GB goes deep on the 24GB tier.

Local model database FAQ

What is the best local LLM in 2026?

It depends on your VRAM. On a 24GB card, Qwen3 32B and Qwen2.5-Coder 32B are the strongest all-rounders, with DeepSeek-R1-Distill Qwen 32B for reasoning. On 12-16GB, Qwen3 14B, Gemma 3 12B and gpt-oss 20B lead. Under 8GB, Qwen3 8B and Llama 3.1 8B are the safe picks. Use the GPU filter above to see exactly which of these fit your card at Q4.

How much VRAM does each model need?

As a rule of thumb the weights take parameters × 4.8 ÷ 8 GB at Q4_K_M (the local sweet spot), plus a little KV-cache and overhead. So an 8B model needs ~5GB, a 32B needs ~20GB, and a 70B needs ~42GB. Mixture-of-Experts models cost their TOTAL parameter count in VRAM even though only the active experts run per token. The VRAM @ Q4 column above is that weight figure; use the VRAM calculator for an exact fit check.

What does the quality tier mean, is it a benchmark score?

No. The S/A/B/C tier is a directional, editorial ranking for LOCAL use as of mid-2026, our honest judgment of how useful each model is on your own hardware, not a measured benchmark. Public leaderboards shift weekly and are often contaminated, so we deliberately do not print fabricated benchmark numbers. Treat the tier as a starting point and verify against your own workload.

Which local models are uncensored or good for roleplay?

Filter by the 'Uncensored' or 'Roleplay' chips above. The strongest current picks are Cydonia 24B and the Dolphin Mistral 24B Venice edition (both 24B, fit a 24GB card), with Rocinante 12B for 12GB and abliterated Qwen3/Llama 8B variants for 8GB. Abliterated models have had their refusal behaviour removed; roleplay finetunes are tuned for character consistency and longer scenes.

Can I use these models commercially?

Most can. Apache-2.0 and MIT models (Qwen, Mistral, Phi, gpt-oss, GLM, Yi, the DeepSeek-R1 distills) allow commercial use. Llama and Gemma models are permissive under their community licenses. The exceptions flagged as non-commercial are Cohere's Command-R / Command-R+ (CC-BY-NC) and Ministral 8B (research license). Filter by license class above and always read the actual license before shipping.

VRAM figures are weight-only footprints computed from published parameter counts; quality tiers are editorial and directional. Specs web-verified 2026-07, verify against the model card before you build. We link to tools we make. No ads, no trackers.