Best local AI model for coding

Coding is the one workload where local models are genuinely competitive - and the category is narrow enough to be honest about it. Six models worth your disk space, verified against Hugging Face on 2026-10-07. General-purpose picks live on best local AI model.

Quick answer

24 GB card: Qwen2.5-Coder-32B-Instruct for quality; Qwen3-Coder-30B-A3B if you want agentic speed (only 3.3B active params).16 GB: Devstral Small 2507, purpose-tuned for software engineering.12 GB: DeepSeek-Coder-V2-Lite (16B MoE, 2.4B active).8 GB: no dedicated coder fits well - use Qwen3 8B as a general model.

The table

ModelParamsLicenseReleased4-bit VRAM (est)Best for
Qwen2.5 Coder 32B Instruct32BApache-2.02024-11~21 GBThe local coding workhorse; top quality under 24 GB
Qwen3 Coder 30B-A3B Instruct30B MoE (3.3B active)Apache-2.02025-07~21 GBAgentic coding; fast for its size thanks to few active params
Devstral Small 250724BApache-2.02025-07~16 GBFine-tuned for agentic software engineering (SWE tasks)
DeepSeek Coder V2 Lite Instruct16B MoE (2.4B active)DeepSeek license2024-06~11.5 GBFast coding MoE; good on mid-range cards
StarCoder2 15B15BBigCode OpenRAIL-M2024-02~11 GBTrained on permissively-licensed code only
Qwen3 8B8BApache-2.02025-04~6.5 GBNot a coder model, but the best 8 GB general fallback

VRAM figures are estimates (0.6-0.7 GB per billion params at 4-bit + ~1.5 GB context overhead). MoE models still need VRAM for all parameters; the "active" count affects speed, not memory.

How to choose

For coding, dense 32B models set the quality ceiling on a single card, but MoE coders punch above their speed class: Qwen3-Coder-30B-A3B only activates ~3.3B parameters per token, so it generates far faster than a dense 30B while needing the same ~21 GB of VRAM. If your workflow is agentic (long tool-use loops, file editing, multi-step tasks), that speed difference is the feature. If you mostly chat about code or want the safest quality, the dense Qwen2.5-Coder-32B remains the reference point and is the most-downloaded coder on this list.

Devstral Small 2507 is worth singling out: Mistral fine-tuned it specifically for agentic software engineering - navigating codebases, editing files, using tools - rather than raw completion quality. On 16 GB it's the strongest option in this table for real repo work.

StarCoder2 is older and weaker, but it is the only model here trained exclusively on permissively-licensed code with governance documentation - relevant if license provenance matters to you. Its OpenRAIL-M license adds use restrictions the Apache-2.0 rows don't have.

Where Ajax fits

PewDiePie's Ajax (a Qwen 3.5 9B fine-tune for the Odysseus workspace, not yet released) is built as a general always-on agent, not a coding specialist. It may end up decent at scripting tasks given its Qwen lineage, but if coding is the job, the models above are trained for it. Ajax's appeal is the Odysseus integration - see Ajax vs Odysseus.

Caveats

Coding ability doesn't reduce to a spec sheet: eval benchmarks measure different things than your stack, your prompts, and your patience. Treat the table as a shortlist - download two candidates on the same day, run both on your actual repo, keep the winner. All of them run in Ollama or LM Studio; see run locally for the general setup path.

Related: best local AI model, Ajax requirements, Ajax status.