Best local AI model for coding
Coding is the one workload where local models are genuinely competitive - and the category is narrow enough to be honest about it. Six models worth your disk space, verified against Hugging Face on 2026-10-07. General-purpose picks live on best local AI model.
Quick answer
24 GB card: Qwen2.5-Coder-32B-Instruct for quality; Qwen3-Coder-30B-A3B if you want agentic speed (only 3.3B active params).16 GB: Devstral Small 2507, purpose-tuned for software engineering.12 GB: DeepSeek-Coder-V2-Lite (16B MoE, 2.4B active).8 GB: no dedicated coder fits well - use Qwen3 8B as a general model.
The table
| Model | Params | License | Released | 4-bit VRAM (est) | Best for |
|---|---|---|---|---|---|
| Qwen2.5 Coder 32B Instruct | 32B | Apache-2.0 | 2024-11 | ~21 GB | The local coding workhorse; top quality under 24 GB |
| Qwen3 Coder 30B-A3B Instruct | 30B MoE (3.3B active) | Apache-2.0 | 2025-07 | ~21 GB | Agentic coding; fast for its size thanks to few active params |
| Devstral Small 2507 | 24B | Apache-2.0 | 2025-07 | ~16 GB | Fine-tuned for agentic software engineering (SWE tasks) |
| DeepSeek Coder V2 Lite Instruct | 16B MoE (2.4B active) | DeepSeek license | 2024-06 | ~11.5 GB | Fast coding MoE; good on mid-range cards |
| StarCoder2 15B | 15B | BigCode OpenRAIL-M | 2024-02 | ~11 GB | Trained on permissively-licensed code only |
| Qwen3 8B | 8B | Apache-2.0 | 2025-04 | ~6.5 GB | Not a coder model, but the best 8 GB general fallback |
VRAM figures are estimates (0.6-0.7 GB per billion params at 4-bit + ~1.5 GB context overhead). MoE models still need VRAM for all parameters; the "active" count affects speed, not memory.
How to choose
For coding, dense 32B models set the quality ceiling on a single card, but MoE coders punch above their speed class: Qwen3-Coder-30B-A3B only activates ~3.3B parameters per token, so it generates far faster than a dense 30B while needing the same ~21 GB of VRAM. If your workflow is agentic (long tool-use loops, file editing, multi-step tasks), that speed difference is the feature. If you mostly chat about code or want the safest quality, the dense Qwen2.5-Coder-32B remains the reference point and is the most-downloaded coder on this list.
Devstral Small 2507 is worth singling out: Mistral fine-tuned it specifically for agentic software engineering - navigating codebases, editing files, using tools - rather than raw completion quality. On 16 GB it's the strongest option in this table for real repo work.
StarCoder2 is older and weaker, but it is the only model here trained exclusively on permissively-licensed code with governance documentation - relevant if license provenance matters to you. Its OpenRAIL-M license adds use restrictions the Apache-2.0 rows don't have.
Where Ajax fits
PewDiePie's Ajax (a Qwen 3.5 9B fine-tune for the Odysseus workspace, not yet released) is built as a general always-on agent, not a coding specialist. It may end up decent at scripting tasks given its Qwen lineage, but if coding is the job, the models above are trained for it. Ajax's appeal is the Odysseus integration - see Ajax vs Odysseus.
Caveats
Coding ability doesn't reduce to a spec sheet: eval benchmarks measure different things than your stack, your prompts, and your patience. Treat the table as a shortlist - download two candidates on the same day, run both on your actual repo, keep the winner. All of them run in Ollama or LM Studio; see run locally for the general setup path.
Related: best local AI model, Ajax requirements, Ajax status.