Models small enough to run on your own laptop, scored on the same benchmarks.

Run it on your own laptop

These models are small enough to download and run yourself with Ollama or vLLM - free, offline, and private. The scores are the same Artificial Analysis benchmarks used everywhere else on this site, so they compare directly with the frontier models.

Measured on your machine: ThinkPad L390 Yoga · Intel Core i5-8365U (4 cores, no discrete GPU) · 64GB RAM

Your memory is not the limit - 64GB is far more than any of these need. The limit is the processor. With no separate graphics card, every word is produced on four CPU cores, and speed falls as the model grows. That is why a 30B model loads fine and is still unusable, while a 4B model is genuinely pleasant.

INSTALL THIS ONE

K2 Horizon 3.7B

3.7Bparameters
11/swords per second
16.2intelligence
26.1coding

It scores 16.2 on intelligence and 26.1 on coding, and runs at about 11 words per second on your machine - fast enough to think with rather than wait for.

If you will tolerate slower: K2 Horizon 7B scores 21 at about 6 words per second.

Best small models · Intelligence

Artificial Analysis Intelligence Index · 14B and under · Higher is better

K2 Horizon 7B21K2 Horizon 3.7B16.2Gemma 4 12B (Reasoning)14.2Qwen3.5 9B (Reasoning)13.7Qwen3.5 9B (Non-reasoning)13.3MiniCPM5-2B13.1Qwen3.5 4B (Reasoning)13.1Granite 4.2 8B11.8G9v3-3B10.8Qwen3.5 4B (Non-reasoning)10.8
K2 Horizon 7B
K2 Horizon 3.7B
Gemma 4 12B (R)
Qwen3.5 9B (R)
Qwen3.5 9B (NR)
MiniCPM5-2B
Qwen3.5 4B (R)
4.2 8B
G9v3-3B
Qwen3.5 4B (NR)

Best small models · Coding

Artificial Analysis Coding Index · 14B and under · Higher is better

K2 Horizon 7B38.6Gemma 4 12B (Reasoning)31Qwen3.5 9B (Reasoning)28.7K2 Horizon 3.7B26.1Qwen3.5 9B (Non-reasoning)23.5Qwen3.5 4B (Reasoning)22.6Granite 4.2 8B22.4Qwen3.5 4B (Non-reasoning)20.3Granite 4.2 3B17.5MiniCPM5-2B14.5
K2 Horizon 7B
Gemma 4 12B (R)
Qwen3.5 9B (R)
K2 Horizon 3.7B
Qwen3.5 9B (NR)
Qwen3.5 4B (R)
4.2 8B
Qwen3.5 4B (NR)
4.2 3B
MiniCPM5-2B

Best small models · Agentic

Agentic score - tool use and multi-step work · 14B and under · Higher is better

K2 Horizon 7B27.6K2 Horizon 3.7B18.4Granite 4.2 3B0.9Qwen3 14B (Reasoning)0.9Gemma 3 12B Instruct0.1
K2 Horizon 7B
K2 Horizon 3.7B
4.2 3B
Qwen3 14B (R)
Gemma 3 12B Instruct
Show up to
ModelSizeSpeed on your laptopbandwidth ÷ memoryMemorysize × 0.6 GBIntelligenceCodingAgentic
Qwen3.8 27B (xhigh)27B2/s Too slow to enjoy17.7 GB33.968.1-
Qwen3.8 27B (medium)27B2/s Too slow to enjoy17.7 GB27.856.1-
Qwen3.8 27B (low)27B2/s Too slow to enjoy17.7 GB26.558.2-
Qwen3.5 27B (Reasoning)27B2/s Too slow to enjoy17.7 GB22.9--
Qwen3.8 27B (Non-reasoning)27B2/s Too slow to enjoy17.7 GB22.444.6-
Qwen3.6 27B (Reasoning)27B2/s Too slow to enjoy17.7 GB21.953.7-
K2 Horizon 7B7B6/s Usable5.7 GB2138.627.6
Qwen3.6 27B (Non-reasoning)27B2/s Too slow to enjoy17.7 GB19.846.6-
Qwen3.5 27B (Non-reasoning)27B2/s Too slow to enjoy17.7 GB19.4--
Gemma 4 26B A4B (Reasoning)26B2/s Too slow to enjoy17.1 GB16.739.3-
K2 Horizon 3.7B3.7B11/s Comfortable3.7 GB16.226.118.4
Gemma 4 31B (Reasoning)31B1/s Too slow to enjoy20.1 GB15.443.4-
Granite 4.2 30B30B1/s Too slow to enjoy19.5 GB14.829.9-
Gemma 4 12B (Reasoning)12B3/s Slow8.7 GB14.231-
Gemma 4 31B (Non-reasoning)31B1/s Too slow to enjoy20.1 GB13.933.2-
Apriel-v1.5-15B-Thinker15B3/s Slow10.5 GB13.8--
Qwen3.5 9B (Reasoning)9B5/s Usable6.9 GB13.728.7-
Apriel-v1.6-15B-Thinker15B3/s Slow10.5 GB13.4--
Qwen3.5 9B (Non-reasoning)9B5/s Usable6.9 GB13.323.5-
MiniCPM5-2B2B21/s Comfortable2.7 GB13.114.5-
Gemma 4 26B A4B (Non-reasoning)26B2/s Too slow to enjoy17.1 GB13.1--
Qwen3.5 4B (Reasoning)4B10/s Comfortable3.9 GB13.122.6-
Qwen3 VL 32B (Reasoning)32B1/s Too slow to enjoy20.7 GB11.9--
Granite 4.2 8B8B5/s Usable6.3 GB11.822.4-
Nemotron Cascade 2 30B A3B30B1/s Too slow to enjoy19.5 GB11.725.3-
HyperCLOVA X SEED Think (32B)32B1/s Too slow to enjoy20.7 GB11.4--
G9v3-3B3B14/s Comfortable3.3 GB10.89.9-
Qwen3.5 4B (Non-reasoning)4B10/s Comfortable3.9 GB10.820.3-
Nemotron 3 Nano Omni 30B A3B Reasoning30B1/s Too slow to enjoy19.5 GB10.313.8-
gpt-oss-20b (low)20B2/s Slow13.5 GB10--
Qwen3 30B A3B 2507 (Reasoning)30B1/s Too slow to enjoy19.5 GB9.812.1-
Tri-21B-think Preview21B2/s Too slow to enjoy14.1 GB9.6--
Qwen3 Coder 30B A3B Instruct30B1/s Too slow to enjoy19.5 GB9.6--
DiffusionGemma 26B A4B26B2/s Too slow to enjoy17.1 GB9.519.7-
QwQ 32B32B1/s Too slow to enjoy20.7 GB9.5--
Qwen3 VL 30B A3B (Reasoning)30B1/s Too slow to enjoy19.5 GB9.5--
Gemma 4 12B (Non-reasoning)12B3/s Slow8.7 GB9.4--
Motif-2-12.7B-Reasoning12.7B3/s Slow9.1 GB9.2--
Granite 4.2 3B3B14/s Comfortable3.3 GB9.117.50.9
gpt-oss-20b (high)20B2/s Slow13.5 GB920.71.4
Tri-21B-Think21B2/s Too slow to enjoy14.1 GB9--
NVIDIA Nemotron 3 Nano 30B A3B (Reasoning)30B1/s Too slow to enjoy19.5 GB8.914.4-
Qwen3 4B 2507 (Reasoning)4B10/s Comfortable3.9 GB8.8--
MiniCPM5-1B (Reasoning)1B42/s Comfortable2.1 GB8.8--
MiniCPM5-1B (Non-reasoning)1B42/s Comfortable2.1 GB8.7--
Nanbeige4.1-3B3B14/s Comfortable3.3 GB8.49.6-
LFM2.5-2.6B2.6B16/s Comfortable3.1 GB8.47.7-
Qwen3 VL 32B Instruct32B1/s Too slow to enjoy20.7 GB8.4--
DeepSeek R1 Distill Qwen 32B32B1/s Too slow to enjoy20.7 GB8.4--
EXAONE 4.0 32B (Reasoning)32B1/s Too slow to enjoy20.7 GB8.2--
Qwen3 VL 8B (Reasoning)8B5/s Usable6.3 GB8.2--
DeepSeek R1 0528 Qwen3 8B8B5/s Usable6.3 GB8.1--
Qwen3 VL 30B A3B Instruct30B1/s Too slow to enjoy19.5 GB7.9--
DeepSeek R1 Distill Qwen 14B14B3/s Slow9.9 GB7.8--
Falcon-H1R-7B7B6/s Usable5.7 GB7.8--
Qwen3 Omni 30B A3B (Reasoning)30B1/s Too slow to enjoy19.5 GB7.8--
Step3 VL 10B10B4/s Usable7.5 GB7.7--
Qwen3 30B A3B (Reasoning)30B1/s Too slow to enjoy19.5 GB7.6--
QwQ 32B-Preview32B1/s Too slow to enjoy20.7 GB7.6--
Qwen3 30B A3B 2507 Instruct30B1/s Too slow to enjoy19.5 GB7.5--
NVIDIA Nemotron Nano 12B v2 VL (Reasoning)12B3/s Slow8.7 GB7.5--
Granite 4.1 30B30B1/s Too slow to enjoy19.5 GB7.410.4-
NVIDIA Nemotron Nano 9B V2 (Reasoning)9B5/s Usable6.9 GB7.4--
NVIDIA Nemotron 3 Nano 4B4B10/s Comfortable3.9 GB7.48-
Qwen3 32B (Non-reasoning)32B1/s Too slow to enjoy20.7 GB7.3--
Llama 3.1 Nemotron Nano 4B v1.1 (Reasoning)4B10/s Comfortable3.9 GB7.3--
Qwen3 VL 8B Instruct8B5/s Usable6.3 GB7.3--
Qwen3 4B (Reasoning)4B10/s Comfortable3.9 GB7.2--
LFM2.5-8B-A1B8B5/s Usable6.3 GB7.2--
Qwen3 32B (Reasoning)32B1/s Too slow to enjoy20.7 GB7.215.3-
Olmo 3.1 32B Think32B1/s Too slow to enjoy20.7 GB7.1--
Qwen3 VL 4B (Reasoning)4B10/s Comfortable3.9 GB7--
Qwen3.5 2B (Reasoning)2B21/s Comfortable2.7 GB6.92.9-
Llama 3.1 Instruct 8B8B5/s Usable6.3 GB6.95.4-
Qwen2.5 Instruct 32B32B1/s Too slow to enjoy20.7 GB6.9--
NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning)30B1/s Too slow to enjoy19.5 GB6.8--
NVIDIA Nemotron Nano 9B V2 (Non-reasoning)9B5/s Usable6.9 GB6.8--
Qwen3 4B 2507 Instruct4B10/s Comfortable3.9 GB6.7--
Qwen2.5 Coder Instruct 32B32B1/s Too slow to enjoy20.7 GB6.7--
Qwen3 14B (Non-reasoning)14B3/s Slow9.9 GB6.7--
Qwen3 30B A3B (Non-reasoning)30B1/s Too slow to enjoy19.5 GB6.6--
Qwen3 4B (Non-reasoning)4B10/s Comfortable3.9 GB6.6--
Granite 4.1 8B8B5/s Usable6.3 GB6.69.5-
Sarvam 30B (high)30B1/s Too slow to enjoy19.5 GB6.6--
Olmo 3.1 32B Instruct32B1/s Too slow to enjoy20.7 GB6.5--
DeepSeek R1 Distill Llama 8B8B5/s Usable6.3 GB6.5--
Olmo 3 32B Think32B1/s Too slow to enjoy20.7 GB6.5--
Qwen3 14B (Reasoning)14B3/s Slow9.9 GB6.413.80.9
EXAONE 4.0 32B (Non-reasoning)32B1/s Too slow to enjoy20.7 GB6.3--
Qwen3.5 2B (Non-reasoning)2B21/s Comfortable2.7 GB6.22.4-
Gemini 1.5 Flash-8B8B5/s Usable6.3 GB6.2--
Qwen3.5 0.8B (Reasoning)0.8B52/s Comfortable2.0 GB6.10-
DeepHermes 3 - Mistral 24B Preview (Non-reasoning)24B2/s Too slow to enjoy15.9 GB6.1--
Ministral 3 14B14B3/s Slow9.9 GB614.4-
Qwen3 Omni 30B A3B Instruct30B1/s Too slow to enjoy19.5 GB6--
Qwen3 8B (Non-reasoning)8B5/s Usable6.3 GB6--
OLMo 2 32B32B1/s Too slow to enjoy20.7 GB6--
LFM2 24B A2B24B2/s Too slow to enjoy15.9 GB5.9--
Granite 4.1 3B3B14/s Comfortable3.3 GB5.94.7-
Phi-3 Mini Instruct 3.8B3.8B11/s Comfortable3.8 GB5.8--
NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning)12B3/s Slow8.7 GB5.8--
Qwen2.5 Coder Instruct 7B 7B6/s Usable5.7 GB5.8--
Llama 2 Chat 7B7B6/s Usable5.7 GB5.7--
Llama 3.2 Instruct 3B3B14/s Comfortable3.3 GB5.7--
MiniCPM-V 4.6 1.3B1.3B32/s Comfortable2.3 GB5.70.7-
Jamba Reasoning 3B3B14/s Comfortable3.3 GB5.7--
Qwen3 VL 4B Instruct4B10/s Comfortable3.9 GB5.7--
Olmo 3 7B Think7B6/s Usable5.7 GB5.6--
OLMo 2 7B7B6/s Usable5.7 GB5.6--
Molmo 7B-D7B6/s Usable5.7 GB5.6--
DeepSeek R1 Distill Qwen 1.5B1.5B28/s Comfortable2.4 GB5.5--
Ministral 3 8B8B5/s Usable6.3 GB5.59.7-
Llama 3.2 Instruct 11B (Vision)11B4/s Slow8.1 GB5.4--
Qwen3.5 0.8B (Non-reasoning)0.8B52/s Comfortable2.0 GB5.41.2-
Llama 2 Chat 13B13B3/s Slow9.3 GB5.3--
Exaone 4.0 1.2B (Reasoning)1.2B35/s Comfortable2.2 GB5.3--
Olmo 3 7B Instruct7B6/s Usable5.7 GB5.2--
Exaone 4.0 1.2B (Non-reasoning)1.2B35/s Comfortable2.2 GB5.2--
LFM2.5-1.2B-Thinking1.2B35/s Comfortable2.2 GB5.2--
LFM2 2.6B2.6B16/s Comfortable3.1 GB5.2--
LFM2.5-1.2B-Instruct1.2B35/s Comfortable2.2 GB5.2--
Granite 4.0 H 1B1B42/s Comfortable2.1 GB5.2--
Qwen3 1.7B (Reasoning)1.7B25/s Comfortable2.5 GB5.2--
Qwen3 8B (Reasoning)8B5/s Usable6.3 GB5.29-
DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning)8B5/s Usable6.3 GB5.1--
Mistral 7B Instruct7B6/s Usable5.7 GB5--
Qwen Chat 14B14B3/s Slow9.9 GB5--
Granite 4.0 1B1B42/s Comfortable2.1 GB5--
Molmo2-8B8B5/s Usable6.3 GB5--
LFM2 8B A1B8B5/s Usable6.3 GB4.9--
Granite 3.3 8B (Non-reasoning)8B5/s Usable6.3 GB4.9--
Qwen3 1.7B (Non-reasoning)1.7B25/s Comfortable2.5 GB4.9--
Gemma 3 27B Instruct27B2/s Too slow to enjoy17.7 GB4.910.10.1
Ministral 3 3B3B14/s Comfortable3.3 GB4.84.8-
Apertus 8B Instruct8B5/s Usable6.3 GB4.8--
Gemma 3 1B Instruct1B42/s Comfortable2.1 GB4.8--
Gemma 3 4B Instruct4B10/s Comfortable3.9 GB4.82.7-
LFM2 1.2B1.2B35/s Comfortable2.2 GB4.8--
LFM2.5-VL-1.6B1.6B26/s Comfortable2.5 GB4.8--
Llama 3 Instruct 8B8B5/s Usable6.3 GB4.8--
Llama 3.2 Instruct 1B1B42/s Comfortable2.1 GB4.8--
Qwen3 0.6B (Non-reasoning)0.6B69/s Comfortable1.9 GB4.8--
Qwen3 0.6B (Reasoning)0.6B69/s Comfortable1.9 GB4.8--
Gemma 3 12B Instruct12B3/s Slow8.7 GB3.85.80.1
K2 Horizon 0.9B0.9B46/s Comfortable2.0 GB33.4-
About the Agentic column: it is nearly empty, and that is the honest picture rather than missing data. Artificial Analysis has scored only three models this size on agentic ability, and all three came out near zero. Models small enough to run on a laptop are genuinely poor at driving tools through a multi-step job - which is exactly the work Claude Code does for you. Use these for writing and answering, not for agent workflows. An earlier version of this site showed flattering agentic numbers here; they were our own arithmetic, ran about 31 points too high against the real Artificial Analysis measure, and have been removed.

Sizes are read from the model name, which is how open-weight models are almost always labelled - Artificial Analysis does not publish parameter counts on its free tier. Speeds are estimated from model size and your memory bandwidth, not measured on your machine, so treat them as a guide. Expect a pause before the first word on long prompts; that part is slower than the numbers here suggest.