Best Local and On-Device LLMs
As of July 2026, the best local LLM to run offline depends almost entirely on how much memory you can give it, so this guide ranks picks by hardware tier.
Explore Small Language Models through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of Small Language Models.
Showing 1-16 of 16 articles
As of July 2026, the best local LLM to run offline depends almost entirely on how much memory you can give it, so this guide ranks picks by hardware tier.
As of July 2026, the two strongest small language models (open models under 15 billion parameters) are Google DeepMind's Gemma 4 12B and Alibaba's Qwen3.5-9B.
Gemma is a family of open-weight models developed by Google DeepMind. The family began in February 2024 with text-only decoder models derived from research used for Google's Gemini systems, then expanded to…
Gemma 2 is a family of open-weights large language models developed by Google DeepMind and released starting June 27, 2024, in three parameter sizes: 2 billion (2B), 9 billion (9B), and 27 billion (27B).
Gemma 3 is a family of open-weight large language models developed by Google DeepMind and released on March 12, 2025.
Gemma 3n is an open, mobile-first multimodal model from Google built to run locally on phones, tablets, laptops, and other resource-constrained hardware, accepting text, image, audio, and video as input and…
MiniCPM is a family of compact, openly licensed language models published by OpenBMB, the shared open-source brand of Tsinghua University's natural language processing lab (THUNLP) and the Beijing company…
MiniCPM5-2B is an open-weights small language model published by OpenBMB on September 7, 2026 under the Apache License 2.0.
Phi is a family of open-weight small language models (SLMs) developed by Microsoft Research, beginning with Phi-1 in June 2023 and spanning thirteen-plus releases through Phi-4-reasoning-vision-15B in March…
Phi-4 is a 14-billion-parameter small language model developed by Microsoft Research and released in December 2024, designed to match or beat models several times its size on reasoning tasks by training…
Phi-4-mini is a 3.8 billion parameter open weight small language model released by Microsoft on February 26, 2025, under the permissive MIT license.
Phi-4-mini-flash-reasoning is a 3.8 billion parameter open weight reasoning model released by Microsoft in July 2025.
SmolLM is a family of small, fully open language models released by Hugging Face on July 16, 2024 in three sizes, 135 million, 360 million, and 1.7 billion parameters, all trained on a curated open dataset…
SmolLM 2 is a family of compact open-weight language models released by Hugging Face on November 1, 2024.
SmolLM 3 is a fully open 3 billion parameter language model released by Hugging Face on July 8, 2025, trained on 11.2 trillion tokens and designed as a small, multilingual, long-context reasoner.