Small Language Models · Efficient AI · Open Models

Phi-3 vs SmolLM2 vs MobileLLM: What Small Models Actually Prove

Below 4B parameters there is no winner: phi-3-mini hits 69 MMLU on curated data, SmolLM2 beats Llama3.2-1B on general benchmarks but trails Qwen2.5-1.5B on math, and MobileLLM adds 2.7 to 4.3 percent at 125M to 350M.

Phi-3 vs SmolLM2 vs MobileLLM: What Small Models Actually Prove

The one-line answer

“Best small language model” is the wrong question, because the papers in this cluster are not competing on one leaderboard. They are answering three different questions. Phi-3 asks how far curated data can push a pocketable model. SmolLM2 asks what happens when you open the entire training pipeline of a 1.7B model. MobileLLM asks whether architecture still matters once you go below one billion parameters. TinyLlama is the earlier open-recipe baseline at 1.1B scale. None of the four papers runs the others’ models under the same harness, so any cross-model table is a table of claims, not a head-to-head. That is exactly why this comparison is worth writing down: the interesting content is in why each number was produced, not in ranking them.

Key numbers

MeasurementPhi-3-miniSmolLM2 (1.7B)MobileLLM (125M / 350M)TinyLlama (1.1B)
Parameters3.8B1.7B125M / 350M1.1B
Training tokens3.3T, filtered web plus synthetic~11T in four staged mixturesn/aOpen recipe; token budget not on this page
Headline claimMMLU 69, vs GPT-3.5 68 and Mixtral 8x7B 70.5Beats Llama3.2-1B everywhere: HellaSwag 68.7 vs 61.2, ARC 60.5 vs 49.2, PIQA 77.6 vs 74.8, MMLU-Pro 19.4 vs 11.7+2.7% / +4.3% over prior 125M / 350M state of the artFully open training recipe at 1.1B scale
Documented weak spotTriviaQA and broad factual recallGSM8K 31.1, MATH 11.6, HumanEval 22.6, far below Qwen2.5-1.5B’s 61.7 / 34.3 / 37.2Block-wise weight sharing adds latency without adding parametersNo comparable leaderboard numbers on this page
Instruction followingMT-bench 8.38 (instruct variant)IFEval 56.7 vs Qwen2.5-1.5B’s 47.4Reports stronger chat results; near LLaMA-v2 7B correctness on API-calling tasksn/a
On-device claim4-bit approx 1.8GB, over 12 tokens/sec offline on iPhone 14 (A16)360M and 135M variants (4T / 2T tokens) for edge deploymentDesigned for sub-billion on-device latencyn/a
What is openWeightsWeights plus FineWeb-Edu (~1.3T tokens), FineMath (up to 54B), Stack-Edu (~125B, 15 languages), SmolTalkn/aOpen training recipe

Every row comes from the paper’s own page on this site; none of the four papers runs the other three under a shared harness, so read the table as four claims side by side, not a ranking.

Recipe one: curate the data (Phi-3)

Phi-3’s architecture is unremarkable, a plain Llama-2-style transformer. The paper’s entire claim is that the scaling law everyone quotes is an artifact of feeding models undifferentiated internet text, and that a 3.8B model trained on 3.3 trillion carefully filtered tokens plus synthetic “textbook-quality” data lands in GPT-3.5 territory: 69 on MMLU against GPT-3.5’s 68 and Mixtral 8x7B’s 70.5, and 8.38 on MT-bench, at roughly a tenth of Mixtral’s active size. The same recipe scales inside the family: phi-3-small (7B) and phi-3-medium (14B), trained on 4.8T tokens, reach 75 and 78 on MMLU and 8.7 and 8.9 on MT-bench.

The cost of that recipe is that it is not a recipe you can copy. The filters, the synthetic-data prompts, and the mixture weights are undisclosed, which is understandable since they are the moat, but it means the claim is supported by benchmark outcomes rather than by anything reproducible. The paper is also candid about the flip side of small capacity: phi-3-mini is weak on trivia-heavy benchmarks like TriviaQA and thin outside English, and the authors themselves suggest pairing it with a search engine. “Matches GPT-3.5” holds for reasoning and chat quality on these benchmarks; it is not a claim about broad knowledge.

Recipe two: open the whole pipeline (SmolLM2)

SmolLM2 takes the opposite trade. It is a 1.7B model from Hugging Face overtrained on roughly 11 trillion tokens, far past the compute-optimal point, and its contribution is not the model but the paperwork: all four training datasets are public. FineWeb-Edu (~1.3T tokens of web text filtered for educational value by a classifier trained on Llama3-70B annotations), FineMath (up to 54B tokens, split into a 10B highest-quality tier and a 34B tier), Stack-Edu (~125B tokens of code across 15 languages), and SmolTalk for instruction tuning. You can rebuild the mixture, not just download the weights.

The training is deliberately staged: broad English web text from 0 to 6T tokens, then math and more code from 6 to 8T, higher-quality specialized data from 8 to 10T, and a learning-rate-decay phase from 10 to 11T that upsamples the best math and code while the model is most sensitive. The documented bet is that when a model sees data matters as much as what it sees.

On general reasoning the results are clean: SmolLM2 beats Meta’s Llama3.2-1B on every reported benchmark, HellaSwag 68.7 vs 61.2, ARC 60.5 vs 49.2, PIQA 77.6 vs 74.8, MMLU-Pro 19.4 vs 11.7, and the instruct variant leads Qwen2.5-1.5B on IFEval, 56.7 vs 47.4.

The number SmolLM2 would rather not lead with

Against Qwen2.5-1.5B, the same SmolLM2 page shows losses too large to wave away: GSM8K 31.1 vs 61.7, MATH 11.6 vs 34.3, HumanEval 22.6 vs 37.2. That is roughly half the math score and two-thirds of the code score, on the same benchmarks. SmolLM2 still leads the general-reasoning averages (HellaSwag 68.7 vs 66.4, MMLU-Pro 19.4 vs 13.7), so the honest summary is: better breadth and full reproducibility, worse math and code. If your workload at ~1.5B is math or code generation, the open-weights-only Qwen model remains the stronger pick despite SmolLM2 winning the transparency argument outright.

Recipe three: redesign the architecture (MobileLLM)

MobileLLM operates where data recipes stop being the bottleneck, below one billion parameters. At 125M and 350M, the paper argues that small models are not shrunken large models: width is expensive and often underused, while depth and weight sharing buy reasoning steps more efficiently. The design is deep-and-thin, more layers with narrower hidden dimensions, combined with embedding sharing and grouped-query attention. Reported gains are +2.7% and +4.3% over the preceding 125M and 350M state of the art, and the weight-sharing variant (MobileLLM-LS, immediate block-wise sharing) adds another +0.7% and +0.8% without adding parameters, at the cost of some latency.

The important framing is what this paper does not claim. It does not chase leaderboard maximalism, and it does not pretend a 350M model replaces a 7B one for broad knowledge. Its deployment-oriented results, stronger chat benchmarks and correctness close to LLaMA-v2 7B specifically on API-calling tasks, are exactly the niche where a sub-billion model is deployable at all: narrow, structured, high-frequency tasks on a phone.

TinyLlama: the open baseline

TinyLlama sits earlier in the same line as SmolLM2: a fully open training recipe at 1.1B scale. Its page on this site documents the scale and the openness rather than benchmark tables, so it is the reference point for “how small can an open recipe go” rather than a competitor on numbers. Read it as the baseline the later open recipes were measured against culturally, if not on a leaderboard.

When to use which

  • You need the strongest pocketable chat model and can accept a closed data pipeline: phi-3-mini. 69 MMLU and 8.38 MT-bench at 3.8B, quantized to ~1.8GB and running offline on an iPhone 14 at over 12 tokens per second, is the strongest single “data beats scale” datapoint in this cluster.
  • You need to audit, retrain, or re-mix the data: SmolLM2, unambiguously. It is the only entry whose entire training pipeline is public, and it beats Llama3.2-1B across the board.
  • Your workload at ~1.5B is math or code: none of the open-recipe papers above win it. Qwen2.5-1.5B’s 61.7 GSM8K vs SmolLM2’s 31.1 is the number to remember.
  • You are below one billion parameters or latency-bound on-device: MobileLLM is the only paper in this set that treats sub-billion architecture as the research question, and its +2.7% to +4.3% at 125M/350M is the evidence that design, not just data, still pays there.
  • You need broad factual knowledge in any of these tiers: none of them. Phi-3’s own paper flags TriviaQA weakness; small capacity is the shared ceiling.

Limits and open questions

The structural limit of this comparison is that no shared harness exists: phi-3 reports MMLU and MT-bench, SmolLM2 reports HellaSwag, ARC, PIQA, MMLU-Pro, GSM8K, MATH, HumanEval and IFEval, and MobileLLM reports relative gains over a prior state of the art it does not headline on this page. Cross-model numbers like “69 vs 68.7” would be meaningless even though both happen to be scores near 70. Benchmark contamination is a standing risk for every data-curated model in this cluster, and only SmolLM2 lets you inspect the data enough to reason about it. And all four papers predate the current wave of sub-3B reasoning models, so treat this as a map of what the training recipes proved, not a shopping list for 2026 deployments.

FAQ

What is the best small language model?

There is no single winner because the papers disagree on purpose. Phi-3-mini shows the strongest curated-data result (3.8B, 69 MMLU, 8.38 MT-bench). SmolLM2 is the most reproducible (1.7B, every dataset public). MobileLLM is the only one engineered for sub-billion on-device tiers. Pick by deployment tier and audit requirements, not by a cross-paper average that no one measured.

Is Phi-3-mini really as good as GPT-3.5?

On the benchmarks its paper reports, close: 69 on MMLU versus GPT-3.5’s 68, and 8.38 on MT-bench. The same paper is candid that a 3.8B model cannot hold broad factual knowledge, so it underperforms on trivia-style tasks; “as good” holds for reasoning and chat quality, not recall.

Does SmolLM2 beat Llama3.2-1B?

On every general-reasoning benchmark reported on this site, yes: HellaSwag 68.7 vs 61.2, ARC 60.5 vs 49.2, PIQA 77.6 vs 74.8, MMLU-Pro 19.4 vs 11.7. It does not beat Qwen2.5-1.5B on math or code (GSM8K 31.1 vs 61.7).

Why does MobileLLM go deep-and-thin instead of wide?

Below one billion parameters, width is expensive and often underused, while depth and block-wise weight sharing buy representational steps more efficiently per parameter. The paper reports +2.7% and +4.3% at 125M and 350M over the prior state of the art, plus another +0.7% and +0.8% from immediate block-wise sharing.

What do these models still fail at?

Broad factual knowledge is the shared ceiling. Phi-3’s own paper flags weak TriviaQA scores and suggests pairing the model with search. Math and code at the 1.5B tier belong to other recipes, per SmolLM2’s losses to Qwen2.5-1.5B. And every score above is a benchmark claim under each paper’s own harness, not a head-to-head result.