Can SmolLM2 1.7B run on iPhone Duo?
Some specifications for iPhone Duo are not confirmed. Compatibility and speed estimates use provisional chip or RAM information. Announced; available October 23, 2026. RAM capacity is not confirmed; results assume a provisional 12GB configuration.
YES — Runs great
Formula estimate
What this means
Our estimate says SmolLM2 1.7B should fit comfortably on your iPhone Duo.
Download the Q4_K_M version, which is about 1.1 GB. We expect it to use about 1.9 GB of the roughly 7.8 GB available to a model on this phone.
At the estimated speed, a roughly 300-word answer may take about 8 seconds to finish.
This result is calculated from the phone and model specifications. It has not been measured on this exact phone and setup.
47.1tokens/s
estimated · instant
1.9GB
needed at Q4_K_M
Q4_K_M
quant checked
needs 1.9 GBusable 7.8 GB
Apple A20 Pro12 GB RAMMetal
See how fast it feels
Estimated at 47.1 tokens/s — instant. A ~300-word reply takes about 8 seconds on this phone.
Where the memory goes
| Component | Detail | GB |
|---|---|---|
| Model weights | Q4_K_M GGUF (1.1 GB) + mmap overhead | 1.2 |
| KV cache | 4K context window | 0.1 |
| Runtime | llama.cpp + app overhead | 0.6 |
| Total needed | at Q4_K_M, 4K context | 1.9 |
| Working budget | 12 GB RAM − iOS reserve (jetsam limit) | 7.8 |
| Headroom | remaining inside the working budget | 5.9 |
Pick your quant
| Quant | Download | Verdict | Speed |
|---|---|---|---|
| Q2_K | 0.7 GB | ✓ Runs great | ~74.1 tokens/s |
| IQ4_XS | 0.9 GB | ✓ Runs great | ~57.6 tokens/s |
| Q3_K_M | 0.9 GB | ✓ Runs great | ~57.6 tokens/s |
| Q4_0 | 1 GB | ✓ Runs great | ~51.8 tokens/s |
| Q4_K_S | 1 GB | ✓ Runs great | ~51.8 tokens/s |
| Q4_K_M BEST HERE | 1.1 GB | ✓ Runs great | ~47.1 tokens/s |
| Q5_K_M | 1.2 GB | ✓ Runs great | ~43.2 tokens/s |
| Q6_K | 1.4 GB | ✓ Runs great | ~37 tokens/s |
| Q8_0 | 1.8 GB | ✓ Runs great | ~28.8 tokens/s |
Get your first offline chat working
Recommended app: PocketPal. Follow the point-and-click steps below. The speed above is an estimate, not a measurement from this exact app and phone.
1
Install or update PocketPal from the App Store. It is free and does not require an account. Use a current version so its loader supports newer model architectures.
2
Open the exact model. In PocketPal, go to Models → + → Add from Hugging Face, then paste
bartowski/SmolLM2-1.7B-Instruct-GGUF.3
Choose the Q4_K_M GGUF file. The download is about 1.1 GB, so use Wi-Fi and keep the app open. Choose the main GGUF weights, not a vision projector, mmproj, or other helper file.
4
Tap Download, then Load. Start with a 4K (4096-token) context. Keep PocketPal’s default Metal acceleration; this does not require Xcode.
5
Send a simple first prompt. Try “Explain why the sky is blue in three sentences.” This page estimates about 47.1 tokens/s, but that number is not a PocketPal measurement unless it carries a ✓ Verified label.
6
Confirm it is really offline. After the first reply, turn on airplane mode and ask a second question. If it still answers, the model is running on your phone.
If it does not work
- Model not listed: update PocketPal and paste the exact repository
bartowski/SmolLM2-1.7B-Instruct-GGUF. - App closes while loading: close other apps, restart the phone, and try 2K context. If it still closes, choose a smaller model.
- No offline reply: confirm that the Q4_K_M GGUF file is loaded in the chat rather than a remote model.
Related checks
More on iPhone Duo
SmolLM2 1.7B on other phones