---
title: "Running a 35B Parameter AI Model on iPhone at 11 Tokens per Second"
slug: "35b-modell-laeuft-auf-dem-iphone-mit-11-tokens-pro-sekunde"
date: 2026-09-17
category: tech-pub
tags: []
language: en
sources_count: 1
featured: false
publisher: AInauten News
url: https://news.ainauten.com/en/story/35b-modell-laeuft-auf-dem-iphone-mit-11-tokens-pro-sekunde
---

# Running a 35B Parameter AI Model on iPhone at 11 Tokens per Second

**Published**: 2026-09-17 | **Category**: tech-pub | **Sources**: 1

---

## TL;DR

Better Stack demonstrates a 35-billion-parameter Mixture of Experts model running directly on an iPhone 17 at about 11 tokens per second.

---

## Summary

Better Stack demonstrates a 35-billion-parameter Mixture of Experts model running directly on an iPhone 17 at about 11 tokens per second. Only 3 billion parameters are active per step, powered by the Flash MoE engine adapted for iOS. Tiered quantization (4-bit for hot experts, 2-bit for cold ones) shrinks the model from 19 GB to 13 GB, with roughly 1.4 GB in RAM and the rest streamed from the SSD on demand. Downsides: the phone heats up noticeably, and setup requires Xcode, a paid Apple developer account and plenty of free storage.

---

## Why it matters

Better Stack demonstrates a 35-billion-parameter Mixture of Experts model running directly on an iPhone 17 at about 11 tokens per second.

---

## Key Points

- Better Stack demonstrates a 35-billion-parameter Mixture of Experts model running directly on an iPhone 17 at about 11 tokens per second.
- Only 3 billion parameters are active per step, powered by the Flash MoE engine adapted for iOS.
- Tiered quantization (4-bit for hot experts, 2-bit for cold ones) shrinks the model from 19 GB to 13 GB, with roughly 1.4 GB in RAM and the rest streamed from the SSD on demand.
- Downsides: the phone heats up noticeably, and setup requires Xcode, a paid Apple developer account and plenty of free storage.

---

## Nauti's Take

Running a 35B model locally on an iPhone is exciting progress for privacy and offline use. The catch is practicality: heat, 13 GB of storage and an Xcode setup are not everyday material. For developers it is a valuable preview of where on-device AI is heading. Everyone else should wait until Apple or app makers package this properly.

---


## FAQ

**Q:** What is Running a 35B Parameter AI Model on iPhone at 11 Tokens per Second about?

**A:** Better Stack demonstrates a 35-billion-parameter Mixture of Experts model running directly on an iPhone 17 at about 11 tokens per second.

**Q:** Why does it matter?

**A:** Better Stack demonstrates a 35-billion-parameter Mixture of Experts model running directly on an iPhone 17 at about 11 tokens per second.

**Q:** What are the key takeaways?

**A:** Better Stack demonstrates a 35-billion-parameter Mixture of Experts model running directly on an iPhone 17 at about 11 tokens per second.. Only 3 billion parameters are active per step, powered by the Flash MoE engine adapted for iOS.. Tiered quantization (4-bit for hot experts, 2-bit for cold ones) shrinks the model from 19 GB to 13 GB, with roughly 1.4 GB in RAM and the rest streamed from the SSD on demand.

---

## Related Topics

- —

---

## Sources

- [Running a 35B Parameter AI Model on iPhone at 11 Tokens per Second](https://www.geeky-gadgets.com/run-35b-ai-model-iphone/) - Geeky Gadgets AI

---

## About This Article

This article is a synthesis of 1 sources, curated and summarized by AInauten News. We aggregate AI news from trusted sources and provide bilingual (German/English) coverage.

**Publisher**: [AInauten](https://www.ainauten.com) | **Site**: [news.ainauten.com](https://news.ainauten.com)

---

*Last Updated: 2026-09-17*
