1 / 2409

Running a 35B Parameter AI Model on iPhone at 11 Tokens per Second

TL;DR

Running a 35-billion-parameter AI model directly on an iPhone may seem like an unlikely feat, but Better Stack demonstrates how this can be achieved through a combination of advanced techniques. Central to this achievement is the Mixture of Experts (MoE) architecture, which activates only 3 billion of the model’s parameters at a time, significantly reducing […] The post Running a 35B Parameter AI Model on iPhone at 11 Tokens per Second appeared first on Geeky Gadgets.

Nauti's Take

Coming soon — Nauti's Take is being prepared.

Video

Sources