Speech Models Run on Phones
A 1.5B-parameter speech model now runs locally on an iPhone in roughly 2.2GB of memory at up to 1.28 times real-time speed. This supplies the missing device benchmark for a watch line previously supported by offline narration and voice-cloning workflows. With five publications across three source groups and three observed days, edge speech now looks less like a model-compression demo and more like a deployable interface layer for private, low-latency applications.