
Whisperkit
On-Device AI Transcription
About
What's New in Whisperkit
1.1
August 28, 2026
Version 1.1 is a redesign of the whole app, plus a substantial rebuild of what happens under it. A NEW LOOK, BUILT FOR THE WORK Whisper has been redesigned around the recording itself. Every session now shows its real waveform — drawn from the actual audio, so silences are gaps and speech has shape — and you can tap anywhere on it to jump straight there. Transcripts are laid out as timestamped lines rather than a wall of text; tap a line and the audio follows. On iPad, past sessions live in a rail down the side, so moving between recordings takes one tap. On iPhone, recent work sits on the home screen. Dark mode is supported properly throughout. EXPORT YOU CAN SEE BEFORE YOU SEND Choose TXT, SRT, VTT, Markdown or JSON and the app shows you the actual file before you share it, with its size. Speaker labels can be included or left out. Subtitle timings have also been corrected — they were rounding a millisecond early. BETTER MODELS, SMALLER DOWNLOADS The model line-up is rebuilt around Neural Engine–optimised builds, with sizes that are measured rather than estimated: • Small — 217 MB. Genuinely multilingual. • Turbo — 646 MB. Recommended for most people, and less than half the size of the model it replaces. • Large V3 — 948 MB, and Large V3 Turbo — 1.1 GB. Previously these pulled over 3 GB each. Downloads show real progress, can be cancelled, and resume where they left off. Long-press a model to delete it. CHINESE, JAPANESE AND KOREAN Transcribing these languages could previously return nothing at all. An internal repetition check misread dense CJK text as a decoding failure and discarded good results. That is fixed. Quiet and far-field recordings are also far less likely to be thrown away. READY WHEN YOU ARE Whisper now runs on the GPU by default and is ready in seconds, rather than spending minutes preparing a model on first use. You can start recording while a model is still loading, and importing a long video no longer stalls. Live text now accumulates as the audio is read, instead of resetting every thirty seconds. Everything still runs entirely on your device. No account, no analytics, and nothing you record ever leaves your iPhone or iPad.
More



