# onLM — Offline AI Assistant — Private AI & LLM Chat Offline

> Private, offline AI assistant for chat, transcription, and image generation on iOS. (iPhone/iPad app by Alexander Kryukov.)

- Source: https://appshunter.io/ios/app/onlm-offline-ai-assistant/id6760297856 (this page in markdown: same URL + `.md`)
- Developer: [Alexander Kryukov](https://appshunter.io/developer/1874309224)
- Category: Utilities, Productivity
- Price: Free
- Rating: 4.00/5 from 9 App Store ratings · 6 written reviews indexed
- Age rating: 17+
- Requires: iOS 26.0 · 43 MB
- Languages: English, French, German, Korean, Russian, Turkish
- Released: 2026-03-24
- Data updated: 2026-08-12
- Monetization: free
- User reviews in markdown: https://appshunter.io/ios/app/onlm-offline-ai-assistant/id6760297856/reviews.md

## What is onLM?

Chat with powerful language models, transcribe voice notes, and generate images — all running directly on your iPhone or iPad. No cloud. No servers. No accounts. Your data never leaves your device.

onLM is a native iOS app that brings state-of-the-art AI to a fully offline workspace. Every message, recording, and image is processed locally using your device's hardware, giving you a genuinely private AI assistant that works without an internet connection and without a subscription.

PRIVATE AI CHAT
Chat with open-source LLMs that run entirely on your device. Conversations are stored only on your iPhone or iPad — no servers, no sign-up, no telemetry. Pick a model that fits your task and switch between them anytime.

VOICE NOTES, TRANSCRIPTION & SUMMARIES
Record audio, transcribe it to text, and summarize long recordings into key points — all processed on-device. Useful for meeting notes, interviews, lectures, or spoken ideas. Recordings are auto-titled based on content for easy browsing.

ON-DEVICE IMAGE GENERATION
On supported devices, generate images from text prompts without sending anything to a remote server. Write your prompt in any supported language and onLM translates it locally before generation. Your images stay in a private gallery on your device.

CHOOSE FROM LEADING OPEN-SOURCE MODELS
- Gemma 4 E2B & E4B — Google's edge models with native audio and vision support
- Gemma 3 (4B, 12B) — strong multilingual capabilities from Google
- Qwen 3.5 (2B, 4B, 9B) — excellent all-around performance
- Llama 3 (3B, 8B) — reliable general-purpose models from Meta
- Phi 4 Mini — optimized for math, logic, and code from Microsoft
- Mistral 7B — versatile European-built model

All downloadable models are 4-bit quantized and run efficiently on mobile hardware via Apple's MLX framework.

APPLE INTELLIGENCE INTEGRATION
On supported devices, onLM can use Apple Intelligence as a built-in chat model with zero setup — no download, instant responses. Switch between Apple Intelligence and open-source models at any time.

BUILT FOR YOUR DEVICE
onLM detects your iPhone or iPad's capabilities and recommends the best models for your hardware. Quality ratings help you balance speed and intelligence, and features that require more memory are only surfaced on devices that can handle them — so you never hit a wall unexpectedly.

SEAMLESS EXPERIENCE
- Real-time streaming — watch responses appear word by word
- Background downloads — continue working while models load
- Smart memory management — stable performance on mobile hardware
- Conversation management — organize, search, and rename chats
- Stop and resume generation anytime
- Disk space checks before every download

NO SUBSCRIPTIONS. NO ADS. NO TRACKING.
Download a model once and use it as much as you want. There are no usage limits, no hidden costs, and no telemetry. onLM is a straightforward native app, not a wrapper around someone else's service.

Whether you need a private AI chatbot, an offline voice transcriber for meetings and ideas, or an on-device image generator — onLM gives you the power of modern AI without compromising your privacy.

## Key features

- Offline AI chat with open-source LLMs
- On-device voice note transcription and summarization
- Local image generation from text prompts
- Integration with Apple Intelligence
- Model selection based on device capabilities
- Real-time response streaming

## Recent user reviews (6 of 6)

All indexed reviews: https://appshunter.io/ios/app/onlm-offline-ai-assistant/id6760297856/reviews.md

### 5/5 — Отличное приложение для запуска локальных LLM

*2026-05-09*

По мне это лучшее из того, что попадалось для запуска локальных LLM. Первые версии крэшились на моем IPhone 17 pro max, но сейчас работает стабильно и транскрибация и текстовые запросы и даже генерация картинок. Интересно было бы добавить голосовой чат и приложение CarPlay, ну и какую нибудь локальную базу знаний ,чтобы решать специфичные задачи в условиях отсутствия интернета

### 4/5 — Отлично но…

*2026-04-13*

Добавьте Gemma 4 E2B для устройств с 6gb ram

**Developer response:** Скачайте обновление, вышла поддержка 14 моделей

### 5/5 — Очень удобно

*2026-04-12*

Пока это лучшее что я нашёл. Спасибо.

### 3/5 — Не получается работать на iPhone 17 pro max .

*2026-04-08*

Что-то пошло не так: MLX Error: [load_safetensors] Invalid json header length file /private/var/mobile/Containers/Data/Application/BF777B31-6424-4DB7-B113-B518FA8A508F/Library/Caches/models/mlx-community/Qwen3.5-4B-MLX-4bit/model.safetensors at /Users/alexanderkryukov/Library/Developer/Xcode/DerivedData/onLM-dqieinjttiyxsgfnbalepjyjawuu/SourcePackages/checkouts/mlx-swift/Source/Cmlx/mlx-c/mlx/c/io.cpp:64

**Developer response:** Спасибо за фидбек. Скачайте пожалуйста обновление и перекачайте модель

### 3/5 — Вылетает

*2026-04-06*

При запросе оффлайн к квэн 3 приложение крашится на вопросе как пожарить шашлыки

**Developer response:** Спасибо за отзыв! Проблема с вылетами уже исправлена — обновите приложение до последней версии. Если после обновления что-то всё ещё не работает, напишите — разберёмся. Приносим извинения за неудобства!

### 1/5 — Constantly crashes on iPhone 17 Pro and latest iOS!

*2026-04-05*

Buggy app. Constantly crashes on iPhone 17 Pro with basic text prompts on any LLM model shown as supported. Fix it please!

**Developer response:** Thanks for the feedback! This issue has been fixed — please update to the latest version. If you still experience any crashes after updating, let me know and I'll look into it. Apologies for the inconvenience!

## Frequently asked questions about onLM

### Does onLM have ads?

No, onLM is completely ad-free. The app explicitly states 'NO ADS' and is designed to provide an uninterrupted user experience without advertisements.

### What is the price of onLM?

onLM is available for free. There are no hidden costs, subscriptions, or in-app purchases mentioned, making it a completely free utility.

### What devices does onLM support?

onLM is designed for Apple devices and is compatible with both iPhone and iPad. It intelligently adapts to the capabilities of your specific device.

### How often is onLM updated?

The latest version of onLM, 1.4.2, was released on May 11, 2026, indicating recent development activity. The app is updated to ensure optimal performance and integration with new AI models.

### What are the privacy features of onLM?

onLM prioritizes user privacy by processing all data locally on your device. Conversations, transcriptions, and generated images never leave your iPhone or iPad, ensuring no cloud storage or server interaction.

### Can I use onLM without an internet connection?

Yes, onLM is built to function entirely offline. All AI models and features run directly on your device, meaning you do not need an internet connection to use it.

### What AI models are available in onLM?

onLM offers a selection of leading open-source models including Gemma (2B, 4B, 12B), Qwen 3.5 (2B, 4B, 9B), Llama 3 (3B, 8B), Phi 4 Mini, and Mistral 7B. It also integrates with Apple Intelligence on supported devices.

### What is the age rating for onLM?

onLM has an age rating of 17+. This is likely due to the advanced nature of AI technology and the potential for complex outputs from the language models.

## Version history (last 3 releases)

### 1.4.2 — 2026-05-11

- 12 GB iPhone stability — capped MLX buffer pool at 20 MB; added per-model KV-cache quantization and output-token caps (Qwen 3.5 9B, Gemma 3/4, Llama 3.1, Mistral) to stop jetsam kills by turn 3.
- Image generation — fixed OOM on 12 GB iPhones (proper hw.memsize gating, quantized SDXL path) and fixed a crash on long or non-Latin prompts (CLIP token sequences now clamped to 77).
- Onboarding — Skip button on the final page once any chat-capable model is ready; downloads keep running in the background.
- Download banner — now shown across all tabs, not just Chat.

### 1.1 — 2026-03-31

- On-device audio transcription powered by Qwen3-ASR
- Record voice notes and transcribe them to text entirely on your device
- No audio data ever leaves your iPhone or iPad

### 1.0 — 2026-03-24

No release notes.

---

*Data collected daily from the US App Store and indexed by [AppsHunter](https://appshunter.io/). User reviews are verbatim App Store reviews. Ratings, prices and chart positions refresh continuously; this snapshot is from 2026-08-12.*
