# LemoLM: Local LLM & AI Chat — Offline AI Models on Device

> Offline AI Models on Device (iPhone/iPad app by PAUL ANTHONY HADFIELD.)

- Source: https://appshunter.io/ios/app/lemolm-local-llm-and-ai-chat/id6768808566 (this page in markdown: same URL + `.md`)
- Developer: [PAUL ANTHONY HADFIELD](https://appshunter.io/developer/1715547572)
- Category: Developer Tools, Productivity
- Price: Free with in-app purchases
- Age rating: 12+
- Requires: iOS 17.6 · 8 MB
- Languages: American English
- Released: 2026-05-14
- Data updated: 2026-09-06
- User reviews in markdown: https://appshunter.io/ios/app/lemolm-local-llm-and-ai-chat/id6768808566/reviews.md

## What is LemoLM?

Run Qwen3, DeepSeek R1, Mistral and Phi-4 directly on your iPhone or
iPad. No account. No server. No internet once a model is downloaded.

LemoLM is a local-first playground for testing and comparing on-device
language models — with real performance numbers, not guesses.

FREE, NO ACCOUNT NEEDED
- Qwen3 1.7B — multilingual, with optional thinking mode
- SmolLM2 1.7B — compact, for older and lower-memory devices
- Apple Foundation Models — on supported devices
- Mock Engine — explore the interface without downloading anything

UNLOCK ALL LOCAL AI MODELS, NO ACCOUNT NEEDED
One purchase, no subscription. Adds nine more GGUF models:
- Mistral 7B Instruct v0.3 — general purpose, multilingual
- DeepSeek R1 Distill Qwen 7B — reasoning, for high-memory devices
- DeepSeek R1 Distill Qwen 1.5B — reasoning, compact
- Qwen3 4B — larger multilingual model with thinking mode
- NVIDIA Nemotron 3 Nano 4B — reasoning and coding
- Phi-4 mini Instruct — instruction following and reasoning
- Ministral 3 3B — chat, coding and instruction following
- SmolLM3 3B — compact general purpose
- TinySwallow 1.5B — Japanese-focused, from Sakana AI

SEE WHAT'S ACTUALLY HAPPENING
Most on-device AI apps hide the numbers. LemoLM shows them.
- Performance Snapshot — generation timing, token estimates, tokens per
  second, memory readings and memory warning status.
- Load Diagnostics — device memory, runtime mode and model file status
  when a model won't load.
- Full model details — size, runtime, licence and source.

CONTROL THE OUTPUT
- Precise, Balanced and Creative presets
- Custom Temperature, Top-p, Top-k and Seed
- Settings persist between sessions, with one-tap reset

COMPARE AND KEEP
- Rate models with your own stars and notes
- Export any conversation as a Markdown file
- LLM Glossary for the concepts behind local AI

PRIVATE BY DESIGN
Conversations are stored on your device, and supported local models
process prompts and responses on-device. Downloading a model requires
internet access and may connect to third-party hosts such as Hugging
Face; after that, the model runs locally. LemoLM collects no data.

BEFORE YOU DOWNLOAD
Performance depends on your device, available memory, iOS version,
thermal state and model size. Compact models suit older or lower-memory
devices. Larger GGUF models are best on recent Pro-class iPhones, M-series iPads 
and Apple silicon Macs. Apple Foundation Models requires a device that supports
Apple Intelligence.

LemoLM also runs on Macs with Apple silicon, where larger models have
room to breathe.

Compact language models can hallucinate, misread prompts, lose context
or state things that are simply wrong. LemoLM is built for
experimentation, learning and comparison. Verify anything that matters.

Model names are trademarks of their respective owners. LemoLM is not
affiliated with or endorsed by them.

## Pricing and in-app purchases

Base price: Free.

**One-time purchases**

| Purchase | Price | Type |
| --- | --- | --- |
| Full Model Playground Unlock | $3.99 | one-time |



## Version history (last 5 releases)

### 1.6.0 — 2026-08-26

New: NVIDIA Nemotron 3 Nano 4B joins the model roster. Qwen3 1.7B is now
free for everyone — multilingual local AI, no account, no server.

### 1.5.0 — 2026-07-26

Version 1.5.0 adds more control, clearer guidance and a new local model:

• Ministral 3 3B joins the Full Model Playground as a capable multilingual model for chat, coding and instruction-following  
• Choose Precise, Balanced or Creative generation presets  
• Fine-tune Temperature, Top-p, Top-k and Seed using Custom settings  
• Learn about local AI concepts in the new LLM Glossary  
• New model labels make recent additions easier to find  
• Additional interface and stability improvements

### 1.4.1 — 2026-07-09

This update improves the model selection experience.

Model categories can now be collapsed and expanded, category headers show model counts, and the current active model is easier to spot. A small polish update to make browsing and managing local models feel cleaner.

### 1.4.0 — 2026-07-01

This update adds TinySwallow-1.5B Instruct a compact Japanese-focused local model from Sakana AI / Swallow Team.

It is useful for testing Japanese conversation, bilingual prompts, and comparing Japanese-language behaviour against other local models.

TinySwallow-1.5B Instruct can be downloaded from the Full Model Playground.

What's New section added to Settings to view the recent updates and model additions.

### 1.3.0 — 2026-06-22

New in LemoLM 1.3.0:

• Export conversations as Markdown files
• Generation and context settings now persist between sessions
• Added reset controls for generation and context settings
• Improved model diagnostics and performance snapshots
• General UI polish and bug fixes

## More apps by PAUL ANTHONY HADFIELD

- [Lemocu Dice](https://appshunter.io/ios/app/lemocu-dice/id6471528561)
- [Moji FlashCards](https://appshunter.io/ios/app/moji-flashcards/id6478466840)
- [Swift Units](https://appshunter.io/ios/app/swift-units/id6739262637)
- [FiDigital](https://appshunter.io/ios/app/fidigital/id6747617113)

All apps by PAUL ANTHONY HADFIELD: https://appshunter.io/developer/1715547572

---

*Data collected daily from the US App Store and indexed by [AppsHunter](https://appshunter.io/). User reviews are verbatim App Store reviews. Ratings, prices and chart positions refresh continuously; this snapshot is from 2026-09-06.*
