
LemoLM: Local LLM & AI Chat
Offline AI Models on Device
0 ratings
Free
In-App Purchases
Rating summary
About
Run Qwen3, DeepSeek R1, Mistral and Phi-4 directly on your iPhone or
iPad. No account. No server. No internet once a model is downloaded.
LemoLM is a local-first playground for testing and comparing on-device
language models — with real performance numbers, not guesses.
FREE, NO ACCOUNT NEEDED
- Qwen3 1.7B — multilingual, with optional thinking mode
- SmolLM2 1.7B — compact, for older and lower-memory devices
- Apple Foundation Models — on supported devices
- Mock Engine — explore the interface without downloading anything
UNLOCK ALL LOCAL AI MODELS, NO ACCOUNT NEEDED
One purchase, no subscription. Adds nine more GGUF models:
- Mistral 7B Instruct v0.3 — general purpose, multilingual
- DeepSeek R1 Distill Qwen 7B — reasoning, for high-memory devices
- DeepSeek R1 Distill Qwen 1.5B — reasoning, compact
- Qwen3 4B — larger multilingual model with thinking mode
- NVIDIA Nemotron 3 Nano 4B — reasoning and coding
- Phi-4 mini Instruct — instruction following and reasoning
- Ministral 3 3B — chat, coding and instruction following
- SmolLM3 3B — compact general purpose
- TinySwallow 1.5B — Japanese-focused, from Sakana AI
SEE WHAT'S ACTUALLY HAPPENING
Most on-device AI apps hide the numbers. LemoLM shows them.
- Performance Snapshot — generation timing, token estimates, tokens per
second, memory readings and memory warning status.
- Load Diagnostics — device memory, runtime mode and model file status
when a model won't load.
- Full model details — size, runtime, licence and source.
CONTROL THE OUTPUT
- Precise, Balanced and Creative presets
- Custom Temperature, Top-p, Top-k and Seed
- Settings persist between sessions, with one-tap reset
COMPARE AND KEEP
- Rate models with your own stars and notes
- Export any conversation as a Markdown file
- LLM Glossary for the concepts behind local AI
PRIVATE BY DESIGN
Conversations are stored on your device, and supported local models
process prompts and responses on-device. Downloading a model requires
internet access and may connect to third-party hosts such as Hugging
Face; after that, the model runs locally. LemoLM collects no data.
BEFORE YOU DOWNLOAD
Performance depends on your device, available memory, iOS version,
thermal state and model size. Compact models suit older or lower-memory
devices. Larger GGUF models are best on recent Pro-class iPhones, M-series iPads
and Apple silicon Macs. Apple Foundation Models requires a device that supports
Apple Intelligence.
LemoLM also runs on Macs with Apple silicon, where larger models have
room to breathe.
Compact language models can hallucinate, misread prompts, lose context
or state things that are simply wrong. LemoLM is built for
experimentation, learning and comparison. Verify anything that matters.
Model names are trademarks of their respective owners. LemoLM is not
affiliated with or endorsed by them.
Show more
What's New in LemoLM
1.6.0
August 26, 2026
New: NVIDIA Nemotron 3 Nano 4B joins the model roster. Qwen3 1.7B is now free for everyone — multilingual local AI, no account, no server.
1
In-App Purchases
$3.99
Full Model Playground Unlock
Unlock the full local model playground







