LemoLM: Local LLM & AI Chat

LemoLM: Local LLM & AI Chat

Offline AI Models on Device

0 ratings
Free
In-App Purchases

Rating summary

Details

  • Released
  • Updated
  • May 14, 2026
  • August 26, 2026

Features

LemoLM: Local LLM & AI Chat screenshot #1 for iPhone
LemoLM: Local LLM & AI Chat screenshot #2 for iPhone
LemoLM: Local LLM & AI Chat screenshot #3 for iPhone
LemoLM: Local LLM & AI Chat screenshot #4 for iPhone
LemoLM: Local LLM & AI Chat screenshot #5 for iPhone
LemoLM: Local LLM & AI Chat screenshot #6 for iPhone
LemoLM: Local LLM & AI Chat screenshot #7 for iPhone
LemoLM: Local LLM & AI Chat screenshot #8 for iPhone
iphone
ipad
🖼️Get Icon
Icons↘︎

About

Run Qwen3, DeepSeek R1, Mistral and Phi-4 directly on your iPhone or iPad. No account. No server. No internet once a model is downloaded. LemoLM is a local-first playground for testing and comparing on-device language models — with real performance numbers, not guesses. FREE, NO ACCOUNT NEEDED - Qwen3 1.7B — multilingual, with optional thinking mode - SmolLM2 1.7B — compact, for older and lower-memory devices - Apple Foundation Models — on supported devices - Mock Engine — explore the interface without downloading anything UNLOCK ALL LOCAL AI MODELS, NO ACCOUNT NEEDED One purchase, no subscription. Adds nine more GGUF models: - Mistral 7B Instruct v0.3 — general purpose, multilingual - DeepSeek R1 Distill Qwen 7B — reasoning, for high-memory devices - DeepSeek R1 Distill Qwen 1.5B — reasoning, compact - Qwen3 4B — larger multilingual model with thinking mode - NVIDIA Nemotron 3 Nano 4B — reasoning and coding - Phi-4 mini Instruct — instruction following and reasoning - Ministral 3 3B — chat, coding and instruction following - SmolLM3 3B — compact general purpose - TinySwallow 1.5B — Japanese-focused, from Sakana AI SEE WHAT'S ACTUALLY HAPPENING Most on-device AI apps hide the numbers. LemoLM shows them. - Performance Snapshot — generation timing, token estimates, tokens per second, memory readings and memory warning status. - Load Diagnostics — device memory, runtime mode and model file status when a model won't load. - Full model details — size, runtime, licence and source. CONTROL THE OUTPUT - Precise, Balanced and Creative presets - Custom Temperature, Top-p, Top-k and Seed - Settings persist between sessions, with one-tap reset COMPARE AND KEEP - Rate models with your own stars and notes - Export any conversation as a Markdown file - LLM Glossary for the concepts behind local AI PRIVATE BY DESIGN Conversations are stored on your device, and supported local models process prompts and responses on-device. Downloading a model requires internet access and may connect to third-party hosts such as Hugging Face; after that, the model runs locally. LemoLM collects no data. BEFORE YOU DOWNLOAD Performance depends on your device, available memory, iOS version, thermal state and model size. Compact models suit older or lower-memory devices. Larger GGUF models are best on recent Pro-class iPhones, M-series iPads and Apple silicon Macs. Apple Foundation Models requires a device that supports Apple Intelligence. LemoLM also runs on Macs with Apple silicon, where larger models have room to breathe. Compact language models can hallucinate, misread prompts, lose context or state things that are simply wrong. LemoLM is built for experimentation, learning and comparison. Verify anything that matters. Model names are trademarks of their respective owners. LemoLM is not affiliated with or endorsed by them.
Show more

What's New in LemoLM

1.6.0

August 26, 2026

New: NVIDIA Nemotron 3 Nano 4B joins the model roster. Qwen3 1.7B is now free for everyone — multilingual local AI, no account, no server.

1

In-App Purchases

$3.99

Full Model Playground Unlock

Unlock the full local model playground

Developer apps