LemoLM: Local LLM & AI Chat

LemoLM: Local LLM & AI Chat

Offline AI Models on Device

0 ratings
Free
In-App Purchases

Rating summary

Details

  • Released
  • Updated
  • May 14, 2026
  • September 21, 2026

Features

LemoLM: Local LLM & AI Chat screenshot #1 for iPhone
LemoLM: Local LLM & AI Chat screenshot #2 for iPhone
LemoLM: Local LLM & AI Chat screenshot #3 for iPhone
LemoLM: Local LLM & AI Chat screenshot #4 for iPhone
LemoLM: Local LLM & AI Chat screenshot #5 for iPhone
LemoLM: Local LLM & AI Chat screenshot #6 for iPhone
LemoLM: Local LLM & AI Chat screenshot #7 for iPhone
LemoLM: Local LLM & AI Chat screenshot #8 for iPhone
iphone
ipad
🖼️Get Icon
Icons↘︎

About

Explore Qwen3, DeepSeek R1, Mistral, Phi-4 and other local models directly on your iPhone or iPad. On a suitable Apple silicon Mac, explore the same model catalogue plus the desktop-class OpenAI gpt-oss-20b. No account. No server. Supported downloaded models can run without an internet connection. LemoLM is a local-first playground for running, testing and comparing on-device language models, with detailed performance information for each generation. FREE, NO ACCOUNT NEEDED • Qwen3 1.7B — multilingual, with optional thinking mode • SmolLM2 1.7B — compact, for older and lower-memory devices • Apple Foundation Models — on supported devices • Mock Engine — explore the interface without downloading anything UNLOCK THE FULL MODEL PLAYGROUND One purchase, no subscription. Adds access to twelve additional supported GGUF models: • OpenAI gpt-oss-20b — desktop-class local reasoning for suitable high-memory Apple silicon Macs • IBM Granite 4.2 3B — compact chat, coding and structured answers • LFM2-350M ENJP-MT — fast English–Japanese and Japanese–English translation • Mistral 7B Instruct v0.3 — general purpose and multilingual • DeepSeek R1 Distill Qwen 7B — reasoning for high-memory devices • DeepSeek R1 Distill Qwen 1.5B — compact reasoning • Qwen3 4B — larger multilingual model with thinking mode • NVIDIA Nemotron 3 Nano 4B — reasoning and coding • Phi-4 mini Instruct — instruction following and reasoning • Ministral 3 3B — chat, coding and instruction following • SmolLM3 3B — compact general-purpose model • TinySwallow 1.5B — Japanese-focused model from Sakana AI SEE WHAT’S ACTUALLY HAPPENING Most on-device AI apps hide the details. LemoLM shows them. • Performance Snapshot — measured loading and generation timing, token and speed estimates, device and operating-system details, memory readings, thermal state, battery and power context • Load Diagnostics — device memory, runtime mode and model file status when a model will not load • Full model details — size, runtime, quantisation, licence and source CONTROL THE OUTPUT • Precise, Balanced and Creative presets • Custom Temperature, Top-p, Top-k and Seed controls • Adjustable reasoning effort on supported models • Persistent settings with one-tap reset COMPARE AND KEEP • Rate models with your own stars and notes • Export a Performance Snapshot and conversation together, or export either separately, as Markdown • Learn about local AI concepts in the LLM Glossary PRIVATE BY DESIGN Conversations are stored on your device, and supported local models process prompts and responses on-device. Downloading a model requires internet access and may connect to third-party hosts such as Hugging Face. Once installed, supported downloadable models run locally. LemoLM collects no data. BEFORE YOU DOWNLOAD Performance depends on your device, available memory, operating system version, thermal state and model size. Compact models suit older or lower-memory devices. Larger GGUF models are best on recent Pro-class iPhones, M-series iPads and Apple silicon Macs. OpenAI gpt-oss-20b is intended for Apple silicon Macs with at least 24 GB of unified memory. Apple Foundation Models require a device that supports Apple Intelligence. Compact language models can hallucinate, misread prompts, lose context or state things that are incorrect. LemoLM is built for experimentation, learning and comparison. Always verify important information. Downloadable models are provided by their respective creators. Model availability, files, licences and download sources may change over time. Model names are trademarks of their respective owners. LemoLM is not affiliated with or endorsed by them.
Show more

What's New in LemoLM

1.7.0

September 21, 2026

Version 1.7.0 expands LemoLM with three new local models and richer performance diagnostics: • OpenAI GPT-OSS 20B brings desktop-class local reasoning to suitable high-memory Apple silicon Macs • IBM Granite 4.2 3B adds a compact model for chat, coding and structured answers • LFM2-350M ENJP-MT provides fast, private English–Japanese and Japanese-English translation • Performance Snapshot now includes device, memory, thermal, battery and power context alongside generation timing • Choose between complete, performance-only and conversation-only Markdown exports • New model badges and an automatic What’s New summary make additions easier to discover This release also includes interface refinements and stability improvements across iPhone, iPad and Apple silicon Mac.

More
1

Subscriptions & In-App Purchases

$3.99

Full Model Playground Unlock

Unlock the full local model playground

Developer apps