
Personal LLM
Private Offline AI Chat
0 ratings
Free
With Ads
Rating summary
About
Chat with an AI that has nowhere to send your data. Personal LLM is a private, offline AI chat app: open AI models run entirely on your phone — no server, no account, and after a one-time model download no internet at all.
YOUR AI, YOUR DEVICE
Cloud chatbots send every word you type to a data centre. Personal LLM computes the answer on your phone's own chip: your chats, your photos and the models themselves stay with you. Uninstall it and everything is gone — nothing was ever anywhere else.
WORKS COMPLETELY OFFLINE
Download a model once, then chat anywhere:
• On flights, subways and trips abroad
• In dead zones and on unreliable connections
• Whenever you want answers that stay private
Switch on airplane mode if you like — the app won't notice.
THE LATEST OPEN MODELS, PHONE-SIZED
• Qwen 3.5 (0.8B, 4B, 9B) — fast, smart and multilingual
• Google Gemma 4 (E2B, E4B) — mobile-first quality, 140+ languages
• GLM 4.6V Flash — flagship-class vision for high-end phones
• Ministral 3 — compact, with a 256k context window
• Bring your own: load any GGUF model by URL (Hugging Face links work)
Models are 0.8–6 GB downloads. A "Fits your device" badge reads your phone's RAM so you know what will run well before you download, the catalog is searchable and filterable, and interrupted downloads resume where they stopped.
BENCHMARK IT ON YOUR OWN PHONE
Run the built-in benchmark on any downloaded model to see exactly how fast it loads and generates on your hardware — tokens per second, GPU or CPU — and pick the model that fits.
CHAT WITH YOUR DOCUMENTS
Attach a PDF, text or Markdown file to any chat and ask about it: the text is extracted and searched on your phone, the model answers from the matching passages and tells you which ones it used. Contracts, papers, manuals, notes — none of it leaves the device.
UNDERSTANDS IMAGES
Every model in the catalog supports vision. Add the optional image-support file, then snap a photo or pick one from your gallery: read documents, translate signs, decode menus, describe scenes.
WATCH IT THINK
Turn on thinking mode and reasoning models show their step-by-step working in a collapsible panel above the answer — great for math, code and anything worth double-checking.
A REAL CHAT APP
• Markdown that renders: code blocks with one-tap copy, tables and lists
• Streaming answers with a live tokens-per-second readout
• Switch models mid-conversation from the chat header
• A system prompt per chat, plus Creative / Balanced / Precise / Thinking presets
• Edit, regenerate, copy, select text — or have any answer read aloud (on-device TTS)
• Search your chats, pin favourites, light and dark themes
• Back up every chat and setting to a single JSON file, and restore anywhere
PRIVATE BY DESIGN
• Conversations and photos never leave your device
• No account, no sign-up, no server
• GPU-accelerated by llama.cpp — Metal on iPhone, OpenCL on Snapdragon Adreno
• Free to use, supported by ads — ads never interrupt an answer, and offline none can load
Your questions are yours. Keep them that way.
Show more
What's New in Personal LLM
1.4.1
September 7, 2026
• Fresh look across the app: smoother animations, new chat, models and settings screens • Benchmark any downloaded model to see how fast it runs on your phone • Search and filter the model catalog • Copy code blocks with one tap and select text inside any reply • Chats grouped by day, with quick-start suggestion cards • Chat with your documents: attach a PDF or text file and ask about it — all on device • Default system prompt (Settings → Behavior) applied to every new chat • Performance and stability improvements
More





