# Local Inference OnDevice Agent — Offline AI chat. No cloud.

> Offline AI chat powered by on-device open AI models with no data collection. (iPhone/iPad app by Sameer Shanbhag.)

- Source: https://appshunter.io/ios/app/local-inference-ondevice-agent/id6794007626 (this page in markdown: same URL + `.md`)
- Developer: [Sameer Shanbhag](https://appshunter.io/developer/6794007628)
- Category: Productivity, Utilities
- Price: Free
- Rating: 5.00/5 from 1 App Store ratings
- Age rating: 17+
- Requires: iOS 16.4 · 52 MB
- Languages: American English
- Released: 2026-08-05
- Data updated: 2026-08-25
- Monetization: free
- User reviews in markdown: https://appshunter.io/ios/app/local-inference-ondevice-agent/id6794007626/reviews.md

## What is Local Inference OnDevice Agent?

Your AI. Your phone. Nobody else's business.

LocalInference runs powerful open AI models entirely on your device. Your conversations never leave your phone — no account, no cloud, no analytics, no tracking. Ever.

PRIVATE BY ARCHITECTURE, NOT BY PROMISE
• Everything runs on-device; works fully offline after a one-time model download
• Zero data collection — we couldn't read your chats if we wanted to
• API keys (optional) stored in the secure keychain, sent only to the provider you chose

A REAL ASSISTANT, NOT JUST A CHATBOT
• Tools: calculator, web search, date & time, webpage reading
• Contacts & calendar integration — every write action asks you first
• Ask about photos with vision models
• Reads answers aloud with on-device voices

BUILT FOR TINKERERS
• 16 curated models from 0.5B to 8B — Qwen3, Llama, Gemma, DeepSeek R1, Phi, SmolVLM
• Only models that fit your device's memory are offered
• Live tok/s, time-to-first-token, and token counts per response
• Thinking models show their reasoning in collapsible blocks
• Bring your own keys: Anthropic, OpenAI, Google — or any OpenAI-compatible server (llama.cpp, LM Studio, Ollama) on your own network

Built with Llama. Gemma models provided under the Gemma Terms of Use.

## Key features

- Runs entirely on-device
- Fully offline functionality
- Zero data collection
- Optional API key integration
- Tools: calculator, web search, webpage reading
- Contacts & calendar integration
- Vision models for photo analysis
- On-device text-to-speech


## Frequently asked questions about Local Inference OnDevice Agent

### What is Local Inference OnDevice Agent?

Local Inference OnDevice Agent is a productivity app that runs AI models entirely on your device. It allows for private, offline AI chat without any data collection or cloud reliance.

### Does Local Inference OnDevice Agent collect my data?

No, Local Inference OnDevice Agent is designed with privacy as a core feature. All AI processing happens on your device, and zero data is collected, ensuring your conversations remain private.

### Can I use Local Inference OnDevice Agent offline?

Yes, the app works fully offline after an initial one-time model download. This means you can use its AI chat capabilities anywhere, without an internet connection.

### What AI models does Local Inference OnDevice Agent support?

It supports 16 curated models from 0.5B to 8B, including Qwen3, Llama, Gemma, DeepSeek R1, Phi, and SmolVLM. Only models that fit your device's memory are offered.

### Does Local Inference OnDevice Agent have ads?

No, Local Inference OnDevice Agent is ad-free. It is a free application with no advertisements to interrupt your experience.

### What devices is Local Inference OnDevice Agent available on?

Local Inference OnDevice Agent is available for iPhone and iPad. It is designed to run efficiently on these Apple mobile devices.

### How often is Local Inference OnDevice Agent updated?

The latest version is 1.1.0, and it was last updated on August 5, 2026. Updates typically bring new features, model support, and performance improvements.

### What is the age rating for Local Inference OnDevice Agent?

Local Inference OnDevice Agent has an age rating of 17+. This is due to the nature of AI models and their potential outputs, which may be more suitable for mature users.

## Version history (last 2 releases)

### 1.1.1 — 2026-08-17

Stability and fixes.

• Much lighter on memory. On phones with less RAM the app was loading the model at the largest possible context window, which left the system paging it in and out and could make the phone feel like it had frozen. It now sizes the model to what your device can actually hold — around 200 MB less memory in use, and far less swapping.
• The app now genuinely frees the model when you switch away, instead of holding it in the background.
• Photos work properly. You can send a picture without typing a caption, retrying a photo question keeps the photo, and attaching one no longer silently turns tools off.
• Tools tell you why they failed instead of just saying "failed", and tools that need the internet are hidden when you're offline rather than timing out.
• Chats can be renamed — press and hold a chat in the list.
• Answers you stop early can now be copied and read aloud.
• A reply interrupted by the app closing is no longer lost; you can retry it.
• The top bar now blends into the background instead of sitting on it as a flat slab.
• Web search results are no longer matched to the wrong page.
• Attached images are cleaned up properly instead of taking up space forever.

### 1.1.0 — 2026-08-05

No release notes.

---

*Data collected daily from the US App Store and indexed by [AppsHunter](https://appshunter.io/). User reviews are verbatim App Store reviews. Ratings, prices and chart positions refresh continuously; this snapshot is from 2026-08-25.*
