# Endemic: Local On-Device AI — Chat with Private Offline LLM

> Local on-device AI chat with private, offline large language models. (iPhone/iPad app by Folding Sky Co.)

- Source: https://appshunter.io/ios/app/endemic-local-on-device-ai/id6760927382 (this page in markdown: same URL + `.md`)
- Developer: [Folding Sky Co](https://appshunter.io/developer/1842155029)
- Category: Utilities, Developer Tools
- Price: Free with in-app purchases
- Rating: 5.00/5 from 4 App Store ratings · 1 written reviews indexed
- Age rating: 17+
- Requires: iOS 18.0 · 67 MB
- Languages: American English, Chinese (Simplified, China)
- Released: 2026-03-30
- Data updated: 2026-08-16
- Monetization: freemium
- User reviews in markdown: https://appshunter.io/ios/app/endemic-local-on-device-ai/id6760927382/reviews.md

## What is Endemic?

Endemic runs large language models directly on your iPhone, iPad, and Mac. No API keys or cloud inference. Models and conversations stay on your device unless you choose to sync conversations through iCloud.

Choose a model from the curated catalog, download it, and chat without an internet connection. Endemic filters the catalog by available memory, so you only see models suited to your hardware.

Endemic includes Qwen 3.5 and Google's Gemma 4, from 0.8B models for recent iPhones to 9B models for Macs and iPad Pros with at least 16GB of memory. Switch models whenever you like.

QWEN 3.5 AND GEMMA 4, ON DEVICE

Qwen 3.5 models are included free, from the compact 0.8B that runs on any supported device to the more capable 9B for devices with at least 16GB of memory. Google's Gemma 4 supports full chat templating and function calling in Endemic. The 4B Gemma works well on a modern phone, while the 9B is better suited to iPad Pro and Mac. Both families support Local Web Search and Local Web Browser for current information.

Gemma 4 is available with Endemic Pro, a small monthly subscription that also includes conversation folders and supports ongoing development by one person. Qwen 3.5 is included free. You can try Pro and cancel any time.

WEB SEARCH AND BROWSING, ON DEVICE

Endemic can search the web and read pages from a conversation. Ask for current information or a specific site, and the model can call local tools to fetch results and summarize them. Requests go directly from your device over your network connection. There is no proxy or Folding Sky server in the middle.

Enable Local Tools in Settings to add the searchWeb and openWebPageLocally tools. Calls and results appear inline, so you can see what was requested and returned.

IMAGES IN, ANSWERS OUT

Vision-capable local models can read images attached to a message. Ask one to describe a photo, transcribe a screenshot, or explain a diagram. The image stays on your device, and the analysis works offline.

WHAT LOCAL INFERENCE ACTUALLY MEANS

During a conversation, the model runs on your device's CPU and GPU. Prompts and responses are not sent to an inference server. The work happens on your hardware.

A 4B model on a phone is useful for many tasks, but it will not match a frontier model running on datacenter hardware. For the strongest responses, use cloud AI. For private, offline work, use Endemic.

MODELS AND DEVICES

Endemic detects your hardware and recommends the strongest model that fits. An iPhone with 8GB of memory runs the 4B comfortably. iPads and Macs with at least 16GB handle the 9B. The 0.8B runs quickly on any supported device.

Models are open-weight GGUF files downloaded from public hosting and stored locally. Endemic excludes them from iCloud backup, so they do not use your iCloud storage quota.

PRIVACY

Endemic has no Folding Sky backend. Optional conversation sync uses your personal iCloud account.

Built by one person in Reno, Nevada.

Privacy Policy: https://folding-sky.com/privacy

Terms of Service: https://folding-sky.com/terms

## Key features

- On-device LLM execution
- Private, offline conversations
- Curated model catalog
- Local web search & browsing
- Image analysis (vision models)
- Device RAM filtering for models

## Pricing and in-app purchases

Base price: Free.

**Subscriptions**

| Subscription | Price | Period |
| --- | --- | --- |
| Endemic Pro | $6.99 | per month |


## Recent user reviews (1 of 1)

All indexed reviews: https://appshunter.io/ios/app/endemic-local-on-device-ai/id6760927382/reviews.md

### 5/5 — So far so good

*2026-04-30*

It's really nice to have a local LLM app from the same developer as Cumbersome. I was able to successfully run the Qwen 9B model on my 17 Pro Max. 

I do want to see support for more models as that's always the main quirk of a lot of these iOS LLM apps. It would be great to have Gemma 4 E2B/E4B. 

I would love to have parameter adjustments, such as temperature, top_P, etc. I'm not a huge fan of the way parameters are adjusted in Cumbersome so I would ideally prefer to see some sort of interface based adjustment system.

Overall, for a brand new app, it's off to a good start. It works, it does what it's supposed to, and it's all on device. Keep up the excellent work!

**Developer response:** Thanks very much for the feedback. Expect more models and features soon! I hear you on the simpler settings interface: always a tricky balance, especially when we have to support multiple models and their different APIs.

## Frequently asked questions about Endemic

### What is Endemic?

Endemic is an application that runs large language models directly on your iPhone, iPad, or Mac. It allows for private, offline AI conversations without needing API keys or cloud services.

### Does Endemic have ads?

No, Endemic does not contain advertisements. The app is free to download and use with the included Qwen 3.5 model.

### What models can I use with Endemic?

Endemic ships with Qwen 3.5 and Google's Gemma 4 models. Models are available in sizes from 0.8B to 9B, filtered by your device's RAM to ensure compatibility.

### Can Endemic access the internet?

Yes, Endemic can perform local web searches and browse web pages. These actions are performed on your device through your network connection, with results summarized inline.

### What are the pricing details for Endemic?

Endemic is free to download and use with the Qwen 3.5 model. An optional subscription, Endemic Pro, unlocks additional features like conversation folders and supports development.

### Can Endemic process images?

Yes, vision-capable local models within Endemic can read images you attach. The image stays on your device, and the model can describe, transcribe, or reason about its content offline.

### How does Endemic compare to cloud AI?

Endemic offers a private, on-device experience where computation happens locally. While a 4B model on a phone is useful, it won't match the performance of frontier models on datacenter hardware.

### What devices are supported by Endemic?

Endemic is supported on iPhone, iPad, and Mac. The app detects your hardware to recommend the strongest model that fits, with larger models like 9B requiring 16GB+ RAM on Macs and iPads.

## Version history (last 5 releases)

### 1.4.0 — 2026-08-11

Bug fixes and improvements.

### 1.3.0 — 2026-05-27

Bug fixes and improvements.

### 1.2.0 — 2026-05-03

Bug fixes and feature improvements.

### 1.1.0 — 2026-04-30

feat(core): embedded tool-call UI and stronger local web-tool prompts
feat(local-tools): Local Tools section, headless mode, Endemic defaults
feat(local-tools): Endemic local LLM tools + searchWeb + shared plumbing
feat(chat): attachment preview share flow and exhaustive commit rules
feat(chat): remove attachments when editing messages
fix(endemic-llm): persist full tool-loop text; strip all think markup in UI
fix(core): normalize assistant content before save to strip empty think tags
fix(endemic-llm): restore streaming on first tool-loop hop
fix(endemic-llm): RAM-based context window cap, sync model loading, clear error messages
fix(endemic-llm): Swift 6 explicit capture for ModelDownloadManager logging
fix(endemic-settings): device-local LLM model selection, readiness fallback, observation chain
fix(core): local LLM tool loop Hermes + headless reliability
perf(endemic-llm): re-enable Metal flash attention and right-size KV cache per turn
docs: document Endemic resource limits and refresh README
docs(endemic-plan): sync local tool plan; fix build metadata after archive
docs(ai-plans): add Endemic local tool use plan

### 1.0.0 — 2026-03-30

No release notes.

## More apps by Folding Sky Co

- [Cumbersome: AI LLM API Client](https://appshunter.io/ios/app/cumbersome-ai-llm-api-client/id6753016821)
- [Dense: AI News Digest Widget](https://appshunter.io/ios/app/dense-ai-news-digest-widget/id6755408172)
- [Alarmist LiDAR Motion Detector](https://appshunter.io/ios/app/alarmist-lidar-motion-detector/id6756851395)
- [Busybody Web Screenshot Widget](https://appshunter.io/ios/app/busybody-web-screenshot-widget/id6757285099)

All apps by Folding Sky Co: https://appshunter.io/developer/1842155029

---

*Data collected daily from the US App Store and indexed by [AppsHunter](https://appshunter.io/). User reviews are verbatim App Store reviews. Ratings, prices and chart positions refresh continuously; this snapshot is from 2026-08-16.*
