# LLM Server — LLM Server & API on your LAN

> On-device AI inference server for large language models with an OpenAI-compatible API. (iPhone/iPad app by Linosec.)

- Source: https://appshunter.io/ios/app/llm-server/id6762212937 (this page in markdown: same URL + `.md`)
- Developer: [Linosec](https://appshunter.io/developer/1771206390)
- Category: Business, Developer Tools
- Price: $17.99
- Age rating: 17+
- Requires: iOS 16.4 · 543 MB
- Languages: American English
- Released: 2026-05-01
- Data updated: 2026-08-24
- Monetization: paid
- User reviews in markdown: https://appshunter.io/ios/app/llm-server/id6762212937/reviews.md

## What is LLM Server?

Your Phone Is Now an AI Server
LLM Server turns your iPhone or iPad into a private AI inference server. Run large language models entirely on-device, expose an OpenAI-compatible API to your local network, and chat with your model from any browser on any device — laptop, desktop, tablet, even another phone. No cloud, no subscription, no data leaving your hardware.

Chat From Any Browser on Your Network
Start the server and open the URL on any device sharing your Wi-Fi. The built-in web interface gives you a clean chat UI instantly — no client app to install, no account to create. Your laptop, your partner's tablet, a colleague's machine: all talking to a model running on your phone.

OpenAI-Compatible API, Drop-In Ready
A standard OpenAI-compatible API means LLM Server slots into any tool you already use — Continue, Open WebUI, LangChain, custom scripts, the lot. Chat completions, text completions, streaming (SSE), and model listing endpoints all supported. Ollama CLI commands work too.

Fully Offline, Fully Yours
- Download a GGUF model once and you're done with the cloud. Inference runs locally on Apple Metal GPU with no internet required. Your prompts, your conversations, your data — none of it leaves the device. Airplane mode works fine.

Any GGUF Model From Hugging Face
- Browse and download directly from Hugging Face with built-in search, or import your own files. LLaMA, Mistral, Phi, Gemma, Qwen, DeepSeek, and every other llama.cpp-supported architecture runs out of the box. Background downloads with progress tracking so you can keep working.
Enterprise-Grade Security

- TLS/HTTPS encryption — generate self-signed certificates or import your own chain and private key.
- API key authentication — Bearer tokens with per-key management. Generate cryptographically secure keys or bring your own.
- Bind control — lock to localhost, open to your LAN, or pin to a specific interface.

Tune Every Knob
Full control over generation: context size up to 32K, temperature, top-p, top-k, repeat penalty, frequency and presence penalties, max tokens, seed, GPU layer offloading, and thread count. Save presets globally or per model.
Smart Resource Management
Your phone stays responsive under load. Real-time thermal monitoring with automatic thread reduction under pressure and request rejection at critical temperatures. Memory-aware model loading with conservative budgeting. Configurable request queues and per-request timeouts.
Built for Developers
- Live API docs with copy-paste curl examples
- Structured logging (debug, info, warning, error)
- One-tap copy for server addresses and API keys

What's Inside

Dashboard with one-tap server control, model manager with download progress, complete settings hub (server, inference, security, API keys, developer tools), and guided onboarding for first-time setup.

## Key features

- On-device AI inference server
- OpenAI-compatible API
- Chat from any browser
- Fully offline operation
- Download GGUF models from Hugging Face
- Enterprise-grade security
- Full generation control

## Chart rankings

- #128 in Top Paid · Business (US App Store)

## Recent user reviews (2 of 2)

All indexed reviews: https://appshunter.io/ios/app/llm-server/id6762212937/reviews.md

### 2/5 — Missing image

*2026-06-22*

Missing importing image on vision models, bought the app for that and it’s not working

**Developer response:** The vision model feature is still being finalized and will land in the next update, coming very soon. If you bought the app specifically for this, email me at contact@linosec.com and I'll make sure you're looked after. Thanks for your patience.

### 5/5 — Super simple et pratique

*2026-05-08*

Vraiment convivial

## Frequently asked questions about LLM Server

### What is LLM Server?

LLM Server turns your iPhone or iPad into a private AI inference server. It allows you to run large language models entirely on-device and exposes an OpenAI-compatible API to your local network.

### How much does LLM Server cost?

LLM Server is a paid application, priced at $6.99. There are no subscriptions or in-app purchases mentioned.

### Does LLM Server have ads?

No, LLM Server does not contain advertisements. It is a paid application designed for a seamless, ad-free user experience.

### What devices support LLM Server?

LLM Server is supported on iPhones and iPads. The app is rated for users aged 17 and older.

### Can I use LLM Server offline?

Yes, LLM Server is designed for fully offline operation. Once a GGUF model is downloaded, inference runs locally on your device's Apple Metal GPU without requiring an internet connection.

### What AI models can I run with LLM Server?

LLM Server supports any GGUF model from Hugging Face, including LLaMA, Mistral, Phi, Gemma, Qwen, and DeepSeek, as well as other llama.cpp-supported architectures.

### How often is LLM Server updated?

The latest version of LLM Server is 1.4, last updated on June 29, 2026. The release date was May 1, 2026.

### What security features does LLM Server offer?

LLM Server provides enterprise-grade security with TLS/HTTPS encryption, API key authentication, and bind control to manage network access.

## Version history (last 5 releases)

### 2.0.7 — 2026-08-09

- Performance & Reliability improvements

### 2.0.5 — 2026-07-23

- stability and performance improvements

### 2.0.4 — 2026-07-08

- Reliability improvements
- Added Apple Intelligence model
- Added Anthropic API support

### 2.0.3 — 2026-07-05

- fixed gemma4 loading

### 2.0.2 — 2026-07-04

- Performance and stability improvements

## More apps by Linosec

- [iFocus Stacking](https://appshunter.io/ios/app/ifocus-stacking/id6752693470)
- [Astro GPS (Offline)](https://appshunter.io/ios/app/astro-gps-offline/id6753069568)
- [iCamera Gear Manager](https://appshunter.io/ios/app/icamera-gear-manager/id6753150898)
- [iTranscription](https://appshunter.io/ios/app/itranscription/id6753682751)
- [Wheel of Emotions](https://appshunter.io/ios/app/wheel-of-emotions/id6757318560)
- [ScreenshotResize](https://appshunter.io/ios/app/screenshotresize/id6758956062)

All apps by Linosec: https://appshunter.io/developer/1771206390

## Related topics

[local llm server pro](https://appshunter.io/ios/topics/local-llm-server-pro) · [llm server](https://appshunter.io/ios/topics/llm-server) · [alpha iota ai open llm client](https://appshunter.io/ios/topics/alpha-iota-ai-open-llm-client)

---

*Data collected daily from the US App Store and indexed by [AppsHunter](https://appshunter.io/). User reviews are verbatim App Store reviews. Ratings, prices and chart positions refresh continuously; this snapshot is from 2026-08-24.*
