
Phantasm
Self-Hosted AI Chat
0 ratings
Free
About
Phantasm is a private AI chat client for the backend you already run. Point it at your own server: Ollama, vLLM, llama.cpp, or any OpenAI-compatible endpoint and chat with your own models over your own network. No account, no middleman, no data collection.
YOUR SERVER, YOUR RULES
• Connect with just a URL and a token — works against a bare Ollama out of the box
• Add the open-source Phantasm orchestrator to unlock web search, image generation, and research modes
• Multiple backend profiles for home network, VPN, or tunnel setups
• Your token lives in the iOS Keychain; conversations are stored only on your phone
A REAL CHAT EXPERIENCE
• Fast streaming responses with markdown, syntax-highlighted code blocks with copy, and charts
• Attach photos and files — vision models can see your images
• Generated images appear inline; save or share them
• Dictate messages with on-device transcription and have replies read aloud
• Long generations survive backgrounding and pick up where they left off
• Ask Phantasm from Siri and Shortcuts
TOOLS ON YOUR TERMS
• Per-chat toggles for web access and image generation
• Optional device tools — location, calendar, Apple Health — are off by default and only shared with your own backend when you enable them
• See the model's reasoning when a thinking model is selected
PRIVATE BY DESIGN
Phantasm operates no servers and collects no data. The app talks exclusively to the backend you configure — nothing else. What you say to your AI is between you and your hardware.
Phantasm requires an OpenAI-compatible server that you provide (for example, Ollama running on your own machine). Feature availability depends on your backend's capabilities.
Show more
What's New in Phantasm
1.2
July 25, 2026
• ### What’s New in Phantasm 1.2 - Generate audio and video with compatible servers. - Play audio in chat with seeking, background playback, and Lock Screen controls. - Watch generated videos in chat and share generated media as files. - View inline and block equations with native math formatting. - Track estimated token use, context limits, response speed, and reasoning time. - Experience smoother performance in long and image-heavy chats. - Use more reliable dictation when switching chats or interrupting recording. - Improved backend compatibility and general stability fixes.
More



