
BastionChat
Chat with Documents Privately
Archived App (Last seen on 30 Jul 2026)
This is an archived listing of the app previously available on the App Store.Although the app is no longer distributed by Apple, you can still view its description, screenshots, version history, ratings, and metadata for reference.
About
What's New in BastionChat
1.1.0
June 7, 2026
This is the biggest update to Bastion Chat yet. VISION & IMAGE UNDERSTANDING Ask questions about photos and images. Multimodal models can now see, describe, and analyze images directly in your conversations. Image history is preserved across sessions. AGENTIC MODE Bastion Chat can now act, not just chat. Enable Agent Mode for multi-step reasoning with autonomous tool use — your AI plans, executes, and delivers results. WEB SEARCH Optionally connect to the web for real-time information. Ground answers in live results without leaving the app or compromising your privacy. WIKIPEDIA IMPORT Pull any Wikipedia article directly into your knowledge base and query it instantly with full RAG. BUILT-IN KNOWLEDGE PACKS Curated, system-managed knowledge packs (first one available: Pet Care) come pre-loaded and searchable — no setup required. SMARTER MODEL BROWSER Redesigned with collapsible sections, family grouping, inline quantization picker, and quality tiers so you always know which model fits your device best. FAVORITES Star your most-used models for instant access at the top of the model list. PER-MODEL SAMPLING PROFILES Fine-tune temperature, context length, and inference settings independently for each model. SMARTER CUSTOM MODEL IMPORT Auto-detects capabilities (reasoning, tool calls, vision) on import. No manual configuration needed. TURBOQUANT KV CACHE New TurboQuant memory modes (3.5-bit and 2-bit) let you run larger models with dramatically less RAM — without sacrificing response quality. FLASH ATTENTION Flash Attention is now enabled by default on all Apple Silicon devices, significantly reducing memory usage during long conversations and improving inference speed. FULL METAL GPU OFFLOAD All model layers now offload to the Metal GPU by default, with device-aware thread and micro-batch tuning for the fastest possible inference on your iPhone, iPad, or Mac. GEMMA 4 THINKING SUPPORT Native support for Gemma 4's thinking block format, including proper stop sequence handling and streamed thinking output. Bug fixes: MCP crash fix, Metal SET_ROWS crash on hybrid GDN models (Qwen 3.6 27B), atomic interrupt flag, download integrity checks, and Qwen context length handling.
More








