Local AI Bench

Local AI Bench

오프라인 온디바이스 AI 벤치마크

0 ratings
Free

Rating summary

Details

  • Released
  • Updated
  • June 16, 2026
  • September 12, 2026

Features

Local AI Bench screenshot #1 for iPhone
Local AI Bench screenshot #2 for iPhone
Local AI Bench screenshot #3 for iPhone
Local AI Bench screenshot #4 for iPhone
Local AI Bench screenshot #5 for iPhone
Local AI Bench screenshot #6 for iPhone
🖼️Get Icon
Icons↘︎

About

LocalAI Bench is an on-device benchmark app that runs AI models directly on your iPhone and measures their performance and capabilities. All inference happens inside your device — your conversations and personal data never leave it. ■ On-Device Execution - Download open AI models (Llama, Qwen, Gemma, Phi, and more) and run them right on your device - Works fully offline once a model is downloaded ■ Real-Time Performance Metrics - Generation speed (tok/s), time-to-first-token (TTFT), peak memory (RAM), and thermal state — displayed live while you chat - Automated benchmark with a standard methodology: repeated measurements of text generation, prompt processing, translation, coding, and math speed ■ Capability Evaluation - Auto-graded tasks: math, coding, multilingual translation (Korean↔English/Japanese/Chinese), summarization, reasoning, and vision - Compare capability, speed, and memory at a glance (choose 2-axis, composite, or efficiency scoring) ■ Vision & Speech Models - Vision: attach photos or recognize objects through the live camera - Speech: compare Whisper (tiny/base/small) transcription speed (RTF) ■ Compare & Manage Models - A/B head-to-head between two models, stress tests (thermal throttling curve) - Search models by capability (vision, coding, reasoning) or vendor and add them, with automatic iOS compatibility checks ■ Web Search (optional) - When enabled, fetches Wikipedia references for the model — ask about recent events - Inference stays 100% on-device; automatically skipped when offline Pure on-device AI — no cloud APIs, no fees, no subscriptions. Runnable models depend on your device's performance; larger models are recommended for devices with more memory.
Show more

What's New in Local AI Bench

1.1.3

September 12, 2026

Release Notes(’26. 9. 12) • Version 1.1.2 New - Latest news search : with web search on, questions like "what's new / today / latest" - or about new products and terms not yet on Wikipedia - pull recent articles (title, source, date). The article titles are shown under the answer - Search API key (optional) : add a Brave Search or Tavily key in Settings to use page excerpts from real web results for more accurate answers. Mention a company's official site ("check Apple's website") and it looks there first. The key stays in this device's Keychain, and news / Wikipedia search keep working without one - Thinking mode toggle : turn reasoning models' thought process on or off. The same setting applies to chat and every benchmark, so model comparisons stay fair - Location (optional) : turn it on to answer "nearby / my area" questions; together with web search, "What's the weather now?" gets current conditions. Only an approximate location is used and nothing is stored - Settings screen : language, web search, thinking mode, location, context length, generation parameters and chat deletion in one place - One chat per model : the chat list shows the model actually used, and picking a different model opens a new chat - Korean items in the capability test : proverbs, honorifics, particles and common knowledge - Model library : sort by newest, with release month shown - New models : Qwen3.5, Kanana 2 (Kakao), Gemma 4 E4B, LFM2.5, Granite 4.1 -> tap "Update" to get them right away Improvements - Web search : fixed search not running with some reasoning models. Referenced document and article titles now appear under the answer - Web search accuracy : the product or name in your question is always kept in the search, and search terms no longer drift toward earlier answers - Fixed reasoning models (Qwen3.5, Gemma 4, etc.) leaking their thought process into the answer, or thinking without ever answering -> it now goes into the collapsible "Reasoning" box Stability - Fixed a rare crash when starting a new chat while an answer was still generating - Fixed cancelling a download letting several queued models download at once (always one at a time now) - Fixed a benchmark failing when re-run right after being cancelled - Live vision : fixed the camera staying on when the screen was closed during the permission prompt - Chat and benchmarks now share a single engine to avoid memory pressure (the send button is briefly disabled during a benchmark)

More

Developer apps