Best local AI runner (for me)
Many conveniences. Runs LM Studio models with no extra steps (unlike Ollama). Loads models faster than other runners, and supports new models sooner. Sometimes the author makes custom LLM versions to enable advanced Inferencer features. Author has xCreate YT channel evaluating new LLMs & demonstrating Inferencer features.
Show more
Response from developer
Thanks so much for the support.
Local LM Studio? Nope
Tried to connect to my tailscale LM Studio by entering in the IP:1234 and various other typical forms typically accepted by local AI apps and it didn’t work.
Response from developer
The Server feature is for hosting or connecting to other Inferencer app servers not other applications.
This App Blew My Mind! The Best AI Experience I’ve Ever Had
This is hands-down the most impressive, enjoyable, and genuinely fun AI tool I’ve used on my iPhone, iPad Pro.
The privacy alone is next-level. Everything happens 100% offline—no data ever leaves my phone.
I can download the absolute latest SOTA models (DeepSeek, Qwen, Kimi, GLM, MiniMax, and more) straight from Hugging Face in minutes, then run them locally with zero cloud dependency.
It’s freaking liberating.
The interface is buttery smooth and intuitive—I can:
• Switch models mid-chat instantly
• Tweak system prompts, temperature, and advanced inference settings on the fly
• Use the Token Entropy inspector to literally peek inside the model’s brain (super cool and nerdy)
• Force structured HTML output or skip the usual AI fluff with one tap
It’s like having a professional AI studio in my pocket, but way more enjoyable than any clunky desktop tool I’ve tried.
If you want real AI power without sacrificing privacy, speed, or fun—this is the one.
I’m genuinely impressed and can’t stop playing with it.
Best AI app I’ve used. Download it now—you’ll thank me later. 🚀
Check out his YouTube channel @xcreate for the latest updates on everything Ai.
The privacy alone is next-level. Everything happens 100% offline—no data ever leaves my phone.
I can download the absolute latest SOTA models (DeepSeek, Qwen, Kimi, GLM, MiniMax, and more) straight from Hugging Face in minutes, then run them locally with zero cloud dependency.
It’s freaking liberating.
The interface is buttery smooth and intuitive—I can:
• Switch models mid-chat instantly
• Tweak system prompts, temperature, and advanced inference settings on the fly
• Use the Token Entropy inspector to literally peek inside the model’s brain (super cool and nerdy)
• Force structured HTML output or skip the usual AI fluff with one tap
It’s like having a professional AI studio in my pocket, but way more enjoyable than any clunky desktop tool I’ve tried.
If you want real AI power without sacrificing privacy, speed, or fun—this is the one.
I’m genuinely impressed and can’t stop playing with it.
Best AI app I’ve used. Download it now—you’ll thank me later. 🚀
Check out his YouTube channel @xcreate for the latest updates on everything Ai.
Show more
Nice features, but issues
As per title, I like the features, but I run into issues with inferencer not loading models that are working with other Ai apps on the exact same phone. While this isn’t surprising it feels a bit odd because I can’t tell why the models fail to load. For example, I can run Bonsai 8b on my old iPhone 13 using Locally AI app, but on my iPhone 15 pro using Inferencer the model just fails to load at all.
Show more
Hugging face models not loading
PEBCAC or this is more like beta
Response from developer
Inferencer supports hundreds of models available on Hugging Face. It's recommended to keep the memory filter enabled to only download models that would fit on the available on-device memory. If you'd like to run larger models, you can use the server feature to connect to models hosted on your Mac. For newly released or niche models you'd like supported feel free to send an email or post a ticket directly on the public roadmap.










