Inferencer User Reviews

Top reviews

Best local AI runner (for me)

Many conveniences. Runs LM Studio models with no extra steps (unlike Ollama). Loads models faster than other runners, and supports new models sooner. Sometimes the author makes custom LLM versions to enable advanced Inferencer features. Author has xCreate YT channel evaluating new LLMs & demonstrating Inferencer features.
Show more

Response from developer

Thanks so much for the support.

Local LM Studio? Nope

Tried to connect to my tailscale LM Studio by entering in the IP:1234 and various other typical forms typically accepted by local AI apps and it didn’t work.

Response from developer

The Server feature is for hosting or connecting to other Inferencer app servers not other applications.

This App Blew My Mind! The Best AI Experience I’ve Ever Had

This is hands-down the most impressive, enjoyable, and genuinely fun AI tool I’ve used on my iPhone, iPad Pro.

The privacy alone is next-level. Everything happens 100% offline—no data ever leaves my phone.

I can download the absolute latest SOTA models (DeepSeek, Qwen, Kimi, GLM, MiniMax, and more) straight from Hugging Face in minutes, then run them locally with zero cloud dependency.

It’s freaking liberating.

The interface is buttery smooth and intuitive—I can:
• Switch models mid-chat instantly
• Tweak system prompts, temperature, and advanced inference settings on the fly
• Use the Token Entropy inspector to literally peek inside the model’s brain (super cool and nerdy)
• Force structured HTML output or skip the usual AI fluff with one tap

It’s like having a professional AI studio in my pocket, but way more enjoyable than any clunky desktop tool I’ve tried.

If you want real AI power without sacrificing privacy, speed, or fun—this is the one.

I’m genuinely impressed and can’t stop playing with it.

Best AI app I’ve used. Download it now—you’ll thank me later. 🚀

Check out his YouTube channel @xcreate for the latest updates on everything Ai.
Show more

Nice features, but issues

As per title, I like the features, but I run into issues with inferencer not loading models that are working with other Ai apps on the exact same phone. While this isn’t surprising it feels a bit odd because I can’t tell why the models fail to load. For example, I can run Bonsai 8b on my old iPhone 13 using Locally AI app, but on my iPhone 15 pro using Inferencer the model just fails to load at all.
Show more

Hugging face models not loading

PEBCAC or this is more like beta

Response from developer

Inferencer supports hundreds of models available on Hugging Face. It's recommended to keep the memory filter enabled to only download models that would fit on the available on-device memory. If you'd like to run larger models, you can use the server feature to connect to models hosted on your Mac. For newly released or niche models you'd like supported feel free to send an email or post a ticket directly on the public roadmap.

QWEN 3.5 9B won’t load

Got this to run QWEN 3.5 but fails. Gives the following error on both versions of the model: Received 333 parameters not in model: language_model.vision_tower.blocks.0.attn.proj.bias,language_model.vision_tower.blocks.0.attn.proj.weight,language_model.vision_tower.blocks.0.attn.qkv.bias,language_model.vision_tower.blocks.0.attn.qkv.weight,la...
Show more

Response from developer

Thanks for the report regarding the community edition of QWEN 3.5 9B, you can follow the progress of the public roadmap. In the meantime you can use the "inferencer qwen 3.5 9b" edition, or email if there's a specific quantization you require.

BEST MLX APP

No comparison in the speed and execution in getting things implemented. this is the best MLX inference app if you’re on this bandwagon.

Response from developer

Thanks for the support, if there's anything in particular you'd like, feel free to reach out via email or add an issue directly to the public roadmap.

Advertised free feature missing

I finally purchased to get the full thing. I really think the free version should have usable per-token probabilities. I am getting into multi-mac inferencing, so now paying for Pro and updating rating.

The free version is supposed to show up to 10 token probabilities per token, but enabling the Inspector shows an empty list every time. I am unemployed I cannot afford to support someone else right now (to upgrade to paying every month for this app). The app itself is very impressive, and I might pay one time for advanced features, but to pay every month is a hard sell for those without income.
The YouTube channel for this product is quite good and I really wanted to see token probabilities, but it is always empty. For reference this is first launch since installation running GLM 4.7 8-bit MLX on 512 GB Mac Studio (thankfully I already have that machine!).
I will re-address my score if circumstances would lead me to do so.
Show more

Response from developer

Hi, thanks so much for the bug report. I'll add the ticket to the public roadmap to investigate and fix. If you can post or email a screenshot of the issue being manifested for you that would be of great help narrowing it down. Kind regards.

Versatile local inference

I’ve been using this app for a few months now for local AI inference and it’s really helped me understand how the models generate. Every update seems to be more useful features, especially in the parallel and optimized inferencing.

Response from developer

Thanks so much for the support. If there's anything you ever need feel free to get in touch or post an issue on the roadmap.

Most Quality Inferencing for OS X

The maker of this app, xCreate is well beyond knowledgeable when it comes to inferencing LLM’s. Though some functions of using it as an agentic coding api endpoint needs work, the creator of this app is updating and fixing things on a literal daily basis. Keep in mind that he is allowing all of this to be used for free.
Show more

Response from developer

Thanks so much for the support. If there's any specific issues you have with the api endpoints, please post them on the issues page or send over an email and we'll get it fixed for you.

Downloading models broken

App would not download models but the developer responded to email with a quick fix, and it should be fixed in next version. This is an interesting program - I haven’t figured out all the bells and whistles.

Response from developer

Hi, an update was submitted to Apple earlier in the week to address an issue with downloading on macOS 26. As a workaround, you can download the models directly from HuggingFace and place them directly in your chosen Downloads folder. Feel free to report any other issues or requests you have on the public roadmap.

Probably good, but…

As someone who has multiple Macs, I was looking forward to using this for distributed inference, but I think that charging a monthly subscription is a really bad business model. I would love to support the development of this program, but monthly fees are the exact kind of thing that I wanted to avoid by investing in local AI hardware. I hope you consider changing this to a one time payment.
Show more

Response from developer

Hi Jack, thanks so much for sharing your perspective.

Like nothing before

This app does something nothing else does: allow you to see alternative tokens in nearly real-time. And to adjust the output to your framing without complicated setup or per-model sidebar configuration. You can do it per conversation. For AI researchers, the curious, and those looking to tweak model outputs? This is a fabulous tool. Highly recommended for the curious and local LLM enthusiasts.
Show more

Response from developer

Thank you so much for the support and feedback, if there's any other features you'd like implemented feel free to add it to the public roadmap.