
MLXHub: Local AI & LLM Server
Private offline AI, Wi-Fi mesh
About
Run powerful AI language models privately on your iPhone or iPad. Your data stays on your device, with options for distributed inference across multiple devices and a local LAN server for API access.
What's New in MLXHub
2.0.0
August 25, 2026
Your personal AI platform reaches version 2.0.0 just two months after launch! We have made a huge number of changes so you can enjoy the app more, without having to make your life complicated trying to understand it. To begin with, the interface has been polished to fit more information into the same space. We have represented concepts visually to reduce cognitive load and improved other areas that needed it. The app now comes to life not only through its orbs —rendered entirely with Metal— but also through more refined animations. An experimental feature has been added: distributed inference! You can now run larger models that do not fit on your device. Have an iPad sitting in a drawer? Add it to the mesh. Your partner’s iPhone? Add that too, and connect devices like puzzle pieces to build an increasingly capable inference server. We admit the previous value proposition was not clear: the MLXHub Plus subscription offered very little, but that changes now. MLXHub Plus now includes LAN Server, distributed-inference hosting, unlimited models, app customization, and future features that will continue adding value. Do not want to subscribe? No problem: there is also a one-time purchase! Version 2.0.0 includes many under-the-hood changes that you will not easily notice, such as fixes for most VLMs that sometimes caused issues, automatic cleanup of partial files generated by the app, improvements to model downloads, more detailed model pages, and many framework improvements that make MLXHub the best app for carrying cutting-edge AI in your pocket: without internet, without model removals, and without the fear that some orange-haired character may feel like blocking your access from an armchair with paper and pen. Enjoy these new features and the immense number of changes still to come. Do not hesitate to contact us with any question or suggestion. Best regards, Juan
MoreCharts
Events
Developer apps
FAQ
What is MLXHub?
MLXHub allows you to run open-source language models and vision-language models directly on your iPhone and iPad. It prioritizes privacy by keeping all model inference and conversations on your device.
Can I run MLXHub models offline?
Yes, MLXHub is designed for offline use. Model inference happens entirely on your device, meaning you do not need an internet connection to chat with downloaded models.
What is distributed inference in MLXHub?
Distributed inference is an experimental feature that lets you pool memory from multiple Apple Silicon devices on the same Wi-Fi network to run larger models than any single device could handle. Hosting this feature requires MLXHub Plus.
How does MLXHub handle privacy?
MLXHub ensures privacy by performing all model inference on your device, meaning your data never leaves. It does not require an account and is ad-free. Anonymous diagnostic data is collected but can be turned off.
Is MLXHub free to use?
MLXHub is free to download and use for running AI models locally and chatting. Features like hosting a distributed inference mesh and the local LAN server are part of MLXHub Plus.
What devices are supported by MLXHub?
MLXHub is available on iPhone and iPad devices. It is built to leverage Apple Silicon for optimal performance, with specific features like Apple Intelligence integration on iPhone 15 Pro/16 and later.
How often is MLXHub updated?
The latest version of MLXHub is 2.0.0, last updated on August 25, 2026. Updates typically focus on improving model compatibility, performance, and adding new features to the local AI experience.
Can I use MLXHub models with other applications?
Yes, MLXHub can function as a local LAN server (part of MLXHub Plus) that exposes an API compatible with /v1/chat/completions. This allows other applications on your local network, such as a Mac running OpenCode or custom scripts, to utilize your device's models.






