To clarify, this MIT-licensed app is from the very same dev, 'Prince Canuma', who maintains the popular MLX-VLM library (<a href="https://github.com/Blaizzy/mlx-vlm" rel="nofollow">https://github.com/Blaizzy/mlx-vlm</a>). MLX-VLM is a long-time dependency of the excellent LM Studio and others because it can provide faster inference on Apple devices than llama.cpp. Historically, MLX is a smaller community than CUDA, but has some of the fastest updates upon the release of new models, particularly in models with modalities beyond text-in, text-out (vision, STT, TTS, video gen). See also (<a href="https://github.com/Blaizzy/mlx-audio-swift" rel="nofollow">https://github.com/Blaizzy/mlx-audio-swift</a>). Would be totally unsurprised if those modalities and models get integrated into this UI.<p>Possibly vibe-coded landing page notwithstanding, the app is mostly written in Swift language. That suggests it will be easy to port this inference stack to iPad and iPhone.
I was so excited when I saw "blaizzy" in the domain, because Prince Canuma's work around MLX has been of such uniquely high quality.
Hi Simon, sorry to spam your comments but 7 months ago you asked me for a media report to back a claim I made and this week it finally arrived. perhaps a little late to be very useful to anyone who maintains that gating function, but here nonetheless:<p><a href="https://ntindependent.com.au/scientist-says-ministers-pole-flip-and-earth-axis-tilt-water-claims-are-false/" rel="nofollow">https://ntindependent.com.au/scientist-says-ministers-pole-f...</a>
<p><pre><code> Would be totally unsurprised if those modalities and models get integrated into this UI.
</code></pre>
Yup, the GitHub repo says:<p><pre><code> Support for dedicated audio-only and image-generation-only models is coming soon.
</code></pre>
Prince Canuma is super-responsive on X and GitHub issues, and I use mlx-audio almost daily with mlx-community/Qwen3-TTS-12Hz-1.7B-Base-bf16 (for voice cloning).
I would also note that for people who want to download these models, you can find MLX versions of just about everything popular on huggingface these days. For instance go look at the "main" page for Qwen 3.6 35B-A3B and then follow the link to quantizations, and pick one of the more popular/reputable MLX variants.<p><a href="https://huggingface.co/Qwen/Qwen3.6-35B-A3B" rel="nofollow">https://huggingface.co/Qwen/Qwen3.6-35B-A3B</a>
Thanks for that. My first question was “What does this do that Unsloth doesn’t?”
LM Studio is trash on Windows / Linux... guess that makes sense...
its a wrapper around mlx, so thats gonna be the portability bottleneck
Is „frontier“ overused? I thought frontier models were the best-of-the-best such as Fable right now. I assume you can’t host these models yourself since you would need many GB of RAM and expensive GPU of is my thinking of „frontier models“ wrong?
I was confused too, but I believe that it refers to the Pareto Frontier: the best set of solutions to a multi-objective problem.<p>Look it up, it’s a bit difficult to explain concisely in words but it is intuitive visually.<p>If we are thinking of intelligence and price, a model will be in the Pareto Frontier if there’s no cheaper model of the same or higher intelligence. Or if there’s no more intelligent model for that price or lower.<p>EDIT: See this chart from Artificial Analysis: <a href="https://artificialanalysis.ai/#intelligence-comparison-tabs" rel="nofollow">https://artificialanalysis.ai/#intelligence-comparison-tabs</a><p>So for example DeepSeek V4 Pro can be considered a frontier model because there's no cheaper model that is as intelligent.<p>For any solution in the Pareto Frontier, there no "no-brainer" alternative, in the sense that there's no other option that is better in some way without giving up something else. It's the best of its "weight-class".
Very interesting rabbit hole for today. Thanks for mentioning Pareto frontier!
The subtitle on their website is: "Nativ puts frontier intelligence on your desk".<p>Also, the title is "Run AI models
locally on your Mac," not "Run frontier open models locally on your Mac."
The frontier is a multidimensional space defined by the “best” combination of traits a model can have in multiple dimensions: parameter size [smaller is better], various task metrics, and relative token generation speed on like hardware [faster is better], and active memory requirements [smaller is better] are all possible dimensions, and a model on the frontier of the current options space is one where getting better on one of those measures cannot be done without getting worse on at least one of the others.
Had the same though. The gap between open-source and the frontier is closing in, especially with Kimi K3, but that is like >2T parameters. The Gemma 4 and other models you can actually run on an average Mac, is not in the same league.
I don't know that I agree with this specific use of frontier because it is confusable as you say.<p>But (off on a tangent) I do think that there are multiple frontiers generally — and I also think the open weights, small local model frontier is <i>by far</i> the most important and exciting one.<p>I keep mucking about with what Gemma 4 12B can do and every time I do I find myself thinking that all the energies in the AI world are going in <i>entirely</i> the wrong direction, because it is small, clever, efficient and remarkable.<p>If all of that research money were to be spent on improving AI models that fit inside a 16GB RAM machine with unified memory, I think really important progress could be made.<p>I enjoy using the Qwen 3.6 models (and BottleCap's new fine tune of the 27B) but the small Gemma 4 models are impressive in a way that I think is going quite unreported.<p>So while I don't think this website <i>should</i> use the word "frontier" here, even referring to Qwen 3.6 27B which is weirdly close, I think it <i>could</i>.
No, your thinking is 100% correct. It's called clickbait.
The frontier is a curve. <a href="https://en.wikipedia.org/wiki/Pareto_front" rel="nofollow">https://en.wikipedia.org/wiki/Pareto_front</a>
That is one sense, but it is also use to refer to the frontier <i>of capability</i>. Especially in this context.
That's not what people are normally referring to when they say "frontier models". It means the most capable models full stop. Not the most capable that you can run locally.
It might be overused elsewhere but I don't think its use here is inappropriate given it's qualified: it doesn't say "frontier models", it says "frontier open models".
+1 came here to say this, I opened the link expecting some technical breakthrough. Misleading click bait title.
I'm surprised that their home page basically acts as if LM Studio and others don't already do this. It's not clear what the difference is from a glance.<p>It also omits Open WebUI. I've been running Deepseek V4 Flash locally on my Macbook Pro for weeks using Open WebUI + DS4.
LM studio is closed source software built ON TOP OF code released by the author of Nativ.
LM Studio seems to do the same thing, except that LM Studio is not open source. So they have a point, they do something more.
I was wondering the same and assuming I'd overlooked something.
"The other “local AI” apps you’ve heard of? They’re proprietary shells built on top of open-source engines they don’t own."<p>This is a roundabout way of addressing LM Studio.
What spec is your macbook? I want to run Deepseek V4 Flash but its too slow for agents on my Strix Halo.
Genuinely curious: what are people using these smaller local models for? They are getting decently capable, but they are still small enough that I don't trust them for "real" work outside of a handful of fun toy projects.<p>Are people actually using them in coding agents? Or are they mostly using them for other things?
We've shipped some code generated by Qwen3.6 27B to production (under OpenCode). It lacks the breadth of knowledge of models like Opus, but if a change is fully inferable from the prompt and the surrounding code, it works very well. It won't be able to write something from scratch that requires niche knowledge (say, a performant inference engine tailored to Blackwell GPUs), but if it's just a PR adding a new use case to an existing project (which is usually just "load from the DB, do some invariant checks, modify the entities, store them back"), it works as well as Sonnet (provided you have the correct configuration, like recommended temperature and top-p settings, the model isn't over-quantized, you have at least 150k tokens of context available, etc.).
I don't use them as coding agents, but they can be very useful for things like text transformation, summarizing, or text extraction.<p>That said, if you have a subscription to a paid model already, you're not necessarily winning out on anything except perhaps privacy, which isn't nothing.
Qwen35ba3b can do a huge amount of data cleaning work on pretty modest hardware. Already have run about 100 billion tokens on it using 2x3090 gpus.
there is plenty of grunt work these smaller models can do. update dependencies, fix merge conflicts, write --help, markdown, or readme files for existing code. etc.<p>sometimes they fail but undo is just a "git restore" or if automated, rejecting a PR and having a better model take a crack at it.
The short answer is that HN is most likely astroturfed by Apple to advertise their hardware to clueless developers.<p>Either that, or I am highly overestimating the already low intelligence of an average Mac user.<p>You can run models like gemma4:31b and qwen3.6:35b locally and get pretty good performance across a lot of areas (especially if you use an agentic framework), but to get them to be actually usefull in the workflow, you need to have something like a dual 3090 setup to get the 100+ tokens per second. The closest you can get with Apple is to buy the $7000 Mac studio that doesn't even include a monitor.<p>Under that, you are limited to running either smaller shittier models that are only good for very small number of tasks at barely usable tok/sec, and if you run them on Macbook, the thing heats up and drains battery life like it had a full discrete gpu.<p>But for some reason, being able to run a model on Mac hardware, no matter how useful it is, somehow makes Mac a better choice for AI.<p>And again, when you consider the price within laptop space, you can get good refurbished laptop with upgradable ram and ssd, run a lightweight distro with i3wm, and get really good battery performance on par with Apple, and then pay for something like Google AI plan with Gemini, and still be way under the cost of an Macbook for 10 years.
Apple would not waste money astroturfing on HN to sell O(1000) Macs.
Wow what a hater. You know what else is thousands of dollars and doesn’t even include a monitor? An ATX case with 2 3090s in it. And that will use a kilowatt or more to do what the Studio (which is excellent) does with about 250 W.
They're great at helping me look up web dev stuff when I don't have internet access.
Looking forward to giving this a try. I have tried MLX using Rapid MLX however the LLM (Qwen) would always have hiccups and get stuck repeating itself.<p>Moving onto llama.cpp I was able to get faster tokens with MTP and a more reliable llm.<p>I wonder what other people's experiences are using MLX vs llama.cpp
FWIW on my M1 Max I have not really seen any advantage at all from MLX.<p>I am fully prepared to believe the benefits accrue more to the M3 and up (because of changes to the Apple Neural Engine).<p>But with the models I've tested, unless I am missing something, the performance of GGUFs in llama.cpp has been better in some cases.<p>I still have not had results from Gemma 4's MTP be really worth it, to be honest; but with the Qwen 3.6 MoE it is measurable. Maybe with newer kit it is more meaningful.<p>(There is every chance that the above is not the experience of anyone who really deeply knows what they are doing; it feels like I am a perpetual novice at this stuff)
This looks like the Prism folks, who are making binary/ternary versions of popular edge models, so that those models will fit on constrained devices like phones. E.g., their Bonsai model derived from Qwen:<p><a href="https://news.ycombinator.com/item?id=48910545">https://news.ycombinator.com/item?id=48910545</a><p>Perhaps they got tired of LM Studio, etc., not being able to run their models properly.
Only interesting thing about this vibe coded runner is the MLX support, as that's still annoying to use in other ones, most still use GGUFs. Unsloth Studio which is an OSS runner I use is still in progress with MLX support although it's still a ways away.
Ironic that the app is named Nativ(e) and yet bundles a full Python runtime. Nonetheless, still less bloated than LM Studio, which bundles a full Python runtime and electron.js (which in turn bundles a whole browser runtime).
The server wont start for me.
"ERROR: Application startup failed. Exiting."
"mlx-vlm-server stopped with status 3"
Is Gemma 4 E2B actually "usable"?
I've been running Gemma 4 12B and it handles everything very well!
But the second I've moved down to E4B it's been unable to perform the simplest of tasks.
So I can't even imagine how E2B would do...
Or am I doing something wrong?
I'm a bit curious why not running DeepSeek V4 on top of <a href="https://github.com/antirez/ds4" rel="nofollow">https://github.com/antirez/ds4</a>. I think the results could be really good.
Has anyone found a model that can run on a normal macbook? I have an M3 Pro with 18GB of memory and whenever I try to run even a basic model the fans goes off and the mac starts to get heated up and becomes so laggy.
There is a Gemma 4 model with 12B parameters which might be worth trying. e.g. <a href="https://huggingface.co/mlx-community/gemma-4-12B-it-qat-4bit" rel="nofollow">https://huggingface.co/mlx-community/gemma-4-12B-it-qat-4bit</a><p>That said, your computer will still get hot!
You need more RAM, plus the models take up a lot of space. 32gb min but I’d recommend 48/64gb, you won’t get close to frontier but it’s still fun to play with, images are very good
If you go small enough it should be no problem. For example Gemma 4 E4B in Q6 or Q4 quantization should run well on your laptop. It shouldn't be too taxing, but would still want to eat 7-9 GB of VRAM or so.<p>Now that model is mostly useful for writing or chatting.
heating up is normal, that cannot be avoided. it should become laggy, but you just have very little RAM (I assume 16 GB?) So most models are too big with other stuff running.
Tbh, the only reason to run a model locally is when you want to be completely safe.<p>From productivity point of view, it doesn’t make sense to have any notebook running a local LLM.<p>We have one life. We should spend it wisely.
Advice: remove all slop and fluff from the website such as "Everything you need.
Nothing you don’t."<p>Just state the information you want to communicate in the plainest and most straightforward way possible.
That phrase itself is such an astounding performative contradiction [0].<p>0: <a href="https://en.wikipedia.org/wiki/Performative_contradiction" rel="nofollow">https://en.wikipedia.org/wiki/Performative_contradiction</a>
Something nobody needs is pointless hot air like "Everything you need. Nothing you don't."
You’d be surprised how hard this actually is. I spent 3 days iterating on a marketing site, where I had very explicit / “well written” copy, and it would just repeatedly rewrite it back to the most awful slop. Over and over again! Ended up adding various AGENTS rules telling it to leave the copy alone
I don’t think anyone ever looked at result. Website if full of overflow bugs. It’s pure slop.
so now we have this, <a href="https://pypi.org/project/rapid-mlx/" rel="nofollow">https://pypi.org/project/rapid-mlx/</a>, <a href="https://mtplx.com" rel="nofollow">https://mtplx.com</a>, and the oldest I could find at <a href="https://omlx.ai" rel="nofollow">https://omlx.ai</a>.
Ok, so we're at the point that even design are just one-shot by AI. This is exact look and feel any time I ask it to present a HTML doc about anything.
Looks like my call [0] for more competitors to Ollama has been answered.<p>We need more like this as well as llama.app, which also has a native mac app.<p>[0] <a href="https://news.ycombinator.com/item?id=48968898">https://news.ycombinator.com/item?id=48968898</a>
Agreed. Just based on this not being Ollama, so I will give it a try.
my thing is kind of an Ollama competitor (surrogate?) too. more for prose/text planning, not so much for coding, at least the harness, but i'm sure someone could set it up to do that: github.com/0gsd/enough
[flagged]
[flagged]
[dead]
How does this compare to LM Studio Bionic?