{"author":"redundantly","children":[{"author":"babblingfish","children":[{"author":"gedy","children":[{"author":"aurareturn","children":[{"author":"gedy","children":[],"created_at":"2026-03-31T05:28:09.000Z","created_at_i":1774934889,"id":47583085,"options":[],"parent_id":47582979,"points":null,"story_id":47582482,"text":"True, but I&#x27;m already producing code&#x2F;features faster than company knows what to do with, (even though every company says &quot;omg we need this <i>yesterday</i>&quot;, etc).  Even coding before AI was basically same.<p>Code tools that free my time up is very nice.","title":null,"type":"comment","url":null},{"author":"dgb23","children":[],"created_at":"2026-03-31T11:41:27.000Z","created_at_i":1774957287,"id":47585890,"options":[],"parent_id":47582979,"points":null,"story_id":47582482,"text":"That&#x27;s not necessarily the case. So far, commercial cloud LLMs have maintained a head-start, but there is no law of nature that prevents us from having competitive open models.<p>In fact the space seems to move at a rapid pace as more and more specialized models come out. There&#x27;s a possible trajectory where open weight models will compete side by side or even be preferable for many use cases, just like what happened with OS&#x27;s and SQL DB&#x27;s.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T05:11:17.000Z","created_at_i":1774933877,"id":47582979,"options":[],"parent_id":47582895,"points":null,"story_id":47582482,"text":"When local LLMs get good enough for you to use delightfully, cloud LLMs will have gotten so much smarter that you&#x27;ll still use it for stuff that needs more intelligence.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T04:53:59.000Z","created_at_i":1774932839,"id":47582895,"options":[],"parent_id":47582826,"points":null,"story_id":47582482,"text":"Man I really hope so, as, as much as I like Claude Code, I hate the company paying for it and tracking your usage, bullshit management control, etc.  I feel like I&#x27;m training my replacement.  Things feel like they are tightening vs more power and freedom.<p>On device I would <i>gladly</i> pay for good hardware - it&#x27;s my machine and I&#x27;m using as I see fit like an IDE.","title":null,"type":"comment","url":null},{"author":"aurareturn","children":[{"author":"AugSun","children":[{"author":"QuantumNomad_","children":[{"author":"Ericson2314","children":[{"author":"AugSun","children":[],"created_at":"2026-03-31T06:30:01.000Z","created_at_i":1774938601,"id":47583481,"options":[],"parent_id":47583190,"points":null,"story_id":47582482,"text":"Thank you for clarifying! (I had no idea it needs to be explained, sorry.)","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T05:45:41.000Z","created_at_i":1774935941,"id":47583190,"options":[],"parent_id":47583139,"points":null,"story_id":47582482,"text":"I figured it out from context clues<p>CC: Claude Code<p>TC: total comp(ensation)","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T05:38:13.000Z","created_at_i":1774935493,"id":47583139,"options":[],"parent_id":47583026,"points":null,"story_id":47582482,"text":"What is CC and TC? I have not heard these abbreviations (except for CC to mean credit card or carbon copy, neither of which is what I think you mean here).","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T05:19:10.000Z","created_at_i":1774934350,"id":47583026,"options":[],"parent_id":47582940,"points":null,"story_id":47582482,"text":"Looking at downvotes I feel good about SDE future in 3-5 years. We will have a swamp of &quot;vibe-experts&quot; who won&#x27;t be able to pay 100K a month to CC. Meanwhile, people who still remember <i>how to code in Vim</i> will (slowly) get back to pre-COVID TC levels.","title":null,"type":"comment","url":null},{"author":"virtue3","children":[{"author":"AndroTux","children":[],"created_at":"2026-03-31T06:20:01.000Z","created_at_i":1774938001,"id":47583419,"options":[],"parent_id":47583084,"points":null,"story_id":47582482,"text":"Sir, ChatGPT 3.5 is more than 3 years old, running on your bleeding edge M4 Pro hardware, and only proves the previous commenters point.","title":null,"type":"comment","url":null},{"author":"AugSun","children":[],"created_at":"2026-03-31T06:35:15.000Z","created_at_i":1774938915,"id":47583518,"options":[],"parent_id":47583084,"points":null,"story_id":47582482,"text":"It works <i>really</i> well for &quot;You&#x27;re helpful assistant &#x2F; Hi &#x2F; Hello there. how may I help you today?&quot; Anything else (esp in non-EN language) and you will see the limitations yourself. just try it.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T05:28:07.000Z","created_at_i":1774934887,"id":47583084,"options":[],"parent_id":47582940,"points":null,"story_id":47582482,"text":"We are 100% there already.  In browser.<p>the webgpu model in my browser on my m4 pro macbook was as good as chatgpt 3.5 and doing 80+ tokens&#x2F;s<p>Local is here.","title":null,"type":"comment","url":null},{"author":"raincole","children":[{"author":"kortilla","children":[],"created_at":"2026-03-31T07:44:11.000Z","created_at_i":1774943051,"id":47583987,"options":[],"parent_id":47583329,"points":null,"story_id":47582482,"text":"Don\u2019t try to draw trend lines for an industry that has existed for &lt;5 years.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T06:06:25.000Z","created_at_i":1774937185,"id":47583329,"options":[],"parent_id":47582940,"points":null,"story_id":47582482,"text":"Yep. People were claiming DeepSeek was &quot;almost as good as SOTA&quot; when it came out. Local will always be one step away like fusion.<p>It&#x27;s just wishful thinking (and hatred towards American megacorps). Old as the hills. Understandable, but not based on reality.","title":null,"type":"comment","url":null},{"author":"mirekrusin","children":[{"author":"aurareturn","children":[{"author":"mirekrusin","children":[],"created_at":"2026-03-31T21:13:09.000Z","created_at_i":1774991589,"id":47593541,"options":[],"parent_id":47584046,"points":null,"story_id":47582482,"text":"Yes, it\u2019s expensive hobby.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T07:53:24.000Z","created_at_i":1774943604,"id":47584046,"options":[],"parent_id":47583659,"points":null,"story_id":47582482,"text":"It&#x27;s a $4,000 GPU with 32GB of VRAM and needs a 1,000 watt PSU. It&#x27;s not realistic for the masses.<p>If it has something like 80GB of VRAM, it&#x27;ll cost $10k.<p>The actual local LLM chip is Apple Silicon starting at the M5 generation with matmul acceleration in the GPU. You can run a good model using an M5 Max 128GB system. Good prompt processing and token generation speeds. Good enough for many things. Apple accidentally stumbled upon a huge advantage in local LLMs through unified memory architecture.<p>Still not for the masses and not cheap and not great though. Going to be years to slowly enable local LLMs on general mass local computers.","title":null,"type":"comment","url":null},{"author":"fredoliveira","children":[{"author":"mirekrusin","children":[],"created_at":"2026-03-31T21:13:55.000Z","created_at_i":1774991635,"id":47593547,"options":[],"parent_id":47588653,"points":null,"story_id":47582482,"text":"Look it up.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T15:19:55.000Z","created_at_i":1774970395,"id":47588653,"options":[],"parent_id":47583659,"points":null,"story_id":47582482,"text":"Crazy thing to say without other contextual information - it obviously depends on a number of factors. Do you have an apples to apples comparison at hand?","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T06:52:46.000Z","created_at_i":1774939966,"id":47583659,"options":[],"parent_id":47582940,"points":null,"story_id":47582482,"text":"Local RTX 5090 is actually faster than A100&#x2F;H100.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T05:03:40.000Z","created_at_i":1774933420,"id":47582940,"options":[],"parent_id":47582826,"points":null,"story_id":47582482,"text":"It isn&#x27;t going to replace cloud LLMs since cloud LLMs will always be faster in throughput and smarter. Cloud and local LLMs will grow together, not replace each other.<p>I&#x27;m not convinced that local LLMs use less electricity either. Per token at the same level of intelligence, cloud LLMs should run circles around local LLMs in efficiency. If it doesn&#x27;t, what are we paying hundreds of billions of dollars for?<p>I think local LLMs will continue to grow and there will be an &quot;ChatGPT&quot; moment for it when good enough models meet good enough hardware. We&#x27;re not there yet though.<p>Note, this is why I&#x27;m big on investing in chip manufacture companies. Not only are they completely maxed out due to cloud LLMs, but soon, they will be double maxed out having to replace local computer chips with ones that are suited for inferencing AI. This is a massive transition and will fuel another chip manufacturing boom.","title":null,"type":"comment","url":null},{"author":"AugSun","children":[{"author":"selcuka","children":[{"author":"asutekku","children":[{"author":"Barbing","children":[{"author":"jychang","children":[],"created_at":"2026-03-31T07:44:32.000Z","created_at_i":1774943072,"id":47583990,"options":[],"parent_id":47583748,"points":null,"story_id":47582482,"text":"The free version of ChatGPT is insanely crippled, so that&#x27;s not surprising.","title":null,"type":"comment","url":null},{"author":"throwaway27448","children":[],"created_at":"2026-03-31T08:46:13.000Z","created_at_i":1774946773,"id":47584425,"options":[],"parent_id":47583748,"points":null,"story_id":47582482,"text":"If someone blindly submits chatbot output they deserve to be embarrassed and fired. But I don&#x27;t think that&#x27;s going to improve.","title":null,"type":"comment","url":null},{"author":"theshrike79","children":[],"created_at":"2026-03-31T08:47:49.000Z","created_at_i":1774946869,"id":47584447,"options":[],"parent_id":47583748,"points":null,"story_id":47582482,"text":"Even the paid version of ChatGPT tends to use a 1000 words when 10 will do.<p>You can try asking it the same question as Claude and compare the answers. I can guarantee you that the ChatGPT answer won&#x27;t fit on a single screen on a 32&quot; 4k monitor.<p>Claude&#x27;s will.","title":null,"type":"comment","url":null},{"author":"PhilipRoman","children":[],"created_at":"2026-03-31T13:00:35.000Z","created_at_i":1774962035,"id":47586740,"options":[],"parent_id":47583748,"points":null,"story_id":47582482,"text":"I use the free version of ChatGPT (without logging in) when I need some one-off question without a huge context. Real world prompt:<p><pre><code>  &quot;when hostapd initializes 80211 iface over nl80211, what attributes correspond to selected standard version like ax or be?&quot;\n</code></pre>\nIt works fine, avoids falling into trap due to misleading question. Probably works even better for more popular technologies. Yeah, it has higher failure rates but it&#x27;s not a dealbreaker for non-autonomous use cases.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T07:05:31.000Z","created_at_i":1774940731,"id":47583748,"options":[],"parent_id":47583546,"points":null,"story_id":47582482,"text":"re: trust-<p>Have you tried the free version of ChatGPT? It is positively appalling. It\u2019s like GPT 3.5 but prompted to write three times as much as necessary to seem useful. I wonder how many people have embarrassed themselves, lost their jobs, and been critically misinformed. All easy with state-of-the-art models but seemingly a guarantee with the bottom sub-slop tier.<p>Is the average person just talking to it about their day or something?","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T06:39:32.000Z","created_at_i":1774939172,"id":47583546,"options":[],"parent_id":47583280,"points":null,"story_id":47582482,"text":"Frontier model has much better knowledge and they usually hallucinate less. It&#x27;s not about the coding capabilities, it&#x27;s about how much you can trust the model.","title":null,"type":"comment","url":null},{"author":"lxgr","children":[{"author":"throwaway27448","children":[{"author":"embedding-shape","children":[],"created_at":"2026-03-31T10:51:26.000Z","created_at_i":1774954286,"id":47585446,"options":[],"parent_id":47584416,"points":null,"story_id":47582482,"text":"They&#x27;re awful and hallucinate a lot, I couldn&#x27;t imagine using it even for prompts about TV shows, even less so for serious work. Repeating the question from the parent, have you tried those yourself? Even compared to ChatGPT Thinking, they&#x27;re short of useless.","title":null,"type":"comment","url":null},{"author":"lxgr","children":[],"created_at":"2026-03-31T13:45:30.000Z","created_at_i":1774964730,"id":47587327,"options":[],"parent_id":47584416,"points":null,"story_id":47582482,"text":"They&#x27;re essentially replying based on vibes, instead of grounding their responses in extensive web searches, which is what the paid models&#x2F;configurations generally do. This makes them wrong more often than they&#x27;re right for anything but the most trivial requests that can be easily responded to out of memorized training data.<p>This is all on top of the (to me) insufferable tone of the non-thinking models, but that might well be how most users prefer to be talked to, and whether that&#x27;s how these models should accordingly talk is a much more nuanced question.<p>Regardless of that, everybody deserves correct answers, even users on the free tier. If this makes the free tier uneconomical to serve for hours on end per user per day, then I&#x27;d much rather they limit the number of turns than dial down the quality like that.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T08:45:16.000Z","created_at_i":1774946716,"id":47584416,"options":[],"parent_id":47583827,"points":null,"story_id":47582482,"text":"Say more. Why do you think this?","title":null,"type":"comment","url":null},{"author":"selcuka","children":[{"author":"lxgr","children":[],"created_at":"2026-04-01T10:39:17.000Z","created_at_i":1775039957,"id":47599103,"options":[],"parent_id":47596093,"points":null,"story_id":47582482,"text":"&gt; Obviously it&#x27;s &quot;good enough&quot; for &quot;most people&quot;. Otherwise nobody would be using the free version of ChatGPT today.<p>I&#x27;d say it&#x27;s better than nothing, which to me is not the same thing at all as &quot;good enough&quot;.<p>For example, I believe most people would be better off with half the allowable queries per day, routed to a better model, but that&#x27;s not an available product.","title":null,"type":"comment","url":null}],"created_at":"2026-04-01T02:34:20.000Z","created_at_i":1775010860,"id":47596093,"options":[],"parent_id":47583827,"points":null,"story_id":47582482,"text":"&gt; I think it\u2019s pretty cynical to assume that this is \u201cgood enough for most people\u201d<p>It&#x27;s a deduction, not an assumption. Obviously it&#x27;s &quot;good enough&quot; for &quot;most people&quot;. Otherwise nobody would be using the free version of ChatGPT today.<p>I pay for a Claude subscription, but even then I sometimes downgrade to Sonnet or even Haiku when I need a quick answer.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T07:17:23.000Z","created_at_i":1774941443,"id":47583827,"options":[],"parent_id":47583280,"points":null,"story_id":47582482,"text":"Have you used GPT instant or mini yourself? I think it\u2019s pretty cynical to assume that this is \u201cgood enough for most people\u201d, even if they don\u2019t know the difference between that and better models.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T05:58:07.000Z","created_at_i":1774936687,"id":47583280,"options":[],"parent_id":47582950,"points":null,"story_id":47582482,"text":"Any citations? Because that was my impression, too. I want frontier model performance for my coding assistant, but &quot;most users&quot; could do with smaller&#x2F;faster models.<p>ChatGPT free falls back to GPT-5.2 Mini after a few interactions.","title":null,"type":"comment","url":null},{"author":"helsinkiandrew","children":[{"author":"drob518","children":[],"created_at":"2026-03-31T12:12:07.000Z","created_at_i":1774959127,"id":47586183,"options":[],"parent_id":47584011,"points":null,"story_id":47582482,"text":"\u201cYou are the smartest high school student that has ever lived and on the college track to Harvard or another Ivy League school. Write a 10 page history term paper about Tiananmen Square and the specific events that took place there. Include a bibliography and use footnotes to cite sources.\u201d","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T07:47:52.000Z","created_at_i":1774943272,"id":47584011,"options":[],"parent_id":47582950,"points":null,"story_id":47582482,"text":"&gt; unfortunately, this is not the case<p>Most users are fixing grammar&#x2F;spelling, summarising&#x2F;converting&#x2F;rewriting text, creating funny icons, and looking up simple facts, this is all far from frontier model performance.<p>I&#x27;ve a feeling that if&#x2F;when Apple release their onboard LLM&#x2F;Siri improvements that can call out if needed, the vast majority of people will be happy with what they get for free that&#x27;s running on their phone.","title":null,"type":"comment","url":null},{"author":"blitzar","children":[],"created_at":"2026-03-31T08:45:25.000Z","created_at_i":1774946725,"id":47584417,"options":[],"parent_id":47582950,"points":null,"story_id":47582482,"text":"&quot;Hey dingus, set timer for 30 minutes&quot;","title":null,"type":"comment","url":null},{"author":"theshrike79","children":[{"author":"ZeroGravitas","children":[{"author":"zozbot234","children":[],"created_at":"2026-03-31T09:31:14.000Z","created_at_i":1774949474,"id":47584802,"options":[],"parent_id":47584671,"points":null,"story_id":47582482,"text":"The Wikipedia folks are now working on implementing a language-independent representation for their encyclopedic content - one that&#x27;s intended to be rigorously compositional and semantics-aware, loosely comparable to Universal Meaning Representation (UMR) as known in the linguistics domain, that - if successful - may end up interacting in very interesting ways with multi-language capable LLMs.  Very early experiments (nowhere near as capable as UMR as of yet, but experimenting with the underlying software infrastructure) are at <a href=\"https:&#x2F;&#x2F;abstract.wikipedia.org\" rel=\"nofollow\">https:&#x2F;&#x2F;abstract.wikipedia.org</a> , whilst a direct comparison of the projected design is given by <a href=\"https:&#x2F;&#x2F;commons.wikimedia.org&#x2F;wiki&#x2F;File:Abstract_Wikipedia_NLG_SIG_Meeting_2025-11.webm\" rel=\"nofollow\">https:&#x2F;&#x2F;commons.wikimedia.org&#x2F;wiki&#x2F;File:Abstract_Wikipedia_N...</a> <a href=\"https:&#x2F;&#x2F;elemwala.toolforge.org&#x2F;static&#x2F;nlgsig-nov2025.html\" rel=\"nofollow\">https:&#x2F;&#x2F;elemwala.toolforge.org&#x2F;static&#x2F;nlgsig-nov2025.html</a>","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T09:14:50.000Z","created_at_i":1774948490,"id":47584671,"options":[],"parent_id":47584434,"points":null,"story_id":47582482,"text":"You can install the complete text of Wikipedia locally too.<p>They&#x27;ve usually been intended for ereader&#x2F;off-grid&#x2F;post-zombie-apocalypse situations but I&#x27;d guess someone is working on an llm friendly way to install it already.<p>Be interesting to know the tradeoffs. The Tienammen square example suggests why you&#x27;d maybe want the knowledge facts to come from a separate source.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T08:46:40.000Z","created_at_i":1774946800,"id":47584434,"options":[],"parent_id":47582950,"points":null,"story_id":47582482,"text":"It depends. If they&#x27;re using a small&#x2F;medium local model as a 1:1 ChatGPT replacement as-is, they&#x27;ll have a bad time. Even ChatGPT refers to external services to get more data.<p>But a local model + good harness with a robust toolset will work for people more often than not.<p>The model itself doesn&#x27;t need to know who was the president of Zambia in 1968, because it has a tool it can use to check it from Wikipedia.","title":null,"type":"comment","url":null},{"author":"cyanydeez","children":[],"created_at":"2026-03-31T10:27:49.000Z","created_at_i":1774952869,"id":47585232,"options":[],"parent_id":47582950,"points":null,"story_id":47582482,"text":"eh, its weird how thetech world wants to build trillions of data centers for...what, escapingthe permanent underclass?<p>I think what &quot;need&quot; youspeak of is a bit of a colored statement.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T05:05:20.000Z","created_at_i":1774933520,"id":47582950,"options":[],"parent_id":47582826,"points":null,"story_id":47582482,"text":"&quot;Most users don&#x27;t need frontier model performance&quot; unfortunately, this is not the case.","title":null,"type":"comment","url":null},{"author":"melvinroest","children":[{"author":"nkzd","children":[{"author":"melvinroest","children":[],"created_at":"2026-03-31T08:46:11.000Z","created_at_i":1774946771,"id":47584424,"options":[],"parent_id":47583531,"points":null,"story_id":47582482,"text":"TL;DR: you don&#x27;t need to do any treasure hunt on your notes by just typing stuff into the search bar. Having your own graphRAG system + LLM on your notes is basically a &quot;Google&quot; but then on your own notes. Any question you have: if you have a note for it, it will bubble up. The annoying thing is that false positives will also bubble up.<p>----<p>Full reaction:<p>Yes but perhaps not in a way you might expect. Qwen&#x27;s reasoning ability isn&#x27;t exactly groundbreaking. But it&#x27;s good enough to weave a story, provided it has some solid facts or notes. GraphRAG is definitely a good way to get some good facts, provided your notes are valuable to you and&#x2F;or contain some good facts.<p>So the added value is that you now have a super charged information retrieval system on your notes with an LLM that can stitch loose facts reasonably well together, like a librarian would. It&#x27;s also very easy to see hallucinations, if you recognize your own writing well, which I do.<p>The second thing is that I have a hard time rereading all my notes. I write a lot of notes, and don&#x27;t have the time to reread any of them. So oftentimes I forget my own advice. Now that I have a super charged information retrieval system on my notes, whenever I ask a question: the graphRAG + LLM search for the most relevant notes related to my question. I&#x27;ve found that 20% of what I wrote is incredibly useful <i>and</i> is stuff that I forgot.<p>And there are nuggets of wisdom in there that are quite nuanced. For me specifically, I&#x27;ve seen insights in how I relate to work that I should do more with. I&#x27;ll probably forget most things again but I can reuse my system and at some point I&#x27;ll remember what I actually need to remember. For example, one thing I read was that work doesn&#x27;t feel like work for me if I get to dive in, zoom out, dive in, zoom out. Because in the way I work as a person: that means I&#x27;m always resting and always have energy for the task that I&#x27;m doing. Another thing that it got me to do was to reboot a small meditation practice by using implementation intentions (e.g. &quot;if I wake up then I meditate for at least a brief amount of time&quot;).<p>What also helps is to have a bit of a back and forth with your notes and then copy&#x2F;paste the whole conversation in Claude to see if Claude has anything in its training data that might give some extra insight. It could also be that it just helps with firing off 10 search queries and finds a blog post that is useful to the conversation that you&#x27;ve had with your local LLM.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T06:36:59.000Z","created_at_i":1774939019,"id":47583531,"options":[],"parent_id":47583014,"points":null,"story_id":47582482,"text":"Did you get any insights about yourself from this process? I am thinking of doing the same","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T05:16:21.000Z","created_at_i":1774934181,"id":47583014,"options":[],"parent_id":47582826,"points":null,"story_id":47582482,"text":"I have journaled digitally for the last 5 years with this expectation.<p>Recently I built a graphRAG app with Qwen 3.5 4b for small tasks like classifying what type of question I am asking or the entity extraction process itself, as graphRAG depends on extracted triplets (entity1, relationship_to, entity2). I used Qwen 3.5 27b for actually answering my questions.<p>It works pretty well. I have to be a bit patient but that\u2019s it. So in that particular use case, I would agree.<p>I used MLX and my M1 64GB device. I found that MLX definitely works faster when it comes to extracting entities and triplets in batches.","title":null,"type":"comment","url":null},{"author":"pezgrande","children":[{"author":"aurareturn","children":[{"author":"spiderfarmer","children":[{"author":"aurareturn","children":[{"author":"seanmcdirmid","children":[{"author":"aurareturn","children":[{"author":"spiderfarmer","children":[{"author":"aurareturn","children":[{"author":"spiderfarmer","children":[],"created_at":"2026-04-01T10:50:51.000Z","created_at_i":1775040651,"id":47599179,"options":[],"parent_id":47591639,"points":null,"story_id":47582482,"text":"Right. When I said &quot;you&#x27;ll always be wrong&quot;, I meant you&#x27;re sometimes right.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T18:39:18.000Z","created_at_i":1774982358,"id":47591639,"options":[],"parent_id":47586750,"points":null,"story_id":47582482,"text":"I&#x27;ve been right far more than wrong on this stuff. :)","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T13:01:39.000Z","created_at_i":1774962099,"id":47586750,"options":[],"parent_id":47583729,"points":null,"story_id":47582482,"text":"You will always be wrong.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T07:02:36.000Z","created_at_i":1774940556,"id":47583729,"options":[],"parent_id":47583472,"points":null,"story_id":47582482,"text":"Right. When I said &quot;they&#x27;ll always be behind&quot;, I meant in the next 5-10 years. They&#x27;re gated by EUV tech. And once they have EUV tech, they need to scale up chip manufacturing.","title":null,"type":"comment","url":null},{"author":"Barbing","children":[{"author":"seanmcdirmid","children":[],"created_at":"2026-03-31T15:59:11.000Z","created_at_i":1774972751,"id":47589345,"options":[],"parent_id":47583761,"points":null,"story_id":47582482,"text":"Both are hard nuts but China is throwing massive amounts of money at the problem. They can already get performance or economy from each, they just need to figure out how to get both at the same time.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T07:07:37.000Z","created_at_i":1774940857,"id":47583761,"options":[],"parent_id":47583472,"points":null,"story_id":47582482,"text":"Which might they master first?","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T06:27:21.000Z","created_at_i":1774938441,"id":47583472,"options":[],"parent_id":47583277,"points":null,"story_id":47582482,"text":"About 2.5 decades from the start of the JVs, but they did it. Semiconductors and jet turbines are really the last two tech trees that China has yet to master.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T05:57:38.000Z","created_at_i":1774936658,"id":47583277,"options":[],"parent_id":47583250,"points":null,"story_id":47582482,"text":"It did take decades to catch and surpass US car makers right?","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T05:53:53.000Z","created_at_i":1774936433,"id":47583250,"options":[],"parent_id":47583042,"points":null,"story_id":47582482,"text":"\u201cThey will always be behind\u201d<p>Car manufacturers said the same.","title":null,"type":"comment","url":null},{"author":"pezgrande","children":[{"author":"aurareturn","children":[{"author":"RALaBarge","children":[{"author":"aurareturn","children":[],"created_at":"2026-03-31T18:38:51.000Z","created_at_i":1774982331,"id":47591635,"options":[],"parent_id":47586590,"points":null,"story_id":47582482,"text":"No need inside line. Just look at chip node tech.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T12:47:53.000Z","created_at_i":1774961273,"id":47586590,"options":[],"parent_id":47583741,"points":null,"story_id":47582482,"text":"You must have an inside line on information for &#x27;China&#x27; -- those are bold predictions!","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T07:04:39.000Z","created_at_i":1774940679,"id":47583741,"options":[],"parent_id":47583302,"points":null,"story_id":47582482,"text":"I highly doubt they can catch up in 3-5 years to Nvidia.<p>Chips take about 3 years to design. Do you think China will have Feymann-level AI systems in 3 years?<p>I think in 3 years, they&#x27;ll have H200-equivalent at home.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T06:02:10.000Z","created_at_i":1774936930,"id":47583302,"options":[],"parent_id":47583042,"points":null,"story_id":47582482,"text":"&gt; have to release free open source models because they distill from OpenAI and Anthropic<p>They dont really have to though, they just need to be good enough and cheaper (even if distilled). That being said, it is true they are gaining a lot of visibility (specially Qwen) because of being open-source(weight).<p>Hardware-wise they seem they will catch-up in 3-5 years (Nvidia is kind of irrelevant, what matters is the node).","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T05:21:18.000Z","created_at_i":1774934478,"id":47583042,"options":[],"parent_id":47583019,"points":null,"story_id":47582482,"text":"I agree. I can totally see in the future that open source LLMs will turn into paying a lumpsum for the model. Many will shut down. Some will turn into closed source labs.<p>When VCs inevitably ask their AI labs to start making money or shut down, those free open source LLMS will cease to be free.<p>Chinese AI labs have to release free open source models because they distill from OpenAI and Anthropic. They will always be behind. Therefore, they can&#x27;t charge the same prices as OpenAI and Anthropic. Free open source is how they can get attention and how they can stay fairly close to OpenAI and Anthropic. They have to distill because they&#x27;re banned from Nvidia chips and TSMC.<p>Before people tell me Chinese AI labs do use Nvidia chips, there is a huge difference between using older gimped Nvidia H100 (called H20) chips or sneaking around Southeast Asia for Blackwell chips and officially being allowed to buy millions of Nvidia&#x27;s latest chips to build massive gigawatt data centers.","title":null,"type":"comment","url":null},{"author":"Lio","children":[],"created_at":"2026-03-31T06:32:29.000Z","created_at_i":1774938749,"id":47583498,"options":[],"parent_id":47583019,"points":null,"story_id":47582482,"text":"This seems to be somewhat similar to web browsers.<p>I could see the model becoming part of the OS.<p>Of course Google and Microsoft will still want you to use their models so that they can continue to spy on you.<p>Apple, AMD and Nvidia would sell hardware to run their own largest models.","title":null,"type":"comment","url":null},{"author":"mirekrusin","children":[],"created_at":"2026-03-31T06:46:39.000Z","created_at_i":1774939599,"id":47583603,"options":[],"parent_id":47583019,"points":null,"story_id":47582482,"text":"You can have viable business model around open weight models where you offer fine tuning at a fee.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T05:17:11.000Z","created_at_i":1774934231,"id":47583019,"options":[],"parent_id":47582826,"points":null,"story_id":47582482,"text":"You could argue that the only reason we have good open-weight models is because companies are trying to undermine the big dogs, and they are spending millions to make sure they dont get too far ahead. If the bubble pops then there wont be incentive to keep doing it.","title":null,"type":"comment","url":null},{"author":"karimf","children":[{"author":"Barbing","children":[{"author":"karimf","children":[{"author":"karimf","children":[],"created_at":"2026-04-02T23:03:32.000Z","created_at_i":1775171012,"id":47621337,"options":[],"parent_id":47583809,"points":null,"story_id":47582482,"text":"Ok it&#x27;s on the app store now: <a href=\"https:&#x2F;&#x2F;apps.apple.com&#x2F;app&#x2F;volocal&#x2F;id6761493288\">https:&#x2F;&#x2F;apps.apple.com&#x2F;app&#x2F;volocal&#x2F;id6761493288</a>","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T07:14:40.000Z","created_at_i":1774941280,"id":47583809,"options":[],"parent_id":47583708,"points":null,"story_id":47582482,"text":"Oh thank you! I wasn\u2019t sure if it was worth submitting to the app store since it was just a research preview, but I could do it if people want it.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T07:00:16.000Z","created_at_i":1774940416,"id":47583708,"options":[],"parent_id":47583585,"points":null,"story_id":47582482,"text":"Brilliant. Hope to see you in the App Store!","title":null,"type":"comment","url":null},{"author":"podlp","children":[{"author":"karimf","children":[{"author":"Patrick_Devine","children":[],"created_at":"2026-03-31T22:31:41.000Z","created_at_i":1774996301,"id":47594338,"options":[],"parent_id":47593375,"points":null,"story_id":47582482,"text":"Try it with mxfp8 or bf16. It&#x27;s a decent model for doing tool calling, but I wouldn&#x27;t recommend using it with 4 bit quantization.","title":null,"type":"comment","url":null},{"author":"podlp","children":[],"created_at":"2026-04-01T14:13:30.000Z","created_at_i":1775052810,"id":47601191,"options":[],"parent_id":47593375,"points":null,"story_id":47582482,"text":"Subjectively, AFM isn\u2019t even close to Qwen. It\u2019s one of the weakest models I\u2019ve used. I\u2019m not even sure how many people have Apple Intelligence enabled. But I agree, there must be a huge onboarding win long-term using (and adapting) a model that\u2019s already optimized for your machine. I\u2019ve learned how to navigate most of its shortcomings, but it\u2019s not the most pleasant to work with.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T20:56:55.000Z","created_at_i":1774990615,"id":47593375,"options":[],"parent_id":47587735,"points":null,"story_id":47582482,"text":"&gt; Do you think it the models you\u2019re using could be quantized more that they could be downloaded on first run using Background Assets?<p>I first tried the Qwen 3.5 0.8B Q4_K_S and the model couldn&#x27;t hold a basic conversation. Although I haven&#x27;t tried lower quants on 2B.<p>I&#x27;m also interested on the Apple Foundation models, and it&#x27;s something I plan to try next. AFAIK it&#x27;s on par with Qwen-3-4B [0]. The biggest upside as you alluded to is that you don&#x27;t need to download it, which is huge for user onboarding.<p>[0] <a href=\"https:&#x2F;&#x2F;machinelearning.apple.com&#x2F;research&#x2F;apple-foundation-models-2025-updates\" rel=\"nofollow\">https:&#x2F;&#x2F;machinelearning.apple.com&#x2F;research&#x2F;apple-foundation-...</a>","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T14:15:26.000Z","created_at_i":1774966526,"id":47587735,"options":[],"parent_id":47583585,"points":null,"story_id":47582482,"text":"That\u2019s awesome! I\u2019ve got a similar project for macOS&#x2F; iOS using the Apple Intelligence models and on-device STT Transcriber APIs. Do you think it the models you\u2019re using could be quantized more that they could be downloaded on first run using Background Assets? Maybe we\u2019re not there yet, but I\u2019m interested in a better, local Siri like this with some sort of \u201cagentic lite\u201d capabilities.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T06:45:04.000Z","created_at_i":1774939504,"id":47583585,"options":[],"parent_id":47582826,"points":null,"story_id":47582482,"text":"Depending on the use case, the future is already here.<p>For example, last week I built a real-time voice AI running locally on iPhone 15.<p>One use case is for people learning speaking english. The STT is quite good and the small LLM is enough for basic conversation.<p><a href=\"https:&#x2F;&#x2F;github.com&#x2F;fikrikarim&#x2F;volocal\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;fikrikarim&#x2F;volocal</a>","title":null,"type":"comment","url":null},{"author":"troad","children":[{"author":"whackernews","children":[{"author":"irusensei","children":[{"author":"zozbot234","children":[],"created_at":"2026-03-31T08:47:04.000Z","created_at_i":1774946824,"id":47584439,"options":[],"parent_id":47584175,"points":null,"story_id":47582482,"text":"ANE-powered inference (at least for prefill, which is a key bottleneck on pre-M5 platforms) is also in the works, per <a href=\"https:&#x2F;&#x2F;github.com&#x2F;ggml-org&#x2F;llama.cpp&#x2F;issues&#x2F;10453#issuecomment-4148905254\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;ggml-org&#x2F;llama.cpp&#x2F;issues&#x2F;10453#issuecomm...</a>","title":null,"type":"comment","url":null},{"author":"OkGoDoIt","children":[{"author":"irusensei","children":[{"author":"drob518","children":[{"author":"irusensei","children":[],"created_at":"2026-03-31T23:02:39.000Z","created_at_i":1774998159,"id":47594626,"options":[],"parent_id":47586081,"points":null,"story_id":47582482,"text":"Yeah I&#x27;m terrible with analogies.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T11:59:47.000Z","created_at_i":1774958387,"id":47586081,"options":[],"parent_id":47585472,"points":null,"story_id":47582482,"text":"But you can always fall back to GGUF while waiting for the world to build a few more MLX restaurants. Or something like that; the analogy is a bit stretched.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T10:54:00.000Z","created_at_i":1774954440,"id":47585472,"options":[],"parent_id":47584520,"points":null,"story_id":47582482,"text":"Depends.<p>MLX is faster because it has better integration with Apple hardware. On the other hand GGUF is a far more popular format so there will be more programs and model variety.<p>So its kinda like having a very specific diet that you swear is better for you but you can only order food from a few restaurants.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T08:58:00.000Z","created_at_i":1774947480,"id":47584520,"options":[],"parent_id":47584175,"points":null,"story_id":47582482,"text":"Is that better or worse?","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T08:12:59.000Z","created_at_i":1774944779,"id":47584175,"options":[],"parent_id":47583845,"points":null,"story_id":47582482,"text":"&gt;Oh does llama.cpp use MLX or whatever?<p>No. It runs on MacOS but uses Metal instead of MLX.","title":null,"type":"comment","url":null},{"author":"LoganDark","children":[],"created_at":"2026-03-31T08:25:57.000Z","created_at_i":1774945557,"id":47584273,"options":[],"parent_id":47583845,"points":null,"story_id":47582482,"text":"llama.cpp uses GGML which uses Metal directly.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T07:20:58.000Z","created_at_i":1774941658,"id":47583845,"options":[],"parent_id":47583593,"points":null,"story_id":47582482,"text":"Oh does llama.cpp use MLX or whatever? I had this question, wonder if you know? A search suggests it doesn\u2019t but I don\u2019t really understand.","title":null,"type":"comment","url":null},{"author":"theshrike79","children":[{"author":"girvo","children":[{"author":"theshrike79","children":[{"author":"zozbot234","children":[{"author":"theshrike79","children":[],"created_at":"2026-03-31T10:18:53.000Z","created_at_i":1774952333,"id":47585171,"options":[],"parent_id":47585028,"points":null,"story_id":47582482,"text":"That&#x27;s the key, it just needs to be smart enough to 1) know it doesn&#x27;t know and 2) &quot;know a guy&quot; as they say =) (call a tool for the exact information)<p>Picking a model that&#x27;s juuust smart enough to know it doesn&#x27;t know is the key.","title":null,"type":"comment","url":null},{"author":"spockz","children":[{"author":"zozbot234","children":[],"created_at":"2026-03-31T10:55:48.000Z","created_at_i":1774954548,"id":47585494,"options":[],"parent_id":47585408,"points":null,"story_id":47582482,"text":"The training gives you a very lossy version of the original data (the smaller the model, the lossier it is; very small models will ultimately output gibberish and word salad that only loosely makes some sort of sense) but it&#x27;s the right format for generalization. So you actually want both, they&#x27;re highly complementary.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T10:47:37.000Z","created_at_i":1774954057,"id":47585408,"options":[],"parent_id":47585028,"points":null,"story_id":47582482,"text":"Ive always wondered where the inflection point lies between on the one hand trying to train the model on all kinds of data such as Wikipedia&#x2F;encyclopedia, versus in the system prompt pointing to your local versions of those data sources, perhaps even through a search like api&#x2F;tool.<p>Is there already some research or experimentation done into this area?","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T10:04:15.000Z","created_at_i":1774951455,"id":47585028,"options":[],"parent_id":47584947,"points":null,"story_id":47582482,"text":"&gt; Yep, having a &quot;stupid&quot; central model with multiple tools is IMO the key to efficient agentic systems.<p>That doesn&#x27;t fix the &quot;you don&#x27;t know what you don&#x27;t know&quot; problem which is huge with smaller models. A bigger model with more world knowledge really is a lot smarter in practice, though at a huge cost in efficiency.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T09:54:15.000Z","created_at_i":1774950855,"id":47584947,"options":[],"parent_id":47584800,"points":null,"story_id":47582482,"text":"Yep, having a &quot;stupid&quot; central model with multiple tools is IMO the key to efficient agentic systems.<p>It needs to be just smart enough to use the tools and distill the responses into something usable. And one of the tools can be &quot;ask claude&#x2F;codex&#x2F;gemini&quot; so the local model itself doesn&#x27;t actually need to do much.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T09:31:06.000Z","created_at_i":1774949466,"id":47584800,"options":[],"parent_id":47584392,"points":null,"story_id":47582482,"text":"I&#x27;d recommend it too, because the knowledge cutoff of all the open weight Chinese models (M2.7, Qwen3.5, GLM-5 etc) is earlier than you&#x27;d think, so giving it web search (I use `ddgr` with a skill) helps a surprising amount","title":null,"type":"comment","url":null},{"author":"troad","children":[{"author":"theshrike79","children":[{"author":"troad","children":[],"created_at":"2026-04-01T02:06:48.000Z","created_at_i":1775009208,"id":47595919,"options":[],"parent_id":47586465,"points":null,"story_id":47582482,"text":"Thanks!","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T12:38:01.000Z","created_at_i":1774960681,"id":47586465,"options":[],"parent_id":47585442,"points":null,"story_id":47582482,"text":"Basically ask any coding agent to create you a simple tool-calling harness for a local model and it&#x27;ll most likely one-shot it.<p>Getting the local weather using a free API like met.no is a good first tool to use.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T10:51:18.000Z","created_at_i":1774954278,"id":47585442,"options":[],"parent_id":47584392,"points":null,"story_id":47582482,"text":"That&#x27;s very cool! I think giving it some research tools might be a nifty thing to try next. This is a fairly new area for me, so pointers or suggestions are welcome, even basic ones. :)<p>Worth adding that I had reasoning on for the Tiananmen question, so I could see the prep for the answer, and it had a pretty strong current of &quot;This is a sensitive question to PRC authorities and I must not answer, or even hint at an answer&quot;. I&#x27;m not sure if a research tool would be sufficient to overcome that censorship, though I guess I&#x27;ll find out!","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T08:42:24.000Z","created_at_i":1774946544,"id":47584392,"options":[],"parent_id":47583593,"points":null,"story_id":47582482,"text":"Qwen3.5 has tool calling, so you can give it a wikipedia tool which it uses to know what happened in Tiananmen Square without issues =)","title":null,"type":"comment","url":null},{"author":"WesolyKubeczek","children":[{"author":"troad","children":[],"created_at":"2026-03-31T10:43:02.000Z","created_at_i":1774953782,"id":47585363,"options":[],"parent_id":47585208,"points":null,"story_id":47582482,"text":"Hey, if Margaret Thatcher&#x27;s son can give it a go, why not you? Believe in yourself and reach for those dreams. *sparkle emoji*","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T10:24:02.000Z","created_at_i":1774952642,"id":47585208,"options":[],"parent_id":47583593,"points":null,"story_id":47582482,"text":"Cool, I always wanted to invade Belgium. Maybe if my plan is good, I could run a successful gofundme?","title":null,"type":"comment","url":null},{"author":"austinthetaco","children":[{"author":"troad","children":[],"created_at":"2026-03-31T23:16:00.000Z","created_at_i":1774998960,"id":47594733,"options":[],"parent_id":47588428,"points":null,"story_id":47582482,"text":"Interesting! Unfortunately, the smallest Hermes 4 model I can see is 14B, which would really strain the limits of my little laptop. The only way I might get acceptable performance would be to run it extremely quantised, but then I probably wouldn&#x27;t see much improvement over the 9B Qwen.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T15:05:59.000Z","created_at_i":1774969559,"id":47588428,"options":[],"parent_id":47583593,"points":null,"story_id":47582482,"text":"Have you played around with any of the Hermes models? they are supposed to be one of the best at non-refusal while keeping sane.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T06:45:50.000Z","created_at_i":1774939550,"id":47583593,"options":[],"parent_id":47582826,"points":null,"story_id":47582482,"text":"I very recently installed llama.cpp on my consumer-grade M4 MBP, and I&#x27;ve been having loads of fun poking and prodding the local models. There&#x27;s now a ChatGPT style interface baked into llama.cpp, which is very handy for quick experimentation. (I&#x27;m not entirely sure what Ollama would get me that llama.cpp doesn&#x27;t, happy to hear suggestions!)<p>There are some surprisingly decent models that happily fit even into a mere 16 gigs of RAM. The recent Qwen 3.5 9B model is pretty good, though it did trip all over itself to avoid telling me what happened on Tiananmen Square in 1989. (But then I tried something called &quot;Qwen3.5-9B-Uncensored-HauhauCS-Aggressive&quot;, which veers so hard the other way that it will happily write up a detailed plan for your upcoming invasion of Belgium, so I guess it all balances out?)","title":null,"type":"comment","url":null},{"author":"overfeed","children":[{"author":"DrScientist","children":[],"created_at":"2026-03-31T10:42:04.000Z","created_at_i":1774953724,"id":47585353,"options":[],"parent_id":47583650,"points":null,"story_id":47582482,"text":"Apple via customers paying for the whole solution ( eg a laptop that can run decent local models )?<p>I think Apple had something in the region of 143 billion in revenue in the last quarter.<p>Not saying it will happen - just that there are a variety of business models out there and in the end it all depends on where consumers put their money.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T06:51:27.000Z","created_at_i":1774939887,"id":47583650,"options":[],"parent_id":47582826,"points":null,"story_id":47582482,"text":"&gt; It&#x27;s just a matter of getting the performance good enough.<p>Who will pay for the ongoing development of (near-)SoTA local models? The good open-weight models are all developed by for-profit companies - you know how that story will end.","title":null,"type":"comment","url":null},{"author":"nikanj","children":[],"created_at":"2026-03-31T06:58:16.000Z","created_at_i":1774940296,"id":47583698,"options":[],"parent_id":47582826,"points":null,"story_id":47582482,"text":"That also means sending every user a copy of the model that you spend billions training. The current model (running the models at the vendor side) makes it much easier to protect that investment","title":null,"type":"comment","url":null},{"author":"jl6","children":[{"author":"TeMPOraL","children":[{"author":"chongli","children":[],"created_at":"2026-03-31T11:53:53.000Z","created_at_i":1774958033,"id":47586028,"options":[],"parent_id":47583713,"points":null,"story_id":47582482,"text":"They do, though I don\u2019t think they max out on energy efficient technology. It\u2019s much easier to cut a deal for cheap electricity with a regional government, much to the chagrin of the locals (who see their power bills go up).","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T07:01:17.000Z","created_at_i":1774940477,"id":47583713,"options":[],"parent_id":47583701,"points":null,"story_id":47582482,"text":"Indeed. Data centers have so many ways and reasons to be much more energy-efficient than local compute it&#x27;s not even funny.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T06:59:15.000Z","created_at_i":1774940355,"id":47583701,"options":[],"parent_id":47582826,"points":null,"story_id":47582482,"text":"Not sure about the using less electricity part. With batching, it\u2019s more efficient to serve multiple users simultaneously.","title":null,"type":"comment","url":null},{"author":"ZeroGravitas","children":[{"author":"tomashubelbauer","children":[{"author":"blitzar","children":[],"created_at":"2026-03-31T08:44:45.000Z","created_at_i":1774946685,"id":47584411,"options":[],"parent_id":47584004,"points":null,"story_id":47582482,"text":"&quot;copilot&quot; seems a good term<p>could also be considered a <i>triage</i> layer","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T07:46:16.000Z","created_at_i":1774943176,"id":47584004,"options":[],"parent_id":47583789,"points":null,"story_id":47582482,"text":"I&#x27;d like to coin the term &quot;user agent&quot; for this","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T07:11:59.000Z","created_at_i":1774941119,"id":47583789,"options":[],"parent_id":47582826,"points":null,"story_id":47582482,"text":"It feels like you&#x27;ll soon need a local llm to intermediate with the remote llm, like an ad blocker for browsers to stop them injecting ads or remind you not to send corporate IP out onto the Internet.","title":null,"type":"comment","url":null},{"author":"miki123211","children":[{"author":"kortilla","children":[],"created_at":"2026-03-31T07:41:04.000Z","created_at_i":1774942864,"id":47583966,"options":[],"parent_id":47583910,"points":null,"story_id":47582482,"text":"Well this is an article about running on hardware I already have in my house. In the winter that\u2019s just a little extra electricity that converts into \u201cfree\u201d resistive heating.","title":null,"type":"comment","url":null},{"author":"ysleepy","children":[],"created_at":"2026-03-31T08:16:47.000Z","created_at_i":1774945007,"id":47584198,"options":[],"parent_id":47583910,"points":null,"story_id":47582482,"text":"I&#x27;m actually not sure that&#x27;s true. Apart from people buying the device with or without the neural accelerator, the perf&#x2F;watt could be on par or better with the big iron. The efficiency sweet-spot is usually below the peak performance point, see big.little architectures etc.","title":null,"type":"comment","url":null},{"author":"zozbot234","children":[{"author":"Tepix","children":[{"author":"zozbot234","children":[],"created_at":"2026-03-31T09:07:51.000Z","created_at_i":1774948071,"id":47584606,"options":[],"parent_id":47584538,"points":null,"story_id":47582482,"text":"&gt; your PC or phone at home is mostly idle<p>If you&#x27;re purely repurposing hardware that you need anyway for other uses, that doesn&#x27;t really matter.<p>(Besides, for that matter, your utilization might actually rise if you&#x27;re making do with potato-class hardware that can only achieve low throughput and high latency. You&#x27;d be running inference in the background, basically at all times.)","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T09:00:32.000Z","created_at_i":1774947632,"id":47584538,"options":[],"parent_id":47584394,"points":null,"story_id":47582482,"text":"Seems doubtful. The utilisation will be super high for data center silicon whereas your PC or phone at home is mostly idle.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T08:42:38.000Z","created_at_i":1774946558,"id":47584394,"options":[],"parent_id":47583910,"points":null,"story_id":47582482,"text":"&gt; LLMs are far more efficient on hardware that simultaneously serves many requests at once.<p>The LLM inference itself may be more efficient (though this may be impacted by different throughput vs. latency tradeoffs; local inference makes it easier to run with higher latency) but making the hardware is not.  The cost for datacenter-class hardware is orders of magnitude higher, and repurposing existing hardware is a real gain in efficiency.","title":null,"type":"comment","url":null},{"author":"woadwarrior01","children":[],"created_at":"2026-03-31T11:19:00.000Z","created_at_i":1774955940,"id":47585721,"options":[],"parent_id":47583910,"points":null,"story_id":47582482,"text":"&gt; Sorry to shatter your bubble, but this is patently false, LLMs are far more efficient on hardware that simultaneously serves many requests at once.<p>You might want to read this: <a href=\"https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2502.05317v2\" rel=\"nofollow\">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2502.05317v2</a>","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T07:30:33.000Z","created_at_i":1774942233,"id":47583910,"options":[],"parent_id":47582826,"points":null,"story_id":47582482,"text":"&gt; would use less electricity<p>Sorry to shatter your bubble, but this is patently false, LLMs are far more efficient on hardware that simultaneously serves many requests at once.<p>There&#x27;s also the (environmental and monetary) cost of producing overpowered devices that sit idle when you&#x27;re not using them, in contrast to a cloud GPU, which can be rented out to whoever needs it at a given moment, potentially at a lower cost during periods of lower demand.<p>Many LLM workloads aren&#x27;t even that latency sensitive, so it&#x27;s far easier to move them closer to renewable energy than to move that energy closer to you.","title":null,"type":"comment","url":null},{"author":"thih9","children":[{"author":"jychang","children":[],"created_at":"2026-03-31T07:41:25.000Z","created_at_i":1774942885,"id":47583967,"options":[],"parent_id":47583942,"points":null,"story_id":47582482,"text":"That&#x27;s completely not true. LLM on device would use MORE electricity.<p>Service providers that do batch&gt;1 inference are a lot more efficient per watt.<p>Local inference can only do batch=1 inference, which is very inefficient.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T07:37:16.000Z","created_at_i":1774942636,"id":47583942,"options":[],"parent_id":47582826,"points":null,"story_id":47582482,"text":"&gt; it also would use less electricity<p>How would it use less electricity? I\u2019d like to learn more.","title":null,"type":"comment","url":null},{"author":"amelius","children":[{"author":"theshrike79","children":[{"author":"angoragoats","children":[{"author":"amelius","children":[{"author":"angoragoats","children":[],"created_at":"2026-03-31T23:55:10.000Z","created_at_i":1775001310,"id":47595043,"options":[],"parent_id":47586563,"points":null,"story_id":47582482,"text":"Sure, but that\u2019s somewhat orthogonal to the point I was making, which is that LLMs are huge in size. Even in the case of a custom \u201cLLM chip,\u201d you\u2019ll need huge amounts of very fast storage of some sort (likely DRAM), which places constraints on the size, power consumption, and cost of such a device. This device, if it existed, would not in any way resemble the Coral TPU product that the GP was referencing; I think in fact it would be closer in size, price, and form factor to a GPU.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T12:45:36.000Z","created_at_i":1774961136,"id":47586563,"options":[],"parent_id":47585969,"points":null,"story_id":47582482,"text":"GPUs are still software programmable.<p>An &quot;LLM chip&quot; does not need that and so can be much more efficient.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T11:48:35.000Z","created_at_i":1774957715,"id":47585969,"options":[],"parent_id":47584408,"points":null,"story_id":47582482,"text":"I think there are drastic differences between computer vision models and LLMs that you\u2019re not considering. LLMs are <i>huge</i> relative to vision models, and require gobs of fast memory. For this reason a little USB dongle isn\u2019t going to cut it.<p>Put another way, there already exist add-in boards like this, and they\u2019re called GPUs.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T08:44:07.000Z","created_at_i":1774946647,"id":47584408,"options":[],"parent_id":47584289,"points":null,"story_id":47582482,"text":"I&#x27;m expecting someone to come up with an LLM version of the Coral USB Accelerator: <a href=\"https:&#x2F;&#x2F;www.coral.ai&#x2F;products&#x2F;accelerator\" rel=\"nofollow\">https:&#x2F;&#x2F;www.coral.ai&#x2F;products&#x2F;accelerator</a><p>Just plug in a stick in your USB-C port or add an M.2 or PCIe board and you&#x27;ll get dramatically faster AI inference.","title":null,"type":"comment","url":null},{"author":"jillesvangurp","children":[{"author":"zozbot234","children":[{"author":"jillesvangurp","children":[],"created_at":"2026-04-01T04:16:32.000Z","created_at_i":1775016992,"id":47596757,"options":[],"parent_id":47585227,"points":null,"story_id":47582482,"text":"Standard pig cycle in economics. Production capacity eventually goes up to meet demand and prices come down again. RAM has been going through cycles like this for decades. People seem to have no memory whatsoever of previous cycles every time it happens. Just wait a few years for it to become cheap again.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T10:26:25.000Z","created_at_i":1774952785,"id":47585227,"options":[],"parent_id":47585181,"points":null,"story_id":47582482,"text":"&gt; This will drive an overdue increase in memory size of phones and laptops.<p>DRAM costs are still skyrocketing, so no, I don&#x27;t think so. It&#x27;s more likely that we&#x27;ll bring back wear-resistant persistent memory as formerly seen with Intel Optane.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T10:20:10.000Z","created_at_i":1774952410,"id":47585181,"options":[],"parent_id":47584289,"points":null,"story_id":47582482,"text":"You can always delegate sub agents to cloud based infrastructure for things that need more intelligence. But the future indeed is to keep the core interaction loop on the local device always ready for your input.<p>A lot of stuff that we ask of these models isn&#x27;t all that hard. Summarize this, parse that, call this tool, look that up, etc. 99.999% really isn&#x27;t about implementing complex algorithms, solving important math problems, working your way through a benchmark of leet programming exercises, etc. You also really don&#x27;t need these models to know everything. It&#x27;s nice if it can hallucinate a decent answer to most questions. But the smarter way is to look up the right answer and then summarize it. Good enough goes a long way. Speed and latency are becoming a key selling point. You need enough capability locally to know when to escalate to something slower and more costly.<p>This will drive an overdue increase in memory size of phones and laptops. Laptops especially have been stuck at the same common base level of 8-16GB for about 15 years now. Apple still sells laptops with just 8GB (their new Neo). I had a 16 GB mac book pro in 2012. At the time that wasn&#x27;t even that special. My current one has 48GB; enough for some of the nicer models. You can get as much as 256GB today.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T08:27:54.000Z","created_at_i":1774945674,"id":47584289,"options":[],"parent_id":47582826,"points":null,"story_id":47582482,"text":"LLM in silicon is the future. It won&#x27;t be long until you can just plug an LLM chip into your computer and talk to it at 100x the speed of current LLMs. Capability will be lower but their speed will make up for it.","title":null,"type":"comment","url":null},{"author":"zozbot234","children":[],"created_at":"2026-03-31T09:00:28.000Z","created_at_i":1774947628,"id":47584537,"options":[],"parent_id":47582826,"points":null,"story_id":47582482,"text":"&gt; Most users don&#x27;t need frontier model performance.<p>SSD weights offload makes it feasible to run SOTA local models on consumer or prosumer&#x2F;enthusiast-class platforms, though with very low throughput (the SSD offload bandwidth is a huge bottleneck, mitigated by having a lot of RAM for caching).  But if you only need SOTA performance rarely and can wait for the answer, it becomes a great option.","title":null,"type":"comment","url":null},{"author":"iNic","children":[{"author":"niek_pas","children":[],"created_at":"2026-03-31T09:38:35.000Z","created_at_i":1774949915,"id":47584840,"options":[],"parent_id":47584561,"points":null,"story_id":47582482,"text":"&gt; For many businesses it will still make sense to have more powerful models and to run them centralized in a datacenter.<p>Agree, and I think of it this way: for a lot of businesses, it already makes sense to have a bunch of more powerful computers and run them centralized in a datacenter. Nevertheless, most people at most companies do most of their work on their Macbook Air or Dell whatever. I think LLMs will follow a similar pattern: local for 90% of use cases, powerful models (either on-site in a datacenter or via a service) for everything else.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T09:03:26.000Z","created_at_i":1774947806,"id":47584561,"options":[],"parent_id":47582826,"points":null,"story_id":47582482,"text":"It will probably be a future. My guess is that for many businesses it will still make sense to have more powerful models and to run them centralized in a datacenter. Also, by batching queries you can get efficiencies at scale that might be hard to replicate locally. I can also see a hybrid approach where local models get good at handing off to cloud models for complex queries.","title":null,"type":"comment","url":null},{"author":"goldenarm","children":[],"created_at":"2026-03-31T09:06:26.000Z","created_at_i":1774947986,"id":47584588,"options":[],"parent_id":47582826,"points":null,"story_id":47582482,"text":"It&#x27;s more secure, but it would make supply much much worse.<p>Data centers use GPU batching, much higher utilisation rates, and more efficient hardware. It&#x27;s borderline two order of magnitude more efficient than your desktop.","title":null,"type":"comment","url":null},{"author":"nbenitezl","children":[{"author":"comboy","children":[],"created_at":"2026-03-31T10:03:14.000Z","created_at_i":1774951394,"id":47585017,"options":[],"parent_id":47584931,"points":null,"story_id":47582482,"text":"As things stand today even when doing research tasks, time spent by model is &gt;&gt; than fetching websites. I don&#x27;t see it changing any time soon, except when some deals happen behind the scenes where agents get to access CF guarded resources that normally get blocked from automated access.","title":null,"type":"comment","url":null},{"author":"Const-me","children":[],"created_at":"2026-03-31T10:03:24.000Z","created_at_i":1774951404,"id":47585020,"options":[],"parent_id":47584931,"points":null,"story_id":47582482,"text":"While data centres indeed have awesome internet connectivity, don\u2019t forget the bandwidth is shared by all clients using a particular server.<p>If you have 100 mbit&#x2F;sec internet connection at home, a computer in a data centre has 10 gbit&#x2F;sec, but the server is serving 200 concurrent clients \u2014 your bandwidth is twice as fast.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T09:51:50.000Z","created_at_i":1774950710,"id":47584931,"options":[],"parent_id":47582826,"points":null,"story_id":47582482,"text":"But when using it on the cloud a LLM can consult 50 websites, which is super fast for their datacenters as they are backbone of internet, instead you&#x27;ll have to wait much more on your device to consult those websites before giving you the LLM response. Am i wrong?","title":null,"type":"comment","url":null},{"author":"dwayne_dibley","children":[],"created_at":"2026-03-31T10:29:34.000Z","created_at_i":1774952974,"id":47585250,"options":[],"parent_id":47582826,"points":null,"story_id":47582482,"text":"This might be how Apple will start to see even more sales, the M series processors are so far ahead of anything else, local LLMs could be their main selling point.","title":null,"type":"comment","url":null},{"author":"konschubert","children":[{"author":"ekianjo","children":[{"author":"konschubert","children":[],"created_at":"2026-03-31T14:12:19.000Z","created_at_i":1774966339,"id":47587682,"options":[],"parent_id":47585568,"points":null,"story_id":47582482,"text":"While not everybody is a professional in YOUR domain, many people are professionals in SOME domain. And even outside of that, they deserve a smart conversation partner, for example on topics like health and politics.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T11:03:10.000Z","created_at_i":1774954990,"id":47585568,"options":[],"parent_id":47585352,"points":null,"story_id":47582482,"text":"&gt; What makes you think that?<p>Looking at actual users of LLMs","title":null,"type":"comment","url":null},{"author":"locknitpicker","children":[{"author":"konschubert","children":[],"created_at":"2026-03-31T14:14:05.000Z","created_at_i":1774966445,"id":47587724,"options":[],"parent_id":47585637,"points":null,"story_id":47582482,"text":"Everybody has difficult decisions to make in their daily lives and in their work.<p>Having access to a model that is drawing from good sources and takes time to think instead of hallucinating a response is important in many domains of life.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T11:11:49.000Z","created_at_i":1774955509,"id":47585637,"options":[],"parent_id":47585352,"points":null,"story_id":47582482,"text":"&gt; What makes you think that?<p>The fact that today&#x27;s and yesterday&#x27;s models are quite capable of handling mundane tasks, and even companies behind frontier models are investing heavily in strategies to manage context instead of blindly plowing through problems with brute-force generalist models.<p>But let&#x27;s flip this around: what on earth even suggests to you that most users need frontier models?","title":null,"type":"comment","url":null},{"author":"dgb23","children":[{"author":"pama","children":[{"author":"dudefeliciano","children":[{"author":"Shorel","children":[],"created_at":"2026-03-31T17:02:48.000Z","created_at_i":1774976568,"id":47590328,"options":[],"parent_id":47587471,"points":null,"story_id":47582482,"text":"I&#x27;m sorry to get into this conversation, but the performance of a model is some orders of magnitude lower (meaning it requires greater amounts of specific computing power) than all the network stack of all the nodes involved in the internet traffic of some particular request.<p>Meaning: these 5000 tokens consume tiny amounts of energy being moved all around from the data center to your PC, but enormous amounts of energy being generated at all.  An equivalent webpage with the same amount of text as these tokens would be perceived as instant in any network configuration. Just some kilobytes of text. Much smaller than most background graphics. The two things can&#x27;t be compared at all.<p>However, just last week there have been huge improvements on the hardware required to run some particular models, thanks to some very clever quantisation. This lowers the memory required 6x in our home hardware, which is great.<p>In the end, we spent more energy playing videogames during the last two decades, than all this AI craze, and it was never a problem. We surely can run models locally, and heat our homes in winter.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T13:56:43.000Z","created_at_i":1774965403,"id":47587471,"options":[],"parent_id":47586190,"points":null,"story_id":47582482,"text":"Aren&#x27;t data centers extremely energy inneficient due to network latency, memory bottlenecks and so on? I mean the models that run on them are extremely powerful compared to what you can run on consumer hardware, but I wouldn&#x27;t call them efficient...","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T12:13:10.000Z","created_at_i":1774959190,"id":47586190,"options":[],"parent_id":47585769,"points":null,"story_id":47582482,"text":"Parallel inference on large compute scales in superlinear ways. There is no way to beat the reduction in memory transfers that a data-center inference model provides with hardware that fits at anything called a home.  It is much more energy efficient to process huge batches of parallel requests compared to having one or a handful of queries running on an accelerator.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T11:25:10.000Z","created_at_i":1774956310,"id":47585769,"options":[],"parent_id":47585352,"points":null,"story_id":47582482,"text":"&gt; False, it creates consumer demand for inference chips, which will be badly utilised.<p>I think the opposite is true. Local inference doesn&#x27;t have to go over the wire and through a bunch of firewalls and what have you. The performance from just regular consumer hardware with local, smaller models is already decent. You&#x27;re utilizing the hardware you already have.<p>&gt; The performance limitations are inherent to the limited compute and memory.<p>When you plug in a local LLM and inference engine into an agent that is built around the assumption of using a cloud&#x2F;frontier model then that&#x27;s true.<p>But agents can be built around local assumptions and more specific workflows and problems. That also includes the model orchestration and model choice per task (or even tool).<p>The Jevons Paradox comes into play with using cloud models. But when you have less resources you are forced to move into more deterministic workflows. That includes tighter control over what the agent can do at any point in time, but also per project&#x2F;session workflows where you generate intermediate programs&#x2F;scripts instead of letting the agent just do what ever it wants.<p>I give you an example:<p>When you ask a cloud based agent to do something and it wants more information, it will often do a series of tool calls to gather what it thinks it needs before proceeding. Very often you can front load that part, by first writing a testable program that gathers most of the necessary information up front and only then moving into an agentic workflow.<p>This approach can produce a bunch of .json, .md files or it can move things into a structured database or you can use embeddings or what have you.<p>This can save you a lot of inference, make things more reusable and you don&#x27;t need a model that is as capable if its context is already available and tailored to a specific task.","title":null,"type":"comment","url":null},{"author":"txdv","children":[{"author":"iknowstuff","children":[],"created_at":"2026-03-31T18:44:34.000Z","created_at_i":1774982674,"id":47591703,"options":[],"parent_id":47591596,"points":null,"story_id":47582482,"text":"Thats the point, they\u2019re better utilized in the cloud","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T18:35:52.000Z","created_at_i":1774982152,"id":47591596,"options":[],"parent_id":47585352,"points":null,"story_id":47582482,"text":"&gt; False, it creates consumer demand for inference chips, which will be badly utilised.<p>There are so many CPUs, GPUs, RAM and SSDs which are underutilized. I have some in my closet doing 5% load at peek times. Why would inference chips be special  once they become commodity hardware?","title":null,"type":"comment","url":null},{"author":"nsonha","children":[],"created_at":"2026-04-01T03:32:56.000Z","created_at_i":1775014376,"id":47596492,"options":[],"parent_id":47585352,"points":null,"story_id":47582482,"text":"&quot;consumer demand for inference chips, which will be badly utilised&quot;<p>why do you assume it will be badly utilised? Can&#x27;t be worse than what we have now which is chips already badly utilised for windows&#x27; bloatware","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T10:41:57.000Z","created_at_i":1774953717,"id":47585352,"options":[],"parent_id":47582826,"points":null,"story_id":47582482,"text":"I disagree with every sentence of this.<p>&gt; solves the problem of too much demand for inference<p>False, it creates consumer demand for inference chips, which will be badly utilised.<p>&gt;  also would use less electricity<p>What makes you think that? (MAYBE you can save power on cooling. But not if the data center is close to a natural heat sink)<p>&gt; It&#x27;s just a matter of getting the performance good enough.<p>The performance limitations are inherent to the limited compute and memory.<p>&gt; Most users don&#x27;t need frontier model performance.<p>What makes you think that?","title":null,"type":"comment","url":null},{"author":"g947o","children":[{"author":"RALaBarge","children":[],"created_at":"2026-03-31T12:45:27.000Z","created_at_i":1774961127,"id":47586560,"options":[],"parent_id":47586433,"points":null,"story_id":47582482,"text":"I agree with you in the sense that if you tried to take any model right now and cram it into an iphone, it wouldnt be a claude-level agent.<p>I run 32b agents locally on a big video card, and smaller ones in CPU, but the lack there isn&#x27;t the logic or reasoning, it is the chain of tooling that Claude Code and other stacks have built in.<p>Doing a lot of testing recently with my own harness, you would not believe the quality improvement you can get from a smaller LLM with really good opening context.<p>Even Microsoft is working on 1-bit LLMs...it sucks right now, but what about in 5 years?<p>But the OP is correct -- everything will have an LLM on it eventually, much sooner than people who do not understand what is going on right now would ever believe is possible.","title":null,"type":"comment","url":null},{"author":"kylehotchkiss","children":[{"author":"g947o","children":[],"created_at":"2026-04-01T01:19:57.000Z","created_at_i":1775006397,"id":47595625,"options":[],"parent_id":47590770,"points":null,"story_id":47582482,"text":"You probably want to double check the comment I was responding to.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T17:31:30.000Z","created_at_i":1774978290,"id":47590770,"options":[],"parent_id":47586433,"points":null,"story_id":47582482,"text":"Yes. I&#x27;ve spent months running Qwen2.5-8B on my barebones 16gb ram M4 Mac mini to handle identifying sites from google search results. It has been rock solid. I&#x27;m not even running this MLX-powered improvement on it yet.<p>Your idea of what people need from Local LLMs and others are different. Not everybody needs a &#x2F;r&#x2F;myboyfriendisai level performance.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T12:35:02.000Z","created_at_i":1774960502,"id":47586433,"options":[],"parent_id":47582826,"points":null,"story_id":47582482,"text":"Have you spent more than 10 min actually running LLM on a local machine?<p>As it stands today, local LLMs don&#x27;t work remotely as well as some people try to picture them, in almost every way -- speed, performance, cost, usability etc. The only upside is privacy.","title":null,"type":"comment","url":null},{"author":"eeixlk","children":[],"created_at":"2026-03-31T13:14:25.000Z","created_at_i":1774962865,"id":47586897,"options":[],"parent_id":47582826,"points":null,"story_id":47582482,"text":"Obviously apple would prefer this.  It would boost demand for more powerful and expensive devices, and align with their privacy marketing.  But they have massively fumbled with siri for a long time and then missed huge deadlines with ai promises.  Despite having billions, they have shown no competency in delivering services or accurately marketing what to expect from ai features.","title":null,"type":"comment","url":null},{"author":"jonhohle","children":[{"author":"theChaparral","children":[],"created_at":"2026-03-31T15:16:12.000Z","created_at_i":1774970172,"id":47588591,"options":[],"parent_id":47587440,"points":null,"story_id":47582482,"text":"It&#x27;s Personal Intelligence in the Gemini settings. I just turned that off last night when it was doing similar things.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T13:54:36.000Z","created_at_i":1774965276,"id":47587440,"options":[],"parent_id":47582826,"points":null,"story_id":47582482,"text":"I\u2019ve been using google search AI and Gemini, which I find generally pretty good. In the past week, Gemini and Search AI have been bringing in various details of previous searches I\u2019ve done and Search AI conversations I\u2019ve had and it\u2019s extremely gross and creepy.<p>I was looking for details about cars and it started interjecting how the safety would affect my children by name in a conversation where I never mention my children. I was asking details about Thunderbolt and modern Ryzen processors and a fresh Gemini chat brought in details about a completely unrelated project I work on. I\u2019ve always thought local LLMs would be important, but whatever Google did in the past few weeks has made that even more clear.","title":null,"type":"comment","url":null},{"author":"Aurornis","children":[],"created_at":"2026-03-31T14:19:25.000Z","created_at_i":1774966765,"id":47587779,"options":[],"parent_id":47582826,"points":null,"story_id":47582482,"text":"&gt; solves the problem of too much demand for inference compared to data center supply<p>Maybe in the distant future when device compute capacity has increased by multiples and efficiency improvements have made smaller LLMs better.<p>The current data center buildouts are using GPU clusters and hybrid compute servers that are so much more powerful than anything you can run at home that they\u2019re not in the same league. Even among the open models that you can run at home if you\u2019re willing to spend $40K on hardware, the prefill and token generation speeds are so slow compared to SOTA served models that you really have to be dedicated to avoiding the cloud to run these.<p>We won\u2019t be in a data center crunch forever. I would not be surprised if we have a period of data center oversupply after this rush to build out capacity.<p>However at the current rate of progress I don\u2019t see local compute catching up to hosted models in quality and usability (speed) before data center capacity catches up to demand. This is coming from someone who spends more than is reasonable  on local compute hardware.","title":null,"type":"comment","url":null},{"author":"babblingfish","children":[],"created_at":"2026-03-31T17:11:19.000Z","created_at_i":1774977079,"id":47590460,"options":[],"parent_id":47582826,"points":null,"story_id":47582482,"text":"I see a lot of people are confused about the electricity claim so I&#x27;ll elaborate on it more. The assumption I&#x27;m making here is that on device people will run smaller models, that can fit on their machines without needing to buy new computers. If everyone ran inference on their machine there would be no need for these massive datacenters which use huge quantities of electricity. It would utilize the machines they already have and the electricity they&#x27;re already using.<p>People are making a comparison of the cost per inference or token or whatever and saying datacenters are more efficient which makes obvious sense. What i&#x27;m saying is if we eliminate the need for building out dozens of gigawatt datacenters completely then we would use less electricity. I feel like this makes intuitive sense. People are getting lost in the details about cost per inference, and performance on different models.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T04:40:12.000Z","created_at_i":1774932012,"id":47582826,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"LLMs on device is the future. It&#x27;s more secure and solves the problem of too much demand for inference compared to data center supply, it also would use less electricity. It&#x27;s just a matter of getting the performance good enough. Most users don&#x27;t need frontier model performance.","title":null,"type":"comment","url":null},{"author":"codelion","children":[],"created_at":"2026-03-31T04:50:16.000Z","created_at_i":1774932616,"id":47582875,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"How does it compare to some of the newer mlx inference engines like optiq that support turboquantization - <a href=\"https:&#x2F;&#x2F;mlx-optiq.pages.dev&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;mlx-optiq.pages.dev&#x2F;</a>","title":null,"type":"comment","url":null},{"author":"dial9-1","children":[{"author":"gedy","children":[{"author":"HDBaseT","children":[{"author":"Foobar8568","children":[{"author":"brcmthrowaway","children":[],"created_at":"2026-03-31T05:28:57.000Z","created_at_i":1774934937,"id":47583093,"options":[],"parent_id":47583067,"points":null,"story_id":47582482,"text":"Just train it better with AGENTS.md","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T05:25:42.000Z","created_at_i":1774934742,"id":47583067,"options":[],"parent_id":47582972,"points":null,"story_id":47582482,"text":"I fully agree, I run that one with Q4 on my MBP, and the performance (including quality of response) is a let down.<p>I am wondering how people rave so much about local &quot;small devices&quot; LLM vs what codex or Claude code are capable of.<p>Sadly there are too much hype on local LLM, they look great for 5min tests and that&#x27;s it.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T05:09:53.000Z","created_at_i":1774933793,"id":47582972,"options":[],"parent_id":47582897,"points":null,"story_id":47582482,"text":"You can run Qwen3.5-35B-A3B on 32GB of RAM sure, although to get &#x27;Claude Code&#x27; performance, which I assume he means Sonnet or Opus level models in 2026, this will likely be a few years away before its runnable locally (with reasonable hardware).","title":null,"type":"comment","url":null},{"author":"Hamuko","children":[],"created_at":"2026-03-31T10:48:24.000Z","created_at_i":1774954104,"id":47585416,"options":[],"parent_id":47582897,"points":null,"story_id":47582482,"text":"I&#x27;m reading &quot;more than 32GB of unified memory&quot; to mean at least a 36 GB model.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T04:54:43.000Z","created_at_i":1774932883,"id":47582897,"options":[],"parent_id":47582878,"points":null,"story_id":47582482,"text":"How close is this?  It says it needs 32GB min?","title":null,"type":"comment","url":null},{"author":"rubymamis","children":[{"author":"g947o","children":[],"created_at":"2026-03-31T12:40:15.000Z","created_at_i":1774960815,"id":47586489,"options":[],"parent_id":47585091,"points":null,"story_id":47582482,"text":"You can, but the quality sucks.<p>Local LLMs don&#x27;t make sense for most people compared to &quot;cloud&quot; services, even more so for coding.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T10:10:25.000Z","created_at_i":1774951825,"id":47585091,"options":[],"parent_id":47582878,"points":null,"story_id":47582482,"text":"Doesn&#x27;t OpenCode supports local models?","title":null,"type":"comment","url":null},{"author":"bearjaws","children":[],"created_at":"2026-03-31T12:43:30.000Z","created_at_i":1774961010,"id":47586523,"options":[],"parent_id":47582878,"points":null,"story_id":47582482,"text":"My super uninformed theory is that local LLM will trail foundation models by about 2 years for practical use.<p>For example right now a lot of work is being done on improving tool calling and agentic workflows, which tool calling was first popping up around end of 2023 for local LLMs.<p>This is putting aside the standard benchmarks which get &quot;benchmaxxed&quot; by local LLMs and show impressive numbers, but when used with OpenCode rarely meet expectations. In theory Qwen3.5-397B-A17B should be nearly a Sonnet 4.6 model but it is not.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T04:51:12.000Z","created_at_i":1774932672,"id":47582878,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"still waiting for the day I can comfortably run Claude Code with local llm&#x27;s on MacOS with only 16gb of ram","title":null,"type":"comment","url":null},{"author":"LuxBennu","children":[{"author":"zozbot234","children":[],"created_at":"2026-03-31T08:37:42.000Z","created_at_i":1774946262,"id":47584360,"options":[],"parent_id":47582925,"points":null,"story_id":47582482,"text":"They initially messed up this launch and overwrote some of the GGUF models in their library, making them non-downloadable on platforms other than Apple Silicon. Hopefully that gets fixed.","title":null,"type":"comment","url":null},{"author":"goldenarm","children":[{"author":"LuxBennu","children":[],"created_at":"2026-04-01T07:21:30.000Z","created_at_i":1775028090,"id":47597902,"options":[],"parent_id":47584599,"points":null,"story_id":47582482,"text":"Roughly 8-12 token&#x2F;s on generation depending on context length. Prompt processing is faster obviously. Haven&#x27;t benchmarked it super carefully though, just eyeballing the llama.cpp output.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T09:07:38.000Z","created_at_i":1774948058,"id":47584599,"options":[],"parent_id":47582925,"points":null,"story_id":47582482,"text":"How many tokens per second?","title":null,"type":"comment","url":null},{"author":"yg1112","children":[{"author":"lioeters","children":[],"created_at":"2026-03-31T20:38:36.000Z","created_at_i":1774989516,"id":47593169,"options":[],"parent_id":47591055,"points":null,"story_id":47582482,"text":"Insightful comment, thanks!","title":null,"type":"comment","url":null},{"author":"LuxBennu","children":[{"author":"zozbot234","children":[],"created_at":"2026-04-01T08:06:37.000Z","created_at_i":1775030797,"id":47598169,"options":[],"parent_id":47597922,"points":null,"story_id":47582482,"text":"If it&#x27;s just about skipping some buffer sync that&#x27;s something that could also be adopted by llama.cpp&#x27;s own Metal backend, at least on Apple Silicon platforms.","title":null,"type":"comment","url":null}],"created_at":"2026-04-01T07:25:48.000Z","created_at_i":1775028348,"id":47597922,"options":[],"parent_id":47591055,"points":null,"story_id":47582482,"text":"that tracks with what i&#x27;ve noticed practically. shorter prompts feel basically the same between llama.cpp metal and what i&#x27;d expect from native mlx, but once context gets longer the overhead starts showing up. would be interesting to see if ollama&#x27;s mlx path actually handles kv cache differently under the hood or if it just skips the buffer sync layer","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T17:50:44.000Z","created_at_i":1774979444,"id":47591055,"options":[],"parent_id":47582925,"points":null,"story_id":47582482,"text":"The key difference is that MLX&#x27;s array model assumes unified memory from the ground up. llama.cpp&#x27;s Metal backend works fine but carries abstractions from the discrete GPU world \u2014 explicit buffer synchronization, command buffer boundaries \u2014 that are unnecessary when CPU and GPU share the same address space. You&#x27;ll notice the gap most at large context lengths where KV cache pressure is highest.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T05:01:03.000Z","created_at_i":1774933263,"id":47582925,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"Already running qwen 70b 4-bit on m2 max 96gb through llama.cpp and it&#x27;s pretty solid for day to day stuff. The mlx switch is interesting because ollama was basically shelling out to llama.cpp on mac before, so native mlx should mean better memory handling on apple silicon. Curious to see how it compares on the bigger models vs the gguf path","title":null,"type":"comment","url":null},{"author":"AugSun","children":[],"created_at":"2026-03-31T05:03:38.000Z","created_at_i":1774933418,"id":47582939,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"&quot;We can run your dumbed down models faster&quot;:<p>#The use of NVFP4 results in a 3.5x reduction in model memory footprint relative to FP16 and a 1.8x reduction compared to FP8, while maintaining model accuracy with less than 1% degradation on <i>key language modeling</i> tasks for <i>some</i> models.","title":null,"type":"comment","url":null},{"author":"brcmthrowaway","children":[{"author":"xiconfjs","children":[{"author":"yard2010","children":[{"author":"xiconfjs","children":[{"author":"brcmthrowaway","children":[{"author":"xiconfjs","children":[],"created_at":"2026-03-31T16:15:36.000Z","created_at_i":1774973736,"id":47589629,"options":[],"parent_id":47588966,"points":null,"story_id":47582482,"text":"I tried man times but at least with its API active, LMStudio has some kind of memory leaks which will slow down the whole system (after ~1-2 days of uptime) even after unloading the model and stopping LMStudio up to a point where even playing a 1080p video results in frame drops. No such issues with Ollama.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T15:37:16.000Z","created_at_i":1774971436,"id":47588966,"options":[],"parent_id":47587671,"points":null,"story_id":47582482,"text":"Seems complicated. Switch to LMStudio","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T14:11:41.000Z","created_at_i":1774966301,"id":47587671,"options":[],"parent_id":47584010,"points":null,"story_id":47582482,"text":"* macOS 26.x on MacBookPro M1 Max 32GB\n* Ollama on macOS, cursor to play around\n* Open WebUI [1] on my Homeserver via API to Ollama (also for remote \u201eA.I.\u201c access)\n* running gpt-oss:20b, qwen3.5:9b with ease, qwen3.5:27b for more complex tasks<p>[1] <a href=\"https:&#x2F;&#x2F;github.com&#x2F;open-webui&#x2F;open-webui\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;open-webui&#x2F;open-webui</a>","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T07:47:52.000Z","created_at_i":1774943272,"id":47584010,"options":[],"parent_id":47583114,"points":null,"story_id":47582482,"text":"Can you please write about your hardware?","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T05:34:13.000Z","created_at_i":1774935253,"id":47583114,"options":[],"parent_id":47583101,"points":null,"story_id":47582482,"text":"Ollama on MacOS is a one-click solution with stable obe-click updates. Happy so far. But the mlx support was the only missing piece for me.","title":null,"type":"comment","url":null},{"author":"benob","children":[{"author":"redmalang","children":[],"created_at":"2026-03-31T07:04:47.000Z","created_at_i":1774940687,"id":47583743,"options":[],"parent_id":47583209,"points":null,"story_id":47582482,"text":"i&#x27;ve found llama.cpp (as i understand it, ollama now uses their own version of this) to work much better in practice, faster and much more flexible.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T05:48:27.000Z","created_at_i":1774936107,"id":47583209,"options":[],"parent_id":47583101,"points":null,"story_id":47582482,"text":"Ollama is a user-friendly UI for LLM inference. It is powered by llama.cpp (or a fork of it) which is more power-user oriented and requires command-line wrangling. GGML is the math library behind llama.cpp and GGUF is the associated file format used for storing LLM weights.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T05:30:24.000Z","created_at_i":1774935024,"id":47583101,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"What is the difference between Ollama, llama.cpp, ggml and gguf?","title":null,"type":"comment","url":null},{"author":"mfa1999","children":[{"author":"solarkraft","children":[{"author":"ysleepy","children":[],"created_at":"2026-03-31T08:36:09.000Z","created_at_i":1774946169,"id":47584348,"options":[],"parent_id":47583445,"points":null,"story_id":47582482,"text":"On my M4 Pro MLX has almost 2x tok&#x2F;s","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T06:23:32.000Z","created_at_i":1774938212,"id":47583445,"options":[],"parent_id":47583159,"points":null,"story_id":47582482,"text":"MLX is a bit faster (low double digit percentage), but uses a bit more RAM. Worthwhile tradeoff for many.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T05:41:00.000Z","created_at_i":1774935660,"id":47583159,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"How does this compare to llama.cpp in terms of performance?","title":null,"type":"comment","url":null},{"author":"puskuruk","children":[],"created_at":"2026-03-31T06:48:34.000Z","created_at_i":1774939714,"id":47583625,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"Finally! My local infra is waiting for it for months!","title":null,"type":"comment","url":null},{"author":"Yukonv","children":[{"author":"davesque","children":[],"created_at":"2026-03-31T20:41:33.000Z","created_at_i":1774989693,"id":47593208,"options":[],"parent_id":47583733,"points":null,"story_id":47582482,"text":"Yeah omlx seems to me like the front runner right now for running MLX models locally in agent workflows (which depend heavily on caching).","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T07:03:48.000Z","created_at_i":1774940628,"id":47583733,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"Good to see Ollama is catching up with the times for inference on Mac. MLX powered inference makes a big difference, especially on M5 as their graphs point out.\nWhat really has been a game changer for my workflow is using <a href=\"https:&#x2F;&#x2F;omlx.ai&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;omlx.ai&#x2F;</a> that has SSD KV cold caching. No longer have to worry about a session falling out of memory and needing to prefill again. Combine that with the M5 Max prefill speed means more time is spend on generation than waiting for 50k+ content window to process.","title":null,"type":"comment","url":null},{"author":"robotswantdata","children":[{"author":"vorticalbox","children":[{"author":"zozbot234","children":[],"created_at":"2026-03-31T10:32:16.000Z","created_at_i":1774953136,"id":47585280,"options":[],"parent_id":47585253,"points":null,"story_id":47582482,"text":"You can also use OpenWebUI locally which should give you a nice friendly UX once you set it up.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T10:29:49.000Z","created_at_i":1774952989,"id":47585253,"options":[],"parent_id":47584020,"points":null,"story_id":47582482,"text":"i like ollama, mostly because the cli is pretty nice. its desktop app has stupid choices like if a model can support tools then the ui should give me the &quot;search&quot; option but it only shows for cloud models.<p>i have ran lmstudio for a while but i don&#x27;t really use local models that much other than to mess about.","title":null,"type":"comment","url":null},{"author":"niek_pas","children":[],"created_at":"2026-03-31T10:56:29.000Z","created_at_i":1774954589,"id":47585503,"options":[],"parent_id":47584020,"points":null,"story_id":47582482,"text":"Serious answer: I don&#x27;t use it that much, it&#x27;s what I happened to download like 1.5 years ago, and it works fine. Happy to see what may be a speed boost, and have little interest in switching to something else (unless my situation changes, of course).","title":null,"type":"comment","url":null},{"author":"eddieroger","children":[],"created_at":"2026-03-31T14:13:51.000Z","created_at_i":1774966431,"id":47587720,"options":[],"parent_id":47584020,"points":null,"story_id":47582482,"text":"`ollama serve` and `ollama run`<p>The devex is great and familiar to folks who have used Docker. Reading through the Lemonade documentation, it seems like a natural migration, but we&#x27;re talking about two steps for getting started versus just one. So I&#x27;d need a reason to make that much change when I&#x27;m happy enough with Ollama.","title":null,"type":"comment","url":null},{"author":"hamdingers","children":[],"created_at":"2026-03-31T14:44:01.000Z","created_at_i":1774968241,"id":47588092,"options":[],"parent_id":47584020,"points":null,"story_id":47582482,"text":"Why not? Also serious.<p>It seems to just work every time I try to use it, the API is easy to work with, the model library is convenient. I&#x27;ve never hit any kind of snag that makes me look elsewhere.","title":null,"type":"comment","url":null},{"author":"fennecfoxy","children":[],"created_at":"2026-04-01T08:58:23.000Z","created_at_i":1775033903,"id":47598498,"options":[],"parent_id":47584020,"points":null,"story_id":47582482,"text":"I agree. In my experience on a macbook as well it can be terrifically hard to get some models to run properly on gpu w&#x2F; ollama, let alone containerised ollama.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T07:48:36.000Z","created_at_i":1774943316,"id":47584020,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"Why are people still using Ollama? Serious.<p>Lemonade or even llama.cpp are much better optimised and arguably just as easy to use.","title":null,"type":"comment","url":null},{"author":"darshanmakwana","children":[],"created_at":"2026-03-31T07:54:56.000Z","created_at_i":1774943696,"id":47584056,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"Really nice to see this!","title":null,"type":"comment","url":null},{"author":"franze","children":[{"author":"AbuAssar","children":[{"author":"franze","children":[],"created_at":"2026-03-31T08:59:50.000Z","created_at_i":1774947590,"id":47584530,"options":[],"parent_id":47584359,"points":null,"story_id":47582482,"text":"good idea","title":null,"type":"comment","url":null},{"author":"woadwarrior01","children":[{"author":"franze","children":[{"author":"jedahan","children":[],"created_at":"2026-03-31T13:27:48.000Z","created_at_i":1774963668,"id":47587063,"options":[],"parent_id":47586408,"points":null,"story_id":47582482,"text":"No need for the extra tap step, this works fine alone:<p><pre><code>    brew install Arthur-Ficial&#x2F;tap&#x2F;apfel</code></pre>","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T12:32:52.000Z","created_at_i":1774960372,"id":47586408,"options":[],"parent_id":47585176,"points":null,"story_id":47582482,"text":"done<p><pre><code>  brew tap Arthur-Ficial&#x2F;tap\n  brew install Arthur-Ficial&#x2F;tap&#x2F;apfel</code></pre>","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T10:19:35.000Z","created_at_i":1774952375,"id":47585176,"options":[],"parent_id":47584359,"points":null,"story_id":47582482,"text":"There&#x27;s a very similar afm CLI that can be installed via Homebrew.<p><a href=\"https:&#x2F;&#x2F;github.com&#x2F;scouzi1966&#x2F;maclocal-api\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;scouzi1966&#x2F;maclocal-api</a>","title":null,"type":"comment","url":null},{"author":"grosswait","children":[],"created_at":"2026-03-31T12:11:37.000Z","created_at_i":1774959097,"id":47586178,"options":[],"parent_id":47584359,"points":null,"story_id":47582482,"text":"Looks like they just added homebrew tap to the instructions","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T08:37:38.000Z","created_at_i":1774946258,"id":47584359,"options":[],"parent_id":47584060,"points":null,"story_id":47582482,"text":"nice project, thanks for sharing.<p>any plans for providing it through brew for easy installation?","title":null,"type":"comment","url":null},{"author":"LeoDaVibeci","children":[{"author":"franze","children":[{"author":"beepbooptheory","children":[{"author":"dgacmu","children":[{"author":"corndoge","children":[{"author":"beepbooptheory","children":[{"author":"solatic","children":[],"created_at":"2026-04-01T05:24:28.000Z","created_at_i":1775021068,"id":47597129,"options":[],"parent_id":47587442,"points":null,"story_id":47582482,"text":"mv &#x2F;Users&#x2F;beepbooptheory &#x2F;Users&#x2F;snappy","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T13:54:54.000Z","created_at_i":1774965294,"id":47587442,"options":[],"parent_id":47587145,"points":null,"story_id":47582482,"text":"But it&#x27;s gotta be just a joke right? Which is why all the examples are just classic things you do with bash&#x2F;unix utilities?<p>I&#x27;ll just say, if not a joke, the bit is appreciated either way!<p>&quot;AI change to the home directory. Make it snappy!&quot;","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T13:33:27.000Z","created_at_i":1774964007,"id":47587145,"options":[],"parent_id":47587026,"points":null,"story_id":47582482,"text":"that part is the system prompt, the script is a function that takes a prompt describing a shell command as an argument","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T13:24:39.000Z","created_at_i":1774963479,"id":47587026,"options":[],"parent_id":47586613,"points":null,"story_id":47582482,"text":"The pile of shell and sed is cleaning up the ai output and then running it in the shell.<p>The instruction to the AI was to create _a_ shell command. So it&#x27;s a random shell command generator (maybe).","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T12:50:15.000Z","created_at_i":1774961415,"id":47586613,"options":[],"parent_id":47584896,"points":null,"story_id":47582482,"text":"What is the AI doing here? Or is this just like being cheeky?","title":null,"type":"comment","url":null},{"author":"jorvi","children":[],"created_at":"2026-03-31T13:31:36.000Z","created_at_i":1774963896,"id":47587113,"options":[],"parent_id":47584896,"points":null,"story_id":47582482,"text":"This really makes me think of A Deepness in the Sky by Vernor Vinge. A loose prequel to A Fire Upon The Deep, and IMO actually the superior story. It plays in the far future of humanity.<p>In part of it, one group tries to take control of a huge ship from another group. They in part do this by trying to bypass all the cybersecurity. But in those far future days, you don&#x27;t interface with all the aeons of layers of command protocols anymore, you just query an AI who does it for you. So, this group has a few tech guys that try the bypass by using the old command protocols directly (in a way the same thing like the iOS exploit that used a vulnerability in a PostScript font library from 90s).<p>Imagine being used to LLM prompting + responses, and suddenly you have to deal with something like<p><pre><code>  sed &#x27;&#x2F;^```&#x2F;d;&#x2F;^#&#x2F;d;s&#x2F;^[[:space:]]\\*&#x2F;&#x2F;;&#x2F;^$&#x2F;d&#x27; | head -1); [[ $r ]]\n</code></pre>\nand generally obtuse terminal output and man pages.<p>:)<p>(offtopic: name your variables, don&#x27;t do <i>local x c r a;</i>. Readability is king, and a few hundred thousand years from now some poor Qeng Ho fellow might thank his lucky stars you did).","title":null,"type":"comment","url":null},{"author":"sn0wf1re","children":[],"created_at":"2026-04-01T21:26:45.000Z","created_at_i":1775078805,"id":47606758,"options":[],"parent_id":47584896,"points":null,"story_id":47582482,"text":"That would be pretty cool if the model was a little more useful, but it isn&#x27;t very good. And the guardrails are hilariously bad.<p><pre><code>     cmd print first and last line from stdin\n    $ echo -n | tail -n 2\n\n\n     cmd print first and last line from stdin using sed\n    error: [guardrail] The request was blocked by Apple&#x27;s safety guardrails. Try rephrasing.\n    no command generated</code></pre>","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T09:47:46.000Z","created_at_i":1774950466,"id":47584896,"options":[],"parent_id":47584603,"points":null,"story_id":47582482,"text":"yeah, it is super limited but also you can now do<p><pre><code>  cmd(){ local x c r a; while [[ $1 == -* ]]; do case $1 in -x)x=1;shift;; -c)c=1;shift;; *)break;; esac; done; r=$(apfel -q -s &#x27;Output only a shell command.&#x27; &quot;$*&quot; | sed &#x27;&#x2F;^```&#x2F;d;&#x2F;^#&#x2F;d;s&#x2F;^[[:space:]]*&#x2F;&#x2F;;&#x2F;^$&#x2F;d&#x27; | head -1); [[ $r ]] || { echo &quot;no command generated&quot;; return 1; }; printf &#x27;\\e[32m$\\e[0m %s\\n&#x27; &quot;$r&quot;; [[ $c ]] &amp;&amp; printf %s &quot;$r&quot; | pbcopy &amp;&amp; echo &quot;(copied)&quot;; [[ $x ]] &amp;&amp; { printf &#x27;Run? [y&#x2F;N] &#x27;; read -r a; [[ $a == y ]] &amp;&amp; eval &quot;$r&quot;; }; return 0; } \n</code></pre>\ncmd find all swift files larger than 1MB<p>cmd -c show disk usage sorted by size<p>cmd -x what process is using port 3000<p>cmd list all git branches merged into main<p>cmd count lines of code by language<p>without calling home or downloading extra local models<p>and well, maybe one day they get their local models .... more powerful, &quot;less afraid&quot; and way more context window.","title":null,"type":"comment","url":null},{"author":"drob518","children":[],"created_at":"2026-03-31T11:46:12.000Z","created_at_i":1774957572,"id":47585939,"options":[],"parent_id":47584603,"points":null,"story_id":47582482,"text":"In Apple\u2019s defense, they did make it do something borderline useful while targeting a baseline of M1 Macs with 8 GB of RAM (and even less in phones).","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T09:07:45.000Z","created_at_i":1774948065,"id":47584603,"options":[],"parent_id":47584060,"points":null,"story_id":47582482,"text":"Honestly I can&#x27;t believe Apple put that foundation model product out the door.  I was so excited about it, but when I tried it, it was such a disappointment.  Glad to hear you calling that out so I know it wasn&#x27;t just me.<p>Looks like they have pivoted completely over to Gemini, thank god.","title":null,"type":"comment","url":null},{"author":"JumpCrisscross","children":[{"author":"franze","children":[{"author":"JumpCrisscross","children":[],"created_at":"2026-03-31T11:39:21.000Z","created_at_i":1774957161,"id":47585873,"options":[],"parent_id":47585800,"points":null,"story_id":47582482,"text":"I thought it was a reference to Wine, the Linux Wine, and then thought of apfelwein. Nvm!","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T11:28:39.000Z","created_at_i":1774956519,"id":47585800,"options":[],"parent_id":47585273,"points":null,"story_id":47582482,"text":"just german for apple, cause reasons","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T10:31:55.000Z","created_at_i":1774953115,"id":47585273,"options":[],"parent_id":47584060,"points":null,"story_id":47582482,"text":"\u2026is it a reference to apfelwein?","title":null,"type":"comment","url":null},{"author":"chid","children":[],"created_at":"2026-03-31T13:33:03.000Z","created_at_i":1774963983,"id":47587139,"options":[],"parent_id":47584060,"points":null,"story_id":47582482,"text":"this is real neat. I&#x27;ll give it a spin.","title":null,"type":"comment","url":null},{"author":"podlp","children":[],"created_at":"2026-03-31T14:11:38.000Z","created_at_i":1774966298,"id":47587670,"options":[],"parent_id":47584060,"points":null,"story_id":47582482,"text":"Neat! I\u2019ve actually been building with AFM, including training some LoRA adapters to help steer the model. With the right feedback mechanisms and guardrails, you can even use it for code generation! Hopefully I\u2019ll have a few apps and tools out soon using AFM. I think embedded AI is the future, and in the next few years more platforms will come around to AI as a local API call, not an authorized HTTP request. That said, AFM is still incredibly premature and I\u2019m experimenting with newer models that perform much better.","title":null,"type":"comment","url":null},{"author":"_doctor_love","children":[],"created_at":"2026-03-31T15:48:50.000Z","created_at_i":1774972130,"id":47589175,"options":[],"parent_id":47584060,"points":null,"story_id":47582482,"text":"Dieser apfel ist sehr lecker!","title":null,"type":"comment","url":null},{"author":"newman314","children":[],"created_at":"2026-03-31T18:02:39.000Z","created_at_i":1774980159,"id":47591203,"options":[],"parent_id":47584060,"points":null,"story_id":47582482,"text":"This is quite interesting. I wonder if AFM is smart enough to do spam classification.","title":null,"type":"comment","url":null},{"author":"Multiplayer","children":[],"created_at":"2026-03-31T21:04:49.000Z","created_at_i":1774991089,"id":47593461,"options":[],"parent_id":47584060,"points":null,"story_id":47582482,"text":"this is great!  Incredibly fast and is working pretty well running loads on my m4 max studio.<p>Weirdly though I&#x27;m getting things like this: Apple FM is fast and free but has a hard limitation \u2014 it can&#x27;t process prompts with\n  Spanish&#x2F;non-English words, which is a dealbreaker for California and Southwest real estate where half the street names are Spanish.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T07:55:31.000Z","created_at_i":1774943731,"id":47584060,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"I created &quot;apfel&quot; <a href=\"https:&#x2F;&#x2F;github.com&#x2F;Arthur-Ficial&#x2F;apfel\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;Arthur-Ficial&#x2F;apfel</a> a CLI for the apple on-device local foundation model (Apple intelligence) yeah its super limited with its 4k context window and super common false positives guardrails (just ask it to describe a color) ... bit still ... using it in bash scripts that just work without calling home &#x2F; out or incurring extra costs feels super powerful.","title":null,"type":"comment","url":null},{"author":"harel","children":[{"author":"sgt","children":[{"author":"harel","children":[{"author":"hu3","children":[{"author":"zozbot234","children":[],"created_at":"2026-03-31T10:14:30.000Z","created_at_i":1774952070,"id":47585131,"options":[],"parent_id":47585011,"points":null,"story_id":47582482,"text":"&gt; These models are dumber and slower than API SoTA models and will always be.<p>Sure but you&#x27;re paying per-token costs on the SoTA models that are roughly an order of magnitude higher than third-party inference on the locally available models. So when you account for per-token cost, the math skews the other way.","title":null,"type":"comment","url":null},{"author":"harel","children":[],"created_at":"2026-03-31T11:16:26.000Z","created_at_i":1774955786,"id":47585688,"options":[],"parent_id":47585011,"points":null,"story_id":47582482,"text":"Actually yes. For example, I run local models for ingested documents, summaries, etc. The local models are fine, and there is no need for me to pay for tokens. Performance is adequate for that purpose as well. There are many other cases where I run at scale, time is flexible so things can move slower, and I rather keep it all in house. I&#x27;m not even getting into areas where data cannot leave the premises for legal reasons. Right now I&#x27;m limited with GPUs mostly. But if that world of local models on Apple silicon is so &quot;good&quot;, there is room to expand it to other fruits...","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T10:02:51.000Z","created_at_i":1774951371,"id":47585011,"options":[],"parent_id":47584749,"points":null,"story_id":47582482,"text":"Is there even enough market for this?<p>These models are dumber and slower than API SoTA models and will always be.<p>My time and sanity is much more expensive than insurance against any risk of sending my garbage code to companies worth hundreds of billions of dollars.<p>For most, it&#x27;s a downgrade to use local models in multiple fronts: total cost of ownership, software maintenance, electricity bill, losing performance on the machine doing the inference, having to deal with more hallucinations&#x2F;bugs&#x2F;lower quality code and slower iteration speed.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T09:24:15.000Z","created_at_i":1774949055,"id":47584749,"options":[],"parent_id":47584247,"points":null,"story_id":47582482,"text":"It&#x27;s odd no manufacturer jumped on this wagon to offer a competitive alternative.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T08:22:40.000Z","created_at_i":1774945360,"id":47584247,"options":[],"parent_id":47584188,"points":null,"story_id":47582482,"text":"Not even close. If you want to run this on PC&#x27;s you need to get a GPU like 5090 but that&#x27;s still not the same cost per token, and it will be less reliable and use a lot more power. Right now the Apple Silicon machines are the most cost effective per token and per watt.","title":null,"type":"comment","url":null},{"author":"theshrike79","children":[{"author":"eigenspace","children":[{"author":"theshrike79","children":[],"created_at":"2026-03-31T12:38:54.000Z","created_at_i":1774960734,"id":47586470,"options":[],"parent_id":47585205,"points":null,"story_id":47582482,"text":"There&#x27;s a reason why it&#x27;s cheaper than the Mac equivalent and it&#x27;s not all because of Apple&#x27;s premium pricing =)<p>But it&#x27;s still the easiest and cleanest way to get decent local AI speeds on a non-Mac.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T10:23:16.000Z","created_at_i":1774952596,"id":47585205,"options":[],"parent_id":47584617,"points":null,"story_id":47582482,"text":"Note though that that a MAX 395 has half the memory bandwidth of a M4 Max chip, and the memory bandwidth is going to be the biggest limiting factor, so you&#x27;ll likely be getting around half the tokens&#x2F;second with that Framework Desktop.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T09:09:02.000Z","created_at_i":1774948142,"id":47584617,"options":[],"parent_id":47584188,"points":null,"story_id":47582482,"text":"Framework Desktop is the closest one with the MAX 385&#x2F;395 chip. It&#x27;s mostly about the memory being fast enough rather than just CPU&#x2F;GPU oomph.<p>The 64GB model is 2240\u20ac base and the 128GB is 3069\u20ac base + all the stuff you need to add to make it an actual computer.<p>As a comparison the 64GB Mac Mini is 2499\u20ac here and a 128GB Mac Studio is 4274\u20ac.","title":null,"type":"comment","url":null},{"author":"dabinat","children":[{"author":"mistercheese","children":[],"created_at":"2026-04-01T04:52:06.000Z","created_at_i":1775019126,"id":47596959,"options":[],"parent_id":47585014,"points":null,"story_id":47582482,"text":"Is it feasible to run LLM inference comparably without CUDA or Rocm? How much of the cost performance goes away?","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T10:03:02.000Z","created_at_i":1774951382,"id":47585014,"options":[],"parent_id":47584188,"points":null,"story_id":47582482,"text":"Intel\u2019s doing interesting things with their Arc GPUs. They\u2019re offering GPUs that aren\u2019t super fast for gaming but are relatively low power and have a boatload of VRAM. The new B70 is half the retail price of a 5090 (probably more like 1&#x2F;3rd or 1&#x2F;4 of actual 5090 selling prices) but has the same amount of memory and half the TDP. So for the same price as a 5090 you could get several and use them together.","title":null,"type":"comment","url":null},{"author":"rubymamis","children":[],"created_at":"2026-03-31T10:08:45.000Z","created_at_i":1774951725,"id":47585076,"options":[],"parent_id":47584188,"points":null,"story_id":47582482,"text":"I wonder if the Snapdragon X Elite already caught up with the Apple&#x27;s M series in that regard - does anybody know?","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T08:15:23.000Z","created_at_i":1774944923,"id":47584188,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"What would be the non Mac computer to run these models locally at the same performance profile? Any similar linux ARM based computers that can reach the same level?","title":null,"type":"comment","url":null},{"author":"janandonly","children":[{"author":"zozbot234","children":[],"created_at":"2026-03-31T09:47:58.000Z","created_at_i":1774950478,"id":47584898,"options":[],"parent_id":47584837,"points":null,"story_id":47582482,"text":"&gt; Please make sure you have a Mac with more than 32GB of unified memory.<p>The lack of proper support for SSD offload (via mmap or otherwise) is really the worst part about this.  There&#x27;s no underlying reason why a 3B-active model shouldn&#x27;t be able to run, however slowly, on a cheap 8GB MacBook Neo with active weights being streamed in from SSD and cached.  (This seems to be in the works for GGML&#x2F;GGUF as part of upgrading to newer upstream versions; no idea whether MLX inference can also support this easily.)","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T09:37:59.000Z","created_at_i":1774949879,"id":47584837,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"&gt; <i>Please make sure you have a Mac with more than 32GB of unified memory.</i><p>Yeah, I can still save money by buying a cheaper device with less RAM and just paying my PPQ.AI or OpenRouter.com fees .","title":null,"type":"comment","url":null},{"author":"daveorzach","children":[],"created_at":"2026-03-31T10:13:33.000Z","created_at_i":1774952013,"id":47585122,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"What are significant differences between Ollama and LM Studio now? I haven\u2019t used Ollama because it was missing MLX when I started using LLM GUIs.","title":null,"type":"comment","url":null},{"author":"domh","children":[{"author":"Octoth0rpe","children":[],"created_at":"2026-03-31T11:11:34.000Z","created_at_i":1774955494,"id":47585631,"options":[],"parent_id":47585599,"points":null,"story_id":47582482,"text":"&gt; it can take anywhere between 6-25 seconds for a response (after lots of thinking) from me asking &quot;Hello world&quot;.<p>That&#x27;s not an unsurprising result given the pretty ambiguous query, hence all the thinking. Asking &quot;write a simple hello world program in python3&quot; results in a much faster response for me (m4 base w&#x2F; 24gb, using qwen3.6:9b).","title":null,"type":"comment","url":null},{"author":"zozbot234","children":[{"author":"drob518","children":[{"author":"Kichererbsen","children":[],"created_at":"2026-03-31T15:15:50.000Z","created_at_i":1774970150,"id":47588584,"options":[],"parent_id":47585802,"points":null,"story_id":47582482,"text":"Solid Terry Pratchett reference right there.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T11:28:59.000Z","created_at_i":1774956539,"id":47585802,"options":[],"parent_id":47585720,"points":null,"story_id":47582482,"text":"Indeed. Qwen doesn\u2019t just second guess itself, it third and fourth guesses itself.","title":null,"type":"comment","url":null},{"author":"domh","children":[],"created_at":"2026-03-31T11:41:10.000Z","created_at_i":1774957270,"id":47585887,"options":[],"parent_id":47585720,"points":null,"story_id":47582482,"text":"OK thanks! That&#x27;s helpful. I ignorantly assumed simpler prompt == faster first response.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T11:18:58.000Z","created_at_i":1774955938,"id":47585720,"options":[],"parent_id":47585599,"points":null,"story_id":47582482,"text":"&gt; it can take anywhere between 6-25 seconds for a response (after lots of thinking) from me asking &quot;Hello world&quot;.<p>Qwen thinking likes to second-guess itself a LOT when faced with simple&#x2F;vague prompts like that. (I&#x27;ll answer it this way. Generating output. Wait, I&#x27;ll answer it that way. Generating output. Wait, I&#x27;ll answer it this way... lather, rinse, repeat.)  I suppose this is their version of &quot;super smart fancy thinking mode&quot;.  Try something more complex instead.","title":null,"type":"comment","url":null},{"author":"xienze","children":[{"author":"domh","children":[],"created_at":"2026-03-31T11:42:13.000Z","created_at_i":1774957333,"id":47585899,"options":[],"parent_id":47585815,"points":null,"story_id":47582482,"text":"Thanks! I assumed simpler == faster, but my ignorance is showing itself.<p>I am using the model they recommended in the blog post - which I assumed was using MLX?","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T11:31:47.000Z","created_at_i":1774956707,"id":47585815,"options":[],"parent_id":47585599,"points":null,"story_id":47582482,"text":"Well, two things. First, \u201chi\u201d isn\u2019t a good prompt for these thinking models. They\u2019ll have an identity crisis trying to answer it. Stupid, but it\u2019s how it is. Stick to real questions.<p>Second, for the best performance on a Mac you want to use an MLX model.","title":null,"type":"comment","url":null},{"author":"fooker","children":[],"created_at":"2026-03-31T13:42:52.000Z","created_at_i":1774964572,"id":47587287,"options":[],"parent_id":47585599,"points":null,"story_id":47582482,"text":"Avoid reasoning models in any situation where you have low tokens&#x2F;second","title":null,"type":"comment","url":null},{"author":"functional_dev","children":[{"author":"duffyjp","children":[{"author":"Patrick_Devine","children":[],"created_at":"2026-03-31T22:41:35.000Z","created_at_i":1774996895,"id":47594425,"options":[],"parent_id":47590251,"points":null,"story_id":47582482,"text":"They are nvidia-fp4 weights, but CUDA support isn&#x27;t _quite_ ready yet, but we&#x27;ve got that cooking.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T16:56:33.000Z","created_at_i":1774976193,"id":47590251,"options":[],"parent_id":47588029,"points":null,"story_id":47582482,"text":"I still don&#x27;t think I understand it.  I saw those nvfp4 models up by chance yesterday and tried them on my Linux PC with a 5060TI 16gb.  Ollama refused to pull them saying they were macOS only.<p>I assumed it was a meta-data bug and posted an issue, but apparently nvfp4 doesn&#x27;t necessarily mean nvidia-fp4.<p><a href=\"https:&#x2F;&#x2F;github.com&#x2F;ollama&#x2F;ollama&#x2F;issues&#x2F;15149\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;ollama&#x2F;ollama&#x2F;issues&#x2F;15149</a>","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T14:39:11.000Z","created_at_i":1774967951,"id":47588029,"options":[],"parent_id":47585599,"points":null,"story_id":47582482,"text":"I did not know, that NVFP4 was handled at the silicon level... until I dug deeper here - <a href=\"https:&#x2F;&#x2F;vectree.io&#x2F;c&#x2F;llm-quantization-from-weights-to-bits-gguf-exl2-nvfp4\" rel=\"nofollow\">https:&#x2F;&#x2F;vectree.io&#x2F;c&#x2F;llm-quantization-from-weights-to-bits-g...</a>","title":null,"type":"comment","url":null},{"author":"EagnaIonat","children":[],"created_at":"2026-03-31T14:47:10.000Z","created_at_i":1774968430,"id":47588141,"options":[],"parent_id":47585599,"points":null,"story_id":47582482,"text":"When MLX comes out you will see a huge difference. I currently moved to LMStudio as it currently supports MLX.","title":null,"type":"comment","url":null},{"author":"kylehotchkiss","children":[],"created_at":"2026-03-31T17:34:28.000Z","created_at_i":1774978468,"id":47590830,"options":[],"parent_id":47585599,"points":null,"story_id":47582482,"text":"I made my M2 Max generate a biryani recipe for me last night with 64gb ram and the baseline qwen3.5:35b model. I used the newest ollama with MLX.<p><a href=\"https:&#x2F;&#x2F;gist.github.com&#x2F;kylehotchkiss&#x2F;8f28e6c75f22a56e8d2d31f2c5c13f2d\" rel=\"nofollow\">https:&#x2F;&#x2F;gist.github.com&#x2F;kylehotchkiss&#x2F;8f28e6c75f22a56e8d2d31...</a><p>Under 3 minutes to get all that. The thinking is amusing, my laptop got quite warm, but for a 35b model on nearly 4 year old hardware, I see the light. This is the future.","title":null,"type":"comment","url":null},{"author":"Patrick_Devine","children":[],"created_at":"2026-03-31T22:39:22.000Z","created_at_i":1774996762,"id":47594403,"options":[],"parent_id":47585599,"points":null,"story_id":47582482,"text":"The 35b-a3b-coding-nvfp4 model has the recommended hyperparameters set for coding, not chatting. If you want to use it to chat you can pull the `35b-a3b-nvfp4` model (it doesn&#x27;t need to re-download the weights again so it will pull quickly) which has the presence penalty turned on which will stop it from thinking so much. You can also try `&#x2F;set nothink` in the CLI which will turn off thinking entirely.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T11:06:58.000Z","created_at_i":1774955218,"id":47585599,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"I have an M4 Max with 48GB RAM. Anyone have any tips for good local models? Context length? Using the model recommended in the blog post (qwen3.5:35b-a3b-coding-nvfp4) with Ollama 0.19.0 and it can take anywhere between 6-25 seconds for a response (after lots of thinking) from me asking &quot;Hello world&quot;. Is this the best that&#x27;s currently achievable with my hardware or is there something that can be configured to get better results?","title":null,"type":"comment","url":null},{"author":"androiddrew","children":[],"created_at":"2026-03-31T11:14:01.000Z","created_at_i":1774955641,"id":47585661,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"Get turboquant 4 bit implemented and this would be game changer.","title":null,"type":"comment","url":null},{"author":"dev_l1x_be","children":[],"created_at":"2026-03-31T11:35:47.000Z","created_at_i":1774956947,"id":47585849,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"&gt; Please make sure you have a Mac with more than 32GB of unified memory.\nTime for an upgrade I guess. If I can run Qwen3.5 locally than it is time to switch over to local first LLM usage.","title":null,"type":"comment","url":null},{"author":"jedisct1","children":[],"created_at":"2026-03-31T11:51:44.000Z","created_at_i":1774957904,"id":47586000,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"Works really great with <a href=\"https:&#x2F;&#x2F;swival.dev\" rel=\"nofollow\">https:&#x2F;&#x2F;swival.dev</a> and qwen3.5.","title":null,"type":"comment","url":null},{"author":"harrouet","children":[{"author":"pram","children":[],"created_at":"2026-03-31T14:02:14.000Z","created_at_i":1774965734,"id":47587551,"options":[],"parent_id":47586154,"points":null,"story_id":47582482,"text":"M4 Max is going to be faster.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T12:08:56.000Z","created_at_i":1774958936,"id":47586154,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"As being on the market for a new mac and comparing refub M4 Max vs M5 _Pro_, I am interested in how much faster the neural engines are -- compared to marketing claims.","title":null,"type":"comment","url":null},{"author":"a-dub","children":[{"author":"Casteil","children":[],"created_at":"2026-03-31T18:55:07.000Z","created_at_i":1774983307,"id":47591833,"options":[],"parent_id":47587303,"points":null,"story_id":47582482,"text":"It&#x27;s gotten significantly better with the advent of local&#x2F;offline MoE models (e.g. qwen3.5:35b-a3b, qwen3:30b-a3b, gpt-oss:20b-3.6b), which offer a good balance of prompt response speed and output quality.<p>&#x27;Dense&#x27; models of yesteryear (e.g. llama:70b, gemma2&#x2F;3:27b) tend to be significantly slower by comparison, therefore, your hardware spends a lot more time &#x27;maxed out&#x27; for a given prompt.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T13:43:48.000Z","created_at_i":1774964628,"id":47587303,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"is local llm inference on modern macbook pros comfortable yet?  when i played with it a year or so ago, it worked fairly ok but definitely produced uncomfortable levels of heat.<p>(regarding mlx, there were toolkits built on mlx that supported qlora fine tuning and inference, but also produced a bunch of heat)","title":null,"type":"comment","url":null},{"author":"abu_ameena","children":[{"author":"raw_anon_1111","children":[{"author":"barelysapient","children":[{"author":"raw_anon_1111","children":[{"author":"innagadadavida","children":[{"author":"zozbot234","children":[],"created_at":"2026-03-31T15:34:37.000Z","created_at_i":1774971277,"id":47588919,"options":[],"parent_id":47588853,"points":null,"story_id":47582482,"text":"ClawBot doesn&#x27;t generally run the model locally, it just talks to remote APIs. No different than any other agentic harness.  You could run a local model on the same Mac Mini as your agent, but it wouldn&#x27;t be very smart and many agentic tasks around computer GUI&#x2F;browser use, etc. would be out of reach.","title":null,"type":"comment","url":null},{"author":"raw_anon_1111","children":[],"created_at":"2026-03-31T15:36:19.000Z","created_at_i":1774971379,"id":47588950,"options":[],"parent_id":47588853,"points":null,"story_id":47582482,"text":"And people using Clawdbot are still not using local inference for the most part\u2026<p>They aren\u2019t buying high end $2000+ Mac Minis.","title":null,"type":"comment","url":null},{"author":"bigyabai","children":[{"author":"JSR_FDED","children":[{"author":"bigyabai","children":[],"created_at":"2026-04-01T17:49:42.000Z","created_at_i":1775065782,"id":47604151,"options":[],"parent_id":47601595,"points":null,"story_id":47582482,"text":"If you refuse to abandon iMessage for $300 in MSRP savings, you are beyond helping.","title":null,"type":"comment","url":null}],"created_at":"2026-04-01T14:39:12.000Z","created_at_i":1775054352,"id":47601595,"options":[],"parent_id":47590797,"points":null,"story_id":47582482,"text":"Not if you want to use Messages to talk to it.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T17:33:28.000Z","created_at_i":1774978408,"id":47590797,"options":[],"parent_id":47588853,"points":null,"story_id":47582482,"text":"&gt; Why aren\u2019t these folks going towards cloud solutions?<p>They are. The majority aren&#x27;t doing inference on a Mac Mini, but instead using it as a local host for cloud-based inference. You could have the same general experience on a $200 Chromebook or $300 Windows box.","title":null,"type":"comment","url":null},{"author":"victorbjorklund","children":[],"created_at":"2026-03-31T18:23:44.000Z","created_at_i":1774981424,"id":47591465,"options":[],"parent_id":47588853,"points":null,"story_id":47582482,"text":"They are running cloud models in almost all cases. Like saying it isn\u2019t cloud when you use the Facebook app on your phone (it is ON your phone and running there).","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T15:31:05.000Z","created_at_i":1774971065,"id":47588853,"options":[],"parent_id":47588299,"points":null,"story_id":47582482,"text":"These are all great statistics, but how do you explain ClawdBot explosion. Even in lower income countries like China. So much demand that Apple can\u2019t keep up production of Mac Minis. Why aren\u2019t these folks going towards cloud solutions? Is it cost or is there some consideration for having more control over their data?","title":null,"type":"comment","url":null},{"author":"JambalayaJimbo","children":[{"author":"raw_anon_1111","children":[],"created_at":"2026-03-31T15:37:36.000Z","created_at_i":1774971456,"id":47588969,"options":[],"parent_id":47588904,"points":null,"story_id":47582482,"text":"And Capital One and Goldman Sachs are both hosted on AWS\u2026","title":null,"type":"comment","url":null},{"author":"Aissen","children":[],"created_at":"2026-04-01T08:28:10.000Z","created_at_i":1775032090,"id":47598324,"options":[],"parent_id":47588904,"points":null,"story_id":47582482,"text":"What&#x27;s the plan to migrate off of it ? <a href=\"https:&#x2F;&#x2F;www.atlassian.com&#x2F;licensing&#x2F;data-center-end-of-life#data-center-eol-general-questions\" rel=\"nofollow\">https:&#x2F;&#x2F;www.atlassian.com&#x2F;licensing&#x2F;data-center-end-of-life#...</a>","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T15:34:05.000Z","created_at_i":1774971245,"id":47588904,"options":[],"parent_id":47588299,"points":null,"story_id":47582482,"text":"The banking industry absolutely does care about privacy of their business data btw.\nWe do use tools like Confluence but they&#x27;re all hosted in our own data centers.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T14:58:45.000Z","created_at_i":1774969125,"id":47588299,"options":[],"parent_id":47588162,"points":null,"story_id":47582482,"text":"70% of the world\u2019s population use at least one Meta property at least once per day. How many of the other 30% are too poor&#x2F;young&#x2F;computer illiterate to be part of an addressable market?<p>Every company has dozens of SaaS products that store their business critical information.  Amazon installs Office on each computer, Slack (they were moving away from Chime when I left), and the sales department uses SalesForce - SA\u2019s and Professional Services (former employee).<p>The addressable market of even companies that care about privacy is not a large addressable market.  How long will it be before computers become cheap enough that can run even GPT 4 level LLMs that companies will give it to all of their developers?","title":null,"type":"comment","url":null},{"author":"amelius","children":[{"author":"woopsn","children":[{"author":"bigyabai","children":[],"created_at":"2026-03-31T17:31:47.000Z","created_at_i":1774978307,"id":47590776,"options":[],"parent_id":47590083,"points":null,"story_id":47582482,"text":"&gt; but they&#x27;ll walled garden the model space somehow for sure.<p>People have said this since Pytorch was published and it&#x27;s not any more true now than it was 10 years ago.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T16:45:00.000Z","created_at_i":1774975500,"id":47590083,"options":[],"parent_id":47589140,"points":null,"story_id":47582482,"text":"It&#x27;s largely out of Meta&#x27;s hands now anyway. The risk here not so much to privacy (it&#x27;s Apple) but they&#x27;ll walled garden the model space somehow for sure.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T15:46:37.000Z","created_at_i":1774971997,"id":47589140,"options":[],"parent_id":47588162,"points":null,"story_id":47582482,"text":"&gt; Different users. Many people care about privacy and aren\u2019t using Meta products.<p>Yeah but if they can rake in 100x as much by making products for people who don&#x27;t care about privacy, then why spend time developing stuff for people who care?<p>There is still a small market left, of course, but that market will not have the billions of R&amp;D behind it.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T14:48:25.000Z","created_at_i":1774968505,"id":47588162,"options":[],"parent_id":47587875,"points":null,"story_id":47582482,"text":"Different users. Many people care about privacy and aren\u2019t using Meta products. And many businesses care about it too and have information policies to protect their IP.","title":null,"type":"comment","url":null},{"author":"abu_ameena","children":[{"author":"raw_anon_1111","children":[{"author":"esseph","children":[{"author":"raw_anon_1111","children":[{"author":"esseph","children":[{"author":"raw_anon_1111","children":[{"author":"esseph","children":[{"author":"raw_anon_1111","children":[{"author":"esseph","children":[],"created_at":"2026-04-01T02:19:08.000Z","created_at_i":1775009948,"id":47595998,"options":[],"parent_id":47594229,"points":null,"story_id":47582482,"text":"That doesn&#x27;t dispute what I said, in fact it agrees with what I said. Read it again.<p>&gt; You have data showing growth in cloud, which I expect and don&#x27;t disagree with. The data I come across shows this too!","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T22:18:06.000Z","created_at_i":1774995486,"id":47594229,"options":[],"parent_id":47593756,"points":null,"story_id":47582482,"text":"Again, anecdotes.  I have public company quarterly statements - you have unsourced quotes. You can quote Geico - I can quote Netflix.  If on prem was really growing, I wouldn\u2019t expect Intel to be in the shitter and I would expect Capex to be focused on Colo centers not cloud.<p>Also when I searched for your quotation the very next paragraph was<p>\u201c This trend does not represent a rejection of cloud computing. Organizations continue investing heavily in cloud services, with Gartner forecasting that global cloud spending will reach approximately $723 billion by the end of 2025.\u201d","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T21:34:15.000Z","created_at_i":1774992855,"id":47593756,"options":[],"parent_id":47589392,"points":null,"story_id":47582482,"text":"&gt; Again - anecdotes is not data. We have data.<p>You have data showing growth in cloud, which I expect and don&#x27;t disagree with. The data I come across shows this too!<p>What I disagree with, from my own experiences and all the data I can seem to find online  is that the growth rate in repatriation is MUCH higher than the growth in cloud.<p>It has flipped over the last 3yr.<p>US Enterprises, Fortune 100, especially. Also a lot of public entities (gov).<p>&quot;In 2025, repatriation is still generally an upward trend. Data from the end of 2024 showed that 86% of CIOs planned to move some public cloud workloads back to private cloud or on-premises \u2014 the highest on record for the Barclays CIO Survey.&quot;<p>&quot;Real examples of cloud repatriation include Dropbox, Adobe, and GEICO. All three companies moved a significant portion of their infrastructure onto public cloud before moving it to a combination of on-premises and hybrid cloud providers.&quot;<p>Noted: SaaS accounts for 46.10% of market revenue, while PaaS is the fastest-growing segment at 21.35% CAGR","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T16:01:59.000Z","created_at_i":1774972919,"id":47589392,"options":[],"parent_id":47589045,"points":null,"story_id":47582482,"text":"Again - anecdotes is not data.  We have data.  That would be about as silly as me citing my own experience as proof that \u201ceveryone is moving to AWS\u201d when I work for a company that is exclusively an AWS partner consulting company.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T15:41:48.000Z","created_at_i":1774971708,"id":47589045,"options":[],"parent_id":47588998,"points":null,"story_id":47582482,"text":"Oh I&#x27;m sure they&#x27;ll continue to have some cloud services, no doubt. But look at VMware for example, even after the insane price increases. Nutanix also seems to be doing quite well. I&#x27;m seeing a fair amount of on-prem bare metal k8s too.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T15:39:04.000Z","created_at_i":1774971544,"id":47588998,"options":[],"parent_id":47588610,"points":null,"story_id":47582482,"text":"<i>Your</i> customers are an anecdote, now compare that to the publicly reported numbers from AWS, GCP and Azure where they all say the only thing keeping them from growing more is the chip shortage.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T15:17:30.000Z","created_at_i":1774970250,"id":47588610,"options":[],"parent_id":47588471,"points":null,"story_id":47582482,"text":"&gt; The world is not moving back to on prem.<p>Lol, you should tell my customers (that are moving back on prem) that!<p>You should also tell Microsoft, who just yesterday said they are going back to focusing on local apps.","title":null,"type":"comment","url":null},{"author":"Aurornis","children":[],"created_at":"2026-03-31T17:19:50.000Z","created_at_i":1774977590,"id":47590592,"options":[],"parent_id":47588471,"points":null,"story_id":47582482,"text":"&gt; Realistically, just looking at Mac prices, the cost of a computer with decent local inference would be around $6000 per person.<p>As someone who has hardware in that price range and plays with local LLMs: The gap between Opus or GPT and the local models is still very large for work beyond simple queries.<p>Self-hosted also starts making my office hot due to all of the power consumption when I use it for anything more than short queries. If you haven&#x27;t heard your Mac&#x27;s fans spin up much yet, running local LLMs will get you acquainted with the sound of their cooling systems at full blast.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T15:07:59.000Z","created_at_i":1774969679,"id":47588471,"options":[],"parent_id":47588343,"points":null,"story_id":47582482,"text":"In the history of cloud computing, prices have mostly only come down especially as inference becomes a commodity.  Realistically, just looking at Mac prices, the cost of a computer with decent local inference would be around $6000 per person.<p>The world is not moving back to on prem.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T15:01:27.000Z","created_at_i":1774969287,"id":47588343,"options":[],"parent_id":47587875,"points":null,"story_id":47582482,"text":"I see it as a long-term tradeoff on user freedom. \nYou pay upfront for a capable hardware, you get your services running locally (you don\u2019t pay subscriptions). \nOr you buy cheap hardware, you still need the same services \u201crunning in some cloud\u201d for $X monthly. X goes up depending on the corporate bottom-line","title":null,"type":"comment","url":null},{"author":"DesiLurker","children":[{"author":"raw_anon_1111","children":[{"author":"nozzlegear","children":[{"author":"raw_anon_1111","children":[],"created_at":"2026-03-31T18:47:43.000Z","created_at_i":1774982863,"id":47591746,"options":[],"parent_id":47590994,"points":null,"story_id":47582482,"text":"And they do have a choice on proactively giving FB more information than just what it infers","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T17:46:23.000Z","created_at_i":1774979183,"id":47590994,"options":[],"parent_id":47588835,"points":null,"story_id":47582482,"text":"People don&#x27;t have a choice between Facebook and not-Facebook-but-still-has-all-of-your-friends-and-family. Abstinence isn&#x27;t a choice here any more than shutting off your cell phone service is a choice; true in the literal sense, but only if you don&#x27;t mind being unreachable to everyone who still has a phone.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T15:30:15.000Z","created_at_i":1774971015,"id":47588835,"options":[],"parent_id":47588816,"points":null,"story_id":47582482,"text":"People A) don\u2019t have to use Meta and B) do have a choice between not using a mobile phone by an ad tech company.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T15:29:01.000Z","created_at_i":1774970941,"id":47588816,"options":[],"parent_id":47587875,"points":null,"story_id":47582482,"text":"you are missing a but &#x27;given a choice&#x27; disclaimer. Meta is pretty much a monopoly in social space. So is Android. given a choice people will absolutely gravitate towards not-always-snooping device. most people with resources anyway, who matter for the AI adoption.<p>Oh an wait till ad companies start selling your healthcare data and you will see how fast things turn &#x27;given a choice&#x27;.","title":null,"type":"comment","url":null},{"author":"roadside_picnic","children":[{"author":"charcircuit","children":[{"author":"KerrAvon","children":[],"created_at":"2026-03-31T18:47:10.000Z","created_at_i":1774982830,"id":47591736,"options":[],"parent_id":47591059,"points":null,"story_id":47582482,"text":"As one of those users: absolutely fucking not.","title":null,"type":"comment","url":null},{"author":"xprnio","children":[],"created_at":"2026-04-01T11:03:54.000Z","created_at_i":1775041434,"id":47599261,"options":[],"parent_id":47591059,"points":null,"story_id":47582482,"text":"You will own nothing, and you will be greatful for it","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T17:51:02.000Z","created_at_i":1774979462,"id":47591059,"options":[],"parent_id":47590247,"points":null,"story_id":47582482,"text":"Those users are addressed by being able to rent their own exclusive machines to run the model on. There will be some compromise that will be made to get access to the best intelligence available.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T16:56:21.000Z","created_at_i":1774976181,"id":47590247,"options":[],"parent_id":47587875,"points":null,"story_id":47582482,"text":"&gt; Users don\u2019t care about \u201cprivacy\u201d.<p>I worked for a research focused AI startup that had a strict &quot;no external LLM&quot; policy for code touching our core research.<p>You&#x27;re right that the average <i>consumer</i> doesn&#x27;t care about privacy, but there are many, many <i>users</i> who do. The average consumer also don&#x27;t have a desktop with GPU or high end Mac Studio, but that doesn&#x27;t mean there aren&#x27;t many people working with AI how <i>do</i> have these things.<p>If we continue to see improvements in running local models, and RAM prices continue to fall as they have in the last month, then suddenly you don&#x27;t have to worry about token counts any more and can be much more trusting of your agents since they are fully under your control.","title":null,"type":"comment","url":null},{"author":"Angostura","children":[],"created_at":"2026-03-31T17:33:55.000Z","created_at_i":1774978435,"id":47590807,"options":[],"parent_id":47587875,"points":null,"story_id":47582482,"text":"It\u2019s not all or nothing there ads trade offs. The fact that Apple still bothers to expend marketing effort on its privacy chops suggests significant numbers of people still <i>do</i> care.","title":null,"type":"comment","url":null},{"author":"ilovecake1984","children":[],"created_at":"2026-03-31T17:37:06.000Z","created_at_i":1774978626,"id":47590863,"options":[],"parent_id":47587875,"points":null,"story_id":47582482,"text":"Users here probably means corporations. I still don\u2019t see much use of LLMs in my personal life, other than one thing. Googling stuff in a foreign language.","title":null,"type":"comment","url":null},{"author":"api","children":[],"created_at":"2026-03-31T17:42:04.000Z","created_at_i":1774978924,"id":47590929,"options":[],"parent_id":47587875,"points":null,"story_id":47582482,"text":"&quot;Users&quot; is a large set of people. Many don&#x27;t care about privacy, but some do. There&#x27;s also a difference between where you post random social media stuff vs what you run with something like OpenClaw and give access to your machine.","title":null,"type":"comment","url":null},{"author":"Nevermark","children":[{"author":"raw_anon_1111","children":[{"author":"Nevermark","children":[{"author":"raw_anon_1111","children":[{"author":"Nevermark","children":[{"author":"raw_anon_1111","children":[{"author":"Nevermark","children":[{"author":"raw_anon_1111","children":[{"author":"Nevermark","children":[{"author":"raw_anon_1111","children":[{"author":"Nevermark","children":[{"author":"raw_anon_1111","children":[{"author":"Nevermark","children":[],"created_at":"2026-04-02T11:28:38.000Z","created_at_i":1775129318,"id":47612973,"options":[],"parent_id":47597453,"points":null,"story_id":47582482,"text":"Of course. Nobody disputes whether.","title":null,"type":"comment","url":null}],"created_at":"2026-04-01T06:15:03.000Z","created_at_i":1775024103,"id":47597453,"options":[],"parent_id":47596621,"points":null,"story_id":47582482,"text":"Because it is irrelevant to whether people are purposefully explicitly sharing their likes and dislikes, and other information to let FB know more about them.","title":null,"type":"comment","url":null}],"created_at":"2026-04-01T03:56:08.000Z","created_at_i":1775015768,"id":47596621,"options":[],"parent_id":47596455,"points":null,"story_id":47582482,"text":"Please reread what i wrote from the start more carefully.<p>I am not claiming what percentage of people care or not. I made a valid point of what is evidence or not, for not caring.<p>You also responded to my E2EE reference without absorbing the example.<p>It isn\u2019t a big deal. We can move on.","title":null,"type":"comment","url":null}],"created_at":"2026-04-01T03:27:29.000Z","created_at_i":1775014049,"id":47596455,"options":[],"parent_id":47596169,"points":null,"story_id":47582482,"text":"I am not confused at all.<p>You\u2019re arguing that people care about their privacy when they are explicitly sharing private information above what is needed to participate in FB.<p>You are completely wrong and your argument is illogical.  People may not know that FB is making a profile of you based on your behavior.  But logically, if I add to my profile that my favorite site is \u201cgrandma-midget-porn.com\u201d [1], that I care that people don\u2019t know I like senior citizen midgets<p>[1] Please don\u2019t let that be a real website.","title":null,"type":"comment","url":null}],"created_at":"2026-04-01T02:46:56.000Z","created_at_i":1775011616,"id":47596169,"options":[],"parent_id":47594921,"points":null,"story_id":47582482,"text":"I already gave you a storage&#x2F;share E2EE example.<p>I suggest that when things keep going over your head, like they did here, just Google the topic.<p>And when people are kind enough to reply to your confusion, read with a little more care.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T23:40:07.000Z","created_at_i":1775000407,"id":47594921,"options":[],"parent_id":47593002,"points":null,"story_id":47582482,"text":"When you post a check in, your relationship status, your pictures without setting your sharing preferences and update your profile - you are specifically doing with the intention to share.  WhatsApp is E2E encrypted","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T20:26:46.000Z","created_at_i":1774988806,"id":47593002,"options":[],"parent_id":47592784,"points":null,"story_id":47582482,"text":"An E2EE system (e.g. as offered by Apple iCloud). Or a terms of service guarantee. (e.g. Dropbox, Anthropic and 1000 other companies that partition sharable user content from non-support divisions.)<p>&gt; Would the button disable them from checking in and updating their profile?<p>No.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T20:09:48.000Z","created_at_i":1774987788,"id":47592784,"options":[],"parent_id":47592521,"points":null,"story_id":47582482,"text":"They are explicitly adding their information to FB why do they need a button to not share the information? Would  the button disable them from checking in and updating their profile?","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T19:50:54.000Z","created_at_i":1774986654,"id":47592521,"options":[],"parent_id":47592348,"points":null,"story_id":47582482,"text":"&gt; You don\u2019t have to share everything I mentioned just to be involved in a group.<p>This is clearly true. There is an implied point here but I am not sure what.<p>They share in their profile what they want other people to see. And often choose to not fill out everything. Nobody signs up to share with Meta, Inc.<p>Most people would love a &quot;[ ] Do not share with Facebook&quot;.<p>People choosing an imperfect option, from imperfect options, are not demonstrating evidence they don&#x27;t care about the imperfections.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T19:35:58.000Z","created_at_i":1774985758,"id":47592348,"options":[],"parent_id":47592002,"points":null,"story_id":47582482,"text":"Wouldn\u2019t the most obvious way for people to protect their privacy while using FB  if they cared and still wanted to use FB be not to proactively give them information?  You don\u2019t have to share everything I mentioned just to be involved in a group.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T19:07:44.000Z","created_at_i":1774984064,"id":47592002,"options":[],"parent_id":47591688,"points":null,"story_id":47582482,"text":"&gt; They are choosing to give Facebook info.<p>Yes, they do. That&#x27;s is exactly the phenomena my comment addressed.<p>But the way you wrote that implies an improbable motivation or choice framing.<p>Perhaps their real motive&#x2F;choice is to share with other people on the site.<p>It is called a network effect.<p>If (1) Facebook had been the surveillance&#x2F;manipulation capital of the world from inception, (2) an equally inviting privacy protecting site took off at the same time, and (3) everyone chose Facebook over E2EE anyway, then sure, we could throw up our hands! Those silly users!<p>The term I have for when people discuss choices involving many-dimensional criteria, as if the choice involved just one or two selected dimensions, is &quot;dimension blindness&quot;. It happens in a lot of heated discussions about phone choices too.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T18:43:42.000Z","created_at_i":1774982622,"id":47591688,"options":[],"parent_id":47591083,"points":null,"story_id":47582482,"text":"Consumers pro actively tell Facebook their age, sexual preference, race, relationship status, likes and dislikes, they check in to where they are and who they are there with\u2026<p>They are choosing to give Facebook info.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T17:53:26.000Z","created_at_i":1774979606,"id":47591083,"options":[],"parent_id":47587875,"points":null,"story_id":47582482,"text":"Have you done A&#x2F;B tests to see if consumers prefer Facebook with or without privacy?<p>No? What? Oh, you can&#x27;t?<p>Neither can consumers. Most consumers are very aware of the lack of privacy, the manipulation, and have very cynical feelings about Facebook and similar companies. But it&#x27;s where their friends and family are.<p>For most people the web is a mine field maze where basic things they want are compromised everywhere. And they are routinely creeped out by ads that reveal they know them far too personally.<p>You are mistaking network capture for preference.<p>Another telling example. Lots of privacy valuing technical people, who would never have a Facebook account, send unencrypted text emails.<p>It is network capture, not preference.","title":null,"type":"comment","url":null},{"author":"barkerja","children":[],"created_at":"2026-03-31T17:57:02.000Z","created_at_i":1774979822,"id":47591137,"options":[],"parent_id":47587875,"points":null,"story_id":47582482,"text":"User&#x27;s care about privacy when they understand the threat and impact. The issue is most user&#x27;s don&#x27;t understand this, especially when it comes to use of products like Meta where on the surface, everything appears harmless.","title":null,"type":"comment","url":null},{"author":"duxup","children":[],"created_at":"2026-03-31T20:46:31.000Z","created_at_i":1774989991,"id":47593273,"options":[],"parent_id":47587875,"points":null,"story_id":47582482,"text":"Yeah I agree, I fear users don\u2019t care \u201cenough\u201d about privacy that it will matter. :(<p>Care at all sure, but enough to make a difference, the history of the web and recent computing history indicates otherwise.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T14:26:11.000Z","created_at_i":1774967171,"id":47587875,"options":[],"parent_id":47587546,"points":null,"story_id":47582482,"text":"Users don\u2019t care about \u201cprivacy\u201d.  If they did, Meta and Alphabet wouldn\u2019t be worth $1T+.<p>Users really don\u2019t matter at all.  The revenue for AI companies will be B2B where the user is not the customer  - including coding agents.  Most people don\u2019t even use computers as their primary \u201ccomputing device\u201d and most people are buying crappy low end Android phones - no I\u2019m not saying all Android phones are crappy.  But that\u2019s what most people are buying with the average selling price of an Android phone being $300.","title":null,"type":"comment","url":null},{"author":"testing22321","children":[{"author":"samuel","children":[{"author":"testing22321","children":[],"created_at":"2026-03-31T20:48:31.000Z","created_at_i":1774990111,"id":47593287,"options":[],"parent_id":47588199,"points":null,"story_id":47582482,"text":"Thanks. What do you do with such an agent? What is the use case?","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T14:50:59.000Z","created_at_i":1774968659,"id":47588199,"options":[],"parent_id":47588093,"points":null,"story_id":47582482,"text":"Chat is certainly an option, but the real deal are agents, which have access to way more sensitive information.","title":null,"type":"comment","url":null},{"author":"dec0dedab0de","children":[{"author":"roboror","children":[],"created_at":"2026-03-31T18:13:16.000Z","created_at_i":1774980796,"id":47591338,"options":[],"parent_id":47588620,"points":null,"story_id":47582482,"text":"What models have you found capable? I was recently recommended Qwen3 Coder Next and I did not find it very successful. I have a good amount of VRAM&#x2F;RAM so would love to run something locally.","title":null,"type":"comment","url":null},{"author":"testing22321","children":[],"created_at":"2026-03-31T21:36:20.000Z","created_at_i":1774992980,"id":47593779,"options":[],"parent_id":47588620,"points":null,"story_id":47582482,"text":"Thanks.<p>I still don\u2019t understand. What are you using this long you\u2019re running locally to actually do?<p>What is the use case?","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T15:18:00.000Z","created_at_i":1774970280,"id":47588620,"options":[],"parent_id":47588093,"points":null,"story_id":47582482,"text":"most of the llm tooling can handle different models.  Ollama makes it easy to install and run different models locally.  So you can configure aider or vscode or whatever you&#x27;re using to connect to chatgpt to point to your local models instead.<p>None of them are as good as the big hosted models, but you might be surprised at how capable they are.  I like running things locally when I can,  and I also like not worrying about accidentally burning through tokens.<p>I think the future is multiple locally run models that call out to hosted models when necessary. I can imagine every device coming with a base model and using loras to learn about the users needs.  With companies and maybe even households having their own shared models that do heavier lifting. while companies like openai and anhtropic continue to host the most powerful and expensive options.","title":null,"type":"comment","url":null},{"author":"svachalek","children":[{"author":"testing22321","children":[],"created_at":"2026-03-31T20:47:26.000Z","created_at_i":1774990046,"id":47593282,"options":[],"parent_id":47589666,"points":null,"story_id":47582482,"text":"Thanks. I understand that.<p>What are you doing with it?<p>Why do you want it?","title":null,"type":"comment","url":null},{"author":"jkl5xx","children":[],"created_at":"2026-03-31T21:38:15.000Z","created_at_i":1774993095,"id":47593805,"options":[],"parent_id":47589666,"points":null,"story_id":47582482,"text":"Good points. What local models have you found work best for your use cases? I feel like if we get to opus 4.6 level intelligence running on local hardware, we\u2019re in the clear for a lot of day to day use cases.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T16:17:25.000Z","created_at_i":1774973845,"id":47589666,"options":[],"parent_id":47588093,"points":null,"story_id":47582482,"text":"1. There are small local models that have the capabilities of frontier models a year ago<p>2. They aren&#x27;t harvesting your data for government files or training purposes<p>3. They won&#x27;t be altered overnight to push advertising or a political agenda<p>4. They won&#x27;t have their pricing raised at will<p>5. They won&#x27;t disappear as soon as their host wants you to switch","title":null,"type":"comment","url":null},{"author":"derangedHorse","children":[],"created_at":"2026-04-01T11:16:42.000Z","created_at_i":1775042202,"id":47599357,"options":[],"parent_id":47588093,"points":null,"story_id":47582482,"text":"Qwen3.5 is like an old version of ChatGPT and I can use it the same way I used GPT4 \u2014 writing emails, reading documentation and answering questions about it, reviewing code, answering trivia, etc.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T14:44:02.000Z","created_at_i":1774968242,"id":47588093,"options":[],"parent_id":47587546,"points":null,"story_id":47582482,"text":"I see all these LLM posts about if a certain model can run locally on certain hardware and I don\u2019t get it.<p>What are you doing with these local models that run at x tokens&#x2F;sec.<p>Do you have the equivalent of ChatGPT running entirely locally? What do you do with it? Why? I honestly don\u2019t understand the point or use case.","title":null,"type":"comment","url":null},{"author":"jesse23","children":[],"created_at":"2026-03-31T16:11:13.000Z","created_at_i":1774973473,"id":47589552,"options":[],"parent_id":47587546,"points":null,"story_id":47582482,"text":"Yes so far do we have a working practice that, with a given local mode, any infra we could use, that provide a good practice that can leverage it for local task?","title":null,"type":"comment","url":null},{"author":"thefourthchime","children":[],"created_at":"2026-03-31T16:25:08.000Z","created_at_i":1774974308,"id":47589786,"options":[],"parent_id":47587546,"points":null,"story_id":47582482,"text":"Maybe some more distant future. For me, I&#x27;m still struggling with the hallucinations and screw-ups that the state-of-the-art models give me.","title":null,"type":"comment","url":null},{"author":"mrinterweb","children":[{"author":"fauigerzigerk","children":[],"created_at":"2026-03-31T19:30:42.000Z","created_at_i":1774985442,"id":47592274,"options":[],"parent_id":47590424,"points":null,"story_id":47582482,"text":"Who will be funding state of the art local models going forward? AI models are never done or good enough. They will have to be trained on new data and eventually with new model architectures. It will remain an expensive exercise.<p>I could be wrong because I&#x27;m not following this too closely, but the open weights future of both Llama and Qwen looks tenuous to me. Yes, there are others, but I don&#x27;t understand the business model.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T17:08:47.000Z","created_at_i":1774976927,"id":47590424,"options":[],"parent_id":47587546,"points":null,"story_id":47582482,"text":"I think two recent advances make your statement more true. The new Qwen 3.5 series has shown a relatively high intelligence density, and Google&#x27;s new turboquant could result in dramatically smaller&#x2F;efficient models without the normal quantization accuracy tradeoff.<p>I would expect consumer inference ASIC chips will emerge when model developments start plateauing, and &quot;baking&quot; a highly capable and dense model to a chip makes economic sense.","title":null,"type":"comment","url":null},{"author":"mgaunard","children":[{"author":"Lucasoato","children":[],"created_at":"2026-03-31T17:18:42.000Z","created_at_i":1774977522,"id":47590580,"options":[],"parent_id":47590554,"points":null,"story_id":47582482,"text":"They will eventually catch up, that\u2019s the hope to avoid a techno feudalism in which too much power is in too few hands.","title":null,"type":"comment","url":null},{"author":"abu_ameena","children":[],"created_at":"2026-03-31T17:42:59.000Z","created_at_i":1774978979,"id":47590941,"options":[],"parent_id":47590554,"points":null,"story_id":47582482,"text":"Yes, but you don\u2019t always want the power&#x2F;expense of these models for the task at hand. A hammer is good enough to push a nail inside a wall. Save the nail gun for when you are building a house.","title":null,"type":"comment","url":null},{"author":"sbassi","children":[],"created_at":"2026-03-31T19:23:15.000Z","created_at_i":1774984995,"id":47592184,"options":[],"parent_id":47590554,"points":null,"story_id":47582482,"text":"It&#x27;s a trade off.","title":null,"type":"comment","url":null},{"author":"anon373839","children":[],"created_at":"2026-03-31T20:52:40.000Z","created_at_i":1774990360,"id":47593330,"options":[],"parent_id":47590554,"points":null,"story_id":47582482,"text":"They\u2019re not far behind, unless you mean for \u201cvibe coding\u201d.  And for probably 85% of queries that people use LLMs for, you can\u2019t even really perceive the difference between frontier and local.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T17:17:22.000Z","created_at_i":1774977442,"id":47590554,"options":[],"parent_id":47587546,"points":null,"story_id":47582482,"text":"These local models are far behind the capabilities of latest Gemini Pro, Claude Opus or GPT.<p>Why waste time with subpar AI?","title":null,"type":"comment","url":null},{"author":"sowbug","children":[{"author":"port11","children":[],"created_at":"2026-04-05T16:05:54.000Z","created_at_i":1775405154,"id":47650797,"options":[],"parent_id":47590694,"points":null,"story_id":47582482,"text":"There\u2019s been some success training models on top of differential privacy.<p>I imagine that with live requests it would be quite challenging but not impossible, assuming you could somehow sanitize all sorts of private data that people throw at these prompts.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T17:26:58.000Z","created_at_i":1774978018,"id":47590694,"options":[],"parent_id":47587546,"points":null,"story_id":47582482,"text":"I am concerned that local models will never benefit from the training on live requests that is surely improving cloud-only models.<p>This might be the cost of privacy, and it might be worth paying, unless cloud models reach an inflection point that make local models archaic.","title":null,"type":"comment","url":null},{"author":"throwawayq3423","children":[],"created_at":"2026-03-31T18:15:27.000Z","created_at_i":1774980927,"id":47591357,"options":[],"parent_id":47587546,"points":null,"story_id":47582482,"text":"Technologists make the same mistake over and over in thinking the better technology will win. vhs vs betamax, etc.<p>Actual consumers not only don&#x27;t care, they will not even be aware of the difference.","title":null,"type":"comment","url":null},{"author":"whazor","children":[{"author":"michaelmior","children":[],"created_at":"2026-03-31T18:21:20.000Z","created_at_i":1774981280,"id":47591442,"options":[],"parent_id":47591436,"points":null,"story_id":47582482,"text":"&gt; no reason why future devices couldn&#x27;t bundle 256GB of mem by default<p>Cost is a pretty big reason.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T18:20:46.000Z","created_at_i":1774981246,"id":47591436,"options":[],"parent_id":47587546,"points":null,"story_id":47582482,"text":"Obviously hardware wise the real blocker is memory cost. But there is no reason why future devices couldn&#x27;t bundle 256GB of mem by default.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T14:01:59.000Z","created_at_i":1774965719,"id":47587546,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"On-device models are the future. Users prefer them. No privacy issues. No dealing with connectivity, tokens, or changes to vendors implementations. I have an app using Foundation Model, and it works great. I only wish I could backport it to pre macOS 26 versions.","title":null,"type":"comment","url":null},{"author":"ranjeethacker","children":[],"created_at":"2026-03-31T14:04:52.000Z","created_at_i":1774965892,"id":47587579,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"I used today, working nicely.","title":null,"type":"comment","url":null},{"author":"braum","children":[{"author":"EagnaIonat","children":[],"created_at":"2026-03-31T14:45:53.000Z","created_at_i":1774968353,"id":47588122,"options":[],"parent_id":47588044,"points":null,"story_id":47582482,"text":"You can create an MCP to call out to Ollama. Then have Claude farm work out to local models where the raw power isn&#x27;t required. You can then have Claude review the work from the model.<p>Its not 100% offline, but there is a dramatic drop in token usage. As long as you can put up with the speed.","title":null,"type":"comment","url":null},{"author":"navigate8310","children":[],"created_at":"2026-03-31T14:50:06.000Z","created_at_i":1774968606,"id":47588188,"options":[],"parent_id":47588044,"points":null,"story_id":47582482,"text":"I believe one can use the CC as the primary model driving local agents that use local models","title":null,"type":"comment","url":null},{"author":"samuel","children":[],"created_at":"2026-03-31T14:54:01.000Z","created_at_i":1774968841,"id":47588240,"options":[],"parent_id":47588044,"points":null,"story_id":47582482,"text":"You can connect it to any anthropic compatible endpoint(kimi allows this) but it&#x27;s a weird choice, given that Open code, pi.dev and others are open source.","title":null,"type":"comment","url":null},{"author":"0xc133","children":[],"created_at":"2026-03-31T15:17:10.000Z","created_at_i":1774970230,"id":47588606,"options":[],"parent_id":47588044,"points":null,"story_id":47582482,"text":"<a href=\"https:&#x2F;&#x2F;docs.ollama.com&#x2F;integrations&#x2F;claude-code\">https:&#x2F;&#x2F;docs.ollama.com&#x2F;integrations&#x2F;claude-code</a><p>You can use models like qwen3.5 running on local hardware in ollama and redirect Claude to use the local ollama API endpoint instead of Anthropic\u2019s servers.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T14:40:08.000Z","created_at_i":1774968008,"id":47588044,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"How does Ollama help with Claude Code? Claude code runs in terminal but AFAIK connects back to anthropic directly and cannot run locally. I hope I&#x27;m missing something obvious.","title":null,"type":"comment","url":null},{"author":"xmddmx","children":[],"created_at":"2026-03-31T14:51:19.000Z","created_at_i":1774968679,"id":47588208,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"On a M4 Pro MacBook Pro with 48GB RAM I did this test:<p>ollama run $model &quot;calculate fibonacci numbers in a one-line bash script&quot; --verbose<p><pre><code>  Model                         PromptEvalRate EvalRate\n  ------------------------------------------------------\n  qwen3.5:35b-a3b-q4_K_M         6.6            30.0\n  qwen3.5:35b-a3b-nvfp4         13.2            66.5\n  qwen3.5:35b-a3b-int4          59.4            84.4\n\n</code></pre>\nI can&#x27;t comment on the quality differences (if any) between these three.","title":null,"type":"comment","url":null},{"author":"rurban","children":[],"created_at":"2026-03-31T15:03:53.000Z","created_at_i":1774969433,"id":47588396,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"Does that mean they are now finally a bit faster than llama.cpp? Cannot believe that.","title":null,"type":"comment","url":null},{"author":"bwfan123","children":[{"author":"xiphias2","children":[],"created_at":"2026-03-31T15:18:24.000Z","created_at_i":1774970304,"id":47588626,"options":[],"parent_id":47588493,"points":null,"story_id":47582482,"text":"It doesn&#x27;t look like RAM, CPU GPU or bandwidth is getting cheaper if that helps you, quite the opposite.","title":null,"type":"comment","url":null},{"author":"KerrickStaley","children":[],"created_at":"2026-03-31T16:58:17.000Z","created_at_i":1774976297,"id":47590271,"options":[],"parent_id":47588493,"points":null,"story_id":47582482,"text":"I think (without having done extensive research) that some sort of Apple hardware is your best bet right now. Apple hasn\u2019t raised RAM upgrade prices [1] (although to be fair their RAM upgrades were hugely inflated before the crunch) and their high memory bandwidth means they do inference faster than most consumer GPUs.<p>I have an M4 MacBook Air with 24 GB RAM and it doesn\u2019t feel sufficient to run a substantial coding model (in addition to all my desktop apps). I\u2019m thinking about upgrading to an M5 MacBook Pro with much more RAM, but I think the capabilities of cloud-hosted models will always run ahead of local models and it might never be that useful to do local inference. In the cloud you can run multiple models in parallel (e.g. to work on different problems in parallel) but locally you only have a fixed amount of memory bandwidth so running multiple model instances in parallel is slower.<p>[1] <a href=\"https:&#x2F;&#x2F;9to5mac.com&#x2F;2026&#x2F;03&#x2F;03&#x2F;apple-macbook-price-increase-ram-same&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;9to5mac.com&#x2F;2026&#x2F;03&#x2F;03&#x2F;apple-macbook-price-increase-...</a>","title":null,"type":"comment","url":null},{"author":"victords","children":[],"created_at":"2026-03-31T19:29:22.000Z","created_at_i":1774985362,"id":47592255,"options":[],"parent_id":47588493,"points":null,"story_id":47582482,"text":"As mentioned before, I think Apple hardware is the best alternative right now.<p>Mac Studio, Mac Mini, MacBook Pro, you can find even some used ones with enough RAM that will run models like Qwen reasonably well.<p>I&#x27;m using a M1 Max MacBook Pro and it runs Qwen 3.5 on Ollama (without MLX) at a decent speed.","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T15:09:13.000Z","created_at_i":1774969753,"id":47588493,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"What is the cheapest usable local rig for coding ? I dont want fancy agents and such, but something purpose built for coders, and fast-enough for my use, and open-source, so I can tweak it to my liking. Things are moving fast, and I am hesitant to put in 3-4K now in the hope that it would be cheaper if i wait.","title":null,"type":"comment","url":null},{"author":"jwr","children":[],"created_at":"2026-03-31T15:53:27.000Z","created_at_i":1774972407,"id":47589251,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"Two things: 1) MLX has been available in LM Studio for a long time now, 2) I found that GGUF produced consistently better results in my benchmarking. The difference isn&#x27;t big, but it&#x27;s there.","title":null,"type":"comment","url":null},{"author":"DevKoan","children":[{"author":"peronperon","children":[{"author":"subarctic","children":[],"created_at":"2026-03-31T18:07:29.000Z","created_at_i":1774980449,"id":47591265,"options":[],"parent_id":47589996,"points":null,"story_id":47582482,"text":"What gave this one away \u2014 just the em dashes?","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T16:39:22.000Z","created_at_i":1774975162,"id":47589996,"options":[],"parent_id":47589820,"points":null,"story_id":47582482,"text":"Don&#x27;t post generated comments or AI-edited comments. HN is for conversation between humans. <a href=\"https:&#x2F;&#x2F;news.ycombinator.com&#x2F;newsguidelines.html#comments\">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;newsguidelines.html#comments</a>","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T16:27:16.000Z","created_at_i":1774974436,"id":47589820,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"The Foundation Model point is real. As an iOS developer, what excites me most isn&#x27;t the performance \u2014 it&#x27;s what on-device inference does to the app architecture.<p>When you&#x27;re not making network calls, you stop thinking in &quot;loading states&quot; and start thinking in &quot;local state machines.&quot; The UX design space opens up completely. Interactions that felt too fast to justify a server round-trip are suddenly viable.<p>The backporting issue is painful though. I&#x27;ve been shipping features wrapped in #available(iOS 26, *) and the fallback UX is basically a different product. It forces you to essentially maintain two app experiences.<p>Still think this is the right direction \u2014 especially for junior devs just learning to ship. Fewer moving parts, less infrastructure to debug.","title":null,"type":"comment","url":null},{"author":"adolph","children":[],"created_at":"2026-03-31T17:06:42.000Z","created_at_i":1774976802,"id":47590391,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"Much of the discussion here is local <i>versus</i> remote. I like seeing things as &quot;and&quot; and &quot;or.&quot; There will be small things I don&#x27;t want to burn my Claude tokens on and other things that I want to access larger compute resources. And along the way checking results from both to understand comparative advantage on an ongoing basis.","title":null,"type":"comment","url":null},{"author":"jiehong","children":[],"created_at":"2026-03-31T19:00:19.000Z","created_at_i":1774983619,"id":47591908,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"This is excellent news!<p>What I&#x27;m waiting for next is MLX supported speech recognition directly from Ollama. I don\u2019t understand why it should be a separate thing entirely.","title":null,"type":"comment","url":null},{"author":"pyinstallwoes","children":[],"created_at":"2026-03-31T23:14:58.000Z","created_at_i":1774998898,"id":47594721,"options":[],"parent_id":47582482,"points":null,"story_id":47582482,"text":"What\u2019s the best local coding model these days?","title":null,"type":"comment","url":null}],"created_at":"2026-03-31T03:40:45.000Z","created_at_i":1774928445,"id":47582482,"options":[],"parent_id":null,"points":648,"story_id":47582482,"text":null,"title":"Ollama is now powered by MLX on Apple Silicon in preview","type":"story","url":"https://ollama.com/blog/mlx"}
