{"author":"pember","children":[{"author":"timpera","children":[{"author":"constantcrying","children":[{"author":"crimsoneer","children":[{"author":"constantcrying","children":[{"author":"Lapel2742","children":[{"author":"constantcrying","children":[{"author":"esafak","children":[],"created_at":"2025-12-02T16:31:45.000Z","created_at_i":1764693105,"id":46122973,"options":[],"parent_id":46122828,"points":null,"story_id":46121889,"text":"Which lightweight models do these compare unfavorably with?","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:22:00.000Z","created_at_i":1764692520,"id":46122828,"options":[],"parent_id":46122671,"points":null,"story_id":46121889,"text":"Come on. Do you just not read posts at all?","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:10:22.000Z","created_at_i":1764691822,"id":46122671,"options":[],"parent_id":46122437,"points":null,"story_id":46121889,"text":"&gt;  that the comparisons would be extremely unfavorable.<p>Why should they compare apples to oranges? Ministral3 Large costs ~1&#x2F;10th of Sonnet 4.5. They clearly target different users. If you want a coding assistant you probably wouldn&#x27;t choose this model for various reasons. There is place for more than only the benchmark king.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T15:50:34.000Z","created_at_i":1764690634,"id":46122437,"options":[],"parent_id":46122266,"points":null,"story_id":46121889,"text":"Completely agree, that there are legitimate reasons to prefer comparison to e.g. deepeek models. But that doesn&#x27;t change my point, we both agree that the comparisons would be extremely unfavorable.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T15:36:59.000Z","created_at_i":1764689819,"id":46122266,"options":[],"parent_id":46122169,"points":null,"story_id":46121889,"text":"If someone is using these models, they probably can&#x27;t or won&#x27;t use the existing SOTA models, so not sure how useful those comparisons actually are.  &quot;Here is a benchmark that makes us look bad from a model  you can&#x27;t use on a task you won&#x27;t be undertaking&quot; isn&#x27;t actually helpful (and definitely not in a press release).","title":null,"type":"comment","url":null},{"author":"tarruda","children":[{"author":"constantcrying","children":[{"author":"tarruda","children":[{"author":"meatmanek","children":[{"author":"kergonath","children":[],"created_at":"2025-12-03T08:11:52.000Z","created_at_i":1764749512,"id":46131545,"options":[],"parent_id":46125735,"points":null,"story_id":46121889,"text":"Thanks. That was not obvious to me either.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T19:40:31.000Z","created_at_i":1764704431,"id":46125735,"options":[],"parent_id":46122678,"points":null,"story_id":46121889,"text":"Generation time is more or less proportional to tokens * model size, so if you can get the same quality result with fewer tokens from the same size of model, then you save time and money.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:10:58.000Z","created_at_i":1764691858,"id":46122678,"options":[],"parent_id":46122599,"points":null,"story_id":46121889,"text":"&gt;  Do you disagree with that?<p>I think that Qwen3 8B and 4B are SOTA for their size. The GPQA Diamond accuracy chart is weird: Both Qwen3 8B and 4B have higher scores, so they used this weid chart where &quot;x&quot; axis shows the number of output tokens. I missed the point of this.","title":null,"type":"comment","url":null},{"author":"saubeidl","children":[{"author":"supermatt","children":[],"created_at":"2025-12-02T17:09:53.000Z","created_at_i":1764695393,"id":46123522,"options":[],"parent_id":46122984,"points":null,"story_id":46121889,"text":"&gt; It&#x27;s a separate league from closed models entirely.<p>To be fair, the SOTA models aren&#x27;t even a single LLM these days. They are doing all manner of tool use and specialised submodel calls behind the scenes - a far cry from in-model MoE.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:32:26.000Z","created_at_i":1764693146,"id":46122984,"options":[],"parent_id":46122599,"points":null,"story_id":46121889,"text":"Those <i>are</i> SOTA for open models. It&#x27;s a separate league from closed models entirely.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:03:15.000Z","created_at_i":1764691395,"id":46122599,"options":[],"parent_id":46122564,"points":null,"story_id":46121889,"text":"And implicit in this is that it compares very poorly to SOTA models. Do you disagree with that? Do you think these Models are beating SOTA and they did not include the benchmarks, because they forgot?","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:00:04.000Z","created_at_i":1764691204,"id":46122564,"options":[],"parent_id":46122169,"points":null,"story_id":46121889,"text":"Here&#x27;s what I understood from the blog post:<p>- Mistral Large 3 is comparable with the previous Deepseek release.<p>- Ministral 3 LLMs are comparable with older open LLMs of similar sizes.","title":null,"type":"comment","url":null},{"author":"popinman322","children":[{"author":"extr","children":[],"created_at":"2025-12-02T16:41:54.000Z","created_at_i":1764693714,"id":46123114,"options":[],"parent_id":46122669,"points":null,"story_id":46121889,"text":"??? Closed US frontier models are vastly more effective than anything OSS right now, the reason they didn\u2019t compare is because they\u2019re a different weight class (and therefore product) and it\u2019s a bit unfair.<p>We\u2019re actually at a unique point right now where the gap is larger than it has been in some time. Consensus since the latest batch of releases is that we haven\u2019t found the wall yet. 5.1 Max, Opus 4.5, and G3 are absolutely astounding models and unless you have unique requirements some way down the price&#x2F;perf curve I would not even look at this release (which is fine!)","title":null,"type":"comment","url":null},{"author":"kalkin","children":[],"created_at":"2025-12-02T16:59:49.000Z","created_at_i":1764694789,"id":46123369,"options":[],"parent_id":46122669,"points":null,"story_id":46121889,"text":"Scale AI wrote a paper a year ago comparing various models performance on benchmarks to performance on similar but held-out questions. Generally the closed source models performed better, and Mistral came out looking pretty badly: <a href=\"https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2405.00332\" rel=\"nofollow\">https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2405.00332</a>","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:10:17.000Z","created_at_i":1764691817,"id":46122669,"options":[],"parent_id":46122169,"points":null,"story_id":46121889,"text":"They&#x27;re comparing against open weights models that are roughly a month away from the frontier. Likely there&#x27;s an implicit open-weights political stance here.<p>There are also plenty of reasons not to use proprietary US models for comparison:\nThe major US models haven&#x27;t been living up to their benchmarks; their releases rarely include training &amp; architectural details; they&#x27;re not terribly cost effective; they often fail to compare with non-US models; and the performance delta between model releases has plateaued.<p>A decent number of users in r&#x2F;LocalLlama have reported that they&#x27;ve switched back from Opus 4.5 to Sonnet 4.5 because Opus&#x27; real world performance was worse. From my vantage point it seems like trust in OpenAI, Anthropic, and Google is waning and this lack of comparison is another symptom.","title":null,"type":"comment","url":null},{"author":"bildung","children":[{"author":"BoorishBears","children":[{"author":"sofixa","children":[{"author":"BoorishBears","children":[{"author":"sofixa","children":[],"created_at":"2025-12-02T20:00:05.000Z","created_at_i":1764705605,"id":46126014,"options":[],"parent_id":46125854,"points":null,"story_id":46121889,"text":"&gt; The fact they would not exist without the leeches and built their business on the leeches is irrelevant.<p>How so?","title":null,"type":"comment","url":null},{"author":"baq","children":[],"created_at":"2025-12-02T21:34:50.000Z","created_at_i":1764711290,"id":46127172,"options":[],"parent_id":46125854,"points":null,"story_id":46121889,"text":"If you want to allocate capital efficiently planet-scale you have to ignore nations to the largest extent possible.","title":null,"type":"comment","url":null},{"author":"Fnoord","children":[{"author":"BoorishBears","children":[{"author":"Fnoord","children":[{"author":"BoorishBears","children":[],"created_at":"2025-12-05T07:03:08.000Z","created_at_i":1764918188,"id":46157567,"options":[],"parent_id":46145772,"points":null,"story_id":46121889,"text":"Thank goodness for that, otherwise all we might have is useless copies of Deepseek.","title":null,"type":"comment","url":null}],"created_at":"2025-12-04T09:58:27.000Z","created_at_i":1764842307,"id":46145772,"options":[],"parent_id":46145455,"points":null,"story_id":46121889,"text":"At the very least it is a step in the right direction. Can&#x27;t say the same for these proprietary models. And guess which country has all these proprietary models? USA.","title":null,"type":"comment","url":null}],"created_at":"2025-12-04T09:14:28.000Z","created_at_i":1764839668,"id":46145455,"options":[],"parent_id":46130647,"points":null,"story_id":46121889,"text":"Ah, so &quot;crawled the web without consent, and then put their LLM in a blackbox without attribution&quot; is not being a leech once you release the weights of an underperforming model using someone else&#x27;s arch.<p>I knew y&#x27;all&#x27;s standards were lower but geez!","title":null,"type":"comment","url":null}],"created_at":"2025-12-03T05:37:19.000Z","created_at_i":1764740239,"id":46130647,"options":[],"parent_id":46125854,"points":null,"story_id":46121889,"text":"Those who crawled the web without consent, and then put their LLM in a blackbox without attribution, with secret prompt and secret weights -- ie. all of this without giving back, while creating tons of Co2. Those are the leeches.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T19:49:11.000Z","created_at_i":1764704951,"id":46125854,"options":[],"parent_id":46125672,"points":null,"story_id":46121889,"text":"The fact they would not exist without the leeches and built their business on the leeches is irrelevant.<p>Pan-nationalism is a hell of a drug: a company that does not know you exist puts out an objectively awful release, and people take frank discussion of it as a personal slight.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T19:37:22.000Z","created_at_i":1764704242,"id":46125672,"options":[],"parent_id":46123924,"points":null,"story_id":46121889,"text":"Mistral are mostly focusing on b2b, and for customers that want to self-host (banks and stuff). So their founders being from Meta, or where their cloud platform are hosted, are entirely irrelevant to the story.","title":null,"type":"comment","url":null},{"author":"troyvit","children":[],"created_at":"2025-12-02T19:55:07.000Z","created_at_i":1764705307,"id":46125942,"options":[],"parent_id":46123924,"points":null,"story_id":46121889,"text":"It&#x27;s wayyyy to early in the game to say who is out-executing whom.<p>I mean why do you think those guys left Meta? It reminds me of a time ten years ago I was sitting on a flight with a guy who works for the natural gas industry. I was (<i>cough</i> still am) a pretty naive environmentalist, so I asked him what he thought of solar, wind, etc. and why should we be investing in natural gas when there are all these other options. His response was simple. Natural gas can serve as a bridge from hydrocarbons to true green energy sources. Leverage that dense energy to springboard the other sources in the mix and you build a path forward to carbon free energy.<p>I see Mistral&#x27;s use of US VCs the same way. Those VCs are hedging their bets and maybe hoping to make a few bucks. A few of them are probably involved because they&#x27;re buddies with the former Meta guys &quot;back in the day.&quot; If Mistral executes on their plan of being a transparent b2b option with solid data protections then they used those VCs the way they deserve to be used and the VCs make a few bucks. If Europe ever catches up to the US in terms of data centers, would Mistral move off of Azure? I&#x27;d bet $5 that they would.","title":null,"type":"comment","url":null},{"author":"bildung","children":[],"created_at":"2025-12-03T13:16:25.000Z","created_at_i":1764767785,"id":46134143,"options":[],"parent_id":46123924,"points":null,"story_id":46121889,"text":"I didn&#x27;t mean to imply US bad EU good. As such, this isn&#x27;t about which passport the VCs have, but about local hosting and open weight models. A closed model from a US company always comes with the risk of data exfiltration either for training or thanks to CLOUD Act etc (i.e. industrial espionage).<p>And personally I don&#x27;t care at all about the performance delta - we are talking about a difference of 6 to at most 12 months here, between closed source SOTA and open weight models.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T17:37:24.000Z","created_at_i":1764697044,"id":46123924,"options":[],"parent_id":46123212,"points":null,"story_id":46121889,"text":"Mistral is founded by multiple Meta engineers, no?<p>Funded mostly by US VCs?<p>Hosted primarily on Azure?<p>Do you really have to go out of your way to start calling their competition &quot;data leeches&quot; for out-executing them?","title":null,"type":"comment","url":null},{"author":"adam_patarino","children":[],"created_at":"2025-12-02T19:08:10.000Z","created_at_i":1764702490,"id":46125213,"options":[],"parent_id":46123212,"points":null,"story_id":46121889,"text":"We&#x27;re seeing the same thing for many companies, even in the US. Exposing your entire codebase to an unreliable third party is not exactly SOC &#x2F; ISO compliant. This is one of the core things that motivated us to develop cortex.build so we could put the model on the developer&#x27;s machine and completely isolate the code without complicated model deployments and maintenance.","title":null,"type":"comment","url":null},{"author":"leobg","children":[],"created_at":"2025-12-03T14:01:46.000Z","created_at_i":1764770506,"id":46134557,"options":[],"parent_id":46123212,"points":null,"story_id":46121889,"text":"Does your company use Microsoft Teams?","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:48:59.000Z","created_at_i":1764694139,"id":46123212,"options":[],"parent_id":46122169,"points":null,"story_id":46121889,"text":"I think people from the US often aren&#x27;t aware <i>how many</i> companies from the EU simply won&#x27;t risk losing their data to the providers you have in mind, OpenAI, Anthropic and Google. They simply are no option at all.<p>The company I work for for example, a mid-sized tech business, currently investigates their local hosting options for LLMs. So Mistral certainly will be an option, among the Qwen familiy and Deepseek.<p>Mistral is positioning themselves for that market, not the one you have in mind. Comparing their models with Claude etc. would mean associating themselves with the data leeches, which they probably try to avoid.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T15:29:16.000Z","created_at_i":1764689356,"id":46122169,"options":[],"parent_id":46121991,"points":null,"story_id":46121889,"text":"The lack of the comparison (which absolutely was done), tells you exactly what you need to know.","title":null,"type":"comment","url":null},{"author":"Youden","children":[{"author":"jampekka","children":[{"author":"supermatt","children":[{"author":"esafak","children":[{"author":"uejfiweun","children":[{"author":"JustFinishedBSG","children":[],"created_at":"2025-12-03T10:10:25.000Z","created_at_i":1764756625,"id":46132657,"options":[],"parent_id":46129202,"points":null,"story_id":46121889,"text":"I wouldn&#x27;t trust LMArena results much. They measure user preference and users are highly skewed by style, tone etc.<p>You can litteraly &quot;improve&quot; your model on LMArena by just adding a bunch of emojis.","title":null,"type":"comment","url":null}],"created_at":"2025-12-03T01:28:18.000Z","created_at_i":1764725298,"id":46129202,"options":[],"parent_id":46123287,"points":null,"story_id":46121889,"text":"Wow. If all the trillions only produces that small of a diff... that&#x27;s shocking. That&#x27;s the sort of knowledge that could pop the bubble.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:54:10.000Z","created_at_i":1764694450,"id":46123287,"options":[],"parent_id":46123245,"points":null,"story_id":46121889,"text":"Yes, of course.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:50:56.000Z","created_at_i":1764694256,"id":46123245,"options":[],"parent_id":46123121,"points":null,"story_id":46121889,"text":"Probably naive questions:<p>Does that also mean that Gemini-3 (the top ranked model) loses to mistral 3 40% of the time?<p>Does that make Gemini 1.5x better, or mistral 2&#x2F;3rd as good as Gemini, or can we not quantify the difference like that?","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:42:34.000Z","created_at_i":1764693754,"id":46123121,"options":[],"parent_id":46122538,"points":null,"story_id":46121889,"text":"1491 vs 1418 ELO means the stronger model wins about 60% of the time.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T15:58:05.000Z","created_at_i":1764691085,"id":46122538,"options":[],"parent_id":46121991,"points":null,"story_id":46121889,"text":"They mentioned LMArena, you can get the results for that here: <a href=\"https:&#x2F;&#x2F;lmarena.ai&#x2F;leaderboard&#x2F;text\" rel=\"nofollow\">https:&#x2F;&#x2F;lmarena.ai&#x2F;leaderboard&#x2F;text</a><p>Mistral Large 3 is ranked 28, behind all the other major SOTA models. The delta between Mistral and the leader is only 1418 vs. 1491 though. I *think* that means the difference is relatively small.","title":null,"type":"comment","url":null},{"author":"qznc","children":[],"created_at":"2025-12-02T16:01:00.000Z","created_at_i":1764691260,"id":46122577,"options":[],"parent_id":46121991,"points":null,"story_id":46121889,"text":"I guess that could be considered comparative advertising then and companies generally try to avoid that scrutiny.","title":null,"type":"comment","url":null},{"author":"rvz","children":[],"created_at":"2025-12-02T16:17:19.000Z","created_at_i":1764692239,"id":46122761,"options":[],"parent_id":46121991,"points":null,"story_id":46121889,"text":"&gt; I just wish they would also include comparisons to SOTA models from OpenAI, Google, and Anthropic in the press release,<p>Why would they? They know they can&#x27;t compete against the heavily closed-source models.<p>They are not even comparing against GPT-OSS.<p>That is absolutely and shockingly bearish.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T15:11:36.000Z","created_at_i":1764688296,"id":46121991,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"Extremely cool! I just wish they would also include comparisons to SOTA models from OpenAI, Google, and Anthropic in the press release, so it&#x27;s easier to know how it fares in the grand scheme of things.","title":null,"type":"comment","url":null},{"author":"codybontecou","children":[{"author":"Y_Y","children":[],"created_at":"2025-12-02T15:36:28.000Z","created_at_i":1764689788,"id":46122261,"options":[],"parent_id":46122091,"points":null,"story_id":46121889,"text":"In principle any model can do these. Tool use is just detecting something like &quot;I should run a db query for pattern X&quot; and structured output is even easier, just reject output tokens that don&#x27;t match the grammar. The only question is how well they&#x27;re trained, and how well your inference environment takes advantage.","title":null,"type":"comment","url":null},{"author":"Ey7NFZ3P0nzAe","children":[],"created_at":"2025-12-02T21:10:18.000Z","created_at_i":1764709818,"id":46126917,"options":[],"parent_id":46122091,"points":null,"story_id":46121889,"text":"Yes they all support tool use at least.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T15:21:40.000Z","created_at_i":1764688900,"id":46122091,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"Do all of these models, regardless of parameters, support tool use and structured output?","title":null,"type":"comment","url":null},{"author":"simgt","children":[{"author":"prodigycorp","children":[],"created_at":"2025-12-02T15:37:32.000Z","created_at_i":1764689852,"id":46122273,"options":[],"parent_id":46122146,"points":null,"story_id":46121889,"text":"gpt-oss are really solid models. by far the best at tool calling, and performant.","title":null,"type":"comment","url":null},{"author":"talliman","children":[{"author":"memming","children":[],"created_at":"2025-12-02T15:54:11.000Z","created_at_i":1764690851,"id":46122483,"options":[],"parent_id":46122340,"points":null,"story_id":46121889,"text":"It\u2019s funny how future money drive the world. Fortunately it\u2019s fueling progress this time around.","title":null,"type":"comment","url":null},{"author":"mirekrusin","children":[{"author":"simgt","children":[],"created_at":"2025-12-02T16:21:24.000Z","created_at_i":1764692484,"id":46122818,"options":[],"parent_id":46122537,"points":null,"story_id":46121889,"text":"I was fully expecting that but it doesn&#x27;t get old ;)","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T15:58:04.000Z","created_at_i":1764691084,"id":46122537,"options":[],"parent_id":46122340,"points":null,"story_id":46121889,"text":"Explained well in this documentary [0].<p>[0] <a href=\"https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=BzAdXyPYKQo\" rel=\"nofollow\">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=BzAdXyPYKQo</a>","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T15:42:48.000Z","created_at_i":1764690168,"id":46122340,"options":[],"parent_id":46122146,"points":null,"story_id":46121889,"text":"Until there is a sustainable, profitable and moat-building business model for generative AI, the competition is not to have the best proprietary model, but rather to raise the most VC money to be well positioned when that business model does arise.<p>Releasing a near stat-of-the-art open model instanly catapults companies to a valuation of several billion dollars, making it possible raise money to acquire GPUs and train more SOTA models.<p>Now, what happens if such a business model does not emerge? I hope we won&#x27;t find out!","title":null,"type":"comment","url":null},{"author":"NitpickLawyer","children":[{"author":"lostmsu","children":[{"author":"NitpickLawyer","children":[{"author":"data-ottawa","children":[],"created_at":"2025-12-02T17:45:54.000Z","created_at_i":1764697554,"id":46124026,"options":[],"parent_id":46123729,"points":null,"story_id":46121889,"text":"I find the qwen3 models spend a ton of thinking tokens which could hamstring them on the runtime limitations. Gpt-oss 120b is much more focused and steerable there.<p>The token use chart in the OP release page demonstrates the Qwen issue well.<p>Token churn does help smaller models on math tasks, but for general purpose stuff it seems to hurt.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T17:23:03.000Z","created_at_i":1764696183,"id":46123729,"options":[],"parent_id":46123552,"points":null,"story_id":46121889,"text":"There is a leaderboard [1] but we&#x27;ll have to wait till april for the competition to end to know what models they&#x27;re using. The current number 3 on there (34&#x2F;50) has mentioned in discussions that they&#x27;re using gpt-oss-120b. There were also some scores shared for gpt-oss-20b, in the 25&#x2F;50 range.<p>The next &quot;public&quot; model is qwen30b-thinking at 23&#x2F;50.<p>Competition is limited to 1 H100 (80GB) and 5h runtime for 50 problems. So larger open models (deepseek, larger qwens) don&#x27;t fit.<p>[1] <a href=\"https:&#x2F;&#x2F;www.kaggle.com&#x2F;competitions&#x2F;ai-mathematical-olympiad-progress-prize-3&#x2F;leaderboard\" rel=\"nofollow\">https:&#x2F;&#x2F;www.kaggle.com&#x2F;competitions&#x2F;ai-mathematical-olympiad...</a>","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T17:11:35.000Z","created_at_i":1764695495,"id":46123552,"options":[],"parent_id":46122435,"points":null,"story_id":46121889,"text":"Are they ahead of all other recent open models? Is there a leaderboard?","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T15:50:27.000Z","created_at_i":1764690627,"id":46122435,"options":[],"parent_id":46122146,"points":null,"story_id":46121889,"text":"&gt; gpt-oss that games the benchmarks just for PR.<p>gpt-oss is killing the ongoing AIME3 competition on kaggle. They&#x27;re using a hidden, new set of problems, IMO level, handcrafted to be &quot;AI hardened&quot;. And gpt-oss submissions are at ~33&#x2F;50 right now, two weeks into the competition. The benchmarks (at least for math) were not gamed at all. They are really good at math.","title":null,"type":"comment","url":null},{"author":"mirekrusin","children":[],"created_at":"2025-12-02T15:54:16.000Z","created_at_i":1764690856,"id":46122485,"options":[],"parent_id":46122146,"points":null,"story_id":46121889,"text":"Because there is no money in making them closed.<p>Open weight means secondary sales channels like their fine tuning service for enterprises [0].<p>They can&#x27;t compete with large proprietary providers but they can erode and potentially collapse them.<p>Open weights and research builds on itself advancing its participants creating environment that has a shot at proprietary services.<p>Transparency, control, privacy, cost etc. do matter to people and corporations.<p>[0] <a href=\"https:&#x2F;&#x2F;mistral.ai&#x2F;solutions&#x2F;custom-model-training\" rel=\"nofollow\">https:&#x2F;&#x2F;mistral.ai&#x2F;solutions&#x2F;custom-model-training</a>","title":null,"type":"comment","url":null},{"author":"nullbio","children":[],"created_at":"2025-12-02T17:00:04.000Z","created_at_i":1764694804,"id":46123373,"options":[],"parent_id":46122146,"points":null,"story_id":46121889,"text":"Google games benchmarks more than anyone, hence Gemini&#x27;s strong bench lead. In reality though, it&#x27;s still garbage for general usage.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T15:27:27.000Z","created_at_i":1764689247,"id":46122146,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"I still don&#x27;t understand what the incentive is for releasing genuinely good model weights. What makes sense however is OpenAI releasing a somewhat generic model like gpt-oss that games the benchmarks just for PR. Or some Chinese companies doing the same to cut the ground from under the feet of American big tech. Are we really hopeful we&#x27;ll still get decent open weights models in the future?","title":null,"type":"comment","url":null},{"author":"yvoschaap","children":[{"author":"sebzim4500","children":[{"author":"GaggiX","children":[{"author":"ot","children":[{"author":"GaggiX","children":[],"created_at":"2025-12-02T15:51:24.000Z","created_at_i":1764690684,"id":46122447,"options":[],"parent_id":46122427,"points":null,"story_id":46121889,"text":"I think you missed the joke","title":null,"type":"comment","url":null},{"author":"usrnm","children":[{"author":"rc1","children":[{"author":"rkomorn","children":[],"created_at":"2025-12-03T19:52:57.000Z","created_at_i":1764791577,"id":46139163,"options":[],"parent_id":46139122,"points":null,"story_id":46121889,"text":"Eurasia is the widely accepted answer.","title":null,"type":"comment","url":null}],"created_at":"2025-12-03T19:50:36.000Z","created_at_i":1764791436,"id":46139122,"options":[],"parent_id":46122489,"points":null,"story_id":46121889,"text":"If Europe isn\u2019t a continent, on what continent are the EU member states sitting on?","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T15:54:41.000Z","created_at_i":1764690881,"id":46122489,"options":[],"parent_id":46122427,"points":null,"story_id":46121889,"text":"Europe isn&#x27;t even a continent and has no real definition (none that would make any sense, anyway), so the whole thing is confusing by design","title":null,"type":"comment","url":null},{"author":"lostmsu","children":[{"author":"TulliusCicero","children":[{"author":"layer8","children":[{"author":"lostmsu","children":[],"created_at":"2025-12-03T02:56:05.000Z","created_at_i":1764730565,"id":46129779,"options":[],"parent_id":46129300,"points":null,"story_id":46121889,"text":"What&#x27;s more interesting is that the comment you are replying to mistakenly asked me instead of asking the parent.","title":null,"type":"comment","url":null}],"created_at":"2025-12-03T01:42:16.000Z","created_at_i":1764726136,"id":46129300,"options":[],"parent_id":46123947,"points":null,"story_id":46121889,"text":"While Japan is part of Asia, and Asia is a continent, Japan is also separated from the Asian continent: <a href=\"https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Geography_of_Japan#Location\" rel=\"nofollow\">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Geography_of_Japan#Location</a>","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T17:39:04.000Z","created_at_i":1764697144,"id":46123947,"options":[],"parent_id":46123499,"points":null,"story_id":46121889,"text":"So I guess Japan isn&#x27;t Asian then?","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T17:08:33.000Z","created_at_i":1764695313,"id":46123499,"options":[],"parent_id":46122427,"points":null,"story_id":46121889,"text":"Isn&#x27;t London on an island, mr. Pedantic?","title":null,"type":"comment","url":null},{"author":"denysvitali","children":[{"author":"MadDemon","children":[],"created_at":"2025-12-02T18:23:50.000Z","created_at_i":1764699830,"id":46124489,"options":[],"parent_id":46123643,"points":null,"story_id":46121889,"text":"Switzerland has such close ties to the EU that I would consider them half in.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T17:17:38.000Z","created_at_i":1764695858,"id":46123643,"options":[],"parent_id":46122427,"points":null,"story_id":46121889,"text":"I honestly think it is.\nThe amount of people who thinks Europe and EU are the same thing is really concerning.<p>And no, it&#x27;s not only americans. I keep hearing this thing from people living in Europe as well (or better, in the EU). \nI also very often hear phrases like &quot;Switzerland is not in Europe&quot; to indicate that the country is not part of the European Union.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T15:49:34.000Z","created_at_i":1764690574,"id":46122427,"options":[],"parent_id":46122388,"points":null,"story_id":46121889,"text":"Is it so hard for people to understand that Europe is a continent, EU is a federation of European countries, and the two are not the same?","title":null,"type":"comment","url":null},{"author":"tmoravec","children":[],"created_at":"2025-12-02T18:36:48.000Z","created_at_i":1764700608,"id":46124705,"options":[],"parent_id":46122388,"points":null,"story_id":46121889,"text":"Drifted to the Caribbean.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T15:45:58.000Z","created_at_i":1764690358,"id":46122388,"options":[],"parent_id":46122364,"points":null,"story_id":46121889,"text":"London is not part of Europe anymore since Brexit &#x2F;s","title":null,"type":"comment","url":null},{"author":"p2detar","children":[],"created_at":"2025-12-02T15:56:59.000Z","created_at_i":1764691019,"id":46122526,"options":[],"parent_id":46122364,"points":null,"story_id":46121889,"text":"That&#x27;s ok. How could they know that there are companies like Aleph Alpha, Helsing or the famous DeepL. European companies are not that vocal, but that doesn&#x27;t mean they aren&#x27;t making progress in the field.<p>edit: typos","title":null,"type":"comment","url":null},{"author":"Glemkloksdjf","children":[{"author":"gishh","children":[{"author":"vintermann","children":[],"created_at":"2025-12-02T16:53:27.000Z","created_at_i":1764694407,"id":46123275,"options":[],"parent_id":46123014,"points":null,"story_id":46121889,"text":"Currency is interchangeable. Location might not be.","title":null,"type":"comment","url":null},{"author":"data-ottawa","children":[],"created_at":"2025-12-02T17:38:46.000Z","created_at_i":1764697126,"id":46123942,"options":[],"parent_id":46123014,"points":null,"story_id":46121889,"text":"Increasingly where the desks and servers are is critical.<p>The cloud act and the current US administration doing things like sanctioning the ICC demonstrate why the locations of those desks is important.","title":null,"type":"comment","url":null},{"author":"cycomanic","children":[],"created_at":"2025-12-02T19:01:34.000Z","created_at_i":1764702094,"id":46125120,"options":[],"parent_id":46123014,"points":null,"story_id":46121889,"text":"That&#x27;s such a silly argument. X, OpenAI and others have large Saudi investments. In the grant scheme of things the US is largely indebted to China and Japan.","title":null,"type":"comment","url":null},{"author":"Glemkloksdjf","children":[],"created_at":"2025-12-03T09:55:18.000Z","created_at_i":1764755718,"id":46132549,"options":[],"parent_id":46123014,"points":null,"story_id":46121889,"text":"An EU Company pays taxes in EU, has a EU mindset (worker laws etc.), focuses more on EU than other countries.<p>And an EU company can&#x27;t be forced by the US Gov to hand over data.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:34:16.000Z","created_at_i":1764693256,"id":46123014,"options":[],"parent_id":46122546,"points":null,"story_id":46121889,"text":"Using US VC dollars. Where their desks are isn\u2019t really important.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T15:58:35.000Z","created_at_i":1764691115,"id":46122546,"options":[],"parent_id":46122364,"points":null,"story_id":46121889,"text":"Thats not the point.<p>Deepmind is not an UK company, its google aka US.<p>Mistral is a real EU based company.","title":null,"type":"comment","url":null},{"author":"colesantiago","children":[],"created_at":"2025-12-02T16:11:50.000Z","created_at_i":1764691910,"id":46122690,"options":[],"parent_id":46122364,"points":null,"story_id":46121889,"text":"Deepmind doesn&#x27;t exist anymore.<p>Google DeepMind does exist.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T15:44:23.000Z","created_at_i":1764690263,"id":46122364,"options":[],"parent_id":46122156,"points":null,"story_id":46121889,"text":"That&#x27;s unfair to Europe. A bunch of AI work is done in London (Deepmind is based here for a start)","title":null,"type":"comment","url":null},{"author":"LunaSea","children":[{"author":"DarmokJalad1701","children":[{"author":"LunaSea","children":[{"author":"DarmokJalad1701","children":[],"created_at":"2025-12-02T23:55:23.000Z","created_at_i":1764719723,"id":46128585,"options":[],"parent_id":46125277,"points":null,"story_id":46121889,"text":"&quot;best effort at Operating Systems development&quot; doesn&#x27;t imply anything about the market share.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T19:11:46.000Z","created_at_i":1764702706,"id":46125277,"options":[],"parent_id":46123707,"points":null,"story_id":46121889,"text":"What&#x27;s the market share of those compared to Windows and Linux?","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T17:21:21.000Z","created_at_i":1764696081,"id":46123707,"options":[],"parent_id":46123295,"points":null,"story_id":46121889,"text":"Wouldn&#x27;t that be macOS? Or BSD? Or Unix? CentOS?","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:54:41.000Z","created_at_i":1764694481,"id":46123295,"options":[],"parent_id":46122156,"points":null,"story_id":46121889,"text":"Upvoting Windows 11 as the US&#x27;s best effort at Operating Systems development.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T15:28:13.000Z","created_at_i":1764689293,"id":46122156,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"Upvoting for Europe&#x27;s best efforts.","title":null,"type":"comment","url":null},{"author":"hnuser123456","children":[{"author":"janpio","children":[],"created_at":"2025-12-02T15:43:30.000Z","created_at_i":1764690210,"id":46122352,"options":[],"parent_id":46122209,"points":null,"story_id":46121889,"text":"Seems fixed now:<p><a href=\"https:&#x2F;&#x2F;huggingface.co&#x2F;collections&#x2F;mistralai&#x2F;mistral-large-3\" rel=\"nofollow\">https:&#x2F;&#x2F;huggingface.co&#x2F;collections&#x2F;mistralai&#x2F;mistral-large-3</a><p><a href=\"https:&#x2F;&#x2F;huggingface.co&#x2F;collections&#x2F;mistralai&#x2F;ministral-3\" rel=\"nofollow\">https:&#x2F;&#x2F;huggingface.co&#x2F;collections&#x2F;mistralai&#x2F;ministral-3</a>","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T15:32:01.000Z","created_at_i":1764689521,"id":46122209,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"Looks like their own HF link is broken or the collection hasn&#x27;t been made public yet. The 14B instruct model is here:<p><a href=\"https:&#x2F;&#x2F;huggingface.co&#x2F;mistralai&#x2F;Ministral-3-14B-Instruct-2512\" rel=\"nofollow\">https:&#x2F;&#x2F;huggingface.co&#x2F;mistralai&#x2F;Ministral-3-14B-Instruct-25...</a><p>The unsloth quants are here:<p><a href=\"https:&#x2F;&#x2F;huggingface.co&#x2F;unsloth&#x2F;Ministral-3-14B-Instruct-2512-GGUF\" rel=\"nofollow\">https:&#x2F;&#x2F;huggingface.co&#x2F;unsloth&#x2F;Ministral-3-14B-Instruct-2512...</a>","title":null,"type":"comment","url":null},{"author":"andhuman","children":[{"author":"yoavm","children":[{"author":"Havoc","children":[{"author":"mesebrec","children":[{"author":"CamperBob2","children":[{"author":"Terretta","children":[],"created_at":"2025-12-05T14:57:20.000Z","created_at_i":1764946640,"id":46162124,"options":[],"parent_id":46127453,"points":null,"story_id":46121889,"text":"You&#x27;d have to ask EU&#x27;s regulators why they wanted Meta to disallow it.<p>Much like you&#x27;d have to ask UK lawmakers why they wanted UK citizens to be unable to keep their own Apple iCloud backups secure.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T21:55:54.000Z","created_at_i":1764712554,"id":46127453,"options":[],"parent_id":46124908,"points":null,"story_id":46121889,"text":"Why does it disallow usage in the EU?","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T18:48:39.000Z","created_at_i":1764701319,"id":46124908,"options":[],"parent_id":46122579,"points":null,"story_id":46121889,"text":"Llama&#x27;s license explicitly disallows its usage in the EU.<p>If that doesn&#x27;t even meet the threshold for &quot;terrible&quot;, then what does?","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:01:10.000Z","created_at_i":1764691270,"id":46122579,"options":[],"parent_id":46122403,"points":null,"story_id":46121889,"text":"Guessing GP commenter considers Apache more &quot;open&quot; than Meta&#x27;s license. Which to be fair isn&#x27;t terrible but also not quite as clean as straight apache","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T15:47:16.000Z","created_at_i":1764690436,"id":46122403,"options":[],"parent_id":46122214,"points":null,"story_id":46121889,"text":"How is this different from Llama 3.2 &quot;vision capabilities&quot;?<p><a href=\"https:&#x2F;&#x2F;www.llama.com&#x2F;docs&#x2F;how-to-guides&#x2F;vision-capabilities&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;www.llama.com&#x2F;docs&#x2F;how-to-guides&#x2F;vision-capabilities...</a>","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T15:32:35.000Z","created_at_i":1764689555,"id":46122214,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"This is big. The first really big open weights model that understands images.","title":null,"type":"comment","url":null},{"author":"Tiberium","children":[],"created_at":"2025-12-02T15:40:26.000Z","created_at_i":1764690026,"id":46122310,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"A bit interesting that they used Deepseek 3&#x27;s architecture for their Large model :)","title":null,"type":"comment","url":null},{"author":"GaggiX","children":[],"created_at":"2025-12-02T15:44:51.000Z","created_at_i":1764690291,"id":46122372,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"The small dense model seems particularly good for their small sizes, I can&#x27;t wait to test them out.","title":null,"type":"comment","url":null},{"author":"tucnak","children":[{"author":"NitpickLawyer","children":[],"created_at":"2025-12-02T16:05:27.000Z","created_at_i":1764691527,"id":46122622,"options":[],"parent_id":46122391,"points":null,"story_id":46121889,"text":"&gt; I wonder why scores on TriviaQA vis-a-vis 14b model lags behind Gemma 12b so much; that one is not a formatting-heavy benchmark.<p>My guess is the vast scale of google data. They&#x27;ve been hoovering data for decades now, and have had curation pipelines (guided by real human interactions) since forever.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T15:46:19.000Z","created_at_i":1764690379,"id":46122391,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"If the claims on multilingual and pretraining performance are accurate, this is huge! This may be the best-in-class multilingual stuff since the more recent Gemma&#x27;s, where they used to be unmatched. I know Americans don&#x27;t care much about the rest of the world, but we&#x27;re still using our native tongues thank you very much; there is a huge issue with i.e. Ukrainian (as opposed to Russian) being underrepresented in many open-weight and weight-available models. Gemma used to be a notable exception, I wonder if it&#x27;s still the case. On a different note: I wonder why scores on TriviaQA vis-a-vis 14b model lags behind Gemma 12b so much; that one is not a formatting-heavy benchmark.","title":null,"type":"comment","url":null},{"author":"arnaudsm","children":[{"author":"jasonjmcghee","children":[{"author":"gishh","children":[{"author":"rdtsc","children":[],"created_at":"2025-12-02T16:43:11.000Z","created_at_i":1764693791,"id":46123128,"options":[],"parent_id":46123028,"points":null,"story_id":46121889,"text":"I always joke that Google pays for a dedicated developer to spend their full time just to make pelicans on bicycles look good. They certainly have the cash to do it.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:35:16.000Z","created_at_i":1764693316,"id":46123028,"options":[],"parent_id":46122741,"points":null,"story_id":46121889,"text":"Gamed tests?","title":null,"type":"comment","url":null},{"author":"arnaudsm","children":[{"author":"netdur","children":[],"created_at":"2025-12-02T19:28:48.000Z","created_at_i":1764703728,"id":46125533,"options":[],"parent_id":46123240,"points":null,"story_id":46121889,"text":"I believe it is the system instructions that make the difference for Gemini, as I use Gemini on AI Studio with my system prompts to get it to do what I need it to do, which is not possible with gemini.google.com&#x27;s gems","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:50:30.000Z","created_at_i":1764694230,"id":46123240,"options":[],"parent_id":46122741,"points":null,"story_id":46121889,"text":"Could be optimized for benchmarks, but Gemini 3 has been stellar for my tasks so far.<p>Maybe an architectural leap?","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:15:33.000Z","created_at_i":1764692133,"id":46122741,"options":[],"parent_id":46122566,"points":null,"story_id":46121889,"text":"How is there such a gap between Gemini 3 vs GPT 5.1&#x2F;Opus 4.5? What is Gemini 3 crushing the others on?","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:00:11.000Z","created_at_i":1764691211,"id":46122566,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"Geometric mean of MMMLU + GPQA-Diamond + SimpleQA + LiveCodeBench :<p>- Gemini 3.0 Pro         : 84.8<p>- DeepSeek 3.2           : 83.6<p>- GPT-5.1                : 69.2<p>- Claude Opus 4.5        : 67.4<p>- Kimi-K2 (1.2T)         : 42.0<p>- Mistral Large 3 (675B) : 41.9<p>- Deepseek-3.1 (670B)    : 39.7<p>The 14B 8B &amp; 3B models are SOTA though, and do not have chinese censorship like Qwen3.","title":null,"type":"comment","url":null},{"author":"barrell","children":[{"author":"metadat","children":[{"author":"barrell","children":[{"author":"barbazoo","children":[{"author":"barrell","children":[{"author":"sandblast","children":[],"created_at":"2025-12-02T17:07:09.000Z","created_at_i":1764695229,"id":46123482,"options":[],"parent_id":46123364,"points":null,"story_id":46121889,"text":"XD XD","title":null,"type":"comment","url":null},{"author":"barbazoo","children":[{"author":"barrell","children":[],"created_at":"2025-12-02T18:11:28.000Z","created_at_i":1764699088,"id":46124343,"options":[],"parent_id":46124025,"points":null,"story_id":46121889,"text":"Heh it&#x27;s a quote from Archer FX (and admittedly a poor machine translation, it&#x27;s a very old expression of mine).<p>And yes, this only happens when I ask it to apply my formatting rules. If you let GPT format itself, I would be surprised if this ever happens.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T17:45:52.000Z","created_at_i":1764697552,"id":46124025,"options":[],"parent_id":46123364,"points":null,"story_id":46121889,"text":"Surely reads like someone&#x27;s brain transformed into a tree :)<p>Impressive, I haven&#x27;t seen that myself yet, I&#x27;ve only used 5 conversationally, not via API yet.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:59:43.000Z","created_at_i":1764694783,"id":46123364,"options":[],"parent_id":46123243,"points":null,"story_id":46121889,"text":"If you wanted examples, you needed only ask :)<p>These are screenshots from that week: <a href=\"https:&#x2F;&#x2F;x.com&#x2F;barrelltech&#x2F;status&#x2F;1995900100174880806\" rel=\"nofollow\">https:&#x2F;&#x2F;x.com&#x2F;barrelltech&#x2F;status&#x2F;1995900100174880806</a><p>I&#x27;m not going to share the prompt because (1) it&#x27;s very long (2) there were dozens of variations and (3) it seems like poor business practices to share the most indefensible part of your business online XD","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:50:40.000Z","created_at_i":1764694240,"id":46123243,"options":[],"parent_id":46123194,"points":null,"story_id":46121889,"text":"Hard to gauge what gibberish is without an example of the data and what you prompted the LLM with.","title":null,"type":"comment","url":null},{"author":"data-ottawa","children":[{"author":"barrell","children":[],"created_at":"2025-12-02T18:25:09.000Z","created_at_i":1764699909,"id":46124511,"options":[],"parent_id":46123798,"points":null,"story_id":46121889,"text":"Reasoning was set to minimal and low (and I think I tried medium at some point). I do not believe the timeouts were due to the reasoning taking to long, although I never streamed the results. I think the model just fails often. It stops producing tokens and eventually the request times out.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T17:28:16.000Z","created_at_i":1764696496,"id":46123798,"options":[],"parent_id":46123194,"points":null,"story_id":46121889,"text":"With gpt5 did you try adjusting the reasoning level to &quot;minimal&quot;?<p>I tried using it for a very small and quick summarization task that needed low latency and any level above that took several seconds to get a response. Using minimal brought that down significantly.<p>Weirdly gpt5&#x27;s reasoning levels don&#x27;t map to the OpenAI api level reasoning effort levels.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:47:33.000Z","created_at_i":1764694053,"id":46123194,"options":[],"parent_id":46122968,"points":null,"story_id":46121889,"text":"Yes. I spent about 3 days trying to optimize the prompt to get gpt-5 to not produce gibberish, to no avail. Completions took several minutes, had an above 50% timeout rate (with a 6 minute timeout mind you), and after retrying they still would return gibberish about 15% of the time (12% on one task, 20% on another task).<p>I then tried multiple models, and they all failed in spectacular ways. Only Grok and Mistral had an acceptable success rate, although Grok did not follow the formatting instructions as well as Mistral.<p>Phrasing is a language learning application, so the formatting is very complicated, with multiple languages and multiple scripts intertwined with markdown formatting. I do include dozens of examples in the prompts, but it&#x27;s something many models struggle with.<p>This was a few months ago, so to be fair, it&#x27;s possible gpt-5.1 or gemini-3 or the new deepseek model may have caught up. I have not had the time or need to compare, as Mistral has been sufficient for my use cases.<p>I mean, I&#x27;d love to get that 0.1% error rate down, but there have always more pressing issues XD","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:31:23.000Z","created_at_i":1764693083,"id":46122968,"options":[],"parent_id":46122612,"points":null,"story_id":46121889,"text":"Are you saying gpt-5 produces gibberish 15% of the time?  Or are you comparing Mistral gibberish production rate to gpt-5.1&#x27;s complex task failure rate?<p>Does Mistral even have a Tool Use model?  That would be awesome to have a new coder entrant beyond OpenAI, Anthropic, Grok, and Qwen.","title":null,"type":"comment","url":null},{"author":"mrtksn","children":[{"author":"barbazoo","children":[{"author":"viking123","children":[],"created_at":"2025-12-03T07:43:12.000Z","created_at_i":1764747792,"id":46131340,"options":[],"parent_id":46123229,"points":null,"story_id":46121889,"text":"For me it&#x27;s just that I am too lazy to start switching from my GPT subscription, I use it with codex and it&#x27;s very good for my use-case. And the price at least here in Asia is not expensive at all for the plus tier. The amount of tokens are so much that I usually cannot even spend the weekly quota, although I use context smartly and know my codebase so I can always point it to right place right away.<p>I feel like at least for normies if they are familiar with ChatGPT, it might be hard to make them switch especially if they are subscribed.","title":null,"type":"comment","url":null},{"author":"b3ing","children":[],"created_at":"2025-12-03T13:54:38.000Z","created_at_i":1764770078,"id":46134483,"options":[],"parent_id":46123229,"points":null,"story_id":46121889,"text":"I estimate at 10% of meetup runs like that","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:49:39.000Z","created_at_i":1764694179,"id":46123229,"options":[],"parent_id":46123090,"points":null,"story_id":46121889,"text":"&gt; I guess they hope I forget to cancel.<p>Business model of most subscription based services.","title":null,"type":"comment","url":null},{"author":"barrell","children":[{"author":"distalx","children":[{"author":"amy_petrik","children":[],"created_at":"2025-12-03T23:09:55.000Z","created_at_i":1764803395,"id":46141580,"options":[],"parent_id":46139400,"points":null,"story_id":46121889,"text":"usually either use Grok to optimize a mistral prompt, or you can use gemini to optimize a chatGPT prompt.  It&#x27;s best to keep those pairs of AIs and not cross streams!","title":null,"type":"comment","url":null}],"created_at":"2025-12-03T20:09:46.000Z","created_at_i":1764792586,"id":46139400,"options":[],"parent_id":46123249,"points":null,"story_id":46121889,"text":"What tools or process do you use to optimize your prompts?","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:51:18.000Z","created_at_i":1764694278,"id":46123249,"options":[],"parent_id":46123090,"points":null,"story_id":46121889,"text":"Yep I spent 3 days optimizing my prompt trying to get gpt-5 to work. Tried a bunch of different models (some Azure some OpenRouter) and got a better success rate with several others without any tailoring of the prompt.<p>Was really plug and play. There are still small nuances to each one, but compared to a year ago prompts are much more portable","title":null,"type":"comment","url":null},{"author":"acuozzo","children":[{"author":"mrtksn","children":[],"created_at":"2025-12-02T18:49:06.000Z","created_at_i":1764701346,"id":46124911,"options":[],"parent_id":46124380,"points":null,"story_id":46121889,"text":"my use case is Google replacement, things that I can do by myself so I can verify and things that are not important so I don\u2019t have to verify.<p>Sure, they produce different output so sometimes I will run the same thing on a few different models when Im not sure or happy but I\u2019d don\u2019t delegate the thinking part actually, I always give a direction in my prompts. I don\u2019t see myself running 30min queries because I will never trust the output and will have to do all the work myself. Instead I like to go step by step together.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T18:14:44.000Z","created_at_i":1764699284,"id":46124380,"options":[],"parent_id":46123090,"points":null,"story_id":46121889,"text":"&gt; because they are interchangeable<p>What is your use-case?<p>Mine is: I use &quot;Pro&quot;&#x2F;&quot;Max&quot;&#x2F;&quot;DeepThink&quot; models to iterate on novel cross-domain applications of existing mathematics.<p>My interaction is: I craft a detailed prompt in my editor, hand it off, come back 20-30 minutes later, review the reply, and then repeat if necessary.<p>My experience is that they&#x27;re all very, very different from one another.","title":null,"type":"comment","url":null},{"author":"giancarlostoro","children":[{"author":"mrtksn","children":[{"author":"ecommerceguy","children":[{"author":"giancarlostoro","children":[],"created_at":"2025-12-03T02:47:33.000Z","created_at_i":1764730053,"id":46129717,"options":[],"parent_id":46128601,"points":null,"story_id":46121889,"text":"Oh man I use Comet nearly daily, I tried setting perplexity as my new tab page on other browsers and for some reason its not the same. I mostly use it that boring way too.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T23:58:09.000Z","created_at_i":1764719889,"id":46128601,"options":[],"parent_id":46124941,"points":null,"story_id":46121889,"text":"I use their browser called Comet for finance related research. Very nice. I use pretty much all of the main ai&#x27;s, chat, deep, gem, claude - all i have found little niche use case that i&#x27;m sure will rotate at some point in an upgrade cycle. there are so many ai&#x27;s i don&#x27;t see the point in paying for one. I&#x27;m convinced they will need ads to survive.<p>excited to add mistral to the rotation!","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T18:51:00.000Z","created_at_i":1764701460,"id":46124941,"options":[],"parent_id":46124386,"points":null,"story_id":46121889,"text":"I like perplexity actually but haven\u2019t been using it since some time. Maybe I should give it a go :)","title":null,"type":"comment","url":null},{"author":"VHRanger","children":[],"created_at":"2025-12-03T02:57:00.000Z","created_at_i":1764730620,"id":46129790,"options":[],"parent_id":46124386,"points":null,"story_id":46121889,"text":"Kagi has Mistral as well","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T18:15:03.000Z","created_at_i":1764699303,"id":46124386,"options":[],"parent_id":46123090,"points":null,"story_id":46121889,"text":"Maybe give Perplexity a shot? It has Grok, ChatGPT, Gemini, Kimi K2, I dont think it has Mistral unfortunately.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:39:56.000Z","created_at_i":1764693596,"id":46123090,"options":[],"parent_id":46122612,"points":null,"story_id":46121889,"text":"Some time ago I canceled all my paid subscriptions to chatbots because they are interchangeable so I just rotate between Grok, ChatGPT, Gemini, Deepseek and Mistral.<p>On the API side of things my experience is that the model behaving as expected is the greatest feature.<p>There I also switched to Openrouter instead of paying directly so I can use whatever model fits best.<p>The recent buzz about ad-based chatbot services is probably because the companies no longer have an edge despite what the benchmarks say, users are noticing it and cancel paid plans. Just today OpenAI offered me 1 month free trial as if I wasn\u2019t using it two months ago. I guess they hope I forget to cancel.","title":null,"type":"comment","url":null},{"author":"druskacik","children":[{"author":"leobg","children":[{"author":"leobg","children":[],"created_at":"2025-12-03T11:01:57.000Z","created_at_i":1764759717,"id":46133036,"options":[],"parent_id":46131910,"points":null,"story_id":46121889,"text":"Answering my own question:<p>Artificial Analysis ranks them close in terms of price (both 0.3 USD&#x2F;1M tokens) and intelligence (27 &#x2F; 29 for gemini&#x2F;mistral), but ranks gemini-2.0-flash-lite higher in terms of speed (189 tokens&#x2F;s vs. 130).<p>So they should be interchangeable. Looking forward to testing this.<p>[0] <a href=\"https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;?models=o3%2Cgemini-2-5-pro%2Cgemini-2-5-flash-lite-preview-09-2025%2Cmistral-small-3-2%2Cdeepseek-r1%2Cgrok-3-mini-reasoning%2Cllama-3-1-nemotron-ultra-253b-v1-reasoning%2Cgpt-4o%2Cgpt-4-1%2Co4-mini%2Cgemini-2-0-flash-lite-preview%2Cgemini-2-0-flash%2Cgemini-2-5-flash-lite-reasoning%2Cgemini-2-0-flash-lite-001%2Cgemini-2-5-flash-reasoning%2Cgrok-3&amp;cost=cost-vs-intelligence#intelligence-vs-price#intelligence-vs-cost-to-run-artificial-analysis-intelligence-index:~:text=Intelligence%20vs.%20Cost%20to%20Run%20Artificial%20Analysis%20Intelligence%20Index\" rel=\"nofollow\">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;?models=o3%2Cgemini-2-5-pro%2C...</a>","title":null,"type":"comment","url":null},{"author":"druskacik","children":[],"created_at":"2025-12-03T22:14:45.000Z","created_at_i":1764800085,"id":46140968,"options":[],"parent_id":46131910,"points":null,"story_id":46121889,"text":"I did some vibe-evals only and it seemed slightly worse for my use case, so I didn&#x27;t change it.","title":null,"type":"comment","url":null}],"created_at":"2025-12-03T08:45:05.000Z","created_at_i":1764751505,"id":46131910,"options":[],"parent_id":46123473,"points":null,"story_id":46121889,"text":"Did you compare it to gemini-2.0-flash-lite?","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T17:06:41.000Z","created_at_i":1764695201,"id":46123473,"options":[],"parent_id":46122612,"points":null,"story_id":46121889,"text":"This is my experience as well. Mistral models may not be the best according to benchmarks and I don&#x27;t use them for personal chats or coding, but for simple tasks with pre-defined scope (such as categorization, summarization, etc.) they are the option I choose. I use <i>mistral-small</i> with batch API and it&#x27;s probably the best cost-efficient option out there.","title":null,"type":"comment","url":null},{"author":"mentalgear","children":[{"author":"barrell","children":[{"author":"basilgohar","children":[{"author":"barrell","children":[],"created_at":"2025-12-02T18:06:33.000Z","created_at_i":1764698793,"id":46124278,"options":[],"parent_id":46124036,"points":null,"story_id":46121889,"text":"Thank you :) and you&#x27;re definitely not the only one.<p>Full transparency, the first backend version of phrasing was &#x27;vibe-coded&#x27; (long before vibe coding was a thing). I didn&#x27;t like the results, I didn&#x27;t like the experience, I didn&#x27;t feel good ethically, and I didn&#x27;t like my own development.<p>I rewrote the application (completely, from scratch, new repo new language new framework) and all of the sudden I liked the results, I loved the process, I had no moral qualms, and I improved leaps and bounds in all areas I worked on.<p>Automation has some amazing use cases (I am building an automation product at the end of the day) but so does doing hard things yourself.<p>Although most important is just to enjoy what you do; or perhaps do something you can be proud of.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T17:46:27.000Z","created_at_i":1764697587,"id":46124036,"options":[],"parent_id":46123904,"points":null,"story_id":46121889,"text":"I admire and respect this stance. I have been very AI-hesitant and while I&#x27;m using it more and more, I have spaces that I want to definitely keep human-only, as this is my preference. I&#x27;m glad to hear I&#x27;m not the only one like this.","title":null,"type":"comment","url":null},{"author":"willlma","children":[],"created_at":"2025-12-07T05:13:17.000Z","created_at_i":1765084397,"id":46179326,"options":[],"parent_id":46123904,"points":null,"story_id":46121889,"text":"It&#x27;s interesting. I&#x27;ve been tinkering with an article summarizing&#x2F;highlighting browser extension, and realized that I don&#x27;t want the end-user to have read AI-generated content because it&#x27;s not as high-quality as I&#x27;d hoped. But on the flip side, I&#x27;m loving having the AI write most of the code for me.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T17:35:56.000Z","created_at_i":1764696956,"id":46123904,"options":[],"parent_id":46123799,"points":null,"story_id":46121889,"text":"I don&#x27;t see the contention. I do not use llms in the design, development, copywriting, marketing, blogging, or any other aspect of the crafting of the application.<p>I labor over every word, every button, every line of code, every blog post. I would say it is as hand-crafted as something digital can be.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T17:28:16.000Z","created_at_i":1764696496,"id":46123799,"options":[],"parent_id":46122612,"points":null,"story_id":46121889,"text":"Thanks for sharing your use case of the mistral models, which are indeed top-notch ! I had a look at phrasing.app, and while a nice website, I found the copy of &quot;Hand-crafted. Phrasing was designed &amp; developed by humans, for humans.&quot; somewhat of a false virtue given your statements here of advanced lllm usage.","title":null,"type":"comment","url":null},{"author":"mbowcut2","children":[{"author":"pants2","children":[{"author":"airstrike","children":[{"author":"pants2","children":[],"created_at":"2025-12-02T20:00:01.000Z","created_at_i":1764705601,"id":46126011,"options":[],"parent_id":46124388,"points":null,"story_id":46121889,"text":"Generally, the easiest:<p>1. Sample a set of prompts &#x2F; answers from historical usage.<p>2. Run that through various frontier models again and if they don&#x27;t agree on some answers, hand-pick what you&#x27;re looking for.<p>3. Test different models using OpenRouter and score each along cost &#x2F; speed &#x2F; accuracy dimensions against your test set.<p>4. Analyze the results and pick the best, then prompt-optimize to make it even better. Repeat as needed.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T18:15:11.000Z","created_at_i":1764699311,"id":46124388,"options":[],"parent_id":46124242,"points":null,"story_id":46121889,"text":"If you and others have any insights to share on structuring that benchmark, I&#x27;m all ears.<p>There a new model seemingly every week so finding a way to evaluate them repeatedly would be nice.<p>The answer may be that it&#x27;s so bespoke you have to handroll every time, but my gut says there&#x27;s a set of best practiced that are generally applicable.","title":null,"type":"comment","url":null},{"author":"dotancohen","children":[{"author":"pants2","children":[{"author":"dotancohen","children":[{"author":"pants2","children":[{"author":"dotancohen","children":[{"author":"pants2","children":[],"created_at":"2025-12-05T19:44:54.000Z","created_at_i":1764963894,"id":46166312,"options":[],"parent_id":46164434,"points":null,"story_id":46121889,"text":"Nothing off the top of my head! If you find anything good let me know. GRPO is a training technique likely not exactly what you&#x27;d do for benchmarking, but it&#x27;s interesting to read about anyway. Glad I cuold help","title":null,"type":"comment","url":null}],"created_at":"2025-12-05T17:30:50.000Z","created_at_i":1764955850,"id":46164434,"options":[],"parent_id":46163875,"points":null,"story_id":46121889,"text":"Thank you. I will google Group Relative Policy Optimization to learn about that and the other training methods. If you have any resources handy that I should be reading, that would be appreciated as well. Have a great weekend.","title":null,"type":"comment","url":null}],"created_at":"2025-12-05T16:53:51.000Z","created_at_i":1764953631,"id":46163875,"options":[],"parent_id":46157848,"points":null,"story_id":46121889,"text":"Yeah - things are easy when you can objectively score an output, otherwise as you said you&#x27;ll probably need another LLM to score it. For summaries you <i>can</i> try to make that somewhat more objective, like length and &quot;8&#x2F;10 key points are covered in this summary.&quot;<p>This is a real training method (like Group Relative Policy Optimization), so it&#x27;s a legitimate approach.","title":null,"type":"comment","url":null}],"created_at":"2025-12-05T07:59:06.000Z","created_at_i":1764921546,"id":46157848,"options":[],"parent_id":46156941,"points":null,"story_id":46121889,"text":"Thank you! I&#x27;ll see about building a test suite.<p>Do you compare models&#x27; output subjectively, manually? Or do you have some objective measures? My use case would be to test diagnostic information summaries - the output is free text, not structured. The only way I can think to automate that would be with another LLM.<p>Advice welcome!","title":null,"type":"comment","url":null}],"created_at":"2025-12-05T04:49:38.000Z","created_at_i":1764910178,"id":46156941,"options":[],"parent_id":46133837,"points":null,"story_id":46121889,"text":"Just grab the top ~30 models on OpenRouter[1] and test them all. If that&#x27;s too expensive make a sample &#x27;screening&#x27; benchmark that&#x27;s just a few of the hardest problems to see if it&#x27;s even worth the full benchmark.<p>1. <a href=\"https:&#x2F;&#x2F;openrouter.ai&#x2F;models?order=top-weekly&amp;fmt=table\" rel=\"nofollow\">https:&#x2F;&#x2F;openrouter.ai&#x2F;models?order=top-weekly&amp;fmt=table</a>","title":null,"type":"comment","url":null}],"created_at":"2025-12-03T12:42:57.000Z","created_at_i":1764765777,"id":46133837,"options":[],"parent_id":46124242,"points":null,"story_id":46121889,"text":"How do you find and decide which obscure models to test? Do you manually review the model card for each new model on Hugging Face? Is there a better resource?","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T18:02:49.000Z","created_at_i":1764698569,"id":46124242,"options":[],"parent_id":46124029,"points":null,"story_id":46121889,"text":"The best benchmark is one that you build for your use-case. I finally did that for a project and I was not expecting the results. Frontier models are generally &quot;good enough&quot; for most use-cases but if you have something specific you&#x27;re optimizing for there&#x27;s probably a more obscure model that just does a better job.","title":null,"type":"comment","url":null},{"author":"pembrook","children":[{"author":"astrange","children":[],"created_at":"2025-12-03T18:16:14.000Z","created_at_i":1764785774,"id":46137924,"options":[],"parent_id":46125401,"points":null,"story_id":46121889,"text":"Americans have an opposing bias via the phenomenon of &quot;safe edgy&quot;, where for obvious reasons they&#x27;re uncomfortable with being biased towards anyone who looks like a US minority, and redirect all that energy towards being racist to the French. So it&#x27;s all balanced.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T19:19:15.000Z","created_at_i":1764703155,"id":46125401,"options":[],"parent_id":46124029,"points":null,"story_id":46121889,"text":"If the models from the big US labs are being overfit to benchmarks, than we also need to account for HN commenters overfitting positive evaluations to Chinese or European models based on their political biases (US big tech = default bad, anything European = default good).<p>Also, we should be aware of people cynically playing into that bias to try to advertise their app, like OP who has managed to spam a link in the first line of a top comment on this popular front page article by telling the audience exactly what they want to hear ;)","title":null,"type":"comment","url":null},{"author":"Legend2440","children":[],"created_at":"2025-12-02T20:53:08.000Z","created_at_i":1764708788,"id":46126698,"options":[],"parent_id":46124029,"points":null,"story_id":46121889,"text":"I don\u2019t think benchmark overfitting is as common as people think. Benchmark scores are highly correlated with the subjective \u201cintelligence\u201d of the model. So is pretraining loss.<p>The only exception I can think of is models trained on synthetic data like Phi.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T17:46:04.000Z","created_at_i":1764697564,"id":46124029,"options":[],"parent_id":46122612,"points":null,"story_id":46121889,"text":"It makes me wonder about the gaps in evaluating LLMs by benchmarks. There almost certainly is overfitting happening which could degrade other use cases. &quot;In practice&quot; evaluation is what inspired the Chatbot Arena right? But then people realized that Chatbot arena over-prioritizes formatting, and maybe sycophancy(?). Makes you wonder what the best evaluation would be. We probably need lots more task-specific models. That&#x27;s seemed to be fruitful for improved coding.","title":null,"type":"comment","url":null},{"author":"acuozzo","children":[{"author":"barrell","children":[{"author":"acuozzo","children":[{"author":"barrell","children":[],"created_at":"2025-12-02T19:50:28.000Z","created_at_i":1764705028,"id":46125871,"options":[],"parent_id":46125339,"points":null,"story_id":46121889,"text":"If I cannot tolerate a failure rate, I do not use LLMs (or and ML models).<p>But in that case the larger the better. If mistral medium can run on your M2 Ultra then it should be up to the task. Should eek out ministral and be just shy of the biggest frontier models.<p>But I wouldn\u2019t even trust GPT-5 or Claude Opus or Gemini 3 Pro to get close to a zero percent success rate, and for a task such as this I would not expect mistral medium to outperform the big boys","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T19:15:09.000Z","created_at_i":1764702909,"id":46125339,"options":[],"parent_id":46124427,"points":null,"story_id":46121889,"text":"I&#x27;d prefer for the error rate to be as close to 0% as possible under the strict requirement of having to use a local model. I have access to nodes with 8xH200, but I&#x27;d prefer to not tie those up with this task. I&#x27;d, instead, prefer to use a model I can run on an M2 Ultra.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T18:18:04.000Z","created_at_i":1764699484,"id":46124427,"options":[],"parent_id":46124304,"points":null,"story_id":46121889,"text":"What&#x27;s your acceptable error rate? Honestly ministral would probably be sufficient if you can tolerate a small failure rate. I feel like medium would be overkill.<p>But I&#x27;m no expert. I can&#x27;t say I&#x27;ve used mistral much outside of my own domain.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T18:08:54.000Z","created_at_i":1764698934,"id":46124304,"options":[],"parent_id":46122612,"points":null,"story_id":46121889,"text":"I have a need to remove loose &quot;signature&quot; lines from the last 10% of a tremendous e-mail dataset. Based on your experience, how do you think mistral-3-medium-0525 would do?","title":null,"type":"comment","url":null},{"author":"mackross","children":[],"created_at":"2025-12-03T12:32:40.000Z","created_at_i":1764765160,"id":46133765,"options":[],"parent_id":46122612,"points":null,"story_id":46121889,"text":"Cool app. I couldn\u2019t see a way to report an error in one of the default expressions.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:04:42.000Z","created_at_i":1764691482,"id":46122612,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"I use large language models in <a href=\"http:&#x2F;&#x2F;phrasing.app\" rel=\"nofollow\">http:&#x2F;&#x2F;phrasing.app</a> to format data I can retrieve in a consistent skimmable manner. I switched to mistral-3-medium-0525 a few months back after struggling to get gpt-5 to stop producing gibberish. It&#x27;s been insanely fast, cheap, reliable, and follows formatting instructions to the letter. I was (and still am) super super impressed. Even if it does not hold up in benchmarks, it still outperformed in practice.<p>I&#x27;m not sure how these new models compare to the biggest and baddest models, but if price, speed, and reliability are a concern for your use cases I cannot recommend Mistral enough.<p>Very excited to try out these new models! To be fair, mistral-3-medium-0525 still occasionally produces gibberish ~0.1% of my use cases (vs gpt-5&#x27;s 15% failure rate). Will report back if that goes up or down with these new models","title":null,"type":"comment","url":null},{"author":"esafak","children":[{"author":"nullbio","children":[],"created_at":"2025-12-02T17:01:30.000Z","created_at_i":1764694890,"id":46123399,"options":[],"parent_id":46122613,"points":null,"story_id":46121889,"text":"Benchmarks are never to be believed, and that has been the case since day 1.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:04:46.000Z","created_at_i":1764691486,"id":46122613,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"Well done to the France&#x27;s Mistral team for closing the gap. If the benchmarks are to be believed, this is a viable model, especially at the edge.","title":null,"type":"comment","url":null},{"author":"mythz","children":[{"author":"rvz","children":[{"author":"crimsoneer","children":[{"author":"kergonath","children":[],"created_at":"2025-12-03T08:05:32.000Z","created_at_i":1764749132,"id":46131495,"options":[],"parent_id":46122768,"points":null,"story_id":46121889,"text":"&gt; I would be shocked if there isn&#x27;t some French gov funding somewhere in the massive mistral pile<p>There is a bit of it, yes, although how much exactly is difficult to know. It\u2019s not all tax breaks and subventions; several public agencies are using it, including in the army so finding out the details is not trivial.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:17:43.000Z","created_at_i":1764692263,"id":46122768,"options":[],"parent_id":46122717,"points":null,"story_id":46121889,"text":"I mean, one is a government, the other are VCs (also, I would be <i>shocked</i> if there isn&#x27;t some French gov funding somewhere in the massive mistral pile).","title":null,"type":"comment","url":null},{"author":"whiplash451","children":[{"author":"apexalpha","children":[{"author":"didibus","children":[{"author":"JumpCrisscross","children":[{"author":"didibus","children":[],"created_at":"2025-12-03T03:19:03.000Z","created_at_i":1764731943,"id":46129921,"options":[],"parent_id":46125113,"points":null,"story_id":46121889,"text":"Interesting, is that still the case? And how is the decision to take those high risk investments made for things like pensions and such?","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T19:01:05.000Z","created_at_i":1764702065,"id":46125113,"options":[],"parent_id":46123054,"points":null,"story_id":46121889,"text":"&gt; <i>and people with too much money?</i><p>No. VC\u2019s historical capital has come from institutional investors. Pensions. Endowments. Foundations.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:37:02.000Z","created_at_i":1764693422,"id":46123054,"options":[],"parent_id":46122947,"points":null,"story_id":46121889,"text":"For VC don&#x27;t you need a lot of capital and people with too much money?<p>Isn&#x27;t that then a chicken and egg?","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:30:18.000Z","created_at_i":1764693018,"id":46122947,"options":[],"parent_id":46122847,"points":null,"story_id":46121889,"text":"1. Big problem<p>2. ASML was propped up by ASM and Philips, stepping in as &quot;VCs&quot;","title":null,"type":"comment","url":null},{"author":"rvz","children":[],"created_at":"2025-12-02T16:36:10.000Z","created_at_i":1764693370,"id":46123048,"options":[],"parent_id":46122847,"points":null,"story_id":46121889,"text":"1. It matters.<p>2. Did ASML invest in Mistral in their first round of venture funding or was it US VCs all along that took that early risk and backed them from the <i>very</i> start?<p>Risk aversion is in the DNA and in almost every plot of land in Europe such that US VCs saw something in Mistral before even the european giants like ASML did.<p>ASML would have passed on Mistral from the start and Mistral would have instead begged to the EU for a grant.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:23:07.000Z","created_at_i":1764692587,"id":46122847,"options":[],"parent_id":46122717,"points":null,"story_id":46121889,"text":"1. so what\n2. asml","title":null,"type":"comment","url":null},{"author":"amarcheschi","children":[],"created_at":"2025-12-02T16:54:59.000Z","created_at_i":1764694499,"id":46123302,"options":[],"parent_id":46122717,"points":null,"story_id":46121889,"text":"Mistral biggest investor is asml, although it became so later than other vcs","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:14:09.000Z","created_at_i":1764692049,"id":46122717,"options":[],"parent_id":46122618,"points":null,"story_id":46121889,"text":"All thanks to the US VCs that acutally have money to fund Mistral&#x27;s entire business.<p>Had they gone to the EU, Mistral would have gotten a miniscule grant from the EU  to train their AI models.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:05:05.000Z","created_at_i":1764691505,"id":46122618,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"Europe&#x27;s bright star has been quiet for a while, great to see them back and good to see them come back to Open Source light with Apache 2.0 licenses - they&#x27;re too far from the SOTA pack that exclusive&#x2F;proprietary models would work in their favor.<p>Mistral had the best small models on consumer GPUs for a while, hopefully Ministral 14B lives up to their benchmarks.","title":null,"type":"comment","url":null},{"author":"lalassu","children":[{"author":"para_parolu","children":[],"created_at":"2025-12-02T16:16:02.000Z","created_at_i":1764692162,"id":46122746,"options":[],"parent_id":46122695,"points":null,"story_id":46121889,"text":"It\u2019s not for users but for businesses. There is demand for inhouse use with data privacy.\nRegular users can\u2019t even run large model due to lack of compute.","title":null,"type":"comment","url":null},{"author":"hopelite","children":[],"created_at":"2025-12-02T16:23:52.000Z","created_at_i":1764692632,"id":46122853,"options":[],"parent_id":46122695,"points":null,"story_id":46121889,"text":"It seems to be a reasonable comparison since that is the primary&#x2F;differentiating characteristic of the model. It\u2019s really common to also and seemingly only ever see the comparison of closed weight&#x2F;proprietary models in a way that seems to act as if all of the non-American and open weight models don\u2019t even exist.<p>I also think most people do not consider open weights as OSS.","title":null,"type":"comment","url":null},{"author":"troyvit","children":[],"created_at":"2025-12-02T20:38:13.000Z","created_at_i":1764707893,"id":46126514,"options":[],"parent_id":46122695,"points":null,"story_id":46121889,"text":"Glad I&#x27;m not most users. I&#x27;m down for 80% of the quality for an open weight model. Hell I&#x27;ve been using Linux for 25 years so I suppose I&#x27;m used to not-the-greatest-but-free.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:12:18.000Z","created_at_i":1764691938,"id":46122695,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"It&#x27;s sad that they only compare to open weight models. I feel most users don&#x27;t care much about OSS&#x2F;not OSS. The value proposition is the quality of the generation for some use case.<p>I guess it says a bit about the state of European AI","title":null,"type":"comment","url":null},{"author":"s_dev","children":[{"author":"shlomo_z","children":[{"author":"s_dev","children":[],"created_at":"2025-12-02T17:45:29.000Z","created_at_i":1764697529,"id":46124019,"options":[],"parent_id":46123576,"points":null,"story_id":46121889,"text":"My critique is more levelled at Mistral and not specifically what they&#x27;ve just released so it could be that some see what I have to say as off topic.<p>Also a lot of Europeans are upset at US tech dominance. It&#x27;s a position we&#x27;ve roped ourselves in to so any commentary that criticises an EU tech success story is seen as being unnecessarily negative.<p>However I do mean it as a warning to others, I got burned even with good intentions.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T17:13:29.000Z","created_at_i":1764695609,"id":46123576,"options":[],"parent_id":46122778,"points":null,"story_id":46121889,"text":"This seems like a legitimate complaint... I wonder why it&#x27;s downvoted","title":null,"type":"comment","url":null},{"author":"cycomanic","children":[{"author":"s_dev","children":[],"created_at":"2025-12-03T09:07:42.000Z","created_at_i":1764752862,"id":46132118,"options":[],"parent_id":46126231,"points":null,"story_id":46121889,"text":"&gt;This sounds like the you expect your subscription to work as an on-demand service?<p>That&#x27;s exactly what it is.<p>&gt;I&#x27;m not sure I understand you correctly,<p>I understand perfectly well, I don&#x27;t agree with that approach is the issue.<p>If I paid for 11&#x2F;12 months I should get 11&#x2F;12 months subscription not 1&#x2F;12 months. They happily just took a years subscription and provided nothing in return. Even if I fixed the outstanding balance they would have provided 2&#x2F;12 months of service at a cost of 12&#x2F;12 months of payment.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T20:15:28.000Z","created_at_i":1764706528,"id":46126231,"options":[],"parent_id":46122778,"points":null,"story_id":46121889,"text":"I&#x27;m not sure I understand you correctly, but it seems you had a subscription missed one payment some time ago, but now expect that your subscription works because the missed month was in the past and &quot;you paid for this month&quot;?<p>This sounds like the you expect your subscription to work as an on-demand service? It seems quite obvious that to be able to use a service you would need to be up to date on your payments, that would be no different in any other subscription&#x2F;lease&#x2F;rental agreement? Now Mistral might certainly look back at their records and see that you actually didn&#x27;t use their service at all for the last few month and waive the missed payment. And that could be good customer service, but they might not even have record that you didn&#x27;t use it, or at least those records would not be available to the billing department?","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:18:24.000Z","created_at_i":1764692304,"id":46122778,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"I was subscribing to these guys purely to support the EU tech scene. So I was on Pro for about 2 years while using ChatGPT and Claude.<p>Went to actually use it, got a message saying that I missed a payment 8 months previously and thus wasn&#x27;t allowed to use Pro despite having paid for Pro for the previous 8 months. The lady I contacted in support simply told me to pay the outstanding balance. You would think if you missed a payment it would relate to simply that month that was missed not all subsequent months.<p>Utterly ridiculous that one missed payment can justify not providing the service (otherwise paid for in full) at all.<p>Basically if you find yourself in this situation you&#x27;re actually better of deleting the account and resigning up again under a different email.<p>We really need to get our shit together in the EU on this sort of stuff, I was a paying customer purely out of sympathy but that sympathy dried up pretty quick with hostile customer service.","title":null,"type":"comment","url":null},{"author":"jasonjmcghee","children":[],"created_at":"2025-12-02T16:19:26.000Z","created_at_i":1764692366,"id":46122788,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"I wish they showed how they compared to models larger&#x2F;better and what the gap is, rather than only models they&#x27;re better than.<p>Like how does 14B compare to Qwen30B-A3B?<p>(Which I think is a lot of people&#x27;s goto or it&#x27;s instruct&#x2F;coding variant, from what I&#x27;ve seen in local model circles)","title":null,"type":"comment","url":null},{"author":"another_twist","children":[{"author":"Rastonbury","children":[],"created_at":"2025-12-02T16:32:55.000Z","created_at_i":1764693175,"id":46122994,"options":[],"parent_id":46122797,"points":null,"story_id":46121889,"text":"Age aside, not sure what Zuck was thinking, seeing as Scale AI was in data labelling and not training models, perhaps he thought he was a good operator? Then again the talent scarcity is in scientists, there are many operators, let alone one worth 14B. Back to age, the people he is managing are likely all several years older than him and Meta long timers, which would make it even more challenging","title":null,"type":"comment","url":null},{"author":"vintagedave","children":[{"author":"another_twist","children":[],"created_at":"2025-12-03T07:10:39.000Z","created_at_i":1764745839,"id":46131116,"options":[],"parent_id":46129026,"points":null,"story_id":46121889,"text":"True no one involved in Scale AI right now is a kid. But, their expertise is in data labelling not cutting edge AI. Compare that to the Mistral team. They launched a new LLM within 6months of founding. They&#x27;re also ex-Meta researchers. But they dont have the distribution coz europe. If we want to tout 13B acqusitions and 100m pay packages, Mistral is the perfect candidate. Its basically plug and play. Compare that to Scale and the shitshow that ensued. MSL lost talent and have to start from scratch given that their head knows nothing about LLMs.","title":null,"type":"comment","url":null}],"created_at":"2025-12-03T01:00:49.000Z","created_at_i":1764723649,"id":46129026,"options":[],"parent_id":46122797,"points":null,"story_id":46121889,"text":"What is this referring to? I googled and the company was founded in 2016. No one involved can to a \u201ckid\u201d?","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:20:10.000Z","created_at_i":1764692410,"id":46122797,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"I am not sure why Meta paid 13B+ to hire some kid vs just hiring back or acquiring these folks. They&#x27;ll easily catch up.","title":null,"type":"comment","url":null},{"author":"msp26","children":[{"author":"make3","children":[],"created_at":"2025-12-02T19:50:03.000Z","created_at_i":1764705003,"id":46125864,"options":[],"parent_id":46123008,"points":null,"story_id":46121889,"text":"Architecture difference wrt vanilla transformers and between modern transformers are a tiny part of what makes a model nowadays","title":null,"type":"comment","url":null},{"author":"Jackson__","children":[{"author":"Ey7NFZ3P0nzAe","children":[],"created_at":"2025-12-02T21:12:29.000Z","created_at_i":1764709949,"id":46126936,"options":[],"parent_id":46126832,"points":null,"story_id":46121889,"text":"Well, behind &quot;models&quot; not &quot;langual models&quot;.<p>Of course models purely made for image stuff will completely wipe it out. The vision language models are useful for their generalist capabilities","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T21:01:54.000Z","created_at_i":1764709314,"id":46126832,"options":[],"parent_id":46123008,"points":null,"story_id":46121889,"text":"So they spent all of their R&amp;D to copy deepseek, leaving none for the singular novel added feature: vision.<p>To quote the hf page:<p>&gt;Behind vision-first models in multimodal tasks: Mistral Large 3 can lag behind models optimized for vision tasks and use cases.","title":null,"type":"comment","url":null},{"author":"halJordan","children":[],"created_at":"2025-12-02T21:43:22.000Z","created_at_i":1764711802,"id":46127282,"options":[],"parent_id":46123008,"points":null,"story_id":46121889,"text":"I don&#x27;t think it&#x27;s fair to demand everything be open and then get mad when they open-ness is used. It&#x27;s an obsessive and harmful double standard.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:33:46.000Z","created_at_i":1764693226,"id":46123008,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"The new large model uses DeepseekV2 architecture. 0 mention on the page lol.<p>It&#x27;s a good thing that open source models use the best arch available. K2 does the same but at least mentions &quot;Kimi K2 was designed to further scale up Moonlight, which employs an architecture similar to DeepSeek-V3&quot;.<p>---<p>vllm&#x2F;model_executor&#x2F;models&#x2F;mistral_large_3.py<p>```<p>from vllm.model_executor.models.deepseek_v2 import DeepseekV3ForCausalLM<p>class MistralLarge3ForCausalLM(DeepseekV3ForCausalLM):<p>```<p>&quot;Science has always thrived on openness and shared discovery.&quot; btw<p>Okay I&#x27;ll stop being snarky now and try the 14B model at home. Vision is good additional functionality on Large.","title":null,"type":"comment","url":null},{"author":"tootyskooty","children":[],"created_at":"2025-12-02T16:52:16.000Z","created_at_i":1764694336,"id":46123261,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"Since no one has mentioned it yet: note that the benchmarks for large are for the base model, not for the instruct model available in the API.<p>Most likely reason is that the instruct model underperforms compared to the open competition (even among non-reasoners like Kimi K2).","title":null,"type":"comment","url":null},{"author":"nullbio","children":[{"author":"apexalpha","children":[{"author":"Synthetic7346","children":[],"created_at":"2025-12-02T17:17:36.000Z","created_at_i":1764695856,"id":46123642,"options":[],"parent_id":46123390,"points":null,"story_id":46121889,"text":"I found gemini 3 to be pretty lackluster for setting up an onprem k8s cluster - sonnet 4.5 was more accurate from the get go, required less handholding","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T17:00:56.000Z","created_at_i":1764694856,"id":46123390,"options":[],"parent_id":46123345,"points":null,"story_id":46121889,"text":"No, I&#x27;ve been using Gemini for help while learning &#x2F; building my onprem k8s cluster and it has been almost spotless.<p>Granted, this is a subject that is very well present in the training data but still.","title":null,"type":"comment","url":null},{"author":"alfalfasprout","children":[],"created_at":"2025-12-02T17:01:10.000Z","created_at_i":1764694870,"id":46123394,"options":[],"parent_id":46123345,"points":null,"story_id":46121889,"text":"If anything it&#x27;s a testament to human intelligence that benchmarks haven&#x27;t really been a good measure of a model&#x27;s competence for some time now. They provide a relative sorting to some degree, within model families, but it feels like we&#x27;ve hit an AI winter.","title":null,"type":"comment","url":null},{"author":"mvkel","children":[{"author":"barrell","children":[{"author":"theshrike79","children":[],"created_at":"2025-12-03T12:18:14.000Z","created_at_i":1764764294,"id":46133655,"options":[],"parent_id":46123477,"points":null,"story_id":46121889,"text":"In my use cases mistral has been next to useless.<p>Granted my uses have been programming related. Mistral prints the answer almost immediately and is also completely and utterly hallucinating everything and producing just something that looks like code but could never even compile...","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T17:06:49.000Z","created_at_i":1764695209,"id":46123477,"options":[],"parent_id":46123411,"points":null,"story_id":46121889,"text":"I can attest to Mistral beating OpenAI in my use cases pretty definitively :)","title":null,"type":"comment","url":null},{"author":"re-thc","children":[{"author":"lowkey_","children":[{"author":"mvkel","children":[],"created_at":"2025-12-03T00:42:13.000Z","created_at_i":1764722533,"id":46128897,"options":[],"parent_id":46124103,"points":null,"story_id":46121889,"text":"I continue to be surprised that the supposed bastion of &quot;safe&quot; AI, anthropic, has a record of being the least-open AI company","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T17:51:44.000Z","created_at_i":1764697904,"id":46124103,"options":[],"parent_id":46123726,"points":null,"story_id":46121889,"text":"Not the above poster, but:<p>OpenAI went closed (despite open literally being in the name) once they had the advantage. Meta also is going closed now that they&#x27;ve caught up.<p>Open-source makes sense to accelerate to catch up, but once ahead, closed will come back to retain advantage.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T17:22:55.000Z","created_at_i":1764696175,"id":46123726,"options":[],"parent_id":46123411,"points":null,"story_id":46121889,"text":"&gt; Open weight LLMs aren&#x27;t supposed to &quot;beat&quot; closed models, and they never will. That isn\u2019t their purpose.<p>Do things ever work that way? What if Google did Open source Gemini. Would you say the same? You never know. There&#x27;s never &quot;supposed&quot; and &quot;purpose&quot; like that.","title":null,"type":"comment","url":null},{"author":"cmrdporcupine","children":[{"author":"troyvit","children":[],"created_at":"2025-12-02T20:35:42.000Z","created_at_i":1764707742,"id":46126478,"options":[],"parent_id":46124051,"points":null,"story_id":46121889,"text":"I think you&#x27;re right, and I feel the same about Mistral. It&#x27;s &quot;good enough&quot;, super cheap, privacy friendly, and doesn&#x27;t burn coal by the shovel-full. No need to pay through the nose for the SOTA models just to get wrapped into the same SaaS games that plague the rest of the industry.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T17:47:37.000Z","created_at_i":1764697657,"id":46124051,"options":[],"parent_id":46123411,"points":null,"story_id":46121889,"text":"This may be the case, but DeepSeek 3.2 is &quot;good enough&quot; that it competes well with Sonnet 4 -- maybe 4.5 -- for about 80% of my use cases, at a fraction of the cost.<p>I feel we&#x27;re only a year or two away from hitting a plateau with the frontier closed models having diminishing returns vs what&#x27;s &quot;open&quot;","title":null,"type":"comment","url":null},{"author":"pants2","children":[{"author":"array_key_first","children":[],"created_at":"2025-12-03T01:08:18.000Z","created_at_i":1764724098,"id":46129078,"options":[],"parent_id":46124332,"points":null,"story_id":46121889,"text":"It kind of does, because the proprietary systems are unacceptable for many usecases because they are proprietary.<p>There&#x27;s a lot of businesses who do not want to hand over their sensitive data to hackers, employees of their competitors, and various world governments. There&#x27;s inherent risk in choosing a propreitary option, and that doesn&#x27;t just go for LLMs. You can get your feet swept up from underneath you.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T18:10:54.000Z","created_at_i":1764699054,"id":46124332,"options":[],"parent_id":46123411,"points":null,"story_id":46121889,"text":"&gt; Their value is as a structural check on the power of proprietary systems<p>Unfortunately that doesn&#x27;t pay the electricity bill","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T17:02:49.000Z","created_at_i":1764694969,"id":46123411,"options":[],"parent_id":46123345,"points":null,"story_id":46121889,"text":"Open weight LLMs aren&#x27;t supposed to &quot;beat&quot; closed models, and they never will. That isn\u2019t their purpose. Their value is as a structural check on the power of proprietary systems; they guarantee a competitive floor. They\u2019re essential to the ecosystem, but they\u2019re not chasing SOTA.","title":null,"type":"comment","url":null},{"author":"mrtksn","children":[{"author":"cmrdporcupine","children":[{"author":"erichocean","children":[],"created_at":"2025-12-07T20:04:45.000Z","created_at_i":1765137885,"id":46184636,"options":[],"parent_id":46124077,"points":null,"story_id":46121889,"text":"I exclusively use Gemini Pro for coding, and it&#x27;s been writing ~100% of the code I produce since July.<p>It&#x27;s great.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T17:49:39.000Z","created_at_i":1764697779,"id":46124077,"options":[],"parent_id":46123820,"points":null,"story_id":46121889,"text":"I think a lot of the hype around Gemini comes down to people who aren&#x27;t using it for coding but for other things maybe.<p>Frankly, I don&#x27;t actually care about or want &quot;general intelligence&quot; -- I want it to make good code, follow instructions, and find bugs. Gemini wasn&#x27;t bad at the last bit, but wasn&#x27;t great at the others.<p>They&#x27;re all trying to make general purpose AI, but I just want really smart augmentation &#x2F; tools.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T17:29:38.000Z","created_at_i":1764696578,"id":46123820,"options":[],"parent_id":46123345,"points":null,"story_id":46121889,"text":"Yep, Gemini is my least favorite and I\u2019m convinced that the hype around it isn\u2019t organic because I don\u2019t see the claimed \u201csuperiority\u201d, quite the opposite.","title":null,"type":"comment","url":null},{"author":"minimaxir","children":[],"created_at":"2025-12-02T17:30:49.000Z","created_at_i":1764696649,"id":46123833,"options":[],"parent_id":46123345,"points":null,"story_id":46121889,"text":"For noncoding tasks, Gemini atleast allows for easier grounding with Google Search.","title":null,"type":"comment","url":null},{"author":"bluecalm","children":[],"created_at":"2025-12-02T17:31:15.000Z","created_at_i":1764696675,"id":46123840,"options":[],"parent_id":46123345,"points":null,"story_id":46121889,"text":"My experience is the opposite although I don&#x27;t use it to write code but to explore&#x2F;learn about algorithms and various programming ideas. It&#x27;s amazing. I am close to cancelling my ChatGPT subscription (I would only use Open Router if it had nicer GUI and dark mode anyway).","title":null,"type":"comment","url":null},{"author":"cmrdporcupine","children":[],"created_at":"2025-12-02T17:37:43.000Z","created_at_i":1764697063,"id":46123927,"options":[],"parent_id":46123345,"points":null,"story_id":46121889,"text":"I also had bad luck when I finally tried Gemini 3 in the gemini CLI coding tool. I am unclear if it&#x27;s the model or their bad tooling&#x2F;prompting. It had, as you said, hallucination problems, and it also had memory issues where it seemed to drop context between prompts here and there.<p>It&#x27;s also slower than both Opus 4.5 and Sonnet.","title":null,"type":"comment","url":null},{"author":"llm_nerd","children":[],"created_at":"2025-12-02T17:39:57.000Z","created_at_i":1764697197,"id":46123958,"options":[],"parent_id":46123345,"points":null,"story_id":46121889,"text":"What does your comment have to do with the submission? What a weird non-sequitur. I even went looking at the linked article to see if it somehow compares with Gemini. It doesn&#x27;t, and only relates to open models.<p>In prior posts you oddly attack &quot;Palantir-partnered Anthropic&quot; as well.<p>Are things that grim at OpenAI that this sort of FUD is necessary? I mean, I know they&#x27;re doing the whole code red thing, but I guarantee that posting nonsense like this on HN isn&#x27;t the way.","title":null,"type":"comment","url":null},{"author":"dchest","children":[],"created_at":"2025-12-02T18:11:26.000Z","created_at_i":1764699086,"id":46124341,"options":[],"parent_id":46123345,"points":null,"story_id":46121889,"text":"Nope, Gemini 3 is hallucinating less than GPT-5.1 for my questions.","title":null,"type":"comment","url":null},{"author":"moffkalast","children":[],"created_at":"2025-12-02T18:15:37.000Z","created_at_i":1764699337,"id":46124393,"options":[],"parent_id":46123345,"points":null,"story_id":46121889,"text":"Yes, and likewise with Kimi K2. Despite being on the top of open source benches it makes up more batshit nonsense than even Llama 3.<p>Trust no one, test your use case yourself is pretty much the only approach, because people either don&#x27;t run benchmarks correctly or have the incentive not to.","title":null,"type":"comment","url":null},{"author":"tootie","children":[],"created_at":"2025-12-02T18:26:21.000Z","created_at_i":1764699981,"id":46124536,"options":[],"parent_id":46123345,"points":null,"story_id":46121889,"text":"No? My recent experience with Gemini was terrific. The last big test I gave of Claude it spun an immaculate web of lies before I forced it to confess.","title":null,"type":"comment","url":null},{"author":"VeejayRampay","children":[],"created_at":"2025-12-03T04:49:42.000Z","created_at_i":1764737382,"id":46130409,"options":[],"parent_id":46123345,"points":null,"story_id":46121889,"text":"no, I find Gemini to be the best","title":null,"type":"comment","url":null},{"author":"gunalx","children":[],"created_at":"2025-12-03T08:56:44.000Z","created_at_i":1764752204,"id":46132046,"options":[],"parent_id":46123345,"points":null,"story_id":46121889,"text":"Have used gemini3 to GEW shot a few problems GPT5 struggled on.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T16:58:16.000Z","created_at_i":1764694696,"id":46123345,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"Anyone else find that despite Gemini performing best on benches, it&#x27;s actually still far worse than ChatGPT and Claude? It seems to hallucinate nonsense far more frequently than any of the others. Feels like Google just bench maxes all day every day. As for Mistral, hopefully OSS can eat all of their lunch soon enough.","title":null,"type":"comment","url":null},{"author":"trvz","children":[{"author":"ThrowawayTestr","children":[{"author":"nikcub","children":[{"author":"ThrowawayTestr","children":[],"created_at":"2025-12-02T21:28:42.000Z","created_at_i":1764710922,"id":46127108,"options":[],"parent_id":46125303,"points":null,"story_id":46121889,"text":"I meant more how do they pay for all that bandwidth. I can download a 20gb model in like 2 minutes","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T19:13:08.000Z","created_at_i":1764702788,"id":46125303,"options":[],"parent_id":46124134,"points":null,"story_id":46121889,"text":"s3 + cloudfront<p><a href=\"https:&#x2F;&#x2F;huggingface.co&#x2F;blog&#x2F;rearchitecting-uploads-and-downloads\" rel=\"nofollow\">https:&#x2F;&#x2F;huggingface.co&#x2F;blog&#x2F;rearchitecting-uploads-and-downl...</a>","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T17:54:11.000Z","created_at_i":1764698051,"id":46124134,"options":[],"parent_id":46123382,"points":null,"story_id":46121889,"text":"How does HF manage to serve such big files?","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T17:00:24.000Z","created_at_i":1764694824,"id":46123382,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"Sad to see they&#x27;ve apparently fully given up on releasing their models via torrent magnet URLs shared on Twitter; those will stay around long after Hugging Face is dead.","title":null,"type":"comment","url":null},{"author":"dmezzetti","children":[],"created_at":"2025-12-02T17:28:43.000Z","created_at_i":1764696523,"id":46123804,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"Looking forward to trying them out. Great to see they are Apache 2.0...always good to have easy-to-understand licensing.","title":null,"type":"comment","url":null},{"author":"RomanPushkin","children":[],"created_at":"2025-12-02T17:30:52.000Z","created_at_i":1764696652,"id":46123835,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"Mistral presented DeepSeek 3.2","title":null,"type":"comment","url":null},{"author":"ThrowawayTestr","children":[],"created_at":"2025-12-02T17:33:06.000Z","created_at_i":1764696786,"id":46123863,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"Awesome! Can&#x27;t wait till someone abliterates them.","title":null,"type":"comment","url":null},{"author":"simonw","children":[{"author":"troyvit","children":[{"author":"GaggiX","children":[{"author":"embedding-shape","children":[],"created_at":"2025-12-02T21:30:37.000Z","created_at_i":1764711037,"id":46127129,"options":[],"parent_id":46126050,"points":null,"story_id":46121889,"text":"Nor would it be describing things as they happen, but instead needing pre-processing, so in the end, very different :)","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T20:02:21.000Z","created_at_i":1764705741,"id":46126050,"options":[],"parent_id":46125693,"points":null,"story_id":46121889,"text":"This is not local but Gemini models can process very long videos and provide description with timestamps if asked for.<p><a href=\"https:&#x2F;&#x2F;ai.google.dev&#x2F;gemini-api&#x2F;docs&#x2F;video-understanding#transcribe-video\" rel=\"nofollow\">https:&#x2F;&#x2F;ai.google.dev&#x2F;gemini-api&#x2F;docs&#x2F;video-understanding#tr...</a>","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T19:38:37.000Z","created_at_i":1764704317,"id":46125693,"options":[],"parent_id":46123976,"points":null,"story_id":46121889,"text":"I&#x27;m reading this post and wondering what kind of crazy accessibility tools one could make. I think it&#x27;s a little off the rails but imagine a tool that describes a web video for a blind user as it happens, not just the speech, but the actual action.","title":null,"type":"comment","url":null},{"author":"user_of_the_wek","children":[],"created_at":"2025-12-03T07:35:30.000Z","created_at_i":1764747330,"id":46131292,"options":[],"parent_id":46123976,"points":null,"story_id":46121889,"text":"&gt; The image depicts and older man...<p>Ouch","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T17:41:10.000Z","created_at_i":1764697270,"id":46123976,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"The 3B vision model runs in the browser (after a 3GB model download). There&#x27;s a very cool demo of that here: <a href=\"https:&#x2F;&#x2F;huggingface.co&#x2F;spaces&#x2F;mistralai&#x2F;Ministral_3B_WebGPU\" rel=\"nofollow\">https:&#x2F;&#x2F;huggingface.co&#x2F;spaces&#x2F;mistralai&#x2F;Ministral_3B_WebGPU</a><p>Pelicans are OK but not earth-shattering: <a href=\"https:&#x2F;&#x2F;simonwillison.net&#x2F;2025&#x2F;Dec&#x2F;2&#x2F;introducing-mistral-3&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;simonwillison.net&#x2F;2025&#x2F;Dec&#x2F;2&#x2F;introducing-mistral-3&#x2F;</a>","title":null,"type":"comment","url":null},{"author":"RYJOX","children":[],"created_at":"2025-12-02T18:03:18.000Z","created_at_i":1764698598,"id":46124246,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"I find that there are too many paid sub models at the minute with non legitimate progress to warrant the money spent. Recently cancelled GPT.","title":null,"type":"comment","url":null},{"author":"domoritz","children":[],"created_at":"2025-12-02T18:04:56.000Z","created_at_i":1764698696,"id":46124261,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"Urg, the bar charts to not start at 0. It&#x27;s making it impossible to compare across model sizes. That&#x27;s a pretty basic chart design principle. I hope they can fix it. At least give me consistent y scales!","title":null,"type":"comment","url":null},{"author":"mrinterweb","children":[{"author":"hiddencost","children":[],"created_at":"2025-12-03T05:11:26.000Z","created_at_i":1764738686,"id":46130538,"options":[],"parent_id":46125893,"points":null,"story_id":46121889,"text":"Idk. They look like they&#x27;re ahead on the saturated benchmarks and behind on the unsaturated ones. Looks more like that over fit to the benchmarks.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T19:51:56.000Z","created_at_i":1764705116,"id":46125893,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"I don&#x27;t like being this guy, but I think Deepseek 3.2 stole all the thunder yesterday. Notice that these comparisons are to Deepseek 3.1. Deepseek 3.2 is a big step up over 3.1, if benchmarks are to be believed. Just unfortunate timing of release. <a href=\"https:&#x2F;&#x2F;api-docs.deepseek.com&#x2F;news&#x2F;news251201\" rel=\"nofollow\">https:&#x2F;&#x2F;api-docs.deepseek.com&#x2F;news&#x2F;news251201</a>","title":null,"type":"comment","url":null},{"author":"Aissen","children":[{"author":"dloss","children":[],"created_at":"2025-12-02T21:05:18.000Z","created_at_i":1764709518,"id":46126865,"options":[],"parent_id":46126045,"points":null,"story_id":46121889,"text":"Yes, the 3B variant, with vLLM 0.11.2. Parameters are given on the HF page. Had to override the temperature to 0.15 though (as suggested on HF) to avoid random looking syllables.","title":null,"type":"comment","url":null},{"author":"Patrick_Devine","children":[],"created_at":"2025-12-02T21:59:13.000Z","created_at_i":1764712753,"id":46127508,"options":[],"parent_id":46126045,"points":null,"story_id":46121889,"text":"The instruct models are available on Ollama (e.g. `ollama run ministral-3:8b`), however the reasoning models still are a wip. I was trying to get them to work last night and it works for single turn, but is still very flakey w&#x2F; multi-turn.","title":null,"type":"comment","url":null},{"author":"Aissen","children":[],"created_at":"2025-12-03T15:43:04.000Z","created_at_i":1764776584,"id":46135771,"options":[],"parent_id":46126045,"points":null,"story_id":46121889,"text":"It now seems to work with the latest vLLM git.","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T20:02:03.000Z","created_at_i":1764705723,"id":46126045,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"Anyone succeed in running it with vLLM?","title":null,"type":"comment","url":null},{"author":"tmaly","children":[{"author":"PhilippGille","children":[],"created_at":"2025-12-02T23:32:59.000Z","created_at_i":1764718379,"id":46128395,"options":[],"parent_id":46126204,"points":null,"story_id":46121889,"text":"<a href=\"https:&#x2F;&#x2F;openrouter.ai&#x2F;mistralai&#x2F;mistral-large-2512\" rel=\"nofollow\">https:&#x2F;&#x2F;openrouter.ai&#x2F;mistralai&#x2F;mistral-large-2512</a>","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T20:13:29.000Z","created_at_i":1764706409,"id":46126204,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"I see several 3.x versions on Openrouter.ai, any idea which of those are the new models?","title":null,"type":"comment","url":null},{"author":"RandyOrion","children":[],"created_at":"2025-12-03T06:27:29.000Z","created_at_i":1764743249,"id":46130875,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"Thank you Mistral for releasing new small parameter-efficient (aka dense) models.","title":null,"type":"comment","url":null},{"author":"mortsnort","children":[],"created_at":"2025-12-03T16:00:56.000Z","created_at_i":1764777656,"id":46136034,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"I use a small model as a chatbot of sorts in a game I&#x27;m making. I was hoping the 3b could replace qwen 4b, but it&#x27;s far worse at following instructions and providing entertaining content. I suppose this is expected given smaller size and their own benchmarks that show Qwen beating it at instruct.","title":null,"type":"comment","url":null},{"author":"accrual","children":[],"created_at":"2025-12-03T16:15:11.000Z","created_at_i":1764778511,"id":46136226,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"Congrats on the release, Mistral team!<p>I haven&#x27;t used Mistral much until today but am impressed. I normally use Gemma 3 27B locally, but after regenerating some responses with Mistral 3 14B, the output quality is very similar despite generating much faster on my hardware.<p>The vision aspect also worked fine, and actually was slightly better on the same inputs versus qwen3 VL 8B.<p>All in all impressive small dense model, looking forward to using it more.","title":null,"type":"comment","url":null},{"author":"pixel_popping","children":[],"created_at":"2025-12-03T17:13:58.000Z","created_at_i":1764782038,"id":46137097,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"fyi Mistral admins, there is no dates showing on your article.","title":null,"type":"comment","url":null},{"author":"Frannky","children":[],"created_at":"2025-12-04T08:18:32.000Z","created_at_i":1764836312,"id":46145066,"options":[],"parent_id":46121889,"points":null,"story_id":46121889,"text":"I haven&#x27;t tried a Mistral model in ages. Llama and Mistral feel like something I was using in another era. Are they good?","title":null,"type":"comment","url":null}],"created_at":"2025-12-02T15:01:53.000Z","created_at_i":1764687713,"id":46121889,"options":[],"parent_id":null,"points":826,"story_id":46121889,"text":null,"title":"Mistral 3 family of models released","type":"story","url":"https://mistral.ai/news/mistral-3"}
