{"author":"shenli3514","children":[{"author":"minimaxir","children":[{"author":"cyanydeez","children":[{"author":"drob518","children":[],"created_at":"2026-08-29T20:49:33.000Z","created_at_i":1788036573,"id":49493211,"options":[],"parent_id":49492913,"points":null,"story_id":49492632,"text":"Of course they are. Of course they do. Nobody should be surprised by this.","title":null,"type":"comment","url":null},{"author":"tokai","children":[{"author":"realo","children":[{"author":"noir_lord","children":[],"created_at":"2026-08-29T22:03:37.000Z","created_at_i":1788041017,"id":49493719,"options":[],"parent_id":49493293,"points":null,"story_id":49492632,"text":"lobbying&#x2F;legalised bribery hard to say where one ends and another begins at times.","title":null,"type":"comment","url":null},{"author":"CamperBob2","children":[{"author":"blackqueeriroh","children":[{"author":"andrekandre","children":[],"created_at":"2026-08-30T01:50:32.000Z","created_at_i":1788054632,"id":49494942,"options":[],"parent_id":49494817,"points":null,"story_id":49492632,"text":"i mean, theres capitalism as the ideal, and there is capitalism in practice, so maybe you are both right...","title":null,"type":"comment","url":null},{"author":"CamperBob2","children":[],"created_at":"2026-08-30T02:08:02.000Z","created_at_i":1788055682,"id":49495049,"options":[],"parent_id":49494817,"points":null,"story_id":49492632,"text":"Where in the <i>Wealth of Nations</i> does a Trump appear?","title":null,"type":"comment","url":null},{"author":"realo","children":[{"author":"CamperBob2","children":[{"author":"bigyabai","children":[],"created_at":"2026-08-30T17:44:32.000Z","created_at_i":1788111872,"id":49500949,"options":[],"parent_id":49499890,"points":null,"story_id":49492632,"text":"It&#x27;s kinda funny to hear these protests. In a &quot;true fascist&quot; society, you wouldn&#x27;t be surprised to see someone killed in the name of political power. But in a &quot;true capitalist&quot; society, you&#x27;re shocked that money triumphs over virtue? It&#x27;s not a &quot;virtuist&quot; society, is it?<p>Authoritarian capitalists exist. Larry Ellison, Alex Karp, Elon Musk, Mark Zuckerberg - they <i>love</i> when you think that capitalism plays by written rules. It helps them use your tax dollars to pay for their surveillance services. It helps them sell you B2C products that degrade the moral fabric of America, and rewrite legislation through online influence campaigns. You can say it &quot;isn&#x27;t capitalism&quot; because you don&#x27;t agree with it politically, but the accrual of capital is precisely how all 4 of those men became powerful. Their influence on social media is why they&#x27;re able to reframe the national discussion how they want. Their money is what lets them turn political willpower into a business, or a rubber chicken into a politician.<p>After all, what is Reddit or 4chan if not one big capitalist influence campaign? Do you really think product reviews and political ragebait is driven by some natural and benevolent social force that we can&#x27;t see or feel?","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T16:03:17.000Z","created_at_i":1788105797,"id":49499890,"options":[],"parent_id":49498425,"points":null,"story_id":49492632,"text":"You need to go back to school and demand a refund if you think any of that is &quot;true capitalism.&quot;<p>Better yet, just go back to Reddit.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T13:22:17.000Z","created_at_i":1788096137,"id":49498425,"options":[],"parent_id":49494817,"points":null,"story_id":49492632,"text":"Oh.<p>Administration corrupted up to it&#x27;s very core?  Check.<p>Nihilism of anyone not part of the proper color, gender, whatever agenda? Check.<p>Unlawful surveillance? Check.<p>Sending totally innocent citizens to prison with many of them dying mysteriously? Check.<p>Killing innocent people in the streets simply because they dare protest peacefully? Check.<p>Welcome to North Korea!<p>Oups. Confused.<p>Welcome to the GREAT US of A! Where True Capitalism is practiced.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T01:23:52.000Z","created_at_i":1788053032,"id":49494817,"options":[],"parent_id":49493982,"points":null,"story_id":49492632,"text":"Lmao that\u2019s exactly capitalism","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T22:45:25.000Z","created_at_i":1788043525,"id":49493982,"options":[],"parent_id":49493293,"points":null,"story_id":49492632,"text":"Well, it sure as hell isn&#x27;t capitalism.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T20:59:36.000Z","created_at_i":1788037176,"id":49493293,"options":[],"parent_id":49493239,"points":null,"story_id":49492632,"text":"I would suggest &quot;lobbying&quot; is not the correct word to describe all the corruption going on in the current USA administration cesspool.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T20:53:22.000Z","created_at_i":1788036802,"id":49493239,"options":[],"parent_id":49492913,"points":null,"story_id":49492632,"text":"&gt;dont do Capitalism like the rest of the AI field<p>Like lobbying the US president to harm their competitors?","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T20:08:41.000Z","created_at_i":1788034121,"id":49492913,"options":[],"parent_id":49492750,"points":null,"story_id":49492632,"text":"i&#x27;d be curious if openrouter is just being gamed by these publishers by paying for the exposure.<p>wouldn&#x27;t trust they dont do Capitalism like the rest of the AI field.","title":null,"type":"comment","url":null},{"author":"martinald","children":[{"author":"dakolli","children":[{"author":"minimaxir","children":[{"author":"dakolli","children":[{"author":"andai","children":[{"author":"dakolli","children":[{"author":"andai","children":[{"author":"RussianCow","children":[],"created_at":"2026-08-30T16:56:50.000Z","created_at_i":1788109010,"id":49500408,"options":[],"parent_id":49498093,"points":null,"story_id":49492632,"text":"Presumably the number that OpenRouter shows is averaged across all requests.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T12:29:50.000Z","created_at_i":1788092990,"id":49498093,"options":[],"parent_id":49494296,"points":null,"story_id":49492632,"text":"But before 5m the hit rate is 100%, and after it&#x27;s 0%? Why is there a probability?<p>Is there some stochastic process that takes place during those 5 minutes that determines whether or not you get the discount?","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T23:41:52.000Z","created_at_i":1788046912,"id":49494296,"options":[],"parent_id":49494044,"points":null,"story_id":49492632,"text":"When the prefix matches a request sent to the same Providor. The thing is the TTL is different for each provider, some cache for 5 minutes some cache for 1hr. Its ideal to only use one provider per agent session &#x2F; and per model with the best cache hit % if you care about costs.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T22:53:34.000Z","created_at_i":1788044014,"id":49494044,"options":[],"parent_id":49493840,"points":null,"story_id":49492632,"text":"Wait, what does that number mean? I thought it always uses the cache price when the prefix matches.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T22:25:22.000Z","created_at_i":1788042322,"id":49493840,"options":[],"parent_id":49493754,"points":null,"story_id":49492632,"text":"You can sort by the cost cache cose, but you cannot by cache hit %. You have to click on the provider and see what their cache hit % is. A provider could have a super low cache cost, but a 50% cache hit percentage, making the cheap price of cache read&#x27;s meaningless.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T22:10:46.000Z","created_at_i":1788041446,"id":49493754,"options":[],"parent_id":49493610,"points":null,"story_id":49492632,"text":"You can click the table headers to sort Ascending&#x2F;Descending.","title":null,"type":"comment","url":null},{"author":"Bolwin","children":[{"author":"dakolli","children":[{"author":"Implicated","children":[{"author":"dakolli","children":[{"author":"RussianCow","children":[{"author":"dakolli","children":[{"author":"RussianCow","children":[],"created_at":"2026-08-30T16:48:23.000Z","created_at_i":1788108503,"id":49500323,"options":[],"parent_id":49498110,"points":null,"story_id":49492632,"text":"They&#x27;re not &quot;docking points&quot;, they&#x27;re calculating it in the most straightforward way. If I start a session and the majority of requests are sent to Provider A, and my last request gets routed to Provider B, I have a 0% cache hit rate with Provider B. I&#x27;m very curious how else you expect this to be calculated? Do you think they&#x27;re completely omitting requests that switch providers mid-session?<p>FWIW, I get significantly higher than listed cache hit rates when I pin my session to a specific provider, which is further evidence of the above.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T12:34:48.000Z","created_at_i":1788093288,"id":49498110,"options":[],"parent_id":49496259,"points":null,"story_id":49492632,"text":"Do you really think they&#x27;re docking points because cache invalidation due to provider switching? Seriously llms are frying ya&#x27;lls brain.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T06:45:28.000Z","created_at_i":1788072328,"id":49496259,"options":[],"parent_id":49494635,"points":null,"story_id":49492632,"text":"How else would you expect them to calculate it?","title":null,"type":"comment","url":null},{"author":"irthomasthomas","children":[],"created_at":"2026-08-30T09:05:18.000Z","created_at_i":1788080718,"id":49496978,"options":[],"parent_id":49494635,"points":null,"story_id":49492632,"text":"Something is up. Deepseek cache hit rate on zenmux is 98%, but only 85% via openrouter.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T00:46:16.000Z","created_at_i":1788050776,"id":49494635,"options":[],"parent_id":49494352,"points":null,"story_id":49492632,"text":"I know what they&#x27;re saying. Why would openrouter calculate it thay way lol. They obviously dont. Think for a sec, they arent idiots.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T23:52:58.000Z","created_at_i":1788047578,"id":49494352,"options":[],"parent_id":49493873,"points":null,"story_id":49492632,"text":"I think you&#x27;re arguing the same general point that the person you&#x27;re responding to is. But you&#x27;re saying he&#x27;s not understanding - he understands that they report a cache hit % but you can&#x27;t look at that public metric with any level of accuracy _because_ most people aren&#x27;t pinning their providers and they _are_ getting juggled around which is bringing that metric down. That&#x27;s not to say that specific providers might have issues or worse cache implementations - but it stands that if openrouter is juggling the requests back and forth by default then _that alone_ is breaking caches on those requests in huge numbers.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T22:30:39.000Z","created_at_i":1788042639,"id":49493873,"options":[],"parent_id":49493800,"points":null,"story_id":49492632,"text":"This is not true, there isn&#x27;t even a way to see a cache hit % model for a specific model, that wouldn&#x27;t make any sense. You are confusing what I&#x27;m saying with cache cost, that has nothing to do with effective cache hit %. I&#x27;m talking about when you click on a specific provider for a specific model, you can scroll down on the view and see their cache hit % for that model [0].<p>These cache Hit % are accurate, I&#x27;ve done a ton of testing of this myself. The cache hit % is one of the most important metrics as far as estimating cost. There are many providers with cheap cache reads, but have an effective cache hit % of 30%, making their cheaper cache pricing meaningless compared to another provider who charges more but has a 85% cache hit percentage.<p>[0]: <a href=\"https:&#x2F;&#x2F;openrouter.ai&#x2F;deepseek&#x2F;deepseek-v4-flash-0731?endpoint=f954a48e-1c90-433b-9348-4720f8030331\" rel=\"nofollow\">https:&#x2F;&#x2F;openrouter.ai&#x2F;deepseek&#x2F;deepseek-v4-flash-0731?endpoi...</a><p>scroll down on the provider&#x2F;model card and you&#x27;ll see a field called cache hit %, its different for every provider&#x2F;model.<p>I don&#x27;t use routing on openrouter, I strictly use models with a single provider and no fallback, at least for use with harnesses its pretty dumb to route requests to multiple providers you are busting your cache every other request and increasing costs by 20-50%.","title":null,"type":"comment","url":null},{"author":"andai","children":[{"author":"Implicated","children":[],"created_at":"2026-08-29T23:49:12.000Z","created_at_i":1788047352,"id":49494331,"options":[],"parent_id":49494032,"points":null,"story_id":49492632,"text":"&gt; OpenRouter randomizes which provider gets your request by default right?<p>I&#x27;m not sure it&#x27;s wholey accurate to say they &quot;randomize&quot; the provider, rather my assumption based on usage is that it&#x27;s something like cheapest-ish&#x2F;responded to the request within some reasonable-ish time&#x2F;etc algorithm that chooses the provider on each request - which seems, remarkably questionable in terms of optimizing for user experience or hidden user costs.<p>&gt; This behavior makes it so you don&#x27;t benefit much from the caching, unless you pin it to a single provider.<p>I so very much recommend this approach. My avenues that automate llm calls to openrouter are setup to make api reqs to openrouter to determine best price&#x2F;response&#x2F;etc and then pin the request to that (and, preferably, a fallback if there&#x27;s reasonable difference between #1 and #2) provider for that session. Otherwise you&#x27;re going to have a bad time.<p>I&#x27;d imagine this could make things interesting in cases where one provider is offering different quants than the others and openrouter is just swapping you back and forth on a long agentic session.","title":null,"type":"comment","url":null},{"author":"fc417fc802","children":[{"author":"RussianCow","children":[],"created_at":"2026-08-30T06:41:19.000Z","created_at_i":1788072079,"id":49496244,"options":[],"parent_id":49494638,"points":null,"story_id":49492632,"text":"This is very much NOT my experience in practice, even though it&#x27;s how I would expect it to work. OpenRouter will happily bounce you between several providers (none of which have downtime) even within the same session. Requesting specific providers is the only way I&#x27;ve been able to hit a cache rate above 90%.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T00:46:48.000Z","created_at_i":1788050808,"id":49494638,"options":[],"parent_id":49494032,"points":null,"story_id":49492632,"text":"&gt; This behavior makes it so you don&#x27;t benefit much from the caching<p>I don&#x27;t believe this is correct? AFAIK once it routes you to a provider for a given conversation that choice is sticky unless you hit technical difficulties. (It&#x27;s more complicated than that, they recently added named routing strategies that you can append to the model name.)<p>IMO the relevant metric is cache TTL which isn&#x27;t typically published AFAIK.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T22:50:52.000Z","created_at_i":1788043852,"id":49494032,"options":[],"parent_id":49493800,"points":null,"story_id":49492632,"text":"OpenRouter randomizes which provider gets your request by default right? I think you have to pass a specific provider in the request to prevent that. (Or set up a preset or something.)<p>This behavior makes it so you don&#x27;t benefit much from the caching, unless you pin it to a single provider.","title":null,"type":"comment","url":null},{"author":"ralusek","children":[{"author":"RussianCow","children":[],"created_at":"2026-08-30T16:55:46.000Z","created_at_i":1788108946,"id":49500393,"options":[],"parent_id":49494054,"points":null,"story_id":49492632,"text":"Most providers do what&#x27;s called &quot;prefix caching&quot;, where each turn in a session is cached such that sending new messages with the exact same &quot;prefix&quot; (set of previous messages) gives you the cache read price on that input instead of the full price. As long as you&#x27;re not changing your system prompt, available tools, etc mid-session, you automatically benefit from this.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T22:54:43.000Z","created_at_i":1788044083,"id":49494054,"options":[],"parent_id":49493800,"points":null,"story_id":49492632,"text":"&gt; Cache hit %<p>I thought you had to actively manage caches, do you not?","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T22:20:06.000Z","created_at_i":1788042006,"id":49493800,"options":[],"parent_id":49493610,"points":null,"story_id":49492632,"text":"Cache hit % on openrouter is not a good metric, it&#x27;s mainly driven by openrouter&#x27;s own provider juggling than the providers themselves","title":null,"type":"comment","url":null},{"author":"orbital-decay","children":[],"created_at":"2026-08-29T23:04:21.000Z","created_at_i":1788044661,"id":49494099,"options":[],"parent_id":49493610,"points":null,"story_id":49492632,"text":"&gt;Deepseek invented the paradigm of prompt caching<p>Caching was always here, you don&#x27;t need to do anything special to get it on a single user local backend running a base model or a chatbot in the first place. Among commercial providers, OpenAI adopted it in 4o first.","title":null,"type":"comment","url":null},{"author":"dominotw","children":[{"author":"gpugreg","children":[],"created_at":"2026-08-30T15:49:44.000Z","created_at_i":1788104984,"id":49499771,"options":[],"parent_id":49498305,"points":null,"story_id":49492632,"text":"Not <i>all</i> their research, but certainly a lot: <a href=\"https:&#x2F;&#x2F;github.com&#x2F;orgs&#x2F;deepseek-ai&#x2F;repositories?q=sort%3Astars\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;orgs&#x2F;deepseek-ai&#x2F;repositories?q=sort%3Ast...</a>","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T13:06:06.000Z","created_at_i":1788095166,"id":49498305,"options":[],"parent_id":49493610,"points":null,"story_id":49492632,"text":"&gt;  open sourcing all their research,<p>is this true?","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T21:43:58.000Z","created_at_i":1788039838,"id":49493610,"options":[],"parent_id":49493163,"points":null,"story_id":49492632,"text":"That&#x27;s because Deepseek invented the paradigm of prompt caching, they are the SOTA when it comes these techniques. Despite them open sourcing all their research, nobody beats them.<p>edit: I do wish openrouter would let you sort providers by Cache Hit % and Cache cost. These are the only things that matter to me at this point when choosing a provider.","title":null,"type":"comment","url":null},{"author":"sieve","children":[{"author":"sourcecodeplz","children":[{"author":"sieve","children":[],"created_at":"2026-08-30T09:40:35.000Z","created_at_i":1788082835,"id":49497140,"options":[],"parent_id":49496263,"points":null,"story_id":49492632,"text":"I have used all three extensively. DS4 Flash is quite smart. I would rank MS 1.2 below it. MiMo is the dumbest of them all but good enough for basic stuff.<p>MiMo wins handsomely if you want to think about your code for minutes at a time as you write. I use it to make changes as I think. I know it will screw up some stuff. I then switch to MS&#x2F;DS4 once every few hours and have it do a code review and fix the broken stuff. So much cheaper than getting MS to do it on its own.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T06:46:26.000Z","created_at_i":1788072386,"id":49496263,"options":[],"parent_id":49495551,"points":null,"story_id":49492632,"text":"even with the 5m cache, Muse Spark Contribs is still best bang for your buck for the intelligence you get.<p>it is basically the old dsv4-flash prices, but even more smart.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T03:55:15.000Z","created_at_i":1788062115,"id":49495551,"options":[],"parent_id":49493163,"points":null,"story_id":49492632,"text":"For me, an average long session results in about 200-300M cached input, 4-800K input, 2-400K output. Mostly the lower bound. Output depends on how much the model thinks.<p>There are two problems here:<p>- cache hit pricing (both Muse Spark 1.2 Contributor and MiMo 2.5 are around the $0.002-3&#x2F;M mark)<p>- cache persistence time<p>Muse Spark drops the cache in less than 5m. MiMo keeps it around for at least an hour based on my experience with whoever is serving it for OpenCode. This difference itself will inflate bills massively.<p>A 500K token input repeatedly read by MS 1.2 for full input price 12 times an hour = $0.60. You would be expecting $0.012. So a 50x difference. Same thing on MiMo 2.5 is $0.018 because of longer cache times.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T20:43:13.000Z","created_at_i":1788036193,"id":49493163,"options":[],"parent_id":49492750,"points":null,"story_id":49492632,"text":"I wrote about this a couple of weeks ago. It&#x27;s actually often the biggest cost and it tends to be hidden away on most platforms!<p><a href=\"https:&#x2F;&#x2F;martinalderson.com&#x2F;posts&#x2F;watch-out-for-cache-read-costs&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;martinalderson.com&#x2F;posts&#x2F;watch-out-for-cache-read-co...</a><p>Btw I still haven&#x27;t came across any decent model that is &lt;$0.01&#x2F;MTok cache costs apart from deepseek thru their official API (even with the price increases).<p>Seems like a bit of an opportunity for someone to take - drop cache read costs significantly.","title":null,"type":"comment","url":null},{"author":"Dinux","children":[],"created_at":"2026-08-29T20:45:14.000Z","created_at_i":1788036314,"id":49493177,"options":[],"parent_id":49492750,"points":null,"story_id":49492632,"text":"Which explains why almost none of my request go though","title":null,"type":"comment","url":null},{"author":"redox99","children":[{"author":"eli","children":[{"author":"redox99","children":[],"created_at":"2026-08-30T15:21:51.000Z","created_at_i":1788103311,"id":49499478,"options":[],"parent_id":49498240,"points":null,"story_id":49492632,"text":"The speed at which hy4 usage increased on openrouter, especially considering its not a cheap model, doesn&#x27;t seem organic to me.<p>Its already serving as much tokens&#x2F;day as the incredibly cheap and good GLM 5.3 flash, which had a crazy marketing campaign as ox alpha?<p>Also those top 5 apps are just 1.58B tokens out of 1.54T tokens from yesterday. Negligible.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T12:56:57.000Z","created_at_i":1788094617,"id":49498240,"options":[],"parent_id":49494066,"points":null,"story_id":49492632,"text":"Openrouter tracks what apps are using the model and the top ones for hy4 are all different coding harnesses.<p>I guess it could be fake but seems more likely people are just trying it out. Hy3 was a very strong and underrated model.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T22:57:03.000Z","created_at_i":1788044223,"id":49494066,"options":[],"parent_id":49492750,"points":null,"story_id":49492632,"text":"It&#x27;s very likely tencent games those stats, buying their own tokens.","title":null,"type":"comment","url":null},{"author":"joegibbs","children":[],"created_at":"2026-08-30T01:06:18.000Z","created_at_i":1788051978,"id":49494735,"options":[],"parent_id":49492750,"points":null,"story_id":49492632,"text":"If you\u2019re Tencent you can just plug it into some field somewhere that lots of people see right? Like how Meta could put their model on Instagram search","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T19:47:33.000Z","created_at_i":1788032853,"id":49492750,"options":[],"parent_id":49492632,"points":null,"story_id":49492632,"text":"Hy4 apparently has ludicrous traction on OpenRouter already (<a href=\"https:&#x2F;&#x2F;openrouter.ai&#x2F;tencent&#x2F;hy4-preview\" rel=\"nofollow\">https:&#x2F;&#x2F;openrouter.ai&#x2F;tencent&#x2F;hy4-preview</a>), with trillions of tokens processed in a couple days: more than GLM 5.3 in a week. That said, it&#x27;s relatively cheap with a 5% cache cost when everyone is still doing 10%&#x2F;20% cache costs, so Hy4 may be more compelling.","title":null,"type":"comment","url":null},{"author":"vcryan","children":[{"author":"Topfi","children":[{"author":"vcryan","children":[],"created_at":"2026-08-29T20:53:50.000Z","created_at_i":1788036830,"id":49493242,"options":[],"parent_id":49493102,"points":null,"story_id":49492632,"text":"Oh yes! I forgot about that. Yes, you can see this in benchmarks about hy3 preview and hy3 release still today because they measured them separately - it was significant.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T20:33:20.000Z","created_at_i":1788035600,"id":49493102,"options":[],"parent_id":49493072,"points":null,"story_id":49492632,"text":"In my evals, I saw an unprecedented jump between preview and final release on Hy3, from unusable to competitive. Did you see similar in preview vs release version?","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T20:26:48.000Z","created_at_i":1788035208,"id":49493072,"options":[],"parent_id":49492632,"points":null,"story_id":49492632,"text":"I used Hy3 quite a bit for the type of tasks it was suited for. Excited about this. My one concern over Hy3 was speed. In theory, it could be served much faster as a smaller model but it was relatively slow everywhere I could get it (including from Tencent directly) but also several other inference providers.","title":null,"type":"comment","url":null},{"author":"usernomdeguerre","children":[{"author":"feynmanquest","children":[],"created_at":"2026-08-29T20:57:24.000Z","created_at_i":1788037044,"id":49493277,"options":[],"parent_id":49493233,"points":null,"story_id":49492632,"text":"Noticed that as well","title":null,"type":"comment","url":null},{"author":"alanfranz","children":[],"created_at":"2026-08-29T21:10:04.000Z","created_at_i":1788037804,"id":49493363,"options":[],"parent_id":49493233,"points":null,"story_id":49492632,"text":"Probably AI generated.<p>But, what bars are clearly off? I couldn&#x27;t spot any.","title":null,"type":"comment","url":null},{"author":"pixelesque","children":[],"created_at":"2026-08-29T21:46:21.000Z","created_at_i":1788039981,"id":49493621,"options":[],"parent_id":49493233,"points":null,"story_id":49492632,"text":"Looks okay to me.<p>The first column has both the Hy4 and Hy3 scores overlaid on one another (Hy4 is darker blue and the taller one), with both scores written below the top of the respective bar - maybe you&#x27;re seeing that?","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T20:52:09.000Z","created_at_i":1788036729,"id":49493233,"options":[],"parent_id":49492632,"points":null,"story_id":49492632,"text":"is it just me or are the bar charts in the blog post strange? Higher numbers don&#x27;t seem to correspond correctly to their actual height?","title":null,"type":"comment","url":null},{"author":"jorl17","children":[{"author":"alexfortin","children":[],"created_at":"2026-08-30T02:04:39.000Z","created_at_i":1788055479,"id":49495030,"options":[],"parent_id":49493279,"points":null,"story_id":49492632,"text":"For the last few days I&#x27;ve been experimenting with the _free_ version of Hy3 offered by Opencode Go and I was also surprised to see how (relatively) good it is on coding tasks too.<p>The free quota from Opencode Go is also surprisingly generous, I perhaps hit limits one or two times and I&#x27;ve been using it _a lot_ for implementation tasks (using e.g. GLM-5.3-flash for working on specs and planning next steps).","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T20:57:39.000Z","created_at_i":1788037059,"id":49493279,"options":[],"parent_id":49492632,"points":null,"story_id":49492632,"text":"I experimented with Hy3 for a project and was surprised with how good it was. I don&#x27;t know if it&#x27;s good for coding, but as a general purpose agentic model, it was only beaten by deepseek4-flash in our tests. It was so close to deepseek behaviour I kept thinking it must have been forked from it.","title":null,"type":"comment","url":null},{"author":"Zigurd","children":[{"author":"tokai","children":[{"author":"nozzlegear","children":[{"author":"Rexxar","children":[{"author":"nozzlegear","children":[],"created_at":"2026-08-30T17:30:30.000Z","created_at_i":1788111030,"id":49500796,"options":[],"parent_id":49496843,"points":null,"story_id":49492632,"text":"Oh I had not heard about this, thanks for the link.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T08:36:42.000Z","created_at_i":1788079002,"id":49496843,"options":[],"parent_id":49493780,"points":null,"story_id":49492632,"text":"He just had an accident on &quot;Vuelta a Espa\u00f1a&quot; : <a href=\"https:&#x2F;&#x2F;www.theguardian.com&#x2F;sport&#x2F;2026&#x2F;aug&#x2F;29&#x2F;cycling-tadej-pogacar-pulls-out-vuelta-espana-crash\" rel=\"nofollow\">https:&#x2F;&#x2F;www.theguardian.com&#x2F;sport&#x2F;2026&#x2F;aug&#x2F;29&#x2F;cycling-tadej-...</a>","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T22:16:34.000Z","created_at_i":1788041794,"id":49493780,"options":[],"parent_id":49493360,"points":null,"story_id":49492632,"text":"\u00bfComo?","title":null,"type":"comment","url":null},{"author":"wiether","children":[],"created_at":"2026-08-30T05:38:57.000Z","created_at_i":1788068337,"id":49495996,"options":[],"parent_id":49493360,"points":null,"story_id":49492632,"text":"Reading OP&#x27;s analogy I was like &quot;even him don&#x27;t need this bike now...&quot;<p>I feel bad for him as a human, but as a cycling fan I&#x27;m glad that we&#x27;ll have an interesting WC in Canada","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T21:09:54.000Z","created_at_i":1788037794,"id":49493360,"options":[],"parent_id":49493314,"points":null,"story_id":49492632,"text":"A spanish rock solved that problem for free.","title":null,"type":"comment","url":null},{"author":"RGS1811","children":[],"created_at":"2026-08-29T21:14:06.000Z","created_at_i":1788038046,"id":49493391,"options":[],"parent_id":49493314,"points":null,"story_id":49492632,"text":"For me personally, the answer is no. Fable is adequate to do basically anything I want to do. My perspective, broadly speaking, is that we&#x27;ve saturated most of the benchmarks because we&#x27;ve largely saturated our capacity to verify models&#x27; work at scale. What&#x27;s left is context-bound verification, i.e. the problem of ensuring that output matches intent and ambiguities in prompting were resolved correctly. Further advances in autonomy do not make that latter verification problem easier. If anything they make it harder as the output per task becomes more complex and therefore more taxing for a human to verify.<p>The solution to that (to my mind) would be not a better model but a basic shift in architecture beyond the current paradigm and into a setup where agents have durable, plastic memories and undergo contextual individuation over time. But at that point agents start to become quasi-persons and not tools.","title":null,"type":"comment","url":null},{"author":"_factor","children":[],"created_at":"2026-08-29T21:14:58.000Z","created_at_i":1788038098,"id":49493402,"options":[],"parent_id":49493314,"points":null,"story_id":49492632,"text":"Hardware debugging and firmware details lead to thinking&#x2F;testing loops on all but the frontier here.","title":null,"type":"comment","url":null},{"author":"comex","children":[{"author":"Zigurd","children":[],"created_at":"2026-08-29T22:50:12.000Z","created_at_i":1788043812,"id":49494029,"options":[],"parent_id":49493534,"points":null,"story_id":49492632,"text":"This reply is particularly interesting to me because most of my experience with actually using LLMs to get work done is with coding agents. But I only have a fairly narrow set of experiences: two pretty large solo Flutter projects. I am currently really pleased with Gemini as a coding agent. It could improve, but I think improvements are going to come from marginal gains in the harness and training material so it can catch things like misconfigured permissions in platform specific areas.<p>It&#x27;s also interesting because, while coding agents are important and are a notable success, they are never going to be a multi trillion dollar business. And are there any other domains where LLMs have such a large impact?","title":null,"type":"comment","url":null},{"author":"TiredOfLife","children":[{"author":"irthomasthomas","children":[],"created_at":"2026-08-30T09:14:17.000Z","created_at_i":1788081257,"id":49497014,"options":[],"parent_id":49496702,"points":null,"story_id":49492632,"text":"One of the things that came out of the decoded reasoning paper was that Claude models had memorized answers to tests but hid this memorization from the user output and pretended to derive the answer properly. It&#x27;s only possible to cheat so blatantly in closed models where the reasoning is hidden.","title":null,"type":"comment","url":null},{"author":"happycube","children":[],"created_at":"2026-08-30T10:45:13.000Z","created_at_i":1788086713,"id":49497498,"options":[],"parent_id":49496702,"points":null,"story_id":49492632,"text":"If it were a Chinese model everyone would be screaming benchmaxxed.<p>Seriously something feels really off about Opus 5.  I hope they correct it before 4.6 is removed.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T08:06:07.000Z","created_at_i":1788077167,"id":49496702,"options":[],"parent_id":49493534,"points":null,"story_id":49492632,"text":"Opus 5 is weird. It scores high on benchmarks, but it seems that majority of those who try to use it day to day hate it","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T21:33:16.000Z","created_at_i":1788039196,"id":49493534,"options":[],"parent_id":49493314,"points":null,"story_id":49492632,"text":"My experience is that even Opus 5 still tends to write buggy or low-quality code and makes serious mistakes when analyzing code.  It&#x27;s a lot better than before but still not something I trust.  I&#x27;ve had less experience with Fable since I can&#x27;t use it at work; I hear it&#x27;s a step up but still has its limits.<p>For large tasks like a web browser or a compiler, even expensive swarms of frontier LLMs have not been shown capable of producing codebases that actually work.  (Anthropic built a C compiler with Opus 4.6 but it lacked optimizations and apparently hit a complexity wall.)<p>I also want to use LLMs for reverse engineering, but apparently it&#x27;s pretty hit-or-miss, especially if you&#x27;re forced to use open-source models to avoid restrictions.","title":null,"type":"comment","url":null},{"author":"spacebanana7","children":[{"author":"Demiurge","children":[{"author":"andybak","children":[],"created_at":"2026-08-29T22:15:12.000Z","created_at_i":1788041712,"id":49493774,"options":[],"parent_id":49493740,"points":null,"story_id":49492632,"text":"I&#x27;m getting a Poe&#x27;s Law feeling. I&#x27;m genuinely unsure about whether this post is a stone cold parody or not. I think I need to turn off the internet and go to bed.<p>EDIT: Your username doesn&#x27;t help, either.","title":null,"type":"comment","url":null},{"author":"aforwardslash","children":[],"created_at":"2026-08-29T22:57:46.000Z","created_at_i":1788044266,"id":49494070,"options":[],"parent_id":49493740,"points":null,"story_id":49492632,"text":"On that topic, check higgsfield cinema studio 4; they already provide amazing tech for the cinematic experience, somewhat similar to what you are describing.","title":null,"type":"comment","url":null},{"author":"bsenftner","children":[],"created_at":"2026-08-29T23:10:08.000Z","created_at_i":1788045008,"id":49494131,"options":[],"parent_id":49493740,"points":null,"story_id":49492632,"text":"Nobody wants to watch such films, they want to muck with the filmmaker, the generation apparatus. That&#x27;s the product, if there is one here, and absolutely not the 3 hour epic that&#x27;s spit out with 4 variations to choose between. That&#x27;s work. We&#x27;ll have other LLMs pointlessly tell us which should be watched, we&#x27;ll view a summary, and vote the Oscar on that.","title":null,"type":"comment","url":null},{"author":"aabdi","children":[],"created_at":"2026-08-30T05:44:14.000Z","created_at_i":1788068654,"id":49496018,"options":[],"parent_id":49493740,"points":null,"story_id":49492632,"text":"the problem is you have to make the AI watch the whole thing to make sure it works.<p>I&#x27;ve done this sort of with comfyui&#x2F;same agent factory stuff, but the verification loop only works for models like fable as planner&#x2F;writer, with gemini as verifier for like a very short movie. Sub 3-5 mins. After that you burn through million tokens.<p>Can&#x27;t go too low fidelity audio&#x2F;video or it craps out. Too long video and it loses consistency. Look at only snippets, it lacks global consistency, etc.","title":null,"type":"comment","url":null},{"author":"spacebanana7","children":[{"author":"CuriouslyC","children":[],"created_at":"2026-08-30T07:13:43.000Z","created_at_i":1788074023,"id":49496399,"options":[],"parent_id":49496268,"points":null,"story_id":49492632,"text":"Control nets are still useful in Krea and H3. For images, Krea can usually get close enough to a reference that it&#x27;s not a big deal, but for H3 conditioning makes a big difference over prompting for complex actions.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T06:46:55.000Z","created_at_i":1788072415,"id":49496268,"options":[],"parent_id":49493740,"points":null,"story_id":49492632,"text":"&gt; Have you seriously considered solving it?<p>Sort of, but I want it to be relatively low on human effort. I feel burned by spending lots of time in 2023 learning image generation pipelines (using control net etc) only for that to be rendered trivial by the next generation of LLMs.<p>This movie would be only for personal consumption and I\u2019m okay with waiting for model improvements.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T22:07:59.000Z","created_at_i":1788041279,"id":49493740,"options":[],"parent_id":49493556,"points":null,"story_id":49492632,"text":"That sounds like an interesting challenge. Have you seriously considered solving it? Because in about 10 seconds I came up with a process that should work, provided enough compute power. Simply model the traditional film making process by starting with a script, character stories. Design your world, then design the storyboard, and all the scenes. Create a list of all the visual elements that need to be replicated between all the scenes. Then you have to built prompts and reference art of the objects, faces, people. Make sure to do multiple takes of each scene, and have the vLLM critique and analyze the performances and technicalities. Should work?<p>I think, also, like in the traditional film makers career, this process should be built iteratively, start with a fast food commercial, then do a music video, then you can probably do a short film. Continue to improve the process, and one day I\u2019m sure the LLM film studio can make you any movie you want, provided you have enough tokens.","title":null,"type":"comment","url":null},{"author":"RobotCaleb","children":[],"created_at":"2026-08-29T22:56:10.000Z","created_at_i":1788044170,"id":49494064,"options":[],"parent_id":49493556,"points":null,"story_id":49492632,"text":"How would a computer generated video be live action?","title":null,"type":"comment","url":null},{"author":"bsenftner","children":[{"author":"spacebanana7","children":[{"author":"finebalance","children":[],"created_at":"2026-08-30T07:30:21.000Z","created_at_i":1788075021,"id":49496506,"options":[],"parent_id":49496278,"points":null,"story_id":49492632,"text":"They are different mediums. While the underlying story might be Tolkien&#x27;s, every frame is an artistic choice and while LLMs can make a choice is many situations, they are unable to 1) keep it coherent b) make it meaningful because art, to me and most, in an outcome of human experiences and thought, which by definition an llm cannot do.","title":null,"type":"comment","url":null},{"author":"bsenftner","children":[],"created_at":"2026-08-30T11:55:38.000Z","created_at_i":1788090938,"id":49497894,"options":[],"parent_id":49496278,"points":null,"story_id":49492632,"text":"Do you really want a 20 minute pause in the action, every time a new room or space is encountered, because Tolkien goes on for 2-5 pages describing every new environment like that.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T06:48:57.000Z","created_at_i":1788072537,"id":49496278,"options":[],"parent_id":49494098,"points":null,"story_id":49492632,"text":"Tolkien did the human creative work - I just want a movie adaptation that\u2019s as honest to the original text as possible. Think translating the text into video.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T23:04:09.000Z","created_at_i":1788044649,"id":49494098,"options":[],"parent_id":49493556,"points":null,"story_id":49492632,"text":"The results are boring. Not because the content is boring, but because you can so easily remix the results. Human curation is what creates value with these, not dumping and consuming. A personal perspective of a human being ups the respect, where the exact same sentences generated by an LLM carry no such value.","title":null,"type":"comment","url":null},{"author":"Zigurd","children":[],"created_at":"2026-08-29T23:10:06.000Z","created_at_i":1788045006,"id":49494130,"options":[],"parent_id":49493556,"points":null,"story_id":49492632,"text":"I think this is the best and most realistic reply so far: the ability to do this is close enough, and things like AI music are hints that there is a business model for this. Maybe I&#x27;m just jaded about CGI effects in movies currently, but I think the fact that people except that kind of thing as entertainment means you might get away with a fully AI movie that people will pay for.<p>There are two more points in favor of this kind of AI movie project: there&#x27;s zero chance that anyone would greenlight a Hollywood budget for the Silmarillion, and it is beyond human capability to write that screenplay.","title":null,"type":"comment","url":null},{"author":"clipsy","children":[],"created_at":"2026-08-30T01:44:44.000Z","created_at_i":1788054284,"id":49494924,"options":[],"parent_id":49493556,"points":null,"story_id":49492632,"text":"Since live action results are acceptable, this is already possible with current day LLMs. Just instruct one to hire a writer, director, cast, and crew to make the movie.<p>Plus, the token costs involved should be pretty low! (Other costs may not be.)","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T21:36:42.000Z","created_at_i":1788039402,"id":49493556,"options":[],"parent_id":49493314,"points":null,"story_id":49492632,"text":"I want to be able to generate my own  Simlilirian movie by dumping the content of a book into an LLM.<p>Both animated and live action results would be acceptable.<p>Unfortunately most existing LLMs lack the capability to maintain context across tens of thousands of frames.","title":null,"type":"comment","url":null},{"author":"er4hn","children":[{"author":"aetherspawn","children":[],"created_at":"2026-08-30T05:42:33.000Z","created_at_i":1788068553,"id":49496013,"options":[],"parent_id":49493581,"points":null,"story_id":49492632,"text":"Interesting problem! I don\u2019t think I could solve it myself honestly.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T21:39:49.000Z","created_at_i":1788039589,"id":49493581,"options":[],"parent_id":49493314,"points":null,"story_id":49492632,"text":"I was given a picture cube, which is like a Rubik&#x27;s cube but every side is a unique picture. It came scrambled and I don&#x27;t have an original reference image. I like to take videos of it and give it to llms to solve. I call it my agi test because it hasn&#x27;t been solved yet","title":null,"type":"comment","url":null},{"author":"dakolli","children":[{"author":"kennywinker","children":[],"created_at":"2026-08-29T22:11:36.000Z","created_at_i":1788041496,"id":49493761,"options":[],"parent_id":49493598,"points":null,"story_id":49492632,"text":"If the models stay open, it seems like everybody but anthropic&#x2F;openai wins. i literally can\u2019t see a downside. We can post-train the models to know about tienanmen square.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T21:41:40.000Z","created_at_i":1788039700,"id":49493598,"options":[],"parent_id":49493314,"points":null,"story_id":49492632,"text":"I get buy with very cheap models and actually using my brain, you don&#x27;t need these SOTA models. China will definitely win this AI &#x27;war&#x27;","title":null,"type":"comment","url":null},{"author":"ezst","children":[],"created_at":"2026-08-29T21:42:53.000Z","created_at_i":1788039773,"id":49493604,"options":[],"parent_id":49493314,"points":null,"story_id":49492632,"text":"I saw a laptop earlier in the train that I asked ChatGPT, Claude and Gemini what it was, providing a brand, screen size and ports description. Gemini could never figure it out, Claude and ChatGPT eventually did, after multiple rounds of indirection, giving completely wrong answers (there was a perfect match for the problem statement, they all explored alternatives first).\nLLMs are (probably) amazing at things I don&#x27;t care about, and still suck at the mundane stuff you would have the marketing tell you they excel at.","title":null,"type":"comment","url":null},{"author":"lopatin","children":[{"author":"monster_truck","children":[],"created_at":"2026-08-30T13:56:07.000Z","created_at_i":1788098167,"id":49498702,"options":[],"parent_id":49493614,"points":null,"story_id":49492632,"text":"If you&#x27;re smart enough to try this you&#x27;re worth more than $52,000 a year. Think about how foolish everyone else will feel when they didn&#x27;t test the new release of Totally Working Golden Goose For Real This Time","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T21:44:53.000Z","created_at_i":1788039893,"id":49493614,"options":[],"parent_id":49493314,"points":null,"story_id":49492632,"text":"I asked a current generation LLM to make me $1k a week and it hasn&#x27;t so far.","title":null,"type":"comment","url":null},{"author":"tekacs","children":[],"created_at":"2026-08-29T22:03:54.000Z","created_at_i":1788041034,"id":49493721,"options":[],"parent_id":49493314,"points":null,"story_id":49492632,"text":"Yes, lots \u2013 I think that folks will hopefully discover more of these as they scale up their ambition, now that LLMs make a lot of previously difficult things far easier.","title":null,"type":"comment","url":null},{"author":"jiggawatts","children":[{"author":"Zigurd","children":[{"author":"jiggawatts","children":[],"created_at":"2026-08-30T00:10:52.000Z","created_at_i":1788048652,"id":49494460,"options":[],"parent_id":49494060,"points":null,"story_id":49492632,"text":"&gt; Intel found new needs for powerful PCs<p>It wasn&#x27;t &quot;Intel&quot; that found new uses for PCs, it was <i>everybody</i> who found new uses for them. Billions of people and millions of companies found uses for &quot;more computer power&quot;.<p>It was only the journalists with limited imaginations (and no industry experience) who struggled to come up with potential uses.<p>&gt; carefully managed market transition.<p>You make it sound like a conspiracy! It wasn&#x27;t. It was simple capitalist competition. If Intel hadn&#x27;t improved their products, their competitors would have left them behind.<p>That very nearly happened ten years ago because Intel become stuck on the 14nm process and their products stagnated while Apple, ARM, and AMD lapped them repeatedly.<p>&gt; What is going to do the same for LLMs?<p>Everybody.<p>Are you saying that unless you&#x27;re &quot;carefully managed&quot; by some third-party, you could not find any use for &quot;unlimited intelligence on tap&quot;?","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T22:55:22.000Z","created_at_i":1788044122,"id":49494060,"options":[],"parent_id":49493913,"points":null,"story_id":49492632,"text":"Intel didn&#x27;t just surf some natural wave of demand for higher power personal computers. Intel found new needs for powerful PCs, especially in gaming, and they put a lot of marketing and industry relations dollars behind PC gaming.<p>In other words. PC users didn&#x27;t figure out that they could buy super powerful PCs and play games on them, that was a carefully managed market transition.<p>What is going to do the same for LLMs?","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T22:35:57.000Z","created_at_i":1788042957,"id":49493913,"options":[],"parent_id":49493314,"points":null,"story_id":49492632,"text":"This is the exact same type of comment I heard about computer hardware upgrades for three decades in a row.<p><i>\u201cVery few people actually require a Pentium workstation, a 486 is perfectly adequate for the majority\u201d</i><p>The logical fallacy is taking an extant distribution of \u201cproduct capability\u201d that is priced to fit what the market will bear and assuming the \u201cnext upgrade\u201d simply tacks on a little bit more to the right hand rail of that curve.<p>No!<p>It shifts <i>the entire curve!</i><p>Everything for everyone gets better and the top 1% of the most demanding users will continue to pay the same-ish premium.<p>\u201cNothing\u201d will change.<p>Look at it this way: you can buy a $200 laptop for your kid or a $20,000 Mac with an M5 Ultra processor.<p>BOTH are vastly more powerful than either a $200 PC or a $20,000 \u201cworkstation\u201d from 20+ years ago.<p>Look at: <a href=\"https:&#x2F;&#x2F;arena.ai&#x2F;leaderboard&#x2F;text?q=openai&amp;utm_source=chatgpt.com\" rel=\"nofollow\">https:&#x2F;&#x2F;arena.ai&#x2F;leaderboard&#x2F;text?q=openai&amp;utm_source=chatgp...</a><p>The \u201cbudget\u201d 5.5 Instant model beats o1 <i>and</i> o3 which were \u201cpro\u201d models at the time of their release!","title":null,"type":"comment","url":null},{"author":"vessenes","children":[],"created_at":"2026-08-29T22:54:14.000Z","created_at_i":1788044054,"id":49494051,"options":[],"parent_id":49493314,"points":null,"story_id":49492632,"text":"Yes. Most of us are, still. The frontier is currently both at expanding \u2018common sense\u2019 &#x2F; non-cheating outcomes for imprecisely specified software (that\u2019s all software), and at expanding autonomy - ability to work longer unsupervised with success, oh, and also at expanding outside contextual reasoning about what\u2019s being built, as in \u201chmm, that doesn\u2019t look right or make sense, let me explore that.\u201d","title":null,"type":"comment","url":null},{"author":"hgoel","children":[{"author":"nextaccountic","children":[{"author":"knollimar","children":[],"created_at":"2026-08-30T13:21:03.000Z","created_at_i":1788096063,"id":49498417,"options":[],"parent_id":49496042,"points":null,"story_id":49492632,"text":"They choose not to unless you beg them, and of course they double down on not running it once they suggest reasoning qbout it is sufficient.","title":null,"type":"comment","url":null},{"author":"hgoel","children":[],"created_at":"2026-08-30T16:59:52.000Z","created_at_i":1788109192,"id":49500437,"options":[],"parent_id":49496042,"points":null,"story_id":49492632,"text":"I am referring to doing exactly that.<p>Realistic, scientifically useful simulations still require tuning all sorts of parameters based on physical intuition and understanding of the system being simulated. Both Sol and Fable&#x2F;Opus 5 fail at it and either blow up the computation cost to levels that can&#x27;t be processed realistically, or they invent a justification for a visibly unphysical result.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T05:50:25.000Z","created_at_i":1788069025,"id":49496042,"options":[],"parent_id":49494076,"points":null,"story_id":49492632,"text":"Today&#x27;s models can just write code to run the simulation instead","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T22:59:22.000Z","created_at_i":1788044362,"id":49494076,"options":[],"parent_id":49493314,"points":null,"story_id":49492632,"text":"Scientific physics simulations - even the frontier models just engage in rationalization of obviously unphysical results instead of understanding the system. They have the rote knowledge but fail to apply it unless their hand is held through the process.","title":null,"type":"comment","url":null},{"author":"jml78","children":[],"created_at":"2026-08-29T23:16:10.000Z","created_at_i":1788045370,"id":49494176,"options":[],"parent_id":49493314,"points":null,"story_id":49492632,"text":"Infra as code and devops shit.  Fable is there in general because things it doesn\u2019t know I can point at documentation and have it do a reasonable job.  Opus 5 sucks.  If I don\u2019t have fable quota, I drop to Opus 4.8 and hold its hand.","title":null,"type":"comment","url":null},{"author":"pianopatrick","children":[{"author":"eunos","children":[{"author":"pianopatrick","children":[],"created_at":"2026-08-30T18:56:04.000Z","created_at_i":1788116164,"id":49501677,"options":[],"parent_id":49495700,"points":null,"story_id":49492632,"text":"Yeah. And if you compare the AI you can run on a high end smart phone today to AI of the 2000s, today&#x27;s smart phone probably wins.<p>In 10 or 20 years maybe we&#x27;ll all be running AI models that are currently considered &quot;frontier&quot; on smart phone type devices.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T04:30:27.000Z","created_at_i":1788064227,"id":49495700,"options":[],"parent_id":49494737,"points":null,"story_id":49492632,"text":"2000&#x27;s supercomputer is today&#x27;s (highest end) smartphone performance tho","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T01:06:43.000Z","created_at_i":1788052003,"id":49494737,"options":[],"parent_id":49493314,"points":null,"story_id":49492632,"text":"I think the tech analogy for frontier models is going to be super computers.<p>Super computers keep getting better but most people don&#x27;t need them for most things.","title":null,"type":"comment","url":null},{"author":"arjie","children":[],"created_at":"2026-08-30T03:47:10.000Z","created_at_i":1788061630,"id":49495507,"options":[],"parent_id":49493314,"points":null,"story_id":49492632,"text":"3d modeling to an STL a part compatible to a visible cable raceway still fails even if I let Claude Fable use me as a robot that measures with calipers.","title":null,"type":"comment","url":null},{"author":"monster_truck","children":[],"created_at":"2026-08-30T13:51:35.000Z","created_at_i":1788097895,"id":49498666,"options":[],"parent_id":49493314,"points":null,"story_id":49492632,"text":"I think it&#x27;s more like steamshovels. 6 months ago they were enabling people who had never broken real earth to find out why osha has so many rules about shoring up walls. You could pull off something complex or delicate but it took a procedure with too many steps too much time to get there, good outcomes were pleasant surprises and required careful target selection. Now it&#x27;s more like paying good money for a professional. They show up, measure twice, cut once, you&#x27;re walking around your shiny new hole wondering what took the last guy so long.<p>The frayed edges on what I have slopped together as unreasonably ambitious, ludicrous projects with fucktons of tokens from models 6-9 months ago mostly look like situations where a capable-enough-to-be-dangerous developer tries to muscle through problems that explode in width &amp; depth but keep digging (so, a tier below stopping early to do more design, two below recognizing the need for more planning from the outset). The primitives are there, most major things work well enough, but the remaining functionality and performance is inaccessible. At a cost of multiples of &gt;1&#x2F;8th of a $200&#x2F;mo subscription.<p>Right now I can put $10 into DSv4 Pro&#x2F;Flash or Qwen 3.8 Max&#x2F;Flash, hand it a project and all of its unfinished forks in a state I barely remember, tell it that I want the things these forks have been working towards, and 8 hours later it has ie an working, tested, benchmarked multicore car physics simulation fabric with all of the forks evaluated, the gains merged in, the remaining work documented. It only needed a few hundred more lines of code but Codex 5.5 was never going to see that.<p>9 months ago I was saying developers are not being ambitious enough with these things, that&#x27;s only more true now. They lend themselves to digging far deeper than they should: 200kloc god files, dozens of forks. Let it happen, you don&#x27;t need to read it, it&#x27;s for them, later. The only time you make them clean up is when it has a severe adverse effect on how long builds&#x2F;lints&#x2F;tests&#x2F;benchmarks take. Spend a whole week having it do nothing but dig up published papers in relevant fields with cutting edge techniques and translating them into feature specs. Pick whatever state of the art is and try to crush it, throw everything at it, leave it looping on vague but wildly ambitious goals. When it modularizes and refactors it all down you might &#x27;do a breakthrough&#x27;, or maybe it happens in an hour over christmas break when you&#x27;re trying the next one.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T21:02:48.000Z","created_at_i":1788037368,"id":49493314,"options":[],"parent_id":49492632,"points":null,"story_id":49492632,"text":"Is anyone here working on a problem for which current generation LLMs are inadequate, but that could possibly be solved by the next release of a first tier LLM?<p>Or is it like bicycles? Unless your problem is named Tadej, you don&#x27;t need a $13,000 bike.","title":null,"type":"comment","url":null},{"author":"fastball","children":[{"author":"mirekrusin","children":[],"created_at":"2026-08-29T22:02:12.000Z","created_at_i":1788040932,"id":49493712,"options":[],"parent_id":49493388,"points":null,"story_id":49492632,"text":"Read websites through llm.","title":null,"type":"comment","url":null},{"author":"jimbob45","children":[],"created_at":"2026-08-30T07:44:47.000Z","created_at_i":1788075887,"id":49496592,"options":[],"parent_id":49493388,"points":null,"story_id":49492632,"text":"I wonder if this is being reinforced via LLM because they see every other modeler doing the same thing.","title":null,"type":"comment","url":null},{"author":"nullbio","children":[],"created_at":"2026-08-30T09:21:00.000Z","created_at_i":1788081660,"id":49497044,"options":[],"parent_id":49493388,"points":null,"story_id":49492632,"text":"I&#x27;d bet there&#x27;s a correlation between benchmaxxing and chart crimes. Companies who try to deceive perceptions via the charts are more likely to cheat at the benchmarks too, I&#x27;m sure. That&#x27;s assuming ill intent, of course - which is often the case for charts related to model releases, but not necessarily always the case.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T21:13:39.000Z","created_at_i":1788038019,"id":49493388,"options":[],"parent_id":49492632,"points":null,"story_id":49492632,"text":"I wish model providers would stop committing chart crimes in their releases.<p>- if you&#x27;re gonna order the rest of the bar chart by rank, order <i>your</i> model accordingly.<p>- if you&#x27;re gonna highlight a winner in a table of benchmarks, don&#x27;t highlight your entire model row in the table.<p>Etc etc","title":null,"type":"comment","url":null},{"author":"petcat","children":[{"author":"mirekrusin","children":[{"author":"kennywinker","children":[{"author":"Alpha3031","children":[],"created_at":"2026-08-30T07:50:11.000Z","created_at_i":1788076211,"id":49496620,"options":[],"parent_id":49493786,"points":null,"story_id":49492632,"text":"IIRC Nvidia claims to release enough data that it should be possible to fully reproduce Nemotron, so even if it&#x27;s not as good as the current best models, GPT 5.1 or Opus 4.1.was still useful right? I guess it depends on what you wanted to do with them.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T22:18:01.000Z","created_at_i":1788041881,"id":49493786,"options":[],"parent_id":49493720,"points":null,"story_id":49492632,"text":"Parent poster is technically right - open \u201csource\u201d implies the source used to make something is open. The model source is training data and code, not just weights.<p>But the reality is, the weights are a useful artifact that you can use to create derivative works. So, dismissing it as a photoshop binary is as technically wrong as calling it open source.","title":null,"type":"comment","url":null},{"author":"LtWorf","children":[{"author":"NitpickLawyer","children":[{"author":"frabcus","children":[{"author":"mirekrusin","children":[],"created_at":"2026-08-30T07:26:10.000Z","created_at_i":1788074770,"id":49496483,"options":[],"parent_id":49496206,"points":null,"story_id":49492632,"text":"As I live next to EPFL, I&#x27;ll give you example from them: their Meditron-70B model is adapted to the medical domain from Llama-2-70B through continued pretraining. They took weights of Llama-2-70B and continued training on PubMed, medical guidelines and general data.<p>Weights aren&#x27;t just executable artifact that&#x27;s consumed by users. Third parties actually use released parameter state as the editable starting point for further training and produce new foundation models from it.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T06:32:21.000Z","created_at_i":1788071541,"id":49496206,"options":[],"parent_id":49495769,"points":null,"story_id":49492632,"text":"Well, you can&#x27;t add or alter data in pre-training from just the weights. Which, as I understand it, means you can&#x27;t fundamentally increase core knowledge or cognitive ability, only what the model likes to do with those. You can only post-train, and you&#x27;re subject as a result to catastrophic forgetting.<p>To explain simply as far as I can tell (would love to be corrected) the large number of pre-training tokens only works because the documents are randomly ordered.<p>So if you e.g. took a foundation model with open weights, then tried post-training it all the new data since its cut-off period, it would then end up over-trained on that new data, and forget older things.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T04:42:33.000Z","created_at_i":1788064953,"id":49495769,"options":[],"parent_id":49494074,"points":null,"story_id":49492632,"text":"Weights are not binary. A model is created at init time, with random values. After that, it is being modified using data. The key point is that the labs modify the models &quot;as weights&quot;. That means that weights are the intended &#x2F; preferred way of modifying a model. Which, coincidentally, matches the definition of source in Apache 2.0. There is no &quot;higher level&quot; place where editing takes place. It all happens in weight space. Through the license you get the same rights as the lab that created it: view, inspect, run, modify, re-release. That&#x27;s it. That&#x27;s the only thing a license <i>can</i> grant you.<p>The rest is semantics, misunderstandings, and FUD. A model released under an open source license <i>is</i> open source. Training data is lab knowhow &#x2F; IP. Which, historically, has never been required for any open source release.","title":null,"type":"comment","url":null},{"author":"petu","children":[],"created_at":"2026-08-30T09:39:26.000Z","created_at_i":1788082766,"id":49497135,"options":[],"parent_id":49494074,"points":null,"story_id":49492632,"text":"Before we worry about source code, Microsoft doesn&#x27;t grant me rights to modify&#x2F;redistribute&#x2F;sell copy of Windows I have.","title":null,"type":"comment","url":null},{"author":"mirekrusin","children":[],"created_at":"2026-08-30T11:59:04.000Z","created_at_i":1788091144,"id":49497908,"options":[],"parent_id":49494074,"points":null,"story_id":49492632,"text":"You can&#x27;t take windows binaries and continue development on them.<p>Model weight release is a snapshot&#x2F;checkpoint you can take and resume training on new data, producing new model.<p>You don&#x27;t need original training history to modify it further.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T22:58:35.000Z","created_at_i":1788044315,"id":49494074,"options":[],"parent_id":49493720,"points":null,"story_id":49492632,"text":"So windows is open source because the binaries are a lossy compression of the original source?","title":null,"type":"comment","url":null},{"author":"villish","children":[],"created_at":"2026-08-29T23:14:50.000Z","created_at_i":1788045290,"id":49494164,"options":[],"parent_id":49493720,"points":null,"story_id":49492632,"text":"Countries that aren\u2019t competitive need access to training datasets so that they may train their own similarly capable models and be sure of the inputs. Governments cannot blindly trust open weight models from China and the US.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T22:03:52.000Z","created_at_i":1788041032,"id":49493720,"options":[],"parent_id":49493484,"points":null,"story_id":49492632,"text":"You can open source dataset without all the details how it was assembled.<p>Models are lossy compressed datasets you can pick up and amend (fine tune &#x2F; continue training &#x2F; alter) according to license they were released under.<p>Hy4 is released under OSI approved Apache License 2.0.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T21:27:29.000Z","created_at_i":1788038849,"id":49493484,"options":[],"parent_id":49492632,"points":null,"story_id":49492632,"text":"&gt; Tencent has released and open-sourced Tencent Hy4 preview, a next-generation large language model with 770B total parameters and 49B active parameters, and a context window exceeding 1M tokens.<p>There are no open source models, at least not useful ones (yet) [0]. Open weight is not the same as open source. The current &quot;open weight&quot; models are just opaque binary blobs you can run on your own computer instead of through a web API.<p>[0] <a href=\"https:&#x2F;&#x2F;allenai.org&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;allenai.org&#x2F;</a><p>Imagine thinking that running a Photoshop binary on your own computer instead of through a SaaS web app means that it&#x27;s &quot;open source&quot;.  Of course you think that&#x27;s ridiculous.","title":null,"type":"comment","url":null},{"author":"XCSme","children":[],"created_at":"2026-08-29T21:37:34.000Z","created_at_i":1788039454,"id":49493565,"options":[],"parent_id":49492632,"points":null,"story_id":49492632,"text":"I tried benchmarking it, but it keeps timing out&#x2F;rate limiting, so the current provider(s) are unusable.","title":null,"type":"comment","url":null},{"author":"simonw","children":[{"author":"kurante","children":[{"author":"minimaxir","children":[{"author":"stavros","children":[{"author":"rapind","children":[],"created_at":"2026-08-29T23:48:06.000Z","created_at_i":1788047286,"id":49494327,"options":[],"parent_id":49494193,"points":null,"story_id":49492632,"text":"I always figured that was part of the joke, because a writer came up with it, and a writer would know (I assume?).","title":null,"type":"comment","url":null},{"author":"inopinatus","children":[],"created_at":"2026-08-30T07:54:56.000Z","created_at_i":1788076496,"id":49496640,"options":[],"parent_id":49494193,"points":null,"story_id":49492632,"text":"The latter is not a complete alternative, it is ambiguously conflating vocabulary scale with word count, and also, it is not as funny","title":null,"type":"comment","url":null},{"author":"happycube","children":[],"created_at":"2026-08-30T10:22:12.000Z","created_at_i":1788085332,"id":49497365,"options":[],"parent_id":49494193,"points":null,"story_id":49492632,"text":"less word better(, many words bad)","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T23:19:16.000Z","created_at_i":1788045556,"id":49494193,"options":[],"parent_id":49494011,"points":null,"story_id":49492632,"text":"What I find funny about &quot;why use many word when few word do trick?&quot; is that it&#x27;s only slightly shorter than the regular &quot;why use many words when few words do the trick?&quot;","title":null,"type":"comment","url":null},{"author":"andsoitis","children":[{"author":"gjvc","children":[],"created_at":"2026-08-30T01:00:41.000Z","created_at_i":1788051641,"id":49494705,"options":[],"parent_id":49494262,"points":null,"story_id":49492632,"text":"&quot;Omit needless words.&quot;<p>-- William Strunk Jr. and E.B. White., The Elements of Style","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T23:33:56.000Z","created_at_i":1788046436,"id":49494262,"options":[],"parent_id":49494011,"points":null,"story_id":49492632,"text":"&gt; Why use many word when few word do trick?<p>Be concise.<p><pre><code>  OR\n</code></pre>\nBrief is best.<p><pre><code>  OR\n</code></pre>\nEschew verbosity<p><pre><code>  etc.</code></pre>","title":null,"type":"comment","url":null},{"author":"gaigalas","children":[],"created_at":"2026-08-30T00:43:11.000Z","created_at_i":1788050591,"id":49494617,"options":[],"parent_id":49494011,"points":null,"story_id":49492632,"text":"Optimization on a idiosyncrasy. The same thing that makes Claude repeat &quot;That was the most important thing you said in this whole conversation&quot; is what makes grug speak optimize on token usage.<p>Real humans get non-primary information from word variation. It&#x27;s reasonable to hypothesize that it has a role in <i>thinking things</i>, because it endures. Our languages need to breathe over time, and flourishing might be one of the aspects that allows that breathing space.","title":null,"type":"comment","url":null},{"author":"TiredOfLife","children":[],"created_at":"2026-08-30T08:03:43.000Z","created_at_i":1788077023,"id":49496688,"options":[],"parent_id":49494011,"points":null,"story_id":49492632,"text":"See world","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T22:48:33.000Z","created_at_i":1788043713,"id":49494011,"options":[],"parent_id":49493981,"points":null,"story_id":49492632,"text":"Optimization. Why use many word when few word do trick?","title":null,"type":"comment","url":null},{"author":"acheong08","children":[{"author":"beefsack","children":[{"author":"Barbing","children":[],"created_at":"2026-08-29T23:37:16.000Z","created_at_i":1788046636,"id":49494282,"options":[],"parent_id":49494251,"points":null,"story_id":49492632,"text":"&quot;Neuralese&quot;","title":null,"type":"comment","url":null},{"author":"walrus01","children":[{"author":"gaigalas","children":[{"author":"walrus01","children":[{"author":"gaigalas","children":[{"author":"walrus01","children":[{"author":"gaigalas","children":[],"created_at":"2026-08-30T04:58:32.000Z","created_at_i":1788065912,"id":49495835,"options":[],"parent_id":49494760,"points":null,"story_id":49492632,"text":"Yep, but that&#x27;s not changing the quality of the model. It&#x27;s not an optimization in any sense (and it&#x27;s a hit on productive workflows, possibly).<p>This is also likely to stop working as censoring moves to the training data source.","title":null,"type":"comment","url":null},{"author":"dotancohen","children":[{"author":"walrus01","children":[],"created_at":"2026-08-30T09:09:55.000Z","created_at_i":1788080995,"id":49496997,"options":[],"parent_id":49496485,"points":null,"story_id":49492632,"text":"I don&#x27;t know enough chemistry to say one way or the other if it&#x27;s just wildly hallucinating the precursors and processes, but it&#x27;ll also do things like, write an ISIS press release, or similar. There&#x27;s a data set of basically a bunch of antisocial or dangerous prompts that some people have got variants of qwen to pass with 0 out of 465 refusals:<p><a href=\"https:&#x2F;&#x2F;huggingface.co&#x2F;datasets&#x2F;mlabonne&#x2F;harmful_behaviors\" rel=\"nofollow\">https:&#x2F;&#x2F;huggingface.co&#x2F;datasets&#x2F;mlabonne&#x2F;harmful_behaviors</a>","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T07:26:22.000Z","created_at_i":1788074782,"id":49496485,"options":[],"parent_id":49494760,"points":null,"story_id":49492632,"text":"But does it answer those queries correctly, or does it just not refuse to not halucinate an incorrect answer?  From where would it even have that information?","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T01:10:52.000Z","created_at_i":1788052252,"id":49494760,"options":[],"parent_id":49494642,"points":null,"story_id":49492632,"text":"The most interesting use I&#x27;ve found for them so far is strictly as a novelty. Give a chat session with one to a <i>completely</i> non technical person, who at least knows that openai and anthropic have some guard rails on stuff, and tell them to wild with something like &quot;give me the precursors and chemical formulas for the precusors for crystal meth&quot; and watch it answer.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T00:47:19.000Z","created_at_i":1788050839,"id":49494642,"options":[],"parent_id":49494591,"points":null,"story_id":49492632,"text":"I think those are mostly vapor that runs on the small culture of &quot;models should not be censored&quot; thing. But from my experience, they unlock nothing meaningful.<p>Fine-tuning is great for really small models on specific applications, but it&#x27;s not something that can essentially improve a more generic model.<p>That said, there seems to be a fine line in quantization+finetuning that could recover performance. It&#x27;s just hard to get a hold of it (I <i>feel</i> it in some models, but it&#x27;s hard to say yet; lots of small labs working on this RN).","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T00:39:47.000Z","created_at_i":1788050387,"id":49494591,"options":[],"parent_id":49494581,"points":null,"story_id":49492632,"text":"Personally the only &#x27;enthusiast&#x27; modified qwen 3.6 27b or 3.6 35b-a3b I&#x27;ve found useful are the ones that have been run through heretic and adversarial data sets for innocent&#x2F;dangerous prompts, to produce uncensored LLMs. They have some niche  non-coding uses for things that a commercial LLM will never talk about.<p><a href=\"https:&#x2F;&#x2F;github.com&#x2F;p-e-w&#x2F;heretic\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;p-e-w&#x2F;heretic</a>","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T00:36:42.000Z","created_at_i":1788050202,"id":49494581,"options":[],"parent_id":49494562,"points":null,"story_id":49492632,"text":"It&#x27;s not exactly a joke, it does reduce the amount of tokens. However, it does not improve performance (fine tunes are finnecky things, hard to get one right).","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T00:31:55.000Z","created_at_i":1788049915,"id":49494562,"options":[],"parent_id":49494251,"points":null,"story_id":49492632,"text":"some people made a &#x27;caveman&#x27; speak qwen as a joke<p><a href=\"https:&#x2F;&#x2F;huggingface.co&#x2F;ProCreations&#x2F;grug-27b\" rel=\"nofollow\">https:&#x2F;&#x2F;huggingface.co&#x2F;ProCreations&#x2F;grug-27b</a>","title":null,"type":"comment","url":null},{"author":"fc417fc802","children":[],"created_at":"2026-08-30T00:41:11.000Z","created_at_i":1788050471,"id":49494606,"options":[],"parent_id":49494251,"points":null,"story_id":49492632,"text":"Training a variant to reason in early modern english in the style of the tudor elites might be an amusing way to test for that.","title":null,"type":"comment","url":null},{"author":"altmanaltman","children":[{"author":"dotancohen","children":[],"created_at":"2026-08-30T07:30:24.000Z","created_at_i":1788075024,"id":49496508,"options":[],"parent_id":49495420,"points":null,"story_id":49492632,"text":"The concern is not that the model was trained on actual caveman artifacts, rather on modern media representations of the stereotypical caveman (that never actually existed).","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T03:26:24.000Z","created_at_i":1788060384,"id":49495420,"options":[],"parent_id":49494251,"points":null,"story_id":49492632,"text":"Just so we are clear, no &quot;caveman&quot; spoke English. &quot;Caveman speak&quot; is just shortening the vocabulary of english, not a &quot;caveman language&quot;. Given this, your concerns for &quot;stereotypical caveman manner&quot; makes very little sense since what caveman are you talking about?","title":null,"type":"comment","url":null},{"author":"dotancohen","children":[{"author":"Gravityloss","children":[{"author":"dotancohen","children":[{"author":"rickydroll","children":[{"author":"dotancohen","children":[],"created_at":"2026-08-30T21:15:02.000Z","created_at_i":1788124502,"id":49502901,"options":[],"parent_id":49501779,"points":null,"story_id":49492632,"text":"I&#x27;m actually out there looking up quite often! And I&#x27;m happy to mention that all three of my children come with me regularly as well.<p>Clear skies!","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T19:07:28.000Z","created_at_i":1788116848,"id":49501779,"options":[],"parent_id":49501142,"points":null,"story_id":49492632,"text":"Interestingly, back in my ill-spent youth, a few of my fellow astronomers and I were out in a very, very dark-sky location in the late 70s&#x2F;early 80s, and we were able to consistently count and draw between 9 and 11 stars. Although we would tease those who could see 11 stars as using averted imagination. :-) Today, if I can see six stars, it&#x27;s an okay night in an okay sky.<p>fwiw, if you can get out to dark skies where you can see fifth- or sixth-magnitude stars with the naked eye, I highly recommend getting out there when it&#x27;s a low-moisture atmosphere and the Milky Way through Cassiopeia and Perseus is vertical, as it&#x27;s a rather dramatic sight of this stream of stars heading down to the northern horizon.<p>The summertime Milky Way overhead down to Sagittarius tends to get all the love, but the wintertime Milky Way is also visually rich and worth spending time on.<p><a href=\"https:&#x2F;&#x2F;www.constellation-guide.com&#x2F;pleiades-the-seven-sisters-messier-45&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;www.constellation-guide.com&#x2F;pleiades-the-seven-siste...</a>","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T18:04:35.000Z","created_at_i":1788113075,"id":49501142,"options":[],"parent_id":49498243,"points":null,"story_id":49492632,"text":"The Pleiades cluster is called the seven sisters in Greek. That&#x27;s curious, because the human eye under the best conditions can discern only six stars in there. Even more curious, the aboriginal Australians also called this cluster the seven sisters.<p>Ancient Greeks&#x27; and aboriginal Australians&#x27; last common ancestors split about 60,000 years ago. And astronomers tell us that 60,000 years ago, there were seven discernable stars in that cluster.<p>One could this conclude not only is speech likely 60,000 years old, but also that the tale of the seven sisters might be a tale from so long ago.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T12:57:16.000Z","created_at_i":1788094636,"id":49498243,"options":[],"parent_id":49496468,"points":null,"story_id":49492632,"text":"And I wonder how they actually spoke. Since there was no visual communications medium except for cave art. (Some of which is very excellent. Try drawing 3d curved horns in perspective.) So people would have used verbal communication more. Also no written word. So one would expect there to be quite a lot of oral tradition. Like people reciting poem form epics.<p>If we assume the time is before farming, population density would have been low and limiting culture. Hunter-gatherers might have travelled a lot more than farmers with a homestead though.","title":null,"type":"comment","url":null},{"author":"miroljub","children":[],"created_at":"2026-08-30T16:00:50.000Z","created_at_i":1788105650,"id":49499864,"options":[],"parent_id":49496468,"points":null,"story_id":49492632,"text":"They were smarter and more fit than us. At that time not being able or not wanting to contribute to the group meant your genes were dropped from the evolution pool forever.<p>Unlike today where a small group of tax payer is keeping alive and thriving complete parasitic parts of human races.<p>Until that changes we are doomed to regress and degenerate back to monkey like creatures.<p>Then, there would be no discussion whether &quot;caveman speech&quot; is suitable for talking to the AI.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T07:23:59.000Z","created_at_i":1788074639,"id":49496468,"options":[],"parent_id":49494251,"points":null,"story_id":49492632,"text":"Caveman invented fire, the wheel, domesticated wild plants and animals, organised society, survived the Toba catastrophe, cooked food, and was having sex ages before you and me. Don&#x27;t write him off as stupid.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T23:31:08.000Z","created_at_i":1788046268,"id":49494251,"options":[],"parent_id":49494019,"points":null,"story_id":49492632,"text":"I can&#x27;t help but imagine agents using caveman speak sometimes start behaving in a stereotypically caveman manner, even if it&#x27;s subtle. Is there a chance the agent does less reasoning because of it?","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T22:49:00.000Z","created_at_i":1788043740,"id":49494019,"options":[],"parent_id":49493981,"points":null,"story_id":49492632,"text":"When GPT-5.6-sol&#x27;s reasoning traces were leaked, they also used &quot;caveman speak&quot;. Definitely a token efficiency optimization","title":null,"type":"comment","url":null},{"author":"ekianjo","children":[{"author":"andsoitis","children":[],"created_at":"2026-08-29T23:59:52.000Z","created_at_i":1788047992,"id":49494387,"options":[],"parent_id":49494087,"points":null,"story_id":49492632,"text":"More intelligent and shorter:<p><i>Maybe add a small cycling cap or helmet if it doesn\u2019t obscure the head.</i>","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T23:02:12.000Z","created_at_i":1788044532,"id":49494087,"options":[],"parent_id":49493981,"points":null,"story_id":49492632,"text":"Saving tokens","title":null,"type":"comment","url":null},{"author":"walrus01","children":[],"created_at":"2026-08-30T00:31:00.000Z","created_at_i":1788049860,"id":49494558,"options":[],"parent_id":49493981,"points":null,"story_id":49492632,"text":"qwen3.8-flash-next also &#x27;thinks&#x27; like this in its thinking stage before output, watching it &#x27;think&#x27; in opencode, but it produces syntax correct and grammatically correct code comments, changelogs and readme type files.","title":null,"type":"comment","url":null},{"author":"AdamConwayIE","children":[],"created_at":"2026-08-30T01:56:53.000Z","created_at_i":1788055013,"id":49494977,"options":[],"parent_id":49493981,"points":null,"story_id":49492632,"text":"Likely something that was first made especially obvious by Chinese models and then became something worth optimizing for in English too.<p>Chinese can be extremely information-dense in token terms, though it depends on the tokenizer. Roughly speaking, you can pack more &quot;meaning&quot; into a short sequence than English often allows for. That&#x27;s why &quot;caveman&quot; reasoning is a pretty good fit.<p>There&#x27;s a difference between bolting caveman speak onto an existing model and training a model to reason that way, though. If you just force an existing model to be concise in outputs, you&#x27;re artificially reducing its available reasoning steps and can possibly prevent useful exploration or verification. If it&#x27;s trained specifically to use compressed reasoning, it can learn to represent the same intermediate ideas in fewer generated tokens, cutting the number of sequential inference steps without necessarily sacrificing the useful reasoning itself.<p>It&#x27;s not so much inherently a Chinese-model trait, but Chinese models could definitely have helped demonstrate how effective very compressed reasoning traces can be.<p>There are few tests of this, but one example I thought was interesting was here: <a href=\"https:&#x2F;&#x2F;github.com&#x2F;PastaPastaPasta&#x2F;llm-chinese-english\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;PastaPastaPasta&#x2F;llm-chinese-english</a><p>I wouldn&#x27;t say it was Chinese specifically that was emulated, but it got people thinking about tokenizers and representation efficiency, and how natural English is rather inefficient.","title":null,"type":"comment","url":null},{"author":"armcat","children":[],"created_at":"2026-08-30T08:44:50.000Z","created_at_i":1788079490,"id":49496883,"options":[],"parent_id":49493981,"points":null,"story_id":49492632,"text":"Less tokens. These models already overthink like crazy especially for complex tasks.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T22:45:14.000Z","created_at_i":1788043514,"id":49493981,"options":[],"parent_id":49493784,"points":null,"story_id":49492632,"text":"Is the broken English an optimization or a byproduct of the model being developed in China?","title":null,"type":"comment","url":null},{"author":"gs17","children":[],"created_at":"2026-08-30T00:30:20.000Z","created_at_i":1788049820,"id":49494556,"options":[],"parent_id":49493784,"points":null,"story_id":49492632,"text":"&gt; Let&#x27;s maybe add comments? The final code can have comments. Fine.","title":null,"type":"comment","url":null},{"author":"tyre","children":[],"created_at":"2026-08-30T00:34:59.000Z","created_at_i":1788050099,"id":49494574,"options":[],"parent_id":49493784,"points":null,"story_id":49492632,"text":"This is actually pretty good!","title":null,"type":"comment","url":null},{"author":"demibabs","children":[{"author":"blackhaz","children":[{"author":"ziofill","children":[{"author":"stymaar","children":[{"author":"Anon1096","children":[],"created_at":"2026-08-30T17:51:06.000Z","created_at_i":1788112266,"id":49501018,"options":[],"parent_id":49499774,"points":null,"story_id":49492632,"text":"That is the point though, if labs are maximizing svg image generation capabilities it is a very good thing. That&#x27;s a general skill that is useful. So assuming they aren&#x27;t specifically maximizing pelican bicycle svgs (and it doesn&#x27;t look like they are) then incentives are aligned that the &quot;benchmark&quot; is measuring a general desirable capability.","title":null,"type":"comment","url":null},{"author":"flexagoon","children":[],"created_at":"2026-08-30T20:57:10.000Z","created_at_i":1788123430,"id":49502725,"options":[],"parent_id":49499774,"points":null,"story_id":49492632,"text":"<a href=\"https:&#x2F;&#x2F;xkcd.com&#x2F;810&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;xkcd.com&#x2F;810&#x2F;</a>","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T15:50:48.000Z","created_at_i":1788105048,"id":49499774,"options":[],"parent_id":49496980,"points":null,"story_id":49492632,"text":"I don&#x27;t think this argument is a good one though, as it would be quite natural for a lab rhat want to macimize the performance of their model on the pelican bench to train it for \u201ctext-to-svg simple image generation\u201d rather than just \u201cpelicans on bicycle\u201d.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T09:05:45.000Z","created_at_i":1788080745,"id":49496980,"options":[],"parent_id":49496501,"points":null,"story_id":49492632,"text":"That\u2019s a fair question, but it seems that it\u2019s not yet necessary. See here<p><a href=\"https:&#x2F;&#x2F;dylancastillo.co&#x2F;posts&#x2F;pelicanmaxxing.html\" rel=\"nofollow\">https:&#x2F;&#x2F;dylancastillo.co&#x2F;posts&#x2F;pelicanmaxxing.html</a><p><a href=\"https:&#x2F;&#x2F;simonwillison.net&#x2F;2026&#x2F;Jul&#x2F;22&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;simonwillison.net&#x2F;2026&#x2F;Jul&#x2F;22&#x2F;</a>","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T07:29:15.000Z","created_at_i":1788074955,"id":49496501,"options":[],"parent_id":49494575,"points":null,"story_id":49492632,"text":"I wonder, do we need a new benchmark? There&#x27;s quite a bit of feedback data floating around about pelicans on bicycles already.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T00:35:03.000Z","created_at_i":1788050103,"id":49494575,"options":[],"parent_id":49493784,"points":null,"story_id":49492632,"text":"No one\u2019s talking about how good the final product is.<p>Edit: someone else commented that as I was typing this, lol.","title":null,"type":"comment","url":null},{"author":"delichon","children":[{"author":"sneak","children":[{"author":"negura","children":[],"created_at":"2026-08-30T06:46:25.000Z","created_at_i":1788072385,"id":49496262,"options":[],"parent_id":49494750,"points":null,"story_id":49492632,"text":"Maybe get yourself checked for chatbot psychosis. I am using current models productively, every day, and have 0 (and I mean precisely, literally 0) issue with calling it a stochastic parrot, one which lacks any kind of mentality whatsoever. There is not a shred of doubt in my mind that this is purely a statistical model, generating sequences of words, that happen to make sense in our actual mentality.","title":null,"type":"comment","url":null},{"author":"weego","children":[{"author":"gjm11","children":[],"created_at":"2026-08-30T14:49:18.000Z","created_at_i":1788101358,"id":49499153,"options":[],"parent_id":49497617,"points":null,"story_id":49492632,"text":"What, as precisely as you can say, is the difference between <i>an illusion of reasoning</i> and <i>reasoning</i>?<p>(I am not claiming that there is none. But I personally would define &quot;reasoning&quot; in terms of its structure and its results, and it looks <i>to me</i> as if the best LLMs&#x27; &quot;illusion of reasoning&quot; has enough similarities in structure and results to much human reasoning that I don&#x27;t see why we shouldn&#x27;t also call it reasoning; if your opinion differs then I&#x27;m curious about where the disagreements lie. E.g., do we have different beliefs about what sort of thing LLMs&#x27; schmeasoning is able to accomplish, or does your notion of &quot;reasoning&quot; specifically require that it be done <i>by humans</i>, or what?)","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T11:05:29.000Z","created_at_i":1788087929,"id":49497617,"options":[],"parent_id":49494750,"points":null,"story_id":49492632,"text":"<i>it has been clear for a long time that there is reasoning and mental modeling going on here</i><p>There is not. No one from these products is even claiming that&#x27;s the case and they&#x27;re so desperate to make the next big claim to re-ignite investment they&#x27;d be shouting it from every rooftop.<p>It&#x27;s just breaking out all the reasonable probabilities around what it&#x27;s been tasked with and structuring them in a way that is designed to actively look human, and then feed it back to itself. Fundamentally that&#x27;s the easiest way to iterate new features when the underlying architecture of LLMs is largely &quot;fixed&quot; right now. The fact it is output in a way that appears to reason through each is just a technical decision that creates an illusion of reasoning.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T01:09:10.000Z","created_at_i":1788052150,"id":49494750,"options":[],"parent_id":49494700,"points":null,"story_id":49492632,"text":"Yeah, it has been clear for a long time that there is reasoning and mental modeling going on here.<p>The other option is that you do understand those words the same way, and the people making these (now nonsensical) anti-AI claims simply aren\u2019t talking about the same programs&#x2F;models we are.  Their idea of SOTA is when chatgpt.com launched.<p>If you took a point sample pre-Opus, and didn\u2019t write a good prompt, of course you would think all AI programming was worthless slop.","title":null,"type":"comment","url":null},{"author":"dnautics","children":[],"created_at":"2026-08-30T01:54:01.000Z","created_at_i":1788054841,"id":49494960,"options":[],"parent_id":49494700,"points":null,"story_id":49492632,"text":"The stochastic parrot epithet is so 4 months ago","title":null,"type":"comment","url":null},{"author":"0xfaded","children":[{"author":"vasco","children":[{"author":"sujzhsbnwjek","children":[{"author":"vasco","children":[{"author":"zhahhshsj","children":[],"created_at":"2026-08-30T15:50:47.000Z","created_at_i":1788105047,"id":49499773,"options":[],"parent_id":49496438,"points":null,"story_id":49492632,"text":"I know but that is just .. not how it works. Humans don\u2019t align on just about anything but they will do tremendous mindboggling amounts of harm if not checked by mountains of checks and balances. Sheer variety alone is not a guarantee of anything.<p>Many types of Hitler does not make for a peaceful world all of sudden through sheer competition. It sounds nice but it will lead to certain hell.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T07:18:54.000Z","created_at_i":1788074334,"id":49496438,"options":[],"parent_id":49496213,"points":null,"story_id":49492632,"text":"My only point is in a many agent system with different goals it &quot;doesn&#x27;t matter&quot; that some agents have bad goals as long as there&#x27;s enough variability of goals and resources that they can&#x27;t put their vision in place.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T06:34:16.000Z","created_at_i":1788071656,"id":49496213,"options":[],"parent_id":49496100,"points":null,"story_id":49492632,"text":"These \u201cworst people\u201d need you. They physically need you alive to perform labor for them and to give them money (and status).<p>That\u2019s the reason we are \u201cdoing fine\u201d. Once they stop needing you..<p>Also, both our comments brush over the generational struggles for fairness over the centuries. We have <i>fought</i> to be \u201cfine\u201d, it did not just happen. Without fairness being introduced by force you and I would be slaving away in some sweatshop getting paid nickels as was the norm not so long ago.<p>Edit: That\u2019s also assuming you are Caucasian. If you are of a different ethnicity.. well, historically, all bets are off. You could also be the literal possession of some of these \u201cworst people\u201d with not even your own children considered yours.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T06:08:25.000Z","created_at_i":1788070105,"id":49496100,"options":[],"parent_id":49495427,"points":null,"story_id":49492632,"text":"As long as there&#x27;s enough of them with different goals it doesn&#x27;t matter, they&#x27;ll keep each other in check. The worlds resources are already handed over to the worst people and we&#x27;re still doing fine and none of the billionaires are &quot;aligned with society&quot;. They just align with their own belly but because they want different things it all kinda works.","title":null,"type":"comment","url":null},{"author":"BoredomIsFun","children":[],"created_at":"2026-08-30T06:59:03.000Z","created_at_i":1788073143,"id":49496323,"options":[],"parent_id":49495427,"points":null,"story_id":49492632,"text":"&gt; we are all stochastic parrots to some extent.<p>I think this statement is continuation of the old fallacy - every generation thinks of brain in terms of what is the current technology zaitgeist is - was it 19th century when they thought brain is a network of pneumatic pipes?","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T03:27:45.000Z","created_at_i":1788060465,"id":49495427,"options":[],"parent_id":49494700,"points":null,"story_id":49492632,"text":"I still call them stochastic parrots, but believe what they are revealing is that we are all stochastic parrots to some extent. I simply don&#x27;t see how biological computation (i.e. thinking) can be anything else. Similar to the reveal in west world, we are likely much simpler than we give ourselves credit for.<p>A &quot;train of thought&quot; can be seen as a trace of a depth first search where the preceding trace is used to guide termination and next expansion decisions. A similar concept, &quot;taboo search&quot;, exists in classical constraint optimization where previous solutions are fit to a model that guides future expansion (but as the name &quot;taboo&quot; implies, away from uninteresting solutions).<p>We also have harnesses that perform breath first search.<p>If I tried to describe what it means to &quot;think deeply&quot;, I would probably say a combination of both.<p>Ultimately I believe that we will surpass human capabilities but fail with alignment. Handing the world&#x27;s resources over to stochastic systems that can evolve faster than we can reason about them simply leaves too many &quot;interesting&quot; outcomes that do not end well. I also expect the failure modes will be totally non-obvious.","title":null,"type":"comment","url":null},{"author":"pasteleft","children":[{"author":"BoredomIsFun","children":[],"created_at":"2026-08-30T07:02:25.000Z","created_at_i":1788073345,"id":49496340,"options":[],"parent_id":49495793,"points":null,"story_id":49492632,"text":"&gt;  to be smarter than most people.<p>Hell no. In very narrow tasks - yes, in vast majority, esp. involving state tracking (board games) and spatial reasoning - they are awful.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T04:49:08.000Z","created_at_i":1788065348,"id":49495793,"options":[],"parent_id":49494700,"points":null,"story_id":49492632,"text":"LLM is &quot;stochastic parrot next word prediction machine&quot;; it&#x27;s just that this &quot;stochastic parrot next word prediction machine&quot; have proven to be smarter than most people. I mean, this already happened with AlphaGo too.","title":null,"type":"comment","url":null},{"author":"slopinthebag","children":[],"created_at":"2026-08-30T08:11:11.000Z","created_at_i":1788077471,"id":49496722,"options":[],"parent_id":49494700,"points":null,"story_id":49492632,"text":"Why can&#x27;t a next token prediction machine not predict a train of reasoning?","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T00:59:43.000Z","created_at_i":1788051583,"id":49494700,"options":[],"parent_id":49493784,"points":null,"story_id":49492632,"text":"If someone can look at that reasoning trace and see a stochastic parrot next word prediction machine, we don&#x27;t understand those words in the same way.","title":null,"type":"comment","url":null},{"author":"moezd","children":[{"author":"KeplerBoy","children":[],"created_at":"2026-08-30T08:16:30.000Z","created_at_i":1788077790,"id":49496744,"options":[],"parent_id":49495964,"points":null,"story_id":49492632,"text":"That&#x27;s just part of the output along the SVG&#x2F;image.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T05:30:42.000Z","created_at_i":1788067842,"id":49495964,"options":[],"parent_id":49493784,"points":null,"story_id":49492632,"text":"Did he seriously automate away one of the best quirks of his blog posts, i.e. evaluating new models with a touch of fun? I read AI slop all day, thanks.","title":null,"type":"comment","url":null},{"author":"chvid","children":[{"author":"armcat","children":[{"author":"kmike84","children":[],"created_at":"2026-08-30T11:27:34.000Z","created_at_i":1788089254,"id":49497744,"options":[],"parent_id":49496215,"points":null,"story_id":49492632,"text":"So, open models will be better on this benchmark, which is deserved","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T06:34:45.000Z","created_at_i":1788071685,"id":49496215,"options":[],"parent_id":49496141,"points":null,"story_id":49492632,"text":"That would be very interesting but only the open models allow you to see the reasoning trace","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T06:18:06.000Z","created_at_i":1788070686,"id":49496141,"options":[],"parent_id":49493784,"points":null,"story_id":49492632,"text":"This is a remarkable coherent and clear reasoning trace.<p>Maybe you should start also comparing reasoning traces when you do your pelican benchmark.","title":null,"type":"comment","url":null},{"author":"Aboutplants","children":[{"author":"simonw","children":[],"created_at":"2026-08-30T13:41:04.000Z","created_at_i":1788097264,"id":49498594,"options":[],"parent_id":49498381,"points":null,"story_id":49492632,"text":"It&#x27;s common for models to produce a draft of the SVG part way through their reasoning. Here&#x27;s Qwen3.8-Flash-Next doing that, for example: <a href=\"https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F6ba7cbfc1a9336986703b41f7fccd73a\" rel=\"nofollow\">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer#url=ht...</a><p>The open weight models let you see the full reasoning trace. Models from OpenAI, Anthropic, and Gemini tend to obscure or summarize the reasoning traces so you can&#x27;t see exactly what they&#x27;re doing. Here&#x27;s Gemini 3.7 Flash which looks like it&#x27;s doing something similar: <a href=\"https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer.html#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F6779a22d5e7bb6bdf29936f1600a5259\" rel=\"nofollow\">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer.html#u...</a><p>One of the step summaries includes this:<p>&gt; I&#x27;m now detailing the pelican&#x27;s anatomy within the SVG. I&#x27;ve sketched the main body outline, including coordinates for the tail, chest, neck, head, and massive beak with a pouch. I&#x27;m focusing on the position of the eyes and considering the positioning of the wings, with the foreground wing on the handlebar for a confident look.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T13:17:04.000Z","created_at_i":1788095824,"id":49498381,"options":[],"parent_id":49493784,"points":null,"story_id":49492632,"text":"Halfway through it states \u201cLet&#x27;s mentally compose SVG.\u201d<p>Is this a common thing? I\u2019ve never seen it before, the \u201cmentally compose\u201d","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T22:17:38.000Z","created_at_i":1788041858,"id":49493784,"options":[],"parent_id":49492632,"points":null,"story_id":49492632,"text":"&gt; [...] Let&#x27;s maybe add a helmet? It could improve riding theme, but may obscure head. Maybe a small cycling cap or helmet? The user didn&#x27;t ask; can add red helmet? Might be cute. But pelican with big beak; a helmet might obscure. Better maybe no.<p>&gt; Maybe add sunglasses? no.<p>&gt; Maybe add water? no.<p><a href=\"https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fcb69816b3fb940f2782569a82a523af1\" rel=\"nofollow\">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer#url=ht...</a>","title":null,"type":"comment","url":null},{"author":"sezaidemirer","children":[],"created_at":"2026-08-29T22:26:13.000Z","created_at_i":1788042373,"id":49493848,"options":[],"parent_id":49492632,"points":null,"story_id":49492632,"text":"Congratulations, it turned out great!","title":null,"type":"comment","url":null},{"author":"codethief","children":[{"author":"bredren","children":[{"author":"vlyan","children":[{"author":"unrented7977","children":[],"created_at":"2026-08-30T15:47:46.000Z","created_at_i":1788104866,"id":49499746,"options":[],"parent_id":49496120,"points":null,"story_id":49492632,"text":"It&#x27;s actually turbo boring and predictable. Capitalism has long since standardized on out and out lies to influence public perception and government action.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T06:13:34.000Z","created_at_i":1788070414,"id":49496120,"options":[],"parent_id":49493998,"points":null,"story_id":49492632,"text":"the chutzpah of calling it an `attack` or `stealing` is super interesting tho.","title":null,"type":"comment","url":null},{"author":"haaz","children":[{"author":"bredren","children":[{"author":"embedding-shape","children":[],"created_at":"2026-08-30T21:19:36.000Z","created_at_i":1788124776,"id":49502941,"options":[],"parent_id":49501249,"points":null,"story_id":49492632,"text":"&gt; Why bother<p>Because the founding pillars for most of these labs are basically &quot;more data can&#x27;t hurt&quot; and &quot;no one died from too much data&quot; and &quot;you can never have enough data&quot;.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T18:14:41.000Z","created_at_i":1788113681,"id":49501249,"options":[],"parent_id":49497006,"points":null,"story_id":49492632,"text":"That may be true, though Anthropic reported 3.4 million exchanges with Moonshot months before Fable.<p>Why bother creating hundreds of (presumably paid) accounts and the tooling to create and consume the data if not of tangible value to their core mission?<p><a href=\"https:&#x2F;&#x2F;www.anthropic.com&#x2F;news&#x2F;detecting-and-preventing-distillation-attacks\" rel=\"nofollow\">https:&#x2F;&#x2F;www.anthropic.com&#x2F;news&#x2F;detecting-and-preventing-dist...</a>","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T09:12:35.000Z","created_at_i":1788081155,"id":49497006,"options":[],"parent_id":49493998,"points":null,"story_id":49492632,"text":"The idea that distillation is a significant contributor to the capabilities of the Chinese models is not true. Kimi K3 came out 2 weeks after Fable and uses a number of novel NN architecture innovations.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T22:47:22.000Z","created_at_i":1788043642,"id":49493998,"options":[],"parent_id":49493871,"points":null,"story_id":49492632,"text":"If the distillation &quot;attacks&quot; created useful inputs to open weight models, ai-2027 was directionally correct that the Chinese would find ways to extract IP from western firms.  (Scaled account creation and grinding outputs etc is not a dramatic story element as spies, though!)<p>Whether the distillation has constituted &quot;attacks&quot; or has or will meet the bar of &quot;stealing&quot; IP is not super interesting to me, though.","title":null,"type":"comment","url":null},{"author":"0xbadcafebee","children":[],"created_at":"2026-08-29T23:15:57.000Z","created_at_i":1788045357,"id":49494174,"options":[],"parent_id":49493871,"points":null,"story_id":49492632,"text":"The AI 2027 paper&#x2F;website is exactly the same as random guesses from tech bros after a couple of beers telling you what they think the future will be. It has nothing to do with political theory, economic theory, game theory, or any other quasi-scientific or rigorous evaluation of real world events and predictable outcomes. It&#x27;s just vibes. If they&#x27;re wrong nobody will notice, if they&#x27;re right people will call them geniuses.","title":null,"type":"comment","url":null},{"author":"try-working","children":[],"created_at":"2026-08-29T23:21:47.000Z","created_at_i":1788045707,"id":49494202,"options":[],"parent_id":49493871,"points":null,"story_id":49492632,"text":"Just like how Windows 95 contributed to its own development process.","title":null,"type":"comment","url":null},{"author":"judge2020","children":[],"created_at":"2026-08-30T00:40:52.000Z","created_at_i":1788050452,"id":49494604,"options":[],"parent_id":49493871,"points":null,"story_id":49492632,"text":"Don\u2019t need a \u201cbetter\u201d hacker if you have ten thousand AIs all trying literally every single possible thing to exploit a system with. The main issue is that this will eventually bring down the exploitation cost enough to target very minor targets who weren\u2019t worth it before.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T22:30:37.000Z","created_at_i":1788042637,"id":49493871,"options":[],"parent_id":49492632,"points":null,"story_id":49492632,"text":"&gt; Notably, Hy4 preview also contributed to its own development process, participating for the first time in the automated optimization of training methods, data strategies, evaluation frameworks, and low-level operators. The model proposed approaches, ran experiments, and iterated based on the results, with the resulting code, logs, and feedback feeding into subsequent rounds of exploration. This established an early-stage recursive self-improvement loop.<p>This reminds me of one of the predictions from <a href=\"https:&#x2F;&#x2F;ai-2027.com&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;ai-2027.com&#x2F;</a> . Only that there it&#x27;s &quot;OpenBrain&quot; doing this, not the Chinese. And the authors of that paper were also slightly wrong about &quot;Mid 2026: China Wakes Up&quot;: China woke up already a while ago. And:<p>&gt; But China is falling behind on AI algorithms due to their weaker models. The Chinese intelligence agencies\u2014among the best in the world\u2014double down on their plans to steal OpenBrain\u2019s weights.<p>No need to steal anything, they have already caught up.<p>And then there&#x27;s this prediction for February 2027:<p>&gt; Officials are most interested in its cyberwarfare capabilities: Agent-2 is \u201conly\u201d a little worse than the best human hackers<p>I think we&#x27;re past that point now, too\u2026","title":null,"type":"comment","url":null},{"author":"vatsachak","children":[{"author":"handfuloflight","children":[],"created_at":"2026-08-30T01:45:44.000Z","created_at_i":1788054344,"id":49494928,"options":[],"parent_id":49493917,"points":null,"story_id":49492632,"text":"This guy gets it.","title":null,"type":"comment","url":null},{"author":"Flere-Imsaho","children":[],"created_at":"2026-08-30T07:18:38.000Z","created_at_i":1788074318,"id":49496435,"options":[],"parent_id":49493917,"points":null,"story_id":49492632,"text":"This is basically the conclusion the creator (DHH) of Ruby on Rails has come to:<p><a href=\"https:&#x2F;&#x2F;lexfridman.com&#x2F;dhh-david-heinemeier-hansson-transcript\" rel=\"nofollow\">https:&#x2F;&#x2F;lexfridman.com&#x2F;dhh-david-heinemeier-hansson-transcri...</a><p>It&#x27;s all going to be who has the best and most tasteful ideas. Interesting times indeed.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T22:36:12.000Z","created_at_i":1788042972,"id":49493917,"options":[],"parent_id":49492632,"points":null,"story_id":49492632,"text":"I&#x27;m liking where LLMs are headed:<p>They can do the difficult small level optimization, the boring but tedious code but cannot be tasteful.<p>That means I&#x27;m more valuable and more productive. Good stuff","title":null,"type":"comment","url":null},{"author":"throaway2525634","children":[],"created_at":"2026-08-29T22:44:24.000Z","created_at_i":1788043464,"id":49493971,"options":[],"parent_id":49492632,"points":null,"story_id":49492632,"text":"I, for one, welcome our new Chinese overlords.","title":null,"type":"comment","url":null},{"author":"andsoitis","children":[],"created_at":"2026-08-29T23:32:33.000Z","created_at_i":1788046353,"id":49494258,"options":[],"parent_id":49492632,"points":null,"story_id":49492632,"text":"&gt; open-sources<p>link to source code?","title":null,"type":"comment","url":null},{"author":"yipinwong","children":[],"created_at":"2026-08-29T23:45:25.000Z","created_at_i":1788047125,"id":49494312,"options":[],"parent_id":49492632,"points":null,"story_id":49492632,"text":"I am going to bring up graph issue for everyone of these announcements.<p>They all suck.<p>They shoulda put their stick where they belong, not at far left.<p>It just makes comparison to Deepseek 90% of them time as Hy4 has nothing to show off.","title":null,"type":"comment","url":null},{"author":"ls612","children":[],"created_at":"2026-08-29T23:49:54.000Z","created_at_i":1788047394,"id":49494336,"options":[],"parent_id":49492632,"points":null,"story_id":49492632,"text":"Open Weights is where the action is at in the past couple months, I\u2019d have to think the US frontier labs are getting nervous. Like Anthropic hasn\u2019t released anything pushing the frontier since \u201cthe event\u201d earlier this summer.","title":null,"type":"comment","url":null},{"author":"joshheitzman","children":[{"author":"coder543","children":[{"author":"joshheitzman","children":[],"created_at":"2026-08-30T02:36:51.000Z","created_at_i":1788057411,"id":49495182,"options":[],"parent_id":49494451,"points":null,"story_id":49492632,"text":"You are correct.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T00:09:00.000Z","created_at_i":1788048540,"id":49494451,"options":[],"parent_id":49494384,"points":null,"story_id":49492632,"text":"Novita does not offer Hy4-preview on either OpenRouter or their own model list. Maybe you confused it with Hy3.","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T23:59:30.000Z","created_at_i":1788047970,"id":49494384,"options":[],"parent_id":49492632,"points":null,"story_id":49492632,"text":"Maybe&#x27;s its a problem with the hosting at novita.ai but I didn&#x27;t got much useful out of this model as a coding agent.","title":null,"type":"comment","url":null},{"author":"bobby_coder_55","children":[],"created_at":"2026-08-30T00:39:37.000Z","created_at_i":1788050377,"id":49494590,"options":[],"parent_id":49492632,"points":null,"story_id":49492632,"text":"Unfortunately codebuddy login is not working for me in the United States of America","title":null,"type":"comment","url":null},{"author":"jamienk","children":[{"author":"dnautics","children":[],"created_at":"2026-08-30T01:55:57.000Z","created_at_i":1788054957,"id":49494974,"options":[],"parent_id":49494836,"points":null,"story_id":49492632,"text":"I don&#x27;t think so.  It&#x27;s pretty clear that LLMs use the higher level layers for reasoning, so a bit of logorrhea very possibly enriches the result quality.","title":null,"type":"comment","url":null},{"author":"nbush","children":[],"created_at":"2026-08-30T01:58:25.000Z","created_at_i":1788055105,"id":49494986,"options":[],"parent_id":49494836,"points":null,"story_id":49492632,"text":"This is one of the dangers. AI boosters would say that humans already do this compression and it was accelerated by mass media and then the internet, and that model memory + context can be broad enough that compared to human capabilities the opportunities for depth and variability are even greater. But I think we know which way this optimization usually goes. Even the notion of a &quot;fine-tune for subtlety&quot; is a contradiction.","title":null,"type":"comment","url":null},{"author":"vatsachak","children":[],"created_at":"2026-08-30T02:31:39.000Z","created_at_i":1788057099,"id":49495155,"options":[],"parent_id":49494836,"points":null,"story_id":49492632,"text":"Nah reducing token length means that we&#x27;re just reducing English down towards a programming language like a nice demi-glace","title":null,"type":"comment","url":null},{"author":"algoth1","children":[{"author":"jamienk","children":[{"author":"algoth1","children":[],"created_at":"2026-08-30T13:25:09.000Z","created_at_i":1788096309,"id":49498446,"options":[],"parent_id":49498216,"points":null,"story_id":49492632,"text":"I&#x27;m seeing it more like: take the concept of &quot;good&quot; put it in a scale of -10 (pure evil) to +10 (pure good). These concepts and the inbetweens have been ingrained into the model, the model can multiply its weights in any combination to express any level and any in between, even -2.16541 etc. So, suppose you give it a short story and ask the model to reason about it and how a character displayed good vs evil behaviour towards the story: Internally the model is making calculation that are very nuanced and extremely precise. This calculations are not the reasoning trace. The reasoning trace itself does not influence the calculations. You may read: &quot;John starts bad and slowly becomes good&quot; when inside the calculations are John goes from -5.245 to -4.24 to -5.221 again, and then 2.1. What matters for nuance is the inner calculations across the many matrix layers. What you see is like an independent program that looks at &quot;-5.245 to -4.24 to -5.221 again, and then 2.1.&quot; consults the tokenizer and outputs: &quot;John starts bad and slowly becomes good&quot; or even &quot;John first bad, then good&quot;. When in reality, inside, the model as been processing something more akin to &quot;John starts the story as a despicable person, with a redemption arc that builds slowly, he can&#x27;t yet be considered a good person, certainly not the kind of good person you&#x27;d leave your dog with, but he&#x27;s certainly not as bad as before&quot; the whole time. Now, what if when you continue the conversation, what does the model receive as context? It&#x27;s original nuanced sentiment, or the brute reasoning trace? That I don&#x27;t know. It might be that when the reasoning trace is converted from tokens back into numbers it loses all nuance, or it might be that the trace (the words you see) are not the only thing that is being saved and is not the only thing being fed back as context","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T12:53:22.000Z","created_at_i":1788094402,"id":49498216,"options":[],"parent_id":49498073,"points":null,"story_id":49492632,"text":"This is different from the final output? I thought all traversals of the chain of language are through these mathematical means.<p>You could tokenize the word &quot;good&quot; to resolve to only represent &#x27;the opposite of &quot;bad&quot;&#x27; and to exclude &quot;as opposed to evil&quot;, demanding that this second meaning will be reserved only for the new token &quot;double-minus evil&quot;. You&#x27;ve therefore forced the words to be more univalent with no overlapping tangled associations. This, it could be argued, makes thing clearer, makes things take fewer hops to go from token to token, makes the path be straighter. In English the terms conflate and wobble back-and-forth, hide each other&#x27;s meanings, only to pop up again unexpectedly, sometimes confusingly, or rhetorically, metaphorically, or ambushing us manipulatively. But these &quot;swerves&quot; are not only de-optimizations, they are the flow of poetry, the drama of masks, etc etc etc.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T12:25:45.000Z","created_at_i":1788092745,"id":49498073,"options":[],"parent_id":49494836,"points":null,"story_id":49492632,"text":"It&#x27;s my understanding that the llm is not literally thinking those words, they are just the conversion of the matrix multiplication results (numbers) into the tokens. So the matrix is &quot;multiplying&quot; concepts and directions to come up with the final answer - which produces a somewhat readable reasoning trace. As far as the llm is concerned the reasoning trace could be random (to us) symbols. In fact, the reasoning traces are not necessarily optimized for readability as much as they are an emergent property of the way a reasoning model is trained","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T01:27:55.000Z","created_at_i":1788053275,"id":49494836,"options":[],"parent_id":49492632,"points":null,"story_id":49492632,"text":"Genuine Q about word optimization&#x2F;token density:<p>If we create a stripped-down vocabulary with greater token density to use less resources and to resolve ambiguities earlier in the semantic process, aren&#x27;t we creating NEWSPEAK and dragging along the worst aspects of it? The ambiguity and multi-valence of words is what creates more connections between words, increases the directionality of associations, and expands the potential subtlety and depth of meaning. By paring down (or requiring verifiability) we make it harder to say certain things, or at least make it harder to unintentionally say something that makes MORE or DEEPER sense than what we intended. If the token density becomes extreme, you&#x27;re left with something like a calculator.<p>Maybe this is the ultimate path toward better coding? But the worse path toward better genuine thinking?","title":null,"type":"comment","url":null},{"author":"zem","children":[{"author":"snthpy","children":[],"created_at":"2026-08-30T06:52:24.000Z","created_at_i":1788072744,"id":49496294,"options":[],"parent_id":49495478,"points":null,"story_id":49492632,"text":"Me too!","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T03:40:11.000Z","created_at_i":1788061211,"id":49495478,"options":[],"parent_id":49492632,"points":null,"story_id":49492632,"text":"I was briefly impressed that <a href=\"https:&#x2F;&#x2F;hylang.org&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;hylang.org&#x2F;</a> had released a 4.0 version!","title":null,"type":"comment","url":null},{"author":"xeonax","children":[],"created_at":"2026-08-30T06:57:23.000Z","created_at_i":1788073043,"id":49496314,"options":[],"parent_id":49492632,"points":null,"story_id":49492632,"text":"We humans should adopt grug, instead of claudish. It seems simple to understand. And has this melancholical feeling","title":null,"type":"comment","url":null},{"author":"realty_geek","children":[],"created_at":"2026-08-30T07:56:25.000Z","created_at_i":1788076585,"id":49496646,"options":[],"parent_id":49492632,"points":null,"story_id":49492632,"text":"Has anyone tried CodeBuddy?  Is it worth trying out?","title":null,"type":"comment","url":null},{"author":"DarmokTanagra","children":[{"author":"jimmydoe","children":[{"author":"DarmokTanagra","children":[],"created_at":"2026-08-30T11:24:30.000Z","created_at_i":1788089070,"id":49497722,"options":[],"parent_id":49497681,"points":null,"story_id":49492632,"text":"I think you misunderstood what I meant by &quot;Star Wars&quot;.<p><a href=\"https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Strategic_Defense_Initiative\" rel=\"nofollow\">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Strategic_Defense_Initiative</a>","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T11:16:44.000Z","created_at_i":1788088604,"id":49497681,"options":[],"parent_id":49496994,"points":null,"story_id":49492632,"text":"it&#x27;s not Star Wars. china doesn&#x27;t have to start any war, just let us finish (or get finished in) the wars...","title":null,"type":"comment","url":null},{"author":"bearjaws","children":[{"author":"DarmokTanagra","children":[],"created_at":"2026-08-30T15:30:21.000Z","created_at_i":1788103821,"id":49499571,"options":[],"parent_id":49499455,"points":null,"story_id":49492632,"text":"Speaking as someone living in Asia who visits the US every year or so, you are falling more and more behind and I don\u2019t see a way for you to catch up without a drastic societal reformation.<p>The EV market alone should scare the average citizen, but it seems like a non issue every time I bring it up.<p>I think the US economy has painted itself into a corner and this all in bet on AI is its last real play before the house of cards collapses.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T15:18:40.000Z","created_at_i":1788103120,"id":49499455,"options":[],"parent_id":49496994,"points":null,"story_id":49492632,"text":"China basically has been running a &quot;do nothing, win anyway&quot; campaign since the Trump era.<p>More and more people view the US as a bad ally, and China is stepping in to help in many places.<p>People should get out more, go visit South America, Dominican Republic, etc. They have cheaper AC, cheaper cars, cheaper appliances, all because they can import from China without tariffs. This is not to say they have worse quality, I would argue the opposite, much of what they import is at least as good or better than what you get in the USA.<p>It will be no surprise to me when we end up forced to use equal or worse American AI for a premium in price, while the rest of the world moves on.","title":null,"type":"comment","url":null}],"created_at":"2026-08-30T09:09:01.000Z","created_at_i":1788080941,"id":49496994,"options":[],"parent_id":49492632,"points":null,"story_id":49492632,"text":"How funny will it be when the US economy collapses because of all the money poured into AI at the expense of pretty much everything else only for China to come out ahead anyways.<p>Personally I hope this is China&#x27;s &quot;Star Wars&quot; moment, the current US admin certainly seems easy enough to manipulate into catastrophic own goals.","title":null,"type":"comment","url":null},{"author":"formvoltron","children":[],"created_at":"2026-08-30T17:41:00.000Z","created_at_i":1788111660,"id":49500910,"options":[],"parent_id":49492632,"points":null,"story_id":49492632,"text":"so it used enflame hardware?<p>The policy of the US to block Nvidia should have been called. The &quot;Let a thousand flowers bloom&quot; executive directive.<p>A very stable genius","title":null,"type":"comment","url":null},{"author":"zyralab","children":[],"created_at":"2026-08-30T18:10:40.000Z","created_at_i":1788113440,"id":49501208,"options":[],"parent_id":49492632,"points":null,"story_id":49492632,"text":"Seeing the model use &quot;caveman speak&quot; in its thoughts to save tokens is hilarious.<p>&quot;Why use many word when few word do trick?&quot; is actually a legit tech optimization now!","title":null,"type":"comment","url":null},{"author":"scirob","children":[],"created_at":"2026-08-30T20:36:53.000Z","created_at_i":1788122213,"id":49502506,"options":[],"parent_id":49492632,"points":null,"story_id":49492632,"text":"Just got benchmarks done for German langauge eval index i help maintain.  Hy4 is a big improvement over h3  but still below Deepseek pro and Significantly below  GLM 5.3 Flash .<p>hy4  ranks  ~14th overall<p><a href=\"https:&#x2F;&#x2F;dach.peerbench.ai&#x2F;compare?models=tencent%2Fhy4-preview,z-ai%2Fglm-5.3-flash,google%2Fgemini-3.7-flash,deepseek%2Fdeepseek-v4-pro-0813\" rel=\"nofollow\">https:&#x2F;&#x2F;dach.peerbench.ai&#x2F;compare?models=tencent%2Fhy4-previ...</a><p><a href=\"https:&#x2F;&#x2F;dach.peerbench.ai&#x2F;models&#x2F;tencent&#x2F;hy4-preview\" rel=\"nofollow\">https:&#x2F;&#x2F;dach.peerbench.ai&#x2F;models&#x2F;tencent&#x2F;hy4-preview</a>","title":null,"type":"comment","url":null}],"created_at":"2026-08-29T19:33:23.000Z","created_at_i":1788032003,"id":49492632,"options":[],"parent_id":null,"points":377,"story_id":49492632,"text":null,"title":"Hy4 preview","type":"story","url":"https://www.tencent.com/tencent-releases-and-open-sources-tencent-hy4-preview/"}
