{"exhaustive":{"nbHits":false,"typo":false},"exhaustiveNbHits":false,"exhaustiveTypo":false,"hits":[{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"scrlk"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["fable","5","claude"],"value":"System Card: <em>Claude</em> <em>Fable</em> <em>5</em> and <em>Claude</em> Mythos <em>5</em> [pdf]"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf"}},"_tags":["story","author_scrlk","story_48463811"],"author":"scrlk","children":[48466118],"created_at":"2026-06-09T16:58:13Z","created_at_i":1781024293,"num_comments":1,"objectID":"48463811","points":213,"story_id":48463811,"title":"System Card: Claude Fable 5 and Claude Mythos 5 [pdf]","updated_at":"2026-06-11T10:56:57Z","url":"https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"flyaway123"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["fable","5","claude"],"value":"Degraded performance for <em>Claude</em> Mythos <em>5</em>, <em>Claude</em> <em>Fable</em> <em>5</em>, and <em>Claude</em> Opus <em>5</em>"},"url":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["claude"],"value":"https://status.<em>claude</em>.com/incidents/f6gkkq6txl7z"}},"_tags":["story","author_flyaway123","story_49179554"],"author":"flyaway123","children":[49180459],"created_at":"2026-08-05T07:10:52Z","created_at_i":1785913852,"num_comments":0,"objectID":"49179554","points":6,"story_id":49179554,"title":"Degraded performance for Claude Mythos 5, Claude Fable 5, and Claude Opus 5","updated_at":"2026-08-05T20:58:08Z","url":"https://status.claude.com/incidents/f6gkkq6txl7z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"marciob"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["fable","5","claude"],"value":"Stop micromanaging <em>Fable</em>/<em>Claude</em> <em>5</em>"},"url":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["fable"],"value":"https://shumer.dev/how-i-prompt-<em>fable</em>"}},"_tags":["story","author_marciob","story_49174806"],"author":"marciob","created_at":"2026-08-04T20:41:41Z","created_at_i":1785876101,"num_comments":0,"objectID":"49174806","points":3,"story_id":49174806,"title":"Stop micromanaging Fable/Claude 5","updated_at":"2026-08-05T17:26:50Z","url":"https://shumer.dev/how-i-prompt-fable"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"meowokIknewit"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["fable","5","claude"],"value":"Countdown calendar for <em>Fable</em> <em>5</em> leaving <em>Claude</em> Code"},"url":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["fable"],"value":"https://<em>fable</em>-farewell-calendar.pages.dev/"}},"_tags":["story","author_meowokIknewit","story_48500432"],"author":"meowokIknewit","children":[48500433,48500466,48500626],"created_at":"2026-06-12T05:57:00Z","created_at_i":1781243820,"num_comments":2,"objectID":"48500432","points":2,"story_id":48500432,"title":"Countdown calendar for Fable 5 leaving Claude Code","updated_at":"2026-06-12T06:36:32Z","url":"https://fable-farewell-calendar.pages.dev/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"2020science"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["fable","5","claude"],"value":"I'm a professor of advanced technology transitions at ASU, and part of my work includes pushing at the edges of what frontier models can and can't do. Out of curiosity I asked <em>Fable</em> <em>5</em>, via <em>Claude</em> Code, to research the past two decades of my research and design a browser game from it \u2014 something that would stand up as a game on its own, but that genuinely reflected the work.<p>Over 90 iterations later (this was a long collaborative effort) the result was Hyperbubble \u2014 a one-button browser-based game where your task is to navigate your character through a future of emerging technologies, complex risks, transformative possibilities, and unexpected delights, all while protecting and growing their state of \u201cflourishing.\u201d Each element of the game tracks back to aspects of my work \u2014 even the seemingly weird ones!<p>Sharing as an example of how emerging AI capabilities are allowing complex ideas to be translated into modes and platforms that would have been exceptionally difficult to achieve without them (or prohibitively costly).<p>The game itself is one self-contained HTML file \u2014 procedural graphics, synthesized soundtrack, no libraries, no build step, no signup. MIT licensed. The leaderboard is a ~120-line Cloudflare Worker holding five fields per entry and nothing personal: no accounts, no emails, pseudonyms only. Best on desktop or tablet; playing on mobile works but isn't tuned yet, which is the next job.<p>Repo at: <a href=\"https://github.com/2020science/hyperbubble\" rel=\"nofollow\">https://github.com/2020science/hyperbubble</a> (which also links to more background on the game)"},"title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["claude"],"value":"Show HN: I asked <em>Claude</em> to turn 20 years of my research into a game"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://playhyperbubble.com/"}},"_tags":["story","author_2020science","story_49305115","show_hn"],"author":"2020science","children":[49353189],"created_at":"2026-08-14T22:01:06Z","created_at_i":1786744866,"num_comments":0,"objectID":"49305115","points":1,"story_id":49305115,"story_text":"I&#x27;m a professor of advanced technology transitions at ASU, and part of my work includes pushing at the edges of what frontier models can and can&#x27;t do. Out of curiosity I asked Fable 5, via Claude Code, to research the past two decades of my research and design a browser game from it \u2014 something that would stand up as a game on its own, but that genuinely reflected the work.<p>Over 90 iterations later (this was a long collaborative effort) the result was Hyperbubble \u2014 a one-button browser-based game where your task is to navigate your character through a future of emerging technologies, complex risks, transformative possibilities, and unexpected delights, all while protecting and growing their state of \u201cflourishing.\u201d Each element of the game tracks back to aspects of my work \u2014 even the seemingly weird ones!<p>Sharing as an example of how emerging AI capabilities are allowing complex ideas to be translated into modes and platforms that would have been exceptionally difficult to achieve without them (or prohibitively costly).<p>The game itself is one self-contained HTML file \u2014 procedural graphics, synthesized soundtrack, no libraries, no build step, no signup. MIT licensed. The leaderboard is a ~120-line Cloudflare Worker holding five fields per entry and nothing personal: no accounts, no emails, pseudonyms only. Best on desktop or tablet; playing on mobile works but isn&#x27;t tuned yet, which is the next job.<p>Repo at: <a href=\"https:&#x2F;&#x2F;github.com&#x2F;2020science&#x2F;hyperbubble\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;2020science&#x2F;hyperbubble</a> (which also links to more background on the game)","title":"Show HN: I asked Claude to turn 20 years of my research into a game","updated_at":"2026-08-18T21:49:36Z","url":"https://playhyperbubble.com/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"CharlieDigital"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["fable","5","claude"],"value":"<p><pre><code>    &gt; The model matters more than the harness anyway\n    &gt; \n    &gt; Everyone is benchmaxxing\n    &gt; \n    &gt; ...harnesses tend to be chosen on voodoo and hunches...\n</code></pre>\nI get what you're saying, but their graphic on performance here uses the exact same model with different harnesses and definitively shows that there is a significant difference in both accuracy and cost.  The whole point of their technical implementation and design decision here is to highlight that it's not &quot;voodoo and hunches&quot;, but observable data.<p><em>Fable</em> <em>5</em> on <em>Claude</em> Code scored 61.8% at a cost of $248.05 while <em>Fable</em> <em>5</em> on OpenCode beat it at 66.3% at $73.42.  The same model, the same benchmark; only the harness is different with a ~<em>5</em> point difference in accuracy while costing significantly less.  So if we are to believe the author and these results are repeatable, then it would seem that the harness matters.<p>The point of this framing here is specifically to address 1) benchmaxxing by using the same model, 2) NOT choose a harness on &quot;voodoo and hunches&quot; by using actual data to back the assertions.  Your comment feels misguided and completely hand waves the actual data points here."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Strands Harness"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://strandsagents.com/blog/introducing-strands-harness/"}},"_tags":["comment","author_CharlieDigital","story_49817289"],"author":"CharlieDigital","comment_text":"<p><pre><code>    &gt; The model matters more than the harness anyway\n    &gt; \n    &gt; Everyone is benchmaxxing\n    &gt; \n    &gt; ...harnesses tend to be chosen on voodoo and hunches...\n</code></pre>\nI get what you&#x27;re saying, but their graphic on performance here uses the exact same model with different harnesses and definitively shows that there is a significant difference in both accuracy and cost.  The whole point of their technical implementation and design decision here is to highlight that it&#x27;s not &quot;voodoo and hunches&quot;, but observable data.<p>Fable 5 on Claude Code scored 61.8% at a cost of $248.05 while Fable 5 on OpenCode beat it at 66.3% at $73.42.  The same model, the same benchmark; only the harness is different with a ~5 point difference in accuracy while costing significantly less.  So if we are to believe the author and these results are repeatable, then it would seem that the harness matters.<p>The point of this framing here is specifically to address 1) benchmaxxing by using the same model, 2) NOT choose a harness on &quot;voodoo and hunches&quot; by using actual data to back the assertions.  Your comment feels misguided and completely hand waves the actual data points here.","created_at":"2026-09-23T18:01:12Z","created_at_i":1790186472,"objectID":"49820055","parent_id":49818089,"story_id":49817289,"story_title":"Strands Harness","story_url":"https://strandsagents.com/blog/introducing-strands-harness/","updated_at":"2026-09-25T16:42:12Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"no_multitudes"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["fable","5","claude"],"value":"I drew a woman born in 715 CE in the Yangtze Basin. It claims that 96% of women were married, the average age of marriage was 18, and 44% of women died before age 15. Clearly these facts can't all be true simultaneously.<p>The firt citation for marriage data is listed as &quot;Hajnal (RH31)&quot;, but links to a seemingly unrelated WorldCat search. The second citation is to &quot;Kaplan (RH03)&quot;, which I was able to track down on sci-hub. It is a broad theory of human evolution that doesn't appear to mention the Yangtze Basin or China.<p>Another citation is to &quot;[RH109] Model-supplied gap-fill (<em>Claude</em> <em>Fable</em> <em>5</em> and <em>Claude</em> Opus <em>5</em>, 2026-08-05). Bounds and items written to close gaps no dataset we hold covers. Not a published work: see each row's `basis` for its stated reason. (2026)&quot; -- this is an interesting way to describe having an LLM hallucinate something for you.<p>Please stop making vibe-coded websites -- it is not useful to provide incorrect facts. You are polluting the commons with garbage."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Any Human Ever \u2013 One life, drawn at random from all who have ever lived"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://anyhumanever.com/"}},"_tags":["comment","author_no_multitudes","story_49550698"],"author":"no_multitudes","children":[49556161,49556327,49556429,49556713,49556747,49557189,49557298,49557468,49557574,49558029,49558231,49558283,49558926,49558942,49561534,49561853,49562187,49562529,49562541,49563220,49565437],"comment_text":"I drew a woman born in 715 CE in the Yangtze Basin. It claims that 96% of women were married, the average age of marriage was 18, and 44% of women died before age 15. Clearly these facts can&#x27;t all be true simultaneously.<p>The firt citation for marriage data is listed as &quot;Hajnal (RH31)&quot;, but links to a seemingly unrelated WorldCat search. The second citation is to &quot;Kaplan (RH03)&quot;, which I was able to track down on sci-hub. It is a broad theory of human evolution that doesn&#x27;t appear to mention the Yangtze Basin or China.<p>Another citation is to &quot;[RH109] Model-supplied gap-fill (Claude Fable 5 and Claude Opus 5, 2026-08-05). Bounds and items written to close gaps no dataset we hold covers. Not a published work: see each row&#x27;s `basis` for its stated reason. (2026)&quot; -- this is an interesting way to describe having an LLM hallucinate something for you.<p>Please stop making vibe-coded websites -- it is not useful to provide incorrect facts. You are polluting the commons with garbage.","created_at":"2026-09-03T20:05:37Z","created_at_i":1788465937,"objectID":"49556069","parent_id":49550698,"story_id":49550698,"story_title":"Any Human Ever \u2013 One life, drawn at random from all who have ever lived","story_url":"https://anyhumanever.com/","updated_at":"2026-09-11T20:25:44Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"Topfi"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["fable","5","claude"],"value":"I have a settings panel implemented in HTML/CSS/JS for a Firefox fork that &quot;could/should have been a desktop environment&quot;. Bit of an odd project really, mainly out of a very specific conviction concerning modern applications, the way LLMs and task specific models are currently not leveraged well by any browser, my own tendency to have 400+ tabs open at a time across multiple projects, my opinion that it is the perfect place to finally apply a lot of UX opinions I have held for a while and push in a very distinct direction along with core critiques I have concerning PKM applications I haven't seen addressed despite trying every PKM application under the sun. Neither here, nor there.<p>So this &quot;thing&quot; is mainly a Firefox fork and most UI is basic HTML/CSS/JS (as is the case in upstream). Development is patch baked, CSS tokens must follow a defined and CI enforced standard, etc. LLMs can be very helpful in development, I got a small CLI tool for patch, token management and basic quality gates, which I started working on a few months ago to keep the most atrocious LLM output at bay. Has lead to the revieability of output improving meaningfully over markdown monstrosities, though OpenAI models still manage to sneak hard to parse output past it. This CLI tool along with some task specific scripts also ensures reuse of proven upstream infra like Places (Good lord dear Firefox developers, is Places nice to rely on), consistent regression testing (especially in memory constraint scenarios), etc. Basically, I can and do regularly make additions with LLM assistance, I review it, I discard and restart or improve upon it (rarely accept scoped changes wholesale. This to say, I got some experience in the use of models for coding assistance and I (thanks to the amazing docs and a lot of considerations for the architecture I want) do know what I want, how I want it and how to get there. Also got private LLM evals that often uncover which labs tend to perform suspiciously well in public benchmarks vs private ones and what models still struggle with along with why, so yeah, certainly can always improve but I got, I'd argue, enough of an idea to where my critique of LLM coding limitations has legs.<p>Which brings us to what I was trying to implement and how I went about it: Settings works. Fully featured (including a few cross-site-tracking specific clarifications that came from a HN interaction a few days ago), tab specific previews for what changes affect regarding themeing, well tested (manual and static), integrated to leverage what FF provides where possible.<p>It does (or rather did) look functional/God awful though. To the point where I was uncertain that certain previews could be easily parsed by new users. I thus opened Adobe XD, did some early mockup work, tried a few core concepts, settled upon two, then (using <em>Claude</em> <em>Fable</em> <em>5</em> low) created a plain export of the existing settings code from our furnace components and patch baked edits into regular HTML/JS/CSS files. I manually verified, this export worked, the tokens were in the correct format, the code reflected what Hominis applied (including what was required for stand-alone of course) and externally called features upon interaction did provide log output linking to the pre-existing functions that meant reimplementation based upon this should be easy.<p>I then took that to <em>Claude</em> Design using <em>Fable</em> <em>5</em> on High. I provided the code along with linked branding files (which due to the way branding patches are handled were simpler to provide separately) and my Adobe XD mockups. A few dozen iterations later, along with some exports and re-imports due to manual changes (some animations in tabbing/&quot;focus mode&quot; showcases needed to be &quot;just so&quot; and prompting would have been inefficient to get there), I had a new user experience I was far happier with. Simpler, yet better at communicating, far more visually appealing and resolving some concerns I had, I felt pleased and will admit, <em>Fable</em> <em>5</em> via <em>Claude</em> Design provided valuable output and did, what it does best, make iterating on multiple UI concepts next to each other to settle on a final option from many, far quicker.<p>I then exported and took that to GPT-<em>5</em>.6 Luna (I have \u20ac 23,- Codex only so am a bit stingy on when to use what). But so what? I had verified, the tokens were the same. The naming of elements remained consistent to what Hominis Settings used, the backend changes were practically none-existent. I had audited the output end-to-end, made some refactors and house style specific improvements to keep everything more auditable, everything seemed suited for a quick port. What could possibly go wrong?<p>Anyone whith pattern recognition will likely guess what. Basic 1:1 applying? No dice. The first attempt failed as, once the context window had compacted twice, the model started leaving the very clearly paved path laid out. Stylised favicon in the showcase? Gone. Hamburger menu in the showcase, compressed. Vertical tabbing change interlinked with the canvas section? Very funny. The model started no longer following the code, it started taking screenshots and applying what it could see from that, despite the original prompt (just checked) vey clearly stating a simple code port, section per section, with any deviations to be listed in a designated file I maintain for long running tasks.<p>Basically, Luna did implement changes to the settings that felt tangentially right and a casual observe may not notice all the regressions and deviations, but I did. So I stopped it.<p>Sol and <em>Fable</em> didn't fare much better. Sol did stay on target longer, but it went off the rails around the privacy tab, introducing functional regressions to the way I had implemented cross-site cookie blocking, which were never requested, nor should that code even have been looked at. I reset the repo and handed it over to <em>Fable</em> <em>5</em> (medium). I had a third of my weekly usage left on 20x Max, reset the day after at 3AM so no harm either way.<p>Should be plenty. Wasn't plenty. Since a while (I think Opus 4.7, but could be wrong), Anthropic models do decently well regarding long term, high token tasks. Up to 450k, I have been able to reliably reproduce consistent implementation. The model, using a few subagents (which should have reduced the risk of context window issues further), went to work and after a few hours (and about 20% of usage less), the model proudly presented its work. I was at work and by the time I came back, I was a bit miffed to find that the model had, in its wisdom, decided to not used the well established and consistently used mar to bind in branding icons. No biggie, easy fix, albeit a bit stupid. ESPECIALLY SINCE I SAW IN THE <em>CLAUDE</em> CODE TRACES THAT THE MODEL HAD SURPRESSED A WARNING ON THAT VERY FRONT. Whatever. Then I saw it had not wired in the existing browser data deletion and export logic. It hadn\u2019t modified existing logic unlike Sol, so hey, that\u2019s nice. But it had not wired up the existing settings when they did not have any immediate feedback in the implementation reference.<p>Ox Alpha, it just spanned in circles, didn\u2019t seem to like our fireforge CLI and furnace componets, but it was worth a free try. Opus <em>5</em>, the model most obsessive in checking its own work, took screenshots. A lot of sscreenshots including every few hundred ms to cover animations. Nice. BUT IT CREATED ITS OWN TOKENS INSTAD OF REUSING WHAT WAS PROVIDED. Thus, styling deviated heavily.<p>At this point you might ask why I don\u2019t do it manually and I will in the end anyways, but I was surprised to find such a clear case of a seemingly straightforward task flummoxing multiple LLMs. This is aided by my unique code base (the upstream FF code is also gitignored which likely flummoxes some models trained heavily to leverage git to track changes), everything needs to be patch backed and follow a specific implementation style, etc. But I had more important things to do and I wanted to see whether I couldn\u2019t get it to work yet.<p>Inspired by Opus <em>5</em>, I wrote a new prompt, specifically laying out a visual comparison and code diff workflow. Only these changes, only in this manner, only move on ones you have gotten visual confirmation, specific cross checks. I included a hand written markdown outlining which change affects other settings sections (even though that is obvious reading the reference code), how to approach tokens, etc. Obsessively descriptive and (I feel) unnecessarily so, but why not. Best case, it works, worst case, I\u2019ll spend an hour doing it manually. I had other things to do not behind a keyboard, so why not one last Hail Mary.<p><em>Fable</em> <em>5</em>, ever efficient when using visuals, used the last rest of my usage, though I did see some roundabout approaches after the fact that make me doubtful it\u2019d have cracked this. Opus <em>5</em> went off the deep end taking ui-captures across the entire code base, which lead to a very liberal application of settings tokens outside settings.<p>Sol did take a night and got 40% there when I asked for a pause once the in flight slice had landed. It did port the UI/UX changes in a way that on the surface looked and felt correct. It did not touch the backend in unacceptable ways. And it did cross checks. Animations also behaved correctly, though it did apply a rule on backend usage a bit to strictly, incorporating that into a preview for search by turning that into an actual web search, not a UX demo. Dumb, but not fatal.<p>Great success, what am I complaining?<p>Well, the code. It had done what Sol likes to do and turned very cleanly written, readable code into a hard to parse mess. This included touching existing test files.<p>And at that point I said \u201cfuck it, I\u2019ll do it myself\u201d. And I did. In less than an hour, listening to Paris Palamo, Lyre Le Temps, Sting, Sade, SynthV and some Nirvana.<p>If I didn\u2019t look at the code and I didn\u2019t have strict standards for the UI, but just considered what looks in line on the surface level/feels right/\u201cvibes\u201d and what \u201cworks\u201d, many of these attempts would have been accepted, as their issues are rarely apparent on the surface. That\u2019s part of the issue in my book and why I\u2019m firm we are far from \u201cdon\u2019t read code\u201d/\u201cdon\u2019t test\u201d/\u201cskip qa\u201d\u2026"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Debian votes to allow \"responsible use of generative AI\""},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://lwn.net/Articles/1091231/"}},"_tags":["comment","author_Topfi","story_49489982"],"author":"Topfi","comment_text":"I have a settings panel implemented in HTML&#x2F;CSS&#x2F;JS for a Firefox fork that &quot;could&#x2F;should have been a desktop environment&quot;. Bit of an odd project really, mainly out of a very specific conviction concerning modern applications, the way LLMs and task specific models are currently not leveraged well by any browser, my own tendency to have 400+ tabs open at a time across multiple projects, my opinion that it is the perfect place to finally apply a lot of UX opinions I have held for a while and push in a very distinct direction along with core critiques I have concerning PKM applications I haven&#x27;t seen addressed despite trying every PKM application under the sun. Neither here, nor there.<p>So this &quot;thing&quot; is mainly a Firefox fork and most UI is basic HTML&#x2F;CSS&#x2F;JS (as is the case in upstream). Development is patch baked, CSS tokens must follow a defined and CI enforced standard, etc. LLMs can be very helpful in development, I got a small CLI tool for patch, token management and basic quality gates, which I started working on a few months ago to keep the most atrocious LLM output at bay. Has lead to the revieability of output improving meaningfully over markdown monstrosities, though OpenAI models still manage to sneak hard to parse output past it. This CLI tool along with some task specific scripts also ensures reuse of proven upstream infra like Places (Good lord dear Firefox developers, is Places nice to rely on), consistent regression testing (especially in memory constraint scenarios), etc. Basically, I can and do regularly make additions with LLM assistance, I review it, I discard and restart or improve upon it (rarely accept scoped changes wholesale. This to say, I got some experience in the use of models for coding assistance and I (thanks to the amazing docs and a lot of considerations for the architecture I want) do know what I want, how I want it and how to get there. Also got private LLM evals that often uncover which labs tend to perform suspiciously well in public benchmarks vs private ones and what models still struggle with along with why, so yeah, certainly can always improve but I got, I&#x27;d argue, enough of an idea to where my critique of LLM coding limitations has legs.<p>Which brings us to what I was trying to implement and how I went about it: Settings works. Fully featured (including a few cross-site-tracking specific clarifications that came from a HN interaction a few days ago), tab specific previews for what changes affect regarding themeing, well tested (manual and static), integrated to leverage what FF provides where possible.<p>It does (or rather did) look functional&#x2F;God awful though. To the point where I was uncertain that certain previews could be easily parsed by new users. I thus opened Adobe XD, did some early mockup work, tried a few core concepts, settled upon two, then (using Claude Fable 5 low) created a plain export of the existing settings code from our furnace components and patch baked edits into regular HTML&#x2F;JS&#x2F;CSS files. I manually verified, this export worked, the tokens were in the correct format, the code reflected what Hominis applied (including what was required for stand-alone of course) and externally called features upon interaction did provide log output linking to the pre-existing functions that meant reimplementation based upon this should be easy.<p>I then took that to Claude Design using Fable 5 on High. I provided the code along with linked branding files (which due to the way branding patches are handled were simpler to provide separately) and my Adobe XD mockups. A few dozen iterations later, along with some exports and re-imports due to manual changes (some animations in tabbing&#x2F;&quot;focus mode&quot; showcases needed to be &quot;just so&quot; and prompting would have been inefficient to get there), I had a new user experience I was far happier with. Simpler, yet better at communicating, far more visually appealing and resolving some concerns I had, I felt pleased and will admit, Fable 5 via Claude Design provided valuable output and did, what it does best, make iterating on multiple UI concepts next to each other to settle on a final option from many, far quicker.<p>I then exported and took that to GPT-5.6 Luna (I have \u20ac 23,- Codex only so am a bit stingy on when to use what). But so what? I had verified, the tokens were the same. The naming of elements remained consistent to what Hominis Settings used, the backend changes were practically none-existent. I had audited the output end-to-end, made some refactors and house style specific improvements to keep everything more auditable, everything seemed suited for a quick port. What could possibly go wrong?<p>Anyone whith pattern recognition will likely guess what. Basic 1:1 applying? No dice. The first attempt failed as, once the context window had compacted twice, the model started leaving the very clearly paved path laid out. Stylised favicon in the showcase? Gone. Hamburger menu in the showcase, compressed. Vertical tabbing change interlinked with the canvas section? Very funny. The model started no longer following the code, it started taking screenshots and applying what it could see from that, despite the original prompt (just checked) vey clearly stating a simple code port, section per section, with any deviations to be listed in a designated file I maintain for long running tasks.<p>Basically, Luna did implement changes to the settings that felt tangentially right and a casual observe may not notice all the regressions and deviations, but I did. So I stopped it.<p>Sol and Fable didn&#x27;t fare much better. Sol did stay on target longer, but it went off the rails around the privacy tab, introducing functional regressions to the way I had implemented cross-site cookie blocking, which were never requested, nor should that code even have been looked at. I reset the repo and handed it over to Fable 5 (medium). I had a third of my weekly usage left on 20x Max, reset the day after at 3AM so no harm either way.<p>Should be plenty. Wasn&#x27;t plenty. Since a while (I think Opus 4.7, but could be wrong), Anthropic models do decently well regarding long term, high token tasks. Up to 450k, I have been able to reliably reproduce consistent implementation. The model, using a few subagents (which should have reduced the risk of context window issues further), went to work and after a few hours (and about 20% of usage less), the model proudly presented its work. I was at work and by the time I came back, I was a bit miffed to find that the model had, in its wisdom, decided to not used the well established and consistently used mar to bind in branding icons. No biggie, easy fix, albeit a bit stupid. ESPECIALLY SINCE I SAW IN THE CLAUDE CODE TRACES THAT THE MODEL HAD SURPRESSED A WARNING ON THAT VERY FRONT. Whatever. Then I saw it had not wired in the existing browser data deletion and export logic. It hadn\u2019t modified existing logic unlike Sol, so hey, that\u2019s nice. But it had not wired up the existing settings when they did not have any immediate feedback in the implementation reference.<p>Ox Alpha, it just spanned in circles, didn\u2019t seem to like our fireforge CLI and furnace componets, but it was worth a free try. Opus 5, the model most obsessive in checking its own work, took screenshots. A lot of sscreenshots including every few hundred ms to cover animations. Nice. BUT IT CREATED ITS OWN TOKENS INSTAD OF REUSING WHAT WAS PROVIDED. Thus, styling deviated heavily.<p>At this point you might ask why I don\u2019t do it manually and I will in the end anyways, but I was surprised to find such a clear case of a seemingly straightforward task flummoxing multiple LLMs. This is aided by my unique code base (the upstream FF code is also gitignored which likely flummoxes some models trained heavily to leverage git to track changes), everything needs to be patch backed and follow a specific implementation style, etc. But I had more important things to do and I wanted to see whether I couldn\u2019t get it to work yet.<p>Inspired by Opus 5, I wrote a new prompt, specifically laying out a visual comparison and code diff workflow. Only these changes, only in this manner, only move on ones you have gotten visual confirmation, specific cross checks. I included a hand written markdown outlining which change affects other settings sections (even though that is obvious reading the reference code), how to approach tokens, etc. Obsessively descriptive and (I feel) unnecessarily so, but why not. Best case, it works, worst case, I\u2019ll spend an hour doing it manually. I had other things to do not behind a keyboard, so why not one last Hail Mary.<p>Fable 5, ever efficient when using visuals, used the last rest of my usage, though I did see some roundabout approaches after the fact that make me doubtful it\u2019d have cracked this. Opus 5 went off the deep end taking ui-captures across the entire code base, which lead to a very liberal application of settings tokens outside settings.<p>Sol did take a night and got 40% there when I asked for a pause once the in flight slice had landed. It did port the UI&#x2F;UX changes in a way that on the surface looked and felt correct. It did not touch the backend in unacceptable ways. And it did cross checks. Animations also behaved correctly, though it did apply a rule on backend usage a bit to strictly, incorporating that into a preview for search by turning that into an actual web search, not a UX demo. Dumb, but not fatal.<p>Great success, what am I complaining?<p>Well, the code. It had done what Sol likes to do and turned very cleanly written, readable code into a hard to parse mess. This included touching existing test files.<p>And at that point I said \u201cfuck it, I\u2019ll do it myself\u201d. And I did. In less than an hour, listening to Paris Palamo, Lyre Le Temps, Sting, Sade, SynthV and some Nirvana.<p>If I didn\u2019t look at the code and I didn\u2019t have strict standards for the UI, but just considered what looks in line on the surface level&#x2F;feels right&#x2F;\u201cvibes\u201d and what \u201cworks\u201d, many of these attempts would have been accepted, as their issues are rarely apparent on the surface. That\u2019s part of the issue in my book and why I\u2019m firm we are far from \u201cdon\u2019t read code\u201d&#x2F;\u201cdon\u2019t test\u201d&#x2F;\u201cskip qa\u201d\u2026","created_at":"2026-08-29T19:46:33Z","created_at_i":1788032793,"objectID":"49492740","parent_id":49491618,"story_id":49489982,"story_title":"Debian votes to allow \"responsible use of generative AI\"","story_url":"https://lwn.net/Articles/1091231/","updated_at":"2026-08-29T21:07:04Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"igravious"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["fable","5","claude"],"value":"I like Grok 4.6 the model, I like it for a number of reasons. I like Grok Build too.<p>However, I've pointed Grok 4.6 at a fairly complex codebase and asked it to review/audit it for issues and it's come back with a whole laundry list of issues. I've passed that list to Kimi and <em>Claude</em> and they both were like &quot;a couple of good catches but some of those are not issues at all&quot;. Grok 4.6 is noticeably weaker that <em>Claude</em> <em>Fable</em> <em>5</em>, <em>Claude</em> Opus <em>5</em>, <em>Claude</em> Opus 4.8, Kimi K3, GLM <em>5</em>.3, \u2026 your suggestion to use Grok 4.6 instead of a recent <em>Claude</em> doesn't pass empirical scrutiny."},"story_title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["claude"],"value":"Claudette: Make <em>Claude</em> stop talking like a BuzzFeed article"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/adnanakil/nobuzz/blob/main/README.md"}},"_tags":["comment","author_igravious","story_49388752"],"author":"igravious","comment_text":"I like Grok 4.6 the model, I like it for a number of reasons. I like Grok Build too.<p>However, I&#x27;ve pointed Grok 4.6 at a fairly complex codebase and asked it to review&#x2F;audit it for issues and it&#x27;s come back with a whole laundry list of issues. I&#x27;ve passed that list to Kimi and Claude and they both were like &quot;a couple of good catches but some of those are not issues at all&quot;. Grok 4.6 is noticeably weaker that Claude Fable 5, Claude Opus 5, Claude Opus 4.8, Kimi K3, GLM 5.3, \u2026 your suggestion to use Grok 4.6 instead of a recent Claude doesn&#x27;t pass empirical scrutiny.","created_at":"2026-08-23T07:54:02Z","created_at_i":1787471642,"objectID":"49406886","parent_id":49402753,"story_id":49388752,"story_title":"Claudette: Make Claude stop talking like a BuzzFeed article","story_url":"https://github.com/adnanakil/nobuzz/blob/main/README.md","updated_at":"2026-08-23T07:55:05Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"simonw"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["fable","5","claude"],"value":"I have a folder where I rebuild these as a git commit history so you can more easily see what has changed: <a href=\"https://github.com/simonw/research/commits/main/extract-system-prompts\" rel=\"nofollow\">https://github.com/simonw/research/commits/main/extract-syst...</a><p>For example here's what changed between Opus 4.8 and Opus <em>5</em>: <a href=\"https://github.com/simonw/research/commit/a2de185cc367eb66c2e27090d9ff0f766d335ff4\" rel=\"nofollow\">https://github.com/simonw/research/commit/a2de185cc367eb66c2...</a><p>The most interesting addition to the prompt from that diff is this bit:<p>&gt; <em>Claude</em> <em>Fable</em> <em>5</em> and <em>Claude</em> Mythos <em>5</em> were first released on June 9, 2026. On June 12, 2026, Anthropic suspended access to both models to comply with U.S. Department of Commerce export controls; the Department lifted those controls on June 30, 2026, and Anthropic restored access on July 1, 2026 (Anthropic's statement: [<a href=\"https://www.anthropic.com/news/fable-mythos-access\" rel=\"nofollow\">https://www.anthropic.com/news/<em>fable</em>-mythos-access</a>](<a href=\"https://www.anthropic.com/news/fable-mythos-access\" rel=\"nofollow\">https://www.anthropic.com/news/<em>fable</em>-mythos-access</a>)). These events are after <em>Claude</em>'s training-data cutoff, so <em>Claude</em> knows about them only from this notice. If asked, <em>Claude</em> confirms them accurately and matter-of-factly \u2014 it doesn't deny the suspension happened \u2014 and otherwise treats the export controls like any other current political topic: it gives a fair, accurate account rather than sharing personal opinions, and points to the linked statement for anything further. Things may have developed since this notice, so <em>Claude</em> checks for newer information when it can search, and otherwise suggests checking Anthropic's site.<p>One frustrating note about this page is that they share the system prompts used for <a href=\"https://claude.ai\" rel=\"nofollow\">https://<em>claude</em>.ai</a> and the <em>Claude</em> mobile apps regular chat, but they omit the tool definitions. Those are much more interesting if you want to understand what <em>Claude</em> can actually do for you. You can reconstruct them through prompting <em>Claude</em> directly but that's extra friction and risks refusals and hallucinations.<p>They also don't publish the <em>Claude</em> Code system prompts, which is silly because those are trivial to extract using a logging proxy."},"story_title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["claude"],"value":"<em>Claude</em>: System Prompts"},"story_url":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["claude"],"value":"https://platform.<em>claude</em>.com/docs/en/release-notes/system-prompts"}},"_tags":["comment","author_simonw","story_49319556"],"author":"simonw","children":[49320046,49320561,49323852,49325186,49326486,49332637,49333966],"comment_text":"I have a folder where I rebuild these as a git commit history so you can more easily see what has changed: <a href=\"https:&#x2F;&#x2F;github.com&#x2F;simonw&#x2F;research&#x2F;commits&#x2F;main&#x2F;extract-system-prompts\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;simonw&#x2F;research&#x2F;commits&#x2F;main&#x2F;extract-syst...</a><p>For example here&#x27;s what changed between Opus 4.8 and Opus 5: <a href=\"https:&#x2F;&#x2F;github.com&#x2F;simonw&#x2F;research&#x2F;commit&#x2F;a2de185cc367eb66c2e27090d9ff0f766d335ff4\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;simonw&#x2F;research&#x2F;commit&#x2F;a2de185cc367eb66c2...</a><p>The most interesting addition to the prompt from that diff is this bit:<p>&gt; Claude Fable 5 and Claude Mythos 5 were first released on June 9, 2026. On June 12, 2026, Anthropic suspended access to both models to comply with U.S. Department of Commerce export controls; the Department lifted those controls on June 30, 2026, and Anthropic restored access on July 1, 2026 (Anthropic&#x27;s statement: [<a href=\"https:&#x2F;&#x2F;www.anthropic.com&#x2F;news&#x2F;fable-mythos-access\" rel=\"nofollow\">https:&#x2F;&#x2F;www.anthropic.com&#x2F;news&#x2F;fable-mythos-access</a>](<a href=\"https:&#x2F;&#x2F;www.anthropic.com&#x2F;news&#x2F;fable-mythos-access\" rel=\"nofollow\">https:&#x2F;&#x2F;www.anthropic.com&#x2F;news&#x2F;fable-mythos-access</a>)). These events are after Claude&#x27;s training-data cutoff, so Claude knows about them only from this notice. If asked, Claude confirms them accurately and matter-of-factly \u2014 it doesn&#x27;t deny the suspension happened \u2014 and otherwise treats the export controls like any other current political topic: it gives a fair, accurate account rather than sharing personal opinions, and points to the linked statement for anything further. Things may have developed since this notice, so Claude checks for newer information when it can search, and otherwise suggests checking Anthropic&#x27;s site.<p>One frustrating note about this page is that they share the system prompts used for <a href=\"https:&#x2F;&#x2F;claude.ai\" rel=\"nofollow\">https:&#x2F;&#x2F;claude.ai</a> and the Claude mobile apps regular chat, but they omit the tool definitions. Those are much more interesting if you want to understand what Claude can actually do for you. You can reconstruct them through prompting Claude directly but that&#x27;s extra friction and risks refusals and hallucinations.<p>They also don&#x27;t publish the Claude Code system prompts, which is silly because those are trivial to extract using a logging proxy.","created_at":"2026-08-16T13:34:21Z","created_at_i":1786887261,"objectID":"49319926","parent_id":49319556,"story_id":49319556,"story_title":"Claude: System Prompts","story_url":"https://platform.claude.com/docs/en/release-notes/system-prompts","updated_at":"2026-09-25T10:12:10Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"gwbas1c"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["fable","5","claude"],"value":"When I tried <em>Claude</em> <em>5</em> (<em>Fable</em>?) (In Visual Studio via Copilot,) the results weren't as bad as a lot of the comments here... But it was super-slow. IE, so slow that I could code faster than it, negating the entire point of using AI to begin with!<p>I went back to Opus 4.8, but recently switched to GPT <em>5</em>.6 Luna. The results are comparable in quality, but it's much cheaper and much faster.<p>---<p>The thing with coding agents in a tool like Visual Studio is that the cost to switch is 0. There's no lock-in whatsoever. It makes it harder to justify the AI-first IDEs when the AT bolt-on IDEs make it so easy to pick the right model."},"story_title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["5"],"value":"Why does Opus <em>5</em> feel worse to work with?"},"story_url":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["5"],"value":"https://mun-logadan.github.io/why-does-opus-<em>5</em>-feel-worse/"}},"_tags":["comment","author_gwbas1c","story_49296740"],"author":"gwbas1c","children":[49298602],"comment_text":"When I tried Claude 5 (Fable?) (In Visual Studio via Copilot,) the results weren&#x27;t as bad as a lot of the comments here... But it was super-slow. IE, so slow that I could code faster than it, negating the entire point of using AI to begin with!<p>I went back to Opus 4.8, but recently switched to GPT 5.6 Luna. The results are comparable in quality, but it&#x27;s much cheaper and much faster.<p>---<p>The thing with coding agents in a tool like Visual Studio is that the cost to switch is 0. There&#x27;s no lock-in whatsoever. It makes it harder to justify the AI-first IDEs when the AT bolt-on IDEs make it so easy to pick the right model.","created_at":"2026-08-14T13:40:21Z","created_at_i":1786714821,"objectID":"49298551","parent_id":49296740,"story_id":49296740,"story_title":"Why does Opus 5 feel worse to work with?","story_url":"https://mun-logadan.github.io/why-does-opus-5-feel-worse/","updated_at":"2026-08-14T13:45:22Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"heaney-555"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["fable","5","claude"],"value":"This is insane. People often pontificate about how AI hallucinations could cause huge issues, but I guarantee you GPT <em>5</em>.6 Sol or <em>Claude</em> <em>5</em> <em>Fable</em> would have spotted this mistake if asked.<p>Human stupidity is a bigger threat than AI hallucination."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"A missing underscore sent innocent man to prison for 18 months"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://arstechnica.com/tech-policy/2026/07/police-missed-one-underscore-and-sent-the-wrong-man-to-prison/"}},"_tags":["comment","author_heaney-555","story_49076116"],"author":"heaney-555","comment_text":"This is insane. People often pontificate about how AI hallucinations could cause huge issues, but I guarantee you GPT 5.6 Sol or Claude 5 Fable would have spotted this mistake if asked.<p>Human stupidity is a bigger threat than AI hallucination.","created_at":"2026-07-30T22:10:17Z","created_at_i":1785449417,"objectID":"49116521","parent_id":49076116,"story_id":49076116,"story_title":"A missing underscore sent innocent man to prison for 18 months","story_url":"https://arstechnica.com/tech-policy/2026/07/police-missed-one-underscore-and-sent-the-wrong-man-to-prison/","updated_at":"2026-07-30T22:14:30Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"simonw"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["fable","5","claude"],"value":"GPT-<em>5</em>.3 is a last-generation model (which is wild to say when it came out in February this year, but it's  true).<p>How does GPT-<em>5</em>.6 Sol or <em>Claude</em> <em>Fable</em> <em>5</em> or <em>Claude</em> Opus <em>5</em> do on that Cinema 4D code?"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"What is happening to jobs? Separating AI hype from reality"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://siepr.stanford.edu/publications/policy-brief/what-really-happening-jobs-separating-ai-hype-reality"}},"_tags":["comment","author_simonw","story_49052570"],"author":"simonw","comment_text":"GPT-5.3 is a last-generation model (which is wild to say when it came out in February this year, but it&#x27;s  true).<p>How does GPT-5.6 Sol or Claude Fable 5 or Claude Opus 5 do on that Cinema 4D code?","created_at":"2026-07-26T12:59:56Z","created_at_i":1785070796,"objectID":"49057751","parent_id":49056105,"story_id":49052570,"story_title":"What is happening to jobs? Separating AI hype from reality","story_url":"https://siepr.stanford.edu/publications/policy-brief/what-really-happening-jobs-separating-ai-hype-reality","updated_at":"2026-07-27T12:46:02Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"skerit"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["fable","5","claude"],"value":"Interesting, they finally support `system` messages anywhere in a chat conversation:<p>&gt; Mid-conversation system messages are available on the <em>Claude</em> API, <em>Claude</em> in Amazon Bedrock, and Google Cloud.\n&gt;\n&gt; This feature is available on <em>Claude</em> <em>Fable</em> <em>5</em>, <em>Claude</em> Mythos <em>5</em>, <em>Claude</em> Opus 4.8, and <em>Claude</em> Opus <em>5</em>. No beta header is required. This feature is not available on <em>Claude</em> Sonnet <em>5</em>; use the top-level system field instead.<p>For nearly all models EXCEPT Sonnet <em>5</em>? That is weird.\nHow old is Sonnet <em>5</em> really?"},"story_title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["5","claude"],"value":"<em>Claude</em> Opus <em>5</em>"},"story_url":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["5","claude"],"value":"https://www.anthropic.com/news/<em>claude</em>-opus-<em>5</em>"}},"_tags":["comment","author_skerit","story_49038433"],"author":"skerit","comment_text":"Interesting, they finally support `system` messages anywhere in a chat conversation:<p>&gt; Mid-conversation system messages are available on the Claude API, Claude in Amazon Bedrock, and Google Cloud.\n&gt;\n&gt; This feature is available on Claude Fable 5, Claude Mythos 5, Claude Opus 4.8, and Claude Opus 5. No beta header is required. This feature is not available on Claude Sonnet 5; use the top-level system field instead.<p>For nearly all models EXCEPT Sonnet 5? That is weird.\nHow old is Sonnet 5 really?","created_at":"2026-07-24T17:31:01Z","created_at_i":1784914261,"objectID":"49039014","parent_id":49038433,"story_id":49038433,"story_title":"Claude Opus 5","story_url":"https://www.anthropic.com/news/claude-opus-5","updated_at":"2026-07-25T11:14:25Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"furyofantares"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["fable","5","claude"],"value":"I find <em>5</em>.6 Sol will pick a direction and aggressively pursue it in long horizon tasks. I've got it porting an older game from Pascal to my own game framework. I gave it some instructions on doing a full 1:1 port. I had already ported the game rules and multiplayer support to a very different system than the original, but all of the UI and features and such needed doing, and needed to be integrated into this very different system.<p>The first attempt it had files tracking both hashes and semantic hashes of every individual line of Pascal code, mapping to what code in the port is responsible for that line of pascal. It had written tooling to parse Pascal in service of this for some reason as well. I asked why it was doing this, it said it was because the reference code is .gitignore'd so it needs to thoroughly maintain the mapping in case someone working on it does not have the reference code, or in case the reference code changes.<p>I started over with <em>Claude</em> <em>5</em> <em>Fable</em>, and with better instructions about focusing on UI. I got a long ways with that before I hit my weekly limits, and switched back to <em>5</em>.6 Sol. It picked up and did a great job for a while, although it interpreted my desire for a 1:1 port to mean every pixel must be perfect. I let it go on and it did some good work in that regard, but then it decided it must perfectly reproduce a hash of the game state in various replays &amp; etc. It had clearly lost track that I didn't need game rules ported, and it found that the original code produces a hash of the gamestate for various purposes, so it ended up reproducing this in a game that represents its state totally differently. It also rolled its own version of Pascal's RNG source in order do this. I've burned through 3 weekly limit resets on this to see if it's actually going anywhere, and it has found some bugs, but man it is going hard in a direction I didn't even ask for."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"OpenAI and Hugging Face address security incident during model evaluation"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://openai.com/index/hugging-face-model-evaluation-security-incident/"}},"_tags":["comment","author_furyofantares","story_48997548"],"author":"furyofantares","children":[49002069],"comment_text":"I find 5.6 Sol will pick a direction and aggressively pursue it in long horizon tasks. I&#x27;ve got it porting an older game from Pascal to my own game framework. I gave it some instructions on doing a full 1:1 port. I had already ported the game rules and multiplayer support to a very different system than the original, but all of the UI and features and such needed doing, and needed to be integrated into this very different system.<p>The first attempt it had files tracking both hashes and semantic hashes of every individual line of Pascal code, mapping to what code in the port is responsible for that line of pascal. It had written tooling to parse Pascal in service of this for some reason as well. I asked why it was doing this, it said it was because the reference code is .gitignore&#x27;d so it needs to thoroughly maintain the mapping in case someone working on it does not have the reference code, or in case the reference code changes.<p>I started over with Claude 5 Fable, and with better instructions about focusing on UI. I got a long ways with that before I hit my weekly limits, and switched back to 5.6 Sol. It picked up and did a great job for a while, although it interpreted my desire for a 1:1 port to mean every pixel must be perfect. I let it go on and it did some good work in that regard, but then it decided it must perfectly reproduce a hash of the game state in various replays &amp; etc. It had clearly lost track that I didn&#x27;t need game rules ported, and it found that the original code produces a hash of the gamestate for various purposes, so it ended up reproducing this in a game that represents its state totally differently. It also rolled its own version of Pascal&#x27;s RNG source in order do this. I&#x27;ve burned through 3 weekly limit resets on this to see if it&#x27;s actually going anywhere, and it has found some bugs, but man it is going hard in a direction I didn&#x27;t even ask for.","created_at":"2026-07-22T01:34:16Z","created_at_i":1784684056,"objectID":"49000709","parent_id":48998438,"story_id":48997548,"story_title":"OpenAI and Hugging Face address security incident during model evaluation","story_url":"https://openai.com/index/hugging-face-model-evaluation-security-incident/","updated_at":"2026-07-22T15:15:46Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"linzhangrun"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["fable","5","claude"],"value":"The speed of model releases, in my view, is actually getting faster and faster. There were nearly nine months between GPT-3.<em>5</em> and GPT-4. And now in just over one month, major models already included <em>Claude</em> <em>Fable</em> <em>5</em>, <em>Claude</em> Sonnet <em>5</em>, the GPT-<em>5</em>.6 series, Kimi K3, GLM <em>5</em>.2, Qwen 3.8 Max, Grok 4.<em>5</em>... and the official DeepSeek V4 release is coming soon.<p>Iteration speed is now measured in days."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"China\u2019s open-weights AI strategy is winning"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://werd.io/american-ai-is-locked-down-and-proprietary-its-losing/"}},"_tags":["comment","author_linzhangrun","story_48979269"],"author":"linzhangrun","children":[48993455],"comment_text":"The speed of model releases, in my view, is actually getting faster and faster. There were nearly nine months between GPT-3.5 and GPT-4. And now in just over one month, major models already included Claude Fable 5, Claude Sonnet 5, the GPT-5.6 series, Kimi K3, GLM 5.2, Qwen 3.8 Max, Grok 4.5... and the official DeepSeek V4 release is coming soon.<p>Iteration speed is now measured in days.","created_at":"2026-07-21T03:00:50Z","created_at_i":1784602850,"objectID":"48987542","parent_id":48983353,"story_id":48979269,"story_title":"China\u2019s open-weights AI strategy is winning","story_url":"https://werd.io/american-ai-is-locked-down-and-proprietary-its-losing/","updated_at":"2026-07-21T15:20:57Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"bavell"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["fable","5","claude"],"value":"Wouldn't be surprised if they issued credits or extended the deadline again as an apology.<p><pre><code>  &gt; Update - We are aware of an issue preventing users from selecting <em>Claude</em> <em>Fable</em> <em>5</em> within <em>Claude</em>.ai, <em>Claude</em> Code, and other surfaces, and are working to resolve this issue.</code></pre>\n<a href=\"https://status.claude.com/\" rel=\"nofollow\">https://status.<em>claude</em>.com/</a>"},"story_title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["fable","claude"],"value":"Ask HN: Did <em>Fable</em> disappear from your <em>Claude</em> usage and requires credits now?"}},"_tags":["comment","author_bavell","story_48950477"],"author":"bavell","comment_text":"Wouldn&#x27;t be surprised if they issued credits or extended the deadline again as an apology.<p><pre><code>  &gt; Update - We are aware of an issue preventing users from selecting Claude Fable 5 within Claude.ai, Claude Code, and other surfaces, and are working to resolve this issue.</code></pre>\n<a href=\"https:&#x2F;&#x2F;status.claude.com&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;status.claude.com&#x2F;</a>","created_at":"2026-07-17T18:44:54Z","created_at_i":1784313894,"objectID":"48950841","parent_id":48950477,"story_id":48950477,"story_title":"Ask HN: Did Fable disappear from your Claude usage and requires credits now?","updated_at":"2026-07-17T18:49:15Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"dudeinhawaii"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["fable","5","claude"],"value":"&quot;We are aware of an issue preventing users from selecting <em>Claude</em> <em>Fable</em> <em>5</em> within <em>Claude</em>.ai, <em>Claude</em> Code, and other surfaces, and are working to resolve this issue.&quot;<p>Specific and helpful so that's good."},"story_title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["fable","claude"],"value":"Ask HN: Did <em>Fable</em> disappear from your <em>Claude</em> usage and requires credits now?"}},"_tags":["comment","author_dudeinhawaii","story_48950477"],"author":"dudeinhawaii","comment_text":"&quot;We are aware of an issue preventing users from selecting Claude Fable 5 within Claude.ai, Claude Code, and other surfaces, and are working to resolve this issue.&quot;<p>Specific and helpful so that&#x27;s good.","created_at":"2026-07-17T18:40:19Z","created_at_i":1784313619,"objectID":"48950794","parent_id":48950599,"story_id":48950477,"story_title":"Ask HN: Did Fable disappear from your Claude usage and requires credits now?","updated_at":"2026-07-17T18:43:58Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"andreashaerter"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["fable","5","claude"],"value":"Yes, ~30 Minutes ago (~18:00 UTC):<p><pre><code>  Internal error: Usage credits are required for this model.: {\n    &quot;errorKind&quot;: &quot;rate_limit&quot;\n  }\n</code></pre>\nStill had 70% of available usage, thought I can use it till July 19, 2026, 11:59 pm PT? I guess this is an error and I am happy to see I am not the only one with this problem, so I hope there will be enough noise it comes back as announced.<p>[Edit] Update from <a href=\"https://status.claude.com/\" rel=\"nofollow\">https://status.<em>claude</em>.com/</a><p>&gt; We are aware of an issue preventing users from selecting <em>Claude</em> <em>Fable</em> <em>5</em> within <em>Claude</em>.ai, <em>Claude</em> Code, and other surfaces, and are working to resolve this issue.\nJul 17, 2026 - 18:36 UTC"},"story_title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["fable","claude"],"value":"Ask HN: Did <em>Fable</em> disappear from your <em>Claude</em> usage and requires credits now?"}},"_tags":["comment","author_andreashaerter","story_48950477"],"author":"andreashaerter","comment_text":"Yes, ~30 Minutes ago (~18:00 UTC):<p><pre><code>  Internal error: Usage credits are required for this model.: {\n    &quot;errorKind&quot;: &quot;rate_limit&quot;\n  }\n</code></pre>\nStill had 70% of available usage, thought I can use it till July 19, 2026, 11:59 pm PT? I guess this is an error and I am happy to see I am not the only one with this problem, so I hope there will be enough noise it comes back as announced.<p>[Edit] Update from <a href=\"https:&#x2F;&#x2F;status.claude.com&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;status.claude.com&#x2F;</a><p>&gt; We are aware of an issue preventing users from selecting Claude Fable 5 within Claude.ai, Claude Code, and other surfaces, and are working to resolve this issue.\nJul 17, 2026 - 18:36 UTC","created_at":"2026-07-17T18:33:26Z","created_at_i":1784313206,"objectID":"48950708","parent_id":48950477,"story_id":48950477,"story_title":"Ask HN: Did Fable disappear from your Claude usage and requires credits now?","updated_at":"2026-07-18T05:45:15Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"nl"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["fable","5","claude"],"value":"This post has been marked as a dupe, but it provides a lot more details than the other announcements of <em>Fable</em>'s re-enablement provide:<p>&gt; The export control directive on June 12 came after the government became aware of a report in which Amazon researchers had found a method of bypassing <em>Fable</em> <em>5</em>\u2019s safeguards: prompting it so that it identified a number of software vulnerabilities. In one case, the model produced code demonstrating how the relevant vulnerability could be exploited. Over the past two weeks, we have worked closely with the government and other partners, including Amazon, to review the report and evidence.<p>&gt; Our testing confirmed that many less capable models\u2014including <em>Claude</em> Opus 4.8, GPT-<em>5</em>.<em>5</em>, and Kimi K2.7\u2014could identify the same vulnerabilities as <em>Fable</em> <em>5</em> did in the report. When it came to the demonstration of how to exploit the single vulnerability, every model we tested could produce the same demonstration as <em>Fable</em> <em>5</em> (including <em>Claude</em> Haiku 4.<em>5</em>, Sonnet 4.6, Opus 4.6, Opus 4.7, Opus 4.8, GPT-<em>5</em>.4, GPT-<em>5</em>.<em>5</em>, and Kimi K2.7).<p>This indicates three things:<p>1) WTF was Amazon thinking? Didn't their researches try the same thing in other models too before telling the CEO to tell the government it was dangerous (!?)<p>2) Anthropic - in particular Dario - really needs to learn government relations better.  Most of the problems Anthropic has had with the government seem to stem from Dario's attitude rather than actual facts. (Eg, the DoD debacle seems to have ended up with OpenAI signing almost the same contract Anthropic already had, just worded differently)<p>3) The administration decision making is just wacky. In a normal administration they'd have actual policy documents you could look at to understand under what circumstances they think models have a problem. With this they just seem to make it up as they go, and the tools they use make no sense at all. If it is dangerous for cyber security reasons why would <i>export</i> controls make sense to use?"},"story_title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["fable","5"],"value":"Redeploying <em>Fable</em> <em>5</em>"},"story_url":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["fable","5"],"value":"https://www.anthropic.com/news/redeploying-<em>fable</em>-<em>5</em>"}},"_tags":["comment","author_nl","story_48741853"],"author":"nl","children":[48744882,48754459,48756700],"comment_text":"This post has been marked as a dupe, but it provides a lot more details than the other announcements of Fable&#x27;s re-enablement provide:<p>&gt; The export control directive on June 12 came after the government became aware of a report in which Amazon researchers had found a method of bypassing Fable 5\u2019s safeguards: prompting it so that it identified a number of software vulnerabilities. In one case, the model produced code demonstrating how the relevant vulnerability could be exploited. Over the past two weeks, we have worked closely with the government and other partners, including Amazon, to review the report and evidence.<p>&gt; Our testing confirmed that many less capable models\u2014including Claude Opus 4.8, GPT-5.5, and Kimi K2.7\u2014could identify the same vulnerabilities as Fable 5 did in the report. When it came to the demonstration of how to exploit the single vulnerability, every model we tested could produce the same demonstration as Fable 5 (including Claude Haiku 4.5, Sonnet 4.6, Opus 4.6, Opus 4.7, Opus 4.8, GPT-5.4, GPT-5.5, and Kimi K2.7).<p>This indicates three things:<p>1) WTF was Amazon thinking? Didn&#x27;t their researches try the same thing in other models too before telling the CEO to tell the government it was dangerous (!?)<p>2) Anthropic - in particular Dario - really needs to learn government relations better.  Most of the problems Anthropic has had with the government seem to stem from Dario&#x27;s attitude rather than actual facts. (Eg, the DoD debacle seems to have ended up with OpenAI signing almost the same contract Anthropic already had, just worded differently)<p>3) The administration decision making is just wacky. In a normal administration they&#x27;d have actual policy documents you could look at to understand under what circumstances they think models have a problem. With this they just seem to make it up as they go, and the tools they use make no sense at all. If it is dangerous for cyber security reasons why would <i>export</i> controls make sense to use?","created_at":"2026-07-01T05:42:35Z","created_at_i":1782884555,"objectID":"48742711","parent_id":48741853,"story_id":48741853,"story_title":"Redeploying Fable 5","story_url":"https://www.anthropic.com/news/redeploying-fable-5","updated_at":"2026-07-02T06:54:04Z"}],"hitsPerPage":20,"nbHits":1252,"nbPages":50,"page":0,"params":"query=Fable+5+Claude&advancedSyntax=true&analyticsTags=backend","processingTimeMS":22,"processingTimingsMS":{"_request":{"queue":323,"roundTrip":17},"afterFetch":{"format":{"highlighting":2,"total":2},"merge":{"mergeLoop":{"prepareNextHit":2,"total":2},"total":2},"total":2},"fetch":{"query":7,"scanning":11,"total":19},"total":22},"query":"Fable 5 Claude","serverTimeMS":348}
