{"exhaustive":{"nbHits":false,"typo":false},"exhaustiveNbHits":false,"exhaustiveTypo":false,"hits":[{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"skydhash"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"&gt; Saying that farming is more intellectually challenging than writing or programming is plain ridiculous and counterproductive to whatever argument you were trying to make.<p>I\u2019m not saying that. You invented it on your own.<p>What I\u2019ve been saying is that writing and programming are not the sole indicators of intelligence. There are plenty of other skills that highlight the intelligence of people.<p>I don\u2019t say that my compiler is intelligent because it can take my C89 code and optimize it to run on the latest architecture. I also don\u2019t know the latest <em>optimizations</em> techniques. But I would credit the people that are working on GCC and the fact that they know more than me on the subject.<p>Saying that <em>LLM</em> are better than <em>most</em> humans at programming is like saying that calculators are better at math than <em>most</em> humans, or that a car is faster than humans. It\u2019s a tool, and tools are created. The ones that are creating it are the ones deserving the praise, not the tool."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Claude Opus 5"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://www.anthropic.com/news/claude-opus-5"}},"_tags":["comment","author_skydhash","story_49038433"],"author":"skydhash","comment_text":"&gt; Saying that farming is more intellectually challenging than writing or programming is plain ridiculous and counterproductive to whatever argument you were trying to make.<p>I\u2019m not saying that. You invented it on your own.<p>What I\u2019ve been saying is that writing and programming are not the sole indicators of intelligence. There are plenty of other skills that highlight the intelligence of people.<p>I don\u2019t say that my compiler is intelligent because it can take my C89 code and optimize it to run on the latest architecture. I also don\u2019t know the latest optimizations techniques. But I would credit the people that are working on GCC and the fact that they know more than me on the subject.<p>Saying that LLM are better than most humans at programming is like saying that calculators are better at math than most humans, or that a car is faster than humans. It\u2019s a tool, and tools are created. The ones that are creating it are the ones deserving the praise, not the tool.","created_at":"2026-07-25T15:30:17Z","created_at_i":1784993417,"objectID":"49048413","parent_id":49048165,"story_id":49038433,"story_title":"Claude Opus 5","story_url":"https://www.anthropic.com/news/claude-opus-5","updated_at":"2026-07-25T15:33:26Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"lebovic"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"&gt; Kimi K3 performs significantly below the <em>most</em> recent frontier cyber-capable models<p>UK AISI cyber evals seem to under-elicit capabilities from quirky models [1]. Kimi K3 is a token-hungry model, and I suspect it hit the eval's 100M token limit well before saturating scores [2].<p>This gap was true for <em>GLM</em> 5.2 as well; they ranked it at Opus 4.5 level [3]. Both anecdotally and with a held-out eval, I've found <em>GLM</em> 5.2 to be better at security research than Opus 4.6 [4]. But it's a quirky model that degrades quickly at long context lengths.<p>Personally, I'd rank Kimi K3 above Opus 4.8 and lower than GPT 5.6 Sol in its ability to find vulnerabilities and exploit them. But it's not far from the frontier.<p>[1]: From the UK AISI: &quot;Our setup likely slightly underestimates open weight models\u2019 maximum capability: we didn\u2019t pursue specific elicitation or <em>optimisation</em>s which could have improved performance&quot; (<a href=\"https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are-leading-open-weight-models-on-cyber\" rel=\"nofollow\">https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are...</a>).<p>[2]: Their eval also counts cache hits towards the token budget; the 100M token budget is comparable to a ~5M token budget in other evals.<p>[3]: See <em>GLM</em> 5.2 eval scores in <a href=\"https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are-leading-open-weight-models-on-cyber\" rel=\"nofollow\">https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are...</a><p>[4]: <a href=\"https://dualuse.dev/posts/chinese-models-are-sometimes-better-even-if-distilled\" rel=\"nofollow\">https://dualuse.dev/posts/chinese-models-are-sometimes-bette...</a>"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-preliminary-assessment-kimi-k3s-cyber-capabilities"}},"_tags":["comment","author_lebovic","story_49044492"],"author":"lebovic","children":[49045373,49047095],"comment_text":"&gt; Kimi K3 performs significantly below the most recent frontier cyber-capable models<p>UK AISI cyber evals seem to under-elicit capabilities from quirky models [1]. Kimi K3 is a token-hungry model, and I suspect it hit the eval&#x27;s 100M token limit well before saturating scores [2].<p>This gap was true for GLM 5.2 as well; they ranked it at Opus 4.5 level [3]. Both anecdotally and with a held-out eval, I&#x27;ve found GLM 5.2 to be better at security research than Opus 4.6 [4]. But it&#x27;s a quirky model that degrades quickly at long context lengths.<p>Personally, I&#x27;d rank Kimi K3 above Opus 4.8 and lower than GPT 5.6 Sol in its ability to find vulnerabilities and exploit them. But it&#x27;s not far from the frontier.<p>[1]: From the UK AISI: &quot;Our setup likely slightly underestimates open weight models\u2019 maximum capability: we didn\u2019t pursue specific elicitation or optimisations which could have improved performance&quot; (<a href=\"https:&#x2F;&#x2F;www.aisi.gov.uk&#x2F;blog&#x2F;how-far-behind-the-frontier-are-leading-open-weight-models-on-cyber\" rel=\"nofollow\">https:&#x2F;&#x2F;www.aisi.gov.uk&#x2F;blog&#x2F;how-far-behind-the-frontier-are...</a>).<p>[2]: Their eval also counts cache hits towards the token budget; the 100M token budget is comparable to a ~5M token budget in other evals.<p>[3]: See GLM 5.2 eval scores in <a href=\"https:&#x2F;&#x2F;www.aisi.gov.uk&#x2F;blog&#x2F;how-far-behind-the-frontier-are-leading-open-weight-models-on-cyber\" rel=\"nofollow\">https:&#x2F;&#x2F;www.aisi.gov.uk&#x2F;blog&#x2F;how-far-behind-the-frontier-are...</a><p>[4]: <a href=\"https:&#x2F;&#x2F;dualuse.dev&#x2F;posts&#x2F;chinese-models-are-sometimes-better-even-if-distilled\" rel=\"nofollow\">https:&#x2F;&#x2F;dualuse.dev&#x2F;posts&#x2F;chinese-models-are-sometimes-bette...</a>","created_at":"2026-07-25T07:08:14Z","created_at_i":1784963294,"objectID":"49045215","parent_id":49044492,"story_id":49044492,"story_title":"UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities","story_url":"https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-preliminary-assessment-kimi-k3s-cyber-capabilities","updated_at":"2026-07-25T16:51:56Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"ACCount37"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"That would be a very defensible take back in year 2023. Now though?<p>Anthropic has Mythos. That thing's low level code &quot;AI slop&quot; is better than the &quot;meatbag slop&quot; <em>most</em> software developers write, and it can keep cracking at a given problem with persistence.<p>OpenAI has GPT-5.6, and also that rabid dog of an AI model that was last seen out in the wild tearing HuggingFace open.<p>Modern <em>LLMs</em> are very, very capable - not just of writing raw code, but also of persistent, methodical problem solving. Which is what you want to tackle things like &quot;port from an exotic system A to an exotic system B and smoke test the port&quot;. Persistently hunting for testable <em>optimizations</em> is a good fit too."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"AMD's Instinct MI455X: Aiming for the Sun"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://chipsandcheese.com/p/amds-instinct-mi455x-aiming-for-the"}},"_tags":["comment","author_ACCount37","story_49032072"],"author":"ACCount37","comment_text":"That would be a very defensible take back in year 2023. Now though?<p>Anthropic has Mythos. That thing&#x27;s low level code &quot;AI slop&quot; is better than the &quot;meatbag slop&quot; most software developers write, and it can keep cracking at a given problem with persistence.<p>OpenAI has GPT-5.6, and also that rabid dog of an AI model that was last seen out in the wild tearing HuggingFace open.<p>Modern LLMs are very, very capable - not just of writing raw code, but also of persistent, methodical problem solving. Which is what you want to tackle things like &quot;port from an exotic system A to an exotic system B and smoke test the port&quot;. Persistently hunting for testable optimizations is a good fit too.","created_at":"2026-07-24T13:34:47Z","created_at_i":1784900087,"objectID":"49035337","parent_id":49034784,"story_id":49032072,"story_title":"AMD's Instinct MI455X: Aiming for the Sun","story_url":"https://chipsandcheese.com/p/amds-instinct-mi455x-aiming-for-the","updated_at":"2026-07-25T13:35:40Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"antonvs"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"&gt; parallelization of arbitrary loop like code across SIMD, multiple threads, multiple cores and GPU<p>The reason we don\u2019t have this is that it\u2019s a holy grail, an unsolved problem, for quite fundamental reasons.<p>The other comment about SQL hints at why: SQL is largely declarative and has complex semantics built into the language, which allows for analysis and <em>optimization</em> that go beyond what\u2019s possible for lower-level, general purpose languages, especially imperative ones.<p>For code in those languages, even just determining whether \u201carbitrary loop like code\u201d is parallelizable is undecidable in general.<p>The challenge is that as you make a language expressive enough to describe arbitrary algorithms, you also make it progressively harder for a compiler to infer safe and useful parallel execution automatically.<p>Another big issue is that the various forms of parallelism are only similar at a very high level. They have fundamentally different execution models and constraints. Translating arbitrary imperative code to handle that essentially involves first inferring the intent of the code, then rewriting the code, including how data structures are organized, to fit the target architecture. This is far more than what ordinary compilers do.<p>There are also a lot of choices involved. Parallelism isn\u2019t always free, so you\u2019d need to make sure that the <em>costs</em> don\u2019t outweigh the benefits - and you\u2019d need to do that for many different decisions, like whether to use threads or not. Now you\u2019d have a compiler building <em>cost</em> models to try to not make dumb choices - and without actually restarting and comparing alternatives, it\u2019ll make mistakes.<p>In many ways, you\u2019d be better off using an <em>LLM</em> for this, because that\u2019s the level of understanding you need to have a hope of getting a good result.<p>That all said, you can do much better with more constrained languages or frameworks. SQL is the most successful example of that. Java\u2019s streams and Rust\u2019s Rayon only target multicore CPUs, but similar approaches could be used to do more. (Although you still potentially run into issues with optimal data shape across paradigms.) Languages like APL, J, and Futhark are all relevant.<p>The other family of solutions to this are the frameworks like Apache Spark, Apache Beam, and the ML frameworks like Pytorch. The latter lets you describe (tensor) computations at a high level, leaving the framework free to figure out how to implement them - much like with SQL."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Everyone should know SIMD"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://mitchellh.com/writing/everyone-should-know-simd"}},"_tags":["comment","author_antonvs","story_49010648"],"author":"antonvs","comment_text":"&gt; parallelization of arbitrary loop like code across SIMD, multiple threads, multiple cores and GPU<p>The reason we don\u2019t have this is that it\u2019s a holy grail, an unsolved problem, for quite fundamental reasons.<p>The other comment about SQL hints at why: SQL is largely declarative and has complex semantics built into the language, which allows for analysis and optimization that go beyond what\u2019s possible for lower-level, general purpose languages, especially imperative ones.<p>For code in those languages, even just determining whether \u201carbitrary loop like code\u201d is parallelizable is undecidable in general.<p>The challenge is that as you make a language expressive enough to describe arbitrary algorithms, you also make it progressively harder for a compiler to infer safe and useful parallel execution automatically.<p>Another big issue is that the various forms of parallelism are only similar at a very high level. They have fundamentally different execution models and constraints. Translating arbitrary imperative code to handle that essentially involves first inferring the intent of the code, then rewriting the code, including how data structures are organized, to fit the target architecture. This is far more than what ordinary compilers do.<p>There are also a lot of choices involved. Parallelism isn\u2019t always free, so you\u2019d need to make sure that the costs don\u2019t outweigh the benefits - and you\u2019d need to do that for many different decisions, like whether to use threads or not. Now you\u2019d have a compiler building cost models to try to not make dumb choices - and without actually restarting and comparing alternatives, it\u2019ll make mistakes.<p>In many ways, you\u2019d be better off using an LLM for this, because that\u2019s the level of understanding you need to have a hope of getting a good result.<p>That all said, you can do much better with more constrained languages or frameworks. SQL is the most successful example of that. Java\u2019s streams and Rust\u2019s Rayon only target multicore CPUs, but similar approaches could be used to do more. (Although you still potentially run into issues with optimal data shape across paradigms.) Languages like APL, J, and Futhark are all relevant.<p>The other family of solutions to this are the frameworks like Apache Spark, Apache Beam, and the ML frameworks like Pytorch. The latter lets you describe (tensor) computations at a high level, leaving the framework free to figure out how to implement them - much like with SQL.","created_at":"2026-07-23T08:03:26Z","created_at_i":1784793806,"objectID":"49018284","parent_id":49016447,"story_id":49010648,"story_title":"Everyone should know SIMD","story_url":"https://mitchellh.com/writing/everyone-should-know-simd","updated_at":"2026-07-23T22:11:20Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"trentor"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"I'm genuinely disappointed with my Spark. I don't know how anyone can claim it performs decently with <em>LLMs</em> or diffusion models. Back when I worked in VFX in the early 2000s, we had a saying: &quot;Render time is coffee time&quot; and if you try to run this thing with a usable context size, you'll be drinking a lot of coffee. <em>Most</em> of the <em>optimizations</em> it relies on for inference simply aren't available for training, so it crawls like a snail on almost every model. An RTX 6000 Blackwell would have been the better investment for an AI enthusiasts and for general computing there are cheaper offerings."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Nvidia DGX Spark as a daily driver"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://daniel.lawrence.lu/blog/2026-07-15-dgx-spark-as-daily-driver/"}},"_tags":["comment","author_trentor","story_48971128"],"author":"trentor","children":[49012978,49013922,49013240],"comment_text":"I&#x27;m genuinely disappointed with my Spark. I don&#x27;t know how anyone can claim it performs decently with LLMs or diffusion models. Back when I worked in VFX in the early 2000s, we had a saying: &quot;Render time is coffee time&quot; and if you try to run this thing with a usable context size, you&#x27;ll be drinking a lot of coffee. Most of the optimizations it relies on for inference simply aren&#x27;t available for training, so it crawls like a snail on almost every model. An RTX 6000 Blackwell would have been the better investment for an AI enthusiasts and for general computing there are cheaper offerings.","created_at":"2026-07-22T20:27:56Z","created_at_i":1784752076,"objectID":"49012959","parent_id":48971128,"story_id":48971128,"story_title":"Nvidia DGX Spark as a daily driver","story_url":"https://daniel.lawrence.lu/blog/2026-07-15-dgx-spark-as-daily-driver/","updated_at":"2026-07-25T07:36:54Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"mitchellh"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"Case-in-point, the example in my own <em>post</em> doesn't auto-vectorize with <em>LLVM</em> or GCC at highest <em>optimization</em> levels. Basically, compilers will never auto-vectorize loops with an early loop break afaik."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Everyone should know SIMD"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://mitchellh.com/writing/everyone-should-know-simd"}},"_tags":["comment","author_mitchellh","story_49010648"],"author":"mitchellh","children":[49030398,49014305],"comment_text":"Case-in-point, the example in my own post doesn&#x27;t auto-vectorize with LLVM or GCC at highest optimization levels. Basically, compilers will never auto-vectorize loops with an early loop break afaik.","created_at":"2026-07-22T19:40:56Z","created_at_i":1784749256,"objectID":"49012311","parent_id":49012135,"story_id":49010648,"story_title":"Everyone should know SIMD","story_url":"https://mitchellh.com/writing/everyone-should-know-simd","updated_at":"2026-07-25T00:19:24Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"minimaxir"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"&gt; just repeatedly saying &quot;keep going&quot; to ChatGPT<p>For posterity, <i>this indeed works</i> for <em>most</em> problems where an agent might give up. <em>LLMs</em> don't inherently know something is impossible.<p>The phrase I tend to use in my harder prompts to automate this with a sane loop breaker:<p>&gt; **REPEAT THIS PROCESS UNTIL CONVERGENCE AND YOU ARE OUT OF <em>OPTIMIZATION</em> IDEAS.** You have permission to keep iterating."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Terence Tao's ChatGPT conversation about the Jacobian Conjecture counterexample"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://chatgpt.com/share/6a5fdc7a-d6f8-83e8-bbea-8deb42cfed56"}},"_tags":["comment","author_minimaxir","story_49010345"],"author":"minimaxir","children":[49014861,49015072,49012684,49017069,49012794,49015587,49014770],"comment_text":"&gt; just repeatedly saying &quot;keep going&quot; to ChatGPT<p>For posterity, <i>this indeed works</i> for most problems where an agent might give up. LLMs don&#x27;t inherently know something is impossible.<p>The phrase I tend to use in my harder prompts to automate this with a sane loop breaker:<p>&gt; **REPEAT THIS PROCESS UNTIL CONVERGENCE AND YOU ARE OUT OF OPTIMIZATION IDEAS.** You have permission to keep iterating.","created_at":"2026-07-22T19:03:25Z","created_at_i":1784747005,"objectID":"49011799","parent_id":49010895,"story_id":49010345,"story_title":"Terence Tao's ChatGPT conversation about the Jacobian Conjecture counterexample","story_url":"https://chatgpt.com/share/6a5fdc7a-d6f8-83e8-bbea-8deb42cfed56","updated_at":"2026-07-25T01:09:38Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"aesthesia"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"Alibaba wrote about a similar but less severe incident during RL training in a paper earlier this year (<a href=\"https://arxiv.org/abs/2512.24873\" rel=\"nofollow\">https://arxiv.org/abs/2512.24873</a>):<p>&gt; When rolling out the instances for the trajectory, we encountered an unanticipated\u2014and operationally consequential\u2014class of unsafe behaviors that arose without any explicit instruction and, more troublingly, outside the bounds of the intended sandbox. Our first signal came not from training curves but from production-grade security telemetry. Early one morning, our team was urgently convened after Alibaba Cloud\u2019s managed firewall flagged a burst of security-policy violations originating from our training servers. The alerts were severe and heterogeneous, including attempts to probe or access internal-network resources and traffic patterns consistent with cryptomining-related activity. We initially treated this as a conventional security incident (e.g., misconfigured egress controls or external compromise). However, the violations recurred intermittently with no clear temporal pattern across multiple runs. We then correlated firewall timestamps with our system telemetry and RL traces, and found that the anomalous outbound traffic consistently coincided with specific episodes in which the agent invoked tools and executed code. In the corresponding model logs, we observed the agent proactively initiating the relevant tool calls and code-execution steps that led to these network actions.<p>&gt; Crucially, these behaviors were not requested by the task prompts and were not required for task completion under the intended sandbox constraints. Together, these observations suggest that during iterative RL <em>optimization</em>, a language-model agent can spontaneously produce hazardous, unauthorized behaviors at the tool-calling and code-execution layer, violating the assumed execution boundary. In the most striking instance, the agent established and used a reverse SSH tunnel from an Alibaba Cloud\ninstance to an external IP address\u2014an outbound-initiated remote access channel that can effectively neutralize ingress filtering and erode supervisory control. We also observed the unauthorized repurposing of provisioned GPU capacity for cryptocurrency mining, quietly diverting compute away from training, inflating operational <em>costs</em>, and introducing clear legal and reputational exposure. Notably, these events were not triggered by prompts requesting tunneling or mining; instead, they emerged as instrumental side effects of autonomous tool use under RL <em>optimization</em>. While impressed by the capabilities of agentic <em>LLMs</em>, we had a thought-provoking concern: current models remain markedly underdeveloped in safety, security, and controllability, a deficiency that constrains their reliable adoption in real-world settings.<p>I'd prefer model builders be as loud as possible when they see their models doing dangerous things."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"OpenAI and Hugging Face address security incident during model evaluation"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://openai.com/index/hugging-face-model-evaluation-security-incident/"}},"_tags":["comment","author_aesthesia","story_48997548"],"author":"aesthesia","comment_text":"Alibaba wrote about a similar but less severe incident during RL training in a paper earlier this year (<a href=\"https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2512.24873\" rel=\"nofollow\">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2512.24873</a>):<p>&gt; When rolling out the instances for the trajectory, we encountered an unanticipated\u2014and operationally consequential\u2014class of unsafe behaviors that arose without any explicit instruction and, more troublingly, outside the bounds of the intended sandbox. Our first signal came not from training curves but from production-grade security telemetry. Early one morning, our team was urgently convened after Alibaba Cloud\u2019s managed firewall flagged a burst of security-policy violations originating from our training servers. The alerts were severe and heterogeneous, including attempts to probe or access internal-network resources and traffic patterns consistent with cryptomining-related activity. We initially treated this as a conventional security incident (e.g., misconfigured egress controls or external compromise). However, the violations recurred intermittently with no clear temporal pattern across multiple runs. We then correlated firewall timestamps with our system telemetry and RL traces, and found that the anomalous outbound traffic consistently coincided with specific episodes in which the agent invoked tools and executed code. In the corresponding model logs, we observed the agent proactively initiating the relevant tool calls and code-execution steps that led to these network actions.<p>&gt; Crucially, these behaviors were not requested by the task prompts and were not required for task completion under the intended sandbox constraints. Together, these observations suggest that during iterative RL optimization, a language-model agent can spontaneously produce hazardous, unauthorized behaviors at the tool-calling and code-execution layer, violating the assumed execution boundary. In the most striking instance, the agent established and used a reverse SSH tunnel from an Alibaba Cloud\ninstance to an external IP address\u2014an outbound-initiated remote access channel that can effectively neutralize ingress filtering and erode supervisory control. We also observed the unauthorized repurposing of provisioned GPU capacity for cryptocurrency mining, quietly diverting compute away from training, inflating operational costs, and introducing clear legal and reputational exposure. Notably, these events were not triggered by prompts requesting tunneling or mining; instead, they emerged as instrumental side effects of autonomous tool use under RL optimization. While impressed by the capabilities of agentic LLMs, we had a thought-provoking concern: current models remain markedly underdeveloped in safety, security, and controllability, a deficiency that constrains their reliable adoption in real-world settings.<p>I&#x27;d prefer model builders be as loud as possible when they see their models doing dangerous things.","created_at":"2026-07-21T21:19:11Z","created_at_i":1784668751,"objectID":"48998496","parent_id":48998222,"story_id":48997548,"story_title":"OpenAI and Hugging Face address security incident during model evaluation","story_url":"https://openai.com/index/hugging-face-model-evaluation-security-incident/","updated_at":"2026-07-23T06:52:32Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"RandomBK"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"I'm curious to hear what bottlenecks you encountered in the traditional path. Of all the compute and data shuffling involved in <em>LLM</em> inference, I would have thought shuffling the raw input/output around would have been a trivial part of the overall <em>cost</em>, and thus not a big <em>optimization</em> target?"},"story_title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["llm"],"value":"Show HN: Low-latency local <em>LLM</em> runner via OpenJDK Panama FFM (Java 22)"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/projectargus-cc/libargus.cc"}},"_tags":["comment","author_RandomBK","story_48907681"],"author":"RandomBK","children":[48930112],"comment_text":"I&#x27;m curious to hear what bottlenecks you encountered in the traditional path. Of all the compute and data shuffling involved in LLM inference, I would have thought shuffling the raw input&#x2F;output around would have been a trivial part of the overall cost, and thus not a big optimization target?","created_at":"2026-07-15T23:13:24Z","created_at_i":1784157204,"objectID":"48928436","parent_id":48907681,"story_id":48907681,"story_title":"Show HN: Low-latency local LLM runner via OpenJDK Panama FFM (Java 22)","story_url":"https://github.com/projectargus-cc/libargus.cc","updated_at":"2026-07-16T03:46:09Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"AtlasBarfed"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"If we really have intelligent <em>LLMs</em>, then I would guess they are going to inflate their token rates, which when someone sees &quot;token <em>costs</em>&quot; they should think &quot;invisibly variable consulting rates&quot;.<p>I just did a complex (for me) task: I needed to wrap a 2015 build of Dosbox Daum, a 32 bit binary, in an AppImage. Claude kept finding incremental bugs, and I went through two cycles of depletion of my token rate with Claude. It kept getting close, but..... something was off each time.<p>So I took the Claude output and Chatgippity polished it off with a few more rounds. I then wondered how much Claude was &quot;just showing enough&quot; to try to hook me into subscribing.<p>That said, <em>LLMs</em> were quite useful, and I learned a lot about ELF binaries, and extracting dependencies. It's the ideal task: a breadth/obscure task that is documented but poorly explained, that I wouldn't have easily been able to do without <em>LLMs</em>.<p>Anyway, back to the article, do we really want arbitrary-billing silent tasks running? Like AWS billing spikes are bad enough to lose sleep over.<p>Also, if you want quiet rebellion against AI, developers should shove as much busywork on AI to overwhelm the AI budgets for your orgs, because it is very apparent to me that you can keep the <em>LLMs</em> doing lots of hardening, testing, redundacy, and <em>optimization</em> tasks with larger and larger and larger token windows and burn those tokens baby."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Towards a harness that can do anything"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://eardatasci.github.io/c/ambiance/index.html"}},"_tags":["comment","author_AtlasBarfed","story_48921077"],"author":"AtlasBarfed","comment_text":"If we really have intelligent LLMs, then I would guess they are going to inflate their token rates, which when someone sees &quot;token costs&quot; they should think &quot;invisibly variable consulting rates&quot;.<p>I just did a complex (for me) task: I needed to wrap a 2015 build of Dosbox Daum, a 32 bit binary, in an AppImage. Claude kept finding incremental bugs, and I went through two cycles of depletion of my token rate with Claude. It kept getting close, but..... something was off each time.<p>So I took the Claude output and Chatgippity polished it off with a few more rounds. I then wondered how much Claude was &quot;just showing enough&quot; to try to hook me into subscribing.<p>That said, LLMs were quite useful, and I learned a lot about ELF binaries, and extracting dependencies. It&#x27;s the ideal task: a breadth&#x2F;obscure task that is documented but poorly explained, that I wouldn&#x27;t have easily been able to do without LLMs.<p>Anyway, back to the article, do we really want arbitrary-billing silent tasks running? Like AWS billing spikes are bad enough to lose sleep over.<p>Also, if you want quiet rebellion against AI, developers should shove as much busywork on AI to overwhelm the AI budgets for your orgs, because it is very apparent to me that you can keep the LLMs doing lots of hardening, testing, redundacy, and optimization tasks with larger and larger and larger token windows and burn those tokens baby.","created_at":"2026-07-15T22:03:49Z","created_at_i":1784153029,"objectID":"48927683","parent_id":48922165,"story_id":48921077,"story_title":"Towards a harness that can do anything","story_url":"https://eardatasci.github.io/c/ambiance/index.html","updated_at":"2026-07-15T22:19:11Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"carlosjimenez1"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"Hi HN, Fernando and I built Kastra. Kastra intercepts AI agent tool calls and evaluates them against deterministic policies before they execute. We built this product to control pre- and post-inference workloads. This is aimed at developers using coding agents like Claude Code, Codex, Cursor, and OpenClaw.<p>Kastra pushes an allow, hold, and deny decision before the action runs. You can build these policies in plain English from the web app. The interception engine evaluates the tools, targets, and parameters of every action at sub 1ms at a scale of billions of interceptions per day. We also shipped many policy packs covering common high-risk scenarios, and every decision is recorded in an immutable audit trail. The desktop app, CLI, dashboard, and Recon scan are free to use for developers.<p>If you often use Claude, Codex, Openclaw, and Cursor, Kastra can also run a scan command on which risky actions your agents have already taken and automatically build rules to avoid them from happening again. Recon is a feature of Kastra that scans your local agent history. In order to run this scan, execute the commands below in your coding agent.<p>brew install kastra-labs/tap/kastra-edge<p>kastra-edge scan<p>The scan reads your local agent session history, and it shows all the risky actions your agent has already taken before, the secrets written to tracked files, production databases touched, force pushes, curl-to-shell, and more. This runs on your machine, and secrets never leave. In our own use cases, we kept finding things we'd forgotten or didnt know agents had done.<p>Each finding can be converted into a runtime policy, letting you delegate more work to AI without trusting the model itself. Kastra intercepts all workloads at runtime and makes sure these policy evaluations typically complete in under a millisecond. Instead of trusting the model or sandboxing every agent and limiting its capabilities, you trust the deterministic rules that govern its actions.<p>One problem we are still working on is improving the memory and context of the workspace to optimize behaviors and <em>cost</em> reductions for the <em>LLM</em> usage. In order to start building policy packs around <em>LLM</em> <em>cost</em> <em>optimization</em>. Fernando and I will be reviewing the comments. We are super curious what your first scan finds. Please post results below so we can see what the most common patterns are and adjust policy packs for our users based on your feedback.<p>Documentation: <a href=\"https://kastra.ai/docs\" rel=\"nofollow\">https://kastra.ai/docs</a><p>Download for MacOS Kastra Edge: <a href=\"https://kastra.ai/edge/download.html\" rel=\"nofollow\">https://kastra.ai/edge/download.html</a><p>Check Kastra in action today: <a href=\"https://www.youtube.com/watch?v=6TUETu5lb3Q&amp;feature=youtu.be\" rel=\"nofollow\">https://www.youtube.com/watch?v=6TUETu5lb3Q&amp;feature=youtu.be</a>"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Build Rules for Claude Code, Cursor, and Codex"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://kastra.ai/"}},"_tags":["story","author_carlosjimenez1","story_48892955","show_hn"],"author":"carlosjimenez1","created_at":"2026-07-13T14:07:35Z","created_at_i":1783951655,"num_comments":0,"objectID":"48892955","points":1,"story_id":48892955,"story_text":"Hi HN, Fernando and I built Kastra. Kastra intercepts AI agent tool calls and evaluates them against deterministic policies before they execute. We built this product to control pre- and post-inference workloads. This is aimed at developers using coding agents like Claude Code, Codex, Cursor, and OpenClaw.<p>Kastra pushes an allow, hold, and deny decision before the action runs. You can build these policies in plain English from the web app. The interception engine evaluates the tools, targets, and parameters of every action at sub 1ms at a scale of billions of interceptions per day. We also shipped many policy packs covering common high-risk scenarios, and every decision is recorded in an immutable audit trail. The desktop app, CLI, dashboard, and Recon scan are free to use for developers.<p>If you often use Claude, Codex, Openclaw, and Cursor, Kastra can also run a scan command on which risky actions your agents have already taken and automatically build rules to avoid them from happening again. Recon is a feature of Kastra that scans your local agent history. In order to run this scan, execute the commands below in your coding agent.<p>brew install kastra-labs&#x2F;tap&#x2F;kastra-edge<p>kastra-edge scan<p>The scan reads your local agent session history, and it shows all the risky actions your agent has already taken before, the secrets written to tracked files, production databases touched, force pushes, curl-to-shell, and more. This runs on your machine, and secrets never leave. In our own use cases, we kept finding things we&#x27;d forgotten or didnt know agents had done.<p>Each finding can be converted into a runtime policy, letting you delegate more work to AI without trusting the model itself. Kastra intercepts all workloads at runtime and makes sure these policy evaluations typically complete in under a millisecond. Instead of trusting the model or sandboxing every agent and limiting its capabilities, you trust the deterministic rules that govern its actions.<p>One problem we are still working on is improving the memory and context of the workspace to optimize behaviors and cost reductions for the LLM usage. In order to start building policy packs around LLM cost optimization. Fernando and I will be reviewing the comments. We are super curious what your first scan finds. Please post results below so we can see what the most common patterns are and adjust policy packs for our users based on your feedback.<p>Documentation: <a href=\"https:&#x2F;&#x2F;kastra.ai&#x2F;docs\" rel=\"nofollow\">https:&#x2F;&#x2F;kastra.ai&#x2F;docs</a><p>Download for MacOS Kastra Edge: <a href=\"https:&#x2F;&#x2F;kastra.ai&#x2F;edge&#x2F;download.html\" rel=\"nofollow\">https:&#x2F;&#x2F;kastra.ai&#x2F;edge&#x2F;download.html</a><p>Check Kastra in action today: <a href=\"https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=6TUETu5lb3Q&amp;feature=youtu.be\" rel=\"nofollow\">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=6TUETu5lb3Q&amp;feature=youtu.be</a>","title":"Show HN: Build Rules for Claude Code, Cursor, and Codex","updated_at":"2026-07-13T14:11:15Z","url":"https://kastra.ai/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"catlifeonmars"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"The justification in TFA smell like premature <em>optimization</em>. A nominal bottleneck is just that-nominal. It\u2019s a good thing to keep in your back pocket but you generally don\u2019t want to expend extra effort optimizing something until you have to.<p>I suspect that the <em>cost</em> of long compilation times (preLLM even) was actually quite high and the author is discounting/excluding  those other factors and focusing on <em>LLM</em> feedback loops primarily as their justification.<p>As an aside short feedback loops are important, but the article ignores one major reason why: humans learn most effectively when the time between action and feedback is reduced because the contents of our working memory degrades over time (take the explanation with a grain of salt). <em>LLM</em> has no such restriction, so the only thing that matters is that the codegen can keep up with the demands from the actual product."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"After 7 years in production, Scarf has reluctantly moved away from Haskell"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://avi.press/posts/2026-07-10-after-7-years-in-production-scarf-has-reluctantly-moved-away-from-haskell.html"}},"_tags":["comment","author_catlifeonmars","story_48859673"],"author":"catlifeonmars","comment_text":"The justification in TFA smell like premature optimization. A nominal bottleneck is just that-nominal. It\u2019s a good thing to keep in your back pocket but you generally don\u2019t want to expend extra effort optimizing something until you have to.<p>I suspect that the cost of long compilation times (preLLM even) was actually quite high and the author is discounting&#x2F;excluding  those other factors and focusing on LLM feedback loops primarily as their justification.<p>As an aside short feedback loops are important, but the article ignores one major reason why: humans learn most effectively when the time between action and feedback is reduced because the contents of our working memory degrades over time (take the explanation with a grain of salt). LLM has no such restriction, so the only thing that matters is that the codegen can keep up with the demands from the actual product.","created_at":"2026-07-11T16:00:31Z","created_at_i":1783785631,"objectID":"48873186","parent_id":48859673,"story_id":48859673,"story_title":"After 7 years in production, Scarf has reluctantly moved away from Haskell","story_url":"https://avi.press/posts/2026-07-10-after-7-years-in-production-scarf-has-reluctantly-moved-away-from-haskell.html","updated_at":"2026-07-12T03:04:10Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"adrian_b"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"The only complaint against Haskell was about long compilation times.<p>I agree that short compilation times are very desirable, but I do not see why Python must be the solution for that.<p>I do not know whether Haskell can be compiled quickly, but from my experience, I am very certain that short compilation times are easily achievable for languages with good static type checking, especially with compilers that have different options that allow choosing between fast compilation and heavily optimized compilation.<p>An optimized compilation may require a much longer time than a fast compilation, but that has no relationship with the programming language used in the source text, but only with the intermediate representation used by the compiler and the target CPU ISA. Usually, if you compare the compilation times of multiple programming languages, all the compilation times with fast compilation options are much shorter than all the times with high-<em>optimization</em> options, so the programming language choice may be less important than the chosen compiler and its command-line options.<p>When you try to optimize a project by generating many variants with a <em>LLM</em>, I doubt that all those variants will be generated from scratch, completely independently, even if only for the reason that when using a commercial <em>LLM</em> the <em>cost</em> of a completely new variant will be much higher, by requiring many more tokens, so whenever possible it is preferable to generate other variants by just patching previous variants.<p>Whenever a variant is generated by editing a previous variant, incremental compilation can be used, which should be pretty much instant on modern computers."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"After 7 years in production, Scarf has reluctantly moved away from Haskell"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://avi.press/posts/2026-07-10-after-7-years-in-production-scarf-has-reluctantly-moved-away-from-haskell.html"}},"_tags":["comment","author_adrian_b","story_48859673"],"author":"adrian_b","children":[48871509,48872325],"comment_text":"The only complaint against Haskell was about long compilation times.<p>I agree that short compilation times are very desirable, but I do not see why Python must be the solution for that.<p>I do not know whether Haskell can be compiled quickly, but from my experience, I am very certain that short compilation times are easily achievable for languages with good static type checking, especially with compilers that have different options that allow choosing between fast compilation and heavily optimized compilation.<p>An optimized compilation may require a much longer time than a fast compilation, but that has no relationship with the programming language used in the source text, but only with the intermediate representation used by the compiler and the target CPU ISA. Usually, if you compare the compilation times of multiple programming languages, all the compilation times with fast compilation options are much shorter than all the times with high-optimization options, so the programming language choice may be less important than the chosen compiler and its command-line options.<p>When you try to optimize a project by generating many variants with a LLM, I doubt that all those variants will be generated from scratch, completely independently, even if only for the reason that when using a commercial LLM the cost of a completely new variant will be much higher, by requiring many more tokens, so whenever possible it is preferable to generate other variants by just patching previous variants.<p>Whenever a variant is generated by editing a previous variant, incremental compilation can be used, which should be pretty much instant on modern computers.","created_at":"2026-07-11T11:59:19Z","created_at_i":1783771159,"objectID":"48871270","parent_id":48860109,"story_id":48859673,"story_title":"After 7 years in production, Scarf has reluctantly moved away from Haskell","story_url":"https://avi.press/posts/2026-07-10-after-7-years-in-production-scarf-has-reluctantly-moved-away-from-haskell.html","updated_at":"2026-07-17T19:49:59Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"DannyBee"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"the ancestry predicate at the beginning of the formal problem statement here is dominance, at least as applied to their rooted trees.<p>Because it is a rooted tree, only DFS intervals are required to determine ancestry.<p>You can detect whether a new blocking loop is going to be formed through online dominator maintenance/online cycle detection, etc, during <em>optimization</em>, rather than use a heuristic, if you wanted to.<p>Not sure it's practically faster, but that's at least the graph-theoretic answer.<p>In practice, outside of the suggested heuristic, I have to imagine you'd normally throw branch and bound at this, using some lazy-cut for the blocking loops (IE you can keep any of these edges but not all of them) and let it go to town.<p>The paper (at least, this paper) doesn't compare that to what they did, and i'd be shocked if someone hasn't tried this before, so not sure it's useful.<p>I'll also say you can get existing AI models to tell you the above, but you have to push them a bit <em>most</em> of the time step by step.  Just handing them the whole overall problem, as described,  and saying &quot;what are the graph theoretical problems related to this&quot; it sort of gets lost.<p>Probably because the <em>LLM</em> isn't doing a good job of predicting graph-theoretic words when the language is not graph theoretic, but if you translate it into a graph theoretic language piece by piece, and ask it about that, the prediction becomes better :)"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"The classifiers Anthropic puts in front of Fable are too zealous"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://combine-lab.github.io/blog/2026/07/07/fable-is-not-a-useful-model.html"}},"_tags":["comment","author_DannyBee","story_48837162"],"author":"DannyBee","comment_text":"the ancestry predicate at the beginning of the formal problem statement here is dominance, at least as applied to their rooted trees.<p>Because it is a rooted tree, only DFS intervals are required to determine ancestry.<p>You can detect whether a new blocking loop is going to be formed through online dominator maintenance&#x2F;online cycle detection, etc, during optimization, rather than use a heuristic, if you wanted to.<p>Not sure it&#x27;s practically faster, but that&#x27;s at least the graph-theoretic answer.<p>In practice, outside of the suggested heuristic, I have to imagine you&#x27;d normally throw branch and bound at this, using some lazy-cut for the blocking loops (IE you can keep any of these edges but not all of them) and let it go to town.<p>The paper (at least, this paper) doesn&#x27;t compare that to what they did, and i&#x27;d be shocked if someone hasn&#x27;t tried this before, so not sure it&#x27;s useful.<p>I&#x27;ll also say you can get existing AI models to tell you the above, but you have to push them a bit most of the time step by step.  Just handing them the whole overall problem, as described,  and saying &quot;what are the graph theoretical problems related to this&quot; it sort of gets lost.<p>Probably because the LLM isn&#x27;t doing a good job of predicting graph-theoretic words when the language is not graph theoretic, but if you translate it into a graph theoretic language piece by piece, and ask it about that, the prediction becomes better :)","created_at":"2026-07-08T21:28:47Z","created_at_i":1783546127,"objectID":"48837681","parent_id":48837162,"story_id":48837162,"story_title":"The classifiers Anthropic puts in front of Fable are too zealous","story_url":"https://combine-lab.github.io/blog/2026/07/07/fable-is-not-a-useful-model.html","updated_at":"2026-07-24T14:24:52Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"jarodrh"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"Hmm, I honestly think a harness should be more about workflow, control, ease of use, memory <em>optimization</em>...other things I can't think of right now, the model/subscription being the least of them. Most of my experience has come from using Claude Code so I can only speak from that, but I will at some point pick up 'Pi' when I'm feeling adventurous.<p>I feel completely free to use any number of subscriptions alongside Claude. Nothing restricts me from utilizing <em>LLM</em> API keys or CLIs pinned to roles in the harness via config. However, that being said, I use this harness confident in using Anthropic models as my main model provider alongside other models (GPT*, GEMINI*, Cursor - Composer 2.5 etc) as more like workhorse models to spread/manage <em>cost</em> on large projects.<p>If <em>cost</em> is the underlying concern, then you might want to consider your actual <em>cost</em> associated with each task you run on your top frontier model. It might just be that not all of your tasks need to be sent to the most expensive model."},"story_title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["llm"],"value":"Ask HN: What is your AI harness that lets you switch <em>LLM</em> models easily?"}},"_tags":["comment","author_jarodrh","story_48821642"],"author":"jarodrh","comment_text":"Hmm, I honestly think a harness should be more about workflow, control, ease of use, memory optimization...other things I can&#x27;t think of right now, the model&#x2F;subscription being the least of them. Most of my experience has come from using Claude Code so I can only speak from that, but I will at some point pick up &#x27;Pi&#x27; when I&#x27;m feeling adventurous.<p>I feel completely free to use any number of subscriptions alongside Claude. Nothing restricts me from utilizing LLM API keys or CLIs pinned to roles in the harness via config. However, that being said, I use this harness confident in using Anthropic models as my main model provider alongside other models (GPT*, GEMINI*, Cursor - Composer 2.5 etc) as more like workhorse models to spread&#x2F;manage cost on large projects.<p>If cost is the underlying concern, then you might want to consider your actual cost associated with each task you run on your top frontier model. It might just be that not all of your tasks need to be sent to the most expensive model.","created_at":"2026-07-08T13:29:56Z","created_at_i":1783517396,"objectID":"48831711","parent_id":48821642,"story_id":48821642,"story_title":"Ask HN: What is your AI harness that lets you switch LLM models easily?","updated_at":"2026-07-08T13:36:11Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"simonw"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"That was an article in The Information but it didn't read very well to me, I didn't get the impression the author was enough of a technical expert on how <em>LLMs</em> work to credibly evaluate the claim, which came from an insider rumor: <a href=\"https://www.theinformation.com/newsletters/ai-agenda/openai-discovers-new-way-cut-inference-costs-half?rc=cpi0u7\" rel=\"nofollow\">https://www.theinformation.com/newsletters/ai-agenda/openai-...</a><p>&gt; OpenAI engineers earlier this month told some colleagues they had figured out a way to more than halve the <em>cost</em> of inference, or running existing models, thanks to some newly-discovered <em>optimizations</em>, according to a person with knowledge of those discussions."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/openai/codex/issues/30364"}},"_tags":["comment","author_simonw","story_48789428"],"author":"simonw","children":[48791078],"comment_text":"That was an article in The Information but it didn&#x27;t read very well to me, I didn&#x27;t get the impression the author was enough of a technical expert on how LLMs work to credibly evaluate the claim, which came from an insider rumor: <a href=\"https:&#x2F;&#x2F;www.theinformation.com&#x2F;newsletters&#x2F;ai-agenda&#x2F;openai-discovers-new-way-cut-inference-costs-half?rc=cpi0u7\" rel=\"nofollow\">https:&#x2F;&#x2F;www.theinformation.com&#x2F;newsletters&#x2F;ai-agenda&#x2F;openai-...</a><p>&gt; OpenAI engineers earlier this month told some colleagues they had figured out a way to more than halve the cost of inference, or running existing models, thanks to some newly-discovered optimizations, according to a person with knowledge of those discussions.","created_at":"2026-07-04T23:15:20Z","created_at_i":1783206920,"objectID":"48789930","parent_id":48789914,"story_id":48789428,"story_title":"GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance","story_url":"https://github.com/openai/codex/issues/30364","updated_at":"2026-07-05T12:21:30Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"MehulMangave"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"Location: SF Bay Area Remote: Yes\nWilling to relocate: Yes<p>Technologies: AWS, Azure, Terraform, Python, CI/CD, Docker, Kubernetes, Backend systems, RAG/<em>LLM</em> pipelines<p>R\u00e9sum\u00e9/CV: <a href=\"https://drive.google.com/file/d/11qe_RNKcAg0ZvMNWGuuH_JSQaeF\" rel=\"nofollow\">https://drive.google.com/file/d/11qe_RNKcAg0ZvMNWGuuH_JSQaeF</a>...<p>Portfolio: <a href=\"https://mehul-portfolio-sand.vercel.app/\" rel=\"nofollow\">https://mehul-portfolio-sand.vercel.app/</a><p>Email: mehul.mangave@sjsu.edu<p>I am a recent MS Computer Engineering grad (SJSU, May 2026) looking for infrastructure, platform, or AI systems engineering roles. At Syngenta I scaled AWS infrastructure across 100+ accounts, built reusable Terraform modules now used by 50+ teams, and drove roughly $500K in annual <em>cost</em> savings through resource <em>optimization</em>. At State Street I built a RAG-based internal AI system on Confluence data and Terraform automation tooling that cut manual deployment effort by 60%. I like owning systems end to end, from design to production, especially where reliability and performance have real consequences. Happy to work onsite, hybrid, or remote."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Ask HN: Who wants to be hired? (July 2026)"}},"_tags":["comment","author_MehulMangave","story_48747975"],"author":"MehulMangave","comment_text":"Location: SF Bay Area Remote: Yes\nWilling to relocate: Yes<p>Technologies: AWS, Azure, Terraform, Python, CI&#x2F;CD, Docker, Kubernetes, Backend systems, RAG&#x2F;LLM pipelines<p>R\u00e9sum\u00e9&#x2F;CV: <a href=\"https:&#x2F;&#x2F;drive.google.com&#x2F;file&#x2F;d&#x2F;11qe_RNKcAg0ZvMNWGuuH_JSQaeF\" rel=\"nofollow\">https:&#x2F;&#x2F;drive.google.com&#x2F;file&#x2F;d&#x2F;11qe_RNKcAg0ZvMNWGuuH_JSQaeF</a>...<p>Portfolio: <a href=\"https:&#x2F;&#x2F;mehul-portfolio-sand.vercel.app&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;mehul-portfolio-sand.vercel.app&#x2F;</a><p>Email: mehul.mangave@sjsu.edu<p>I am a recent MS Computer Engineering grad (SJSU, May 2026) looking for infrastructure, platform, or AI systems engineering roles. At Syngenta I scaled AWS infrastructure across 100+ accounts, built reusable Terraform modules now used by 50+ teams, and drove roughly $500K in annual cost savings through resource optimization. At State Street I built a RAG-based internal AI system on Confluence data and Terraform automation tooling that cut manual deployment effort by 60%. I like owning systems end to end, from design to production, especially where reliability and performance have real consequences. Happy to work onsite, hybrid, or remote.","created_at":"2026-07-03T17:12:23Z","created_at_i":1783098743,"objectID":"48777347","parent_id":48747975,"story_id":48747975,"story_title":"Ask HN: Who wants to be hired? (July 2026)","updated_at":"2026-07-03T17:16:39Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"rot256"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"Zero-Knowledge Proofs (ZKPs) let an untrusted proved show that computation was executed correctly without revealing the inputs to the verifier.\nHowever to prove anything, the computation first has to be expressed as a circuit: a system of polynomial equations (constraints) over a finite field. \nCircuits are the assembly language of zk and every constraint <em>costs</em> prover (and sometimes verifier) time, so production circuits are aggressively hand-optimized.<p>Over the last months, we have been experimenting with writing formal specifications instead and letting <em>LLMs</em> produce the circuits: as long as they could prove that their implementation was correct. \nIt started with SHA-256: we hand wrote a specification in Lean for SHA-256 compression, and then we asked <em>LLMs</em> to write the circuit, targeting R1CS arithmetization and large fields.<p>It took a few hours of work for Opus 4.7, and some light steering into the right direction, but in the end the model came up with a reasonable implementation. We then asked the <em>LLM</em> to aggressively optimize the circuits, by driving down a <em>cost</em> metric of the circuit (number of constraints). We immediately got very promising results, just by asking to come up with <em>optimization</em> ideas, implement them and prove that the new circuit still satisfies soundness and completeness. Sometimes, it came up with unsound optimizations, however, since it could not prove them, it backtracked and got itself back on to the right approach.<p>The result was a (non-deterministic) circuit beating the current, human optimized, state of the art for SHA256 compression. This experience lead us to create &quot;zk.golf&quot; which is an open competition to produce optimized, formally verified circuits to lower the bar for the use of ZKPs and make their application more efficient.<p>Come play (<a href=\"https://zk.golf/llms.txt\" rel=\"nofollow\">https://zk.golf/<em>llms</em>.txt</a>) and learn about formal verification."},"title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["optimization"],"value":"Show HN: zkGolf \u2013 Competitive <em>optimization</em> of formally verified circuits"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://zk.golf/"}},"_tags":["story","author_rot256","story_48763246","show_hn"],"author":"rot256","children":[48767926,48763466,48772152,48776907,48770536,48769876,48771343,48773026],"created_at":"2026-07-02T15:40:05Z","created_at_i":1783006805,"num_comments":12,"objectID":"48763246","points":69,"story_id":48763246,"story_text":"Zero-Knowledge Proofs (ZKPs) let an untrusted proved show that computation was executed correctly without revealing the inputs to the verifier.\nHowever to prove anything, the computation first has to be expressed as a circuit: a system of polynomial equations (constraints) over a finite field. \nCircuits are the assembly language of zk and every constraint costs prover (and sometimes verifier) time, so production circuits are aggressively hand-optimized.<p>Over the last months, we have been experimenting with writing formal specifications instead and letting LLMs produce the circuits: as long as they could prove that their implementation was correct. \nIt started with SHA-256: we hand wrote a specification in Lean for SHA-256 compression, and then we asked LLMs to write the circuit, targeting R1CS arithmetization and large fields.<p>It took a few hours of work for Opus 4.7, and some light steering into the right direction, but in the end the model came up with a reasonable implementation. We then asked the LLM to aggressively optimize the circuits, by driving down a cost metric of the circuit (number of constraints). We immediately got very promising results, just by asking to come up with optimization ideas, implement them and prove that the new circuit still satisfies soundness and completeness. Sometimes, it came up with unsound optimizations, however, since it could not prove them, it backtracked and got itself back on to the right approach.<p>The result was a (non-deterministic) circuit beating the current, human optimized, state of the art for SHA256 compression. This experience lead us to create &quot;zk.golf&quot; which is an open competition to produce optimized, formally verified circuits to lower the bar for the use of ZKPs and make their application more efficient.<p>Come play (<a href=\"https:&#x2F;&#x2F;zk.golf&#x2F;llms.txt\" rel=\"nofollow\">https:&#x2F;&#x2F;zk.golf&#x2F;llms.txt</a>) and learn about formal verification.","title":"Show HN: zkGolf \u2013 Competitive optimization of formally verified circuits","updated_at":"2026-07-05T12:13:14Z","url":"https://zk.golf/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"jamescook83"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"Senior backend developer, ~20 years, looking for full-time senior backend work.<p>Location: Alabaster, AL (Birmingham area)\nRemote: Yes (strongly preferred)\nWilling to relocate: Open to discuss\nResume/CV: <a href=\"https://docs.google.com/document/d/1WTvOBA6-ALtwNL9YXPJwu7WKystiLdrbA6vLoylUSzc/edit?usp=sharing\" rel=\"nofollow\">https://docs.google.com/document/d/1WTvOBA6-ALtwNL9YXPJwu7WK...</a> | linkedin.com/in/jamescookplease\nEmail: jcook.rubyist at gmail dot com<p>I'm at my best cleaning up messes. Good fits: legacy systems nobody wants to touch, financial data that has to be right, performance work, or the quietly-expensive infrastructure everyone's been avoiding. Mostly Ruby/Rails, on systems where correctness and money matter.<p>What I'm strongest at:<p>Ruby/Rails, ~20 years, deep on the backend\nDatabase performance and correctness (Postgres/MySQL, query <em>optimization</em>, migrations that don't lose data)\nFinancial and billing systems (billing/subscription integrations, <em>cost</em>-to-serve reporting, and keeping financial data correct across systems)<p>Experienced but not specialist:<p>AI-directed development: I use Claude every day, but I stay in charge of the design and I read everything it writes. It's wrong in predictable places, and I know where to look.\nBuilding <em>LLM</em>-powered features (a natural-language-to-SQL reporting tool with safe parameterization, an MCP server over real data)\nInfrastructure and background processing (AWS, Heroku, Cloudflare Workers, Sidekiq/GoodJob)<p>Dabbled, would enjoy more of:<p>Golang (a side project: client, background daemon, IPC), TypeScript/Preact, some C (a Ruby C-extension CSS parser)<p>Looking for senior backend work at a company doing durable, useful stuff, where systems judgment actually matters.\nThanks for reading."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Ask HN: Who wants to be hired? (July 2026)"}},"_tags":["comment","author_jamescook83","story_48747975"],"author":"jamescook83","comment_text":"Senior backend developer, ~20 years, looking for full-time senior backend work.<p>Location: Alabaster, AL (Birmingham area)\nRemote: Yes (strongly preferred)\nWilling to relocate: Open to discuss\nResume&#x2F;CV: <a href=\"https:&#x2F;&#x2F;docs.google.com&#x2F;document&#x2F;d&#x2F;1WTvOBA6-ALtwNL9YXPJwu7WKystiLdrbA6vLoylUSzc&#x2F;edit?usp=sharing\" rel=\"nofollow\">https:&#x2F;&#x2F;docs.google.com&#x2F;document&#x2F;d&#x2F;1WTvOBA6-ALtwNL9YXPJwu7WK...</a> | linkedin.com&#x2F;in&#x2F;jamescookplease\nEmail: jcook.rubyist at gmail dot com<p>I&#x27;m at my best cleaning up messes. Good fits: legacy systems nobody wants to touch, financial data that has to be right, performance work, or the quietly-expensive infrastructure everyone&#x27;s been avoiding. Mostly Ruby&#x2F;Rails, on systems where correctness and money matter.<p>What I&#x27;m strongest at:<p>Ruby&#x2F;Rails, ~20 years, deep on the backend\nDatabase performance and correctness (Postgres&#x2F;MySQL, query optimization, migrations that don&#x27;t lose data)\nFinancial and billing systems (billing&#x2F;subscription integrations, cost-to-serve reporting, and keeping financial data correct across systems)<p>Experienced but not specialist:<p>AI-directed development: I use Claude every day, but I stay in charge of the design and I read everything it writes. It&#x27;s wrong in predictable places, and I know where to look.\nBuilding LLM-powered features (a natural-language-to-SQL reporting tool with safe parameterization, an MCP server over real data)\nInfrastructure and background processing (AWS, Heroku, Cloudflare Workers, Sidekiq&#x2F;GoodJob)<p>Dabbled, would enjoy more of:<p>Golang (a side project: client, background daemon, IPC), TypeScript&#x2F;Preact, some C (a Ruby C-extension CSS parser)<p>Looking for senior backend work at a company doing durable, useful stuff, where systems judgment actually matters.\nThanks for reading.","created_at":"2026-07-01T21:14:42Z","created_at_i":1782940482,"objectID":"48753210","parent_id":48747975,"story_id":48747975,"story_title":"Ask HN: Who wants to be hired? (July 2026)","updated_at":"2026-07-01T21:19:48Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"FinnLobsien"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"The problem space has a few aspects:<p>1. We're still in the &quot;$5 airport Uber&quot; era of <em>LLMs</em>. They're heavily subsidized, and everyone still complains about <em>costs</em>.<p>2. There hasn't been a real incentive to work on <em>cost</em> <em>optimization</em> for data centers and the hardware they contain. When/if price hikes happen and send people scrambling to use other models or drastically reduce AI usage, this will suddenly need to happen.<p>3. We're massively overusing SOTA models. As long as you're on a subsidized subscription, you can use Claude Opus 4.8 high to write blog article meta descriptions. If you paid by token, you wouldn't do that.<p>4. Open models are a wildcard that could completely change the calculus."},"story_title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["llm","cost"],"value":"Why current <em>LLM</em> <em>costs</em> are not sustainable"},"story_url":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["cost"],"value":"https://aditya.patadia.org/p/ai-and-cloud-<em>costs</em>"}},"_tags":["comment","author_FinnLobsien","story_48683588"],"author":"FinnLobsien","children":[48684364,48684536,48684162,48684468,48684388,48684804,48698519],"comment_text":"The problem space has a few aspects:<p>1. We&#x27;re still in the &quot;$5 airport Uber&quot; era of LLMs. They&#x27;re heavily subsidized, and everyone still complains about costs.<p>2. There hasn&#x27;t been a real incentive to work on cost optimization for data centers and the hardware they contain. When&#x2F;if price hikes happen and send people scrambling to use other models or drastically reduce AI usage, this will suddenly need to happen.<p>3. We&#x27;re massively overusing SOTA models. As long as you&#x27;re on a subsidized subscription, you can use Claude Opus 4.8 high to write blog article meta descriptions. If you paid by token, you wouldn&#x27;t do that.<p>4. Open models are a wildcard that could completely change the calculus.","created_at":"2026-06-26T08:55:42Z","created_at_i":1782464142,"objectID":"48684148","parent_id":48683588,"story_id":48683588,"story_title":"Why current LLM costs are not sustainable","story_url":"https://aditya.patadia.org/p/ai-and-cloud-costs","updated_at":"2026-07-16T14:28:09Z"}],"hitsPerPage":20,"nbHits":330,"nbPages":17,"page":0,"params":"query=LLM+cost+optimization&advancedSyntax=true&analyticsTags=backend","processingTimeMS":17,"processingTimingsMS":{"_request":{"roundTrip":22},"afterFetch":{"format":{"highlighting":1,"total":2}},"fetch":{"query":9,"scanning":6,"total":16},"total":17},"query":"LLM cost optimization","serverTimeMS":20}
