{"exhaustive":{"nbHits":false,"typo":false},"exhaustiveNbHits":false,"exhaustiveTypo":false,"hits":[{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"tgies"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"Location: Omaha, NE metro<p>Remote: Yes (remote only, US; async-friendly)<p>Willing to relocate: Yes, but high bar<p>Technologies: C, C++, Rust, TypeScript/JavaScript (Node, browser, Workers), Python, C#/.NET, SQL (Postgres/MSSQL), AWS, audio/DSP (formant + FM synthesis, AudioWorklet, embedded realtime), WebAssembly, x86/ARM assembly, MCP servers and agentic dev tooling<p>R\u00e9sum\u00e9/CV: <a href=\"https://crashunited.com/resume.pdf\" rel=\"nofollow\">https://crashunited.com/resume.pdf</a><p>Email: tony.gies@crashunited.com<p>Senior generalist, 18 years. Been doing a lot of FOSS work lately; 62 PRs merged upstream this year, into ScummVM, 86Box, chocolate-doom, dsda-doom, woof, DOSBox-X, dosbox-staging, DOSBox Pure, SillyTavern, CHIRP, cross-rs, DefinitelyTyped, and nodejs/node. Literally none of my PRs have been rejected by maintainer in the past year.<p>Recent public work:<p>- klattsch (github.com/tgies/klattsch): singing parallel-formant speech synthesizer I wrote from scratch based on papers from 1980. Went viral in May (8M+ impressions, X trending, etc). Now a paid desktop/Android app on itch.io and Google Play, a built-in chip in the Furnace tracker, and in Mikoto Studio, a commercial Vocaloid-like vocal synth studio. Ported the core speech synth engine to TS, C++, Rust, and C#.<p>- Nuked-OPL3-fast: <em>Optimization</em> fork of the industry-standard OPL3 emulator, adopted upstream by ScummVM, dosbox-staging, DOSBox-X, DOSBox Pure, chocolate-doom, Even runs realtime at full fidelity on one RP2350 core.<p>- copy-fail-c: Cross-platform PoC for CVE-2026-31431, the #2 <em>most</em>-starred PoC for the CVE on GitHub (444 stars).<p>- client-certificate-auth: Node mTLS middleware, maintained 13 years, recommended in AWS's API Gateway docs, used in BAE Systems products.<p>- GPG Guide 2026: 376-page guide to modern GnuPG, YubiKey, and OpenPGP v6, self-published solo, with a bespoke typesetting and publishing toolchain and a harness that verifies every command in the book.<p>Previously: 11 years at a .NET CRM vendor (C#, MSSQL, sole AWS keyholder, built and ran the serverless ETL platform), then a year rescuing an embedded livestock-sorting system (Python/ARM/Postgres/Laravel) into its first working field deployment.<p>Looking for: contract / fractional / fixed-scope work through my consultancy, Crash United <em>LLC</em>, or senior IC roles. Deep daily agentic-tooling practice incl. a bunch of custom MCP servers I wrote."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Ask HN: Who wants to be hired? (September 2026)"}},"_tags":["comment","author_tgies","story_49522896"],"author":"tgies","comment_text":"Location: Omaha, NE metro<p>Remote: Yes (remote only, US; async-friendly)<p>Willing to relocate: Yes, but high bar<p>Technologies: C, C++, Rust, TypeScript&#x2F;JavaScript (Node, browser, Workers), Python, C#&#x2F;.NET, SQL (Postgres&#x2F;MSSQL), AWS, audio&#x2F;DSP (formant + FM synthesis, AudioWorklet, embedded realtime), WebAssembly, x86&#x2F;ARM assembly, MCP servers and agentic dev tooling<p>R\u00e9sum\u00e9&#x2F;CV: <a href=\"https:&#x2F;&#x2F;crashunited.com&#x2F;resume.pdf\" rel=\"nofollow\">https:&#x2F;&#x2F;crashunited.com&#x2F;resume.pdf</a><p>Email: tony.gies@crashunited.com<p>Senior generalist, 18 years. Been doing a lot of FOSS work lately; 62 PRs merged upstream this year, into ScummVM, 86Box, chocolate-doom, dsda-doom, woof, DOSBox-X, dosbox-staging, DOSBox Pure, SillyTavern, CHIRP, cross-rs, DefinitelyTyped, and nodejs&#x2F;node. Literally none of my PRs have been rejected by maintainer in the past year.<p>Recent public work:<p>- klattsch (github.com&#x2F;tgies&#x2F;klattsch): singing parallel-formant speech synthesizer I wrote from scratch based on papers from 1980. Went viral in May (8M+ impressions, X trending, etc). Now a paid desktop&#x2F;Android app on itch.io and Google Play, a built-in chip in the Furnace tracker, and in Mikoto Studio, a commercial Vocaloid-like vocal synth studio. Ported the core speech synth engine to TS, C++, Rust, and C#.<p>- Nuked-OPL3-fast: Optimization fork of the industry-standard OPL3 emulator, adopted upstream by ScummVM, dosbox-staging, DOSBox-X, DOSBox Pure, chocolate-doom, Even runs realtime at full fidelity on one RP2350 core.<p>- copy-fail-c: Cross-platform PoC for CVE-2026-31431, the #2 most-starred PoC for the CVE on GitHub (444 stars).<p>- client-certificate-auth: Node mTLS middleware, maintained 13 years, recommended in AWS&#x27;s API Gateway docs, used in BAE Systems products.<p>- GPG Guide 2026: 376-page guide to modern GnuPG, YubiKey, and OpenPGP v6, self-published solo, with a bespoke typesetting and publishing toolchain and a harness that verifies every command in the book.<p>Previously: 11 years at a .NET CRM vendor (C#, MSSQL, sole AWS keyholder, built and ran the serverless ETL platform), then a year rescuing an embedded livestock-sorting system (Python&#x2F;ARM&#x2F;Postgres&#x2F;Laravel) into its first working field deployment.<p>Looking for: contract &#x2F; fractional &#x2F; fixed-scope work through my consultancy, Crash United LLC, or senior IC roles. Deep daily agentic-tooling practice incl. a bunch of custom MCP servers I wrote.","created_at":"2026-09-08T17:41:17Z","created_at_i":1788889277,"objectID":"49613754","parent_id":49522896,"story_id":49522896,"story_title":"Ask HN: Who wants to be hired? (September 2026)","updated_at":"2026-09-08T17:42:04Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"jeremyjh"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"These are all important points and I love the analogy. But there is an even bigger issue with having <em>LLMs</em> write for you:<p>Writing is thinking. Thinking and deciding. There have been many times when I start out writing something substantial - could be an email, a blog <em>post</em>, a software design document, anything - when my own views substantially changed during the writing process. Writing forces you to serialize your thoughts - and you can't always trust the gestalt.<p>Reviewing gives you the chance to ensure the arguments connect solidly, that references are accurate (even informal references) and gives you the time to consider counter-arguments you aren't addressing.<p>None of this matters much on LinkedIn, but it matters a lot in our work. You cannot outsource your <i>understanding</i> to AI. They are powerful tools but they do not have any human understanding - that isn't their <em>optimization</em> target."},"story_title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["llm","cost"],"value":"Your intellectual fly is open when you use an <em>LLM</em> to author a <em>post</em> (2025)"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://bcantrill.dtrace.org/2025/12/05/your-intellectual-fly-is-open/"}},"_tags":["comment","author_jeremyjh","story_49585644"],"author":"jeremyjh","children":[49586453,49588034,49586720,49593099,49587124,49590833,49591291,49589080,49590980,49586805,49589051,49586950,49588552,49592768,49588836,49589916,49591014,49589206,49587046,49591364,49586972,49586609,49591499,49588787],"comment_text":"These are all important points and I love the analogy. But there is an even bigger issue with having LLMs write for you:<p>Writing is thinking. Thinking and deciding. There have been many times when I start out writing something substantial - could be an email, a blog post, a software design document, anything - when my own views substantially changed during the writing process. Writing forces you to serialize your thoughts - and you can&#x27;t always trust the gestalt.<p>Reviewing gives you the chance to ensure the arguments connect solidly, that references are accurate (even informal references) and gives you the time to consider counter-arguments you aren&#x27;t addressing.<p>None of this matters much on LinkedIn, but it matters a lot in our work. You cannot outsource your <i>understanding</i> to AI. They are powerful tools but they do not have any human understanding - that isn&#x27;t their optimization target.","created_at":"2026-09-06T13:26:14Z","created_at_i":1788701174,"objectID":"49586344","parent_id":49585644,"story_id":49585644,"story_title":"Your intellectual fly is open when you use an LLM to author a post (2025)","story_url":"https://bcantrill.dtrace.org/2025/12/05/your-intellectual-fly-is-open/","updated_at":"2026-09-08T20:56:33Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"camdenreslink"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"You can imagine with more operations being available to be done more cheaply and quickly the <em>LLM</em> doesn't need to &quot;one shot&quot; a solution. It could try many solutions, test them, throw some away, wiggle some of the parameters like a genetic algorithm, see how that changes the result, and converge on an optimal solution (based on whatever the <em>cost</em> function is). Basically producing a good result could become like an <em>optimization</em> problem. That would be way too expensive and slow right now."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"OpenAI begins rolling out GPT-6 Astra"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://www.cnbc.com/2026/09/03/open-ai-astra-gpt-6-cyber.html"}},"_tags":["comment","author_camdenreslink","story_49554273"],"author":"camdenreslink","comment_text":"You can imagine with more operations being available to be done more cheaply and quickly the LLM doesn&#x27;t need to &quot;one shot&quot; a solution. It could try many solutions, test them, throw some away, wiggle some of the parameters like a genetic algorithm, see how that changes the result, and converge on an optimal solution (based on whatever the cost function is). Basically producing a good result could become like an optimization problem. That would be way too expensive and slow right now.","created_at":"2026-09-03T19:35:01Z","created_at_i":1788464101,"objectID":"49555521","parent_id":49555128,"story_id":49554273,"story_title":"OpenAI begins rolling out GPT-6 Astra","story_url":"https://www.cnbc.com/2026/09/03/open-ai-astra-gpt-6-cyber.html","updated_at":"2026-09-03T19:45:36Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"entrope"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"This article describes what is usually called a Pareto frontier: the best known achievable trade-offs between two (or more) <em>optimization</em> goals. &quot;Efficient frontier&quot; in common usage seems to be specifically a Pareto frontier for financial risk versus return of an investment portfolio.  Even outside of finance, points on a Pareto frontier are called Pareto optimal or Pareto efficient.  A Pareto frontier is sometimes shown with more than two dimensions, although usually people will pick just two for simplicity.<p>Within <em>LLMs</em>, and even inference naturally, there are many other potential parameters that one might optimize: Unsloth typically shows a Pareto frontier for size of a quantized model versus KL divergence.  Others trade total concurrent tok/s against single-stream tok/s. KV cache size, context length and context coherency are other trade-offs that are closely related to inference. Total intelligence is usually a defining characteristic of a &quot;frontier model&quot;, with <em>cost</em> (per token or task) as a salient trade-off.  <em>Cost</em> is one parameter that is implicitly fixed by the &quot;throughput versus latency&quot; analysis: using a GB300 versus Radeon R9700 moves the curve enormously and probably changes the shape of it.  Lots of threads here argue over local vs cloud inference regarding <em>cost</em> efficiency, often with privacy and control as competing objectives."},"story_title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["llm"],"value":"The efficient frontier of <em>LLM</em> inference"},"story_url":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["llm"],"value":"https://www.baseten.co/blog/the-efficient-frontier-of-<em>llm</em>-inference/"}},"_tags":["comment","author_entrope","story_49529898"],"author":"entrope","comment_text":"This article describes what is usually called a Pareto frontier: the best known achievable trade-offs between two (or more) optimization goals. &quot;Efficient frontier&quot; in common usage seems to be specifically a Pareto frontier for financial risk versus return of an investment portfolio.  Even outside of finance, points on a Pareto frontier are called Pareto optimal or Pareto efficient.  A Pareto frontier is sometimes shown with more than two dimensions, although usually people will pick just two for simplicity.<p>Within LLMs, and even inference naturally, there are many other potential parameters that one might optimize: Unsloth typically shows a Pareto frontier for size of a quantized model versus KL divergence.  Others trade total concurrent tok&#x2F;s against single-stream tok&#x2F;s. KV cache size, context length and context coherency are other trade-offs that are closely related to inference. Total intelligence is usually a defining characteristic of a &quot;frontier model&quot;, with cost (per token or task) as a salient trade-off.  Cost is one parameter that is implicitly fixed by the &quot;throughput versus latency&quot; analysis: using a GB300 versus Radeon R9700 moves the curve enormously and probably changes the shape of it.  Lots of threads here argue over local vs cloud inference regarding cost efficiency, often with privacy and control as competing objectives.","created_at":"2026-09-02T10:59:38Z","created_at_i":1788346778,"objectID":"49534515","parent_id":49529898,"story_id":49529898,"story_title":"The efficient frontier of LLM inference","story_url":"https://www.baseten.co/blog/the-efficient-frontier-of-llm-inference/","updated_at":"2026-09-03T05:22:15Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"divsh17"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"Location: Bellevue, WA, USA<p>Remote: Yes. Prefer in person<p>Willing to relocate: YES<p>Technologies: Python, TypeScript, FastAPI, Next.js, React, PostgreSQL, pgvector, Redis, Apache Kafka, AWS, GCP, PyTorch, <em>LLMs</em>, RAG, Prompt Engineering<p>R\u00e9sum\u00e9/CV: <a href=\"https://drive.google.com/file/d/1dhWtipC3WauUodrj3K0abG8rZn5RjGcv/view?usp=sharing\" rel=\"nofollow\">https://drive.google.com/file/d/1dhWtipC3WauUodrj3K0abG8rZn5...</a><p>Email: divyanshusharma17.work@gmail.com<p>Links: <a href=\"https://divyanshusharma.com\" rel=\"nofollow\">https://divyanshusharma.com</a> | <a href=\"https://github.com/divyanshusharma1709\" rel=\"nofollow\">https://github.com/divyanshusharma1709</a><p>I am an AI/Full-stack Engineer with 3+ YoE focused on production AI pipelines, inference <em>optimization</em>, and high-precision RAG. Recently as a Founding Engineer at Cascade Intelligence (a16z speedrun), I re-architected a vector matching pipeline that tripled RFP match precision from 20% to 60%. I also built Elliot AI, a proactive companion utilizing hybrid RAG to guarantee &lt;300ms context retrieval across 6+ months of history, and previously engineered an AI narration system for VR surgery syncing <em>LLM</em> inference with frame-accurate playback under 200ms latency.<p>My foundational infrastructure background includes scaling an async queue-based email system from 400K to 1.5M+ daily sends and driving $2.1M+ in annualized revenue impact through microservice A/B experimentation at Amazon. I am actively seeking senior or founding AI engineering roles to tackle massive system design challenges and build core agentic frameworks, and I require a STEM-OPT transfer (no sponsorship <em>cost</em>)."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Ask HN: Who wants to be hired? (September 2026)"}},"_tags":["comment","author_divsh17","story_49522896"],"author":"divsh17","comment_text":"Location: Bellevue, WA, USA<p>Remote: Yes. Prefer in person<p>Willing to relocate: YES<p>Technologies: Python, TypeScript, FastAPI, Next.js, React, PostgreSQL, pgvector, Redis, Apache Kafka, AWS, GCP, PyTorch, LLMs, RAG, Prompt Engineering<p>R\u00e9sum\u00e9&#x2F;CV: <a href=\"https:&#x2F;&#x2F;drive.google.com&#x2F;file&#x2F;d&#x2F;1dhWtipC3WauUodrj3K0abG8rZn5RjGcv&#x2F;view?usp=sharing\" rel=\"nofollow\">https:&#x2F;&#x2F;drive.google.com&#x2F;file&#x2F;d&#x2F;1dhWtipC3WauUodrj3K0abG8rZn5...</a><p>Email: divyanshusharma17.work@gmail.com<p>Links: <a href=\"https:&#x2F;&#x2F;divyanshusharma.com\" rel=\"nofollow\">https:&#x2F;&#x2F;divyanshusharma.com</a> | <a href=\"https:&#x2F;&#x2F;github.com&#x2F;divyanshusharma1709\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;divyanshusharma1709</a><p>I am an AI&#x2F;Full-stack Engineer with 3+ YoE focused on production AI pipelines, inference optimization, and high-precision RAG. Recently as a Founding Engineer at Cascade Intelligence (a16z speedrun), I re-architected a vector matching pipeline that tripled RFP match precision from 20% to 60%. I also built Elliot AI, a proactive companion utilizing hybrid RAG to guarantee &lt;300ms context retrieval across 6+ months of history, and previously engineered an AI narration system for VR surgery syncing LLM inference with frame-accurate playback under 200ms latency.<p>My foundational infrastructure background includes scaling an async queue-based email system from 400K to 1.5M+ daily sends and driving $2.1M+ in annualized revenue impact through microservice A&#x2F;B experimentation at Amazon. I am actively seeking senior or founding AI engineering roles to tackle massive system design challenges and build core agentic frameworks, and I require a STEM-OPT transfer (no sponsorship cost).","created_at":"2026-09-01T20:28:34Z","created_at_i":1788294514,"objectID":"49527711","parent_id":49522896,"story_id":49522896,"story_title":"Ask HN: Who wants to be hired? (September 2026)","updated_at":"2026-09-01T20:32:28Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"paimapi"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"in the few books I've read about technological breakthroughs, the vast majority come from academia. I'm thinking about Eckert's adoption of IBM's ticker tape for commands rather than just rote data [0], all the inventions that came out of the NDRC [1] (for better or worse), the ENIAC, FORTRAN, how reliant solid state devices were on understanding quantum theory [2], and so many other examples. Paul E. Ceruzzi's &quot;Computing: A Concise History&quot; goes over much of this (including making note of the US military's heavy usage of <em>LLMs</em> for the Iraq War pre-empted the consumer push for <em>LLMs</em>)<p>most of the things that private enterprise contributes is <em>cost</em> reduction and manufacturing capacity (important in their own right for ubiquity) but having extremely well-funded R&amp;D and universities still seem like the primary driver for truly evolutionary and revolutionary breakthroughs rather than useful but not necessary <em>optimizations</em><p>IP law very much aims to protect the latter so that breakthroughs can be capitalized on but aren't a necessary motivator for researchers to continue frontier work<p>(0)<a href=\"https://en.wikipedia.org/wiki/IBM_SSEC\" rel=\"nofollow\">https://en.wikipedia.org/wiki/IBM_SSEC</a>\n(1)<a href=\"https://en.wikipedia.org/wiki/National_Defense_Research_Committee\" rel=\"nofollow\">https://en.wikipedia.org/wiki/National_Defense_Research_Comm...</a>\n(2)<a href=\"https://www.allaboutcircuits.com/textbook/semiconductors/chpt-2/quantum-devices/\" rel=\"nofollow\">https://www.allaboutcircuits.com/textbook/semiconductors/chp...</a>"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"EFF to Courts: Don't Rewrite Copyright over AI Hype"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://www.eff.org/deeplinks/2026/08/eff-courts-dont-rewrite-copyright-over-ai-hype"}},"_tags":["comment","author_paimapi","story_49521315"],"author":"paimapi","comment_text":"in the few books I&#x27;ve read about technological breakthroughs, the vast majority come from academia. I&#x27;m thinking about Eckert&#x27;s adoption of IBM&#x27;s ticker tape for commands rather than just rote data [0], all the inventions that came out of the NDRC [1] (for better or worse), the ENIAC, FORTRAN, how reliant solid state devices were on understanding quantum theory [2], and so many other examples. Paul E. Ceruzzi&#x27;s &quot;Computing: A Concise History&quot; goes over much of this (including making note of the US military&#x27;s heavy usage of LLMs for the Iraq War pre-empted the consumer push for LLMs)<p>most of the things that private enterprise contributes is cost reduction and manufacturing capacity (important in their own right for ubiquity) but having extremely well-funded R&amp;D and universities still seem like the primary driver for truly evolutionary and revolutionary breakthroughs rather than useful but not necessary optimizations<p>IP law very much aims to protect the latter so that breakthroughs can be capitalized on but aren&#x27;t a necessary motivator for researchers to continue frontier work<p>(0)<a href=\"https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;IBM_SSEC\" rel=\"nofollow\">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;IBM_SSEC</a>\n(1)<a href=\"https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;National_Defense_Research_Committee\" rel=\"nofollow\">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;National_Defense_Research_Comm...</a>\n(2)<a href=\"https:&#x2F;&#x2F;www.allaboutcircuits.com&#x2F;textbook&#x2F;semiconductors&#x2F;chpt-2&#x2F;quantum-devices&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;www.allaboutcircuits.com&#x2F;textbook&#x2F;semiconductors&#x2F;chp...</a>","created_at":"2026-09-01T15:14:53Z","created_at_i":1788275693,"objectID":"49523096","parent_id":49521718,"story_id":49521315,"story_title":"EFF to Courts: Don't Rewrite Copyright over AI Hype","story_url":"https://www.eff.org/deeplinks/2026/08/eff-courts-dont-rewrite-copyright-over-ai-hype","updated_at":"2026-09-01T15:33:27Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"zbentley"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"&gt; the C10k challenge is solved since a long time<p>This has nothing to do with that.<p>Any Node.JS application will happily accept 100K connections. They'll all wait for the under-resourced database behind it. That application &quot;solved&quot; the C10K challenge, but it's still overwhelmed.<p>&gt; it might be possible that the &quot;<em>LLM</em> agent&quot; is requesting the commit contents to &quot;understand&quot; or refer or explain them. Is it a bad thing if it helps users?<p>The article describes random algorithmically-generated traffic arriving in batched waves from laundered residential proxy IP addresses, a few unrelated hits in a group then gone. That's not the pattern you'd see if end users were asking their agents for help.<p>&gt; it is a shame that such talented people would not be able to have a proper <em>optimization</em><p>It's mostly <i>not</i> static content in the sense that you're implying.<p>Routes that access a single commit can be cached. But most of the routes scrapers are hitting are e.g. computing diffs between arbitrary pairs of commits, or other computed-on-the-fly views into history.<p>I'm sure they're already caching their useful-to-real-humans data. As the article said, the vast majority of their traffic is bots hitting those arbitrary, permuted URLs. So whatever cache they're using is probably a) missed almost every time, and b) constantly getting evicted to make room for data served to bots (unless they eschew caching to avoid this--fair--and are thus back to the original issue regardless).<p>There is no &quot;proper <em>optimization</em>&quot; here. It's not <i>slow</i> to go compute the diff between a random pair of refs, render that into pretty HTML, and serve it. But it <em>costs</em> something more than a cache hit, and doing that dozens-to-hundreds of times a second constantly consumes resources."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Creepy Crawlies"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://people.kernel.org/monsieuricon/creepy-crawlies"}},"_tags":["comment","author_zbentley","story_49491791"],"author":"zbentley","children":[49507007],"comment_text":"&gt; the C10k challenge is solved since a long time<p>This has nothing to do with that.<p>Any Node.JS application will happily accept 100K connections. They&#x27;ll all wait for the under-resourced database behind it. That application &quot;solved&quot; the C10K challenge, but it&#x27;s still overwhelmed.<p>&gt; it might be possible that the &quot;LLM agent&quot; is requesting the commit contents to &quot;understand&quot; or refer or explain them. Is it a bad thing if it helps users?<p>The article describes random algorithmically-generated traffic arriving in batched waves from laundered residential proxy IP addresses, a few unrelated hits in a group then gone. That&#x27;s not the pattern you&#x27;d see if end users were asking their agents for help.<p>&gt; it is a shame that such talented people would not be able to have a proper optimization<p>It&#x27;s mostly <i>not</i> static content in the sense that you&#x27;re implying.<p>Routes that access a single commit can be cached. But most of the routes scrapers are hitting are e.g. computing diffs between arbitrary pairs of commits, or other computed-on-the-fly views into history.<p>I&#x27;m sure they&#x27;re already caching their useful-to-real-humans data. As the article said, the vast majority of their traffic is bots hitting those arbitrary, permuted URLs. So whatever cache they&#x27;re using is probably a) missed almost every time, and b) constantly getting evicted to make room for data served to bots (unless they eschew caching to avoid this--fair--and are thus back to the original issue regardless).<p>There is no &quot;proper optimization&quot; here. It&#x27;s not <i>slow</i> to go compute the diff between a random pair of refs, render that into pretty HTML, and serve it. But it costs something more than a cache hit, and doing that dozens-to-hundreds of times a second constantly consumes resources.","created_at":"2026-08-31T02:49:06Z","created_at_i":1788144546,"objectID":"49505057","parent_id":49504056,"story_id":49491791,"story_title":"Creepy Crawlies","story_url":"https://people.kernel.org/monsieuricon/creepy-crawlies","updated_at":"2026-08-31T08:50:08Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"SilenN"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"Hi HN, we built an open source model gateway. It's a single place to manage our own self hosted, frontier, and open source models in one place.<p>It\u2019s is rust native, built for concurrency, and implements all the config quirks across models and providers (streaming formats, tool calls, model parameters, rate limits, and different error behavior).<p>The gateway adds under 1 ms for BYOK requests and under 2 ms when Experiential supplies the provider key. It has every major inference provider, and 1000+ models refreshed daily via a codex agent that opens a PR.<p>Compared to other similar projects we\u2019re open source, take no markup, allow you to mix local models with a marketplace, and use your traffic to (opt in) train you a model. Simple routing doesn\u2019t warrant a 10% token markup.<p>The way we do this is given standardized OTel traces, we mine representative real tasks, use text world models to simulate rollouts for various models, apply an <em>LLM</em> judge, and fit a nearest neighbor classifier on top of an embedding of a prompt to decide the optimal model for each request. Usually this can map out a better pareto curve on <em>cost</em>/quality than just calling single models but it\u2019s not perfect.<p>Using these simulations we can also do things like suggesting cache hit <em>optimizations</em>, new model suggestions, and training models.<p>It\u2019s open source, so you can deploy it on your own infrastructure, use our hosted version with 0 markup, or read how we design for maximum availability on our website."},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: We built open OpenRouter that turns usage into a better model"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/experientiallabs/experiential"}},"_tags":["story","author_SilenN","story_49471407","show_hn"],"author":"SilenN","children":[49471977,49477102,49478084,49476304,49476057,49473195,49473165,49472738,49478018,49472268,49473410,49476194,49474138,49473874,49474132,49484242,49482736,49471998,49475767,49540041,49485872,49503144,49527839,49473371,49476654,49474604],"created_at":"2026-08-27T21:18:35Z","created_at_i":1787865515,"num_comments":47,"objectID":"49471407","points":222,"story_id":49471407,"story_text":"Hi HN, we built an open source model gateway. It&#x27;s a single place to manage our own self hosted, frontier, and open source models in one place.<p>It\u2019s is rust native, built for concurrency, and implements all the config quirks across models and providers (streaming formats, tool calls, model parameters, rate limits, and different error behavior).<p>The gateway adds under 1 ms for BYOK requests and under 2 ms when Experiential supplies the provider key. It has every major inference provider, and 1000+ models refreshed daily via a codex agent that opens a PR.<p>Compared to other similar projects we\u2019re open source, take no markup, allow you to mix local models with a marketplace, and use your traffic to (opt in) train you a model. Simple routing doesn\u2019t warrant a 10% token markup.<p>The way we do this is given standardized OTel traces, we mine representative real tasks, use text world models to simulate rollouts for various models, apply an LLM judge, and fit a nearest neighbor classifier on top of an embedding of a prompt to decide the optimal model for each request. Usually this can map out a better pareto curve on cost&#x2F;quality than just calling single models but it\u2019s not perfect.<p>Using these simulations we can also do things like suggesting cache hit optimizations, new model suggestions, and training models.<p>It\u2019s open source, so you can deploy it on your own infrastructure, use our hosted version with 0 markup, or read how we design for maximum availability on our website.","title":"Show HN: We built open OpenRouter that turns usage into a better model","updated_at":"2026-09-06T16:08:56Z","url":"https://github.com/experientiallabs/experiential"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"loveparade"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"Horrible advice. This may have been good advice 10 years ago, but not today. There are no positions for people who &quot;kind of understand how toy <em>LLMs</em> work&quot; because so many engineers do these days. Most of the real <em>LLM</em> <em>optimization</em> work is at the edge of research and highly proprietary and not something you could ever do without infra that <em>costs</em> millions.<p>But of course, 10 years ago this wasn't obvious."},"story_title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["llm"],"value":"I were 17, I'd learn how to build <em>LLMs</em> from scratch"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://twitter.com/paulg/status/2091544343589060625"}},"_tags":["comment","author_loveparade","story_49412396"],"author":"loveparade","children":[49441978,49416796,49416374,49416868,49436988],"comment_text":"Horrible advice. This may have been good advice 10 years ago, but not today. There are no positions for people who &quot;kind of understand how toy LLMs work&quot; because so many engineers do these days. Most of the real LLM optimization work is at the edge of research and highly proprietary and not something you could ever do without infra that costs millions.<p>But of course, 10 years ago this wasn&#x27;t obvious.","created_at":"2026-08-24T07:36:22Z","created_at_i":1787556982,"objectID":"49416354","parent_id":49412396,"story_id":49412396,"story_title":"I were 17, I'd learn how to build LLMs from scratch","story_url":"https://twitter.com/paulg/status/2091544343589060625","updated_at":"2026-08-28T20:40:44Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"jephs"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"The underlying <em>LLMs</em> <i>do</i>, but we choose not to use the capability because it's expensive and doesn't quite work as well as we'd like it to, or quite in the way that we'd like it to.<p>We are perfectly capable of running <em>LLMs</em> in a way that does a backward pass to update some or all of its weights after every user message. But, naively implemented, you only get partial, fragmentary absorption of the info in those messages, it <em>costs</em> three times as much compute, and you lose out on the ability to implement a ton of <em>optimizations</em> that making modern <em>LLM</em> serving economical.<p>If you want to do it, though, ask your friendly neighborhood robot to get it working with a tiny model (whose full precision weights fit several-times-over on your machine's resources)."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Sol loves to cheat"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://jumploops.com/blog/sol-loves-to-cheat/"}},"_tags":["comment","author_jephs","story_49348189"],"author":"jephs","children":[49373915],"comment_text":"The underlying LLMs <i>do</i>, but we choose not to use the capability because it&#x27;s expensive and doesn&#x27;t quite work as well as we&#x27;d like it to, or quite in the way that we&#x27;d like it to.<p>We are perfectly capable of running LLMs in a way that does a backward pass to update some or all of its weights after every user message. But, naively implemented, you only get partial, fragmentary absorption of the info in those messages, it costs three times as much compute, and you lose out on the ability to implement a ton of optimizations that making modern LLM serving economical.<p>If you want to do it, though, ask your friendly neighborhood robot to get it working with a tiny model (whose full precision weights fit several-times-over on your machine&#x27;s resources).","created_at":"2026-08-20T12:22:26Z","created_at_i":1787228546,"objectID":"49373647","parent_id":49372400,"story_id":49348189,"story_title":"Sol loves to cheat","story_url":"https://jumploops.com/blog/sol-loves-to-cheat/","updated_at":"2026-08-20T13:22:12Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"abdik"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"Hi HN! I'm Bek, founder of Speko, a platform that finds an optimal combination of speech-to-text, <em>LLM</em>, and text-to-speech models, given your constraints, among all our public benchmarked options, and tells you why.<p>Demo: <a href=\"https://www.youtube.com/watch?v=no2LY2gRh-c\" rel=\"nofollow\">https://www.youtube.com/watch?v=no2LY2gRh-c</a><p>Typical production voice agent is an ensemble of three models: STT, an <em>LLM</em>, and TTS.<p>Each of those layers offers a dozen credible vendors, and each month there are new models on the market. Almost everyone evaluates once, picks a stack of their choice, and never rechecks because switching from a vendor to another involves yet another integration and arguments about the numbers.<p>The result is that you use voice agents running last quarter's models while better and cheaper options are available.<p>Before founding Speko, I spent four years as cofounder and CTO building voice agents for enterprises across Asia in 10+ languages. Each time a new speech model would arrive, we repeated the same ritual: hire native-speaking raters, benchmark it against our existing stack, and update production if it improved. Speko turns this process into an API. A team running thousands of calls a day told us: &quot;we can literally go to this dashboard, switch the model, and it will do it for us.&quot;<p>How it works: you send a request with your <em>optimization</em> criteria (accuracy, latency, <em>cost</em> or balanced), language and region. The router filters to models which we measured for the given combination of constraints, benchmarks them, selects the winner, and returns a response with headers containing provider, model names, and the scores. The gateway prefetches signed session plans, so a new session dials the provider straight from memory; no control-plane round trip while a caller waits.<p>Failover happens only during connection setup stage: if the provider refuses the connection attempt, we start connecting to the runners-up.<p>Some of the customer stories: one founder came to us not knowing what to pick at all: he gave us his use case and now routes everything through the platform. A property management AI runs LiveKit in Python and had not updated STT or TTS since launch: they did not know their STT had high error rates on their calls, better options existed, and swapping always looked like an R&amp;D project. One team did not know which models to pick for Spanish. A medical team did not know which STT handles medical vocabulary best. In every case we helped find the right stack from the benchmarks, and now they route through us.<p>The measuring part is public: we pass the same inputs to every model in one region in different dated runs and we publish the boards, including those where our selections perform worse than alternatives. A launch demo answers which 30-second clip sounds better; production asks which model survives minute eight, so we test spontaneous speech, money and dates, ten-minute takes, and the rankings change. We trained an automatic scorer for TTS naturalness on our blind head-to-head listening votes; on providers it has never seen a vote for, it picks the same winner our raters do about as often as raters agree with each other.<p>We don't train or sell models ourselves, that's precisely how we keep our rankings impartial.<p>We also open sourced the gateway for teams who want to avoid an extra network hop on the audio path and don't want to share keys with our cloud (<a href=\"https://github.com/SpekoAI/gateway\" rel=\"nofollow\">https://github.com/SpekoAI/gateway</a>, MIT): one Go binary, which is running as a sidecar in your agent's container, speaks one local protocol over Unix socket, pins provider hosts and attaches your keys. In BYOK mode it doesn't communicate with us at all.<p>Notice that the anonymous, content-free telemetry is enabled by default, and one env var disables it.<p><em>Cost</em>: the gateway and BYOK setup will be free forever, we charge for the hosted router and managed keys with consolidated billing. Since we started the batch in late June, external usage has grown about 25 percent per week on average, front-loaded toward the launch weeks.<p>I would love feedback from the community: how do you pick speech models now, and what makes you trust the third-party benchmark?<p><a href=\"https://speko.ai/\">https://speko.ai/</a>"},"title":{"matchLevel":"none","matchedWords":[],"value":"Launch HN: Speko (YC S26) \u2013 OpenRouter for Voice AI"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://speko.ai/"}},"_tags":["story","author_abdik","story_49332751","launch_hn"],"author":"abdik","children":[49363247,49353640,49370038,49334259,49333835,49345850,49340337,49334453,49335012,49406203,49333279,49334769,49333588,49356537,49334290,49335320,49333760,49337578,49341187,49333353,49335174,49333830,49342326,49333218,49336512,49340627],"created_at":"2026-08-17T15:36:18Z","created_at_i":1786980978,"num_comments":69,"objectID":"49332751","points":118,"story_id":49332751,"story_text":"Hi HN! I&#x27;m Bek, founder of Speko, a platform that finds an optimal combination of speech-to-text, LLM, and text-to-speech models, given your constraints, among all our public benchmarked options, and tells you why.<p>Demo: <a href=\"https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=no2LY2gRh-c\" rel=\"nofollow\">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=no2LY2gRh-c</a><p>Typical production voice agent is an ensemble of three models: STT, an LLM, and TTS.<p>Each of those layers offers a dozen credible vendors, and each month there are new models on the market. Almost everyone evaluates once, picks a stack of their choice, and never rechecks because switching from a vendor to another involves yet another integration and arguments about the numbers.<p>The result is that you use voice agents running last quarter&#x27;s models while better and cheaper options are available.<p>Before founding Speko, I spent four years as cofounder and CTO building voice agents for enterprises across Asia in 10+ languages. Each time a new speech model would arrive, we repeated the same ritual: hire native-speaking raters, benchmark it against our existing stack, and update production if it improved. Speko turns this process into an API. A team running thousands of calls a day told us: &quot;we can literally go to this dashboard, switch the model, and it will do it for us.&quot;<p>How it works: you send a request with your optimization criteria (accuracy, latency, cost or balanced), language and region. The router filters to models which we measured for the given combination of constraints, benchmarks them, selects the winner, and returns a response with headers containing provider, model names, and the scores. The gateway prefetches signed session plans, so a new session dials the provider straight from memory; no control-plane round trip while a caller waits.<p>Failover happens only during connection setup stage: if the provider refuses the connection attempt, we start connecting to the runners-up.<p>Some of the customer stories: one founder came to us not knowing what to pick at all: he gave us his use case and now routes everything through the platform. A property management AI runs LiveKit in Python and had not updated STT or TTS since launch: they did not know their STT had high error rates on their calls, better options existed, and swapping always looked like an R&amp;D project. One team did not know which models to pick for Spanish. A medical team did not know which STT handles medical vocabulary best. In every case we helped find the right stack from the benchmarks, and now they route through us.<p>The measuring part is public: we pass the same inputs to every model in one region in different dated runs and we publish the boards, including those where our selections perform worse than alternatives. A launch demo answers which 30-second clip sounds better; production asks which model survives minute eight, so we test spontaneous speech, money and dates, ten-minute takes, and the rankings change. We trained an automatic scorer for TTS naturalness on our blind head-to-head listening votes; on providers it has never seen a vote for, it picks the same winner our raters do about as often as raters agree with each other.<p>We don&#x27;t train or sell models ourselves, that&#x27;s precisely how we keep our rankings impartial.<p>We also open sourced the gateway for teams who want to avoid an extra network hop on the audio path and don&#x27;t want to share keys with our cloud (<a href=\"https:&#x2F;&#x2F;github.com&#x2F;SpekoAI&#x2F;gateway\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;SpekoAI&#x2F;gateway</a>, MIT): one Go binary, which is running as a sidecar in your agent&#x27;s container, speaks one local protocol over Unix socket, pins provider hosts and attaches your keys. In BYOK mode it doesn&#x27;t communicate with us at all.<p>Notice that the anonymous, content-free telemetry is enabled by default, and one env var disables it.<p>Cost: the gateway and BYOK setup will be free forever, we charge for the hosted router and managed keys with consolidated billing. Since we started the batch in late June, external usage has grown about 25 percent per week on average, front-loaded toward the launch weeks.<p>I would love feedback from the community: how do you pick speech models now, and what makes you trust the third-party benchmark?<p><a href=\"https:&#x2F;&#x2F;speko.ai&#x2F;\">https:&#x2F;&#x2F;speko.ai&#x2F;</a>","title":"Launch HN: Speko (YC S26) \u2013 OpenRouter for Voice AI","updated_at":"2026-09-03T11:59:16Z","url":"https://speko.ai/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"applfanboysbgon"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"No. This conveys a fundamental misunderstanding of how anything pertaining to programming works. This will never happen, <i>ever</i>. For example, take...<p><pre><code>  printf(&quot;Hello, world&quot;);\n</code></pre>\nvs. a plausible illustration of how it might be compiled down to machine code...<p><pre><code>  48 65 6C 6C 6F 2C 20 77 6F 72 6C 64\n  48 83 EC 28\n  48 8D 0D F5 0F 00 00\n  E8 F0 00 00 00\n  33 C0\n  48 83 C4 28\n  C3\n</code></pre>\nThe latter now takes up 10x as many tokens (= 10x the <em>cost</em>/time, + context penalties), and is now architecture-specific, impossible to apply non-brittle program-wide <em>optimizations</em> to, etc. There is absolutely zero reason to ever have the <em>LLM</em> act as a compiler no matter how fast it is. Even if you believe LLMs will reach a state where they can actually generate good code at this level, you would be better off having them generate the compiler they would use."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Accelerating GPT-5.6 Sol Ultrafast"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultrafast-with-openai"}},"_tags":["comment","author_applfanboysbgon","story_49289844"],"author":"applfanboysbgon","children":[49291233,49290935],"comment_text":"No. This conveys a fundamental misunderstanding of how anything pertaining to programming works. This will never happen, <i>ever</i>. For example, take...<p><pre><code>  printf(&quot;Hello, world&quot;);\n</code></pre>\nvs. a plausible illustration of how it might be compiled down to machine code...<p><pre><code>  48 65 6C 6C 6F 2C 20 77 6F 72 6C 64\n  48 83 EC 28\n  48 8D 0D F5 0F 00 00\n  E8 F0 00 00 00\n  33 C0\n  48 83 C4 28\n  C3\n</code></pre>\nThe latter now takes up 10x as many tokens (= 10x the cost&#x2F;time, + context penalties), and is now architecture-specific, impossible to apply non-brittle program-wide optimizations to, etc. There is absolutely zero reason to ever have the LLM act as a compiler no matter how fast it is. Even if you believe LLMs will reach a state where they can actually generate good code at this level, you would be better off having them generate the compiler they would use.","created_at":"2026-08-13T19:26:30Z","created_at_i":1786649190,"objectID":"49290779","parent_id":49290612,"story_id":49289844,"story_title":"Accelerating GPT-5.6 Sol Ultrafast","story_url":"https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultrafast-with-openai","updated_at":"2026-08-14T12:10:20Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"oliveratarcform"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"Arcforma AI (arcforma.ai) | AI Engineer (Marketing / Construction / Fashion &amp; Luxury) | REMOTE (US only) or ONSITE NYC | Full-Time $80k-$120k + equity, or Contract $50-$150/hr | jobs@arcforma.ai<p>We build production AI systems inside other people's companies: hotel groups, automotive retail, creative agencies, consumer brands. One industry at a time, sitting next to the operator whose job the system changes. Very little of it is generic <em>LLM</em> plumbing.<p>Two ways in.<p>CONTRACT, PROJECT-BASED ($50-$150/hr). For people who already have the domain. A scoped first project, typically three to six weeks. This is most of what we need right now, and it is how we would rather start with anyone senior.<p>FULL-TIME ($80k-$120k + equity). For strong engineers earlier on who want to go deep into one industry with us instead of staying a generalist. We teach you the vertical. You bring the engineering.<p>Pick one track. Depth in one beats a tour of all three.<p>MARKETING. Lifecycle and CRM automation, brief-to-asset pipelines, creative variants with brand guardrails that hold, attribution. You have fought HubSpot, Klaviyo, and the Meta and Google ad APIs.<p>CONSTRUCTION. Invoices in, mapped to contracts, schedule of values, and <em>cost</em> codes, back out as pay applications. Then RFIs, submittals, change orders, schedule impacts. You know why generic invoice parsers stall when half the input is handwritten. Weighted toward ex construction tech and people who have run operations at a GC or a sub.<p>FASHION AND LUXURY. Merchandising and assortment planning, tech packs and PLM, Shopify Plus and PIM, allocation, replenishment, demand forecasting, size curves, markdown <em>optimization</em>, clienteling, creative workflows.<p>ALL TRACKS: You own a system end to end (data in, evals, an interface the client actually opens), and you can sit with a non-technical operator and leave with the right spec.<p>TO APPLY: email jobs@arcforma.ai with the subject line<p>HN | TRACK | Your Name | your applicable background in one short line<p>TRACK is MKT, CON, or FASH. For example:<p>HN | FASH | Jane Okafor | 6 yrs building allocation and planning systems at a DTC apparel brand<p>The body: what you shipped that is applicable to the role, a link we can open (GitHub, a writeup, a demo), LinkedIn, and your timezone plus whether you want contract or full-time and at what rate.<p>Every application gets a reply within five business days. 20 minute call, paid trial on real work, then a project or an offer."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Ask HN: Who is hiring? (August 2026)"}},"_tags":["comment","author_oliveratarcform","story_49156683"],"author":"oliveratarcform","children":[49408353,49250174,49250180],"comment_text":"Arcforma AI (arcforma.ai) | AI Engineer (Marketing &#x2F; Construction &#x2F; Fashion &amp; Luxury) | REMOTE (US only) or ONSITE NYC | Full-Time $80k-$120k + equity, or Contract $50-$150&#x2F;hr | jobs@arcforma.ai<p>We build production AI systems inside other people&#x27;s companies: hotel groups, automotive retail, creative agencies, consumer brands. One industry at a time, sitting next to the operator whose job the system changes. Very little of it is generic LLM plumbing.<p>Two ways in.<p>CONTRACT, PROJECT-BASED ($50-$150&#x2F;hr). For people who already have the domain. A scoped first project, typically three to six weeks. This is most of what we need right now, and it is how we would rather start with anyone senior.<p>FULL-TIME ($80k-$120k + equity). For strong engineers earlier on who want to go deep into one industry with us instead of staying a generalist. We teach you the vertical. You bring the engineering.<p>Pick one track. Depth in one beats a tour of all three.<p>MARKETING. Lifecycle and CRM automation, brief-to-asset pipelines, creative variants with brand guardrails that hold, attribution. You have fought HubSpot, Klaviyo, and the Meta and Google ad APIs.<p>CONSTRUCTION. Invoices in, mapped to contracts, schedule of values, and cost codes, back out as pay applications. Then RFIs, submittals, change orders, schedule impacts. You know why generic invoice parsers stall when half the input is handwritten. Weighted toward ex construction tech and people who have run operations at a GC or a sub.<p>FASHION AND LUXURY. Merchandising and assortment planning, tech packs and PLM, Shopify Plus and PIM, allocation, replenishment, demand forecasting, size curves, markdown optimization, clienteling, creative workflows.<p>ALL TRACKS: You own a system end to end (data in, evals, an interface the client actually opens), and you can sit with a non-technical operator and leave with the right spec.<p>TO APPLY: email jobs@arcforma.ai with the subject line<p>HN | TRACK | Your Name | your applicable background in one short line<p>TRACK is MKT, CON, or FASH. For example:<p>HN | FASH | Jane Okafor | 6 yrs building allocation and planning systems at a DTC apparel brand<p>The body: what you shipped that is applicable to the role, a link we can open (GitHub, a writeup, a demo), LinkedIn, and your timezone plus whether you want contract or full-time and at what rate.<p>Every application gets a reply within five business days. 20 minute call, paid trial on real work, then a project or an offer.","created_at":"2026-08-10T18:55:58Z","created_at_i":1786388158,"objectID":"49248055","parent_id":49156683,"story_id":49156683,"story_title":"Ask HN: Who is hiring? (August 2026)","updated_at":"2026-08-23T12:37:16Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"bormisov1"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"Hello!<p>I\u2019m interested in your Senior Fullstack Engineer position. I believe my experience is a strong match:<p>- 7 years of fullstack development, primarily with TypeScript and Node.js,<p>- built a React frontend at Reelay for an AI voice assistant analyzing customer calls,<p>- the product processed audio through transcription, chunking, embeddings, Weaviate retrieval, and <em>LLM</em> analysis to generate summaries, agreements, action items, and key questions,<p>- developed the frontend against GraphQL APIs while also working on retrieval logic, structured <em>LLM</em> outputs, hallucination reduction, and latency/<em>cost</em> <em>optimization</em>,<p>- designed high-load, real-time systems using WebSockets, Centrifugo, RabbitMQ, Redis, PostgreSQL, and ClickHouse,<p>- comfortable taking ownership: previously led three developers, distributed tasks, reviewed code, and mentored teammates,<p>- experienced working in small international teams and communicating in English.<p>I would be glad to discuss the role and the React product in more detail.<p>Best regards,<p>Vlad<p>bormisov1@gmail.com"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Ask HN: Who is hiring? (August 2026)"}},"_tags":["comment","author_bormisov1","story_49156683"],"author":"bormisov1","comment_text":"Hello!<p>I\u2019m interested in your Senior Fullstack Engineer position. I believe my experience is a strong match:<p>- 7 years of fullstack development, primarily with TypeScript and Node.js,<p>- built a React frontend at Reelay for an AI voice assistant analyzing customer calls,<p>- the product processed audio through transcription, chunking, embeddings, Weaviate retrieval, and LLM analysis to generate summaries, agreements, action items, and key questions,<p>- developed the frontend against GraphQL APIs while also working on retrieval logic, structured LLM outputs, hallucination reduction, and latency&#x2F;cost optimization,<p>- designed high-load, real-time systems using WebSockets, Centrifugo, RabbitMQ, Redis, PostgreSQL, and ClickHouse,<p>- comfortable taking ownership: previously led three developers, distributed tasks, reviewed code, and mentored teammates,<p>- experienced working in small international teams and communicating in English.<p>I would be glad to discuss the role and the React product in more detail.<p>Best regards,<p>Vlad<p>bormisov1@gmail.com","created_at":"2026-08-10T09:00:18Z","created_at_i":1786352418,"objectID":"49241158","parent_id":49234667,"story_id":49156683,"story_title":"Ask HN: Who is hiring? (August 2026)","updated_at":"2026-08-12T12:14:44Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"johnsmith1840"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"More like <em>LLMs</em> are near commodity at the common tier the ones who wins are those who can inference the cheapest and google is easily thr best positioned to do that with many years of custom chip model <em>optimization</em>.<p>Google is going to be able to push <em>cost</em> down further and for longer than anyone else. They also rightfully believe this is a company ending gamble if they fail they could be a second tier player for decades.<p>So better inference margins, far better operation experience serving cheap AI at massive scale, a massive warchest, and the fear of complete company irrelevence.<p>That's going to be one hell of a company to beat for free tier <em>LLMs</em>. Also chatgpt is impressive but it's notthing like google search quite yet.  On top of that AI labs must use google's product for their AI.<p>They also own more data by an exponential margin. Anything but dominance of free AI on google's side would mean they are just so incompetent they deserve to fail."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Improving GPT\u20115.6 Sol in ChatGPT, expanding GPT\u20115.6 Luna access for free users"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://openai.com/index/improving-gpt-5-6-sol-in-chatgpt/"}},"_tags":["comment","author_johnsmith1840","story_49199357"],"author":"johnsmith1840","comment_text":"More like LLMs are near commodity at the common tier the ones who wins are those who can inference the cheapest and google is easily thr best positioned to do that with many years of custom chip model optimization.<p>Google is going to be able to push cost down further and for longer than anyone else. They also rightfully believe this is a company ending gamble if they fail they could be a second tier player for decades.<p>So better inference margins, far better operation experience serving cheap AI at massive scale, a massive warchest, and the fear of complete company irrelevence.<p>That&#x27;s going to be one hell of a company to beat for free tier LLMs. Also chatgpt is impressive but it&#x27;s notthing like google search quite yet.  On top of that AI labs must use google&#x27;s product for their AI.<p>They also own more data by an exponential margin. Anything but dominance of free AI on google&#x27;s side would mean they are just so incompetent they deserve to fail.","created_at":"2026-08-07T01:49:21Z","created_at_i":1786067361,"objectID":"49205013","parent_id":49202959,"story_id":49199357,"story_title":"Improving GPT\u20115.6 Sol in ChatGPT, expanding GPT\u20115.6 Luna access for free users","story_url":"https://openai.com/index/improving-gpt-5-6-sol-in-chatgpt/","updated_at":"2026-08-07T23:00:29Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"imvivekvermaaa"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"Location: Haridwar, India<p>Remote: Yes (worldwide)<p>Willing to relocate: Yes<p>Technologies: TypeScript, React, Next.js, Node.js, Python, PostgreSQL, AWS, React Native/Expo, <em>LLM</em> apps (RAG, structured extraction, agents), vLLM<p>R\u00e9sum\u00e9: <a href=\"https://imvivekvermaa.in/vivek-verma-resume.pdf\" rel=\"nofollow\">https://imvivekvermaa.in/vivek-verma-resume.pdf</a><p>GitHub: <a href=\"https://github.com/imvivekvermaa\" rel=\"nofollow\">https://github.com/imvivekvermaa</a><p>Website: <a href=\"https://imvivekvermaa.in\" rel=\"nofollow\">https://imvivekvermaa.in</a><p>Email: imvivekvermaa@gmail.com<p>I ship 0-to-1 systems, then go a layer below them.<p>A real-estate platform I built and still work on serves ~30K daily visitors.\nI architected an offline-first POS that keeps 11 terminals selling through\ninternet outages. And I traced multi-minute reports in a 15-year-old database\ndown to missing keys and absent indexes \u2014 got them under 3 seconds.<p>Right now I'm going deep on <em>LLM</em> inference performance: what a token actually\n<em>costs</em>, where the latency goes, and when an <em>optimisation</em> stops paying for\nitself. I publish every measurement, including the ones that didn't work:\n<a href=\"https://imvivekvermaa.in/log\" rel=\"nofollow\">https://imvivekvermaa.in/log</a>"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Ask HN: Who wants to be hired? (August 2026)"}},"_tags":["comment","author_imvivekvermaaa","story_49156682"],"author":"imvivekvermaaa","comment_text":"Location: Haridwar, India<p>Remote: Yes (worldwide)<p>Willing to relocate: Yes<p>Technologies: TypeScript, React, Next.js, Node.js, Python, PostgreSQL, AWS, React Native&#x2F;Expo, LLM apps (RAG, structured extraction, agents), vLLM<p>R\u00e9sum\u00e9: <a href=\"https:&#x2F;&#x2F;imvivekvermaa.in&#x2F;vivek-verma-resume.pdf\" rel=\"nofollow\">https:&#x2F;&#x2F;imvivekvermaa.in&#x2F;vivek-verma-resume.pdf</a><p>GitHub: <a href=\"https:&#x2F;&#x2F;github.com&#x2F;imvivekvermaa\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;imvivekvermaa</a><p>Website: <a href=\"https:&#x2F;&#x2F;imvivekvermaa.in\" rel=\"nofollow\">https:&#x2F;&#x2F;imvivekvermaa.in</a><p>Email: imvivekvermaa@gmail.com<p>I ship 0-to-1 systems, then go a layer below them.<p>A real-estate platform I built and still work on serves ~30K daily visitors.\nI architected an offline-first POS that keeps 11 terminals selling through\ninternet outages. And I traced multi-minute reports in a 15-year-old database\ndown to missing keys and absent indexes \u2014 got them under 3 seconds.<p>Right now I&#x27;m going deep on LLM inference performance: what a token actually\ncosts, where the latency goes, and when an optimisation stops paying for\nitself. I publish every measurement, including the ones that didn&#x27;t work:\n<a href=\"https:&#x2F;&#x2F;imvivekvermaa.in&#x2F;log\" rel=\"nofollow\">https:&#x2F;&#x2F;imvivekvermaa.in&#x2F;log</a>","created_at":"2026-08-05T12:31:47Z","created_at_i":1785933107,"objectID":"49181943","parent_id":49156682,"story_id":49156682,"story_title":"Ask HN: Who wants to be hired? (August 2026)","updated_at":"2026-08-05T12:33:35Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"MehulMangave"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"Location: SF Bay Area<p>Remote: Yes<p>Willing to relocate: Yes<p>Technologies: AWS, Azure, Terraform, Python, CI/CD, Docker, Kubernetes, Backend systems, RAG/<em>LLM</em> pipelines, AI agents/MCP<p>R\u00e9sum\u00e9/CV: <a href=\"https://drive.google.com/file/d/1QS9N0TTYbmhfTP9U_Dzx70T_95_-SZdT/view?usp=sharing\" rel=\"nofollow\">https://drive.google.com/file/d/1QS9N0TTYbmhfTP9U_Dzx70T_95_...</a><p>Portfolio: <a href=\"https://mehul-portfolio-sand.vercel.app/\" rel=\"nofollow\">https://mehul-portfolio-sand.vercel.app/</a><p>Email: mehul.mangave@sjsu.edu<p>I am a recent MS Computer Engineering grad (SJSU, May 2026) looking for infrastructure, platform, DevOps, or AI systems engineering roles. At Syngenta I scaled AWS infrastructure across 100+ accounts, built reusable Terraform modules now used by 50+ teams, and drove roughly $500K in annual <em>cost</em> savings through resource <em>optimization</em>. At State Street I built a RAG-based internal AI system on Confluence data plus Terraform automation tooling that cut manual deployment effort by 60%. More recently I built Magpie, a multi-agent tool that won Most Innovative Use of Agents at the Harness hackathon at the AWS Builder Loft, and a working MCP server from scratch. I like owning systems end to end, from design to production, especially where reliability and performance have real consequences. Happy to work onsite, hybrid, or remote."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Ask HN: Who wants to be hired? (August 2026)"}},"_tags":["comment","author_MehulMangave","story_49156682"],"author":"MehulMangave","comment_text":"Location: SF Bay Area<p>Remote: Yes<p>Willing to relocate: Yes<p>Technologies: AWS, Azure, Terraform, Python, CI&#x2F;CD, Docker, Kubernetes, Backend systems, RAG&#x2F;LLM pipelines, AI agents&#x2F;MCP<p>R\u00e9sum\u00e9&#x2F;CV: <a href=\"https:&#x2F;&#x2F;drive.google.com&#x2F;file&#x2F;d&#x2F;1QS9N0TTYbmhfTP9U_Dzx70T_95_-SZdT&#x2F;view?usp=sharing\" rel=\"nofollow\">https:&#x2F;&#x2F;drive.google.com&#x2F;file&#x2F;d&#x2F;1QS9N0TTYbmhfTP9U_Dzx70T_95_...</a><p>Portfolio: <a href=\"https:&#x2F;&#x2F;mehul-portfolio-sand.vercel.app&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;mehul-portfolio-sand.vercel.app&#x2F;</a><p>Email: mehul.mangave@sjsu.edu<p>I am a recent MS Computer Engineering grad (SJSU, May 2026) looking for infrastructure, platform, DevOps, or AI systems engineering roles. At Syngenta I scaled AWS infrastructure across 100+ accounts, built reusable Terraform modules now used by 50+ teams, and drove roughly $500K in annual cost savings through resource optimization. At State Street I built a RAG-based internal AI system on Confluence data plus Terraform automation tooling that cut manual deployment effort by 60%. More recently I built Magpie, a multi-agent tool that won Most Innovative Use of Agents at the Harness hackathon at the AWS Builder Loft, and a working MCP server from scratch. I like owning systems end to end, from design to production, especially where reliability and performance have real consequences. Happy to work onsite, hybrid, or remote.","created_at":"2026-08-05T01:30:36Z","created_at_i":1785893436,"objectID":"49177532","parent_id":49156682,"story_id":49156682,"story_title":"Ask HN: Who wants to be hired? (August 2026)","updated_at":"2026-08-05T01:31:32Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"py4"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"I am not sure. Why do you need domain expertise beyond being able to draft a verifier for the problem? Once you have the verifier, it's just a matter of compute. You can argue that being a domain expert allows you to narrow down the search space for the <em>LLM</em> and save time/compute <em>costs</em>. This is partially true, but LLMs are getting better and better at search (they can already do end-to-end performance <em>optimization</em> faster than performance experts at a FAANG company I work at), and compute to maintain the same intelligence level is getting cheaper.<p>A concrete example: GPU performance <em>optimization</em> for a kernel. This was (and still is) a very niche domain with not many top-notch experts. But kernel performance and characteristics are easily verifiable. You can run the agent in a closed loop for it to improve iteratively (and people are already doing it, coming up with kernels better than human-written ones).<p>You see Tao's example because:<p>1. He is curious (so he asks detailed questions, which are not necessarily needed in a closed-loop <em>optimization</em>).<p>2. Verification in math is harder. Many math tasks used in RL are easily verifiable. But for advanced open conjectures that require long proofs, you cannot trust the proof directly from the <em>LLM</em> (so it's not as easily verifiable as basic math problems or code). The model needs to write it in Lean, and you still need to make sure the Lean implementation correctly captures the specification of the problem. So you still need a human for verification in advanced math. But I don't see why you would need this in domains like performance improvement."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"LLMs reward expertise"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://www.seangoedecke.com/llms-reward-expertise/"}},"_tags":["comment","author_py4","story_49161518"],"author":"py4","comment_text":"I am not sure. Why do you need domain expertise beyond being able to draft a verifier for the problem? Once you have the verifier, it&#x27;s just a matter of compute. You can argue that being a domain expert allows you to narrow down the search space for the LLM and save time&#x2F;compute costs. This is partially true, but LLMs are getting better and better at search (they can already do end-to-end performance optimization faster than performance experts at a FAANG company I work at), and compute to maintain the same intelligence level is getting cheaper.<p>A concrete example: GPU performance optimization for a kernel. This was (and still is) a very niche domain with not many top-notch experts. But kernel performance and characteristics are easily verifiable. You can run the agent in a closed loop for it to improve iteratively (and people are already doing it, coming up with kernels better than human-written ones).<p>You see Tao&#x27;s example because:<p>1. He is curious (so he asks detailed questions, which are not necessarily needed in a closed-loop optimization).<p>2. Verification in math is harder. Many math tasks used in RL are easily verifiable. But for advanced open conjectures that require long proofs, you cannot trust the proof directly from the LLM (so it&#x27;s not as easily verifiable as basic math problems or code). The model needs to write it in Lean, and you still need to make sure the Lean implementation correctly captures the specification of the problem. So you still need a human for verification in advanced math. But I don&#x27;t see why you would need this in domains like performance improvement.","created_at":"2026-08-04T20:14:31Z","created_at_i":1785874471,"objectID":"49174420","parent_id":49161518,"story_id":49161518,"story_title":"LLMs reward expertise","story_url":"https://www.seangoedecke.com/llms-reward-expertise/","updated_at":"2026-08-04T20:48:04Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"unscaled"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"I don't think it was true even in 2023.<p>This sounds like tackling the problems of C++ in the early 2000s.<p>1. Casey Muratori also that DRY shouldn't doesn't have to result in non-performant code.<p>2. Smaller functions, functions that do one-thing: Modern compiler can inline those. There are some edge cases where inlining may make less efficient use of states and loops but I don't think that's a main problem nowadays. I also wouldn't say the extreme version of this idea (very small functions) is still popular. The strongest proponent of this was Uncle Bob, and the last time I've heard him speak about code, he said he now lets the <em>LLM</em> write everything and he only reviews the module hierarchy and maybe the modules' public interfaces.<p>3. Polymorphism instead of ifs and switches was a big fad in the late 1990s until the late 2000s and had some holdouts in the 2010s. It was only ever popular in the Enterprise Java and C++ world (and maybe in Enterprise Smalltalk, never hard). Overuse of runtime polymorphism widely considered bad form in newer static languages like Go and Rust and in most dynamic languages there was always a tacit understanding of &quot;use mostly conditions, add polymorphism if you need extensibility&quot;.<p>In functional languages (or languages heavily influenced by functional programming like Rust, Swift and Kotlin[1]), the classic approach for the type of scenario in this example is to use a sum type, and run a safe exhaustive match/switch on all the variants.<p>4. Hiding internals: The sum type example is telling of modern best-practices. Sum type fields are generally made public. Some languages (e.g. Rust and most pure functional languages) do not support private fields in sum types at all! Other languages (e.g. Kotlin)\n but immutable, so it's easy to maintain invariants without hiding information. Sometimes we do want to hide the type details and wrap it with public-facing type (this is a common pattern with internal error enums in Rust for example). Even in this case, there is no impact since we do not use runtime polymorphism or indirection (that would be Box&lt;T&gt; in Rust).<p>Due to compiler <em>optimizations</em>, hiding internals has marginal performance <em>cost</em> (if any) unless you require runtime polymorphism to achieve it. But why should you?<p>I feel like the performance costs lamented in this article mostly have to do with runtime polymorphism in static languages. And I fully agree here: runtime polymorphism is something that should be avoided when you don't need it[2]. But that's the thing: if you're looking at modern static language codebases, runtime polymorphism is not as hyped as it used to be in the past. Some languages still require heavy use of runtime polymorphism (Go is a good example of this), but other languages more often rely on static polymorphism (Rust) or compile time duck-typing (Zig and you could argue C++ template meta-programming used to do that, albeit quite awkwardly).<p>Even with all the issues you get with polymorphism, I don't think it's the main cause of slow application performance. It be very much the culprit in tight loops inside games, but if you look at the performance issues plaguing everyday apps, I think the two major culprits are endless layers of abstraction (the most quintessential example is basically every sluggish Electron app out there) and blocking the user on slow actions (like network loads).<p>---<p>[1] Even Java had sealed record types for a while now, and I'm sure will see Enterprise frameworks encouraging them in 20 years, when the rest of the world has moved on to spacefaring super-intelligent LLMs. But Enterprise frameworks also don't encourage you to write DRY code or keep your functions short.<p>[2] But do keep in mind that in Java it could be almost zero-<em>cost</em> in many cases. The JIT will monomorphize or bimorphize your classes if you always use the same class at the same callsite. The pointer indirection is not an extra <em>cost</em>, since every non-primitive that doesn't undergo Scalar Replacement[3] lives on the heap, and has a pointer.<p>[3] <a href=\"https://shipilev.net/jvm/anatomy-quarks/18-scalar-replacement/\" rel=\"nofollow\">https://shipilev.net/jvm/anatomy-quarks/18-scalar-replacemen...</a>"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"\"Clean\" Code, Horrible Performance (2023)"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://www.computerenhance.com/p/clean-code-horrible-performance"}},"_tags":["comment","author_unscaled","story_49166331"],"author":"unscaled","children":[49170216],"comment_text":"I don&#x27;t think it was true even in 2023.<p>This sounds like tackling the problems of C++ in the early 2000s.<p>1. Casey Muratori also that DRY shouldn&#x27;t doesn&#x27;t have to result in non-performant code.<p>2. Smaller functions, functions that do one-thing: Modern compiler can inline those. There are some edge cases where inlining may make less efficient use of states and loops but I don&#x27;t think that&#x27;s a main problem nowadays. I also wouldn&#x27;t say the extreme version of this idea (very small functions) is still popular. The strongest proponent of this was Uncle Bob, and the last time I&#x27;ve heard him speak about code, he said he now lets the LLM write everything and he only reviews the module hierarchy and maybe the modules&#x27; public interfaces.<p>3. Polymorphism instead of ifs and switches was a big fad in the late 1990s until the late 2000s and had some holdouts in the 2010s. It was only ever popular in the Enterprise Java and C++ world (and maybe in Enterprise Smalltalk, never hard). Overuse of runtime polymorphism widely considered bad form in newer static languages like Go and Rust and in most dynamic languages there was always a tacit understanding of &quot;use mostly conditions, add polymorphism if you need extensibility&quot;.<p>In functional languages (or languages heavily influenced by functional programming like Rust, Swift and Kotlin[1]), the classic approach for the type of scenario in this example is to use a sum type, and run a safe exhaustive match&#x2F;switch on all the variants.<p>4. Hiding internals: The sum type example is telling of modern best-practices. Sum type fields are generally made public. Some languages (e.g. Rust and most pure functional languages) do not support private fields in sum types at all! Other languages (e.g. Kotlin)\n but immutable, so it&#x27;s easy to maintain invariants without hiding information. Sometimes we do want to hide the type details and wrap it with public-facing type (this is a common pattern with internal error enums in Rust for example). Even in this case, there is no impact since we do not use runtime polymorphism or indirection (that would be Box&lt;T&gt; in Rust).<p>Due to compiler optimizations, hiding internals has marginal performance cost (if any) unless you require runtime polymorphism to achieve it. But why should you?<p>I feel like the performance costs lamented in this article mostly have to do with runtime polymorphism in static languages. And I fully agree here: runtime polymorphism is something that should be avoided when you don&#x27;t need it[2]. But that&#x27;s the thing: if you&#x27;re looking at modern static language codebases, runtime polymorphism is not as hyped as it used to be in the past. Some languages still require heavy use of runtime polymorphism (Go is a good example of this), but other languages more often rely on static polymorphism (Rust) or compile time duck-typing (Zig and you could argue C++ template meta-programming used to do that, albeit quite awkwardly).<p>Even with all the issues you get with polymorphism, I don&#x27;t think it&#x27;s the main cause of slow application performance. It be very much the culprit in tight loops inside games, but if you look at the performance issues plaguing everyday apps, I think the two major culprits are endless layers of abstraction (the most quintessential example is basically every sluggish Electron app out there) and blocking the user on slow actions (like network loads).<p>---<p>[1] Even Java had sealed record types for a while now, and I&#x27;m sure will see Enterprise frameworks encouraging them in 20 years, when the rest of the world has moved on to spacefaring super-intelligent LLMs. But Enterprise frameworks also don&#x27;t encourage you to write DRY code or keep your functions short.<p>[2] But do keep in mind that in Java it could be almost zero-cost in many cases. The JIT will monomorphize or bimorphize your classes if you always use the same class at the same callsite. The pointer indirection is not an extra cost, since every non-primitive that doesn&#x27;t undergo Scalar Replacement[3] lives on the heap, and has a pointer.<p>[3] <a href=\"https:&#x2F;&#x2F;shipilev.net&#x2F;jvm&#x2F;anatomy-quarks&#x2F;18-scalar-replacement&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;shipilev.net&#x2F;jvm&#x2F;anatomy-quarks&#x2F;18-scalar-replacemen...</a>","created_at":"2026-08-04T14:58:06Z","created_at_i":1785855486,"objectID":"49169906","parent_id":49168029,"story_id":49166331,"story_title":"\"Clean\" Code, Horrible Performance (2023)","story_url":"https://www.computerenhance.com/p/clean-code-horrible-performance","updated_at":"2026-08-05T00:35:33Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"supersour"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["llm","cost","optimization"],"value":"Yes this is true to some extent. I've been using <em>LLMs</em> to run some computer vision tests, and I've certainly noticed myself running into the trap of &quot;just one more AI experiment&quot;, or &quot;just one more change&quot; while neglecting to actually properly integrate the learnings into my mental model.<p>However, without an agent running its own experiments on a cloud GPU, would I realistically have invested my limited work hours and tried evaluating 10 different models, each with 10 different tuned parameters, to solve my specific use case?<p>Or would I have tried 1-2 models and spent my time trying to optimize those models?<p>I think there is some merit to the spray and pray approach when one is in the exploration phase of the solution space.<p>Also, on more than one occasion now, I have had fable halve the inference latency of a model simply because the original implementation from an academic included unnecessary GPU-to-CPU-to-GPU transfers or similarly inefficient operations. Those <em>optimizations</em> came at essentially 0 time <em>cost</em> to me and I can verify that the outputs are byte-identical. Pretty sweet!"},"story_title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["llm"],"value":"2x, not 10x: coding with <em>LLMs</em> in 2026"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://obryant.dev/p/2x-not-10x/"}},"_tags":["comment","author_supersour","story_49047839"],"author":"supersour","comment_text":"Yes this is true to some extent. I&#x27;ve been using LLMs to run some computer vision tests, and I&#x27;ve certainly noticed myself running into the trap of &quot;just one more AI experiment&quot;, or &quot;just one more change&quot; while neglecting to actually properly integrate the learnings into my mental model.<p>However, without an agent running its own experiments on a cloud GPU, would I realistically have invested my limited work hours and tried evaluating 10 different models, each with 10 different tuned parameters, to solve my specific use case?<p>Or would I have tried 1-2 models and spent my time trying to optimize those models?<p>I think there is some merit to the spray and pray approach when one is in the exploration phase of the solution space.<p>Also, on more than one occasion now, I have had fable halve the inference latency of a model simply because the original implementation from an academic included unnecessary GPU-to-CPU-to-GPU transfers or similarly inefficient operations. Those optimizations came at essentially 0 time cost to me and I can verify that the outputs are byte-identical. Pretty sweet!","created_at":"2026-07-30T21:15:56Z","created_at_i":1785446156,"objectID":"49115908","parent_id":49115707,"story_id":49047839,"story_title":"2x, not 10x: coding with LLMs in 2026","story_url":"https://obryant.dev/p/2x-not-10x/","updated_at":"2026-07-30T23:59:45Z"}],"hitsPerPage":20,"nbHits":344,"nbPages":18,"page":0,"params":"query=LLM+cost+optimization&advancedSyntax=true&analyticsTags=backend","processingTimeMS":17,"processingTimingsMS":{"_request":{"roundTrip":21},"afterFetch":{"format":{"highlighting":2,"total":2}},"fetch":{"query":9,"scanning":7,"total":17},"total":18},"query":"LLM cost optimization","serverTimeMS":20}
