{"exhaustive":{"nbHits":false,"typo":false},"exhaustiveNbHits":false,"exhaustiveTypo":false,"hits":[{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"michaelgdwn"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["cost","per","task","llm","benchmark"],"value":"The Nano tier is the one I'm watching. For agent workflows where you're making dozens of <em>LLM</em> calls <em>per</em> <em>task</em>, the <em>cost</em> <em>per</em> call matters more than peak capability. Would be interesting to see <em>benchmarks</em> on function calling latency specifically \u2014 that's what matters for agents."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"GPT\u20115.4 Mini and Nano"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://openai.com/index/introducing-gpt-5-4-mini-and-nano"}},"_tags":["comment","author_michaelgdwn","story_47415441"],"author":"michaelgdwn","comment_text":"The Nano tier is the one I&#x27;m watching. For agent workflows where you&#x27;re making dozens of LLM calls per task, the cost per call matters more than peak capability. Would be interesting to see benchmarks on function calling latency specifically \u2014 that&#x27;s what matters for agents.","created_at":"2026-03-17T22:47:03Z","created_at_i":1773787623,"objectID":"47419406","parent_id":47415441,"story_id":47415441,"story_title":"GPT\u20115.4 Mini and Nano","story_url":"https://openai.com/index/introducing-gpt-5-4-mini-and-nano","updated_at":"2026-03-18T02:21:12Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"napowderly"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["cost","per","task","llm","benchmark"],"value":"Hi HN \u2014 we build atopile, an open-source language and compiler for designing circuit boards in code. People kept asking us &quot;can AI just design the board?&quot; \u2014 so we measured it.<p>EEBench is a set of original EE <em>tasks</em>: design-from-spec, datasheet comprehension, design review, picking real parts that actually exist. Agents submit design source, not prose. The grader runs compiler checks and simulations, and scores measured behavior \u2014 transients, thresholds, tolerance corners \u2014 against the spec. No human graders, no <em>LLM</em>-as-judge. When a score says the rail sagged below 3.0 V during brown-out, that's a waveform, not an opinion.<p>Things that surprised us:<p>The frontier cluster spans roughly 45\u201372% (Claude Fable 5 on top at 71.7%).\nOpen-weight models that do well on coding <em>benchmarks</em> score 5\u201313% here.\n<em>Cost</em> <em>per</em> <em>task</em> varies 15x between models with similar scores.\nHonest limitations: the <em>task</em> set is curated from a larger pack to separate the frontier models, and it's growing. The pack is held out to keep it out of training data, so you can't self-run it yet. If you want a model on the board, email us and we'll run it.<p>Happy to go deep on the grading harness \u2014 ask us anything."},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: EEBench \u2013 AI agents design real circuits, graded by physics simulation"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://www.eebench.org/"}},"_tags":["story","author_napowderly","story_48766320","show_hn"],"author":"napowderly","created_at":"2026-07-02T19:37:25Z","created_at_i":1783021045,"num_comments":0,"objectID":"48766320","points":1,"story_id":48766320,"story_text":"Hi HN \u2014 we build atopile, an open-source language and compiler for designing circuit boards in code. People kept asking us &quot;can AI just design the board?&quot; \u2014 so we measured it.<p>EEBench is a set of original EE tasks: design-from-spec, datasheet comprehension, design review, picking real parts that actually exist. Agents submit design source, not prose. The grader runs compiler checks and simulations, and scores measured behavior \u2014 transients, thresholds, tolerance corners \u2014 against the spec. No human graders, no LLM-as-judge. When a score says the rail sagged below 3.0 V during brown-out, that&#x27;s a waveform, not an opinion.<p>Things that surprised us:<p>The frontier cluster spans roughly 45\u201372% (Claude Fable 5 on top at 71.7%).\nOpen-weight models that do well on coding benchmarks score 5\u201313% here.\nCost per task varies 15x between models with similar scores.\nHonest limitations: the task set is curated from a larger pack to separate the frontier models, and it&#x27;s growing. The pack is held out to keep it out of training data, so you can&#x27;t self-run it yet. If you want a model on the board, email us and we&#x27;ll run it.<p>Happy to go deep on the grading harness \u2014 ask us anything.","title":"Show HN: EEBench \u2013 AI agents design real circuits, graded by physics simulation","updated_at":"2026-07-03T17:50:23Z","url":"https://www.eebench.org/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"ben_w"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["cost","per","task","llm","benchmark"],"value":"&gt; Imagine being able to run GPT Sol at 1k tokens/second a year from now, at a 100x lower <em>cost</em> <em>per</em> token than now. Would that be useful?<p>Given the current rate of change, it would be hard to guess either way. By some measures the <em>cost</em> at fixed quality score goes down vastly faster than that:<p><pre><code>  A similar trend is evident in the <em>cost</em> of models scoring above 50% on GPQA, a substantially more challenging <em>benchmark</em> than MMLU. There, inference costs declined from $15 <em>per</em> million tokens in May 2024 to $0.12 <em>per</em> million tokens by December 2024 (Phi 4).\n</code></pre>\n- <a href=\"https://hai.stanford.edu/assets/files/hai_ai-index-report-2025_chapter1_final.pdf\" rel=\"nofollow\">https://hai.stanford.edu/assets/files/hai_ai-index-report-20...</a><p>15/0.12 -&gt; factor of 125 <em>cost</em> reduction in 7 months.<p>But that may well be an extreme case. To show how broad the range is, another quote from the same publication:<p><pre><code>  Depending on the <em>task</em>, <em>LLM</em> inference prices have fallen anywhere from 9 to 900 times <em>per</em> year.</code></pre>"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"OpenAI Jalape\u00f1o: Better than Nvidia Blackwell"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://newsletter.semianalysis.com/p/openai-jalapeno-better-than-nvidia"}},"_tags":["comment","author_ben_w","story_49434378"],"author":"ben_w","children":[49448586],"comment_text":"&gt; Imagine being able to run GPT Sol at 1k tokens&#x2F;second a year from now, at a 100x lower cost per token than now. Would that be useful?<p>Given the current rate of change, it would be hard to guess either way. By some measures the cost at fixed quality score goes down vastly faster than that:<p><pre><code>  A similar trend is evident in the cost of models scoring above 50% on GPQA, a substantially more challenging benchmark than MMLU. There, inference costs declined from $15 per million tokens in May 2024 to $0.12 per million tokens by December 2024 (Phi 4).\n</code></pre>\n- <a href=\"https:&#x2F;&#x2F;hai.stanford.edu&#x2F;assets&#x2F;files&#x2F;hai_ai-index-report-2025_chapter1_final.pdf\" rel=\"nofollow\">https:&#x2F;&#x2F;hai.stanford.edu&#x2F;assets&#x2F;files&#x2F;hai_ai-index-report-20...</a><p>15&#x2F;0.12 -&gt; factor of 125 cost reduction in 7 months.<p>But that may well be an extreme case. To show how broad the range is, another quote from the same publication:<p><pre><code>  Depending on the task, LLM inference prices have fallen anywhere from 9 to 900 times per year.</code></pre>","created_at":"2026-08-26T12:37:43Z","created_at_i":1787747863,"objectID":"49447966","parent_id":49444773,"story_id":49434378,"story_title":"OpenAI Jalape\u00f1o: Better than Nvidia Blackwell","story_url":"https://newsletter.semianalysis.com/p/openai-jalapeno-better-than-nvidia","updated_at":"2026-08-27T06:19:09Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"nipah"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["cost","per","task","llm","benchmark"],"value":"No, I think I saw the graphs on someone's channel, but maybe I misinterpreted the results. But to be fair, my point never depended on 100% of the participants being right 100% of the questions, there are innumerous factors that could affect your performance on those tests, including the pressure. The AI also had access to lenient conventions, so it should be &quot;fair&quot; in this sense.<p>Either way, there's something fishy about this presentation, it says:\n&quot;ARC-AGI-1 WAS EASILY BRUTE-FORCIBLE&quot;, but when o3 initially &quot;solved&quot; most of it the co-founder or ARC-PRIZE said:\n&quot;Despite the significant <em>cost</em> <em>per</em> <em>task</em>, these numbers aren't just the result of applying brute force compute to the <em>benchmark</em>. OpenAI's new o3 model represents a significant leap forward in AI's ability to adapt to novel tasks. This is not merely incremental improvement, but a genuine breakthrough, marking a qualitative shift in AI capabilities compared to the prior limitations of <em>LLMs</em>. o3 is a system capable of adapting to tasks it has never encountered before, arguably approaching human-level performance in the ARC-AGI domain.&quot;, he  was saying confidently that it would not be a result of brute-forcing the problems.\nAnd it was not the first time,\n&quot;ARC-AGI-1 consists of 800 puzzle-like tasks, designed as grid-based visual reasoning problems. These tasks, trivial for humans but challenging for machines, typically provide only a small number of example input-output pairs (usually around three). This requires the test taker (human or AI) to deduce underlying rules through abstraction, inference, and prior knowledge rather than brute-force or extensive training.&quot;<p>Now they are saying ARC-AGI-2 is not bruteforcible, what is happening there? They didn't provided any reasoning for why one was bruteforcible and the other not, nor how they are so sure about that.\nThey &quot;recognized&quot; that it could be brute-forced before, but in a way less expressive manner, by explicitly stating it would need &quot;unlimited resources and time&quot; to solve. And they are using the non-bruteforceability in this presentation as a point for it.<p>---\nAlso, I mentioned mammals because those problems are of an order that mammals and even other animals would need to solve in reality for a diversity of cases. I'm not saying that they would literally be able to take the test and solve it, nor to understand this is a test, but that they would need to solve problems of similar nature in reality. Naturally this point has it's own limits, but it's not easily discarded as you tried to do."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"OpenAI o3-pro"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://help.openai.com/en/articles/9624314-model-release-notes"}},"_tags":["comment","author_nipah","story_44240999"],"author":"nipah","children":[44259128],"comment_text":"No, I think I saw the graphs on someone&#x27;s channel, but maybe I misinterpreted the results. But to be fair, my point never depended on 100% of the participants being right 100% of the questions, there are innumerous factors that could affect your performance on those tests, including the pressure. The AI also had access to lenient conventions, so it should be &quot;fair&quot; in this sense.<p>Either way, there&#x27;s something fishy about this presentation, it says:\n&quot;ARC-AGI-1 WAS EASILY BRUTE-FORCIBLE&quot;, but when o3 initially &quot;solved&quot; most of it the co-founder or ARC-PRIZE said:\n&quot;Despite the significant cost per task, these numbers aren&#x27;t just the result of applying brute force compute to the benchmark. OpenAI&#x27;s new o3 model represents a significant leap forward in AI&#x27;s ability to adapt to novel tasks. This is not merely incremental improvement, but a genuine breakthrough, marking a qualitative shift in AI capabilities compared to the prior limitations of LLMs. o3 is a system capable of adapting to tasks it has never encountered before, arguably approaching human-level performance in the ARC-AGI domain.&quot;, he  was saying confidently that it would not be a result of brute-forcing the problems.\nAnd it was not the first time,\n&quot;ARC-AGI-1 consists of 800 puzzle-like tasks, designed as grid-based visual reasoning problems. These tasks, trivial for humans but challenging for machines, typically provide only a small number of example input-output pairs (usually around three). This requires the test taker (human or AI) to deduce underlying rules through abstraction, inference, and prior knowledge rather than brute-force or extensive training.&quot;<p>Now they are saying ARC-AGI-2 is not bruteforcible, what is happening there? They didn&#x27;t provided any reasoning for why one was bruteforcible and the other not, nor how they are so sure about that.\nThey &quot;recognized&quot; that it could be brute-forced before, but in a way less expressive manner, by explicitly stating it would need &quot;unlimited resources and time&quot; to solve. And they are using the non-bruteforceability in this presentation as a point for it.<p>---\nAlso, I mentioned mammals because those problems are of an order that mammals and even other animals would need to solve in reality for a diversity of cases. I&#x27;m not saying that they would literally be able to take the test and solve it, nor to understand this is a test, but that they would need to solve problems of similar nature in reality. Naturally this point has it&#x27;s own limits, but it&#x27;s not easily discarded as you tried to do.","created_at":"2025-06-12T14:10:08Z","created_at_i":1749737408,"objectID":"44258024","parent_id":44255902,"story_id":44240999,"story_title":"OpenAI o3-pro","story_url":"https://help.openai.com/en/articles/9624314-model-release-notes","updated_at":"2025-06-12T15:50:07Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"ojosilva"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["cost","per","task","llm","benchmark"],"value":"To @simonw and all the coding agent and <em>LLM</em> <em>benchmark</em>ers out there: please, always publish the elapsed time for the <em>task</em> to complete successfully! I know this was just a &quot;it works straight in claude.ai&quot; post, but still, nowhere in the transcript there's a timestamp of any kind. Durations seem to be COMPLETELY missing from the <em>LLM</em> coding leaderboards everywhere [1] [2] [3]<p>There's a huge difference in time-to-completion from model to model, platform to platform, and if, like me, you are into trial-and-error, rebooting the session over and over to get the prompt right or &quot;one-shot&quot;, it's important how reasoning efforts, provider's tokens/s, coding agent tooling efficiency, <em>costs</em> and overall model intelligence play together to get the <em>task</em> done. Same thing applies to the coding agent, when applicable.<p>Grok Code Fast and Cerebras Code (qwen) are 2 examples of how models can be very competitive without being the top-notch intelligence. Running inference at 10x speed really allows for a leaner experience in AI-assisted coding and more <em>task</em> completion <em>per</em> day than a sluggish, but more correct AI. Darn, I feel like a corporate butt-head right now.<p>1. <a href=\"https://www.swebench.com/\" rel=\"nofollow\">https://www.swebench.com/</a><p>2. <a href=\"https://www.tbench.ai/leaderboard\" rel=\"nofollow\">https://www.tbench.ai/leaderboard</a><p>3. <a href=\"https://gosuevals.com/agents.html\" rel=\"nofollow\">https://gosuevals.com/agents.html</a>"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Claude Sonnet 4.5"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://www.anthropic.com/news/claude-sonnet-4-5"}},"_tags":["comment","author_ojosilva","story_45415962"],"author":"ojosilva","children":[45419176,45419366,45421807],"comment_text":"To @simonw and all the coding agent and LLM benchmarkers out there: please, always publish the elapsed time for the task to complete successfully! I know this was just a &quot;it works straight in claude.ai&quot; post, but still, nowhere in the transcript there&#x27;s a timestamp of any kind. Durations seem to be COMPLETELY missing from the LLM coding leaderboards everywhere [1] [2] [3]<p>There&#x27;s a huge difference in time-to-completion from model to model, platform to platform, and if, like me, you are into trial-and-error, rebooting the session over and over to get the prompt right or &quot;one-shot&quot;, it&#x27;s important how reasoning efforts, provider&#x27;s tokens&#x2F;s, coding agent tooling efficiency, costs and overall model intelligence play together to get the task done. Same thing applies to the coding agent, when applicable.<p>Grok Code Fast and Cerebras Code (qwen) are 2 examples of how models can be very competitive without being the top-notch intelligence. Running inference at 10x speed really allows for a leaner experience in AI-assisted coding and more task completion per day than a sluggish, but more correct AI. Darn, I feel like a corporate butt-head right now.<p>1. <a href=\"https:&#x2F;&#x2F;www.swebench.com&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;www.swebench.com&#x2F;</a><p>2. <a href=\"https:&#x2F;&#x2F;www.tbench.ai&#x2F;leaderboard\" rel=\"nofollow\">https:&#x2F;&#x2F;www.tbench.ai&#x2F;leaderboard</a><p>3. <a href=\"https:&#x2F;&#x2F;gosuevals.com&#x2F;agents.html\" rel=\"nofollow\">https:&#x2F;&#x2F;gosuevals.com&#x2F;agents.html</a>","created_at":"2025-09-29T21:29:38Z","created_at_i":1759181378,"objectID":"45418990","parent_id":45415962,"story_id":45415962,"story_title":"Claude Sonnet 4.5","story_url":"https://www.anthropic.com/news/claude-sonnet-4-5","updated_at":"2026-03-05T22:46:36Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"rsaha7"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["cost","per","task","llm","benchmark"],"value":"Hi HN community, I have been working on <em>benchmark</em>ing publicly available <em>LLMs</em> these past couple of weeks. More precisely, I am interested on the finetuning piece since a lot of businesses are starting to entertain the idea of self-hosting <em>LLMs</em> trained on their proprietary data rather than relying on third party APIs.<p>To this point, I am tracking the following 4 pillars of evaluation that businesses are typically look into:\n- Performance\n- Time to train an <em>LLM</em>\n- <em>Cost</em> to train an <em>LLM</em>\n- Inference (throughput / latency / <em>cost</em> <em>per</em> token)<p>For each <em>LLM</em>, my aim is to <em>benchmark</em> them for popular <em>tasks</em>, i.e., classification and summarization. Moreover, I would like to compare them against each other.<p>So far, I have <em>benchmark</em>ed Flan-T5-Large, Falcon-7B and RedPajama and have found them to be very efficient in low-data situations, i.e., when there are very few annotated samples. Llama2-7B/13B and Writer\u2019s Palmyra are in the pipeline.<p>But there\u2019s so many <em>LLMs</em> out there! In case this work interests you, would be great to join forces.<p>GitHub repo attached \u2014 feedback is always welcome :)<p>Happy hacking!"},"title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["llm"],"value":"Show HN: finetune <em>LLMs</em> via the Finetuning Hub"},"url":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["llm"],"value":"https://github.com/georgian-io/<em>LLM</em>-Finetuning-Hub"}},"_tags":["story","author_rsaha7","story_37381296","show_hn"],"author":"rsaha7","children":[37382956,37383225,37384606,37385295,37413547],"created_at":"2023-09-04T15:16:42Z","created_at_i":1693840602,"num_comments":17,"objectID":"37381296","points":80,"story_id":37381296,"story_text":"Hi HN community, I have been working on benchmarking publicly available LLMs these past couple of weeks. More precisely, I am interested on the finetuning piece since a lot of businesses are starting to entertain the idea of self-hosting LLMs trained on their proprietary data rather than relying on third party APIs.<p>To this point, I am tracking the following 4 pillars of evaluation that businesses are typically look into:\n- Performance\n- Time to train an LLM\n- Cost to train an LLM\n- Inference (throughput &#x2F; latency &#x2F; cost per token)<p>For each LLM, my aim is to benchmark them for popular tasks, i.e., classification and summarization. Moreover, I would like to compare them against each other.<p>So far, I have benchmarked Flan-T5-Large, Falcon-7B and RedPajama and have found them to be very efficient in low-data situations, i.e., when there are very few annotated samples. Llama2-7B&#x2F;13B and Writer\u2019s Palmyra are in the pipeline.<p>But there\u2019s so many LLMs out there! In case this work interests you, would be great to join forces.<p>GitHub repo attached \u2014 feedback is always welcome :)<p>Happy hacking!","title":"Show HN: finetune LLMs via the Finetuning Hub","updated_at":"2026-05-13T06:25:33Z","url":"https://github.com/georgian-io/LLM-Finetuning-Hub"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"fazlerocks"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["cost","per","task","llm","benchmark"],"value":"We're running <em>LLMs</em> in production for content generation, customer support, and code review assistance. Been trying to build a proper evaluation pipeline for months but every tool we've tested has significant limitations.<p>What we've evaluated:<p>- OpenAI's Evals framework: Works well for <em>benchmark</em>ing but challenging for custom use cases. Configuration through YAML files can be complex and extending functionality requires diving deep into their codebase. Primarily designed for batch processing rather than real-time monitoring.<p>- LangSmith: Strong tracing capabilities but eval features feel secondary to their observability focus. Pricing starts at $0.50 <em>per</em> 1k traces after the free tier, which adds up quickly with high volume. UI can be slow with larger datasets.<p>- Weights &amp; Biases: Powerful platform but designed primarily for traditional ML experiment tracking. Setup is complex and requires significant ML expertise. Our product team struggles to use it effectively.<p>- Humanloop: Clean interface focused on prompt versioning with basic evaluation capabilities. Limited eval types available and pricing is steep for the feature set.<p>- Braintrust: Interesting approach to evaluation but feels like an early-stage product. Documentation is sparse and integration options are limited.<p>What we actually need:\n- Real-time eval monitoring (not just batch)\n- Custom eval functions that don't require PhD-level setup\n- Human-in-the-loop workflows for subjective <em>tasks</em>\n- <em>Cost</em> tracking <em>per</em> model/prompt\n- Integration with our existing observability stack\n- Something our product team can actually use<p>Current solution:<p>Custom scripts + monitoring dashboards for basic metrics. Weekly manual reviews in spreadsheets. It works but doesn't scale and we miss edge cases.<p>Has anyone found tools that handle production <em>LLM</em> evaluation well? Are we expecting too much or is the tooling genuinely immature? Especially interested in hearing from teams without dedicated ML engineers."},"title":{"matchLevel":"none","matchedWords":[],"value":"Ask HN: What tools are you using for AI evals? Everything feels half-baked"}},"_tags":["story","author_fazlerocks","story_44194187","ask_hn"],"author":"fazlerocks","children":[44194536,44194659,44194728,44200047,44246104],"created_at":"2025-06-05T18:11:53Z","created_at_i":1749147113,"num_comments":3,"objectID":"44194187","points":6,"story_id":44194187,"story_text":"We&#x27;re running LLMs in production for content generation, customer support, and code review assistance. Been trying to build a proper evaluation pipeline for months but every tool we&#x27;ve tested has significant limitations.<p>What we&#x27;ve evaluated:<p>- OpenAI&#x27;s Evals framework: Works well for benchmarking but challenging for custom use cases. Configuration through YAML files can be complex and extending functionality requires diving deep into their codebase. Primarily designed for batch processing rather than real-time monitoring.<p>- LangSmith: Strong tracing capabilities but eval features feel secondary to their observability focus. Pricing starts at $0.50 per 1k traces after the free tier, which adds up quickly with high volume. UI can be slow with larger datasets.<p>- Weights &amp; Biases: Powerful platform but designed primarily for traditional ML experiment tracking. Setup is complex and requires significant ML expertise. Our product team struggles to use it effectively.<p>- Humanloop: Clean interface focused on prompt versioning with basic evaluation capabilities. Limited eval types available and pricing is steep for the feature set.<p>- Braintrust: Interesting approach to evaluation but feels like an early-stage product. Documentation is sparse and integration options are limited.<p>What we actually need:\n- Real-time eval monitoring (not just batch)\n- Custom eval functions that don&#x27;t require PhD-level setup\n- Human-in-the-loop workflows for subjective tasks\n- Cost tracking per model&#x2F;prompt\n- Integration with our existing observability stack\n- Something our product team can actually use<p>Current solution:<p>Custom scripts + monitoring dashboards for basic metrics. Weekly manual reviews in spreadsheets. It works but doesn&#x27;t scale and we miss edge cases.<p>Has anyone found tools that handle production LLM evaluation well? Are we expecting too much or is the tooling genuinely immature? Especially interested in hearing from teams without dedicated ML engineers.","title":"Ask HN: What tools are you using for AI evals? Everything feels half-baked","updated_at":"2025-10-08T20:39:40Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"ayushranjan99"},"story_text":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["cost","per","task","llm"],"value":"I\u2019ve been experimenting with running small LLMs directly on mobile hardware (low-range Android devices), without relying on cloud inference. This is a summary of what worked, what didn\u2019t, and why.<p>Cloud-based <em>LLM</em> APIs are convenient, but come with:<p>-latency from network round-trips\n-unpredictable API <em>costs</em>\n-privacy concerns (content leaving device)\n-the need for connectivity<p>For simple <em>tasks</em> like news summarization, small models seem \u201cgood enough,\u201d so I tested whether a ~270M parameter model gemma3-270m could run entirely on-device.<p>Model - Gemma3-270M INT8 Quantized\nRuntime - Cactus SDK (Android NPU/GPU acceleration)\nApp Framework - Flutter\nDevice - Mediatek 7300 with 8GB RAM<p>Architecture\n- User shares a URL to the app (Android share sheet).\n- App fetches article HTML \u2192 extracts readable text.\n- Local model generates a summary.\n- device TTS reads the summary.\nEverything runs offline except the initial page fetch.<p>Performace\n- ~450\u2013900ms Latency for a short summary (100\u2013200 tokens).\n- On devices without NPU acceleration, CPU-only inference takes 2\u20133\u00d7 longer.\n- Peak RAM: ~350\u2013450MB<p>Limitation\n-Quality is noticeably worse than GPT-5 for complex articles.\n-Long-form summarization (&gt;1k words) gets inconsistent.\n-Web scraping is fragile for JS-heavy or paywalled sites.\n-Some low-end phones throttle CPU/GPU aggressively.<p>| Metric  | Local (Gemma 270M)   | GPT-4o Cloud         |\n| ------- | -------------------- | -------------------- |\n| Latency | 0.5\u20131.5s             | 0.7\u20131.5s + network   |\n| <em>Cost</em>    | 0                    | API <em>cost</em> <em>per</em> request |\n| Privacy | Text stays on device | Sent over network    |\n| Quality | Medium               | High                 |<p>Github - https://github.com/ayusrjn/briefly<p>Running small LLMs on-device is viable for narrow <em>tasks</em> like summarization. For more complex reasoning <em>tasks</em>, cloud models still outperform by a large margin, but the \u201clocal-first\u201d approach seems promising for privacy-sensitive or offline-first applications.\nCactus SDK does a pretty good job for handling the model and accelarations.<p>Happy to answer Questions :)"},"title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["llm","benchmark"],"value":"Running a 270M <em>LLM</em> on Android (architecture and <em>benchmarks</em>)"}},"_tags":["story","author_ayushranjan99","story_46018506","ask_hn"],"author":"ayushranjan99","created_at":"2025-11-22T21:41:56Z","created_at_i":1763847716,"num_comments":0,"objectID":"46018506","points":2,"story_id":46018506,"story_text":"I\u2019ve been experimenting with running small LLMs directly on mobile hardware (low-range Android devices), without relying on cloud inference. This is a summary of what worked, what didn\u2019t, and why.<p>Cloud-based LLM APIs are convenient, but come with:<p>-latency from network round-trips\n-unpredictable API costs\n-privacy concerns (content leaving device)\n-the need for connectivity<p>For simple tasks like news summarization, small models seem \u201cgood enough,\u201d so I tested whether a ~270M parameter model gemma3-270m could run entirely on-device.<p>Model - Gemma3-270M INT8 Quantized\nRuntime - Cactus SDK (Android NPU&#x2F;GPU acceleration)\nApp Framework - Flutter\nDevice - Mediatek 7300 with 8GB RAM<p>Architecture\n- User shares a URL to the app (Android share sheet).\n- App fetches article HTML \u2192 extracts readable text.\n- Local model generates a summary.\n- device TTS reads the summary.\nEverything runs offline except the initial page fetch.<p>Performace\n- ~450\u2013900ms Latency for a short summary (100\u2013200 tokens).\n- On devices without NPU acceleration, CPU-only inference takes 2\u20133\u00d7 longer.\n- Peak RAM: ~350\u2013450MB<p>Limitation\n-Quality is noticeably worse than GPT-5 for complex articles.\n-Long-form summarization (&gt;1k words) gets inconsistent.\n-Web scraping is fragile for JS-heavy or paywalled sites.\n-Some low-end phones throttle CPU&#x2F;GPU aggressively.<p>| Metric  | Local (Gemma 270M)   | GPT-4o Cloud         |\n| ------- | -------------------- | -------------------- |\n| Latency | 0.5\u20131.5s             | 0.7\u20131.5s + network   |\n| Cost    | 0                    | API cost per request |\n| Privacy | Text stays on device | Sent over network    |\n| Quality | Medium               | High                 |<p>Github - https:&#x2F;&#x2F;github.com&#x2F;ayusrjn&#x2F;briefly<p>Running small LLMs on-device is viable for narrow tasks like summarization. For more complex reasoning tasks, cloud models still outperform by a large margin, but the \u201clocal-first\u201d approach seems promising for privacy-sensitive or offline-first applications.\nCactus SDK does a pretty good job for handling the model and accelarations.<p>Happy to answer Questions :)","title":"Running a 270M LLM on Android (architecture and benchmarks)","updated_at":"2026-03-05T23:05:26Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"TimoKerr"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["cost","per","task","llm","benchmark"],"value":"Hi HN,<p>TLDR: Cheap (and sometimes old) models perform on par, or better than flagship models on standard OCR <em>tasks</em>, at a fraction of the <em>cost</em>. \nThis conclusion comes from a <em>benchmark</em> we ran on 18 models and over 7k+ <em>LLM</em> calls. Leaderboard and <em>benchmark</em> repo completely open-source.<p>Too many teams are either stuck in legacy OCR pipelines, or are overpaying badly for <em>LLM</em> calls by defaulting to the newest/ biggest model.<p>So we investigated the topic and open-sourced everything, including a free tool to check your own documents.<p>We ran 18 models from OpenAI, Anthropic, Google, and Mistral on 42 real-world documents (invoices, receipts, bills of lading, transport orders). \nEach model ran 10 times <em>per</em> document to measure reliability, not just one-shot accuracy; 7,560 API calls total.<p>The finding: for standard document extraction, mid-tier and older models match or beat state-of-the-art, at a fraction of the <em>cost</em>. \nIn some cases the <em>cost</em> difference is multiple orders of magnitude for equivalent accuracy.<p>We also track pass^n (how reliability degrades over repeated runs, see tau-bench), <em>cost</em>-<em>per</em>-success (not just <em>cost</em>-<em>per</em>-token), \nand critical field accuracy. Full methodology and dataset are open source.<p>Leaderboard: &lt;<a href=\"https://www.arbitrhq.ai/leaderboards/\" rel=\"nofollow\">https://www.arbitrhq.ai/leaderboards/</a>&gt;<p>Dataset + framework (GitHub): &lt;<a href=\"https://github.com/ArbitrHq/ocr-mini-bench\" rel=\"nofollow\">https://github.com/ArbitrHq/ocr-mini-bench</a>&gt;<p>Or test your own documents for free: &lt;<a href=\"https://app.arbitrhq.ai/benchmark-free\" rel=\"nofollow\">https://app.arbitrhq.ai/<em>benchmark</em>-free</a>&gt;<p>Built by two founders in Antwerp. \nVery curious if other people have similar conclusions or if you've seen specific edge cases where the flagships still justify their price tag?"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: We benchmarked 18 LLMs on OCR (7K+ calls) \u2013 cheaper models win"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://www.arbitrhq.ai/leaderboards/"}},"_tags":["comment","author_TimoKerr","story_47860393"],"author":"TimoKerr","comment_text":"Hi HN,<p>TLDR: Cheap (and sometimes old) models perform on par, or better than flagship models on standard OCR tasks, at a fraction of the cost. \nThis conclusion comes from a benchmark we ran on 18 models and over 7k+ LLM calls. Leaderboard and benchmark repo completely open-source.<p>Too many teams are either stuck in legacy OCR pipelines, or are overpaying badly for LLM calls by defaulting to the newest&#x2F; biggest model.<p>So we investigated the topic and open-sourced everything, including a free tool to check your own documents.<p>We ran 18 models from OpenAI, Anthropic, Google, and Mistral on 42 real-world documents (invoices, receipts, bills of lading, transport orders). \nEach model ran 10 times per document to measure reliability, not just one-shot accuracy; 7,560 API calls total.<p>The finding: for standard document extraction, mid-tier and older models match or beat state-of-the-art, at a fraction of the cost. \nIn some cases the cost difference is multiple orders of magnitude for equivalent accuracy.<p>We also track pass^n (how reliability degrades over repeated runs, see tau-bench), cost-per-success (not just cost-per-token), \nand critical field accuracy. Full methodology and dataset are open source.<p>Leaderboard: &lt;<a href=\"https:&#x2F;&#x2F;www.arbitrhq.ai&#x2F;leaderboards&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;www.arbitrhq.ai&#x2F;leaderboards&#x2F;</a>&gt;<p>Dataset + framework (GitHub): &lt;<a href=\"https:&#x2F;&#x2F;github.com&#x2F;ArbitrHq&#x2F;ocr-mini-bench\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;ArbitrHq&#x2F;ocr-mini-bench</a>&gt;<p>Or test your own documents for free: &lt;<a href=\"https:&#x2F;&#x2F;app.arbitrhq.ai&#x2F;benchmark-free\" rel=\"nofollow\">https:&#x2F;&#x2F;app.arbitrhq.ai&#x2F;benchmark-free</a>&gt;<p>Built by two founders in Antwerp. \nVery curious if other people have similar conclusions or if you&#x27;ve seen specific edge cases where the flagships still justify their price tag?","created_at":"2026-04-22T07:48:34Z","created_at_i":1776844114,"objectID":"47860396","parent_id":47860393,"story_id":47860393,"story_title":"Show HN: We benchmarked 18 LLMs on OCR (7K+ calls) \u2013 cheaper models win","story_url":"https://www.arbitrhq.ai/leaderboards/","updated_at":"2026-04-22T07:50:52Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"HawtAds"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["cost","per","task","llm","benchmark"],"value":"Okay, here's the tl;dr:<p>Attention based neural network architectures (on which the majority of <em>LLMs</em> are built) has a unit economic <em>cost</em> that scales (roughly) n^2 i.e. quadratic (for both memory and compute). In other words, the longer the context window, the more expensive it is for the upstream provider. That's one <em>cost</em>.<p>The second <em>cost</em> is that you have to resend the entire context every time you send a new message. So the context is basically (where a, b, and c are messages): first context: a, second context window: a-&gt;b, third context window: a-&gt;b-&gt;c. It's a mostly stateless (there are some short term caching mechanisms, YMMV based on provider, it's why &quot;cached&quot; messages, especially system prompts are cheaper) process from the point of view of the developer, the state i.e. context window string is managed by the end user application (in other words, the coding agent, the IDE, the ChatGPT UI client etc.)<p>The <em>per</em> token <em>cost</em> is an <i>amortized</i> (averaged) <em>cost</em> of memory+compute, the actual <em>cost</em> is mostly quadratic with respect to each marginal token. The longer the context window the more expensive things are. \nBecause of the above, AI agent providers (especially those that charge flat fee subscription plans) are incentivized to keep costs low by limiting the maximum context window size.<p>(And if you think about it carefully, your AI API costs are a quadratic <em>cost</em> curve projected into a linear line (flat fee <em>per</em> token, so the model hosting provider in some cases may make more profit if users send in shorter contexts, versus if they constantly saturate the window. YMMV of course, but it's a race to the bottom right now for <em>LLM</em> unit economics)<p>They do this by interrupting a <em>task</em> halfway through and generating a &quot;summary&quot; of the <em>task</em> progress, then they prompt the <em>LLM</em> again with a fresh prompt and the &quot;summary&quot; so far and the <em>LLM</em> will restart the <em>task</em> from where it left of. Of course text is a poor representation of the <em>LLM</em>'s internal state but it's the best option so far for AI application to keep costs low.<p>Another thing to keep in mind is that <em>LLMs</em> have poorer performance the larger the input size. This is due to a variety of factors (mostly because you don't have enough training data to saturate the massive context window sizes I think).<p>The general graph for <em>LLM</em> context performance looks something like this:\n<a href=\"https://cobusgreyling.medium.com/llm-context-rot-28a6d0399655\" rel=\"nofollow\">https://cobusgreyling.medium.com/<em>llm</em>-context-rot-28a6d039965...</a>\n<a href=\"https://research.trychroma.com/context-rot\" rel=\"nofollow\">https://research.trychroma.com/context-rot</a><p>There are a bunch of tests and <em>benchmarks</em> (commonly referred to as &quot;needle in a haystack&quot;) to improve the <em>LLM</em> performance at large context window sizes, but it's still an open area of research.<p><a href=\"https://cloud.google.com/blog/products/ai-machine-learning/the-needle-in-the-haystack-test-and-how-gemini-pro-solves-it\" rel=\"nofollow\">https://cloud.google.com/blog/products/ai-machine-learning/t...</a><p>The thing is, <i>generally speaking</i>, you will get a slightly better performance if you can squeeze all your code and problem into the context window, because the <em>LLM</em> can get a &quot;whole picture&quot; view of your codebase/problem, instead of a bunch of broken telephone summaries every dozen of thousands of tokens. Take this with a grain of salt as the field is changing rapidly so it might not be valid in a month or two.<p>Keep in mind that if the problem you are solving requires you to saturate the entire context window of the <em>LLM</em>, a <i>single</i> request can <em>cost</em> you dollars. And if you are using 1M+ context window model like gemini, you can rack up costs fairly rapidly."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Opus 4.5 is not the normal AI agent experience that I have had thus far"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://burkeholland.github.io/posts/opus-4-5-change-everything/"}},"_tags":["comment","author_HawtAds","story_46515696"],"author":"HawtAds","children":[46553225],"comment_text":"Okay, here&#x27;s the tl;dr:<p>Attention based neural network architectures (on which the majority of LLMs are built) has a unit economic cost that scales (roughly) n^2 i.e. quadratic (for both memory and compute). In other words, the longer the context window, the more expensive it is for the upstream provider. That&#x27;s one cost.<p>The second cost is that you have to resend the entire context every time you send a new message. So the context is basically (where a, b, and c are messages): first context: a, second context window: a-&gt;b, third context window: a-&gt;b-&gt;c. It&#x27;s a mostly stateless (there are some short term caching mechanisms, YMMV based on provider, it&#x27;s why &quot;cached&quot; messages, especially system prompts are cheaper) process from the point of view of the developer, the state i.e. context window string is managed by the end user application (in other words, the coding agent, the IDE, the ChatGPT UI client etc.)<p>The per token cost is an <i>amortized</i> (averaged) cost of memory+compute, the actual cost is mostly quadratic with respect to each marginal token. The longer the context window the more expensive things are. \nBecause of the above, AI agent providers (especially those that charge flat fee subscription plans) are incentivized to keep costs low by limiting the maximum context window size.<p>(And if you think about it carefully, your AI API costs are a quadratic cost curve projected into a linear line (flat fee per token, so the model hosting provider in some cases may make more profit if users send in shorter contexts, versus if they constantly saturate the window. YMMV of course, but it&#x27;s a race to the bottom right now for LLM unit economics)<p>They do this by interrupting a task halfway through and generating a &quot;summary&quot; of the task progress, then they prompt the LLM again with a fresh prompt and the &quot;summary&quot; so far and the LLM will restart the task from where it left of. Of course text is a poor representation of the LLM&#x27;s internal state but it&#x27;s the best option so far for AI application to keep costs low.<p>Another thing to keep in mind is that LLMs have poorer performance the larger the input size. This is due to a variety of factors (mostly because you don&#x27;t have enough training data to saturate the massive context window sizes I think).<p>The general graph for LLM context performance looks something like this:\n<a href=\"https:&#x2F;&#x2F;cobusgreyling.medium.com&#x2F;llm-context-rot-28a6d0399655\" rel=\"nofollow\">https:&#x2F;&#x2F;cobusgreyling.medium.com&#x2F;llm-context-rot-28a6d039965...</a>\n<a href=\"https:&#x2F;&#x2F;research.trychroma.com&#x2F;context-rot\" rel=\"nofollow\">https:&#x2F;&#x2F;research.trychroma.com&#x2F;context-rot</a><p>There are a bunch of tests and benchmarks (commonly referred to as &quot;needle in a haystack&quot;) to improve the LLM performance at large context window sizes, but it&#x27;s still an open area of research.<p><a href=\"https:&#x2F;&#x2F;cloud.google.com&#x2F;blog&#x2F;products&#x2F;ai-machine-learning&#x2F;the-needle-in-the-haystack-test-and-how-gemini-pro-solves-it\" rel=\"nofollow\">https:&#x2F;&#x2F;cloud.google.com&#x2F;blog&#x2F;products&#x2F;ai-machine-learning&#x2F;t...</a><p>The thing is, <i>generally speaking</i>, you will get a slightly better performance if you can squeeze all your code and problem into the context window, because the LLM can get a &quot;whole picture&quot; view of your codebase&#x2F;problem, instead of a bunch of broken telephone summaries every dozen of thousands of tokens. Take this with a grain of salt as the field is changing rapidly so it might not be valid in a month or two.<p>Keep in mind that if the problem you are solving requires you to saturate the entire context window of the LLM, a <i>single</i> request can cost you dollars. And if you are using 1M+ context window model like gemini, you can rack up costs fairly rapidly.","created_at":"2026-01-08T19:30:46Z","created_at_i":1767900646,"objectID":"46545344","parent_id":46543380,"story_id":46515696,"story_title":"Opus 4.5 is not the normal AI agent experience that I have had thus far","story_url":"https://burkeholland.github.io/posts/opus-4-5-change-everything/","updated_at":"2026-07-31T17:58:05Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"stanleycyang"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["cost","per","task","llm","benchmark"],"value":"Hey HN,<p>I built yardstiq because I got tired of the copy-paste workflow for comparing <em>LLM</em> responses when developing apps. Every time I wanted to see how Claude vs GPT vs Gemini handled the same prompt, I'd open three tabs, paste the same thing, and try to eyeball the differences. It's 2026 and we have 40+ models worth considering \u2014 that doesn't scale.<p>yardstiq is a CLI tool that sends one prompt to multiple models simultaneously and streams the responses side-by-side in your terminal. It also tracks performance metrics (time to first token, tokens/sec, <em>cost</em>) and optionally runs an AI judge to score the outputs.<p>```\nnpx yardstiq &quot;Explain quicksort in 3 sentences&quot; -m claude-sonnet -m gpt-4o\n```<p>What it does:<p>- Streams responses from multiple models in parallel, rendered in columns\n- Shows TTFT, throughput (tok/s), token counts, and <em>cost</em> <em>per</em> request\n- AI judge mode: have a model evaluate and score the responses\n- Export to JSON, Markdown, or self-contained HTML reports\n- Run YAML-defined <em>benchmark</em> suites across models with aggregate scoring\n- Works with Ollama for local model comparisons (zero API <em>cost</em>)\n- Supports 40+ models via direct provider keys or Vercel AI Gateway<p>I built this mostly for my own workflow \u2014 picking models for different <em>tasks</em>, testing prompt variations, and running quick benchmarks without setting up a whole evaluation framework. It's not trying to replace serious eval platforms, just make the &quot;which model is better for X?&quot; question answerable in 10 seconds.<p>MIT licensed, written in TypeScript: <a href=\"https://github.com/stanleycyang/yardstiq\" rel=\"nofollow\">https://github.com/stanleycyang/yardstiq</a><p>Happy to answer questions about the architecture or benchmarking approach."},"title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["llm"],"value":"Show HN: Yardstiq \u2013 Compare <em>LLM</em> outputs side-by-side in your terminal"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://www.yardstiq.sh"}},"_tags":["story","author_stanleycyang","story_47235222","show_hn"],"author":"stanleycyang","children":[47236630,47246254,47265845],"created_at":"2026-03-03T16:54:46Z","created_at_i":1772556886,"num_comments":0,"objectID":"47235222","points":2,"story_id":47235222,"story_text":"Hey HN,<p>I built yardstiq because I got tired of the copy-paste workflow for comparing LLM responses when developing apps. Every time I wanted to see how Claude vs GPT vs Gemini handled the same prompt, I&#x27;d open three tabs, paste the same thing, and try to eyeball the differences. It&#x27;s 2026 and we have 40+ models worth considering \u2014 that doesn&#x27;t scale.<p>yardstiq is a CLI tool that sends one prompt to multiple models simultaneously and streams the responses side-by-side in your terminal. It also tracks performance metrics (time to first token, tokens&#x2F;sec, cost) and optionally runs an AI judge to score the outputs.<p>```\nnpx yardstiq &quot;Explain quicksort in 3 sentences&quot; -m claude-sonnet -m gpt-4o\n```<p>What it does:<p>- Streams responses from multiple models in parallel, rendered in columns\n- Shows TTFT, throughput (tok&#x2F;s), token counts, and cost per request\n- AI judge mode: have a model evaluate and score the responses\n- Export to JSON, Markdown, or self-contained HTML reports\n- Run YAML-defined benchmark suites across models with aggregate scoring\n- Works with Ollama for local model comparisons (zero API cost)\n- Supports 40+ models via direct provider keys or Vercel AI Gateway<p>I built this mostly for my own workflow \u2014 picking models for different tasks, testing prompt variations, and running quick benchmarks without setting up a whole evaluation framework. It&#x27;s not trying to replace serious eval platforms, just make the &quot;which model is better for X?&quot; question answerable in 10 seconds.<p>MIT licensed, written in TypeScript: <a href=\"https:&#x2F;&#x2F;github.com&#x2F;stanleycyang&#x2F;yardstiq\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;stanleycyang&#x2F;yardstiq</a><p>Happy to answer questions about the architecture or benchmarking approach.","title":"Show HN: Yardstiq \u2013 Compare LLM outputs side-by-side in your terminal","updated_at":"2026-03-08T20:27:38Z","url":"https://www.yardstiq.sh"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"nicola_alessi"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["cost","per","task","llm","benchmark"],"value":"I ran a controlled <em>benchmark</em> on AI coding agents (42 runs, FastAPI, Claude Sonnet 4.6) and found something that broke my mental model of <em>LLM</em> costs.\nThe setup: I built an MCP server that pre-indexes a codebase into a dependency graph and serves pre-ranked context to the agent in a single call, instead of letting the agent explore files on its own.\nThe expected result: less input context \u2192 lower <em>cost</em>. Straightforward.\nThe actual result: total tokens processed went UP 20% (23.4M vs 19.6M) while total <em>cost</em> went DOWN 58% ($6.89 vs $16.29).\nThe explanation is in how Anthropic prices tokens. There are three pricing tiers:<p>Output tokens: most expensive (3-5x input price)\nInput tokens (cache miss): full price\nInput tokens (cache hit): 90% discount<p>The agent with pre-indexed context processes more total tokens because the structured context payload is injected every turn. But the token MIX shifts dramatically:\nOutput tokens:     10,588 \u2192 3,965  (-63%)\nCache read rate:   93.8% \u2192 95.3%\nCache creation:    6.1% \u2192 4.6%\nOutput tokens dominate the <em>cost</em> equation. When the agent receives 40K tokens of unfiltered context, it generates verbose orientation narration (&quot;let me look at this file... I can see that...&quot;). When it receives 8K tokens of graph-ranked context, it skips straight to the answer. 504 output tokens <em>per</em> <em>task</em> \u2192 189.\nThe cache effect compounds this: structured, consistent context across turns hits the cache more reliably than ad-hoc file reads that change every turn. So the additional input tokens <em>cost</em> almost nothing (90% discount) while the output token reduction saves the most expensive tokens.\nThe general principle: with tiered token pricing, optimizing for total token count is wrong. You should optimize for token mix \u2014 push volume from expensive tiers (output, cache miss) to cheap tiers (cache hit). More total tokens can <em>cost</em> less if you shift the distribution.\nThis seems obvious in retrospect but I haven't seen it discussed much. Most context engineering work focuses on reducing input tokens. The bigger lever might be reducing output tokens by improving input signal-to-noise ratio \u2014 the model writes less when it doesn't have to think out loud about what it's reading.<p>The tool is vexp (https://vexp.dev) \u2014 local-first context engine, Rust + tree-sitter + SQLite. Free tier available."},"title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["cost"],"value":"More tokens, less <em>cost</em>: why optimizing for token count is wrong"}},"_tags":["story","author_nicola_alessi","story_47326918","ask_hn"],"author":"nicola_alessi","children":[47326980,47327028,47328079,47344045],"created_at":"2026-03-10T18:20:13Z","created_at_i":1773166813,"num_comments":9,"objectID":"47326918","points":1,"story_id":47326918,"story_text":"I ran a controlled benchmark on AI coding agents (42 runs, FastAPI, Claude Sonnet 4.6) and found something that broke my mental model of LLM costs.\nThe setup: I built an MCP server that pre-indexes a codebase into a dependency graph and serves pre-ranked context to the agent in a single call, instead of letting the agent explore files on its own.\nThe expected result: less input context \u2192 lower cost. Straightforward.\nThe actual result: total tokens processed went UP 20% (23.4M vs 19.6M) while total cost went DOWN 58% ($6.89 vs $16.29).\nThe explanation is in how Anthropic prices tokens. There are three pricing tiers:<p>Output tokens: most expensive (3-5x input price)\nInput tokens (cache miss): full price\nInput tokens (cache hit): 90% discount<p>The agent with pre-indexed context processes more total tokens because the structured context payload is injected every turn. But the token MIX shifts dramatically:\nOutput tokens:     10,588 \u2192 3,965  (-63%)\nCache read rate:   93.8% \u2192 95.3%\nCache creation:    6.1% \u2192 4.6%\nOutput tokens dominate the cost equation. When the agent receives 40K tokens of unfiltered context, it generates verbose orientation narration (&quot;let me look at this file... I can see that...&quot;). When it receives 8K tokens of graph-ranked context, it skips straight to the answer. 504 output tokens per task \u2192 189.\nThe cache effect compounds this: structured, consistent context across turns hits the cache more reliably than ad-hoc file reads that change every turn. So the additional input tokens cost almost nothing (90% discount) while the output token reduction saves the most expensive tokens.\nThe general principle: with tiered token pricing, optimizing for total token count is wrong. You should optimize for token mix \u2014 push volume from expensive tiers (output, cache miss) to cheap tiers (cache hit). More total tokens can cost less if you shift the distribution.\nThis seems obvious in retrospect but I haven&#x27;t seen it discussed much. Most context engineering work focuses on reducing input tokens. The bigger lever might be reducing output tokens by improving input signal-to-noise ratio \u2014 the model writes less when it doesn&#x27;t have to think out loud about what it&#x27;s reading.<p>The tool is vexp (https:&#x2F;&#x2F;vexp.dev) \u2014 local-first context engine, Rust + tree-sitter + SQLite. Free tier available.","title":"More tokens, less cost: why optimizing for token count is wrong","updated_at":"2026-04-28T03:44:13Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"robinbanner"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["cost","per","task","llm","benchmark"],"value":"I got frustrated paying $60/M tokens for reasoning queries when a $0.80/M model gives comparable results for most of them. So I built Komilion \u2014 a model router that classifies each API request and routes it to a cheaper model that fits.<p>- Drop-in replacement for the OpenAI SDK (change one line: base_url)\n- Each query gets classified (regex fast path + lightweight <em>LLM</em> classifier) and matched against ~390 models\n- Three tiers (Frugal/Balanced/Premium) to control the quality-<em>cost</em> tradeoff\n- Automatic failover if a provider goes down\n- <em>Cost</em> metadata in every response<p>The routing logic is <em>benchmark</em>-driven (LMArena, Artificial Analysis), not ML-based \u2014 simpler to debug and reason about. The regex fast path handles ~60% of requests in under 5ms with zero API calls.<p>Example: a customer support bot doing 10K conversations/month went from ~$250/mo (everything pinned to Opus 4.6) to ~$40/mo with routing. Most conversations were FAQ-level questions that a smaller model handled fine.<p>Stack: Next.js, Vercel, Neon PostgreSQL, OpenRouter upstream. Hosting <em>cost</em>: ~$20/month.<p>We ran a head-to-head <em>benchmark</em>: same 15 prompts through Opus, GPT-4o, Gemini Pro, and the router. Simple <em>tasks</em> <em>cost</em> 66% less with routing. Complex <em>tasks</em> produced 2x more detailed output because the router picked specialized models <em>per</em> <em>task</em> type. Full data: <a href=\"https://dev.to/robinbanner/we-benchmarked-4-ai-api-strategies-with-real-money-the-results-changed-how-we-think-about-model-5coa\" rel=\"nofollow\">https://dev.to/robinbanner/we-benchmarked-4-ai-api-strategie...</a><p>Architecture writeup: <a href=\"https://dev.to/robinbanner/inside-komilions-architecture-how-we-route-ai-requests-across-394-models-3e5g\" rel=\"nofollow\">https://dev.to/robinbanner/inside-komilions-architecture-how...</a> \u2014 there's a free tier if you want to try it."},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: API router that picks the cheapest model that fits each query"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://www.komilion.com/"}},"_tags":["story","author_robinbanner","story_47036011","show_hn"],"author":"robinbanner","children":[47036018,47043320],"created_at":"2026-02-16T15:11:25Z","created_at_i":1771254685,"num_comments":1,"objectID":"47036011","points":1,"story_id":47036011,"story_text":"I got frustrated paying $60&#x2F;M tokens for reasoning queries when a $0.80&#x2F;M model gives comparable results for most of them. So I built Komilion \u2014 a model router that classifies each API request and routes it to a cheaper model that fits.<p>- Drop-in replacement for the OpenAI SDK (change one line: base_url)\n- Each query gets classified (regex fast path + lightweight LLM classifier) and matched against ~390 models\n- Three tiers (Frugal&#x2F;Balanced&#x2F;Premium) to control the quality-cost tradeoff\n- Automatic failover if a provider goes down\n- Cost metadata in every response<p>The routing logic is benchmark-driven (LMArena, Artificial Analysis), not ML-based \u2014 simpler to debug and reason about. The regex fast path handles ~60% of requests in under 5ms with zero API calls.<p>Example: a customer support bot doing 10K conversations&#x2F;month went from ~$250&#x2F;mo (everything pinned to Opus 4.6) to ~$40&#x2F;mo with routing. Most conversations were FAQ-level questions that a smaller model handled fine.<p>Stack: Next.js, Vercel, Neon PostgreSQL, OpenRouter upstream. Hosting cost: ~$20&#x2F;month.<p>We ran a head-to-head benchmark: same 15 prompts through Opus, GPT-4o, Gemini Pro, and the router. Simple tasks cost 66% less with routing. Complex tasks produced 2x more detailed output because the router picked specialized models per task type. Full data: <a href=\"https:&#x2F;&#x2F;dev.to&#x2F;robinbanner&#x2F;we-benchmarked-4-ai-api-strategies-with-real-money-the-results-changed-how-we-think-about-model-5coa\" rel=\"nofollow\">https:&#x2F;&#x2F;dev.to&#x2F;robinbanner&#x2F;we-benchmarked-4-ai-api-strategie...</a><p>Architecture writeup: <a href=\"https:&#x2F;&#x2F;dev.to&#x2F;robinbanner&#x2F;inside-komilions-architecture-how-we-route-ai-requests-across-394-models-3e5g\" rel=\"nofollow\">https:&#x2F;&#x2F;dev.to&#x2F;robinbanner&#x2F;inside-komilions-architecture-how...</a> \u2014 there&#x27;s a free tier if you want to try it.","title":"Show HN: API router that picks the cheapest model that fits each query","updated_at":"2026-03-05T23:33:47Z","url":"https://www.komilion.com/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"isoldex"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["cost","per","task","llm","benchmark"],"value":"Hi HN - author of Sentinel here.<p>I built Sentinel after using browser-use and Stagehand on a client\nproject and hitting two recurring issues: flaky reliability on\nmulti-step flows, and token <em>costs</em> that ate the budget on anything\nnon-trivial. I suspected the root cause was architectural - both\nlean on the <em>LLM</em> re-reading large portions of the page each step -\nand tried Chrome's Accessibility Object Model (AOM) as the\nobservation layer instead.<p>To check whether that architectural choice actually mattered, I\nbuilt a 9-<em>task</em> <em>benchmark</em> comparing Sentinel, Stagehand, and\nbrowser-use against the same Gemini 3 Flash Preview model, same\nprompts, same programmatic validators, 5 runs <em>per</em> <em>task</em>-tool combo.\nRaw <em>per</em>-run JSON is committed so you can recompute or challenge\nevery number.<p>Headline numbers:\n - Tokens: Sentinel uses 3.1x-56.9x fewer than browser-use,\n   1.4x-13.3x fewer than Stagehand.\n - Reliability: Sentinel 100% (45/45), browser-use 100% (45/45),\n   Stagehand 86.7% (39/45).\n - Speed: Sentinel is fastest on 5 of 9 tasks.\n - The harder the <em>task</em>, the bigger the token gap.<p>Caveats up front:\n - I built Sentinel - treat this as a starting point for your own\n   verification, not an impartial survey. README has a full\n   known-limitations section.\n - Single model (Gemini 3 Flash Preview, which is also Stagehand's\n   documented recommendation).\n - 9 tasks is small; raw JSON is there if you want to add tasks\n   or rerun on a different model.\n - Each framework is used with its idiomatic API (Sentinel/Stagehand:\n   discrete act()/extract(); browser-use: agent-loop prompt).\n   Forcing them into the same call pattern would disadvantage\n   whichever is optimized for the other.<p>Sentinel is already in production with paying clients (all\nself-hosted), which covers development <em>costs</em>.\nA managed offering is on the table\nif there's real demand: you'd pay infra + model usage at <em>cost</em>, no\nmargin. Drop a comment if that would unblock you, otherwise I'd\nrather not maintain hosting nobody needs."},"story_title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["benchmark"],"value":"Show HN: Sentinel \u2013 browser agent using 3x+ fewer tokens (open <em>benchmark</em>)"},"story_url":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["benchmark"],"value":"https://github.com/ArasHuseyin/browser-agent-<em>benchmark</em>"}},"_tags":["comment","author_isoldex","story_48196521"],"author":"isoldex","comment_text":"Hi HN - author of Sentinel here.<p>I built Sentinel after using browser-use and Stagehand on a client\nproject and hitting two recurring issues: flaky reliability on\nmulti-step flows, and token costs that ate the budget on anything\nnon-trivial. I suspected the root cause was architectural - both\nlean on the LLM re-reading large portions of the page each step -\nand tried Chrome&#x27;s Accessibility Object Model (AOM) as the\nobservation layer instead.<p>To check whether that architectural choice actually mattered, I\nbuilt a 9-task benchmark comparing Sentinel, Stagehand, and\nbrowser-use against the same Gemini 3 Flash Preview model, same\nprompts, same programmatic validators, 5 runs per task-tool combo.\nRaw per-run JSON is committed so you can recompute or challenge\nevery number.<p>Headline numbers:\n - Tokens: Sentinel uses 3.1x-56.9x fewer than browser-use,\n   1.4x-13.3x fewer than Stagehand.\n - Reliability: Sentinel 100% (45&#x2F;45), browser-use 100% (45&#x2F;45),\n   Stagehand 86.7% (39&#x2F;45).\n - Speed: Sentinel is fastest on 5 of 9 tasks.\n - The harder the task, the bigger the token gap.<p>Caveats up front:\n - I built Sentinel - treat this as a starting point for your own\n   verification, not an impartial survey. README has a full\n   known-limitations section.\n - Single model (Gemini 3 Flash Preview, which is also Stagehand&#x27;s\n   documented recommendation).\n - 9 tasks is small; raw JSON is there if you want to add tasks\n   or rerun on a different model.\n - Each framework is used with its idiomatic API (Sentinel&#x2F;Stagehand:\n   discrete act()&#x2F;extract(); browser-use: agent-loop prompt).\n   Forcing them into the same call pattern would disadvantage\n   whichever is optimized for the other.<p>Sentinel is already in production with paying clients (all\nself-hosted), which covers development costs.\nA managed offering is on the table\nif there&#x27;s real demand: you&#x27;d pay infra + model usage at cost, no\nmargin. Drop a comment if that would unblock you, otherwise I&#x27;d\nrather not maintain hosting nobody needs.","created_at":"2026-05-19T17:40:52Z","created_at_i":1779212452,"objectID":"48196527","parent_id":48196521,"story_id":48196521,"story_title":"Show HN: Sentinel \u2013 browser agent using 3x+ fewer tokens (open benchmark)","story_url":"https://github.com/ArasHuseyin/browser-agent-benchmark","updated_at":"2026-05-19T17:44:18Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"Jweb_Guru"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["cost","per","task","llm","benchmark"],"value":"I'm mostly surprised that people found the output quality of Opus 4.6 good enough... 4.7 so far is a pretty sizable improvement for the stuff I care about.  I don't really care how cheap 4.6 was <em>per</em> <em>task</em> when 90% of the tasks weren't actually being done correctly.  Or maybe it's that people like the <em>LLM</em> agreeing with them blindly while sneakily doing something else under the hood?  Did people enjoy Claude routinely disregarding their instructions?  Not really sure I understand, I truly found 4.6 immensely frustrating (from the getgo, not just the &quot;pre-nerf&quot; version, whatever that means).  4.7 is a buggy mess, it's slow, and it <em>costs</em> a lot <em>per</em> token.  It's also a huge breath of fresh air because it actually seems to make a good faith effort at doing the thing you asked it to do, and doesn't waste your time with irrelevant nonsense just to make it look busy or because it thinks you want that nonsense (I mean, it still does all of these things to some extent, but so far it seems like it does them much less than 4.6 did).<p>Disclaimer: I'm always running on max and don't really have token limits so I am in a position not to care about <em>cost</em> <em>per</em> token.  But I am not surprised by the improved <em>benchmark</em> results at all, 4.6 was really not nearly as strong of a model as people seem to remember it being."},"story_title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["cost"],"value":"Measuring Claude 4.7's tokenizer <em>costs</em>"},"story_url":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["cost"],"value":"https://www.claudecodecamp.com/p/i-measured-claude-4-7-s-new-tokenizer-here-s-what-it-<em>costs</em>-you"}},"_tags":["comment","author_Jweb_Guru","story_47807006"],"author":"Jweb_Guru","comment_text":"I&#x27;m mostly surprised that people found the output quality of Opus 4.6 good enough... 4.7 so far is a pretty sizable improvement for the stuff I care about.  I don&#x27;t really care how cheap 4.6 was per task when 90% of the tasks weren&#x27;t actually being done correctly.  Or maybe it&#x27;s that people like the LLM agreeing with them blindly while sneakily doing something else under the hood?  Did people enjoy Claude routinely disregarding their instructions?  Not really sure I understand, I truly found 4.6 immensely frustrating (from the getgo, not just the &quot;pre-nerf&quot; version, whatever that means).  4.7 is a buggy mess, it&#x27;s slow, and it costs a lot per token.  It&#x27;s also a huge breath of fresh air because it actually seems to make a good faith effort at doing the thing you asked it to do, and doesn&#x27;t waste your time with irrelevant nonsense just to make it look busy or because it thinks you want that nonsense (I mean, it still does all of these things to some extent, but so far it seems like it does them much less than 4.6 did).<p>Disclaimer: I&#x27;m always running on max and don&#x27;t really have token limits so I am in a position not to care about cost per token.  But I am not surprised by the improved benchmark results at all, 4.6 was really not nearly as strong of a model as people seem to remember it being.","created_at":"2026-04-18T03:43:47Z","created_at_i":1776483827,"objectID":"47812958","parent_id":47807839,"story_id":47807006,"story_title":"Measuring Claude 4.7's tokenizer costs","story_url":"https://www.claudecodecamp.com/p/i-measured-claude-4-7-s-new-tokenizer-here-s-what-it-costs-you","updated_at":"2026-04-18T03:48:06Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"robinbanner"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["cost","per","task","llm","benchmark"],"value":"Backstory: I was building a customer support AI for a client last year. We started with Claude Opus for everything because it worked great. The bill was $250/month for maybe 10K conversations.<p>Then I looked at the actual queries. 70% were things like &quot;what are your hours?&quot; and &quot;how do I return something?&quot; \u2014 questions where a $0.80/M-token model gives the same answer as a $15/M-token model. But about 5% were genuinely complex (multi-step troubleshooting, product comparisons requiring reasoning) where Opus was noticeably better.<p>I started manually routing: simple patterns to a cheap model, everything else to Opus. The bill dropped to $40/month with no quality complaints from users. But maintaining the routing logic across projects got tedious \u2014 every new app needed the same classification + model selection + failover logic.<p>So I built Komilion to package it up. The classification runs in two stages:<p>1. A regex fast path catches ~60% of requests instantly (greetings, FAQ patterns, simple classification <em>tasks</em>). Zero API calls, under 5ms.<p>2. For the rest, a lightweight <em>LLM</em> classifier determines <em>task</em> type and complexity, then matches against a routing table built from LMArena and Artificial Analysis <em>benchmark</em> data.<p>What surprised me in the <em>benchmark</em> data: complex <em>tasks</em> through the router actually produced MORE detailed output than any single pinned model (6,614 chars avg vs 3,573 for Opus). The router selects specialized models <em>per</em> <em>task</em> type rather than using a generalist model for everything.<p>Stack: Next.js on Vercel, Neon PostgreSQL, OpenRouter upstream. Total hosting <em>cost</em> ~$20/month. It's a solo project.<p>The thing I'd do differently: I should have started with the <em>benchmark</em> data instead of building the product first. The numbers make the case better than any feature list.<p>Happy to answer technical questions about the routing logic, <em>benchmark</em> methodology, or anything else."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: API router that picks the cheapest model that fits each query"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://www.komilion.com/"}},"_tags":["comment","author_robinbanner","story_47036011"],"author":"robinbanner","comment_text":"Backstory: I was building a customer support AI for a client last year. We started with Claude Opus for everything because it worked great. The bill was $250&#x2F;month for maybe 10K conversations.<p>Then I looked at the actual queries. 70% were things like &quot;what are your hours?&quot; and &quot;how do I return something?&quot; \u2014 questions where a $0.80&#x2F;M-token model gives the same answer as a $15&#x2F;M-token model. But about 5% were genuinely complex (multi-step troubleshooting, product comparisons requiring reasoning) where Opus was noticeably better.<p>I started manually routing: simple patterns to a cheap model, everything else to Opus. The bill dropped to $40&#x2F;month with no quality complaints from users. But maintaining the routing logic across projects got tedious \u2014 every new app needed the same classification + model selection + failover logic.<p>So I built Komilion to package it up. The classification runs in two stages:<p>1. A regex fast path catches ~60% of requests instantly (greetings, FAQ patterns, simple classification tasks). Zero API calls, under 5ms.<p>2. For the rest, a lightweight LLM classifier determines task type and complexity, then matches against a routing table built from LMArena and Artificial Analysis benchmark data.<p>What surprised me in the benchmark data: complex tasks through the router actually produced MORE detailed output than any single pinned model (6,614 chars avg vs 3,573 for Opus). The router selects specialized models per task type rather than using a generalist model for everything.<p>Stack: Next.js on Vercel, Neon PostgreSQL, OpenRouter upstream. Total hosting cost ~$20&#x2F;month. It&#x27;s a solo project.<p>The thing I&#x27;d do differently: I should have started with the benchmark data instead of building the product first. The numbers make the case better than any feature list.<p>Happy to answer technical questions about the routing logic, benchmark methodology, or anything else.","created_at":"2026-02-16T15:11:36Z","created_at_i":1771254696,"objectID":"47036018","parent_id":47036011,"story_id":47036011,"story_title":"Show HN: API router that picks the cheapest model that fits each query","story_url":"https://www.komilion.com/","updated_at":"2026-03-05T23:33:47Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"faizshah"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["cost","per","task","llm","benchmark"],"value":"The argument has never changed the argument has always been the same.<p><em>LLMs</em> do not think, they do not perform logic they are approximating thought. The reason why CoT works is because of the main feature of <em>LLMs</em>, they are extremely good at picking reasonable next tokens based on the context.<p><em>LLM</em> are good and always have been good at three types of <em>tasks</em>:<p>- Closed form problems where the answer is in the prompt (CoT, Prompt Engineering, RAG)<p>- Recall from the training set as the Parameter space increases (15B -&gt; 70B -&gt; almost 1T now)<p>- Generalization and Zero shot <em>tasks</em> as a result of the first two (this is also what causes hallucinations which is a feature not a bug, we want the <em>LLM</em> to imitate thought not be a Q&amp;A expert system from 1990)<p>If you keep being fooled by <em>LLM</em> thinking they are AGI after every impressive <em>benchmark</em> and everyone keeps telling you that in practice <em>LLM</em> are not good at <em>tasks</em> that are poorly defined, require niche knowledge, or require a special mental model that is on you.<p>I use <em>LLM</em> every day I speed up many <em>tasks</em> that would take 5-15 mins down to 10-120 seconds (worst case for re-prompts). Many times my <em>tasks</em> take longer than if I had done it myself because it\u2019s not my work im just copying it. But overall I am more productive because of <em>LLM</em>.<p>Does <em>LLM</em> speeding up your work mean that <em>LLM</em> can replace Humans?<p>Personally I still don\u2019t think <em>LLM</em> can replace Humans <i>at the same level of quality</i> because they are imitating thought not actually thinking. Now the question among the corporate overlords is will you reduce operating <em>costs</em> by XX% <em>per</em> year (wages) but reducing the quality of service for customers. The last 50 years have shown us the answer\u2026"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Promising results from DeepSeek R1 for code"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://simonwillison.net/2025/Jan/27/llamacpp-pr/"}},"_tags":["comment","author_faizshah","story_42852866"],"author":"faizshah","comment_text":"The argument has never changed the argument has always been the same.<p>LLMs do not think, they do not perform logic they are approximating thought. The reason why CoT works is because of the main feature of LLMs, they are extremely good at picking reasonable next tokens based on the context.<p>LLM are good and always have been good at three types of tasks:<p>- Closed form problems where the answer is in the prompt (CoT, Prompt Engineering, RAG)<p>- Recall from the training set as the Parameter space increases (15B -&gt; 70B -&gt; almost 1T now)<p>- Generalization and Zero shot tasks as a result of the first two (this is also what causes hallucinations which is a feature not a bug, we want the LLM to imitate thought not be a Q&amp;A expert system from 1990)<p>If you keep being fooled by LLM thinking they are AGI after every impressive benchmark and everyone keeps telling you that in practice LLM are not good at tasks that are poorly defined, require niche knowledge, or require a special mental model that is on you.<p>I use LLM every day I speed up many tasks that would take 5-15 mins down to 10-120 seconds (worst case for re-prompts). Many times my tasks take longer than if I had done it myself because it\u2019s not my work im just copying it. But overall I am more productive because of LLM.<p>Does LLM speeding up your work mean that LLM can replace Humans?<p>Personally I still don\u2019t think LLM can replace Humans <i>at the same level of quality</i> because they are imitating thought not actually thinking. Now the question among the corporate overlords is will you reduce operating costs by XX% per year (wages) but reducing the quality of service for customers. The last 50 years have shown us the answer\u2026","created_at":"2025-01-28T17:32:18Z","created_at_i":1738085538,"objectID":"42855218","parent_id":42854411,"story_id":42852866,"story_title":"Promising results from DeepSeek R1 for code","story_url":"https://simonwillison.net/2025/Jan/27/llamacpp-pr/","updated_at":"2025-02-03T12:48:31Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"popthetopnow"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["cost","per","task","llm","benchmark"],"value":"Below is a shortened summary (under 4000 characters) capturing the key points from the Hacker News thread:<p>What is ARC-AGI?<p>A puzzle-based <em>benchmark</em> (\u201ceasy for humans, hard for AI\u201d) using visual/grid <em>tasks</em> akin to Raven\u2019s Matrices.\nProposed by Fran\u00e7ois Chollet et al. to measure genuine generalization, not just memorization.\nConsists of public, semi-private, and private (hidden) test sets.\nHistorically, large language models (<em>LLMs</em>) struggled on it, cited as evidence they lacked \u201ctrue\u201d reasoning.\nWhat did \u201co3\u201d accomplish?<p>OpenAI\u2019s \u201co3\u201d scored 87.5\u201391.5% on the ARC-AGI public/semi-private sets, surpassing the ~85% threshold for the ARC prize.\nHuman benchmarks range ~64\u201376% for average Turkers, &gt;95% for STEM grads.\nTwo modes:\n\u201co3-low\u201d: Cheaper (~$17\u2013$20 <em>per</em> puzzle), achieves ~75\u201382% accuracy.\n\u201co3-high\u201d: 172\u00d7 more compute, ~$1k\u2013$3k+ <em>per</em> puzzle, scoring 87\u201391.5%.\nOpenAI may have spent hundreds of thousands of dollars on 400 test puzzles in high-compute mode.\nReactions<p>Impressive but expensive: Thousands of dollars <em>per</em> puzzle is <em>cost</em>-prohibitive (human solvers might <em>cost</em> $5/puzzle).\nNot \u201cAGI\u201d: Even ARC\u2019s creators say passing ARC doesn\u2019t prove AGI; Chollet notes \u201co3\u201d fails some trivial puzzles. A new \u201cARC-AGI-2\u201d could slash its score to ~30%.\nLikely uses massive search at inference: Possibly a \u201ctree-of-thought\u201d approach (many sub-reasoning steps, a verifier, then a final answer), similar to AlphaZero\u2019s search.\nOther benchmarks: \u201co3\u201d jumped from ~48% to 70+% on SWE-Bench (code-fixing) and from single-digit to ~25% on Frontier Math (Olympiad-level problems).\nEmployment &amp; Future<p>Fears of rapid AI displacement of programmers and knowledge workers. Others see a parallel to the personal computer revolution, with humans still in the loop.\nReal-world <em>tasks</em> need maintainable solutions and large contexts, not just puzzle-solving.\nCurrent compute costs (~$1k\u2013$3k <em>per</em> puzzle) are impractical. However, costs have declined steadily, suggesting future viability.\nGoalpost-Shifting<p>Benchmarks lose perceived importance once models exceed them (as with Chess and Go).\nARC authors never claimed it was a definitive AGI test\u2014just a measure of \u201cgeneralization.\u201d\nInference-time vs. Training-time<p>Models used to rely on big training sets + single forward passes; now we see iterative chain-of-thought plus large-scale search at inference.\nCosts could decrease with more efficient \u201cinternal\u201d search methods or session-based learning that amortizes expense.\nMain Takeaways<p>OpenAI\u2019s \u201co3\u201d demonstrates near-human (or higher) performance on ARC-AGI by combining an <em>LLM</em> with a heavy search strategy.\nThis approach is incredibly costly but shows that, if you throw enough compute and methodical search at puzzles, <em>LLMs</em> can outperform average humans.\nPassing ARC doesn\u2019t confirm AGI; \u201co3\u201d still struggles on some simple puzzles, and a new version of the <em>benchmark</em> aims to challenge it further.\nLong term, decreasing costs and continuing progress could enable practical AI-driven automation in coding, advanced math, research, and more."},"title":{"matchLevel":"none","matchedWords":[],"value":"I fed all the comments from \"OpenAI O3 breakthrough\" HT post to o1 pro mode"}},"_tags":["story","author_popthetopnow","story_42477484","ask_hn"],"author":"popthetopnow","children":[42481265],"created_at":"2024-12-21T04:07:44Z","created_at_i":1734754064,"num_comments":1,"objectID":"42477484","points":1,"story_id":42477484,"story_text":"Below is a shortened summary (under 4000 characters) capturing the key points from the Hacker News thread:<p>What is ARC-AGI?<p>A puzzle-based benchmark (\u201ceasy for humans, hard for AI\u201d) using visual&#x2F;grid tasks akin to Raven\u2019s Matrices.\nProposed by Fran\u00e7ois Chollet et al. to measure genuine generalization, not just memorization.\nConsists of public, semi-private, and private (hidden) test sets.\nHistorically, large language models (LLMs) struggled on it, cited as evidence they lacked \u201ctrue\u201d reasoning.\nWhat did \u201co3\u201d accomplish?<p>OpenAI\u2019s \u201co3\u201d scored 87.5\u201391.5% on the ARC-AGI public&#x2F;semi-private sets, surpassing the ~85% threshold for the ARC prize.\nHuman benchmarks range ~64\u201376% for average Turkers, &gt;95% for STEM grads.\nTwo modes:\n\u201co3-low\u201d: Cheaper (~$17\u2013$20 per puzzle), achieves ~75\u201382% accuracy.\n\u201co3-high\u201d: 172\u00d7 more compute, ~$1k\u2013$3k+ per puzzle, scoring 87\u201391.5%.\nOpenAI may have spent hundreds of thousands of dollars on 400 test puzzles in high-compute mode.\nReactions<p>Impressive but expensive: Thousands of dollars per puzzle is cost-prohibitive (human solvers might cost $5&#x2F;puzzle).\nNot \u201cAGI\u201d: Even ARC\u2019s creators say passing ARC doesn\u2019t prove AGI; Chollet notes \u201co3\u201d fails some trivial puzzles. A new \u201cARC-AGI-2\u201d could slash its score to ~30%.\nLikely uses massive search at inference: Possibly a \u201ctree-of-thought\u201d approach (many sub-reasoning steps, a verifier, then a final answer), similar to AlphaZero\u2019s search.\nOther benchmarks: \u201co3\u201d jumped from ~48% to 70+% on SWE-Bench (code-fixing) and from single-digit to ~25% on Frontier Math (Olympiad-level problems).\nEmployment &amp; Future<p>Fears of rapid AI displacement of programmers and knowledge workers. Others see a parallel to the personal computer revolution, with humans still in the loop.\nReal-world tasks need maintainable solutions and large contexts, not just puzzle-solving.\nCurrent compute costs (~$1k\u2013$3k per puzzle) are impractical. However, costs have declined steadily, suggesting future viability.\nGoalpost-Shifting<p>Benchmarks lose perceived importance once models exceed them (as with Chess and Go).\nARC authors never claimed it was a definitive AGI test\u2014just a measure of \u201cgeneralization.\u201d\nInference-time vs. Training-time<p>Models used to rely on big training sets + single forward passes; now we see iterative chain-of-thought plus large-scale search at inference.\nCosts could decrease with more efficient \u201cinternal\u201d search methods or session-based learning that amortizes expense.\nMain Takeaways<p>OpenAI\u2019s \u201co3\u201d demonstrates near-human (or higher) performance on ARC-AGI by combining an LLM with a heavy search strategy.\nThis approach is incredibly costly but shows that, if you throw enough compute and methodical search at puzzles, LLMs can outperform average humans.\nPassing ARC doesn\u2019t confirm AGI; \u201co3\u201d still struggles on some simple puzzles, and a new version of the benchmark aims to challenge it further.\nLong term, decreasing costs and continuing progress could enable practical AI-driven automation in coding, advanced math, research, and more.","title":"I fed all the comments from \"OpenAI O3 breakthrough\" HT post to o1 pro mode","updated_at":"2025-01-07T06:50:41Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"gauravvij137"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["cost","per","task","llm","benchmark"],"value":"I got tired of picking LLMs based on vibes and leaderboards that don't reflect real workloads, so I built this.<p>You describe a <em>task</em> in plain English. The tool generates a test suite for that specific <em>task</em>, discovers candidate models via OpenRouter, <em>benchmarks</em> them in parallel, and uses a Judge <em>LLM</em> to score every response across 5 dimensions: accuracy, hallucination, grounding, tool-calling, and clarity.<p>Output is a ranked top 3 with average latency <em>per</em> model and a <em>task</em>-specific system prompt optimized for the winner.<p>A few things I learned while building it:<p>- Score and latency rarely correlate. The best model for accuracy on coding tasks was almost never the fastest. This tradeoff is completely <em>task</em>-dependent and impossible to see from <em>benchmarks</em> that don't reflect your workload.\n- The Judge <em>LLM</em> approach is surprisingly consistent but introduces positional and familiarity bias. Using one model to score others isn't perfect, but it's far more reproducible than manual eval. Open to ideas on how to reduce judge bias without blowing up the <em>cost</em>.\n- Model discovery matters more than I expected. The top performers on generic <em>benchmarks</em> often weren't the top performers on narrow tasks.<p>Stack: Python, OpenRouter for model access, MIT licensed.<p><a href=\"https://github.com/gauravvij/llm-evaluator\" rel=\"nofollow\">https://github.com/gauravvij/<em>llm</em>-evaluator</a><p>Happy to answer questions on the design decisions."},"title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["task","llm"],"value":"Show HN: Auto <em>LLM</em> Ranker \u2013 Describe a <em>task</em> in English and get ranked models"},"url":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["llm"],"value":"https://github.com/gauravvij/<em>llm</em>-evaluator"}},"_tags":["story","author_gauravvij137","story_47308325","show_hn"],"author":"gauravvij137","children":[47334718],"created_at":"2026-03-09T12:44:46Z","created_at_i":1773060286,"num_comments":0,"objectID":"47308325","points":3,"story_id":47308325,"story_text":"I got tired of picking LLMs based on vibes and leaderboards that don&#x27;t reflect real workloads, so I built this.<p>You describe a task in plain English. The tool generates a test suite for that specific task, discovers candidate models via OpenRouter, benchmarks them in parallel, and uses a Judge LLM to score every response across 5 dimensions: accuracy, hallucination, grounding, tool-calling, and clarity.<p>Output is a ranked top 3 with average latency per model and a task-specific system prompt optimized for the winner.<p>A few things I learned while building it:<p>- Score and latency rarely correlate. The best model for accuracy on coding tasks was almost never the fastest. This tradeoff is completely task-dependent and impossible to see from benchmarks that don&#x27;t reflect your workload.\n- The Judge LLM approach is surprisingly consistent but introduces positional and familiarity bias. Using one model to score others isn&#x27;t perfect, but it&#x27;s far more reproducible than manual eval. Open to ideas on how to reduce judge bias without blowing up the cost.\n- Model discovery matters more than I expected. The top performers on generic benchmarks often weren&#x27;t the top performers on narrow tasks.<p>Stack: Python, OpenRouter for model access, MIT licensed.<p><a href=\"https:&#x2F;&#x2F;github.com&#x2F;gauravvij&#x2F;llm-evaluator\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;gauravvij&#x2F;llm-evaluator</a><p>Happy to answer questions on the design decisions.","title":"Show HN: Auto LLM Ranker \u2013 Describe a task in English and get ranked models","updated_at":"2026-03-11T12:30:32Z","url":"https://github.com/gauravvij/llm-evaluator"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"p-s-v"},"story_text":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["cost","per","task"],"value":"Hi HN,<p>I\u2019m building *new.knife.day* (<a href=\"https://new.knife.day\" rel=\"nofollow\">https://new.knife.day</a>), a crowd-sourced\ndatabase of every cutlery maker\u2014from Al Mar to brands so small they barely\nshow up on Google.  That means I need an automated way to fetch each brand\u2019s\n<i>official</i> website, even for fringe names like \u201cActilam\u201d or \u201cAiorosu Knives\u201d.<p>So I threw the <em>task</em> at eight web-enabled LLMs via OpenRouter:<p><pre><code>  \u2022 gpt-4o and gpt-4o-mini\n  \u2022 claude-sonnet-4\n  \u2022 gemini-2.5-pro and gemini-2.0-flash\n  \u2022 llama-3.1-70b\n  \u2022 qwen-2.5-72b\n  \u2022 perplexity sonar-deep-research\n</code></pre>\nPrompt:  Return *only* JSON  { brand, official_url, confidence }\nData set: 10 obscure knife brands\nScoring:  exact domain = correct; \u201cno official site\u201d (with reason) = correct\n<em>Costs</em>:   OpenRouter prices on 31 May 2025 (Perplexity billed separately)<p>Highlights\n----------<p><pre><code>  \u2022 Perplexity hit 10/10 but <em>cost</em> $9.42 (860 k tokens!).\n  \u2022 GPT-4o-mini &amp; Llama-3.1-70B got 9/10 for ~2 \u00a2 <em>per</em> correct URL.\n  \u2022 Gemini Flash managed 7/10 for $0.001 total\u2014great if you can QA the misses.\n  \u2022 Half of Gemini 2.5 Pro\u2019s replies were HTML tables my parser rejected.\n</code></pre>\nFull table, code, and raw logs are in the post (and on GitHub).<p>Take-aways\n----------<p><pre><code>  1. 90 % accuracy + quick human review often beats 100 % accuracy that <em>costs</em>\n     45\u00d7 more.\n  2. Structured output is part of model quality\u2014validate JSON on arrival.\n  3. Promo pricing moves fast; always ping the price API before large runs.\n</code></pre>\nNext step: wire GPT-4o-mini into *new.knife.day* so visitors get verified\nmanufacturer links.  Crawling ~250 brands now <em>costs</em> under $5.<p>Curious what you\u2019d improve, and which model you\u2019d bet on for similar\n\u201cfind the canonical URL\u201d tasks.  AMA on the setup, prompts, or results!"},"title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["llm","benchmark"],"value":"Show HN: Which <em>LLM</em> Finds Obscure Knife-Brand URLs Cheapest? (8-Model <em>Benchmark</em>)"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://new.knife.day/blog/using-llms-for-knife-brand-research"}},"_tags":["story","author_p-s-v","story_44180576","show_hn"],"author":"p-s-v","created_at":"2025-06-04T13:34:59Z","created_at_i":1749044099,"num_comments":0,"objectID":"44180576","points":2,"story_id":44180576,"story_text":"Hi HN,<p>I\u2019m building *new.knife.day* (<a href=\"https:&#x2F;&#x2F;new.knife.day\" rel=\"nofollow\">https:&#x2F;&#x2F;new.knife.day</a>), a crowd-sourced\ndatabase of every cutlery maker\u2014from Al Mar to brands so small they barely\nshow up on Google.  That means I need an automated way to fetch each brand\u2019s\n<i>official</i> website, even for fringe names like \u201cActilam\u201d or \u201cAiorosu Knives\u201d.<p>So I threw the task at eight web-enabled LLMs via OpenRouter:<p><pre><code>  \u2022 gpt-4o and gpt-4o-mini\n  \u2022 claude-sonnet-4\n  \u2022 gemini-2.5-pro and gemini-2.0-flash\n  \u2022 llama-3.1-70b\n  \u2022 qwen-2.5-72b\n  \u2022 perplexity sonar-deep-research\n</code></pre>\nPrompt:  Return *only* JSON  { brand, official_url, confidence }\nData set: 10 obscure knife brands\nScoring:  exact domain = correct; \u201cno official site\u201d (with reason) = correct\nCosts:   OpenRouter prices on 31 May 2025 (Perplexity billed separately)<p>Highlights\n----------<p><pre><code>  \u2022 Perplexity hit 10&#x2F;10 but cost $9.42 (860 k tokens!).\n  \u2022 GPT-4o-mini &amp; Llama-3.1-70B got 9&#x2F;10 for ~2 \u00a2 per correct URL.\n  \u2022 Gemini Flash managed 7&#x2F;10 for $0.001 total\u2014great if you can QA the misses.\n  \u2022 Half of Gemini 2.5 Pro\u2019s replies were HTML tables my parser rejected.\n</code></pre>\nFull table, code, and raw logs are in the post (and on GitHub).<p>Take-aways\n----------<p><pre><code>  1. 90 % accuracy + quick human review often beats 100 % accuracy that costs\n     45\u00d7 more.\n  2. Structured output is part of model quality\u2014validate JSON on arrival.\n  3. Promo pricing moves fast; always ping the price API before large runs.\n</code></pre>\nNext step: wire GPT-4o-mini into *new.knife.day* so visitors get verified\nmanufacturer links.  Crawling ~250 brands now costs under $5.<p>Curious what you\u2019d improve, and which model you\u2019d bet on for similar\n\u201cfind the canonical URL\u201d tasks.  AMA on the setup, prompts, or results!","title":"Show HN: Which LLM Finds Obscure Knife-Brand URLs Cheapest? (8-Model Benchmark)","updated_at":"2025-06-04T13:57:30Z","url":"https://new.knife.day/blog/using-llms-for-knife-brand-research"}],"hitsPerPage":20,"nbHits":35,"nbPages":2,"page":0,"params":"query=cost+per+task+LLM+benchmark&advancedSyntax=true&analyticsTags=backend","processingTimeMS":31,"processingTimingsMS":{"_request":{"roundTrip":17},"afterFetch":{"format":{"highlighting":4,"total":5},"merge":{"total":1},"total":1},"fetch":{"query":15,"scanning":13,"total":29},"total":31},"query":"cost per task LLM benchmark","serverTimeMS":37}
