{"exhaustive":{"nbHits":false,"typo":false},"exhaustiveNbHits":false,"exhaustiveTypo":false,"hits":[{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"randomtoast"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["a100","h100","inference","2026"],"value":"The following is all guess work:<p>Since the start of their partnership in 2019, OpenAI has primarily utilized Microsoft's Azure data centers for training its models. In <em>202</em>3, Microsoft acquired approximately 150,000 <em>H100</em> GPUs. [1]<p>The initial version of GPT-4 ran on a cluster of <em>A100</em> GPUs. It is likely that GPT-5 will run on the newly acquired <em>H100</em> GPUs, and it is plausible that GPT-4 Turbo and GPT-4o also utilize this infrastructure. The <em>inference</em> speed of GPT-5 should not be significantly slower than that of GPT-4 to ensure it remains practical for most applications.<p>Assuming the <em>H100</em> is 4.6 times faster for <em>inference</em> than the <em>A100</em> [2], this gives us a lower bound for performance expectations. I anticipate GPT-5 to be at least five times larger in terms of model parameters. Given that both <em>A100</em> and <em>H100</em> have a maximum capacity of 80GB, it is unlikely we will see a single gigantic model. Instead, we can expect an increase in the number of experts. If GPT-4 operates as a mixture of experts with 8x220 billion parameters, then GPT-5 might scale up to something like 40x220 billion parameters. However, the exact release date, safety measures, and benchmark performance of GPT-5 remain uncertain.<p>[1]: <a href=\"https://www.tomshardware.com/tech-industry/nvidia-ai-and-hpc-gpu-sales-reportedly-approached-half-a-million-units-in-q3-thanks-to-meta-facebook\" rel=\"nofollow\">https://www.tomshardware.com/tech-industry/nvidia-ai-and-hpc...</a><p>[2]: <a href=\"https://nvidia.github.io/TensorRT-LLM/blogs/H100vsA100.html\" rel=\"nofollow\">https://nvidia.github.io/TensorRT-LLM/blogs/H100vsA100.html</a>"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Ask HN: Predictions for when GPT-5 will be released and how safe it will be?"}},"_tags":["comment","author_randomtoast","story_40580928"],"author":"randomtoast","comment_text":"The following is all guess work:<p>Since the start of their partnership in 2019, OpenAI has primarily utilized Microsoft&#x27;s Azure data centers for training its models. In 2023, Microsoft acquired approximately 150,000 H100 GPUs. [1]<p>The initial version of GPT-4 ran on a cluster of A100 GPUs. It is likely that GPT-5 will run on the newly acquired H100 GPUs, and it is plausible that GPT-4 Turbo and GPT-4o also utilize this infrastructure. The inference speed of GPT-5 should not be significantly slower than that of GPT-4 to ensure it remains practical for most applications.<p>Assuming the H100 is 4.6 times faster for inference than the A100 [2], this gives us a lower bound for performance expectations. I anticipate GPT-5 to be at least five times larger in terms of model parameters. Given that both A100 and H100 have a maximum capacity of 80GB, it is unlikely we will see a single gigantic model. Instead, we can expect an increase in the number of experts. If GPT-4 operates as a mixture of experts with 8x220 billion parameters, then GPT-5 might scale up to something like 40x220 billion parameters. However, the exact release date, safety measures, and benchmark performance of GPT-5 remain uncertain.<p>[1]: <a href=\"https:&#x2F;&#x2F;www.tomshardware.com&#x2F;tech-industry&#x2F;nvidia-ai-and-hpc-gpu-sales-reportedly-approached-half-a-million-units-in-q3-thanks-to-meta-facebook\" rel=\"nofollow\">https:&#x2F;&#x2F;www.tomshardware.com&#x2F;tech-industry&#x2F;nvidia-ai-and-hpc...</a><p>[2]: <a href=\"https:&#x2F;&#x2F;nvidia.github.io&#x2F;TensorRT-LLM&#x2F;blogs&#x2F;H100vsA100.html\" rel=\"nofollow\">https:&#x2F;&#x2F;nvidia.github.io&#x2F;TensorRT-LLM&#x2F;blogs&#x2F;H100vsA100.html</a>","created_at":"2024-06-05T07:54:00Z","created_at_i":1717574040,"objectID":"40582456","parent_id":40580928,"story_id":40580928,"story_title":"Ask HN: Predictions for when GPT-5 will be released and how safe it will be?","updated_at":"2024-09-20T17:15:36Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"kkielhofner"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["a100","h100","inference","2026"],"value":"I should have been more clear on this - in terms of total install base I don't know that it's going to be &quot;significant&quot; in terms of customer count.<p>However, I do think it will be at least &quot;noticeable&quot; in terms of individual customers with large spend. Total GCP revenue in <em>202</em>2 was roughly 65 billion and Snap leaving alone is 1.5% of total revenue.<p>Especially looking at ML cases where cloud GPU pricing is wildly expensive - retail on-demand instance <em>A100</em> pricing is at least $3/hr which practically speaking with the AWS pricing model is can be twice that all-in. This is for an instance with 32GB of RAM and 8 VCPUs - which for a lot of <em>A100</em> use cases is useless. Need 32 vCPU and 256 GB of RAM? That's more like $20/hr.<p>A single <em>A100</em> machine that's above and beyond more capable can be had from Dell for roughly $50k, which even factoring in hosting based on colo pricing I've seen has an ROI of ~15 months for constant usage. For the equivalent hardware (and still vastly improved performance - 32vCPU and 256GB of RAM) that ROI gets to less than six months.<p>Yes, the <em>A100</em> is typically used for training (and cloud definitely still makes sense there) but more and more models require the performance and VRAM of a V100/<em>H100</em> for <em>inference</em> (24/365 availability). Do it at any kind of scale/redundancy and ROI catches up even faster. An equivalent to this approach is reserved pricing, which over the 1-3yr term of a lease vs. reserved instance self-hosting becomes almost comically more cost and performance effective. With the extra benefit of actually being more flexible.<p>Financing and leasing is readily available and with various tax incentives (like Section 179 leasing) you can pretty quickly pay for a FTE to manage the infra for you - which is probably a wash anyway because at any kind of &quot;real&quot; scale or complexity you almost certainly already have dedicated human resources just to manage cloud. You don't even ever need for an employee to go to the hosting facility because most will rack and provision your hardware for free. Combined with remote hands and standard warranty support any (in my experience very rare) hardware failures just get handled.<p>I should note that this model almost eliminates the tendency for cloud spend to balloon to many X anticipated/budgeted spend - the all too common story of &quot;sticker shock&quot; from clouds on bandwidth alone that cloud has ridiculous markups on. Colo pricing and leases are fixed cost (with all you can eat port speed bandwidth included or so cheap at 95th percentile billing it's practically a rounding error).<p>I have significant experience at CTO level with both approaches (and hybrid, of course). In many situations the benefits of &quot;self-hosting&quot; vs cloud are dramatic.<p>The extremely effective marketing that has created and perpetuated an industry wide fear of self-hosting and hardware (especially with the &quot;always cloud always&quot; generation) is fading. I think the uptime and reliability promises of cloud are also fading - this thread started off with discussion of yet-another cloud outage. My background is in healthcare and telecom and I'm shocked at the cavalier attitude of just accepting these outages and being down while standing around helpless wondering when your big cloud will acknowledge, communicate, and resolve them. A few machines in a few rack units of space across a couple of facilities generally trounces cloud in reliability and uptime.<p>I love HN and the overall knowledge and quality of discussion here but when it comes to hardware and self-hosting many have completely drunk the cloud Kool-Aid and have zero experience with self-hosting - so no idea what they're talking about. Not saying you personally, just generally."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Ongoing Incident in Google Cloud"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://status.cloud.google.com/incidents/LnvJwfYu3TCyUrcrP7yf#RP1d9aZLNFZEJmTBk8e1"}},"_tags":["comment","author_kkielhofner","story_34956319"],"author":"kkielhofner","comment_text":"I should have been more clear on this - in terms of total install base I don&#x27;t know that it&#x27;s going to be &quot;significant&quot; in terms of customer count.<p>However, I do think it will be at least &quot;noticeable&quot; in terms of individual customers with large spend. Total GCP revenue in 2022 was roughly 65 billion and Snap leaving alone is 1.5% of total revenue.<p>Especially looking at ML cases where cloud GPU pricing is wildly expensive - retail on-demand instance A100 pricing is at least $3&#x2F;hr which practically speaking with the AWS pricing model is can be twice that all-in. This is for an instance with 32GB of RAM and 8 VCPUs - which for a lot of A100 use cases is useless. Need 32 vCPU and 256 GB of RAM? That&#x27;s more like $20&#x2F;hr.<p>A single A100 machine that&#x27;s above and beyond more capable can be had from Dell for roughly $50k, which even factoring in hosting based on colo pricing I&#x27;ve seen has an ROI of ~15 months for constant usage. For the equivalent hardware (and still vastly improved performance - 32vCPU and 256GB of RAM) that ROI gets to less than six months.<p>Yes, the A100 is typically used for training (and cloud definitely still makes sense there) but more and more models require the performance and VRAM of a V100&#x2F;H100 for inference (24&#x2F;365 availability). Do it at any kind of scale&#x2F;redundancy and ROI catches up even faster. An equivalent to this approach is reserved pricing, which over the 1-3yr term of a lease vs. reserved instance self-hosting becomes almost comically more cost and performance effective. With the extra benefit of actually being more flexible.<p>Financing and leasing is readily available and with various tax incentives (like Section 179 leasing) you can pretty quickly pay for a FTE to manage the infra for you - which is probably a wash anyway because at any kind of &quot;real&quot; scale or complexity you almost certainly already have dedicated human resources just to manage cloud. You don&#x27;t even ever need for an employee to go to the hosting facility because most will rack and provision your hardware for free. Combined with remote hands and standard warranty support any (in my experience very rare) hardware failures just get handled.<p>I should note that this model almost eliminates the tendency for cloud spend to balloon to many X anticipated&#x2F;budgeted spend - the all too common story of &quot;sticker shock&quot; from clouds on bandwidth alone that cloud has ridiculous markups on. Colo pricing and leases are fixed cost (with all you can eat port speed bandwidth included or so cheap at 95th percentile billing it&#x27;s practically a rounding error).<p>I have significant experience at CTO level with both approaches (and hybrid, of course). In many situations the benefits of &quot;self-hosting&quot; vs cloud are dramatic.<p>The extremely effective marketing that has created and perpetuated an industry wide fear of self-hosting and hardware (especially with the &quot;always cloud always&quot; generation) is fading. I think the uptime and reliability promises of cloud are also fading - this thread started off with discussion of yet-another cloud outage. My background is in healthcare and telecom and I&#x27;m shocked at the cavalier attitude of just accepting these outages and being down while standing around helpless wondering when your big cloud will acknowledge, communicate, and resolve them. A few machines in a few rack units of space across a couple of facilities generally trounces cloud in reliability and uptime.<p>I love HN and the overall knowledge and quality of discussion here but when it comes to hardware and self-hosting many have completely drunk the cloud Kool-Aid and have zero experience with self-hosting - so no idea what they&#x27;re talking about. Not saying you personally, just generally.","created_at":"2023-02-28T14:52:19Z","created_at_i":1677595939,"objectID":"34969893","parent_id":34961656,"story_id":34956319,"story_title":"Ongoing Incident in Google Cloud","story_url":"https://status.cloud.google.com/incidents/LnvJwfYu3TCyUrcrP7yf#RP1d9aZLNFZEJmTBk8e1","updated_at":"2024-09-20T13:29:22Z"},{"_highlightResult":{"author":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["2026"],"value":"kevin-<em>202</em>5"},"story_text":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["a100","h100","inference"],"value":"I built this to answer &quot;what-if&quot; questions about LLM deployment without spinning up expensive infrastructure.<p>The tool models <em>inference</em> physics - latency, bandwidth saturation, and PCIe bottlenecks for large MoE models like DeepSeek-V3 (671B), Mixtral 8x7B, Qwen2.5-MoE, and Grok-1.<p>Key features:<p>- Independent Prefill vs Decode parallelism config (TP/PP/SP/DP)\n- Hardware modeling: <em>H100</em>, B200, <em>A100</em>, NVLink topologies, IB vs RoCE\n- Optimizations: Paged KV Cache, DualPipe, FP8/INT4 quantization\n- Experimental: Memory Pooling (TPP, tiered storage) and Near-Memory Computing - offload cold experts and cold/warm KV-cache to system RAM, node-shared or global-shared memory pool<p>Live demo: <a href=\"https://llm-inference-performance-calculator-1066033662468.us-west1.run.app/\" rel=\"nofollow\">https://llm-<em>inference</em>-performance-calculator-1066033662468.u...</a><p>Built with React, TypeScript, Tailwind, and Vite.<p>Disclaimer: I've calibrated the math models but they're not perfect. Feedback and PRs welcome."},"title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["inference"],"value":"Show HN: LLM <em>Inference</em> Performance Analytic Tool for Moe Models (DeepSeek/etc.)"},"url":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["inference"],"value":"https://github.com/kevinyuan/llm-<em>inference</em>-perf-model"}},"_tags":["story","author_kevin-2025","story_46071557","show_hn"],"author":"kevin-2025","created_at":"2025-11-27T17:47:09Z","created_at_i":1764265629,"num_comments":0,"objectID":"46071557","points":1,"story_id":46071557,"story_text":"I built this to answer &quot;what-if&quot; questions about LLM deployment without spinning up expensive infrastructure.<p>The tool models inference physics - latency, bandwidth saturation, and PCIe bottlenecks for large MoE models like DeepSeek-V3 (671B), Mixtral 8x7B, Qwen2.5-MoE, and Grok-1.<p>Key features:<p>- Independent Prefill vs Decode parallelism config (TP&#x2F;PP&#x2F;SP&#x2F;DP)\n- Hardware modeling: H100, B200, A100, NVLink topologies, IB vs RoCE\n- Optimizations: Paged KV Cache, DualPipe, FP8&#x2F;INT4 quantization\n- Experimental: Memory Pooling (TPP, tiered storage) and Near-Memory Computing - offload cold experts and cold&#x2F;warm KV-cache to system RAM, node-shared or global-shared memory pool<p>Live demo: <a href=\"https:&#x2F;&#x2F;llm-inference-performance-calculator-1066033662468.us-west1.run.app&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;llm-inference-performance-calculator-1066033662468.u...</a><p>Built with React, TypeScript, Tailwind, and Vite.<p>Disclaimer: I&#x27;ve calibrated the math models but they&#x27;re not perfect. Feedback and PRs welcome.","title":"Show HN: LLM Inference Performance Analytic Tool for Moe Models (DeepSeek/etc.)","updated_at":"2026-03-05T23:09:09Z","url":"https://github.com/kevinyuan/llm-inference-perf-model"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"abediaz"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["a100","h100","inference","2026"],"value":"There's a number that should be on every model builder's whiteboard right now, and almost nobody is talking about it:<p>The maximum model size that fits on the next generation of consumer unified-memory chips.<p>When the leading consumer silicon vendor drops its next lineup (and it's coming soon), millions of developers and power users are going to buy in. Not for the marketing. Because these chips offer something no cloud GPU can: unified memory that runs serious models locally, privately, on your own machine.<p>Here's what should make model builders pay attention: there's going to be a full year between this generation and the next. A year where the new chip is the ceiling. A year where &quot;fits on the latest consumer silicon&quot; is the line between usable and irrelevant.<p>Something has shifted. People don't just want to use AI; they want to own their AI. Models on their hardware, data on their machine. No API costs. No rate limits. No terms of service that change overnight. Ollama, LM Studio, llama.cpp, OpenClaw; these aren't niche experiments anymore. They're how a growing segment of technical users interact with AI every day. And every single one is constrained by the same thing: how much model fits in memory.<p>This matters even more for social impact organizations. NGOs and humanitarian teams often work in low-connectivity environments with sensitive data; refugee records, health information, disaster response intel. Sending that to a cloud API isn't just inconvenient, it's a non-starter. A model that runs on a consumer laptop means an aid worker in a field office with no internet still gets AI assistance, privately, on hardware their grant budget can actually afford.<p>If your model only runs well on an <em>H100</em> cluster, you've made a choice. Maybe the right one. But you've also made yourself invisible to every person with a high-end laptop who wants to run it at a coffee shop, or every nonprofit that can't justify cloud compute costs.<p>The teams that win the local AI race will treat consumer hardware constraints as a design target, not an afterthought:<p>1. Quantization-first thinking. Not &quot;can we quantize it later?&quot; but &quot;what's the best model we can build that fits in 48GB unified memory at Q4?&quot;<p>2. Architecture choices that favor <em>inference</em> on consumer silicon. Not every architecture runs well on today's consumer GPU frameworks. The ones that do will have an unfair advantage.<p>3. Benchmarking on real hardware. Not <em>A100</em> throughput numbers that mean nothing to someone on an ultrabook.<p>Say the next-gen Pro chip tops out at 48GB unified memory. Factor in OS overhead, context window, and KV cache; you're looking at 35-38GB usable. That's your target. The model that delivers the best quality within that envelope, with fast <em>inference</em> and real-world usability, becomes the default local model for millions of users. For a full year. That's not a technical milestone. That's a market position.<p>To every model maker reading this, especially in open source:<p>Find out the next-gen chip's memory ceiling. Build your best model to fit inside it. Make it sing on consumer unified-memory hardware. The people who do this will own the local AI market for the next year; at this pace, that's like three years in <em>202</em>0. The people who don't will wonder why nobody's downloading their model. Pair it with something like OpenClaw and you've got a product people actually want.<p>Build for the hardware people actually own."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"An Open Letter to Model Makers"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://abediaz.substack.com/p/the-next-gen-cpu-ceiling-an-open"}},"_tags":["comment","author_abediaz","story_47033215"],"author":"abediaz","comment_text":"There&#x27;s a number that should be on every model builder&#x27;s whiteboard right now, and almost nobody is talking about it:<p>The maximum model size that fits on the next generation of consumer unified-memory chips.<p>When the leading consumer silicon vendor drops its next lineup (and it&#x27;s coming soon), millions of developers and power users are going to buy in. Not for the marketing. Because these chips offer something no cloud GPU can: unified memory that runs serious models locally, privately, on your own machine.<p>Here&#x27;s what should make model builders pay attention: there&#x27;s going to be a full year between this generation and the next. A year where the new chip is the ceiling. A year where &quot;fits on the latest consumer silicon&quot; is the line between usable and irrelevant.<p>Something has shifted. People don&#x27;t just want to use AI; they want to own their AI. Models on their hardware, data on their machine. No API costs. No rate limits. No terms of service that change overnight. Ollama, LM Studio, llama.cpp, OpenClaw; these aren&#x27;t niche experiments anymore. They&#x27;re how a growing segment of technical users interact with AI every day. And every single one is constrained by the same thing: how much model fits in memory.<p>This matters even more for social impact organizations. NGOs and humanitarian teams often work in low-connectivity environments with sensitive data; refugee records, health information, disaster response intel. Sending that to a cloud API isn&#x27;t just inconvenient, it&#x27;s a non-starter. A model that runs on a consumer laptop means an aid worker in a field office with no internet still gets AI assistance, privately, on hardware their grant budget can actually afford.<p>If your model only runs well on an H100 cluster, you&#x27;ve made a choice. Maybe the right one. But you&#x27;ve also made yourself invisible to every person with a high-end laptop who wants to run it at a coffee shop, or every nonprofit that can&#x27;t justify cloud compute costs.<p>The teams that win the local AI race will treat consumer hardware constraints as a design target, not an afterthought:<p>1. Quantization-first thinking. Not &quot;can we quantize it later?&quot; but &quot;what&#x27;s the best model we can build that fits in 48GB unified memory at Q4?&quot;<p>2. Architecture choices that favor inference on consumer silicon. Not every architecture runs well on today&#x27;s consumer GPU frameworks. The ones that do will have an unfair advantage.<p>3. Benchmarking on real hardware. Not A100 throughput numbers that mean nothing to someone on an ultrabook.<p>Say the next-gen Pro chip tops out at 48GB unified memory. Factor in OS overhead, context window, and KV cache; you&#x27;re looking at 35-38GB usable. That&#x27;s your target. The model that delivers the best quality within that envelope, with fast inference and real-world usability, becomes the default local model for millions of users. For a full year. That&#x27;s not a technical milestone. That&#x27;s a market position.<p>To every model maker reading this, especially in open source:<p>Find out the next-gen chip&#x27;s memory ceiling. Build your best model to fit inside it. Make it sing on consumer unified-memory hardware. The people who do this will own the local AI market for the next year; at this pace, that&#x27;s like three years in 2020. The people who don&#x27;t will wonder why nobody&#x27;s downloading their model. Pair it with something like OpenClaw and you&#x27;ve got a product people actually want.<p>Build for the hardware people actually own.","created_at":"2026-02-16T10:12:31Z","created_at_i":1771236751,"objectID":"47033216","parent_id":47033215,"story_id":47033215,"story_title":"An Open Letter to Model Makers","story_url":"https://abediaz.substack.com/p/the-next-gen-cpu-ceiling-an-open","updated_at":"2026-03-05T23:33:37Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"lhl"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["a100","h100","inference","2026"],"value":"I published this as a comment as well, but it's probably worth nothing that the ChatGPT water/power numbers cited (the one that is most widely cited in these discussions) comes from an April <em>202</em>3 paper (Li et al, arXiv:2304.03271) that estimates water/power usage based off of GPT-3 (175B dense model) numbers published from OpenAI's original <i>GPT-3 <em>202</em>1 paper</i>. From Section 3.3.2 <em>Inference</em>:<p>&gt; As a representative usage scenario for an LLM, we consider a conversation task, which typically includes a CPU-intensive prompt phase that processes the user\u2019s input (a.k.a., prompt) and a memory-intensive token phase that produces outputs [37]. More specifically, we consider a medium-sized request, each with approximately \u2264800 words of input and 150 \u2013 300 words of output [37]. The official estimate shows that GPT-3 consumes an order of 0.4 kWh electricity to generate 100 pages of content (e.g., roughly 0.004 kWh per page) [18]. Thus, we consider 0.004 kWh as the per-request server energy consumption for our conversation task. The PUE, WUE, and EWIF are the same as those used for estimating the training water consumption.<p>There is a slightly newer paper (Oct <em>202</em>3) that directly measured power usage on a Llama 65B (on V100/<em>A100</em> hardware) that showed a 14X better efficiency. [2] Ethan Mollick linked to it recently and got me curious since I've recently been running my own <em>inference</em> (performance) testing and it'd be easy enough to just calculate power usage. My results [3] on the latest stable vLLM from last week on a standard <em>H100</em> node w/ Llama 3.3 70B FP8 was almost a 10X better token/joule than the <em>202</em>3 V100/<em>A100</em> testing, which seems about right to me. This is without fancy look-ahead, speculative decode, prefix caching taken into account, just raw token generation. This is 120X more efficient than the commonly cited &quot;ChatGPT&quot; numbers and 250X more efficient than the Llama-3-70B numbers cited in the latest version (v4, <em>202</em>5-01-15) of that same paper.<p>For those interested in a full analysis/table with all the citations (including my full testing results) see this o1 chat that calculated the relative efficiency differences and made a nice results table for me: <a href=\"https://chatgpt.com/share/678b55bb-336c-8012-97cc-b94f70919daa\" rel=\"nofollow\">https://chatgpt.com/share/678b55bb-336c-8012-97cc-b94f70919d...</a><p>(It's worth point out that that used 45s of TTC, which is a point that is not lost on me!)<p>[1] <a href=\"https://arxiv.org/abs/2304.03271\" rel=\"nofollow\">https://arxiv.org/abs/2304.03271</a><p>[2] <a href=\"https://arxiv.org/abs/2310.03003\" rel=\"nofollow\">https://arxiv.org/abs/2310.03003</a><p>[3] <a href=\"https://gist.github.com/lhl/bf81a9c7dfc4244c974335e1605dcf22\" rel=\"nofollow\">https://gist.github.com/lhl/bf81a9c7dfc4244c974335e1605dcf22</a>"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Using ChatGPT is not bad for the environment"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://andymasley.substack.com/p/individual-ai-use-is-not-bad-for"}},"_tags":["comment","author_lhl","story_42745847"],"author":"lhl","comment_text":"I published this as a comment as well, but it&#x27;s probably worth nothing that the ChatGPT water&#x2F;power numbers cited (the one that is most widely cited in these discussions) comes from an April 2023 paper (Li et al, arXiv:2304.03271) that estimates water&#x2F;power usage based off of GPT-3 (175B dense model) numbers published from OpenAI&#x27;s original <i>GPT-3 2021 paper</i>. From Section 3.3.2 Inference:<p>&gt; As a representative usage scenario for an LLM, we consider a conversation task, which typically includes a CPU-intensive prompt phase that processes the user\u2019s input (a.k.a., prompt) and a memory-intensive token phase that produces outputs [37]. More specifically, we consider a medium-sized request, each with approximately \u2264800 words of input and 150 \u2013 300 words of output [37]. The official estimate shows that GPT-3 consumes an order of 0.4 kWh electricity to generate 100 pages of content (e.g., roughly 0.004 kWh per page) [18]. Thus, we consider 0.004 kWh as the per-request server energy consumption for our conversation task. The PUE, WUE, and EWIF are the same as those used for estimating the training water consumption.<p>There is a slightly newer paper (Oct 2023) that directly measured power usage on a Llama 65B (on V100&#x2F;A100 hardware) that showed a 14X better efficiency. [2] Ethan Mollick linked to it recently and got me curious since I&#x27;ve recently been running my own inference (performance) testing and it&#x27;d be easy enough to just calculate power usage. My results [3] on the latest stable vLLM from last week on a standard H100 node w&#x2F; Llama 3.3 70B FP8 was almost a 10X better token&#x2F;joule than the 2023 V100&#x2F;A100 testing, which seems about right to me. This is without fancy look-ahead, speculative decode, prefix caching taken into account, just raw token generation. This is 120X more efficient than the commonly cited &quot;ChatGPT&quot; numbers and 250X more efficient than the Llama-3-70B numbers cited in the latest version (v4, 2025-01-15) of that same paper.<p>For those interested in a full analysis&#x2F;table with all the citations (including my full testing results) see this o1 chat that calculated the relative efficiency differences and made a nice results table for me: <a href=\"https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;678b55bb-336c-8012-97cc-b94f70919daa\" rel=\"nofollow\">https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;678b55bb-336c-8012-97cc-b94f70919d...</a><p>(It&#x27;s worth point out that that used 45s of TTC, which is a point that is not lost on me!)<p>[1] <a href=\"https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2304.03271\" rel=\"nofollow\">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2304.03271</a><p>[2] <a href=\"https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2310.03003\" rel=\"nofollow\">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2310.03003</a><p>[3] <a href=\"https:&#x2F;&#x2F;gist.github.com&#x2F;lhl&#x2F;bf81a9c7dfc4244c974335e1605dcf22\" rel=\"nofollow\">https:&#x2F;&#x2F;gist.github.com&#x2F;lhl&#x2F;bf81a9c7dfc4244c974335e1605dcf22</a>","created_at":"2025-01-18T07:48:22Z","created_at_i":1737186502,"objectID":"42746644","parent_id":42745847,"story_id":42745847,"story_title":"Using ChatGPT is not bad for the environment","story_url":"https://andymasley.substack.com/p/individual-ai-use-is-not-bad-for","updated_at":"2025-01-20T14:00:06Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"kofdai"},"comment_text":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["a100","inference","2026"],"value":"Title: Clarification on my development workflow (re: jaen)<p>You\u2019re right to be skeptical of the speed, and I realize I was incomplete in describing my process. I should have been more transparent: I am using Claude Code as a &quot;pair-programmer&quot; to implement and review the logic I design.<p>While the Verantyx engine itself remains <em>a 100</em>% static, symbolic solver at test-time (no LLM calls during <em>inference</em>), the rapid score jumps from 20.1% to 22.4% are indeed accelerated by an AI-assisted workflow.<p>My role is to identify the geometric pattern in the failed tasks and design the DSL primitive (the &quot;what&quot;). I then use Claude Code to scaffold the implementation, check for regressions across the 1,000 tasks, and refine the code (the &quot;how&quot;).<p>This is why I can commit 30-80 lines of verified geometric logic in minutes rather than hours. The &quot;thinking&quot; and the &quot;logic design&quot; are human-led, but the &quot;implementation&quot; is AI-augmented.<p>My apologies if my previous comments made it sound like I was manually typing every single one of those 26K lines without help. In <em>2026</em>, I believe this &quot;Human-Architect / AI-Builder&quot; model is the most effective way to tackle benchmarks like ARC.<p>I\u2019d love to hear your thoughts on this hybrid approach to symbolic AI development."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"I beat Grok 4 on ARC-AGI-2 using a CPU-only symbolic engine (18.1% score)"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/Ag3497120/verantyx-v6"}},"_tags":["comment","author_kofdai","story_47147113"],"author":"kofdai","comment_text":"Title: Clarification on my development workflow (re: jaen)<p>You\u2019re right to be skeptical of the speed, and I realize I was incomplete in describing my process. I should have been more transparent: I am using Claude Code as a &quot;pair-programmer&quot; to implement and review the logic I design.<p>While the Verantyx engine itself remains a 100% static, symbolic solver at test-time (no LLM calls during inference), the rapid score jumps from 20.1% to 22.4% are indeed accelerated by an AI-assisted workflow.<p>My role is to identify the geometric pattern in the failed tasks and design the DSL primitive (the &quot;what&quot;). I then use Claude Code to scaffold the implementation, check for regressions across the 1,000 tasks, and refine the code (the &quot;how&quot;).<p>This is why I can commit 30-80 lines of verified geometric logic in minutes rather than hours. The &quot;thinking&quot; and the &quot;logic design&quot; are human-led, but the &quot;implementation&quot; is AI-augmented.<p>My apologies if my previous comments made it sound like I was manually typing every single one of those 26K lines without help. In 2026, I believe this &quot;Human-Architect &#x2F; AI-Builder&quot; model is the most effective way to tackle benchmarks like ARC.<p>I\u2019d love to hear your thoughts on this hybrid approach to symbolic AI development.","created_at":"2026-02-27T01:03:52Z","created_at_i":1772154232,"objectID":"47174874","parent_id":47164080,"story_id":47147113,"story_title":"I beat Grok 4 on ARC-AGI-2 using a CPU-only symbolic engine (18.1% score)","story_url":"https://github.com/Ag3497120/verantyx-v6","updated_at":"2026-03-05T23:38:12Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"keeda"},"comment_text":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["a100","inference","2026"],"value":"<i>&gt; These things are so hideously inefficient.</i><p>Quite the opposite, really. I did some napkin math for energy and water consumption, and compared to humans these things are very resource efficient.<p>If LLMs improve productivity by even 5% (studies actually peg productivity gains across various professions at 15 - 30%, and these are from <em>202</em>4!) the resource savings by accelerating all knowledge workers are significant.<p>Simplistically, during 8 hours of work a human would consume 10 kWH of electricity + 27 gallons of water. Sped up by 5%, that drops by 0.5kWH and 1.35 gallons. Even assuming a higher end of resources used by LLMs, <em>a 100</em> large prompts (~1 every 5 minutes) would only consume 0.25 kWH + 0.3 gallons. So we're still saving ~0.25 kWH + 1 gallon overall per day!<p>That is, humans + LLMs are way more efficient than humans alone. As such, the more knowledge workers adopt LLMs, the more efficiently they can achieve the same work output!<p>If we assume a conservative 10% productivity speed up, adoption across all ~100M knowledge work in the US will recoup the resource cost of a full training run in a few business days, even after accounting for the <em>inference</em> costs!<p>Additional reading with more useful numbers (independent of my napkin math):<p><a href=\"https://www.nature.com/articles/s41598-024-76682-6\" rel=\"nofollow\">https://www.nature.com/articles/s41598-024-76682-6</a><p><a href=\"https://cacm.acm.org/blogcacm/the-energy-footprint-of-humans-and-large-language-models/\" rel=\"nofollow\">https://cacm.acm.org/blogcacm/the-energy-footprint-of-humans...</a>"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Covering electricity price increases from our data centers"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://www.anthropic.com/news/covering-electricity-price-increases"}},"_tags":["comment","author_keeda","story_46981058"],"author":"keeda","children":[46984729,46984752,46985353,46989616],"comment_text":"<i>&gt; These things are so hideously inefficient.</i><p>Quite the opposite, really. I did some napkin math for energy and water consumption, and compared to humans these things are very resource efficient.<p>If LLMs improve productivity by even 5% (studies actually peg productivity gains across various professions at 15 - 30%, and these are from 2024!) the resource savings by accelerating all knowledge workers are significant.<p>Simplistically, during 8 hours of work a human would consume 10 kWH of electricity + 27 gallons of water. Sped up by 5%, that drops by 0.5kWH and 1.35 gallons. Even assuming a higher end of resources used by LLMs, a 100 large prompts (~1 every 5 minutes) would only consume 0.25 kWH + 0.3 gallons. So we&#x27;re still saving ~0.25 kWH + 1 gallon overall per day!<p>That is, humans + LLMs are way more efficient than humans alone. As such, the more knowledge workers adopt LLMs, the more efficiently they can achieve the same work output!<p>If we assume a conservative 10% productivity speed up, adoption across all ~100M knowledge work in the US will recoup the resource cost of a full training run in a few business days, even after accounting for the inference costs!<p>Additional reading with more useful numbers (independent of my napkin math):<p><a href=\"https:&#x2F;&#x2F;www.nature.com&#x2F;articles&#x2F;s41598-024-76682-6\" rel=\"nofollow\">https:&#x2F;&#x2F;www.nature.com&#x2F;articles&#x2F;s41598-024-76682-6</a><p><a href=\"https:&#x2F;&#x2F;cacm.acm.org&#x2F;blogcacm&#x2F;the-energy-footprint-of-humans-and-large-language-models&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;cacm.acm.org&#x2F;blogcacm&#x2F;the-energy-footprint-of-humans...</a>","created_at":"2026-02-12T03:42:35Z","created_at_i":1770867755,"objectID":"46984659","parent_id":46983926,"story_id":46981058,"story_title":"Covering electricity price increases from our data centers","story_url":"https://www.anthropic.com/news/covering-electricity-price-increases","updated_at":"2026-03-05T23:35:12Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"cdelsolar"},"comment_text":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["a100","inference","2026"],"value":"I wrote an article about it here a while back: <a href=\"https://cesardelsolar.com/posts/2022-02-13_scrabble-is-nowhere-close-to-a-solved-game/\" rel=\"nofollow\">https://cesardelsolar.com/posts/<em>202</em>2-02-13_scrabble-is-nowhe...</a><p>Having written what I believe is the best bot out there, since that article (BestBot on <a href=\"https://woogles.io\" rel=\"nofollow\">https://woogles.io</a>), it beats many top players around 55% of the time roughly, but still has so many fundamental issues that we have not solved yet. I believe if we matched it against the top player or two in the world that it would be roughly even. I would not put money on it beating Nigel Richards over <em>a 100</em>-game series. We have a while to go until we build a truly superhuman bot; this is not the case for Chess/Go.<p>As to what you're possibly missing here - it's that Scrabble is more than just about finding the top-scoring word. There's a lot of board shape considerations, <em>inferences</em>, volatility, etc that current engines, including mine, don't yet take into account."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"After AI beat them, professional Go players got better and more creative"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://www.henrikkarlsson.xyz/p/go"}},"_tags":["comment","author_cdelsolar","story_39972990"],"author":"cdelsolar","comment_text":"I wrote an article about it here a while back: <a href=\"https:&#x2F;&#x2F;cesardelsolar.com&#x2F;posts&#x2F;2022-02-13_scrabble-is-nowhere-close-to-a-solved-game&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;cesardelsolar.com&#x2F;posts&#x2F;2022-02-13_scrabble-is-nowhe...</a><p>Having written what I believe is the best bot out there, since that article (BestBot on <a href=\"https:&#x2F;&#x2F;woogles.io\" rel=\"nofollow\">https:&#x2F;&#x2F;woogles.io</a>), it beats many top players around 55% of the time roughly, but still has so many fundamental issues that we have not solved yet. I believe if we matched it against the top player or two in the world that it would be roughly even. I would not put money on it beating Nigel Richards over a 100-game series. We have a while to go until we build a truly superhuman bot; this is not the case for Chess&#x2F;Go.<p>As to what you&#x27;re possibly missing here - it&#x27;s that Scrabble is more than just about finding the top-scoring word. There&#x27;s a lot of board shape considerations, inferences, volatility, etc that current engines, including mine, don&#x27;t yet take into account.","created_at":"2024-04-10T18:54:44Z","created_at_i":1712775284,"objectID":"39994406","parent_id":39973973,"story_id":39972990,"story_title":"After AI beat them, professional Go players got better and more creative","story_url":"https://www.henrikkarlsson.xyz/p/go","updated_at":"2024-09-20T16:56:21Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"s_country"},"comment_text":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["a100","inference","2026"],"value":"Dear Engineer, Our Tech Lead Role in  Palo Alto is open! As Harmony\u2019s founder, I\u2019d like to tell you below about the project, your role, and what we are looking for from you.<p>Ideally you specialize in performance, security and systems \u2013 with 3 years of working experience, a degree in computer science, and regular practice of open-source development. You may have built a kernel, a compiler, or a database from scratch before; you enjoy Hacker News, functional idioms, and productivity tools every day.<p>This \u201cTech Lead\u201d role is full-time, permanent, onsite in our Palo Alto headquarters in California. The compensation, with a base salary and vesting equity, may range from $160K to $210K. We offer full benefits but not a working visa. You should live within a 30-minute commute to work in office every day; you must love the long work hours and the chaotic grind of startups over years.<p>We have 6 full-time teammates in office, 15 more across Europe and Asia. You will be working directly with the leadership team of myself, Casey, and Aaron. I lead the team in platform vision, market strategy, and product roadmap. Casey leads in protocol research, network system, and developer community. Aaron leads in backend architecture, product design, and user applications.<p>You may start with our current initiatives on AI bots, Telegram wallets, domain auctions\u2026 then explore our broader roadmap of decentralized protocols, financial products, or cryptographic primitives. Harmony\u2019s mission is \u201cto scale trust and create radically fair economy\u201d. Our next milestones include crypto payments for AI models and products on Telegram, social wallets without separate passwords or devices, and sustainable collectible auctions for .country domains.<p>Harmony is among the first proof-of-stake, sharding blockchains. Our 2-second transaction finality is one of the fastest; our 1000-slot validator network is one of the most diverse. The project started 5 years ago \u2013 launching the network and the ONE token in 2019, staking and delegation in <em>202</em>0, DeFi and NFT in <em>202</em>1, DAO and ZK Proofs in <em>202</em>2, and now Web3 domains and AI bots in <em>202</em>3. Harmony is venture-backed with its network utility tokens available on crypto exchanges and its treasury of more than 3 years of runway.<p>You will be challenged to launch features within the first few weeks, and soon to lead a team of 2 to 4 engineers. You will develop an open blockchain of hundreds of network nodes with hundreds of millions of assets. You will pioneer how applying generative models, low-rank adaptation and edge-device <em>inference</em> will bring productivity as well as entertainment to millions of users. You will be an industry leader, engaging in public forums or with our developer ecosystem, not only for your technical expertise but also your empathetic value.<p>Our backends, written in Go, run on Amazon and other clouds. Our middlewares, written in Solidity, deploy as Ethereum smart contracts. Our frontends, written in Javascript, deploy as React applications. As our commitment to open development and community blockchain, we share all code as public repos on Github and we share development meetings as public videos on YouTube.<p>You will first have a full-hour interview each with Li, Kushagra and myself in office. Li leads on operations, while Kushagra advises and consults on human resources. You will send in advance your past sample work of a full-page writing and <em>a 100</em>-line code snippet.The next steps are technical interviews with Aaron and Casey on Zoom, another day of onsite sessions, and finally a 2-day paid joint-work session. The decision and the offer usually come within 24 hours.<p>You may learn more from the long posts \u201cPerfect Match for Hiring\u201d and \u201cRadically Fair for 1.country\u201d on my Substack blog.s.country. I'm at s@harmony.one or Telegram @stephentse.<p><a href=\"https://.s.country/p/dear-engineer-our-tech-lead-role\" rel=\"nofollow noreferrer\">https://.s.country/p/dear-engineer-our-tech-lead-role</a><p><a href=\"https://.s.country/p/perfect-match-finding-a-partner-in\" rel=\"nofollow noreferrer\">https://.s.country/p/perfect-match-finding-a-partner-in</a>"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"[dead]"}},"_tags":["comment","author_s_country","story_37066026"],"author":"s_country","comment_text":"Dear Engineer, Our Tech Lead Role in  Palo Alto is open! As Harmony\u2019s founder, I\u2019d like to tell you below about the project, your role, and what we are looking for from you.<p>Ideally you specialize in performance, security and systems \u2013 with 3 years of working experience, a degree in computer science, and regular practice of open-source development. You may have built a kernel, a compiler, or a database from scratch before; you enjoy Hacker News, functional idioms, and productivity tools every day.<p>This \u201cTech Lead\u201d role is full-time, permanent, onsite in our Palo Alto headquarters in California. The compensation, with a base salary and vesting equity, may range from $160K to $210K. We offer full benefits but not a working visa. You should live within a 30-minute commute to work in office every day; you must love the long work hours and the chaotic grind of startups over years.<p>We have 6 full-time teammates in office, 15 more across Europe and Asia. You will be working directly with the leadership team of myself, Casey, and Aaron. I lead the team in platform vision, market strategy, and product roadmap. Casey leads in protocol research, network system, and developer community. Aaron leads in backend architecture, product design, and user applications.<p>You may start with our current initiatives on AI bots, Telegram wallets, domain auctions\u2026 then explore our broader roadmap of decentralized protocols, financial products, or cryptographic primitives. Harmony\u2019s mission is \u201cto scale trust and create radically fair economy\u201d. Our next milestones include crypto payments for AI models and products on Telegram, social wallets without separate passwords or devices, and sustainable collectible auctions for .country domains.<p>Harmony is among the first proof-of-stake, sharding blockchains. Our 2-second transaction finality is one of the fastest; our 1000-slot validator network is one of the most diverse. The project started 5 years ago \u2013 launching the network and the ONE token in 2019, staking and delegation in 2020, DeFi and NFT in 2021, DAO and ZK Proofs in 2022, and now Web3 domains and AI bots in 2023. Harmony is venture-backed with its network utility tokens available on crypto exchanges and its treasury of more than 3 years of runway.<p>You will be challenged to launch features within the first few weeks, and soon to lead a team of 2 to 4 engineers. You will develop an open blockchain of hundreds of network nodes with hundreds of millions of assets. You will pioneer how applying generative models, low-rank adaptation and edge-device inference will bring productivity as well as entertainment to millions of users. You will be an industry leader, engaging in public forums or with our developer ecosystem, not only for your technical expertise but also your empathetic value.<p>Our backends, written in Go, run on Amazon and other clouds. Our middlewares, written in Solidity, deploy as Ethereum smart contracts. Our frontends, written in Javascript, deploy as React applications. As our commitment to open development and community blockchain, we share all code as public repos on Github and we share development meetings as public videos on YouTube.<p>You will first have a full-hour interview each with Li, Kushagra and myself in office. Li leads on operations, while Kushagra advises and consults on human resources. You will send in advance your past sample work of a full-page writing and a 100-line code snippet.The next steps are technical interviews with Aaron and Casey on Zoom, another day of onsite sessions, and finally a 2-day paid joint-work session. The decision and the offer usually come within 24 hours.<p>You may learn more from the long posts \u201cPerfect Match for Hiring\u201d and \u201cRadically Fair for 1.country\u201d on my Substack blog.s.country. I&#x27;m at s@harmony.one or Telegram @stephentse.<p><a href=\"https:&#x2F;&#x2F;.s.country&#x2F;p&#x2F;dear-engineer-our-tech-lead-role\" rel=\"nofollow noreferrer\">https:&#x2F;&#x2F;.s.country&#x2F;p&#x2F;dear-engineer-our-tech-lead-role</a><p><a href=\"https:&#x2F;&#x2F;.s.country&#x2F;p&#x2F;perfect-match-finding-a-partner-in\" rel=\"nofollow noreferrer\">https:&#x2F;&#x2F;.s.country&#x2F;p&#x2F;perfect-match-finding-a-partner-in</a>","created_at":"2023-08-09T17:33:47Z","created_at_i":1691602427,"objectID":"37066027","parent_id":37066026,"story_id":37066026,"story_title":"[dead]","updated_at":"2024-09-20T14:51:48Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"joshjob42"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["a100","h100","inference","2026"],"value":"I think generally the expectation is that there are around 100T synapses in the brain, and of course it's probably not a 1:1 correspondence with neural networks, but it doesn't seem infeasible at all to me that a dense-equivalent 100T parameter model would be able to rival the best humans if trained properly.<p>If basically a transformer, that means it needs at <em>inference</em> time ~200T flops per token. The paper assumes humans &quot;think&quot; at ~15 tokens/second which is about 10 words, similar to the reading speed of a college graduate. So that would be ~3 petaflops of compute per second.<p>Assuming that's fp8, an <em>H100</em> could do ~4 petaflops, and the authors of AI <em>202</em>7 guesstimate that purpose wafer scale <em>inference</em> chips circa late <em>202</em>7 should be able to do ~400petaflops for <em>inference</em>, ~<em>100</em> H100s worth, for ~$600k each for fabrication and installation into a datacenter.<p>Rounding that basically means ~$6k would buy you the compute to &quot;think&quot; at 10 words/second. Generally speaking that'd probably work out to maybe $3k/yr after depreciation and electricity costs, or ~30-50\u00a2/hr of &quot;human thought equivalent&quot; 10 words/second. Running an AI at 50x human speed 24/7 would cost ~$23k/yr, so 1 OpenBrain researcher's salary could give them a team of ~10-20 such AIs running flat out all the time. Even if you think the AI would need an &quot;extra&quot; 10 or even 100x in terms of tokens/second to match humans, that still puts you at genius level AIs in principle runnable at human speed for 0.1 to 1x the median US income.<p>There's an open question whether training such a model is feasible in a few years, but the raw compute capability at the chip level to plausibly run a model that large at enormous speed at low cost is already existent (at the street price of B200's it'd cost ~$2-4/hr-human-equivalent)."},"story_title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["2026"],"value":"AI <em>202</em>7"},"story_url":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["2026"],"value":"https://ai-<em>202</em>7.com/"}},"_tags":["comment","author_joshjob42","story_43571851"],"author":"joshjob42","children":[43581156],"comment_text":"I think generally the expectation is that there are around 100T synapses in the brain, and of course it&#x27;s probably not a 1:1 correspondence with neural networks, but it doesn&#x27;t seem infeasible at all to me that a dense-equivalent 100T parameter model would be able to rival the best humans if trained properly.<p>If basically a transformer, that means it needs at inference time ~200T flops per token. The paper assumes humans &quot;think&quot; at ~15 tokens&#x2F;second which is about 10 words, similar to the reading speed of a college graduate. So that would be ~3 petaflops of compute per second.<p>Assuming that&#x27;s fp8, an H100 could do ~4 petaflops, and the authors of AI 2027 guesstimate that purpose wafer scale inference chips circa late 2027 should be able to do ~400petaflops for inference, ~100 H100s worth, for ~$600k each for fabrication and installation into a datacenter.<p>Rounding that basically means ~$6k would buy you the compute to &quot;think&quot; at 10 words&#x2F;second. Generally speaking that&#x27;d probably work out to maybe $3k&#x2F;yr after depreciation and electricity costs, or ~30-50\u00a2&#x2F;hr of &quot;human thought equivalent&quot; 10 words&#x2F;second. Running an AI at 50x human speed 24&#x2F;7 would cost ~$23k&#x2F;yr, so 1 OpenBrain researcher&#x27;s salary could give them a team of ~10-20 such AIs running flat out all the time. Even if you think the AI would need an &quot;extra&quot; 10 or even 100x in terms of tokens&#x2F;second to match humans, that still puts you at genius level AIs in principle runnable at human speed for 0.1 to 1x the median US income.<p>There&#x27;s an open question whether training such a model is feasible in a few years, but the raw compute capability at the chip level to plausibly run a model that large at enormous speed at low cost is already existent (at the street price of B200&#x27;s it&#x27;d cost ~$2-4&#x2F;hr-human-equivalent).","created_at":"2025-04-04T04:28:46Z","created_at_i":1743740926,"objectID":"43578329","parent_id":43577330,"story_id":43571851,"story_title":"AI 2027","story_url":"https://ai-2027.com/","updated_at":"2025-04-06T16:59:45Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"jmyeet"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["a100","h100","inference","2026"],"value":"This is a good article. I've been curious about how this is going to play out. A couple of data points:<p>1. An enthusiast had a project to get a V100 working on his PC [1]. This was a ~$10k GPU 10 years ago. It's now sold for scrap;<p>2. The <em>A100</em> came out in <em>202</em>0 and cannot run a large model like DeepSeek v4 Pro. It can run Flash. You need a 16xH100 cluster to run Pro and that's a ~4 year old GPU and AFAICT 8xB100 or 4xB200;<p>3. We're about to roll out R100/R200s.<p>I'm surprised that NVidia is moving to a 1 year product cycle (per this article) because the big question I've had is what's that going to do to existing investments in GPUs. Why? Because if 4xR100 can do the work of 32xH100 then that's a massive advantage in performance-per-Watt, which I think is going to be the only metric that ends up mattering.<p>In addition to raw power, new capabilities are developed and come online. For example, certain smaller, more efficient quantization methods just didn't exist on older hardware.<p>Oh, another thought from this: a 9% annual failure rate just goes to show you how ridiculous the idea of orbital data centers really is. Orbital DCs were always just a pump-and-dump scheme for SpaceX's IPO.<p>Currently it gets expensive to run models larger than ~31B locally. You start to need some pretty expensive hardware. That's going to change. I don't expect we'll be running 1T+ models on a Macbook Pro within 5 years (at reasonable <em>inference</em> rates) but I think people today will be shocked at what's being run locally in 5 years and that'll easily be <em>100</em>-200B+ models.<p>[1]: <a href=\"https://www.hackster.io/news/hacking-a-server-grade-nvidia-gpu-into-a-home-desktop-674ada7032a7\" rel=\"nofollow\">https://www.hackster.io/news/hacking-a-server-grade-nvidia-g...</a>"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Nobody knows what a used GPU cluster is worth"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://ciphertalk.substack.com/p/nobody-knows-what-a-used-gpu-cluster"}},"_tags":["comment","author_jmyeet","story_48917135"],"author":"jmyeet","comment_text":"This is a good article. I&#x27;ve been curious about how this is going to play out. A couple of data points:<p>1. An enthusiast had a project to get a V100 working on his PC [1]. This was a ~$10k GPU 10 years ago. It&#x27;s now sold for scrap;<p>2. The A100 came out in 2020 and cannot run a large model like DeepSeek v4 Pro. It can run Flash. You need a 16xH100 cluster to run Pro and that&#x27;s a ~4 year old GPU and AFAICT 8xB100 or 4xB200;<p>3. We&#x27;re about to roll out R100&#x2F;R200s.<p>I&#x27;m surprised that NVidia is moving to a 1 year product cycle (per this article) because the big question I&#x27;ve had is what&#x27;s that going to do to existing investments in GPUs. Why? Because if 4xR100 can do the work of 32xH100 then that&#x27;s a massive advantage in performance-per-Watt, which I think is going to be the only metric that ends up mattering.<p>In addition to raw power, new capabilities are developed and come online. For example, certain smaller, more efficient quantization methods just didn&#x27;t exist on older hardware.<p>Oh, another thought from this: a 9% annual failure rate just goes to show you how ridiculous the idea of orbital data centers really is. Orbital DCs were always just a pump-and-dump scheme for SpaceX&#x27;s IPO.<p>Currently it gets expensive to run models larger than ~31B locally. You start to need some pretty expensive hardware. That&#x27;s going to change. I don&#x27;t expect we&#x27;ll be running 1T+ models on a Macbook Pro within 5 years (at reasonable inference rates) but I think people today will be shocked at what&#x27;s being run locally in 5 years and that&#x27;ll easily be 100-200B+ models.<p>[1]: <a href=\"https:&#x2F;&#x2F;www.hackster.io&#x2F;news&#x2F;hacking-a-server-grade-nvidia-gpu-into-a-home-desktop-674ada7032a7\" rel=\"nofollow\">https:&#x2F;&#x2F;www.hackster.io&#x2F;news&#x2F;hacking-a-server-grade-nvidia-g...</a>","created_at":"2026-07-22T21:31:40Z","created_at_i":1784755900,"objectID":"49013750","parent_id":48917135,"story_id":48917135,"story_title":"Nobody knows what a used GPU cluster is worth","story_url":"https://ciphertalk.substack.com/p/nobody-knows-what-a-used-gpu-cluster","updated_at":"2026-07-23T12:56:19Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"pella"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["a100","h100","inference","2026"],"value":"<a href=\"https://x.com/SemiAnalysis_/status/1953206064486519238\" rel=\"nofollow\">https://x.com/SemiAnalysis_/status/1953206064486519238</a><p><pre><code>  &quot;AMD\u2019s Q2 CY2025 earnings came in below expectations, but this was mostly expected. For the past six months, we have been highlighting that demand for the MI325X was soft, the launch was delayed and ended up landing a couple of months after Nvidia\u2019s B200, even though MI325X was initially positioned as an <em>H200</em> competitor.  The MI355X was not broadly available during the quarter either, so it did not move the revenue needle much. Looking ahead, MI355X should be reasonably competitive with B200 on <em>inference</em> performance per TCO, but it will not match GB200, the MI355X scales to a world size of just 8, while GB200 goes up to 72.   We expect AMD to close much of the system-level hardware gap by late <em>2026</em> or early 2027 with MI400 UALoE72.  On the software side, there are massive improvements across training and inferencing since our December 2024 article, though there is still a long road ahead. AMD needs to close gaps around production-ready multi-node disaggregated prefill inferencing, WideEP multi-node <em>inference</em>, ROCm support for DeepEP MoE dispatch, and cleaning up the <em>100</em>+ unit tests in PyTorch currently tagged with @skipIfRocm or @cudaOnly. AMD\u2019s AI lead, @AnushElangovan, is actively working on this every day. Instead of previous stock buybacks, we expect AMD to continue increasing Opex as it reinvests in AI, growing software headcount, boosting talent compensation, and renting back MI325X/MI355X clusters from GPU cloud partners like OCI, Azure, TensorWave, DigitalOcean, Vultr, and Crusoe for internal R&amp;D.&quot;</code></pre>"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Ask HN: Is AMD's AI making progress against Nvidia?"}},"_tags":["comment","author_pella","story_44835374"],"author":"pella","comment_text":"<a href=\"https:&#x2F;&#x2F;x.com&#x2F;SemiAnalysis_&#x2F;status&#x2F;1953206064486519238\" rel=\"nofollow\">https:&#x2F;&#x2F;x.com&#x2F;SemiAnalysis_&#x2F;status&#x2F;1953206064486519238</a><p><pre><code>  &quot;AMD\u2019s Q2 CY2025 earnings came in below expectations, but this was mostly expected. For the past six months, we have been highlighting that demand for the MI325X was soft, the launch was delayed and ended up landing a couple of months after Nvidia\u2019s B200, even though MI325X was initially positioned as an H200 competitor.  The MI355X was not broadly available during the quarter either, so it did not move the revenue needle much. Looking ahead, MI355X should be reasonably competitive with B200 on inference performance per TCO, but it will not match GB200, the MI355X scales to a world size of just 8, while GB200 goes up to 72.   We expect AMD to close much of the system-level hardware gap by late 2026 or early 2027 with MI400 UALoE72.  On the software side, there are massive improvements across training and inferencing since our December 2024 article, though there is still a long road ahead. AMD needs to close gaps around production-ready multi-node disaggregated prefill inferencing, WideEP multi-node inference, ROCm support for DeepEP MoE dispatch, and cleaning up the 100+ unit tests in PyTorch currently tagged with @skipIfRocm or @cudaOnly. AMD\u2019s AI lead, @AnushElangovan, is actively working on this every day. Instead of previous stock buybacks, we expect AMD to continue increasing Opex as it reinvests in AI, growing software headcount, boosting talent compensation, and renting back MI325X&#x2F;MI355X clusters from GPU cloud partners like OCI, Azure, TensorWave, DigitalOcean, Vultr, and Crusoe for internal R&amp;D.&quot;</code></pre>","created_at":"2025-08-08T12:55:38Z","created_at_i":1754657738,"objectID":"44836442","parent_id":44835374,"story_id":44835374,"story_title":"Ask HN: Is AMD's AI making progress against Nvidia?","updated_at":"2026-03-05T22:30:36Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"oskarkk"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["a100","h100","inference","2026"],"value":"&gt; While xAI's model isn't a substantial leap beyond Deepseek R1, it utilizes <em>100</em> times more compute.<p>I'm not sure if it's close to 100x more. xAI had 100K Nvidia H100s, while this is what SemiAnalysis writes about DeepSeek:<p>&gt; We believe they have access to around 50,000 Hopper GPUs, which is not the same as 50,000 <em>H100</em>, as some have claimed. There are different variations of the <em>H100</em> that Nvidia made in compliance to different regulations (H800, H20), with only the H20 being currently available to Chinese model providers today. Note that H800s have the same computational power as H100s, but lower network bandwidth.<p>&gt; We believe DeepSeek has access to around 10,000 of these H800s and about 10,000 H100s. Furthermore they have orders for many more H20\u2019s, with Nvidia having produced over 1 million of the China specific GPU in the last 9 months. These GPUs are shared between High-Flyer and DeepSeek and geographically distributed to an extent. They are used for trading, <em>inference</em>, training, and research. For more specific detailed analysis, please refer to our Accelerator Model.<p>&gt; Our analysis shows that the total server CapEx for DeepSeek is ~$1.6B, with a considerable cost of $944M associated with operating such clusters. Similarly, all AI Labs and Hyperscalers have many more GPUs for various tasks including research and training then they they commit to an individual training run due to centralization of resources being a challenge. X.AI is unique as an AI lab with all their GPUs in 1 location.<p><a href=\"https://semianalysis.com/2025/01/31/deepseek-debates/\" rel=\"nofollow\">https://semianalysis.com/<em>202</em>5/01/31/deepseek-debates/</a><p>I don't know how much slower are these GPUs that they have, but if they have 50K of them, that doesn't sound like 100x less compute to me. Also, a company that has N GPUs and trains AI on them for 2 months can achieve the same results as a company that has 2N GPUs and trains for 1 month. So DeepSeek could spend a longer time training to offset the fact that have less GPUs than competitors."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Grok 3: Another win for the bitter lesson"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://www.thealgorithmicbridge.com/p/grok-3-another-win-for-the-bitter"}},"_tags":["comment","author_oskarkk","story_43111963"],"author":"oskarkk","children":[43117987],"comment_text":"&gt; While xAI&#x27;s model isn&#x27;t a substantial leap beyond Deepseek R1, it utilizes 100 times more compute.<p>I&#x27;m not sure if it&#x27;s close to 100x more. xAI had 100K Nvidia H100s, while this is what SemiAnalysis writes about DeepSeek:<p>&gt; We believe they have access to around 50,000 Hopper GPUs, which is not the same as 50,000 H100, as some have claimed. There are different variations of the H100 that Nvidia made in compliance to different regulations (H800, H20), with only the H20 being currently available to Chinese model providers today. Note that H800s have the same computational power as H100s, but lower network bandwidth.<p>&gt; We believe DeepSeek has access to around 10,000 of these H800s and about 10,000 H100s. Furthermore they have orders for many more H20\u2019s, with Nvidia having produced over 1 million of the China specific GPU in the last 9 months. These GPUs are shared between High-Flyer and DeepSeek and geographically distributed to an extent. They are used for trading, inference, training, and research. For more specific detailed analysis, please refer to our Accelerator Model.<p>&gt; Our analysis shows that the total server CapEx for DeepSeek is ~$1.6B, with a considerable cost of $944M associated with operating such clusters. Similarly, all AI Labs and Hyperscalers have many more GPUs for various tasks including research and training then they they commit to an individual training run due to centralization of resources being a challenge. X.AI is unique as an AI lab with all their GPUs in 1 location.<p><a href=\"https:&#x2F;&#x2F;semianalysis.com&#x2F;2025&#x2F;01&#x2F;31&#x2F;deepseek-debates&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;semianalysis.com&#x2F;2025&#x2F;01&#x2F;31&#x2F;deepseek-debates&#x2F;</a><p>I don&#x27;t know how much slower are these GPUs that they have, but if they have 50K of them, that doesn&#x27;t sound like 100x less compute to me. Also, a company that has N GPUs and trains AI on them for 2 months can achieve the same results as a company that has 2N GPUs and trains for 1 month. So DeepSeek could spend a longer time training to offset the fact that have less GPUs than competitors.","created_at":"2025-02-20T16:23:23Z","created_at_i":1740068603,"objectID":"43116618","parent_id":43112235,"story_id":43111963,"story_title":"Grok 3: Another win for the bitter lesson","story_url":"https://www.thealgorithmicbridge.com/p/grok-3-another-win-for-the-bitter","updated_at":"2025-10-06T20:25:29Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"cbosi"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["a100","h100","inference","2026"],"value":"MUNICH, Germany and PARIS, France, July 9th, <em>202</em>4. Genesis Cloud and Photoroom are excited to announce a partnership that will revolutionize AI-powered photo editing. Since its launch in 2019, Photoroom has achieved over 150 million downloads and maintains a 4.8-star rating from over a million reviews. By utilizing Genesis Cloud's top-tier infrastructure, Photoroom can significantly improve the efficiency, accuracy, and operational performance of their AI-powered photo editing services.\n&quot;We are thrilled to partner with Photoroom, a leader in AI photo editing. This collaboration showcases the strength of Genesis Cloud\u2019s high-performance cloud infrastructure in supporting advanced AI training workloads,&quot; said Dr. Stefan Schiefer, CEO of Genesis Cloud.\nThe partnership offers several benefits:\nBest-in-Class Compute: Best-in-class computational capabilities through a multi-node cluster of NVIDIA HGX <em>H100</em> GPUs, 3.2 Tbps InfiniBand networking, and high-speed file storage from VAST Data.\nExpert Technical Support: A direct line of contact with Genesis engineers ensures fast and high-quality support to minimize setup time and maintain reliable operations at scale.\nReduced Carbon Footprint: Genesis Cloud\u2019s data centers run <em>100</em>% on green energy, helping Photoroom avoid 1300 tons of CO2 per year compared to legacy providers.\nImportantly, Genesis Cloud decided not to rely on carbon offsetting programs, focusing instead on direct connection of the data center to renewable energy. This aligns perfectly with Photoroom\u2019s commitment to sustainability.\n\u201cIt\u2019s important to be honest about the fact that AI models burn a huge amount of energy, so we can tackle this issue head-on. One reason why we wanted to work with Genesis Cloud is that Photoroom believes we can positively impact the AI industry\u2019s energy mix by choosing the right providers\u201d, said Eliot Andres, CTO at Photoroom.\n\u201cPartnering with Genesis Cloud has enabled us at Photoroom to achieve top performance in training our models and significantly enhance our computational capabilities,\u201d added Andres. \u201cThe direct line of contact with Genesis\u2019s competent engineering team has been tremendously useful during the setup and operation of the GPU cluster.\u201d\nThis partnership highlights the commitment of both companies to innovation and sustainability. With this collaboration, users can expect stunning AI photo editing features, reinforcing Photoroom's position as a leader in the AI photo editing industry.<p>About Genesis Cloud\nGenesis Cloud is pioneering best-in-class GPU cloud solutions at scale for EnterpriseAI, GenAI, ML workloads &amp; rendering since 2018. Headquartered in Europe, we're building the next generation public cloud tailored to the needs of modern AI training and <em>inference</em>. Over 20\u2019000 customers utilize our services in our modern, green data centers to enjoy high performance with minimal environmental impact. Our mission is to help AI startups grow and empower businesses to harness the full potential of their data. Genesis enables AI innovation with best-in-class, green and affordable cloud infrastructure.<p>For more information on Genesis Cloud, visit www.genesiscloud.com<p>About Photoroom\nPhotoroom was founded in 2019, and over the past 4 years has carved out a niche in the commerce photography space. Photoroom first found success with its best-in-class background remover. It has now expanded its offering to include a batch photo editor, and generative AI offerings: AI Backgrounds and AI Shadows. Processing over 5 billion images a year, and downloaded over 150 million times, Photoroom is now the world's #1 AI photo-editing app, available across mobile, web and via an API in over 180 countries. Photoroom is headquartered in Paris with a global team of over 50 employees.<p>For more information on Photoroom, visit www.photoroom.com"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"[dead]"}},"_tags":["comment","author_cbosi","story_40913666"],"author":"cbosi","comment_text":"MUNICH, Germany and PARIS, France, July 9th, 2024. Genesis Cloud and Photoroom are excited to announce a partnership that will revolutionize AI-powered photo editing. Since its launch in 2019, Photoroom has achieved over 150 million downloads and maintains a 4.8-star rating from over a million reviews. By utilizing Genesis Cloud&#x27;s top-tier infrastructure, Photoroom can significantly improve the efficiency, accuracy, and operational performance of their AI-powered photo editing services.\n&quot;We are thrilled to partner with Photoroom, a leader in AI photo editing. This collaboration showcases the strength of Genesis Cloud\u2019s high-performance cloud infrastructure in supporting advanced AI training workloads,&quot; said Dr. Stefan Schiefer, CEO of Genesis Cloud.\nThe partnership offers several benefits:\nBest-in-Class Compute: Best-in-class computational capabilities through a multi-node cluster of NVIDIA HGX H100 GPUs, 3.2 Tbps InfiniBand networking, and high-speed file storage from VAST Data.\nExpert Technical Support: A direct line of contact with Genesis engineers ensures fast and high-quality support to minimize setup time and maintain reliable operations at scale.\nReduced Carbon Footprint: Genesis Cloud\u2019s data centers run 100% on green energy, helping Photoroom avoid 1300 tons of CO2 per year compared to legacy providers.\nImportantly, Genesis Cloud decided not to rely on carbon offsetting programs, focusing instead on direct connection of the data center to renewable energy. This aligns perfectly with Photoroom\u2019s commitment to sustainability.\n\u201cIt\u2019s important to be honest about the fact that AI models burn a huge amount of energy, so we can tackle this issue head-on. One reason why we wanted to work with Genesis Cloud is that Photoroom believes we can positively impact the AI industry\u2019s energy mix by choosing the right providers\u201d, said Eliot Andres, CTO at Photoroom.\n\u201cPartnering with Genesis Cloud has enabled us at Photoroom to achieve top performance in training our models and significantly enhance our computational capabilities,\u201d added Andres. \u201cThe direct line of contact with Genesis\u2019s competent engineering team has been tremendously useful during the setup and operation of the GPU cluster.\u201d\nThis partnership highlights the commitment of both companies to innovation and sustainability. With this collaboration, users can expect stunning AI photo editing features, reinforcing Photoroom&#x27;s position as a leader in the AI photo editing industry.<p>About Genesis Cloud\nGenesis Cloud is pioneering best-in-class GPU cloud solutions at scale for EnterpriseAI, GenAI, ML workloads &amp; rendering since 2018. Headquartered in Europe, we&#x27;re building the next generation public cloud tailored to the needs of modern AI training and inference. Over 20\u2019000 customers utilize our services in our modern, green data centers to enjoy high performance with minimal environmental impact. Our mission is to help AI startups grow and empower businesses to harness the full potential of their data. Genesis enables AI innovation with best-in-class, green and affordable cloud infrastructure.<p>For more information on Genesis Cloud, visit www.genesiscloud.com<p>About Photoroom\nPhotoroom was founded in 2019, and over the past 4 years has carved out a niche in the commerce photography space. Photoroom first found success with its best-in-class background remover. It has now expanded its offering to include a batch photo editor, and generative AI offerings: AI Backgrounds and AI Shadows. Processing over 5 billion images a year, and downloaded over 150 million times, Photoroom is now the world&#x27;s #1 AI photo-editing app, available across mobile, web and via an API in over 180 countries. Photoroom is headquartered in Paris with a global team of over 50 employees.<p>For more information on Photoroom, visit www.photoroom.com","created_at":"2024-07-09T08:15:02Z","created_at_i":1720512902,"objectID":"40913667","parent_id":40913666,"story_id":40913666,"story_title":"[dead]","updated_at":"2024-09-20T17:22:49Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"kkielhofner"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["a100","h100","inference","2026"],"value":"Article is hit and miss.<p>Mostly correct about <em>inference</em>.<p>Snapchat has been running small ML on mobile devices for years for stuff like face filters, etc. Same for those features on iOS that do facial recognition to match pictures of friends with your contacts.<p>Start to pay attention and you\u2019ll realize your phone is doing things like object detection, classification, face tracking, voice assistant wake word detection, some on device speech and command recognition, and a wild array of other ML tasks.<p>LLMs are sucking all of the oxygen out of the room but the overwhelming majority of end-user use of AI isn\u2019t generative and won\u2019t be for a long time if ever. The article is correct in saying <em>inference</em>, <em>inference</em>, <em>inference</em>.<p>The future of <em>inference</em> is smaller application and use-case specific models deployed to edge. Many of these applications just don\u2019t work with the latency of networks to cloud. Imagine face tracking for a Snapchat filter if it involved streaming video to a cloud for <em>inference</em>. Yeah, not happening.<p>The hosting costs are also astronomical, big <em>inference</em> hardware is only getting harder to get and Nvidia only has so much manufacturing capacity.<p>Leave the <em>H100s</em> up to Meta, OpenAI, etc that are training massive multi-billion parameter LLMs from scratch. Or people renting them in small batches to do finetuning of \u201csmaller\u201d models, etc.<p>This is also getting chipped at - with the unified memory of Apple Silicon you can get the RAM of 2 A/<em>H100s</em> today for less than the cost of used 80GB <em>A100s</em>. With an entire computer (Mac Pro), new and under warranty.<p>Nvidia still wins on TFLOPS but expect M3/M4/whatever to close the gap on this by leaps and bounds. Again, not going up against Meta\u2019s 15k <em>H100s</em> but all anyone else will ever need.<p>Back to the mobile/edge strategy, Xcode includes what is basically ML training and tuning functionality built in. You can literally train an object recognition model by dragging and dropping pictures, encrypt it, and bundle with your app all within Xcode. App developers are doing ML and barely even noticing. This is in latest Xcode, you can bet your bottom dollar Apple will be putting their significant resources to embracing all of this.<p>Train your model on your Mac, bundle it with your app, scale to infinitely for $0 in hosting costs because the model is running on the user\u2019s device. No Nvidia in sight.<p>In terms of architecture you can still offload the big stuff to datacenter but at increasingly receding rates.<p>I personally think the capability and demand for ML/AI will quickly reach a point where Nvidia and clouds just cannot meet demand for the user base and breadth and scope of the functionality they will increasingly expect.<p>ChatGPT has an estimated 100 MAU. Very impressive but Snapchat alone is 1 billion. Capacity, hardware advancements, and the economics of \u201chost everything on big Nvidia\u201d just doesn\u2019t work out.<p>Google has been putting TPU (lite) silicon  in Pixel devices since roughly <em>202</em>1. Apple with neural engine since the iPhone X in 2017\u2026<p>If you\u2019ve been paying attention to these moves from Google and Apple over the last several years you would have seen this coming. They have not been caught flat-footed on this as so many people, press, etc think.<p>Granted there will always be demand for the big datacenter stuff and Nvidia won\u2019t be hurting anytime soon but expect to see demand for Nvidia hardware and cloud GPU usage to drop more and more as this approach eats more and more."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Nvidia\u2019s AI supremacy is only temporary"},"story_url":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["2026"],"value":"https://petewarden.com/<em>202</em>3/09/10/why-nvidias-ai-supremacy-is-only-temporary/"}},"_tags":["comment","author_kkielhofner","story_37467585"],"author":"kkielhofner","comment_text":"Article is hit and miss.<p>Mostly correct about inference.<p>Snapchat has been running small ML on mobile devices for years for stuff like face filters, etc. Same for those features on iOS that do facial recognition to match pictures of friends with your contacts.<p>Start to pay attention and you\u2019ll realize your phone is doing things like object detection, classification, face tracking, voice assistant wake word detection, some on device speech and command recognition, and a wild array of other ML tasks.<p>LLMs are sucking all of the oxygen out of the room but the overwhelming majority of end-user use of AI isn\u2019t generative and won\u2019t be for a long time if ever. The article is correct in saying inference, inference, inference.<p>The future of inference is smaller application and use-case specific models deployed to edge. Many of these applications just don\u2019t work with the latency of networks to cloud. Imagine face tracking for a Snapchat filter if it involved streaming video to a cloud for inference. Yeah, not happening.<p>The hosting costs are also astronomical, big inference hardware is only getting harder to get and Nvidia only has so much manufacturing capacity.<p>Leave the H100s up to Meta, OpenAI, etc that are training massive multi-billion parameter LLMs from scratch. Or people renting them in small batches to do finetuning of \u201csmaller\u201d models, etc.<p>This is also getting chipped at - with the unified memory of Apple Silicon you can get the RAM of 2 A&#x2F;H100s today for less than the cost of used 80GB A100s. With an entire computer (Mac Pro), new and under warranty.<p>Nvidia still wins on TFLOPS but expect M3&#x2F;M4&#x2F;whatever to close the gap on this by leaps and bounds. Again, not going up against Meta\u2019s 15k H100s but all anyone else will ever need.<p>Back to the mobile&#x2F;edge strategy, Xcode includes what is basically ML training and tuning functionality built in. You can literally train an object recognition model by dragging and dropping pictures, encrypt it, and bundle with your app all within Xcode. App developers are doing ML and barely even noticing. This is in latest Xcode, you can bet your bottom dollar Apple will be putting their significant resources to embracing all of this.<p>Train your model on your Mac, bundle it with your app, scale to infinitely for $0 in hosting costs because the model is running on the user\u2019s device. No Nvidia in sight.<p>In terms of architecture you can still offload the big stuff to datacenter but at increasingly receding rates.<p>I personally think the capability and demand for ML&#x2F;AI will quickly reach a point where Nvidia and clouds just cannot meet demand for the user base and breadth and scope of the functionality they will increasingly expect.<p>ChatGPT has an estimated 100 MAU. Very impressive but Snapchat alone is 1 billion. Capacity, hardware advancements, and the economics of \u201chost everything on big Nvidia\u201d just doesn\u2019t work out.<p>Google has been putting TPU (lite) silicon  in Pixel devices since roughly 2021. Apple with neural engine since the iPhone X in 2017\u2026<p>If you\u2019ve been paying attention to these moves from Google and Apple over the last several years you would have seen this coming. They have not been caught flat-footed on this as so many people, press, etc think.<p>Granted there will always be demand for the big datacenter stuff and Nvidia won\u2019t be hurting anytime soon but expect to see demand for Nvidia hardware and cloud GPU usage to drop more and more as this approach eats more and more.","created_at":"2023-09-11T17:20:33Z","created_at_i":1694452833,"objectID":"37470431","parent_id":37467585,"story_id":37467585,"story_title":"Nvidia\u2019s AI supremacy is only temporary","story_url":"https://petewarden.com/2023/09/10/why-nvidias-ai-supremacy-is-only-temporary/","updated_at":"2024-09-20T15:05:27Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"ozgune"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["a100","h100","inference","2026"],"value":"In March, vLLM picked up some of the improvements in the DeepSeek paper. Through these, vLLM v0.7.3's DeepSeek performance jumped to about 3x+ of what it was before [1].<p>What's exciting is that there's still so much room for improvement. We benchmark around 5K total tokens/s with the sharegpt dataset and 12K total token/s with random 2000/<em>100</em>, using vLLM and under high concurrency.<p>DeepSeek-V3/R1 <em>Inference</em> System Overview [2] quotes &quot;Each <em>H800</em> node delivers an average throughput of 73.7k tokens/s input (including cache hits) during prefilling <i>or</i> 14.8k tokens/s output during decoding.&quot;<p>Yes, DeepSeek deploys a different <em>inference</em> architecture. But this goes onto show just how much room there is for improvement. Looking forward to more open source!<p>[1] <a href=\"https://developers.redhat.com/articles/2025/03/19/how-we-optimized-vllm-deepseek-r1\" rel=\"nofollow\">https://developers.redhat.com/articles/<em>202</em>5/03/19/how-we-opt...</a><p>[2] <a href=\"https://github.com/deepseek-ai/open-infra-index/blob/main/202502OpenSourceWeek/day_6_one_more_thing_deepseekV3R1_inference_system_overview.md\">https://github.com/deepseek-ai/open-infra-index/blob/main/20...</a>"},"story_title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["inference"],"value":"The path to open-sourcing the DeepSeek <em>inference</em> engine"},"story_url":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["inference"],"value":"https://github.com/deepseek-ai/open-infra-index/tree/main/OpenSourcing_DeepSeek_<em>Inference</em>_Engine"}},"_tags":["comment","author_ozgune","story_43682088"],"author":"ozgune","comment_text":"In March, vLLM picked up some of the improvements in the DeepSeek paper. Through these, vLLM v0.7.3&#x27;s DeepSeek performance jumped to about 3x+ of what it was before [1].<p>What&#x27;s exciting is that there&#x27;s still so much room for improvement. We benchmark around 5K total tokens&#x2F;s with the sharegpt dataset and 12K total token&#x2F;s with random 2000&#x2F;100, using vLLM and under high concurrency.<p>DeepSeek-V3&#x2F;R1 Inference System Overview [2] quotes &quot;Each H800 node delivers an average throughput of 73.7k tokens&#x2F;s input (including cache hits) during prefilling <i>or</i> 14.8k tokens&#x2F;s output during decoding.&quot;<p>Yes, DeepSeek deploys a different inference architecture. But this goes onto show just how much room there is for improvement. Looking forward to more open source!<p>[1] <a href=\"https:&#x2F;&#x2F;developers.redhat.com&#x2F;articles&#x2F;2025&#x2F;03&#x2F;19&#x2F;how-we-optimized-vllm-deepseek-r1\" rel=\"nofollow\">https:&#x2F;&#x2F;developers.redhat.com&#x2F;articles&#x2F;2025&#x2F;03&#x2F;19&#x2F;how-we-opt...</a><p>[2] <a href=\"https:&#x2F;&#x2F;github.com&#x2F;deepseek-ai&#x2F;open-infra-index&#x2F;blob&#x2F;main&#x2F;202502OpenSourceWeek&#x2F;day_6_one_more_thing_deepseekV3R1_inference_system_overview.md\">https:&#x2F;&#x2F;github.com&#x2F;deepseek-ai&#x2F;open-infra-index&#x2F;blob&#x2F;main&#x2F;20...</a>","created_at":"2025-04-14T18:04:41Z","created_at_i":1744653881,"objectID":"43684294","parent_id":43682088,"story_id":43682088,"story_title":"The path to open-sourcing the DeepSeek inference engine","story_url":"https://github.com/deepseek-ai/open-infra-index/tree/main/OpenSourcing_DeepSeek_Inference_Engine","updated_at":"2025-04-20T18:43:39Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"zambelli"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["a100","h100","inference","2026"],"value":"Hi HN, I'm Antoine Zambelli, AI Director at Texas Instruments.<p>I built Forge, an open-source reliability layer for self-hosted LLM tool-calling.<p>What it does:<p>- Adds domain-and-tool-agnostic guardrails (retry nudges, step enforcement, error recovery, VRAM-aware context management) to local models running on consumer hardware<p>- Takes an 8B model from ~53% to ~99% on multi-step agentic workflows without changing the model - just the system around it<p>- Ships with an eval harness and interactive dashboard so you can reproduce every number<p>I wanted to run a handful of always-on agentic systems for my portfolio, didn't want to pay cloud frontier costs, and immediately hit the compounding math problem on local models. 90% per-step accuracy sounds great, but with a 5-step workflow that's a 40% failure rate. No existing framework seemed to address this mechanical reliability issue - they all seemed tailor-made for cloud frontier.<p>Demo video: <a href=\"https://youtu.be/MzRgJoJAXGc\" rel=\"nofollow\">https://youtu.be/MzRgJoJAXGc</a> (side-by-side: same model, same task, with and without Forge guardrails)<p>The paper (accepted to ACM CAIS '26, presenting May 26-29 in San Jose) covers the peer-reviewed findings across 97 model/backend configurations, 18 scenarios, 50 runs each. Key numbers:<p>- Ministral 8B with Forge: 99.3%. Claude Sonnet with Forge: <em>100</em>%. The gap between a free local 8B model on a $600 GPU and a frontier API is less than 1 point.<p>- The same 8B local model with Forge (99.3%) outperforms Claude Sonnet without guardrails (87.2%) - an 8B model with framework support beats the best result you can get through frontier API alone.<p>- Error recovery scores 0% for every model tested - local and frontier - without the retry mechanism. Not a capability gap, an architectural absence.<p>I'm currently using this for my home assistant running on Ministral 14B-Reasoning, and for my locally hosted agentic coding harness (8B managed to contribute to the codebase!).<p>The guardrail stack has five layers, each independently toggleable. The two that carry the most weight (per ablation study with McNemar's test): retry nudges (24-49 point drops when disabled) and error recovery (~10 point drops, significant for every model tested). Step enforcement is situational - only fires for models with weaker sequencing discipline. Rescue parsing and context compaction showed no significance in the eval but are retained for production workloads where they activate once in a while.<p>One thing I really didn't expect: the serving backend matters. Same Mistral-Nemo 12B weights produce 7% accuracy on llama-server with native function calling and 83% on Llamafile in prompt mode. A 75-point swing from infrastructure alone. I don't think anyone's published this because standard benchmarks don't control for serving backend.<p>Another surprise: there's no distinction in current LLM tool-calling between &quot;the tool ran successfully and returned data&quot; and &quot;the tool ran successfully but found nothing.&quot; Both return a value, the orchestrator marks the step complete, and bad data cascades downstream. It's the equivalent of HTTP having 200 but no 404. Forge adds this as a new exception class (ToolResolutionError) - the model sees the error and can retry instead of silently passing garbage forward.<p>Biggest technical challenge was context compaction for memory-constrained hardware. Both Ollama and Llamafile silently fall back to CPU when the model exceeds VRAM - no warning, no error, just 10-100x slower <em>inference</em>. Forge queries nvidia-smi at startup and derives a token budget to prevent this.<p>How to try it:<p>- Clone the repo, run the eval harness on a model I haven't tested. If you get interesting results I'll add them to the dashboard.<p>- Try the proxy server mode - point any OpenAI-compatible client at Forge and it handles guardrails transparently. It's the newest model and I'd love more eyes on it.<p>- Dogfooding led me to optimize model parameters in v0.6.0. The harder eval suite (26 scenarios) is designed to raise the ceiling so no one sits at <em>100</em>%. Several that did on the original suite can't sweep it - including Opus 4.6. Curious if anyone finds scenarios that expose gaps I haven't thought of. Paper numbers based on pre v0.6.0 code.<p>Background: prior ML publication in unsupervised learning (83 citations). This paper accepted to ACM CAIS '26 - presenting May 26-29.<p>Repo: <a href=\"https://github.com/antoinezambelli/forge\" rel=\"nofollow\">https://github.com/antoinezambelli/forge</a><p>Paper: <a href=\"https://www.caisconf.org/program/2026/demos/forge-agentic-reliability/\" rel=\"nofollow\">https://www.caisconf.org/program/<em>2026</em>/demos/forge-agentic-re...</a> <a href=\"https://github.com/antoinezambelli/forge/blob/main/docs/forge_ieee_preprint.pdf\" rel=\"nofollow\">https://github.com/antoinezambelli/forge/blob/main/docs/forg...</a><p>Dashboard: <a href=\"https://github.com/antoinezambelli/forge/docs/results/dashboard.html\" rel=\"nofollow\">https://github.com/antoinezambelli/forge/docs/results/dashbo...</a>"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Forge \u2013 Guardrails take an 8B model from 53% to 99% on agentic tasks"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/antoinezambelli/forge"}},"_tags":["story","author_zambelli","story_48192383","show_hn"],"author":"zambelli","children":[48197802,48198479,48198514,48198634,48198706,48198875,48199115,48199142,48199417,48199806,48199948,48200002,48200090,48200124,48200234,48200286,48200304,48200330,48200359,48200762,48200832,48200953,48201036,48201398,48201421,48201422,48201622,48201768,48201888,48202111,48202167,48202227,48202361,48202380,48202590,48203121,48203145,48203298,48203387,48203439,48203590,48203596,48203742,48203774,48203782,48203854,48203968,48203974,48204143,48204259,48204391,48204464,48204487,48204493,48204515,48204556,48205229,48205302,48205733,48205781,48206044,48206436,48206492,48206574,48208442,48208854,48208866,48209241,48209380,48209894,48209974,48210301,48211054,48211145,48213619,48217110,48217211,48221239,48222249,48222279,48223509,48235392,48238617,48238932,48249395,48250938,48264115,48277770,48279760,48304747],"created_at":"2026-05-19T12:23:07Z","created_at_i":1779193387,"num_comments":252,"objectID":"48192383","points":687,"story_id":48192383,"story_text":"Hi HN, I&#x27;m Antoine Zambelli, AI Director at Texas Instruments.<p>I built Forge, an open-source reliability layer for self-hosted LLM tool-calling.<p>What it does:<p>- Adds domain-and-tool-agnostic guardrails (retry nudges, step enforcement, error recovery, VRAM-aware context management) to local models running on consumer hardware<p>- Takes an 8B model from ~53% to ~99% on multi-step agentic workflows without changing the model - just the system around it<p>- Ships with an eval harness and interactive dashboard so you can reproduce every number<p>I wanted to run a handful of always-on agentic systems for my portfolio, didn&#x27;t want to pay cloud frontier costs, and immediately hit the compounding math problem on local models. 90% per-step accuracy sounds great, but with a 5-step workflow that&#x27;s a 40% failure rate. No existing framework seemed to address this mechanical reliability issue - they all seemed tailor-made for cloud frontier.<p>Demo video: <a href=\"https:&#x2F;&#x2F;youtu.be&#x2F;MzRgJoJAXGc\" rel=\"nofollow\">https:&#x2F;&#x2F;youtu.be&#x2F;MzRgJoJAXGc</a> (side-by-side: same model, same task, with and without Forge guardrails)<p>The paper (accepted to ACM CAIS &#x27;26, presenting May 26-29 in San Jose) covers the peer-reviewed findings across 97 model&#x2F;backend configurations, 18 scenarios, 50 runs each. Key numbers:<p>- Ministral 8B with Forge: 99.3%. Claude Sonnet with Forge: 100%. The gap between a free local 8B model on a $600 GPU and a frontier API is less than 1 point.<p>- The same 8B local model with Forge (99.3%) outperforms Claude Sonnet without guardrails (87.2%) - an 8B model with framework support beats the best result you can get through frontier API alone.<p>- Error recovery scores 0% for every model tested - local and frontier - without the retry mechanism. Not a capability gap, an architectural absence.<p>I&#x27;m currently using this for my home assistant running on Ministral 14B-Reasoning, and for my locally hosted agentic coding harness (8B managed to contribute to the codebase!).<p>The guardrail stack has five layers, each independently toggleable. The two that carry the most weight (per ablation study with McNemar&#x27;s test): retry nudges (24-49 point drops when disabled) and error recovery (~10 point drops, significant for every model tested). Step enforcement is situational - only fires for models with weaker sequencing discipline. Rescue parsing and context compaction showed no significance in the eval but are retained for production workloads where they activate once in a while.<p>One thing I really didn&#x27;t expect: the serving backend matters. Same Mistral-Nemo 12B weights produce 7% accuracy on llama-server with native function calling and 83% on Llamafile in prompt mode. A 75-point swing from infrastructure alone. I don&#x27;t think anyone&#x27;s published this because standard benchmarks don&#x27;t control for serving backend.<p>Another surprise: there&#x27;s no distinction in current LLM tool-calling between &quot;the tool ran successfully and returned data&quot; and &quot;the tool ran successfully but found nothing.&quot; Both return a value, the orchestrator marks the step complete, and bad data cascades downstream. It&#x27;s the equivalent of HTTP having 200 but no 404. Forge adds this as a new exception class (ToolResolutionError) - the model sees the error and can retry instead of silently passing garbage forward.<p>Biggest technical challenge was context compaction for memory-constrained hardware. Both Ollama and Llamafile silently fall back to CPU when the model exceeds VRAM - no warning, no error, just 10-100x slower inference. Forge queries nvidia-smi at startup and derives a token budget to prevent this.<p>How to try it:<p>- Clone the repo, run the eval harness on a model I haven&#x27;t tested. If you get interesting results I&#x27;ll add them to the dashboard.<p>- Try the proxy server mode - point any OpenAI-compatible client at Forge and it handles guardrails transparently. It&#x27;s the newest model and I&#x27;d love more eyes on it.<p>- Dogfooding led me to optimize model parameters in v0.6.0. The harder eval suite (26 scenarios) is designed to raise the ceiling so no one sits at 100%. Several that did on the original suite can&#x27;t sweep it - including Opus 4.6. Curious if anyone finds scenarios that expose gaps I haven&#x27;t thought of. Paper numbers based on pre v0.6.0 code.<p>Background: prior ML publication in unsupervised learning (83 citations). This paper accepted to ACM CAIS &#x27;26 - presenting May 26-29.<p>Repo: <a href=\"https:&#x2F;&#x2F;github.com&#x2F;antoinezambelli&#x2F;forge\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;antoinezambelli&#x2F;forge</a><p>Paper: <a href=\"https:&#x2F;&#x2F;www.caisconf.org&#x2F;program&#x2F;2026&#x2F;demos&#x2F;forge-agentic-reliability&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;www.caisconf.org&#x2F;program&#x2F;2026&#x2F;demos&#x2F;forge-agentic-re...</a> <a href=\"https:&#x2F;&#x2F;github.com&#x2F;antoinezambelli&#x2F;forge&#x2F;blob&#x2F;main&#x2F;docs&#x2F;forge_ieee_preprint.pdf\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;antoinezambelli&#x2F;forge&#x2F;blob&#x2F;main&#x2F;docs&#x2F;forg...</a><p>Dashboard: <a href=\"https:&#x2F;&#x2F;github.com&#x2F;antoinezambelli&#x2F;forge&#x2F;docs&#x2F;results&#x2F;dashboard.html\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;antoinezambelli&#x2F;forge&#x2F;docs&#x2F;results&#x2F;dashbo...</a>","title":"Show HN: Forge \u2013 Guardrails take an 8B model from 53% to 99% on agentic tasks","updated_at":"2026-08-25T15:19:50Z","url":"https://github.com/antoinezambelli/forge"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"tommy_mcclung"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["a100","h100","inference","2026"],"value":"Hello Hacker News! We\u2019re Erik, Tommy, and David, the founders of Release (<a href=\"https://release.ai/\" rel=\"nofollow\">https://release.ai/</a>). We launched on HN in <em>202</em>0 (<a href=\"https://news.ycombinator.com/item?id=22486031\">https://news.ycombinator.com/item?id=22486031</a>) after leaving TrueCar, where we managed a 300 person development team. Our original focus was making staging environments easier with ephemeral environments, but along the way AI applications started to emerge as an important and critical component of distributed applications. As we talked to customers using our original product, we realized we had built the underlying platform needed to address the needs of orchestrating AI applications and infrastructure. So here we are and we\u2019re excited to share Release.ai with HN.<p>Here\u2019s a video showcasing the platform and demonstrating how to easily manage new data and changes using the RAG stack of your choice: <a href=\"https://www.youtube.com/watch?v=-OdWRxMX1iA\" rel=\"nofollow\">https://www.youtube.com/watch?v=-OdWRxMX1iA</a><p>If you want to try release.ai out, we\u2019re offering a sandbox account with limited free GPU cycles so you can play around and get a feel for Release.ai: <a href=\"https://release.ai\" rel=\"nofollow\">https://release.ai</a>. We suggest playing around with some of the RAG AI templates and adding custom workflows like in the demo video. The sandbox comes with 5 free compute hours on an Amazon g5.2xlarge instance (<em>A10</em> with 24GB VRAM, 8vCPUs and 32GB). You will also get 16 GB and 4vCPUs for cpu workloads such as web servers. You will be able to run an <em>inference</em> engine plus things like an api server, etc.<p>After the sandbox expires, you can switch to our free plan, which requires a credit card and associating an AWS/GCP account with Release to manage the compute in your cloud account. The free account provides <em>100</em> free managed environment hours a month. If you never go over, you never pay us anything. If you do, our pricing is here: <a href=\"https://release.com/pricing\">https://release.com/pricing</a>.<p>For those that like to read more, here\u2019s the deeper background.<p>It\u2019s clear that open source AI and AI privacy are going to be big. Yes, many developers are going to choose SaaS offerings like OpenAI to build their AI applications, but as open source frameworks and models improve, we\u2019re seeing a shift to open source running on cloud. Security and privacy is a top concern of companies leveraging these SaaS solutions, which forces them to look at running infrastructure themselves. That\u2019s where we hope to come in: we\u2019ve built Release.ai so all your data, models and infrastructure stay in your cloud account and open source frameworks are first class citizens.<p>Orchestration - Integrating AI applications into a software development workflow and orchestrating their lifecycle is a new and different challenge than traditional web application development. Release also makes it possible to manage and integrate your web and AI apps using a single application and methodology.<p>To make orchestrating AI applications easier, we built a workflow engine that can create the complex workflows that AI applications require. For example, you can automate the redeployment of an AI <em>inference</em> server easily when underlying data changes using webhooks and our workflow engine.<p>Cost and expertise - Managing and scaling the hardware required to run AI workloads is hard and can be incredibly expensive. Release.ai lets you manage GPU compute resources across multiple clouds with different instance/node groups for various jobs within a single admin interface. We use K8s under the covers to pull this off. With over 5 years of building and running K8s infrastructure our customers have told us this is how it should be done.<p>Getting started with AI frameworks is time consuming and requires some pretty in-depth expertise. We built out a library of AI templates (<a href=\"https://docs.release.com/release.ai/release.ai-templates\">https://docs.release.com/release.ai/release.ai-templates</a>) using our Application Template format (which is kind of a super docker-compose: <a href=\"https://docs.release.com/reference-documentation/application-settings/application-template\">https://docs.release.com/reference-documentation/application...</a>) for common open source frameworks to make it easy to get started developing AI applications. Setting up and getting these frameworks running is a hassle, so we made it one click to launch and deploy.<p>We currently have over 20 templates including temples for RAG applications, fine tuning and useful tools like Juypter notebooks, Promptfoo, etc. We worked closely with Docker and Nvidia to support their frameworks: GenAI and Nvidia NEMO/Nims. We plan to launch community templates soon after launch. If you have suggestions for more templates we should support, please let us know in the comments.<p>We\u2019re thrilled to share Release.ai with you and would love to get your feedback. We hope you\u2019ll  try it out, and please let us know what you think!"},"title":{"matchLevel":"none","matchedWords":[],"value":"Launch HN: Release (YC W20) \u2013 Orchestrate AI Infrastructure and Applications"}},"_tags":["story","author_tommy_mcclung","story_41182055","launch_hn"],"author":"tommy_mcclung","children":[41182480,41183046,41183413,41183637,41183752,41183753,41183759,41183909,41183918,41184126,41184347,41184640,41184812,41185162,41185839,41186787,41187944,41190854],"created_at":"2024-08-07T14:50:11Z","created_at_i":1723042211,"num_comments":39,"objectID":"41182055","points":73,"story_id":41182055,"story_text":"Hello Hacker News! We\u2019re Erik, Tommy, and David, the founders of Release (<a href=\"https:&#x2F;&#x2F;release.ai&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;release.ai&#x2F;</a>). We launched on HN in 2020 (<a href=\"https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=22486031\">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=22486031</a>) after leaving TrueCar, where we managed a 300 person development team. Our original focus was making staging environments easier with ephemeral environments, but along the way AI applications started to emerge as an important and critical component of distributed applications. As we talked to customers using our original product, we realized we had built the underlying platform needed to address the needs of orchestrating AI applications and infrastructure. So here we are and we\u2019re excited to share Release.ai with HN.<p>Here\u2019s a video showcasing the platform and demonstrating how to easily manage new data and changes using the RAG stack of your choice: <a href=\"https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=-OdWRxMX1iA\" rel=\"nofollow\">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=-OdWRxMX1iA</a><p>If you want to try release.ai out, we\u2019re offering a sandbox account with limited free GPU cycles so you can play around and get a feel for Release.ai: <a href=\"https:&#x2F;&#x2F;release.ai\" rel=\"nofollow\">https:&#x2F;&#x2F;release.ai</a>. We suggest playing around with some of the RAG AI templates and adding custom workflows like in the demo video. The sandbox comes with 5 free compute hours on an Amazon g5.2xlarge instance (A10 with 24GB VRAM, 8vCPUs and 32GB). You will also get 16 GB and 4vCPUs for cpu workloads such as web servers. You will be able to run an inference engine plus things like an api server, etc.<p>After the sandbox expires, you can switch to our free plan, which requires a credit card and associating an AWS&#x2F;GCP account with Release to manage the compute in your cloud account. The free account provides 100 free managed environment hours a month. If you never go over, you never pay us anything. If you do, our pricing is here: <a href=\"https:&#x2F;&#x2F;release.com&#x2F;pricing\">https:&#x2F;&#x2F;release.com&#x2F;pricing</a>.<p>For those that like to read more, here\u2019s the deeper background.<p>It\u2019s clear that open source AI and AI privacy are going to be big. Yes, many developers are going to choose SaaS offerings like OpenAI to build their AI applications, but as open source frameworks and models improve, we\u2019re seeing a shift to open source running on cloud. Security and privacy is a top concern of companies leveraging these SaaS solutions, which forces them to look at running infrastructure themselves. That\u2019s where we hope to come in: we\u2019ve built Release.ai so all your data, models and infrastructure stay in your cloud account and open source frameworks are first class citizens.<p>Orchestration - Integrating AI applications into a software development workflow and orchestrating their lifecycle is a new and different challenge than traditional web application development. Release also makes it possible to manage and integrate your web and AI apps using a single application and methodology.<p>To make orchestrating AI applications easier, we built a workflow engine that can create the complex workflows that AI applications require. For example, you can automate the redeployment of an AI inference server easily when underlying data changes using webhooks and our workflow engine.<p>Cost and expertise - Managing and scaling the hardware required to run AI workloads is hard and can be incredibly expensive. Release.ai lets you manage GPU compute resources across multiple clouds with different instance&#x2F;node groups for various jobs within a single admin interface. We use K8s under the covers to pull this off. With over 5 years of building and running K8s infrastructure our customers have told us this is how it should be done.<p>Getting started with AI frameworks is time consuming and requires some pretty in-depth expertise. We built out a library of AI templates (<a href=\"https:&#x2F;&#x2F;docs.release.com&#x2F;release.ai&#x2F;release.ai-templates\">https:&#x2F;&#x2F;docs.release.com&#x2F;release.ai&#x2F;release.ai-templates</a>) using our Application Template format (which is kind of a super docker-compose: <a href=\"https:&#x2F;&#x2F;docs.release.com&#x2F;reference-documentation&#x2F;application-settings&#x2F;application-template\">https:&#x2F;&#x2F;docs.release.com&#x2F;reference-documentation&#x2F;application...</a>) for common open source frameworks to make it easy to get started developing AI applications. Setting up and getting these frameworks running is a hassle, so we made it one click to launch and deploy.<p>We currently have over 20 templates including temples for RAG applications, fine tuning and useful tools like Juypter notebooks, Promptfoo, etc. We worked closely with Docker and Nvidia to support their frameworks: GenAI and Nvidia NEMO&#x2F;Nims. We plan to launch community templates soon after launch. If you have suggestions for more templates we should support, please let us know in the comments.<p>We\u2019re thrilled to share Release.ai with you and would love to get your feedback. We hope you\u2019ll  try it out, and please let us know what you think!","title":"Launch HN: Release (YC W20) \u2013 Orchestrate AI Infrastructure and Applications","updated_at":"2024-09-20T17:37:31Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"crimeacs"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["a100","h100","inference","2026"],"value":"Author here - quick context so this doesn\u2019t read like a hot take.<p>Dataset: 148,421 public HN stories from Algolia since 2007, filtered to score \u22655. Split is strictly chronological: train &lt; Jul 2025, val Aug\u2013Dec 2025, holdout Jan <em>2026</em>+. Random splits are misleading here because kNN features leak future neighbors.<p>Model: LightGBM with 4 heads: median, p10, p90, and score \u2265<em>100</em> classifier with isotonic calibration. Compiled to plain JS via m2cgen and runs inside a Vercel function \u2014 no Python/ONNX/runtime. ~10 MB bundle, sub-ms <em>inference</em>.<p>Holdout:<p>* Spearman \u03c1 = 0.33 on log_score<p>* MAE log = 1.65, roughly ~5x off in raw points<p>* AUC for score \u2265<em>100</em> = 0.67<p>* Precision@30 = 0.83<p>So: not magic. About one-third of the signal seems recoverable from title/context. AUC is below ontology2\u2019s 2014 title-only baseline, around/above recent BERT fine-tunes I found.<p>Two things I haven\u2019t seen elsewhere:<p>1. Comment simulator grounds every fake comment in a real top comment from a kNN neighbor, with `[src]`.\n2. `/predictions` runs a live calibration ledger against actual HN top 30 every 10 min, so the model can\u2019t hide behind a static benchmark.<p>Open source, MIT, training scripts included:\n<a href=\"https://github.com/crimeacs/foresyn-hackernews\" rel=\"nofollow\">https://github.com/crimeacs/foresyn-hackernews</a><p>I ran the submitted title through the model first. It predicted 32/99 virality and ~12 points. The ledger will soon tell us whether it was calibrated.<p>Roast away."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"148,421 posts later: A model to predict the Hacker News front page"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://hackernews.foresyn.ai/"}},"_tags":["comment","author_crimeacs","story_48109136"],"author":"crimeacs","children":[48116024],"comment_text":"Author here - quick context so this doesn\u2019t read like a hot take.<p>Dataset: 148,421 public HN stories from Algolia since 2007, filtered to score \u22655. Split is strictly chronological: train &lt; Jul 2025, val Aug\u2013Dec 2025, holdout Jan 2026+. Random splits are misleading here because kNN features leak future neighbors.<p>Model: LightGBM with 4 heads: median, p10, p90, and score \u2265100 classifier with isotonic calibration. Compiled to plain JS via m2cgen and runs inside a Vercel function \u2014 no Python&#x2F;ONNX&#x2F;runtime. ~10 MB bundle, sub-ms inference.<p>Holdout:<p>* Spearman \u03c1 = 0.33 on log_score<p>* MAE log = 1.65, roughly ~5x off in raw points<p>* AUC for score \u2265100 = 0.67<p>* Precision@30 = 0.83<p>So: not magic. About one-third of the signal seems recoverable from title&#x2F;context. AUC is below ontology2\u2019s 2014 title-only baseline, around&#x2F;above recent BERT fine-tunes I found.<p>Two things I haven\u2019t seen elsewhere:<p>1. Comment simulator grounds every fake comment in a real top comment from a kNN neighbor, with `[src]`.\n2. `&#x2F;predictions` runs a live calibration ledger against actual HN top 30 every 10 min, so the model can\u2019t hide behind a static benchmark.<p>Open source, MIT, training scripts included:\n<a href=\"https:&#x2F;&#x2F;github.com&#x2F;crimeacs&#x2F;foresyn-hackernews\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;crimeacs&#x2F;foresyn-hackernews</a><p>I ran the submitted title through the model first. It predicted 32&#x2F;99 virality and ~12 points. The ledger will soon tell us whether it was calibrated.<p>Roast away.","created_at":"2026-05-12T15:00:08Z","created_at_i":1778598008,"objectID":"48109317","parent_id":48109136,"story_id":48109136,"story_title":"148,421 posts later: A model to predict the Hacker News front page","story_url":"https://hackernews.foresyn.ai/","updated_at":"2026-05-12T23:42:33Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"TulioKBR"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["a100","h100","inference","2026"],"value":"Great question - I'll be direct.<p>It's not that Gemini &amp; Sonnet are excluded. They're architecture-ready (we built the abstraction layer), but they're *not in v1 for 3 hard technical reasons:*<p>*1. Code Generation Consistency*\nFor *enterprise TypeScript code generation*, you need deterministic output. Gemini &amp; Sonnet show 12-18% variance on repeated prompts (same input, different implementations). Perplexity + Claude stabilize at 3-5%, Groq at 2%. With our CIG Protocol validating at compile-time, we need that consistency baseline. Once Google &amp; Anthropic stabilize their fine-tuning for code tasks, we'll enable them.<p>*2. Long-Context Cost Economics*\nEnterprise prompts for ORUS average 18K tokens (blueprint + requirements + patterns). At current pricing:\n- Perplexity: $3/1M input tokens (~$0.054 per generation)\n- Claude 3.5: $3/1M input (~$0.054 per generation)  \n- Groq: $0.05/1M input (~$0.0009 per generation)\n- Gemini 2.0 Flash: pricing TBA, likely $0.075/1M\n- Sonnet 4.5: $3/1M (~$0.054)<p>For customers running <em>100</em> generations daily, the margin between Groq + Perplexity vs Gemini/Sonnet = $50-<em>100</em>/month difference. We *can't ignore cost* when targeting startups.<p>*3. API Stability During Code Generation*\nThis is the real blocker:\n- Perplexity: 99.8% uptime, code-optimized endpoints\n- Claude: 99.7% uptime, fine-tuning controls  \n- Groq: 99.9% uptime, lightweight <em>inference</em>\n- Gemini: Recent instability (Nov 2025 API timeouts)\n- Sonnet: Good, but new version (4.5) still stabilizing<p>When generating production code, a timeout mid-stream = corrupted output. We can't ship that in v1.<p>*Here's the honest roadmap:*\n- *v1 (now)*: Perplexity + Claude + Groq (battle-tested)\n- *v1.2 (Jan <em>2026</em>)*: Gemini 2.0 (when pricing finalizes &amp; API stabilizes)\n- *v1.3 (Feb <em>2026</em>)*: Sonnet 4.5 (fine-tuning for code generation confirmed)\n- *v2 (Q2 <em>2026</em>)*: All models with fallback switching (if one fails, auto-retry on another)<p>*Why be conservative in v1?* We have 400+ enterprise users waiting for open-source release. One corrupted generation costs us 5+ years of credibility. Better to add models post-launch when we have production telemetry.<p>If you want Gemini/Sonnet support pre-launch, you can self-enable it - our provider abstraction supports any OpenAI-compatible API in ~10 lines of code."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: I built an AI that generates full-stack apps in 30 seconds"}},"_tags":["comment","author_TulioKBR","story_45798328"],"author":"TulioKBR","children":[45798793],"comment_text":"Great question - I&#x27;ll be direct.<p>It&#x27;s not that Gemini &amp; Sonnet are excluded. They&#x27;re architecture-ready (we built the abstraction layer), but they&#x27;re *not in v1 for 3 hard technical reasons:*<p>*1. Code Generation Consistency*\nFor *enterprise TypeScript code generation*, you need deterministic output. Gemini &amp; Sonnet show 12-18% variance on repeated prompts (same input, different implementations). Perplexity + Claude stabilize at 3-5%, Groq at 2%. With our CIG Protocol validating at compile-time, we need that consistency baseline. Once Google &amp; Anthropic stabilize their fine-tuning for code tasks, we&#x27;ll enable them.<p>*2. Long-Context Cost Economics*\nEnterprise prompts for ORUS average 18K tokens (blueprint + requirements + patterns). At current pricing:\n- Perplexity: $3&#x2F;1M input tokens (~$0.054 per generation)\n- Claude 3.5: $3&#x2F;1M input (~$0.054 per generation)  \n- Groq: $0.05&#x2F;1M input (~$0.0009 per generation)\n- Gemini 2.0 Flash: pricing TBA, likely $0.075&#x2F;1M\n- Sonnet 4.5: $3&#x2F;1M (~$0.054)<p>For customers running 100 generations daily, the margin between Groq + Perplexity vs Gemini&#x2F;Sonnet = $50-100&#x2F;month difference. We *can&#x27;t ignore cost* when targeting startups.<p>*3. API Stability During Code Generation*\nThis is the real blocker:\n- Perplexity: 99.8% uptime, code-optimized endpoints\n- Claude: 99.7% uptime, fine-tuning controls  \n- Groq: 99.9% uptime, lightweight inference\n- Gemini: Recent instability (Nov 2025 API timeouts)\n- Sonnet: Good, but new version (4.5) still stabilizing<p>When generating production code, a timeout mid-stream = corrupted output. We can&#x27;t ship that in v1.<p>*Here&#x27;s the honest roadmap:*\n- *v1 (now)*: Perplexity + Claude + Groq (battle-tested)\n- *v1.2 (Jan 2026)*: Gemini 2.0 (when pricing finalizes &amp; API stabilizes)\n- *v1.3 (Feb 2026)*: Sonnet 4.5 (fine-tuning for code generation confirmed)\n- *v2 (Q2 2026)*: All models with fallback switching (if one fails, auto-retry on another)<p>*Why be conservative in v1?* We have 400+ enterprise users waiting for open-source release. One corrupted generation costs us 5+ years of credibility. Better to add models post-launch when we have production telemetry.<p>If you want Gemini&#x2F;Sonnet support pre-launch, you can self-enable it - our provider abstraction supports any OpenAI-compatible API in ~10 lines of code.","created_at":"2025-11-03T13:30:27Z","created_at_i":1762176627,"objectID":"45798754","parent_id":45798678,"story_id":45798328,"story_title":"Show HN: I built an AI that generates full-stack apps in 30 seconds","updated_at":"2026-03-05T23:01:03Z"}],"hitsPerPage":20,"nbHits":38,"nbPages":2,"page":0,"params":"query=a100+h100+inference+2026&advancedSyntax=true&analyticsTags=backend","processingTimeMS":38,"processingTimingsMS":{"_request":{"roundTrip":15},"afterFetch":{"format":{"highlighting":3,"total":4},"merge":{"mergeLoop":{"prepareNextHit":4,"total":4},"total":4},"total":5},"fetch":{"query":12,"scanning":19,"total":32},"total":38},"query":"a100 h100 inference 2026","serverTimeMS":43}
