{"exhaustive":{"nbHits":false,"typo":false},"exhaustiveNbHits":false,"exhaustiveTypo":false,"hits":[{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"llm_nerd"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["h100","sxm","price","performance","inference"],"value":"For those who wonder where this fits relative to platforms we are usually more accustomed to, they sell these to banks and financial services that have a long history of mainframes and are basically just upgrading in place within minimal change. With recent iterations they've added some AI processing on the silicon, offering baby steps to imbuing solutions like fraud detection with neural nets on chip.<p>But to put it in context, the 24 TOPS that they advertise -- the <em>inference</em> <em>performance</em> of their AI module on their Telum II -- doesn't even match an M4's neural engine (40 TOPS). And of course compared to a dedicated device like an <em>H100</em> <em>SXM</em> that can hit 4000 TOPS (yes, 166x more). Of course for both the M4 and <em>H100</em> chip I'm giving quantized numbers, but presumably the Telum II is as it boasts about its quantized support.<p>Massive caches. Tonnes of memory support. Neat device. A &quot;you won't get fired for leasing this&quot; solution for a few of the Fortune 500.<p>But you can almost certainly build a magnitudes faster device in just about every dimension using more traditional hardware stacks and for a fraction of the <em>price</em>."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"IBM Announces the Z17 Mainframe Powered by Telum II Processors"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://www.phoronix.com/news/IBM-z17-Telum-2-Announced"}},"_tags":["comment","author_llm_nerd","story_43620257"],"author":"llm_nerd","children":[43621680,43622067,43622129],"comment_text":"For those who wonder where this fits relative to platforms we are usually more accustomed to, they sell these to banks and financial services that have a long history of mainframes and are basically just upgrading in place within minimal change. With recent iterations they&#x27;ve added some AI processing on the silicon, offering baby steps to imbuing solutions like fraud detection with neural nets on chip.<p>But to put it in context, the 24 TOPS that they advertise -- the inference performance of their AI module on their Telum II -- doesn&#x27;t even match an M4&#x27;s neural engine (40 TOPS). And of course compared to a dedicated device like an H100 SXM that can hit 4000 TOPS (yes, 166x more). Of course for both the M4 and H100 chip I&#x27;m giving quantized numbers, but presumably the Telum II is as it boasts about its quantized support.<p>Massive caches. Tonnes of memory support. Neat device. A &quot;you won&#x27;t get fired for leasing this&quot; solution for a few of the Fortune 500.<p>But you can almost certainly build a magnitudes faster device in just about every dimension using more traditional hardware stacks and for a fraction of the price.","created_at":"2025-04-08T12:37:11Z","created_at_i":1744115831,"objectID":"43621012","parent_id":43620257,"story_id":43620257,"story_title":"IBM Announces the Z17 Mainframe Powered by Telum II Processors","story_url":"https://www.phoronix.com/news/IBM-z17-Telum-2-Announced","updated_at":"2025-04-26T10:12:24Z"}],"hitsPerPage":20,"nbHits":1,"nbPages":1,"page":0,"params":"query=H100+SXM+price+performance+inference&advancedSyntax=true&analyticsTags=backend","processingTimeMS":19,"processingTimingsMS":{"_request":{"roundTrip":18},"fetch":{"query":17,"total":18},"total":19},"query":"H100 SXM price performance inference","serverTimeMS":19}
