{"author":"hardmaru","children":[{"author":"brucethemoose2","children":[{"author":"gkaul","children":[],"created_at":"2023-08-10T02:11:14.000Z","created_at_i":1691633474,"id":37071081,"options":[],"parent_id":37009325,"points":null,"story_id":37009272,"text":"Could you please paste a reference to the Microsoft paper? Thanks","title":null,"type":"comment","url":null}],"created_at":"2023-08-05T05:41:46.000Z","created_at_i":1691214106,"id":37009325,"options":[],"parent_id":37009272,"points":null,"story_id":37009272,"text":"What is the architecture&#x27;s shape? A chain of cheap SRAM heavy chips as proposed in the Microsoft paper?<p>I feel like AMD&#x27;s 7900 XTX strategy would be good too: a bunch of small, cheap (LPDDRX?) memory controller dies to form a massive bus for a central compute tile.","title":null,"type":"comment","url":null},{"author":"klysm","children":[],"created_at":"2023-08-05T06:11:02.000Z","created_at_i":1691215862,"id":37009439,"options":[],"parent_id":37009272,"points":null,"story_id":37009272,"text":"I remember seeing discussion a couple months back about how long it would take to get ASICS for LLMs. I guess the answer is  this long","title":null,"type":"comment","url":null},{"author":"fnordpiglet","children":[{"author":"cool-RR","children":[{"author":"fnordpiglet","children":[],"created_at":"2023-08-05T07:11:53.000Z","created_at_i":1691219513,"id":37009733,"options":[],"parent_id":37009609,"points":null,"story_id":37009272,"text":"In the aquihires I\u2019ve been involved in investors did well, as owners they\u2019re compensated as part of the acquisition as well.  You see these sorts of specialized team acquisitions all the time in big tech companies where a foundational tech team is built that develops some key technology that they don\u2019t have the critical mass, capital, brand, or vertical ability to bring to market.  They\u2019re acquired primarily for the team assembled and expertise, but the amounts paid can be extraordinary depending on how advanced their technology is and how crucial it is.","title":null,"type":"comment","url":null}],"created_at":"2023-08-05T06:48:30.000Z","created_at_i":1691218110,"id":37009609,"options":[],"parent_id":37009561,"points":null,"story_id":37009272,"text":"Would investors really be interested in an acquihire play? Aren&#x27;t they looking for a big multiplier on their investment?","title":null,"type":"comment","url":null}],"created_at":"2023-08-05T06:38:04.000Z","created_at_i":1691217484,"id":37009561,"options":[],"parent_id":37009272,"points":null,"story_id":37009272,"text":"I\u2019ll wager this is a aquirhire play, with no product based exit but rather offering a asic dev team and prototype to an established hardware company.","title":null,"type":"comment","url":null},{"author":"rock_artist","children":[{"author":"7moritz7","children":[],"created_at":"2023-08-05T06:41:48.000Z","created_at_i":1691217708,"id":37009578,"options":[],"parent_id":37009563,"points":null,"story_id":37009272,"text":"Historically speaking there is a good chance the transformers approach will be replaced by a new approach or a substantially altered transformers approach in the next five years. Efficiency will increase both on the software and the hardware front simultaneously, there is a lot to be optimized with the same parameter count.<p>What looks state of the art will probably look to people in 20 years how the 1885 Mercedes Benz car looks like to us in terms of efficiency, right now it&#x27;s a bit of a bruteforce approach in general","title":null,"type":"comment","url":null},{"author":"throwawayadvsec","children":[{"author":"DoingIsLearning","children":[],"created_at":"2023-08-05T07:00:24.000Z","created_at_i":1691218824,"id":37009674,"options":[],"parent_id":37009594,"points":null,"story_id":37009272,"text":"&gt; they can always pivot in a year<p>I think you are downplaying a bit the timescale of going from ideation to tape-out in ASIC design.","title":null,"type":"comment","url":null}],"created_at":"2023-08-05T06:45:11.000Z","created_at_i":1691217911,"id":37009594,"options":[],"parent_id":37009563,"points":null,"story_id":37009272,"text":"there are already dozens of use cases that would become (more) profitable&#x2F;interesting if we had chips optimized for LLMs<p>they were asic for BTC at a way earlier stage<p>+they can always pivot in a year if the market changes too much","title":null,"type":"comment","url":null},{"author":"aperrien","children":[],"created_at":"2023-08-05T06:46:44.000Z","created_at_i":1691218004,"id":37009602,"options":[],"parent_id":37009563,"points":null,"story_id":37009272,"text":"Perhaps some sort of memristor crossbar array that could process large array calculations efficiently? Something like that could be pretty universal, until the tech moves to spiking neural networks. And even then the knowledge you gain from advancing on that path would likely transfer over.","title":null,"type":"comment","url":null}],"created_at":"2023-08-05T06:38:17.000Z","created_at_i":1691217497,"id":37009563,"options":[],"parent_id":37009272,"points":null,"story_id":37009272,"text":"I&#x27;m not in this field as others. so I&#x27;d really appreciate a comment from some  experienced HN folks.<p>I wonder what&#x27;s the end goal.<p>The ML&#x2F;AI world seems to be changing fast. So one model approach might become &quot;old tech&quot; if something better comes.<p>Is our current state of the models is stable so the LLM  glory would be the same in 1-2 years? Or suddenly there&#x27;ll be new approach making this &quot;tech&quot; obsolete?","title":null,"type":"comment","url":null},{"author":"psychphysic","children":[],"created_at":"2023-08-05T07:00:36.000Z","created_at_i":1691218836,"id":37009677,"options":[],"parent_id":37009272,"points":null,"story_id":37009272,"text":"They need to get ChatGPT to write a cooler website for them.","title":null,"type":"comment","url":null},{"author":"mmaunder","children":[],"created_at":"2023-08-05T07:00:57.000Z","created_at_i":1691218857,"id":37009679,"options":[],"parent_id":37009272,"points":null,"story_id":37009272,"text":"It\u2019s like a startup optimizing Java applet performance as the web was taking off.","title":null,"type":"comment","url":null},{"author":"PeterStuer","children":[],"created_at":"2023-08-05T07:01:30.000Z","created_at_i":1691218890,"id":37009681,"options":[],"parent_id":37009272,"points":null,"story_id":37009272,"text":"&quot;Be the compute platform for AGI&quot; and &quot;we target just LLMs&quot; feel to me to be contradicting statements.","title":null,"type":"comment","url":null},{"author":"cpgxiii","children":[{"author":"brucethemoose2","children":[],"created_at":"2023-08-05T14:25:45.000Z","created_at_i":1691245545,"id":37012308,"options":[],"parent_id":37009703,"points":null,"story_id":37009272,"text":"&gt; So long as Pytorch only practically works with Nvidia GPUs, everything else is little more than a rounding error.<p>This is changing.<p><a href=\"https:&#x2F;&#x2F;github.com&#x2F;merrymercy&#x2F;awesome-tensor-compilers\">https:&#x2F;&#x2F;github.com&#x2F;merrymercy&#x2F;awesome-tensor-compilers</a><p>There are more and better projects that can compile an existing PyTorch codebase into a more optimized format for a range of devices. Triton (which is part of PyTorch) TVM and the MLIR based efforts (like torch-MLIR or IREE) are big ones, but there are smaller fish like GGML and Tinygrad, or more narrowly focused projects like Meta&#x27;s AITemplate (which works on AMD datacenter GPUs).<p>Hardware is in a strange place now... It feels like everyone but Cerebras and AMD&#x2F;Intel was squeezed out, but with all the money pouring in, I think this is temporary.","title":null,"type":"comment","url":null}],"created_at":"2023-08-05T07:05:27.000Z","created_at_i":1691219127,"id":37009703,"options":[],"parent_id":37009272,"points":null,"story_id":37009272,"text":"I don&#x27;t want to be too cynical about the state of hardware for ML, but I don&#x27;t see where this is going. Nvidia does not lack for competitors trying (and sometimes nominally succeeding) to build faster&#x2F;cheaper&#x2F;more efficient hardware. Yet still Nvidia is overwhelmingly the vendor of choice because the software story works. So long as Pytorch only practically works with Nvidia GPUs, everything else is little more than a rounding error.<p>I don&#x27;t see MatX ending up any different than the legion of startups that have come already - either they get acquired by a bigger player, or they fade into obscurity.","title":null,"type":"comment","url":null},{"author":"dang","children":[{"author":"modeless","children":[{"author":"dang","children":[],"created_at":"2023-08-06T01:31:23.000Z","created_at_i":1691285483,"id":37018078,"options":[],"parent_id":37009810,"points":null,"story_id":37009272,"text":"I agree, it&#x27;s a nice layout.","title":null,"type":"comment","url":null}],"created_at":"2023-08-05T07:25:54.000Z","created_at_i":1691220354,"id":37009810,"options":[],"parent_id":37009727,"points":null,"story_id":37009272,"text":"I love the presentation though. This is actually the best startup homepage I&#x27;ve ever seen. Hand-typed HTML, 20 lines of CSS, no JavaScript, not even any images. Just some powerful bullet points that actually contain relevant information about the company.<p>A lack of information about the technology is understandable as it presumably doesn&#x27;t exist yet. But I see that it&#x27;s Google TPU and software people striking out on their own with the backing of a bunch of VCs and prominent researchers, and it seems pretty clear where it&#x27;s going. Likely an evolution of or progression from what they were working on at Google, positioned for easy acquisition by any number of competitors.<p>I think George Hotz is right that ML ASIC companies need to invest in software far more than they typically do. I guess the plan here is to sidestep the need for generic software by focusing solely on transformers for text, maybe even only a couple of specific architectures. I don&#x27;t really think that&#x27;s a good strategy but it seems to be what they&#x27;re describing.","title":null,"type":"comment","url":null},{"author":"Tuna-Fish","children":[{"author":"dang","children":[],"created_at":"2023-08-06T01:29:40.000Z","created_at_i":1691285380,"id":37018065,"options":[],"parent_id":37010229,"points":null,"story_id":37009272,"text":"That&#x27;s not enough information to make the post count as substantive. This is not a borderline call!<p>The OP is something between a landing page and a job ad, not an in-depth article. This is not a borderline call!","title":null,"type":"comment","url":null}],"created_at":"2023-08-05T08:52:04.000Z","created_at_i":1691225524,"id":37010229,"options":[],"parent_id":37009727,"points":null,"story_id":37009272,"text":"The how is on the page, after no more than 25 words:<p>&gt; Approach:<p>&gt; * We target just LLMs, whereas GPUs target all ML models.\n     LLMs are different. Our hardware and software can be much simpler.<p>&gt; * We combine deep domain experience, a few key ideas, and a lot of careful engineering.","title":null,"type":"comment","url":null}],"created_at":"2023-08-05T07:10:38.000Z","created_at_i":1691219438,"id":37009727,"options":[],"parent_id":37009272,"points":null,"story_id":37009272,"text":"This post doesn&#x27;t have enough information to count as a good HN submission. Because it doesn&#x27;t have enough information, the comments are almost all generic. Generic comments don&#x27;t make for good threads\u2014there needs to be specific information for people to sink their teeth into.<p>It would be be better to write a post about <i>how</i> you&#x27;re making faster chips for LLMs, that everybody can learn something from.","title":null,"type":"comment","url":null},{"author":"DeathArrow","children":[],"created_at":"2023-08-05T07:11:50.000Z","created_at_i":1691219510,"id":37009732,"options":[],"parent_id":37009272,"points":null,"story_id":37009272,"text":"It is possible to build ASICS for a specific LLM? Would that be cost effective?","title":null,"type":"comment","url":null},{"author":"fragmede","children":[{"author":"ilaksh","children":[],"created_at":"2023-08-05T13:04:21.000Z","created_at_i":1691240661,"id":37011622,"options":[],"parent_id":37009868,"points":null,"story_id":37009272,"text":"AMD drivers are a higher priority but he also made tinygrad <a href=\"https:&#x2F;&#x2F;github.com&#x2F;tinygrad&#x2F;tinygrad\">https:&#x2F;&#x2F;github.com&#x2F;tinygrad&#x2F;tinygrad</a>\nwhich is basically designed to minimize the software complexity and operations that a company like MatX needs to target. It&#x27;s actually kind of a perfect complement to tinygrad.","title":null,"type":"comment","url":null}],"created_at":"2023-08-05T07:37:05.000Z","created_at_i":1691221025,"id":37009868,"options":[],"parent_id":37009272,"points":null,"story_id":37009272,"text":"It&#x27;s interesting to contrast this with tinycorp, geohot&#x27;s company, and his claim that basically you would have to spend a lot to make a better chip, so the more optimal, less capital intensive play, is to write better drivers for AMD cards.<p><a href=\"https:&#x2F;&#x2F;geohot.github.io&#x2F;blog&#x2F;jekyll&#x2F;update&#x2F;2023&#x2F;05&#x2F;24&#x2F;the-tiny-corp-raised-5M.html\" rel=\"nofollow noreferrer\">https:&#x2F;&#x2F;geohot.github.io&#x2F;blog&#x2F;jekyll&#x2F;update&#x2F;2023&#x2F;05&#x2F;24&#x2F;the-t...</a><p>He&#x27;s still on the AMD drivers train, judging by his Twitter post from 4 days ago, so we&#x27;ll see where things go.<p><a href=\"https:&#x2F;&#x2F;twitter.com&#x2F;realgeorgehotz&#x2F;status&#x2F;1686165811386597377\" rel=\"nofollow noreferrer\">https:&#x2F;&#x2F;twitter.com&#x2F;realgeorgehotz&#x2F;status&#x2F;168616581138659737...</a>","title":null,"type":"comment","url":null}],"created_at":"2023-08-05T05:30:18.000Z","created_at_i":1691213418,"id":37009272,"options":[],"parent_id":null,"points":55,"story_id":37009272,"text":null,"title":"MatX: Faster Chips for LLMs","type":"story","url":"https://matx.com"}
