{"author":"djha-skin","children":[{"author":"manojlds","children":[{"author":"dang","children":[],"created_at":"2023-12-01T21:12:19.000Z","created_at_i":1701465139,"id":38492473,"options":[],"parent_id":38490421,"points":null,"story_id":38489533,"text":"Thanks! Macroexpanded:<p><i>Llamafile lets you distribute and run LLMs with a single file</i> - <a href=\"https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=38464057\">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=38464057</a> - Nov 2023 (273 comments)","title":null,"type":"comment","url":null}],"created_at":"2023-12-01T18:33:55.000Z","created_at_i":1701455635,"id":38490421,"options":[],"parent_id":38489533,"points":null,"story_id":38489533,"text":"Previous: <a href=\"https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=38464057\">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=38464057</a>","title":null,"type":"comment","url":null},{"author":"behnamoh","children":[],"created_at":"2023-12-01T18:50:24.000Z","created_at_i":1701456624,"id":38490616,"options":[],"parent_id":38489533,"points":null,"story_id":38489533,"text":"dupe","title":null,"type":"comment","url":null},{"author":"bilsbie","children":[{"author":"_neil","children":[{"author":"robterrell","children":[],"created_at":"2023-12-01T19:41:32.000Z","created_at_i":1701459692,"id":38491226,"options":[],"parent_id":38490925,"points":null,"story_id":38489533,"text":"Don&#x27;t skimp on RAM.","title":null,"type":"comment","url":null}],"created_at":"2023-12-01T19:16:43.000Z","created_at_i":1701458203,"id":38490925,"options":[],"parent_id":38490716,"points":null,"story_id":38489533,"text":"Maybe an M2 Mac mini?","title":null,"type":"comment","url":null},{"author":"superkuh","children":[],"created_at":"2023-12-01T19:26:31.000Z","created_at_i":1701458791,"id":38491043,"options":[],"parent_id":38490716,"points":null,"story_id":38489533,"text":"Pretty much any modern machine will do you fine for ~5 tokens&#x2F;s on a 7B (small end) model like Mistral-7B or llama2-chat-7B (or any of their respective fine-tunes). The computer you already have can probably do this.","title":null,"type":"comment","url":null},{"author":"StillBored","children":[{"author":"nielsole","children":[{"author":"bilsbie","children":[],"created_at":"2023-12-02T02:42:39.000Z","created_at_i":1701484959,"id":38495373,"options":[],"parent_id":38491736,"points":null,"story_id":38489533,"text":"Do you specify the 5 bit at runtime or it\u2019s a certain download?","title":null,"type":"comment","url":null}],"created_at":"2023-12-01T20:18:08.000Z","created_at_i":1701461888,"id":38491736,"options":[],"parent_id":38491117,"points":null,"story_id":38489533,"text":"For reference: I run 7B 5bit quantized models on a ryzen 7 5700G with 64Gb at 8 tokens&#x2F;second CPU only.\nIt&#x27;s not close to what you can get with a high end graphics card but for every day use it is alright and has headroom for bigger models. Upgrading CPU, RAM and Mainboard in a 10+year old PC cost me just 400\u20ac","title":null,"type":"comment","url":null}],"created_at":"2023-12-01T19:32:04.000Z","created_at_i":1701459124,"id":38491117,"options":[],"parent_id":38490716,"points":null,"story_id":38489533,"text":"It can run on CPU cores if you have enough RAM for the model. It feels like you have time warped back to the early 1990&#x27;s and are talking to someone on a BBS (AKA the words appear slowly), but it is entirely functional if you have 32+GB of RAM.","title":null,"type":"comment","url":null},{"author":"tarruda","children":[{"author":"mratsim","children":[],"created_at":"2023-12-01T21:04:57.000Z","created_at_i":1701464697,"id":38492377,"options":[],"parent_id":38492115,"points":null,"story_id":38489533,"text":"Hopefully, AMD wakes up and allows ROCm &#x2F; HIP on their 7040 series, that would be a killer feature.<p>And AFAIK the Ryzen pro do have an &quot;AI coprocessor&quot; but unsure of the API to use them.","title":null,"type":"comment","url":null}],"created_at":"2023-12-01T20:45:40.000Z","created_at_i":1701463540,"id":38492115,"options":[],"parent_id":38490716,"points":null,"story_id":38489533,"text":"I haven&#x27;t tested, but I suspect you might be able to get good CPU inference on this mini pc: <a href=\"https:&#x2F;&#x2F;www.aliexpress.com&#x2F;item&#x2F;1005005825981362.html\" rel=\"nofollow noreferrer\">https:&#x2F;&#x2F;www.aliexpress.com&#x2F;item&#x2F;1005005825981362.html</a>. It costs ~$800 in the maxed configuration with 64GB RAM clocked at 5600Mhz","title":null,"type":"comment","url":null},{"author":"swuecho","children":[],"created_at":"2023-12-01T20:51:38.000Z","created_at_i":1701463898,"id":38492201,"options":[],"parent_id":38490716,"points":null,"story_id":38489533,"text":"mac mini 8G runs at 10tokens&#x2F;sec","title":null,"type":"comment","url":null}],"created_at":"2023-12-01T18:59:28.000Z","created_at_i":1701457168,"id":38490716,"options":[],"parent_id":38489533,"points":null,"story_id":38489533,"text":"Is there a good machine to buy in the $500-1000 range that is able to run some of this stuff?","title":null,"type":"comment","url":null},{"author":"lxe","children":[{"author":"ComputerGuru","children":[{"author":"idonotknowwhy","children":[{"author":"Rastonbury","children":[{"author":"idonotknowwhy","children":[{"author":"ComputerGuru","children":[{"author":"idonotknowwhy","children":[{"author":"ComputerGuru","children":[],"created_at":"2023-12-04T07:11:52.000Z","created_at_i":1701673912,"id":38514479,"options":[],"parent_id":38503513,"points":null,"story_id":38489533,"text":"Thanks, friend!","title":null,"type":"comment","url":null}],"created_at":"2023-12-03T00:10:55.000Z","created_at_i":1701562255,"id":38503513,"options":[],"parent_id":38500124,"points":null,"story_id":38489533,"text":"Most of the time it&#x27;s readily available eg:<p>Panchovix&#x2F;goliath-120b-exl2 (there&#x27;s a different branch for each size)<p>Some of them I&#x27;ve had to do myself eg. I wanted a Q2 GGUF of Falcon 180b<p>There&#x27;s a guy on huggingface called &quot;TheBloke&quot; who does GGUF, AWQ and GPTQ for most models. For exl2, you can usually just search for exl2 and find them.","title":null,"type":"comment","url":null}],"created_at":"2023-12-02T17:21:36.000Z","created_at_i":1701537696,"id":38500124,"options":[],"parent_id":38497787,"points":null,"story_id":38489533,"text":"Did you have to quantize it yourself to 4.75bpw and 3bpw or are they readily available for download?","title":null,"type":"comment","url":null}],"created_at":"2023-12-02T11:17:15.000Z","created_at_i":1701515835,"id":38497787,"options":[],"parent_id":38496442,"points":null,"story_id":38489533,"text":"It is. I have 48GB of VRAM. But exl2 is more efficient, and can be quantized to partial bits. So you can run things like 4.75bpw, etc.<p>I can run 120b models at 3bpw.<p>The larger models like this are less affected (increased perplexity) by the quantization.","title":null,"type":"comment","url":null}],"created_at":"2023-12-02T06:27:05.000Z","created_at_i":1701498425,"id":38496442,"options":[],"parent_id":38492428,"points":null,"story_id":38489533,"text":"Wait, so using large models isn&#x27;t limited by VRAM anymore?","title":null,"type":"comment","url":null},{"author":"idonotknowwhy","children":[],"created_at":"2023-12-02T11:17:48.000Z","created_at_i":1701515868,"id":38497791,"options":[],"parent_id":38492428,"points":null,"story_id":38489533,"text":"Typo: I meant &quot;almost none if you already have python installed&quot;","title":null,"type":"comment","url":null}],"created_at":"2023-12-01T21:09:22.000Z","created_at_i":1701464962,"id":38492428,"options":[],"parent_id":38491877,"points":null,"story_id":38489533,"text":"Almost none of you already have python. Download exl2, exui from github and run a few terminal commands. This let&#x27;s me run the 120b param models, which won&#x27;t fit in vram if I use llamacpp","title":null,"type":"comment","url":null}],"created_at":"2023-12-01T20:28:24.000Z","created_at_i":1701462504,"id":38491877,"options":[],"parent_id":38490938,"points":null,"story_id":38489533,"text":"How much more work is it to get those up and running?","title":null,"type":"comment","url":null}],"created_at":"2023-12-01T19:17:51.000Z","created_at_i":1701458271,"id":38490938,"options":[],"parent_id":38489533,"points":null,"story_id":38489533,"text":"Keep in mind that gguf&#x2F;llama.cpp, although highly performant and portable, is not the best performing way to launch certain models if you have a GPU. (even though llama.cpp does support GPU acceleration)<p>ExLLAmA v2 + elx2 quantization, and maybe tensorrt-llm might be the contender for the top performer","title":null,"type":"comment","url":null},{"author":"superkuh","children":[{"author":"breckenedge","children":[{"author":"dragonwriter","children":[],"created_at":"2023-12-02T00:17:04.000Z","created_at_i":1701476224,"id":38494429,"options":[],"parent_id":38491520,"points":null,"story_id":38489533,"text":"Or you can download LM Studio or oobabooga and execute any model compiled to GGUF format, and also models in a wide variety of formats that aren&#x27;t GGUF.","title":null,"type":"comment","url":null}],"created_at":"2023-12-01T20:02:50.000Z","created_at_i":1701460970,"id":38491520,"options":[],"parent_id":38490942,"points":null,"story_id":38489533,"text":"Per the article:<p>&gt; You can also download a much smaller llamafile binary from their releases, which can then execute any model that has been compiled to GGUF format","title":null,"type":"comment","url":null},{"author":"rodrigobellusci","children":[],"created_at":"2023-12-01T20:07:28.000Z","created_at_i":1701461248,"id":38491577,"options":[],"parent_id":38490942,"points":null,"story_id":38489533,"text":"From what I&#x27;ve read llamafile can also load other models if you pass a flag with a path to them. There&#x27;s also LM Studio for anyone who&#x27;d like to play with a ChatGPT-like GUI.","title":null,"type":"comment","url":null},{"author":"brandall10","children":[],"created_at":"2023-12-01T20:16:37.000Z","created_at_i":1701461797,"id":38491711,"options":[],"parent_id":38490942,"points":null,"story_id":38489533,"text":"I&#x27;d argue LM Studio is the best way by far.<p>- Download and launch the app<p>- Use their interface to huggingface models, download the one you want<p>- Go to the chat window and select load<p>- Chat away w&#x2F; a ChatGPT like interface that includes markdown processing and easy copying of code blocks, etc<p>Outside of the time to download a model, it&#x27;s about 30 seconds of work to get up and running","title":null,"type":"comment","url":null}],"created_at":"2023-12-01T19:18:24.000Z","created_at_i":1701458304,"id":38490942,"options":[],"parent_id":38489533,"points":null,"story_id":38489533,"text":"It&#x27;s not the best way. It&#x27;s a really cool and technically interesting way. But embedding the model with the executable is terrible for anything beyond a demo. The best way would be just running llama.cpp (what llamafile uses) and loading external models. I get that compiling is too difficult for some people and llamafiles are great for them. Maybe even the best for them.<p>But it&#x27;s far from best for people that actually want to explore LLM and play.","title":null,"type":"comment","url":null},{"author":"leonidbelyaev","children":[{"author":"SkyMarshal","children":[{"author":"psanford","children":[{"author":"SkyMarshal","children":[],"created_at":"2023-12-01T22:35:44.000Z","created_at_i":1701470144,"id":38493504,"options":[],"parent_id":38492312,"points":null,"story_id":38489533,"text":"Oh lol, I&#x27;ve never heard of &#x2F;bin&#x2F;mkdir even on regular FHS linuxes.  Is that actually a thing?","title":null,"type":"comment","url":null}],"created_at":"2023-12-01T21:01:11.000Z","created_at_i":1701464471,"id":38492312,"options":[],"parent_id":38491991,"points":null,"story_id":38489533,"text":"It fails because it hardcodes the path to certain binaries instead of looking in PATH: `line 60: &#x2F;bin&#x2F;mkdir: No such file or directory`","title":null,"type":"comment","url":null}],"created_at":"2023-12-01T20:37:12.000Z","created_at_i":1701463032,"id":38491991,"options":[],"parent_id":38491696,"points":null,"story_id":38489533,"text":"Why not, what\u2019s the error, and how did you set it up?  Any chance you tried it in a nix shell or nix vm or other isolated environment?","title":null,"type":"comment","url":null},{"author":"FragenAntworten","children":[],"created_at":"2023-12-01T21:05:03.000Z","created_at_i":1701464703,"id":38492380,"options":[],"parent_id":38491696,"points":null,"story_id":38489533,"text":"See this comment from the other discussion: <a href=\"https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=38469749\">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=38469749</a>","title":null,"type":"comment","url":null}],"created_at":"2023-12-01T20:15:36.000Z","created_at_i":1701461736,"id":38491696,"options":[],"parent_id":38489533,"points":null,"story_id":38489533,"text":"Doesn&#x27;t work on my NixOS workstation.","title":null,"type":"comment","url":null},{"author":"rasengan","children":[],"created_at":"2023-12-01T20:16:58.000Z","created_at_i":1701461818,"id":38491714,"options":[],"parent_id":38489533,"points":null,"story_id":38489533,"text":"This is probably true for certain scenarios, but it&#x27;s absolutely not true for all scenarios.","title":null,"type":"comment","url":null},{"author":"Akashic101","children":[{"author":"bonniemuffin","children":[],"created_at":"2023-12-01T20:35:52.000Z","created_at_i":1701462952,"id":38491980,"options":[],"parent_id":38491726,"points":null,"story_id":38489533,"text":"How did you choose between training a model from scratch vs using retrieval augmented generation with an existing off-the-shelf model? From what I&#x27;ve observed, RAG + off-the-shelf model seems to be the more common approach for use cases like &quot;create LLM that answers questions about my company&#x27;s internal documentation&quot;, particularly because the iteration&#x2F;improvement cycle is much shorter-- it&#x27;s much easier to iterate on RAG&#x2F;prompts vs. training a whole new model to improve it.\n(If the answer is &quot;I just wanted to try training a whole new llm&quot;, I won&#x27;t fault you for that! :) )","title":null,"type":"comment","url":null},{"author":"euroderf","children":[],"created_at":"2023-12-02T12:03:43.000Z","created_at_i":1701518623,"id":38498006,"options":[],"parent_id":38491726,"points":null,"story_id":38489533,"text":"I would like to be able to train an LLM on absolutely everything I have starred at Github.","title":null,"type":"comment","url":null},{"author":"mud_dauber","children":[],"created_at":"2023-12-03T13:18:15.000Z","created_at_i":1701609495,"id":38506839,"options":[],"parent_id":38491726,"points":null,"story_id":38489533,"text":"Seconded. My copy launched without problem - now I want to learn how to ingest my library of blog posts, PDFs &amp; images.","title":null,"type":"comment","url":null}],"created_at":"2023-12-01T20:17:32.000Z","created_at_i":1701461852,"id":38491726,"options":[],"parent_id":38489533,"points":null,"story_id":38489533,"text":"I am currently planning my own small LLM trained on documents we use internally for work. Does anyone have any tips and tricks on how to make this work the best? Could a project like Llamafile help me with this, even if it is just for testing?","title":null,"type":"comment","url":null},{"author":"youniverse","children":[{"author":"buffington","children":[],"created_at":"2023-12-01T21:13:16.000Z","created_at_i":1701465196,"id":38492486,"options":[],"parent_id":38492228,"points":null,"story_id":38489533,"text":"Without knowing what you consider to be of worth, that&#x27;s difficult to answer.<p>The good news is that the article describes in detail how to determine for yourself if it&#x27;s worth it. The author uses an M2 as well, so you can at least know it&#x27;ll likely work.","title":null,"type":"comment","url":null},{"author":"MPSimmons","children":[],"created_at":"2023-12-01T21:14:10.000Z","created_at_i":1701465250,"id":38492497,"options":[],"parent_id":38492228,"points":null,"story_id":38489533,"text":"It depends on what you want them to do and what you want to get out of them. A lot of people use less machine than a M2 Air for different things.","title":null,"type":"comment","url":null},{"author":"brandall10","children":[{"author":"iJohnDoe","children":[],"created_at":"2023-12-02T03:09:37.000Z","created_at_i":1701486577,"id":38495525,"options":[],"parent_id":38492605,"points":null,"story_id":38489533,"text":"Is there any reason to run this if you already have LM Studio installed?<p>Which LLM to download for general questions&#x2F;info that has good accuracy?<p>Separately, I couldn\u2019t figure out how to use the NVIDIA GPU with LM Studio. Only wants to use the Intel card. I use NVIDIA with Diffusion with no problems if that makes any difference.","title":null,"type":"comment","url":null}],"created_at":"2023-12-01T21:21:26.000Z","created_at_i":1701465686,"id":38492605,"options":[],"parent_id":38492228,"points":null,"story_id":38489533,"text":"If you have 8GB you can play around with heavily quantized 7B models, or up to moderately quantized 13GB models w&#x2F; 16GB.<p>Something like a Mistral-7B Dolphin finetune actually is surprisingly useful, like GPT-3.5 in some respects. I imagine it would render ~10 tok&#x2F;sec on an M2 for at least short bursts until throttling sets in.<p>As I mentioned in another post, try this out with LM Studio. Super simple GUI that even a non-tech person could probably figure out for finding&#x2F;downloading&#x2F;loading models w&#x2F; a ChatGPT like interface for chatting.<p><a href=\"https:&#x2F;&#x2F;lmstudio.ai&#x2F;\" rel=\"nofollow noreferrer\">https:&#x2F;&#x2F;lmstudio.ai&#x2F;</a>","title":null,"type":"comment","url":null},{"author":"dataveg","children":[],"created_at":"2023-12-01T23:16:17.000Z","created_at_i":1701472577,"id":38493909,"options":[],"parent_id":38492228,"points":null,"story_id":38489533,"text":"I have Ollama running llama 7B on an 32GB Mac M1 Pro and its fast enough to be useful.","title":null,"type":"comment","url":null}],"created_at":"2023-12-01T20:54:17.000Z","created_at_i":1701464057,"id":38492228,"options":[],"parent_id":38489533,"points":null,"story_id":38489533,"text":"Anyone know if this (or any LLM) is worth running locally on a Macbook M2 Air?<p>I&#x27;ll give it a shot over the weekend but if anyone knows I&#x27;m curious!","title":null,"type":"comment","url":null},{"author":"dragonwriter","children":[],"created_at":"2023-12-01T21:02:34.000Z","created_at_i":1701464554,"id":38492334,"options":[],"parent_id":38489533,"points":null,"story_id":38489533,"text":"Best in what sense? Seems by description a whole lot less convenient (and more limited) if you want to use more than one model (and especially more than just models available as GGUFs) than oobabooga&#x2F;text-generation-webui or other similar tools that do things like bundle multiple backends for different LLM architectures, support downloading models from huggingface, present a common web UI for LLM configuration, managing prompts, and actually doing inference (in chat, notebook, and other styles), and also supports presenting an OpenAI-compatible API endpoint backed with a local LLM to support other frontends.","title":null,"type":"comment","url":null},{"author":"rismay","children":[],"created_at":"2023-12-01T21:46:04.000Z","created_at_i":1701467164,"id":38492916,"options":[],"parent_id":38489533,"points":null,"story_id":38489533,"text":"Will this approach work with image generation models?","title":null,"type":"comment","url":null},{"author":"jawilson","children":[],"created_at":"2023-12-05T06:01:53.000Z","created_at_i":1701756113,"id":38527594,"options":[],"parent_id":38489533,"points":null,"story_id":38489533,"text":"It would be nice if input could be taken from a command line argument or better yet, stdin so that it is fully scriptable.<p>ollama has a way to do this and lets you play with a bunch of models without being very smart (just do ollama pull &lt;model-name&gt; and it downloads a model and makes it available to ollama run&#x2F;ollama serve).<p>(Sounds like I also need to play with LM Studio).","title":null,"type":"comment","url":null},{"author":"maxloo1976","children":[],"created_at":"2023-12-12T18:12:33.000Z","created_at_i":1702404753,"id":38616014,"options":[],"parent_id":38489533,"points":null,"story_id":38489533,"text":"I&#x27;ve tried running Llamafile on my Lenovo Legion Pro 5 laptop with 8GB VRAM, but it has a dashboard that shows the GPU and CPU utilisation in real time, and almost all the processing is done on the CPU.  Is there a way to shift the processing to the GPU like using gpt-fast?<p><a href=\"https:&#x2F;&#x2F;pytorch.org&#x2F;blog&#x2F;accelerating-generative-ai-2&#x2F;\" rel=\"nofollow noreferrer\">https:&#x2F;&#x2F;pytorch.org&#x2F;blog&#x2F;accelerating-generative-ai-2&#x2F;</a>","title":null,"type":"comment","url":null}],"created_at":"2023-12-01T17:36:50.000Z","created_at_i":1701452210,"id":38489533,"options":[],"parent_id":null,"points":195,"story_id":38489533,"text":null,"title":"Llamafile is the new best way to run a LLM on your own computer","type":"story","url":"http://simonwillison.net/2023/Nov/29/llamafile/"}
