{"exhaustive":{"nbHits":false,"typo":false},"exhaustiveNbHits":false,"exhaustiveTypo":false,"hits":[{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"mmaunder"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vllm","fp8","benchmark"],"value":"Nah. Try <em>vLLM</em> and 405B <em>FP8</em> on that hardware. And make sure you\u2019re <em>benchmark</em>ing with some concurrency for max TPS."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Llama 3.1 405B now runs at 969 tokens/s on Cerebras Inference"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://cerebras.ai/blog/llama-405b-inference"}},"_tags":["comment","author_mmaunder","story_42178761"],"author":"mmaunder","children":[42189943],"comment_text":"Nah. Try vLLM and 405B FP8 on that hardware. And make sure you\u2019re benchmarking with some concurrency for max TPS.","created_at":"2024-11-19T03:16:24Z","created_at_i":1731986184,"objectID":"42179794","parent_id":42179476,"story_id":42178761,"story_title":"Llama 3.1 405B now runs at 969 tokens/s on Cerebras Inference","story_url":"https://cerebras.ai/blog/llama-405b-inference","updated_at":"2024-11-20T01:28:22Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"anonova"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vllm","fp8","benchmark"],"value":"<em>vLLM</em>'s study also concluded that &quot;<em>FP8</em> can deliver meaningful latency and capacity gains with small or negligible accuracy loss&quot;. Their <em>benchmarks</em> include LiveCodeBench 6.<p><a href=\"https://vllm-project.github.io/2026/04/22/fp8-kvcache.html\" rel=\"nofollow\">https://<em>vllm</em>-project.github.io/2026/04/22/<em>fp8</em>-kvcache.html</a>"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Smaller, faster, safer: running Kimi and GLM at scale"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://blog.cloudflare.com/smaller-faster-safer-models/"}},"_tags":["comment","author_anonova","story_49158581"],"author":"anonova","children":[49162302,49165158],"comment_text":"vLLM&#x27;s study also concluded that &quot;FP8 can deliver meaningful latency and capacity gains with small or negligible accuracy loss&quot;. Their benchmarks include LiveCodeBench 6.<p><a href=\"https:&#x2F;&#x2F;vllm-project.github.io&#x2F;2026&#x2F;04&#x2F;22&#x2F;fp8-kvcache.html\" rel=\"nofollow\">https:&#x2F;&#x2F;vllm-project.github.io&#x2F;2026&#x2F;04&#x2F;22&#x2F;fp8-kvcache.html</a>","created_at":"2026-08-03T20:40:57Z","created_at_i":1785789657,"objectID":"49161132","parent_id":49160216,"story_id":49158581,"story_title":"Smaller, faster, safer: running Kimi and GLM at scale","story_url":"https://blog.cloudflare.com/smaller-faster-safer-models/","updated_at":"2026-08-04T12:30:01Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"reissbaker"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vllm","fp8","benchmark"],"value":"Basically no one uses FP32 at inference time. BF16/FP16 is typically considered unquantized, whereas <em>FP8</em> is lightly quantized. That being said there's pretty minimal quality loss at <em>FP8</em> compared to 16-bit typically; Llama 3.1 405b, for example, only <em>benchmarks</em> around ~1% worse when run at <em>FP8</em>: <a href=\"https://blog.vllm.ai/2024/07/23/llama31.html\" rel=\"nofollow\">https://blog.<em>vllm</em>.ai/2024/07/23/llama31.html</a><p>Every major inference provider other than Hyperbolic Labs runs Llama 3.1 405b at <em>FP8</em>, FWIW (e.g. Together, Fireworks, Lepton), so to compare against FP32 is misleading to say the least. Even Hyperbolic runs it at BF16.<p>Pretraining is typically done in FP32, although some labs (e.g. Character AI, RIP) apparently train in INT8: <a href=\"https://research.character.ai/optimizing-inference/\" rel=\"nofollow\">https://research.character.ai/optimizing-inference/</a>"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Addition Is All You Need for Energy-Efficient Language Models"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://arxiv.org/abs/2410.00907"}},"_tags":["comment","author_reissbaker","story_41784591"],"author":"reissbaker","children":[41796568],"comment_text":"Basically no one uses FP32 at inference time. BF16&#x2F;FP16 is typically considered unquantized, whereas FP8 is lightly quantized. That being said there&#x27;s pretty minimal quality loss at FP8 compared to 16-bit typically; Llama 3.1 405b, for example, only benchmarks around ~1% worse when run at FP8: <a href=\"https:&#x2F;&#x2F;blog.vllm.ai&#x2F;2024&#x2F;07&#x2F;23&#x2F;llama31.html\" rel=\"nofollow\">https:&#x2F;&#x2F;blog.vllm.ai&#x2F;2024&#x2F;07&#x2F;23&#x2F;llama31.html</a><p>Every major inference provider other than Hyperbolic Labs runs Llama 3.1 405b at FP8, FWIW (e.g. Together, Fireworks, Lepton), so to compare against FP32 is misleading to say the least. Even Hyperbolic runs it at BF16.<p>Pretraining is typically done in FP32, although some labs (e.g. Character AI, RIP) apparently train in INT8: <a href=\"https:&#x2F;&#x2F;research.character.ai&#x2F;optimizing-inference&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;research.character.ai&#x2F;optimizing-inference&#x2F;</a>","created_at":"2024-10-09T13:55:06Z","created_at_i":1728482106,"objectID":"41788030","parent_id":41787469,"story_id":41784591,"story_title":"Addition Is All You Need for Energy-Efficient Language Models","story_url":"https://arxiv.org/abs/2410.00907","updated_at":"2024-10-12T16:38:08Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"Palmik"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vllm","fp8","benchmark"],"value":"It would be great to have real world inference <em>benchmarks</em> for LLMs. These aren't it.<p>That means e.g. 8xH100 with TensorRT-LLM / <em>vLLM</em> vs 8xMI300X with <em>vLLM</em> running many concurrent requests with reasonable # of input and output tokens. Ran both in <em>fp8</em> and fp16.<p>Most of the <em>benchmarks</em> I've seen had setups that no one would use in production. For example running on a single MI300X or 2xH100 -- this will likely be memory bound, you need to go to higher batch sizes (more VRAM) to be compute bound to properly utilize these. Or benchmarking requests with unrealistically low # of input tokens."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Testing AMD's Giant MI300X"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://chipsandcheese.com/2024/06/25/testing-amds-giant-mi300x/"}},"_tags":["comment","author_Palmik","story_40789919"],"author":"Palmik","children":[40798028],"comment_text":"It would be great to have real world inference benchmarks for LLMs. These aren&#x27;t it.<p>That means e.g. 8xH100 with TensorRT-LLM &#x2F; vLLM vs 8xMI300X with vLLM running many concurrent requests with reasonable # of input and output tokens. Ran both in fp8 and fp16.<p>Most of the benchmarks I&#x27;ve seen had setups that no one would use in production. For example running on a single MI300X or 2xH100 -- this will likely be memory bound, you need to go to higher batch sizes (more VRAM) to be compute bound to properly utilize these. Or benchmarking requests with unrealistically low # of input tokens.","created_at":"2024-06-26T04:39:24Z","created_at_i":1719376764,"objectID":"40796431","parent_id":40789919,"story_id":40789919,"story_title":"Testing AMD's Giant MI300X","story_url":"https://chipsandcheese.com/2024/06/25/testing-amds-giant-mi300x/","updated_at":"2024-09-20T17:21:45Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"juliensalinas"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vllm","fp8","benchmark"],"value":"Hey everyone, I\u2019ve been diving into the world of generative AI inference engines for quite some time at NLP Cloud, and I wanted to share some insights from a comparison I put together. I looked at four popular options\u2014NVIDIA\u2019s TensorRT-LLM, <em>vLLM</em>, Hugging Face\u2019s Text Generation Inference (TGI), and LMDeploy\u2014and ran some <em>benchmarks</em> to see how they stack up for real-world use cases. Thought this might spark some discussion here since I know a lot of you are working with LLMs or optimizing inference pipelines:<p>TensorRT-LLM<p>------------<p>NVIDIA\u2019s beast for GPU-accelerated inference. Built on TensorRT, it optimizes models with layer fusion, precision tuning (FP16, INT8, even <em>FP8</em>), and custom CUDA kernels.<p>Pros: Blazing fast on NVIDIA GPUs\u2014think sub-50ms latency for single requests on an A100 and ~700 tokens/sec at 100 concurrent users for LLaMA-3 70B Q4 (per BentoML <em>benchmarks</em>). Dynamic batching and tight integration with Triton Inference Server make it a throughput monster.<p>Cons: Setup can be complex if you\u2019re not already in the NVIDIA ecosystem. You need to deal with model compilation, and it\u2019s not super flexible for quick prototyping.<p><em>vLLM</em><p>----<p>Open-source champion for high-throughput inference. Uses PagedAttention to manage KV caches in chunks, cutting memory waste and boosting speed.<p>Pros: Easy to spin up (pip install, Python-friendly), and it\u2019s flexible\u2014runs on NVIDIA, AMD, even CPU. Throughput is solid (~600-650 tokens/sec at 100 users for LLaMA-3 70B Q4), and dynamic batching keeps it humming. Latency\u2019s decent at 60-80ms solo.<p>Cons: It\u2019s less optimized for single-request latency, so if you\u2019re building a chatbot with one user at a time, it might not shine as much. Also, it\u2019s still maturing\u2014some edge cases (like exotic model architectures) might not be supported.<p>Hugging Face TGI<p>----------------<p>Hugging Face\u2019s production-ready inference tool. Ties into their model hub (BERT, GPT, etc.) and uses Rust for speed, with continuous batching to keep GPUs busy.<p>Pros: Docker setup is quick, and it scales well. Latency\u2019s 50-70ms, throughput matches <em>vLLM</em> (~600-650 tokens/sec at 100 users). Bonus: built-in output filtering for safety. Perfect if you\u2019re already in the HF ecosystem.<p>Cons: Less raw speed than TensorRT-LLM, and memory can bloat with big batches. Feels a bit restrictive outside HF\u2019s world.<p>LMDeploy<p>--------<p>This Toolkit from the MMRazor/MMDeploy crew, focused on fast, efficient LLM deployment. Features TurboMind (a high-performance engine) and a PyTorch fallback, with persistent batching and blocked KV caching for speed.<p>Pros: Decoding speed is nuts\u2014up to 1.8x more requests/sec than <em>vLLM</em> on an A100. TurboMind pushes 4-bit inference 2.4x faster than FP16, hitting ~700 tokens/sec at 100 users (LLaMA-3 70B Q4). Low latency (40-60ms), easy one-command server setup, and it even handles multi-round chats efficiently by caching history.<p>Cons: TurboMind\u2019s picky\u2014doesn\u2019t support sliding window attention (e.g., Mistral) yet. Non-NVIDIA users get stuck with the slower PyTorch engine. Still, on NVIDIA GPUs, it\u2019s a performance beast.<p>What\u2019s your experience with these tools? Any hidden issues I missed? Or are there other inference engines that should be mentioned? Would love to hear your thoughts!<p>Julien"},"title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["vllm"],"value":"Comparing GenAI Inference Engines: TensorRT-LLM, <em>VLLM</em>, HF TGI, and LMDeploy"}},"_tags":["story","author_juliensalinas","story_43620472","ask_hn"],"author":"juliensalinas","children":[43620491],"created_at":"2025-04-08T11:32:20Z","created_at_i":1744111940,"num_comments":1,"objectID":"43620472","points":1,"story_id":43620472,"story_text":"Hey everyone, I\u2019ve been diving into the world of generative AI inference engines for quite some time at NLP Cloud, and I wanted to share some insights from a comparison I put together. I looked at four popular options\u2014NVIDIA\u2019s TensorRT-LLM, vLLM, Hugging Face\u2019s Text Generation Inference (TGI), and LMDeploy\u2014and ran some benchmarks to see how they stack up for real-world use cases. Thought this might spark some discussion here since I know a lot of you are working with LLMs or optimizing inference pipelines:<p>TensorRT-LLM<p>------------<p>NVIDIA\u2019s beast for GPU-accelerated inference. Built on TensorRT, it optimizes models with layer fusion, precision tuning (FP16, INT8, even FP8), and custom CUDA kernels.<p>Pros: Blazing fast on NVIDIA GPUs\u2014think sub-50ms latency for single requests on an A100 and ~700 tokens&#x2F;sec at 100 concurrent users for LLaMA-3 70B Q4 (per BentoML benchmarks). Dynamic batching and tight integration with Triton Inference Server make it a throughput monster.<p>Cons: Setup can be complex if you\u2019re not already in the NVIDIA ecosystem. You need to deal with model compilation, and it\u2019s not super flexible for quick prototyping.<p>vLLM<p>----<p>Open-source champion for high-throughput inference. Uses PagedAttention to manage KV caches in chunks, cutting memory waste and boosting speed.<p>Pros: Easy to spin up (pip install, Python-friendly), and it\u2019s flexible\u2014runs on NVIDIA, AMD, even CPU. Throughput is solid (~600-650 tokens&#x2F;sec at 100 users for LLaMA-3 70B Q4), and dynamic batching keeps it humming. Latency\u2019s decent at 60-80ms solo.<p>Cons: It\u2019s less optimized for single-request latency, so if you\u2019re building a chatbot with one user at a time, it might not shine as much. Also, it\u2019s still maturing\u2014some edge cases (like exotic model architectures) might not be supported.<p>Hugging Face TGI<p>----------------<p>Hugging Face\u2019s production-ready inference tool. Ties into their model hub (BERT, GPT, etc.) and uses Rust for speed, with continuous batching to keep GPUs busy.<p>Pros: Docker setup is quick, and it scales well. Latency\u2019s 50-70ms, throughput matches vLLM (~600-650 tokens&#x2F;sec at 100 users). Bonus: built-in output filtering for safety. Perfect if you\u2019re already in the HF ecosystem.<p>Cons: Less raw speed than TensorRT-LLM, and memory can bloat with big batches. Feels a bit restrictive outside HF\u2019s world.<p>LMDeploy<p>--------<p>This Toolkit from the MMRazor&#x2F;MMDeploy crew, focused on fast, efficient LLM deployment. Features TurboMind (a high-performance engine) and a PyTorch fallback, with persistent batching and blocked KV caching for speed.<p>Pros: Decoding speed is nuts\u2014up to 1.8x more requests&#x2F;sec than vLLM on an A100. TurboMind pushes 4-bit inference 2.4x faster than FP16, hitting ~700 tokens&#x2F;sec at 100 users (LLaMA-3 70B Q4). Low latency (40-60ms), easy one-command server setup, and it even handles multi-round chats efficiently by caching history.<p>Cons: TurboMind\u2019s picky\u2014doesn\u2019t support sliding window attention (e.g., Mistral) yet. Non-NVIDIA users get stuck with the slower PyTorch engine. Still, on NVIDIA GPUs, it\u2019s a performance beast.<p>What\u2019s your experience with these tools? Any hidden issues I missed? Or are there other inference engines that should be mentioned? Would love to hear your thoughts!<p>Julien","title":"Comparing GenAI Inference Engines: TensorRT-LLM, VLLM, HF TGI, and LMDeploy","updated_at":"2025-08-03T12:43:10Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"Facingsouth"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vllm","fp8","benchmark"],"value":"COMMANDS:<p># Install &amp; setup\npip install terradev-cli\nterradev configure --provider runpod<p># Price discovery (19 clouds)\nterradev quote -g H100\nterradev quote -g A100 --max-price 2.50<p># Provision with auto topology optimization\nterradev provision -g H100 -n 4 --parallel 6\nterradev provision -g A100 --dry-run\nterradev run --gpu H100 --image pytorch/pytorch:latest<p># Instance management\nterradev status --live\nterradev manage -i &lt;id&gt; -a stop\nterradev analytics --days 30\nterradev optimize\nTraining Pipeline\nbash\n# Pre-flight validation\nterradev preflight<p># Launch training\nterradev train --script train.py --from-provision latest\nterradev train-status\nterradev monitor --job my-job\nterradev checkpoint list --job my-job\nInference Optimization\nbash\n# <em>vLLM</em> auto-tuning (6 critical knobs)\nterradev <em>vllm</em> auto-optimize -s workload.json -m meta-llama/Llama-2-7b-hf -g 4\nterradev <em>vllm</em> analyze -e http://localhost:8000\nterradev <em>vllm</em> <em>benchmark</em> -e http://localhost:8000 -c 10<p># MoE deployment with auto-optimizations\nterradev provision --task clusters/moe-template/task.yaml \\\n  --set model_id=Qwen/Qwen3.5-397B-A17B<p># Disaggregated prefill/decode\nterradev ml ray --deploy-pd --model zai-org/GLM-5-<em>FP8</em> \\\n  --prefill-tp 8 --decode-tp 1 --decode-dp 24<p># LoRA adapters (hot-load on running endpoint)\nterradev lora add -e http://endpoint:8000 -n customer-a -p /adapters/a\nterradev lora list -e http://endpoint:8000\nterradev lora remove -e http://endpoint:8000 -n customer-a\nKubernetes\nbash\n# Topology-optimized clusters\nterradev k8s create my-cluster --gpu H100 --count 8 --prefer-spot\nterradev k8s list\nterradev k8s info my-cluster\nterradev k8s destroy my-cluster\nSecondary Features\nbash\n# HF Spaces (one-click deployment)\nterradev hf-space my-llama --model-id meta-llama/Llama-2-7b-hf --template llm<p># InferX serverless (&lt;2s cold starts)\nterradev inferx deploy --endpoint my-api --model-id meta-llama/Llama-2-7b-hf\nterradev inferx status --endpoint my-api<p># Observability &amp; Safety\nterradev phoenix deploy --project my-inference\nterradev phoenix spans --project my-inference --limit 100\nterradev qdrant create-collection --name docs --vector-size 1536\nterradev guardrails generate-config --enable-topical --enable-pii<p># GitOps automation\nterradev gitops init --provider github --repo my-org/infra --tool argocd\nterradev gitops sync --cluster production<p># Integrations (BYOAPI)\nterradev configure --provider wandb --api-key $WANDB_KEY\nterradev configure --provider prometheus --api-key $PROMETHEUS_URL\nQuick Workflows\nbash\n# 5-minute GPU setup\npip install terradev-cli &amp;&amp; terradev setup runpod --quick &amp;&amp; \\\nterradev quote -g H100 &amp;&amp; terradev run --gpu H100 --image pytorch/pytorch:latest<p># Production RAG pipeline\nterradev qdrant k8s --namespace rag &amp;&amp; \\\nterradev qdrant create-collection --name kb --vector-size 1536 &amp;&amp; \\\nterradev provision --task clusters/moe-template/task.yaml \\\n  --set model_id=Qwen/Qwen3.5-397B-A17B &amp;&amp; \\\nterradev phoenix deploy --project rag-pipeline<p># Multi-cloud cost optimization\nterradev analytics --days 30 &amp;&amp; terradev optimize<p>Key Features: Auto NUMA/RDMA topology optimization, 19-cloud price comparison, <em>vLLM</em> auto-tuning, disaggregated P/D, LoRA hot-loading, BYOAPI security model."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Terradev: A next-gen slash command CLI for GPU provisioning and management"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/theoddden/Terradev"}},"_tags":["comment","author_Facingsouth","story_47269668"],"author":"Facingsouth","comment_text":"COMMANDS:<p># Install &amp; setup\npip install terradev-cli\nterradev configure --provider runpod<p># Price discovery (19 clouds)\nterradev quote -g H100\nterradev quote -g A100 --max-price 2.50<p># Provision with auto topology optimization\nterradev provision -g H100 -n 4 --parallel 6\nterradev provision -g A100 --dry-run\nterradev run --gpu H100 --image pytorch&#x2F;pytorch:latest<p># Instance management\nterradev status --live\nterradev manage -i &lt;id&gt; -a stop\nterradev analytics --days 30\nterradev optimize\nTraining Pipeline\nbash\n# Pre-flight validation\nterradev preflight<p># Launch training\nterradev train --script train.py --from-provision latest\nterradev train-status\nterradev monitor --job my-job\nterradev checkpoint list --job my-job\nInference Optimization\nbash\n# vLLM auto-tuning (6 critical knobs)\nterradev vllm auto-optimize -s workload.json -m meta-llama&#x2F;Llama-2-7b-hf -g 4\nterradev vllm analyze -e http:&#x2F;&#x2F;localhost:8000\nterradev vllm benchmark -e http:&#x2F;&#x2F;localhost:8000 -c 10<p># MoE deployment with auto-optimizations\nterradev provision --task clusters&#x2F;moe-template&#x2F;task.yaml \\\n  --set model_id=Qwen&#x2F;Qwen3.5-397B-A17B<p># Disaggregated prefill&#x2F;decode\nterradev ml ray --deploy-pd --model zai-org&#x2F;GLM-5-FP8 \\\n  --prefill-tp 8 --decode-tp 1 --decode-dp 24<p># LoRA adapters (hot-load on running endpoint)\nterradev lora add -e http:&#x2F;&#x2F;endpoint:8000 -n customer-a -p &#x2F;adapters&#x2F;a\nterradev lora list -e http:&#x2F;&#x2F;endpoint:8000\nterradev lora remove -e http:&#x2F;&#x2F;endpoint:8000 -n customer-a\nKubernetes\nbash\n# Topology-optimized clusters\nterradev k8s create my-cluster --gpu H100 --count 8 --prefer-spot\nterradev k8s list\nterradev k8s info my-cluster\nterradev k8s destroy my-cluster\nSecondary Features\nbash\n# HF Spaces (one-click deployment)\nterradev hf-space my-llama --model-id meta-llama&#x2F;Llama-2-7b-hf --template llm<p># InferX serverless (&lt;2s cold starts)\nterradev inferx deploy --endpoint my-api --model-id meta-llama&#x2F;Llama-2-7b-hf\nterradev inferx status --endpoint my-api<p># Observability &amp; Safety\nterradev phoenix deploy --project my-inference\nterradev phoenix spans --project my-inference --limit 100\nterradev qdrant create-collection --name docs --vector-size 1536\nterradev guardrails generate-config --enable-topical --enable-pii<p># GitOps automation\nterradev gitops init --provider github --repo my-org&#x2F;infra --tool argocd\nterradev gitops sync --cluster production<p># Integrations (BYOAPI)\nterradev configure --provider wandb --api-key $WANDB_KEY\nterradev configure --provider prometheus --api-key $PROMETHEUS_URL\nQuick Workflows\nbash\n# 5-minute GPU setup\npip install terradev-cli &amp;&amp; terradev setup runpod --quick &amp;&amp; \\\nterradev quote -g H100 &amp;&amp; terradev run --gpu H100 --image pytorch&#x2F;pytorch:latest<p># Production RAG pipeline\nterradev qdrant k8s --namespace rag &amp;&amp; \\\nterradev qdrant create-collection --name kb --vector-size 1536 &amp;&amp; \\\nterradev provision --task clusters&#x2F;moe-template&#x2F;task.yaml \\\n  --set model_id=Qwen&#x2F;Qwen3.5-397B-A17B &amp;&amp; \\\nterradev phoenix deploy --project rag-pipeline<p># Multi-cloud cost optimization\nterradev analytics --days 30 &amp;&amp; terradev optimize<p>Key Features: Auto NUMA&#x2F;RDMA topology optimization, 19-cloud price comparison, vLLM auto-tuning, disaggregated P&#x2F;D, LoRA hot-loading, BYOAPI security model.","created_at":"2026-03-06T01:30:55Z","created_at_i":1772760655,"objectID":"47269669","parent_id":47269668,"story_id":47269668,"story_title":"Terradev: A next-gen slash command CLI for GPU provisioning and management","story_url":"https://github.com/theoddden/Terradev","updated_at":"2026-03-06T02:21:57Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"dikobraz"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vllm","fp8","benchmark"],"value":"LLM inference throughput <em>benchmark</em> for RTX PRO 6000 SE vs H100, H200, and B200 GPUs, based on the <em>vllm</em> serve and <em>vllm</em> bench serve benchmarking tools, to understand the cost-efficiency of various datacenter GPU options.<p>Benchmarking Setup:\nThe <em>benchmark</em> is optimized for throughput. <em>VLLM</em> serves models. The model is split across multiple GPUs using the --tensor-parallel-size <em>VLLM</em> option, if needed. Multiple <em>VLLM</em> instances serve the model; an NGINX load balancer on top distributes requests across them, maximizing throughput (replica parallelism). For example, if only 4 GPUs are required to run the model on an 8-GPU machine, two <em>VLLM</em> instances are launched with --tensor-parallel-size=4, and an NGINX load balancer is used. If all eight GPUs are required, then a single <em>VLLM</em> instance with --tensor-parallel-size=8 is used.<p>The <em>vllm</em> bench serve tool is used for benchmarking with random data and a sequence length of 1000. The number of concurrent requests is set to 64-256 to ensure the LLM's token-generation capacity is saturated.<p>Three models are benchmarked to better understand the effect of PCIe communication on the 8xPro6000 server vs. NVLink on the H100/H200/B200.<p>Here is the model selection and the logic behind it:\n- GLM-4.5-Air-AWQ-4bit (fits 80GB). Testing single-GPU performance and maximum throughput with replica scaling on 8 GPU setups. No PCIE bottleneck.\n- Qwen3-Coder-480B-A35B-Instruct-AWQ (fits 320GB). This 4-bit-quantized model fits into 4 GPUs. Some PCIe communication overhead in Pro 6000 setups may reduce performance relative to NVLink-enabled datacenter GPUs.\n- GLM-4.6-<em>FP8</em> (fits 640GB). This model requires all eight GPUs. PCIe communication overhead expected. The H100 and H200 configurations should have an advantage.<p>Besides raw throughput, graphs show the serving cost per million tokens for each model on its respective hardware. The rental price is set at $0.93 for Pro6000, $1.91 for H100, $2.06 for H200, and $2.68 for B200.<p>Results:\n- B200 wins on throughput, with the largest gap on the most communication-heavy workload \u2013 GLM-4.6-<em>FP8</em> (8-way TP): B200 is 4.87x faster than PRO 6000 (8,036.71 vs 1,651.67 tok/s) \u2013 Qwen3-Coder-480B (4-way TP): B200 is 4.02x faster than PRO 6000 (6,438.43 vs 1,602.96 tok/s) \u2013 GLM-4.5-Air (single-GPU replicas): B200 is 4.22x faster than PRO 6000 (9,675.24 vs 2,290.69 tok/s)\n- B200 is also the cost efficiency leader under updated run-cost estimates. B200\u2019s throughput advantage more than compensates for its higher hourly cost.\n- PRO 6000 is an attractive low-capex option. It beats H100 on cost per across all models and is on par with H200 on GLM-4.5-Air.\n- H200 is a major step up over H100. H200 delivers ~1.83x to 2.14x H100 throughput across the three models.\n- H100 looked worse than expected in this specific setup. It\u2019s on par with PRO 6000 in throughput on GLM-4.5-Air and behind all other contenders in cost per token across all workloads."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"[dead]"}},"_tags":["comment","author_dikobraz","story_46970757"],"author":"dikobraz","comment_text":"LLM inference throughput benchmark for RTX PRO 6000 SE vs H100, H200, and B200 GPUs, based on the vllm serve and vllm bench serve benchmarking tools, to understand the cost-efficiency of various datacenter GPU options.<p>Benchmarking Setup:\nThe benchmark is optimized for throughput. VLLM serves models. The model is split across multiple GPUs using the --tensor-parallel-size VLLM option, if needed. Multiple VLLM instances serve the model; an NGINX load balancer on top distributes requests across them, maximizing throughput (replica parallelism). For example, if only 4 GPUs are required to run the model on an 8-GPU machine, two VLLM instances are launched with --tensor-parallel-size=4, and an NGINX load balancer is used. If all eight GPUs are required, then a single VLLM instance with --tensor-parallel-size=8 is used.<p>The vllm bench serve tool is used for benchmarking with random data and a sequence length of 1000. The number of concurrent requests is set to 64-256 to ensure the LLM&#x27;s token-generation capacity is saturated.<p>Three models are benchmarked to better understand the effect of PCIe communication on the 8xPro6000 server vs. NVLink on the H100&#x2F;H200&#x2F;B200.<p>Here is the model selection and the logic behind it:\n- GLM-4.5-Air-AWQ-4bit (fits 80GB). Testing single-GPU performance and maximum throughput with replica scaling on 8 GPU setups. No PCIE bottleneck.\n- Qwen3-Coder-480B-A35B-Instruct-AWQ (fits 320GB). This 4-bit-quantized model fits into 4 GPUs. Some PCIe communication overhead in Pro 6000 setups may reduce performance relative to NVLink-enabled datacenter GPUs.\n- GLM-4.6-FP8 (fits 640GB). This model requires all eight GPUs. PCIe communication overhead expected. The H100 and H200 configurations should have an advantage.<p>Besides raw throughput, graphs show the serving cost per million tokens for each model on its respective hardware. The rental price is set at $0.93 for Pro6000, $1.91 for H100, $2.06 for H200, and $2.68 for B200.<p>Results:\n- B200 wins on throughput, with the largest gap on the most communication-heavy workload \u2013 GLM-4.6-FP8 (8-way TP): B200 is 4.87x faster than PRO 6000 (8,036.71 vs 1,651.67 tok&#x2F;s) \u2013 Qwen3-Coder-480B (4-way TP): B200 is 4.02x faster than PRO 6000 (6,438.43 vs 1,602.96 tok&#x2F;s) \u2013 GLM-4.5-Air (single-GPU replicas): B200 is 4.22x faster than PRO 6000 (9,675.24 vs 2,290.69 tok&#x2F;s)\n- B200 is also the cost efficiency leader under updated run-cost estimates. B200\u2019s throughput advantage more than compensates for its higher hourly cost.\n- PRO 6000 is an attractive low-capex option. It beats H100 on cost per across all models and is on par with H200 on GLM-4.5-Air.\n- H200 is a major step up over H100. H200 delivers ~1.83x to 2.14x H100 throughput across the three models.\n- H100 looked worse than expected in this specific setup. It\u2019s on par with PRO 6000 in throughput on GLM-4.5-Air and behind all other contenders in cost per token across all workloads.","created_at":"2026-02-11T04:16:15Z","created_at_i":1770783375,"objectID":"46970758","parent_id":46970757,"story_id":46970757,"story_title":"[dead]","updated_at":"2026-03-05T23:34:29Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"adefa"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vllm","fp8","benchmark"],"value":"<em>Benchmarks</em> using DGX Spark on <em>vLLM</em> 0.15.1.dev0+gf17644344<p><pre><code>  <em>FP8</em>: https://huggingface.co/Qwen/Qwen3-Coder-Next-<em>FP8</em>\n\n  Sequential (single request)\n\n    Prompt     Gen     Prompt Processing    Token Gen\n    Tokens     Tokens  (tokens/sec)         (tokens/sec)\n    ------     ------  -----------------    -----------\n       521        49            3,157            44.2\n     1,033        83            3,917            43.7\n     2,057        77            3,937            43.6\n     4,105        77            4,453            43.2\n     8,201        77            4,710            42.2\n\n  Parallel (concurrent requests)\n\n    pp4096+tg128 (4K context, 128 gen):\n\n     n    t/s\n    --    ----\n     1    28.5\n     2    39.0\n     4    50.4\n     8    57.5\n    16    61.4\n    32    62.0\n\n    pp8192+tg128 (8K context, 128 gen):\n\n     n    t/s\n    --    ----\n     1    21.6\n     2    27.1\n     4    31.9\n     8    32.7\n    16    33.7\n    32    31.7</code></pre>"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Qwen3-Coder-Next"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://qwen.ai/blog?id=qwen3-coder-next"}},"_tags":["comment","author_adefa","story_46872706"],"author":"adefa","children":[46879295],"comment_text":"Benchmarks using DGX Spark on vLLM 0.15.1.dev0+gf17644344<p><pre><code>  FP8: https:&#x2F;&#x2F;huggingface.co&#x2F;Qwen&#x2F;Qwen3-Coder-Next-FP8\n\n  Sequential (single request)\n\n    Prompt     Gen     Prompt Processing    Token Gen\n    Tokens     Tokens  (tokens&#x2F;sec)         (tokens&#x2F;sec)\n    ------     ------  -----------------    -----------\n       521        49            3,157            44.2\n     1,033        83            3,917            43.7\n     2,057        77            3,937            43.6\n     4,105        77            4,453            43.2\n     8,201        77            4,710            42.2\n\n  Parallel (concurrent requests)\n\n    pp4096+tg128 (4K context, 128 gen):\n\n     n    t&#x2F;s\n    --    ----\n     1    28.5\n     2    39.0\n     4    50.4\n     8    57.5\n    16    61.4\n    32    62.0\n\n    pp8192+tg128 (8K context, 128 gen):\n\n     n    t&#x2F;s\n    --    ----\n     1    21.6\n     2    27.1\n     4    31.9\n     8    32.7\n    16    33.7\n    32    31.7</code></pre>","created_at":"2026-02-03T22:18:34Z","created_at_i":1770157114,"objectID":"46878131","parent_id":46872706,"story_id":46872706,"story_title":"Qwen3-Coder-Next","story_url":"https://qwen.ai/blog?id=qwen3-coder-next","updated_at":"2026-03-05T23:30:46Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"btbuildem"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vllm","fp8","benchmark"],"value":"I'm not sure, I did not run any <em>benchmarks</em>. As a ballpark figure -- with both cards throttled down to 250W, running a Qwen-30B <em>FP8</em> model (variant depending on task), I get upwards of 60 tok/sec. It feels on par with the premium models, tbh.<p>Of course this is in a single-user environment, with <em>vLLM</em> keeping the model warm."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Tongyi DeepResearch \u2013 open-source 30B MoE Model that rivals OpenAI DeepResearch"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://tongyi-agent.github.io/blog/introducing-tongyi-deep-research/"}},"_tags":["comment","author_btbuildem","story_45789602"],"author":"btbuildem","comment_text":"I&#x27;m not sure, I did not run any benchmarks. As a ballpark figure -- with both cards throttled down to 250W, running a Qwen-30B FP8 model (variant depending on task), I get upwards of 60 tok&#x2F;sec. It feels on par with the premium models, tbh.<p>Of course this is in a single-user environment, with vLLM keeping the model warm.","created_at":"2025-11-02T19:47:48Z","created_at_i":1762112868,"objectID":"45792855","parent_id":45792614,"story_id":45789602,"story_title":"Tongyi DeepResearch \u2013 open-source 30B MoE Model that rivals OpenAI DeepResearch","story_url":"https://tongyi-agent.github.io/blog/introducing-tongyi-deep-research/","updated_at":"2026-03-05T23:00:37Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"NitpickLawyer"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vllm","fp8","benchmark"],"value":"Cool article! The stack (and results) are impressive, but I also appreciate the article in itself, starting from basics and getting to the point in a clear and slowly expanding way. Easy to follow and appreciate.<p>On a bit of a tangent rant, this kind of writing is slowly going away, taken over by LLM slop (and I'm a huge fan of LLMs, just not the people who write those kinds of articles). I was recently looking for real world <em>benchmarks</em> for <em>vllm</em>/sglang deployments of DeepSeek3 on a 8x 96GB pod, to see if the model fits into the amount of RAM, with kv cache and context length, what numbers to people get, etc.<p>Of the ~20 articles that google surfaced on various attempts of keywords, <i>none</i> were what I was looking for. The excerpts seemed promising, some even offered tables &amp; stuff related to ds3 and RAM usage, but <i>all</i> were LLM crap. All were written in that simple style - intro - bla bla - conclusion, some even had RAM requirements that made no sense (running a model trained in <em>FP8</em> in 16bit, something noone would do, etc.)"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"We clone a running VM in 2 seconds"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://codesandbox.io/blog/how-we-clone-a-running-vm-in-2-seconds"}},"_tags":["comment","author_NitpickLawyer","story_43622514"],"author":"NitpickLawyer","children":[43653624,43654145],"comment_text":"Cool article! The stack (and results) are impressive, but I also appreciate the article in itself, starting from basics and getting to the point in a clear and slowly expanding way. Easy to follow and appreciate.<p>On a bit of a tangent rant, this kind of writing is slowly going away, taken over by LLM slop (and I&#x27;m a huge fan of LLMs, just not the people who write those kinds of articles). I was recently looking for real world benchmarks for vllm&#x2F;sglang deployments of DeepSeek3 on a 8x 96GB pod, to see if the model fits into the amount of RAM, with kv cache and context length, what numbers to people get, etc.<p>Of the ~20 articles that google surfaced on various attempts of keywords, <i>none</i> were what I was looking for. The excerpts seemed promising, some even offered tables &amp; stuff related to ds3 and RAM usage, but <i>all</i> were LLM crap. All were written in that simple style - intro - bla bla - conclusion, some even had RAM requirements that made no sense (running a model trained in FP8 in 16bit, something noone would do, etc.)","created_at":"2025-04-11T13:20:26Z","created_at_i":1744377626,"objectID":"43653452","parent_id":43622514,"story_id":43622514,"story_title":"We clone a running VM in 2 seconds","story_url":"https://codesandbox.io/blog/how-we-clone-a-running-vm-in-2-seconds","updated_at":"2025-04-11T15:18:12Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"reissbaker"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vllm","fp8","benchmark"],"value":"By &quot;lower&quot; you mean cheaper/better?<p>I suspect it's much higher throughput than <em>vLLM</em>, which in turn is much higher throughput than llama.cpp. The MLA kernel they just open-sourced seems to indicate that, although we'll see how it does in third party <em>benchmarks</em> on non-hobbled GPUs vs FlashAttention. They only released the BF16 version \u2014 whereas most people, including DeepSeek themselves, serve in <em>FP8</em> \u2014 so it might not be immediately useful to most companies quite yet, although I imagine there'll be <em>FP8</em> ports soon enough."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"DeepSeek Open Source FlashMLA \u2013 MLA Decoding Kernel for Hopper GPUs"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/deepseek-ai/FlashMLA"}},"_tags":["comment","author_reissbaker","story_43155023"],"author":"reissbaker","children":[43156392],"comment_text":"By &quot;lower&quot; you mean cheaper&#x2F;better?<p>I suspect it&#x27;s much higher throughput than vLLM, which in turn is much higher throughput than llama.cpp. The MLA kernel they just open-sourced seems to indicate that, although we&#x27;ll see how it does in third party benchmarks on non-hobbled GPUs vs FlashAttention. They only released the BF16 version \u2014 whereas most people, including DeepSeek themselves, serve in FP8 \u2014 so it might not be immediately useful to most companies quite yet, although I imagine there&#x27;ll be FP8 ports soon enough.","created_at":"2025-02-24T03:57:49Z","created_at_i":1740369469,"objectID":"43155707","parent_id":43155406,"story_id":43155023,"story_title":"DeepSeek Open Source FlashMLA \u2013 MLA Decoding Kernel for Hopper GPUs","story_url":"https://github.com/deepseek-ai/FlashMLA","updated_at":"2025-02-24T15:04:14Z"},{"_highlightResult":{"author":{"fullyHighlighted":true,"matchLevel":"partial","matchedWords":["vllm"],"value":"<em>VLM</em>"},"comment_text":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["fp8","benchmark"],"value":"No its not directly comparable at all.<p>The softcore is pretty small, takes up roughly a tenth of a small cheap (by fpga standards) dev board.  There are bigger, and smaller, soft cores.<p>To compare, you could be running <em>benchmarks</em> on about ten simultaneous systems vs one FPGA if you insist on &quot;one chip&quot; vs &quot;one chip&quot; comparisons so obviously the fpga is ten times faster than it appears.  The advantage a FPGA provides is really smart custom peripherals.  So if for whatever reason you need to do lots of floating point divides in your app, or perhaps in your <em>benchmark</em>, you stick 100 hardware <em>FP</em> dividers on the chip and suddenly your division <em>benchmark</em> absolutely smokes the ARM which I believe has only one hardware <em>FP</em> division unit (or was it two?)"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: A 16-bit Forth machine written in VHDL"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/inforichland/yafc"}},"_tags":["comment","author_VLM","story_8129936"],"author":"VLM","comment_text":"No its not directly comparable at all.<p>The softcore is pretty small, takes up roughly a tenth of a small cheap (by fpga standards) dev board.  There are bigger, and smaller, soft cores.<p>To compare, you could be running benchmarks on about ten simultaneous systems vs one FPGA if you insist on &quot;one chip&quot; vs &quot;one chip&quot; comparisons so obviously the fpga is ten times faster than it appears.  The advantage a FPGA provides is really smart custom peripherals.  So if for whatever reason you need to do lots of floating point divides in your app, or perhaps in your benchmark, you stick 100 hardware FP dividers on the chip and suddenly your division benchmark absolutely smokes the ARM which I believe has only one hardware FP division unit (or was it two?)","created_at":"2014-08-04T14:03:29Z","created_at_i":1407161009,"objectID":"8131744","parent_id":8130070,"story_id":8129936,"story_title":"Show HN: A 16-bit Forth machine written in VHDL","story_url":"https://github.com/inforichland/yafc","updated_at":"2024-09-19T21:04:22Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"gaeld"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vllm","fp8","benchmark"],"value":"Great points.<p>We strived to be fair as possible in the <em>benchmark</em>, but it's indeed not perfect.\nTaalas should have been added in the dedicated hardware section, even though they use 3-bit quantization when we are on FP16 (to be fair in both directions) and they burn the model directly on the card.<p>Our tech preview is about the speed (hence the small dense model, it was easier to implement).<p>The math checks out though to allow support for large frontier MoE models at similar speeds:\n- At batch size 1, GPT-OSS-120B has 5.1B active parameters - in <em>FP8</em>, it's in the same size ballpark than our 2B model in FP16 (5.1 GB vs 4GB).\n- DeepSeek V4 Flash has 13B in mixed FP4/<em>FP8</em>, so let's say ballpark around 3x bigger than 4GB - so in theory we could reach &gt;1,000 tok/s on it with MI300X/H200 and up to 4k on next generation GPUs.<p>Check out the math at the end of our blog post:<p><a href=\"https://blog.kog.ai/real-time-llm-inference-on-standard-gpus-3-000-tokens-s-per-request/\" rel=\"nofollow\">https://blog.kog.ai/real-time-<em>llm</em>-inference-on-standard-gpus...</a>"},"story_title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["vllm"],"value":"Real-time <em>LLM</em> Inference on Standard GPUs: 3k tokens/s per request"},"story_url":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["vllm"],"value":"https://blog.kog.ai/real-time-<em>llm</em>-inference-on-standard-gpus-3-000-tokens-s-per-request/"}},"_tags":["comment","author_gaeld","story_48321076"],"author":"gaeld","children":[48322596],"comment_text":"Great points.<p>We strived to be fair as possible in the benchmark, but it&#x27;s indeed not perfect.\nTaalas should have been added in the dedicated hardware section, even though they use 3-bit quantization when we are on FP16 (to be fair in both directions) and they burn the model directly on the card.<p>Our tech preview is about the speed (hence the small dense model, it was easier to implement).<p>The math checks out though to allow support for large frontier MoE models at similar speeds:\n- At batch size 1, GPT-OSS-120B has 5.1B active parameters - in FP8, it&#x27;s in the same size ballpark than our 2B model in FP16 (5.1 GB vs 4GB).\n- DeepSeek V4 Flash has 13B in mixed FP4&#x2F;FP8, so let&#x27;s say ballpark around 3x bigger than 4GB - so in theory we could reach &gt;1,000 tok&#x2F;s on it with MI300X&#x2F;H200 and up to 4k on next generation GPUs.<p>Check out the math at the end of our blog post:<p><a href=\"https:&#x2F;&#x2F;blog.kog.ai&#x2F;real-time-llm-inference-on-standard-gpus-3-000-tokens-s-per-request&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;blog.kog.ai&#x2F;real-time-llm-inference-on-standard-gpus...</a>","created_at":"2026-05-29T10:59:35Z","created_at_i":1780052375,"objectID":"48321544","parent_id":48321384,"story_id":48321076,"story_title":"Real-time LLM Inference on Standard GPUs: 3k tokens/s per request","story_url":"https://blog.kog.ai/real-time-llm-inference-on-standard-gpus-3-000-tokens-s-per-request/","updated_at":"2026-06-05T12:24:22Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"KronisLV"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vllm","fp8","benchmark"],"value":"&gt; As of May 2026, how much money do I need to spend to buy hardware to have a local model that is 80% as good as SOTA services for assisting me in writing code?<p><a href=\"https://llm-stats.com/benchmarks/swe-bench-verified\">https://<em>llm</em>-stats.com/<em>benchmarks</em>/swe-bench-verified</a><p>SOTA (public proprietary models) would be Opus 4.7 at 0.876<p>80% of that would be around 0.7.<p>These models qualify, and are upwards of 90% as good in <em>benchmarks</em>:<p><pre><code>  DeepSeek-V4-Pro-Max - 1.6T (HuggingFace shows 862B, huh) - 0.806\n  Kimi K2.6 - 1.1T - 0.802\n  MiniMax M2.5 - 229B - 0.802\n  DeepSeek-V4-Flash-Max - 284B (HuggingFace shows 158B as well) - 0.790\n</code></pre>\nThese are 80-90% as good, which is also where you see the smaller ones:<p><pre><code>  GLM-5 - 754B - 0.778\n  Qwen3.6-27B - 27B - 0.772\n  Kimi K2.5 - 1.1T - 0.768\n  Qwen3.5-397B-A17B - 397B - 0.764\n  Step-3.5-Flash - 199B - 0.744\n  GLM-4.7 - 358B - 0.738\n  MiMo-V2-Flash - 310B - 0.734\n  Qwen3.6-35B-A3B - 35B - 0.734\n  DeepSeek-V3.2 - 685B - 0.731\n  DeepSeek-V3.2-Speciale - 685B - 0.731\n  DeepSeek-V3.2 (Thinking) - 685B - 0.731\n  Qwen3.5-27B - 27B - 0.724\n  Qwen3.5-122B-A10B - 125B - 0.720\n  Kimi K2-Thinking-0905 - 1T - 0.713\n  LongCat-Flash-Thinking-2601 - 562B - 0.700\n</code></pre>\nOut of those, the most modest one you could get is Qwen3.6-35B-A3B because the MoE nature makes it faster across more varied hardware.<p>I currently run the Unsloth 8bit quants on-prem (on a bunch of Nvidia L4 GPUs, since low TDP, long story), some people swear by more quantized versions but with the small models the impact is felt more: <a href=\"https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF\" rel=\"nofollow\">https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF</a><p>So essentially you need up to 39 GB for the model itself and then some for the KV cache and whatever context size you want. Ideally I'd aim for 64 GB of memory for that, though if really pressed for resources, could get a heavily quantized version within 32 GB (but very little memory for context and kinda shit).<p>Personally, I think that you need about 45-60 tokens/second for decent usability - even comparatively modest hardware (including those L4) can run the model, though on the lower end options you will not be running parallel sub-agents etc.<p>Some random results for when you don't want a traditional multi-GPU setup:<p><pre><code>  Mac Mini - about 1999 USD, gets you somewhere upwards of 30 tokens/second (depends on quantization and how you run it)\n  Framework Desktop - about 2500 USD, gets you somewhere upwards of 25 tokens/second https://community.frame.work/t/framework-desktop-for-local-ai/80880/5\n  DGX Spark - about 3500 USD, gets you somewhere upwards of 50 tokens/second https://forums.developer.nvidia.com/t/qwen-qwen3-6-35b-a3b-and-<em>fp8</em>-has-landed/366822/27\n</code></pre>\nSome random results from pulling up random shops and approx. <em>benchmarks</em>, for dual GPU setups (not necessarily NVLink etc.):<p><pre><code>  2x Intel Arc Pro B70 - about 1900 USD, gets you around 36 tokens/second, borderline usable, I blame their software stack\n  2x Radeon AI PRO R9700 - about 3000 USD, gets you somewhere upwards of 60 tokens/second, usable\n  2x Radeon PRO W7800 - about 5400 USD, same as above\n  2x NVIDIA RTX 5090 - about 7600 USD, same as above\n  2x NVIDIA RTX 5000 Ada - about 9200 USD, same as above\n</code></pre>\nOf course, for those models, some of those cards are way overkill, but you definitely can get <i>something</i> for running local models without too many compromises involved. That said, you definitely will get a worse experience than SOTA cloud models at that 80% and will have to rework stuff quite a bit often, as my own experience with the Qwen model shows - okay for simple tasks, breaks down on complex stuff. For that, you'd want at least some of the 90% category models and would probably need to consider how much memory you can realistically get.<p>At least it's not hopeless!"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"I am worried about Bun"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://wwj.dev/posts/i-am-worried-about-bun/"}},"_tags":["comment","author_KronisLV","story_48011184"],"author":"KronisLV","comment_text":"&gt; As of May 2026, how much money do I need to spend to buy hardware to have a local model that is 80% as good as SOTA services for assisting me in writing code?<p><a href=\"https:&#x2F;&#x2F;llm-stats.com&#x2F;benchmarks&#x2F;swe-bench-verified\">https:&#x2F;&#x2F;llm-stats.com&#x2F;benchmarks&#x2F;swe-bench-verified</a><p>SOTA (public proprietary models) would be Opus 4.7 at 0.876<p>80% of that would be around 0.7.<p>These models qualify, and are upwards of 90% as good in benchmarks:<p><pre><code>  DeepSeek-V4-Pro-Max - 1.6T (HuggingFace shows 862B, huh) - 0.806\n  Kimi K2.6 - 1.1T - 0.802\n  MiniMax M2.5 - 229B - 0.802\n  DeepSeek-V4-Flash-Max - 284B (HuggingFace shows 158B as well) - 0.790\n</code></pre>\nThese are 80-90% as good, which is also where you see the smaller ones:<p><pre><code>  GLM-5 - 754B - 0.778\n  Qwen3.6-27B - 27B - 0.772\n  Kimi K2.5 - 1.1T - 0.768\n  Qwen3.5-397B-A17B - 397B - 0.764\n  Step-3.5-Flash - 199B - 0.744\n  GLM-4.7 - 358B - 0.738\n  MiMo-V2-Flash - 310B - 0.734\n  Qwen3.6-35B-A3B - 35B - 0.734\n  DeepSeek-V3.2 - 685B - 0.731\n  DeepSeek-V3.2-Speciale - 685B - 0.731\n  DeepSeek-V3.2 (Thinking) - 685B - 0.731\n  Qwen3.5-27B - 27B - 0.724\n  Qwen3.5-122B-A10B - 125B - 0.720\n  Kimi K2-Thinking-0905 - 1T - 0.713\n  LongCat-Flash-Thinking-2601 - 562B - 0.700\n</code></pre>\nOut of those, the most modest one you could get is Qwen3.6-35B-A3B because the MoE nature makes it faster across more varied hardware.<p>I currently run the Unsloth 8bit quants on-prem (on a bunch of Nvidia L4 GPUs, since low TDP, long story), some people swear by more quantized versions but with the small models the impact is felt more: <a href=\"https:&#x2F;&#x2F;huggingface.co&#x2F;unsloth&#x2F;Qwen3.6-35B-A3B-GGUF\" rel=\"nofollow\">https:&#x2F;&#x2F;huggingface.co&#x2F;unsloth&#x2F;Qwen3.6-35B-A3B-GGUF</a><p>So essentially you need up to 39 GB for the model itself and then some for the KV cache and whatever context size you want. Ideally I&#x27;d aim for 64 GB of memory for that, though if really pressed for resources, could get a heavily quantized version within 32 GB (but very little memory for context and kinda shit).<p>Personally, I think that you need about 45-60 tokens&#x2F;second for decent usability - even comparatively modest hardware (including those L4) can run the model, though on the lower end options you will not be running parallel sub-agents etc.<p>Some random results for when you don&#x27;t want a traditional multi-GPU setup:<p><pre><code>  Mac Mini - about 1999 USD, gets you somewhere upwards of 30 tokens&#x2F;second (depends on quantization and how you run it)\n  Framework Desktop - about 2500 USD, gets you somewhere upwards of 25 tokens&#x2F;second https:&#x2F;&#x2F;community.frame.work&#x2F;t&#x2F;framework-desktop-for-local-ai&#x2F;80880&#x2F;5\n  DGX Spark - about 3500 USD, gets you somewhere upwards of 50 tokens&#x2F;second https:&#x2F;&#x2F;forums.developer.nvidia.com&#x2F;t&#x2F;qwen-qwen3-6-35b-a3b-and-fp8-has-landed&#x2F;366822&#x2F;27\n</code></pre>\nSome random results from pulling up random shops and approx. benchmarks, for dual GPU setups (not necessarily NVLink etc.):<p><pre><code>  2x Intel Arc Pro B70 - about 1900 USD, gets you around 36 tokens&#x2F;second, borderline usable, I blame their software stack\n  2x Radeon AI PRO R9700 - about 3000 USD, gets you somewhere upwards of 60 tokens&#x2F;second, usable\n  2x Radeon PRO W7800 - about 5400 USD, same as above\n  2x NVIDIA RTX 5090 - about 7600 USD, same as above\n  2x NVIDIA RTX 5000 Ada - about 9200 USD, same as above\n</code></pre>\nOf course, for those models, some of those cards are way overkill, but you definitely can get <i>something</i> for running local models without too many compromises involved. That said, you definitely will get a worse experience than SOTA cloud models at that 80% and will have to rework stuff quite a bit often, as my own experience with the Qwen model shows - okay for simple tasks, breaks down on complex stuff. For that, you&#x27;d want at least some of the 90% category models and would probably need to consider how much memory you can realistically get.<p>At least it&#x27;s not hopeless!","created_at":"2026-05-05T15:24:24Z","created_at_i":1777994664,"objectID":"48023822","parent_id":48018093,"story_id":48011184,"story_title":"I am worried about Bun","story_url":"https://wwj.dev/posts/i-am-worried-about-bun/","updated_at":"2026-05-06T11:57:26Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"byte123"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vllm","fp8","benchmark"],"value":"A lightning-fast text generation API running on Cloudflare's global network\nAccess to 60+ premium AI models including the latest Llama 4 Scout, DeepSeek R1 Distill, QwQ reasoning models, and more\n100,000 FREE daily requests (worth $1000+ on other platforms)\nFunction calling capabilities for advanced AI applications\nMultimodal support for text and image processing\nProduction-ready deployment in under 10 minutes<p>Featured AI Models (All FREE):<p>Meta Llama 4 Scout 17B - Latest multimodal model with 16 experts\nDeepSeek R1 Distill 32B - Outperforms OpenAI o1-mini across <em>benchmarks</em>\nQwQ 32B - Advanced reasoning model from Qwen series\nMistral Small 3.1 24B - State-of-the-art vision understanding\nLlama 3.3 70B <em>FP8</em> Fast - Optimized for speed and performance\nGemma 3 12B IT - Google's latest multimodal model\nQwen 2.5 Coder 32B - Specialized coding assistant\nWhisper Large V3 Turbo - Speech-to-text processing\nFLUX.1 Schnell - 12B parameter image generation\nand many more in Cloudflare models docs.<p>Why This Matters in 2025:\nWith AI costs skyrocketing and major providers limiting free tiers, having your own unlimited AI API is a game-changer. Whether you're building chatbots, content generators, coding assistants, or multimodal applications, this setup gives you enterprise-grade AI capabilities without the enterprise price tag.\nPerfect For:<p>&quot;free ai api&quot;\n&quot;openai alternative&quot;\n&quot;cloudflare workers ai&quot;\n&quot;free text generation&quot;\n&quot;llama 4&quot;\n&quot;deepseek r1&quot;\n&quot;free <em>llm</em>&quot;\n&quot;ai tutorial&quot;\n&quot;serverless ai&quot;\n&quot;how to build free ai api&quot;\n&quot;cloudflare workers ai tutorial&quot;\n&quot;free openai alternative 2025&quot;\n&quot;100k free ai requests&quot;\n&quot;llama 4 scout free access&quot;\n&quot;deepseek r1 free api&quot;\n&quot;llama 4 scout multimodal&quot;\n&quot;deepseek r1 distill performance&quot;\n&quot;qwq reasoning model&quot;\n&quot;mistral small 3.1 vision&quot;\n&quot;gemma 3 multimodal&quot;\n&quot;ai model comparison 2025&quot;"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"[dead]"}},"_tags":["comment","author_byte123","story_45039312"],"author":"byte123","comment_text":"A lightning-fast text generation API running on Cloudflare&#x27;s global network\nAccess to 60+ premium AI models including the latest Llama 4 Scout, DeepSeek R1 Distill, QwQ reasoning models, and more\n100,000 FREE daily requests (worth $1000+ on other platforms)\nFunction calling capabilities for advanced AI applications\nMultimodal support for text and image processing\nProduction-ready deployment in under 10 minutes<p>Featured AI Models (All FREE):<p>Meta Llama 4 Scout 17B - Latest multimodal model with 16 experts\nDeepSeek R1 Distill 32B - Outperforms OpenAI o1-mini across benchmarks\nQwQ 32B - Advanced reasoning model from Qwen series\nMistral Small 3.1 24B - State-of-the-art vision understanding\nLlama 3.3 70B FP8 Fast - Optimized for speed and performance\nGemma 3 12B IT - Google&#x27;s latest multimodal model\nQwen 2.5 Coder 32B - Specialized coding assistant\nWhisper Large V3 Turbo - Speech-to-text processing\nFLUX.1 Schnell - 12B parameter image generation\nand many more in Cloudflare models docs.<p>Why This Matters in 2025:\nWith AI costs skyrocketing and major providers limiting free tiers, having your own unlimited AI API is a game-changer. Whether you&#x27;re building chatbots, content generators, coding assistants, or multimodal applications, this setup gives you enterprise-grade AI capabilities without the enterprise price tag.\nPerfect For:<p>&quot;free ai api&quot;\n&quot;openai alternative&quot;\n&quot;cloudflare workers ai&quot;\n&quot;free text generation&quot;\n&quot;llama 4&quot;\n&quot;deepseek r1&quot;\n&quot;free llm&quot;\n&quot;ai tutorial&quot;\n&quot;serverless ai&quot;\n&quot;how to build free ai api&quot;\n&quot;cloudflare workers ai tutorial&quot;\n&quot;free openai alternative 2025&quot;\n&quot;100k free ai requests&quot;\n&quot;llama 4 scout free access&quot;\n&quot;deepseek r1 free api&quot;\n&quot;llama 4 scout multimodal&quot;\n&quot;deepseek r1 distill performance&quot;\n&quot;qwq reasoning model&quot;\n&quot;mistral small 3.1 vision&quot;\n&quot;gemma 3 multimodal&quot;\n&quot;ai model comparison 2025&quot;","created_at":"2025-08-27T13:20:11Z","created_at_i":1756300811,"objectID":"45039313","parent_id":45039312,"story_id":45039312,"story_title":"[dead]","updated_at":"2026-03-05T22:32:07Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"fxtentacle"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vllm","fp8","benchmark"],"value":"While I fully agree with you on the absence of good <em>benchmarks</em> and the growing <em>LLM</em> slop ...<p>&quot;running a model trained in <em>FP8</em> in 16bit, something noone would do, etc&quot;<p>I did that because on the RTX 3090 - which can be a good bang per buck for inference - the <em>FP8</em> support is nerfed at the driver level. So a kernel that upscales <em>FP8</em> to FP16 inside SRAM, then does the matmul, then downscales to <em>FP8</em> again can bring massive performance benefits on those consumer cards.<p>BTW, you can run a good DeepSeek3 quant on a single H200."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"We clone a running VM in 2 seconds (2022)"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://codesandbox.io/blog/how-we-clone-a-running-vm-in-2-seconds"}},"_tags":["comment","author_fxtentacle","story_43622514"],"author":"fxtentacle","children":[43653977],"comment_text":"While I fully agree with you on the absence of good benchmarks and the growing LLM slop ...<p>&quot;running a model trained in FP8 in 16bit, something noone would do, etc&quot;<p>I did that because on the RTX 3090 - which can be a good bang per buck for inference - the FP8 support is nerfed at the driver level. So a kernel that upscales FP8 to FP16 inside SRAM, then does the matmul, then downscales to FP8 again can bring massive performance benefits on those consumer cards.<p>BTW, you can run a good DeepSeek3 quant on a single H200.","created_at":"2025-04-11T13:36:10Z","created_at_i":1744378570,"objectID":"43653624","parent_id":43653452,"story_id":43622514,"story_title":"We clone a running VM in 2 seconds (2022)","story_url":"https://codesandbox.io/blog/how-we-clone-a-running-vm-in-2-seconds","updated_at":"2025-04-12T02:10:56Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"andersa"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vllm","fp8","benchmark"],"value":"If you really want to know an exact number for a specific use case, you can rent an 8xH100 node on RunPod and <em>benchmark</em> it.<p>You should expect somewhere around 30t/s for a single response, if running the <em>FP8</em> rowwise quant that would typically be used on such a node, with TensorRT-<em>LLM</em>. Massively more in total with batching.<p>That quant is twice the size as the 4.5bpw one used on the Mac though. A lower quality one would be faster."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Forget ChatGPT: why researchers now run small AIs on their laptops"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://www.nature.com/articles/d41586-024-02998-y"}},"_tags":["comment","author_andersa","story_41609393"],"author":"andersa","comment_text":"If you really want to know an exact number for a specific use case, you can rent an 8xH100 node on RunPod and benchmark it.<p>You should expect somewhere around 30t&#x2F;s for a single response, if running the FP8 rowwise quant that would typically be used on such a node, with TensorRT-LLM. Massively more in total with batching.<p>That quant is twice the size as the 4.5bpw one used on the Mac though. A lower quality one would be faster.","created_at":"2024-09-22T02:32:02Z","created_at_i":1726972322,"objectID":"41614221","parent_id":41611273,"story_id":41609393,"story_title":"Forget ChatGPT: why researchers now run small AIs on their laptops","story_url":"https://www.nature.com/articles/d41586-024-02998-y","updated_at":"2024-09-22T12:28:50Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"lhl"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vllm","fp8","benchmark"],"value":"Isn't SXM5 higher bandwidth? It's 900 GB/s of bidirectional bandwidth per GPU across 18 NVLink 4 channels. The NVL's are on PCIe 5, and even w/ NVLink only get to 600 GB/s of bandwidth across 3 NVLink bridges (across only pairs of cards)?<p>I haven't done a head to head and I suppose it depends on whether tensor parallelism actually scales linearly or not, but my understanding is since the NVL's are just PCIe/NVLink paired H100s, you're not really getting much if any benefit on something like <em>vLLM</em>.<p>I think the more interesting thing critique might be the slightly odd choice of Mixtral 8x7B vs say a more standard Llama2/3 70B (or just test multiple models including some big ones like 8x22B or DBRX.<p>Also, while I don't have a problem w/ <em>vLLM</em>, as TensorRT gets easier to set up, it might become a factor in comparisons (since they punted on <em>FP8</em>/AMP in this tests). Inferless published a shootoff a couple months ago comparing a few different inference engines: <a href=\"https://www.inferless.com/learn/exploring-llms-speed-benchmarks-independent-analysis\" rel=\"nofollow\">https://www.inferless.com/learn/exploring-llms-speed-<em>benchma</em>...</a><p>Price/perf does tell a story, but I think it's one that's mostly about Nvidia's platform dominance and profit margins more than intrinsic hardware advantages. On the spec sheet MI300X has a memory bandwidth and even raw FLOPS advantage but so far it has lacked proper software optimization/support and wide availability (has anyone besides hyperscalers and select partners been able to get them?)"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"AMD's MI300X Outperforms Nvidia's H100 for LLM Inference"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://www.blog.tensorwave.com/amds-mi300x-outperforms-nvidias-h100-for-llm-inference/"}},"_tags":["comment","author_lhl","story_40667102"],"author":"lhl","children":[40669135],"comment_text":"Isn&#x27;t SXM5 higher bandwidth? It&#x27;s 900 GB&#x2F;s of bidirectional bandwidth per GPU across 18 NVLink 4 channels. The NVL&#x27;s are on PCIe 5, and even w&#x2F; NVLink only get to 600 GB&#x2F;s of bandwidth across 3 NVLink bridges (across only pairs of cards)?<p>I haven&#x27;t done a head to head and I suppose it depends on whether tensor parallelism actually scales linearly or not, but my understanding is since the NVL&#x27;s are just PCIe&#x2F;NVLink paired H100s, you&#x27;re not really getting much if any benefit on something like vLLM.<p>I think the more interesting thing critique might be the slightly odd choice of Mixtral 8x7B vs say a more standard Llama2&#x2F;3 70B (or just test multiple models including some big ones like 8x22B or DBRX.<p>Also, while I don&#x27;t have a problem w&#x2F; vLLM, as TensorRT gets easier to set up, it might become a factor in comparisons (since they punted on FP8&#x2F;AMP in this tests). Inferless published a shootoff a couple months ago comparing a few different inference engines: <a href=\"https:&#x2F;&#x2F;www.inferless.com&#x2F;learn&#x2F;exploring-llms-speed-benchmarks-independent-analysis\" rel=\"nofollow\">https:&#x2F;&#x2F;www.inferless.com&#x2F;learn&#x2F;exploring-llms-speed-benchma...</a><p>Price&#x2F;perf does tell a story, but I think it&#x27;s one that&#x27;s mostly about Nvidia&#x27;s platform dominance and profit margins more than intrinsic hardware advantages. On the spec sheet MI300X has a memory bandwidth and even raw FLOPS advantage but so far it has lacked proper software optimization&#x2F;support and wide availability (has anyone besides hyperscalers and select partners been able to get them?)","created_at":"2024-06-13T12:49:37Z","created_at_i":1718282977,"objectID":"40669012","parent_id":40668588,"story_id":40667102,"story_title":"AMD's MI300X Outperforms Nvidia's H100 for LLM Inference","story_url":"https://www.blog.tensorwave.com/amds-mi300x-outperforms-nvidias-h100-for-llm-inference/","updated_at":"2024-09-20T17:17:15Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"mlgoatherder"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vllm","fp8","benchmark"],"value":"I've done some experiments here with Llama 13B, in my subjective experience the original fp16 model is significantly better (particularly on coding tasks). There are a bunch of synthetic <em>benchmarks</em> such a wikitext2 PPL and all the whiz bang quantization schemes seem to score well but subjectively something is missing.<p>I've been able to compare 4 bit GPTQ, naive int8, <em>LLM</em>.int8, fp16, and fp32. <em>LLM</em>.int8 does impressively well but inference is 4-5x slower than native fp16.<p>Oddly I recently ran a fork of the model on the ONNX runtime, I'm convinced that the model performed better than pytorch/transformers, perhaps subtle differences in floating point behavior etc between kernels on different hardware significantly influence performance.<p>The most promising next step in the quantization space IMO has to be <em>fp8</em>, there's a lot of hardware vendors adding support, and there's a lot of reasons to believe <em>fp8</em> will outperform most current quantization schemes [1][2]. Particularly when combined with quantization aware training / fine tuning (I think OpenAI did something similar for GPT3.5 &quot;turbo&quot;).<p>If anybody is interested I'm currently working on an open source <em>fp8</em> emulation library for pytorch, hoping to build something equivalent to bitsandbytes. If you are interested in collaborating my email is in my profile.<p>1. <a href=\"https://arxiv.org/abs/2208.09225\" rel=\"nofollow\">https://arxiv.org/abs/2208.09225</a>\n2. <a href=\"https://arxiv.org/abs/2209.05433\" rel=\"nofollow\">https://arxiv.org/abs/2209.05433</a>"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Llama.cpp 30B runs with only 6GB of RAM now"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/ggerganov/llama.cpp/pull/613"}},"_tags":["comment","author_mlgoatherder","story_35393284"],"author":"mlgoatherder","children":[35395400,35404063],"comment_text":"I&#x27;ve done some experiments here with Llama 13B, in my subjective experience the original fp16 model is significantly better (particularly on coding tasks). There are a bunch of synthetic benchmarks such a wikitext2 PPL and all the whiz bang quantization schemes seem to score well but subjectively something is missing.<p>I&#x27;ve been able to compare 4 bit GPTQ, naive int8, LLM.int8, fp16, and fp32. LLM.int8 does impressively well but inference is 4-5x slower than native fp16.<p>Oddly I recently ran a fork of the model on the ONNX runtime, I&#x27;m convinced that the model performed better than pytorch&#x2F;transformers, perhaps subtle differences in floating point behavior etc between kernels on different hardware significantly influence performance.<p>The most promising next step in the quantization space IMO has to be fp8, there&#x27;s a lot of hardware vendors adding support, and there&#x27;s a lot of reasons to believe fp8 will outperform most current quantization schemes [1][2]. Particularly when combined with quantization aware training &#x2F; fine tuning (I think OpenAI did something similar for GPT3.5 &quot;turbo&quot;).<p>If anybody is interested I&#x27;m currently working on an open source fp8 emulation library for pytorch, hoping to build something equivalent to bitsandbytes. If you are interested in collaborating my email is in my profile.<p>1. <a href=\"https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2208.09225\" rel=\"nofollow\">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2208.09225</a>\n2. <a href=\"https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2209.05433\" rel=\"nofollow\">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2209.05433</a>","created_at":"2023-03-31T21:37:30Z","created_at_i":1680298650,"objectID":"35394006","parent_id":35393652,"story_id":35393284,"story_title":"Llama.cpp 30B runs with only 6GB of RAM now","story_url":"https://github.com/ggerganov/llama.cpp/pull/613","updated_at":"2024-09-20T13:43:36Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"fswd"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vllm","fp8","benchmark"],"value":"For <em>LLM</em>, INT8 is old news but still exciting.  <em>FP8</em> would definitely be an improvement.  However the new coolness is INT4.<p>&gt; Excitingly, we manage to reach the INT4 weight quantization for GLM-130B while existing successes have thus far only come to the INT8 level. Memory-wise, by comparing to INT8, the INT4\nversion helps additionally save half of the required GPU memory to 70GB, thus allowing GLM130B inference on 4 \u00d7 RTX 3090 Ti (24G) or 8 \u00d7 RTX 2080 Ti (11G). Performance-wise, Table 2\nleft indicates that without post-training at all, the INT4-version GLM-130B experiences almost no\nperformance degradation, thus maintaining the advantages over GPT-3 on common <em>benchmarks</em>.<p>Page 7 <a href=\"https://arxiv.org/pdf/2210.02414.pdf\" rel=\"nofollow\">https://arxiv.org/pdf/2210.02414.pdf</a>"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Will Floating Point 8 Solve AI/ML Overhead?"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://semiengineering.com/will-floating-point-8-solve-ai-ml-overhead/"}},"_tags":["comment","author_fswd","story_34393903"],"author":"fswd","children":[34404904,34405231,34407378,34408845],"comment_text":"For LLM, INT8 is old news but still exciting.  FP8 would definitely be an improvement.  However the new coolness is INT4.<p>&gt; Excitingly, we manage to reach the INT4 weight quantization for GLM-130B while existing successes have thus far only come to the INT8 level. Memory-wise, by comparing to INT8, the INT4\nversion helps additionally save half of the required GPU memory to 70GB, thus allowing GLM130B inference on 4 \u00d7 RTX 3090 Ti (24G) or 8 \u00d7 RTX 2080 Ti (11G). Performance-wise, Table 2\nleft indicates that without post-training at all, the INT4-version GLM-130B experiences almost no\nperformance degradation, thus maintaining the advantages over GPT-3 on common benchmarks.<p>Page 7 <a href=\"https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2210.02414.pdf\" rel=\"nofollow\">https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2210.02414.pdf</a>","created_at":"2023-01-16T20:03:11Z","created_at_i":1673899391,"objectID":"34404859","parent_id":34393903,"story_id":34393903,"story_title":"Will Floating Point 8 Solve AI/ML Overhead?","story_url":"https://semiengineering.com/will-floating-point-8-solve-ai-ml-overhead/","updated_at":"2024-09-20T12:58:22Z"}],"hitsPerPage":20,"nbHits":28,"nbPages":2,"page":0,"params":"query=vllm+fp8+benchmark&advancedSyntax=true&analyticsTags=backend","processingTimeMS":17,"processingTimingsMS":{"_request":{"roundTrip":22},"afterFetch":{"format":{"highlighting":2,"total":2},"merge":{"mergeLoop":{"prepareNextHit":4,"total":4},"total":4},"total":4},"fetch":{"query":9,"scanning":2,"total":12},"total":17},"query":"vllm fp8 benchmark","serverTimeMS":20}
