{"author":"dskuldeep","children":[],"created_at":"2025-12-09T10:03:34.000Z","created_at_i":1765274614,"id":46203228,"options":[],"parent_id":null,"points":3,"story_id":46203228,"text":"We built Bifrost because we found existing Python-based gateways struggled with high concurrency in production. We wanted something that treated LLM infra like high-availability software.<p>We ran side-by-side benchmarks against LiteLLM on a single t3.medium instance (using a mock LLM with 1.5s fixed latency) to test pure gateway overhead.<p>The Results:<p>p99 Latency: 90.72s (LiteLLM) vs 1.68s (Bifrost)<p>Throughput: 44 req&#x2F;sec vs 424 req&#x2F;sec<p>Memory: ~3x lighter usage in Go.<p>It\u2019s a drop-in replacement (OpenAI compatible) designed for teams needing semantic caching, failover, and observability without the overhead.<p>We\u2019d love to hear your feedback.","title":"Show HN: Bifrost \u2013 open-source LLM Gateway (50x lower latency than LiteLLM)","type":"story","url":"https://github.com/maximhq/bifrost"}
