{"exhaustive":{"nbHits":false,"typo":false},"exhaustiveNbHits":false,"exhaustiveTypo":false,"hits":[{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"honorable_coder"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["katanemo"],"value":"Hey HN!<p>This is Adil, Salman and Jose and and we\u2019re behind archgw [1]. An intelligent proxy server designed as an edge and AI gateway for agents - one that natively know how to handle prompts, not just network traffic. We\u2019ve made several sweeping changes so sharing the project again.<p>A bit of background on why we\u2019ve built this project. Building AI agent demos is easy, but to create something production-ready there is a lot of repeat low-level plumbing work that everyone is doing. You\u2019re applying guardrails to make sure unsafe or off-topic requests don\u2019t get through. You\u2019re clarifying vague input so agents don\u2019t make mistakes. You\u2019re routing prompts to the right expert agent based on context or task type. You\u2019re writing integration code to quickly and safely add support for new LLMs. And every time a new framework hits the market or is updated, you\u2019re validating or re-implementing that same logic\u2014again and again.<p>Putting all the low-level plumbing code in a framework gets messy to manage, harder to update and scale. Low-level work isn't business logic. That\u2019s why we built archgw - an intelligent proxy server that handles prompts during ingress and egress and offers several related capabilities from a single software service. It lives outside your app runtime, so you can keep your business logic clean and focus on what matters. Think of it like a service mesh, but for AI agents.<p>Prior to building archgw, the team spent time building Envoy [2] at Lyft, API Gateway at AWS, specialized NLP models at Microsoft Research and worked on safety at Meta. archgw was born out of the belief that rule-based, single-purpose tools that handle the work around resiliency, processing and routing prompts should move into a dedicated infrastructure layer for agents, but built on the battle-tested foundational of Envoy Proxy.<p>The intelligence in archgw comes from our fast Task-specific LLMs [3] that can handle things like agent routing and hand off, guardrails and preference-based intelligent LLM calling. Here are some additional details about the open source project. archgw is written in rust, and the request path has three main parts:<p>* Listener subsystem which handles downstream (ingress) and upstream (egress) request processing.\n* Prompt handler subsystem. This is where archgw makes decisions on the safety of the incoming request via its prompt_guard hooks and identifies where to forward the conversation to via its prompt_target primitive.\n* Model serving subsystem is the interface that hosts all the lightweight LLMs engineered in archgw and offers a framework for things like hallucination detection of our these models<p>We loved building this open source project, and our belief is that this infra primitive would help developers build faster, safer and more personalized agents without all the manual prompt engineering and systems integration work needed to get there. We hope to invite other developers to use and improve Arch. Please give it a shot and leave feedback here, or at our discord channel [4]\nAlso here is a quick demo of the project in action [5]. You can check out our public docs here at [6]. Our models are also available here [7].<p>[1] <a href=\"https://github.com/katanemo/archgw\">https://github.com/<em>katanemo</em>/archgw</a>\n[2] <a href=\"https://www.envoyproxy.io/\" rel=\"nofollow\">https://www.envoyproxy.io/</a>\n[3] <a href=\"https://huggingface.co/collections/katanemo/arch-function-66\" rel=\"nofollow\">https://huggingface.co/collections/<em>katanemo</em>/arch-function-66</a>...\n[4] <a href=\"https://discord.com/channels/1292630766827737088/12926307682\" rel=\"nofollow\">https://discord.com/channels/1292630766827737088/12926307682</a>...\n[5] <a href=\"https://www.youtube.com/watch?v=I4Lbhr-NNXk\" rel=\"nofollow\">https://www.youtube.com/watch?v=I4Lbhr-NNXk</a>\n[6] <a href=\"https://docs.archgw.com/\" rel=\"nofollow\">https://docs.archgw.com/</a>\n[7] <a href=\"https://huggingface.co/katanemo\" rel=\"nofollow\">https://huggingface.co/<em>katanemo</em></a>"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: ArchGW \u2013 An intelligent edge and service proxy for agents"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["katanemo"],"value":"https://github.com/<em>katanemo</em>/archgw/"}},"_tags":["story","author_honorable_coder","story_44546265","show_hn"],"author":"honorable_coder","children":[44546662,44550846,44551188,44551466,44558678],"created_at":"2025-07-12T23:55:39Z","created_at_i":1752364539,"num_comments":15,"objectID":"44546265","points":118,"story_id":44546265,"story_text":"Hey HN!<p>This is Adil, Salman and Jose and and we\u2019re behind archgw [1]. An intelligent proxy server designed as an edge and AI gateway for agents - one that natively know how to handle prompts, not just network traffic. We\u2019ve made several sweeping changes so sharing the project again.<p>A bit of background on why we\u2019ve built this project. Building AI agent demos is easy, but to create something production-ready there is a lot of repeat low-level plumbing work that everyone is doing. You\u2019re applying guardrails to make sure unsafe or off-topic requests don\u2019t get through. You\u2019re clarifying vague input so agents don\u2019t make mistakes. You\u2019re routing prompts to the right expert agent based on context or task type. You\u2019re writing integration code to quickly and safely add support for new LLMs. And every time a new framework hits the market or is updated, you\u2019re validating or re-implementing that same logic\u2014again and again.<p>Putting all the low-level plumbing code in a framework gets messy to manage, harder to update and scale. Low-level work isn&#x27;t business logic. That\u2019s why we built archgw - an intelligent proxy server that handles prompts during ingress and egress and offers several related capabilities from a single software service. It lives outside your app runtime, so you can keep your business logic clean and focus on what matters. Think of it like a service mesh, but for AI agents.<p>Prior to building archgw, the team spent time building Envoy [2] at Lyft, API Gateway at AWS, specialized NLP models at Microsoft Research and worked on safety at Meta. archgw was born out of the belief that rule-based, single-purpose tools that handle the work around resiliency, processing and routing prompts should move into a dedicated infrastructure layer for agents, but built on the battle-tested foundational of Envoy Proxy.<p>The intelligence in archgw comes from our fast Task-specific LLMs [3] that can handle things like agent routing and hand off, guardrails and preference-based intelligent LLM calling. Here are some additional details about the open source project. archgw is written in rust, and the request path has three main parts:<p>* Listener subsystem which handles downstream (ingress) and upstream (egress) request processing.\n* Prompt handler subsystem. This is where archgw makes decisions on the safety of the incoming request via its prompt_guard hooks and identifies where to forward the conversation to via its prompt_target primitive.\n* Model serving subsystem is the interface that hosts all the lightweight LLMs engineered in archgw and offers a framework for things like hallucination detection of our these models<p>We loved building this open source project, and our belief is that this infra primitive would help developers build faster, safer and more personalized agents without all the manual prompt engineering and systems integration work needed to get there. We hope to invite other developers to use and improve Arch. Please give it a shot and leave feedback here, or at our discord channel [4]\nAlso here is a quick demo of the project in action [5]. You can check out our public docs here at [6]. Our models are also available here [7].<p>[1] <a href=\"https:&#x2F;&#x2F;github.com&#x2F;katanemo&#x2F;archgw\">https:&#x2F;&#x2F;github.com&#x2F;katanemo&#x2F;archgw</a>\n[2] <a href=\"https:&#x2F;&#x2F;www.envoyproxy.io&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;www.envoyproxy.io&#x2F;</a>\n[3] <a href=\"https:&#x2F;&#x2F;huggingface.co&#x2F;collections&#x2F;katanemo&#x2F;arch-function-66\" rel=\"nofollow\">https:&#x2F;&#x2F;huggingface.co&#x2F;collections&#x2F;katanemo&#x2F;arch-function-66</a>...\n[4] <a href=\"https:&#x2F;&#x2F;discord.com&#x2F;channels&#x2F;1292630766827737088&#x2F;12926307682\" rel=\"nofollow\">https:&#x2F;&#x2F;discord.com&#x2F;channels&#x2F;1292630766827737088&#x2F;12926307682</a>...\n[5] <a href=\"https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=I4Lbhr-NNXk\" rel=\"nofollow\">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=I4Lbhr-NNXk</a>\n[6] <a href=\"https:&#x2F;&#x2F;docs.archgw.com&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;docs.archgw.com&#x2F;</a>\n[7] <a href=\"https:&#x2F;&#x2F;huggingface.co&#x2F;katanemo\" rel=\"nofollow\">https:&#x2F;&#x2F;huggingface.co&#x2F;katanemo</a>","title":"Show HN: ArchGW \u2013 An intelligent edge and service proxy for agents","updated_at":"2025-12-03T21:23:18Z","url":"https://github.com/katanemo/archgw/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"adilhafeez"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["katanemo"],"value":"Hi HN \u2014 we're the team behind Arch (<a href=\"https://github.com/katanemo/archgw\">https://github.com/<em>katanemo</em>/archgw</a>), an open-source proxy for LLMs written in Rust. Today we're releasing Arch-Router (<a href=\"https://huggingface.co/katanemo/Arch-Router-1.5B\" rel=\"nofollow\">https://huggingface.co/<em>katanemo</em>/Arch-Router-1.5B</a>), a 1.5B router model for preference-based routing, now integrated into the proxy. As teams integrate multiple LLMs - each with different strengths, styles, or cost/latency profiles \u2014 routing the right prompt to the right model becomes a critical part of the application design. But it's still an open problem. Most routing systems fall into two camps:<p>- Embedding-based routers use intent classifiers \u2014 label a prompt as \u201csupport,\u201d \u201cSQL,\u201d or \u201cmath,\u201d then route to a matching model. This works for simple tasks but breaks down in real conversations. Users shift topics mid-conversation, task boundaries blur, and product changes require retraining classifiers.<p>- Performance-based routers pick models based on benchmarks like MMLU or MT-Bench, or based on latency or cost curves. But benchmarks often miss what matters in production: domain-specific quality or subjective preferences like \u201cWill legal accept this clause?\u201d<p>Arch-Router takes a different approach: route by preferences written in plain language. You write rules like \u201ccontract clauses \u2192 GPT-4o\u201d or \u201cquick travel tips \u2192 Gemini Flash.\u201d The router maps the prompt (and conversation context) to those rules using a lightweight 1.5B autoregressive model. No retraining, no fragile if/else chains. We built this with input from teams at Twilio and Atlassian. It handles intent drift, supports multi-turn conversations, and lets you swap in or out models with a one-line change to the routing policy. Full details are in our paper (<a href=\"https://arxiv.org/abs/2506.16655\" rel=\"nofollow\">https://arxiv.org/abs/2506.16655</a>), but here's a snapshot:<p>Specs:<p>- 1.5B params \u2014 runs on a single GPU (or CPU for testing)<p>- No retraining needed \u2014 point it at any mix of LLMs<p>- Cost and latency aware \u2014 route heavy tasks to expensive models, light tasks to faster/cheaper ones<p>- Outperforms larger closed models on our conversational routing benchmarks (details in the paper)<p>Links:<p>- Arch Proxy (open source): <a href=\"https://github.com/katanemo/archgw\">https://github.com/<em>katanemo</em>/archgw</a><p>- Model + code: <a href=\"https://huggingface.co/katanemo/Arch-Router-1.5B\" rel=\"nofollow\">https://huggingface.co/<em>katanemo</em>/Arch-Router-1.5B</a><p>- Paper: <a href=\"https://arxiv.org/abs/2506.16655\" rel=\"nofollow\">https://arxiv.org/abs/2506.16655</a>"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Arch-Router \u2013 1.5B model for LLM routing by preferences, not benchmarks"}},"_tags":["story","author_adilhafeez","story_44436031","show_hn"],"author":"adilhafeez","children":[44436070,44437340,44437954,44438177,44438786,44439077,44445753],"created_at":"2025-07-01T17:13:11Z","created_at_i":1751389991,"num_comments":15,"objectID":"44436031","points":66,"story_id":44436031,"story_text":"Hi HN \u2014 we&#x27;re the team behind Arch (<a href=\"https:&#x2F;&#x2F;github.com&#x2F;katanemo&#x2F;archgw\">https:&#x2F;&#x2F;github.com&#x2F;katanemo&#x2F;archgw</a>), an open-source proxy for LLMs written in Rust. Today we&#x27;re releasing Arch-Router (<a href=\"https:&#x2F;&#x2F;huggingface.co&#x2F;katanemo&#x2F;Arch-Router-1.5B\" rel=\"nofollow\">https:&#x2F;&#x2F;huggingface.co&#x2F;katanemo&#x2F;Arch-Router-1.5B</a>), a 1.5B router model for preference-based routing, now integrated into the proxy. As teams integrate multiple LLMs - each with different strengths, styles, or cost&#x2F;latency profiles \u2014 routing the right prompt to the right model becomes a critical part of the application design. But it&#x27;s still an open problem. Most routing systems fall into two camps:<p>- Embedding-based routers use intent classifiers \u2014 label a prompt as \u201csupport,\u201d \u201cSQL,\u201d or \u201cmath,\u201d then route to a matching model. This works for simple tasks but breaks down in real conversations. Users shift topics mid-conversation, task boundaries blur, and product changes require retraining classifiers.<p>- Performance-based routers pick models based on benchmarks like MMLU or MT-Bench, or based on latency or cost curves. But benchmarks often miss what matters in production: domain-specific quality or subjective preferences like \u201cWill legal accept this clause?\u201d<p>Arch-Router takes a different approach: route by preferences written in plain language. You write rules like \u201ccontract clauses \u2192 GPT-4o\u201d or \u201cquick travel tips \u2192 Gemini Flash.\u201d The router maps the prompt (and conversation context) to those rules using a lightweight 1.5B autoregressive model. No retraining, no fragile if&#x2F;else chains. We built this with input from teams at Twilio and Atlassian. It handles intent drift, supports multi-turn conversations, and lets you swap in or out models with a one-line change to the routing policy. Full details are in our paper (<a href=\"https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2506.16655\" rel=\"nofollow\">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2506.16655</a>), but here&#x27;s a snapshot:<p>Specs:<p>- 1.5B params \u2014 runs on a single GPU (or CPU for testing)<p>- No retraining needed \u2014 point it at any mix of LLMs<p>- Cost and latency aware \u2014 route heavy tasks to expensive models, light tasks to faster&#x2F;cheaper ones<p>- Outperforms larger closed models on our conversational routing benchmarks (details in the paper)<p>Links:<p>- Arch Proxy (open source): <a href=\"https:&#x2F;&#x2F;github.com&#x2F;katanemo&#x2F;archgw\">https:&#x2F;&#x2F;github.com&#x2F;katanemo&#x2F;archgw</a><p>- Model + code: <a href=\"https:&#x2F;&#x2F;huggingface.co&#x2F;katanemo&#x2F;Arch-Router-1.5B\" rel=\"nofollow\">https:&#x2F;&#x2F;huggingface.co&#x2F;katanemo&#x2F;Arch-Router-1.5B</a><p>- Paper: <a href=\"https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2506.16655\" rel=\"nofollow\">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2506.16655</a>","title":"Show HN: Arch-Router \u2013 1.5B model for LLM routing by preferences, not benchmarks","updated_at":"2025-12-29T14:06:44Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"sparacha"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["katanemo"],"value":"Hi HN! This is Salman, Adil, Shuguang and Co working on ArchGW[1] - an open-source lightweight proxy server for prompts - written in Rust and built on top of Envoy[2]. Arch moves the critical but pesky handling and processing of prompts: task understanding, prompt routing, safety, and observability - outside business logic. Its an edge and egress proxy for agentic apps.<p>We've talked to 100s of developers at places like Twilio, GE Healthcare, Redhat, Square, etc and there was a consistent theme in building AI apps: to move past a nascent demo they are left to their own devices in building out middle ware capabilities so that developers can move faster and ship with confidence.<p>Today, the approach to building an enterprise-ready AI app is cobbling together a large set of mono-functional tools, adding LLM-based preprocessing steps to determine safety (e.g. applying governance and guardrails), ask clarifying questions to improve task performance, support common agentic operations by packaging and managing function calling scenarios manually, etc. Not to mention, all the undifferentiated work in incorporating different LLM models and versions, and managing resiliency, retries and fallback logic.<p>ArchGW was built with the belief that prompts are nuanced and opaque user requests, which require the same capabilities as traditional HTTP requests including secure handling, intelligent routing, robust observability, and integration with backend (API) systems for personalization \u2013 outside business logic. We help built Envoy while at Lyft and think its offers a great foundation to build a proxy to manage traffic for prompts.<p>Here are some additional details about the open source project. ArchGW is written in rust, and the request path has three main parts:<p>* Listener subsystem which handles downstream (ingress) and upstream (egress) request processing.<p>* Prompt handler subsystem. This is where ArchGW makes decisions on the safety of the incoming request via its prompt_guard primitive and identifies where to forward the conversation to via its prompt_target primitive.<p>* Model serving subsystem is the interface that hosts all the lightweight LLMs[3] engineered in ArchGW and offers a framework for things like hallucination detection of our these models<p>We loved building this open source project, and our belief is that this infrastructure primitive would help developers build faster, safer and more personalized agents without all the manual prompt engineering and systems integration work needed to get there. We hope to invite other developers to use and improve Arch. Please give it a shot and leave feedback here, or at our discord channel [4]<p>Also here is a quick demo of the project in action [5]. You can check out our public docs here at [6]. Our models are also available here [7].<p>[1] <a href=\"https://github.com/katanemo/archgw\">https://github.com/<em>katanemo</em>/archgw</a><p>[2] <a href=\"https://www.envoyproxy.io/\" rel=\"nofollow\">https://www.envoyproxy.io/</a><p>[3] <a href=\"https://huggingface.co/collections/katanemo/arch-function-66\" rel=\"nofollow\">https://huggingface.co/collections/<em>katanemo</em>/arch-function-66</a>...<p>[4] <a href=\"https://discord.com/channels/1292630766827737088/12926307682\" rel=\"nofollow\">https://discord.com/channels/1292630766827737088/12926307682</a>...<p>[5] <a href=\"https://www.youtube.com/watch?v=I4Lbhr-NNXk\" rel=\"nofollow\">https://www.youtube.com/watch?v=I4Lbhr-NNXk</a><p>[6] <a href=\"https://docs.archgw.com/\" rel=\"nofollow\">https://docs.archgw.com/</a><p>[7] <a href=\"https://huggingface.co/katanemo\" rel=\"nofollow\">https://huggingface.co/<em>katanemo</em></a>"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: ArchGW \u2013 An open-source intelligent proxy server for prompts"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["katanemo"],"value":"https://github.com/<em>katanemo</em>/archgw"}},"_tags":["story","author_sparacha","story_43259862","show_hn"],"author":"sparacha","children":[43260545,43260629,43261567,43270877,43274671,43274975],"created_at":"2025-03-04T21:14:56Z","created_at_i":1741122896,"num_comments":7,"objectID":"43259862","points":39,"story_id":43259862,"story_text":"Hi HN! This is Salman, Adil, Shuguang and Co working on ArchGW[1] - an open-source lightweight proxy server for prompts - written in Rust and built on top of Envoy[2]. Arch moves the critical but pesky handling and processing of prompts: task understanding, prompt routing, safety, and observability - outside business logic. Its an edge and egress proxy for agentic apps.<p>We&#x27;ve talked to 100s of developers at places like Twilio, GE Healthcare, Redhat, Square, etc and there was a consistent theme in building AI apps: to move past a nascent demo they are left to their own devices in building out middle ware capabilities so that developers can move faster and ship with confidence.<p>Today, the approach to building an enterprise-ready AI app is cobbling together a large set of mono-functional tools, adding LLM-based preprocessing steps to determine safety (e.g. applying governance and guardrails), ask clarifying questions to improve task performance, support common agentic operations by packaging and managing function calling scenarios manually, etc. Not to mention, all the undifferentiated work in incorporating different LLM models and versions, and managing resiliency, retries and fallback logic.<p>ArchGW was built with the belief that prompts are nuanced and opaque user requests, which require the same capabilities as traditional HTTP requests including secure handling, intelligent routing, robust observability, and integration with backend (API) systems for personalization \u2013 outside business logic. We help built Envoy while at Lyft and think its offers a great foundation to build a proxy to manage traffic for prompts.<p>Here are some additional details about the open source project. ArchGW is written in rust, and the request path has three main parts:<p>* Listener subsystem which handles downstream (ingress) and upstream (egress) request processing.<p>* Prompt handler subsystem. This is where ArchGW makes decisions on the safety of the incoming request via its prompt_guard primitive and identifies where to forward the conversation to via its prompt_target primitive.<p>* Model serving subsystem is the interface that hosts all the lightweight LLMs[3] engineered in ArchGW and offers a framework for things like hallucination detection of our these models<p>We loved building this open source project, and our belief is that this infrastructure primitive would help developers build faster, safer and more personalized agents without all the manual prompt engineering and systems integration work needed to get there. We hope to invite other developers to use and improve Arch. Please give it a shot and leave feedback here, or at our discord channel [4]<p>Also here is a quick demo of the project in action [5]. You can check out our public docs here at [6]. Our models are also available here [7].<p>[1] <a href=\"https:&#x2F;&#x2F;github.com&#x2F;katanemo&#x2F;archgw\">https:&#x2F;&#x2F;github.com&#x2F;katanemo&#x2F;archgw</a><p>[2] <a href=\"https:&#x2F;&#x2F;www.envoyproxy.io&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;www.envoyproxy.io&#x2F;</a><p>[3] <a href=\"https:&#x2F;&#x2F;huggingface.co&#x2F;collections&#x2F;katanemo&#x2F;arch-function-66\" rel=\"nofollow\">https:&#x2F;&#x2F;huggingface.co&#x2F;collections&#x2F;katanemo&#x2F;arch-function-66</a>...<p>[4] <a href=\"https:&#x2F;&#x2F;discord.com&#x2F;channels&#x2F;1292630766827737088&#x2F;12926307682\" rel=\"nofollow\">https:&#x2F;&#x2F;discord.com&#x2F;channels&#x2F;1292630766827737088&#x2F;12926307682</a>...<p>[5] <a href=\"https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=I4Lbhr-NNXk\" rel=\"nofollow\">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=I4Lbhr-NNXk</a><p>[6] <a href=\"https:&#x2F;&#x2F;docs.archgw.com&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;docs.archgw.com&#x2F;</a><p>[7] <a href=\"https:&#x2F;&#x2F;huggingface.co&#x2F;katanemo\" rel=\"nofollow\">https:&#x2F;&#x2F;huggingface.co&#x2F;katanemo</a>","title":"Show HN: ArchGW \u2013 An open-source intelligent proxy server for prompts","updated_at":"2025-07-12T23:49:16Z","url":"https://github.com/katanemo/archgw"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"adilhafeez"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["katanemo"],"value":"Hi HN! My name is Adil Hafeez, and I am the Co-Founder at <em>Katanemo</em> and the lead developer behind Arch - an open source project for developers to build faster, generative AI apps. Previously I worked on Envoy at Lyft.<p>Engineered with purpose-built LLMs, Arch handles the critical but undifferentiated tasks related to the handling and processing of prompts, including detecting and rejecting jailbreak attempts, intelligently calling \u201cbackend\u201d APIs to fulfill the user\u2019s request represented in a prompt, routing to and offering disaster recovery between upstream LLMs, and managing the observability of prompts and LLM interactions in a centralized way - all outside business logic.<p>Here are some additional key details of the project,<p>* Built on top of Envoy and is written in rust. It runs alongside application servers, and uses Envoy's proven HTTP management and scalability features to handle traffic related to prompts and LLMs.<p>* Function calling for fast agentic and RAG apps. Engineered with purpose-built fast LLMs to handle fast, cost-effective, and accurate prompt-based tasks like function/API calling, and parameter extraction from prompts.<p>* Prompt guardrails to prevent jailbreak attempts and ensure safe user interactions without writing a single line of code.<p>* Manages LLM calls, offering smart retries, automatic cutover, and resilient upstream connections for continuous availability.<p>* Uses the W3C Trace Context standard to enable complete request tracing across applications, ensuring compatibility with observability tools, and provides metrics to monitor latency, token usage, and error rates, helping optimize AI application performance.<p>This is our first release, and would love to build alongside the community. We are just getting started on reinventing what we could do at the networking layer for prompts.<p>Do check it out on GitHub at <a href=\"https://github.com/katanemo/arch/\">https://github.com/<em>katanemo</em>/arch/</a>.<p>Please leave a comment or feedback here and I will be happy to answer!"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Arch \u2013 an intelligent prompt gateway built on Envoy"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["katanemo"],"value":"https://github.com/<em>katanemo</em>/arch"}},"_tags":["story","author_adilhafeez","story_41801315","show_hn"],"author":"adilhafeez","children":[41864015],"created_at":"2024-10-10T17:45:48Z","created_at_i":1728582348,"num_comments":1,"objectID":"41801315","points":31,"story_id":41801315,"story_text":"Hi HN! My name is Adil Hafeez, and I am the Co-Founder at Katanemo and the lead developer behind Arch - an open source project for developers to build faster, generative AI apps. Previously I worked on Envoy at Lyft.<p>Engineered with purpose-built LLMs, Arch handles the critical but undifferentiated tasks related to the handling and processing of prompts, including detecting and rejecting jailbreak attempts, intelligently calling \u201cbackend\u201d APIs to fulfill the user\u2019s request represented in a prompt, routing to and offering disaster recovery between upstream LLMs, and managing the observability of prompts and LLM interactions in a centralized way - all outside business logic.<p>Here are some additional key details of the project,<p>* Built on top of Envoy and is written in rust. It runs alongside application servers, and uses Envoy&#x27;s proven HTTP management and scalability features to handle traffic related to prompts and LLMs.<p>* Function calling for fast agentic and RAG apps. Engineered with purpose-built fast LLMs to handle fast, cost-effective, and accurate prompt-based tasks like function&#x2F;API calling, and parameter extraction from prompts.<p>* Prompt guardrails to prevent jailbreak attempts and ensure safe user interactions without writing a single line of code.<p>* Manages LLM calls, offering smart retries, automatic cutover, and resilient upstream connections for continuous availability.<p>* Uses the W3C Trace Context standard to enable complete request tracing across applications, ensuring compatibility with observability tools, and provides metrics to monitor latency, token usage, and error rates, helping optimize AI application performance.<p>This is our first release, and would love to build alongside the community. We are just getting started on reinventing what we could do at the networking layer for prompts.<p>Do check it out on GitHub at <a href=\"https:&#x2F;&#x2F;github.com&#x2F;katanemo&#x2F;arch&#x2F;\">https:&#x2F;&#x2F;github.com&#x2F;katanemo&#x2F;arch&#x2F;</a>.<p>Please leave a comment or feedback here and I will be happy to answer!","title":"Show HN: Arch \u2013 an intelligent prompt gateway built on Envoy","updated_at":"2024-10-16T21:28:40Z","url":"https://github.com/katanemo/arch"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"adilhafeez"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["katanemo"],"value":"Hi HN! My name is Adil Hafeez, and I am the Co-Founder at <em>Katanemo</em> and the lead developer behind Arch - an open source project for developers to build faster, generative AI apps. Previously I worked on Envoy at Lyft.<p>Engineered with purpose-built LLMs, Arch handles the critical but undifferentiated tasks related to the handling and processing of prompts, including detecting and rejecting jailbreak attempts, intelligently calling \u201cbackend\u201d APIs to fulfill the user\u2019s request represented in a prompt, routing to and offering disaster recovery between upstream LLMs, and managing the observability of prompts and LLM interactions in a centralized way - all outside business logic.<p>Here are some additional key details of the project,<p>* Built on top of Envoy and is written in rust. It runs alongside application servers, and uses Envoy's proven HTTP management and scalability features to handle traffic related to prompts and LLMs.<p>* Function calling for fast agentic and RAG apps. Engineered with purpose-built fast LLMs to handle fast, cost-effective, and accurate prompt-based tasks like function/API calling, and parameter extraction from prompts.<p>* Prompt guardrails to prevent jailbreak attempts and ensure safe user interactions without writing a single line of code.<p>* Manages LLM calls, offering smart retries, automatic cutover, and resilient upstream connections for continuous availability.<p>* Uses the W3C Trace Context standard to enable complete request tracing across applications, ensuring compatibility with observability tools, and provides metrics to monitor latency, token usage, and error rates, helping optimize AI application performance.<p>This is our first release, and would love to build alongside the community. We are just getting started on reinventing what we could do at the networking layer for prompts.<p>Do check it out on GitHub at <a href=\"https://github.com/katanemo/arch/\">https://github.com/<em>katanemo</em>/arch/</a>.<p>Please leave a comment or feedback here and I will be happy to answer!"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Arch \u2013 an intelligent prompt gateway built on Envoy"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["katanemo"],"value":"https://github.com/<em>katanemo</em>/arch"}},"_tags":["story","author_adilhafeez","story_41864014","show_hn"],"author":"adilhafeez","children":[41801341,41851210,41853593,41853887,41853978,41854960,41864018,41864396],"created_at":"2024-10-16T21:27:27Z","created_at_i":1729114047,"num_comments":19,"objectID":"41864014","points":21,"story_id":41864014,"story_text":"Hi HN! My name is Adil Hafeez, and I am the Co-Founder at Katanemo and the lead developer behind Arch - an open source project for developers to build faster, generative AI apps. Previously I worked on Envoy at Lyft.<p>Engineered with purpose-built LLMs, Arch handles the critical but undifferentiated tasks related to the handling and processing of prompts, including detecting and rejecting jailbreak attempts, intelligently calling \u201cbackend\u201d APIs to fulfill the user\u2019s request represented in a prompt, routing to and offering disaster recovery between upstream LLMs, and managing the observability of prompts and LLM interactions in a centralized way - all outside business logic.<p>Here are some additional key details of the project,<p>* Built on top of Envoy and is written in rust. It runs alongside application servers, and uses Envoy&#x27;s proven HTTP management and scalability features to handle traffic related to prompts and LLMs.<p>* Function calling for fast agentic and RAG apps. Engineered with purpose-built fast LLMs to handle fast, cost-effective, and accurate prompt-based tasks like function&#x2F;API calling, and parameter extraction from prompts.<p>* Prompt guardrails to prevent jailbreak attempts and ensure safe user interactions without writing a single line of code.<p>* Manages LLM calls, offering smart retries, automatic cutover, and resilient upstream connections for continuous availability.<p>* Uses the W3C Trace Context standard to enable complete request tracing across applications, ensuring compatibility with observability tools, and provides metrics to monitor latency, token usage, and error rates, helping optimize AI application performance.<p>This is our first release, and would love to build alongside the community. We are just getting started on reinventing what we could do at the networking layer for prompts.<p>Do check it out on GitHub at <a href=\"https:&#x2F;&#x2F;github.com&#x2F;katanemo&#x2F;arch&#x2F;\">https:&#x2F;&#x2F;github.com&#x2F;katanemo&#x2F;arch&#x2F;</a>.<p>Please leave a comment or feedback here and I will be happy to answer!","title":"Show HN: Arch \u2013 an intelligent prompt gateway built on Envoy","updated_at":"2025-08-28T06:47:26Z","url":"https://github.com/katanemo/arch"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"adilhafeez"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["katanemo"],"value":"Hi HN! This is Adil, Salman, Co and Shuguang and we're excited to introduce archgw [1], an open source intelligent proxy for agents built on Envoy [2]. Arch moves the critical but crufty work around safety, observability, and routing of prompts outside business logic. Arch is a uniquely intelligent infrastructure primitive, engineered with purpose-built fast LLMs [3] for tasks like intent detection over multi-turn, parameter identification and extraction, triggering single/multiple function calls, and offers convenience features to auto dispatch LLM calls for summarization based on data from your APIs via system prompts configured in archgw.<p>Today, the approach to build a smart production-ready agent is weaving together a large set of mono-functional opinionated libraries, adding extra layers like LLM-based preprocessing to determine things like relevance and safety of the user's prompt (e.g. applying governance and guardrails). Once past that stage, developers must extract relevant information from the user prompt to determine intent, extract parameters as necessary, package relevant tools calls to an LLM to trigger a backend API to execute particular domain-specific task. etc. After all that is done then only are developers ready to trigger an LLM call for summarization and must manage upstream error handling and retry logic themselves. Not to mention, if they want to experiment with multiple LLMs or move between LLM versions, they have to write crufty undifferentiated code. This entire experience is slow, error prone, cumbersome, and not specifically unique.<p>Prior to building archgw, the team spent time building Envoy [2] at Lyft, API Gateway at AWS, specialized search and intent models at Microsoft Research and worked on safety at Meta. archgw was born out of the belief that several rules based mono-functional tools should be converged into a multi-functional infrastructure primitive designed for prompts and agents. We built archgw on the highly popular, battle-tested open source proxy Envoy and re-imagined it for prompts and agents. For this we had to build blazing fast LLMs [3] that can handle crufty, ahead-in-the-request-path type of work in handling and processing prompts that are sent to an agent, so that developers can focus on what matters most: building fast personalized agents without the unnecessary prompt engineering and systems integration work needed to get there.<p>Here are some additional details about the open source project. arghw is written in rust, and the request path has three main parts:<p>* Listener subsystem which handles downstream (ingress) and upstream (egress) request processing.<p>* Prompt handler subsystem. This is where archgw makes decisions on the safety of the incoming request via its prompt_guard primitive and identifies where to forward the conversation to via its prompt_target primitive.<p>* Model serving subsystem is the interface that hosts all the lightweight LLMs engineered in archgw and offers a framework for things like hallucination detection of our these models<p>We loved building this open source project, and our belief is that this infra primitive would help developers build faster, safer and more personalized agents without all the manual prompt engineering and systems integration work needed to get there. We hope to invite other developers to use and improve Arch. Please give it a shot and leave feedback here, or at our discord channel [4]<p>Also here is a quick demo of the project in action [5]. You can check out our public docs here at [6]. Our models are also available here [7].<p>[1] <a href=\"https://github.com/katanemo/archgw\">https://github.com/<em>katanemo</em>/archgw</a><p>[2] <a href=\"https://www.envoyproxy.io/\" rel=\"nofollow\">https://www.envoyproxy.io/</a><p>[3] <a href=\"https://huggingface.co/collections/katanemo/arch-function-66f209a693ea8df14317ad68\" rel=\"nofollow\">https://huggingface.co/collections/<em>katanemo</em>/arch-function-66...</a><p>[4] <a href=\"https://discord.com/channels/1292630766827737088/1292630768283029638\" rel=\"nofollow\">https://discord.com/channels/1292630766827737088/12926307682...</a><p>[5] <a href=\"https://www.youtube.com/watch?v=I4Lbhr-NNXk\" rel=\"nofollow\">https://www.youtube.com/watch?v=I4Lbhr-NNXk</a><p>[6] <a href=\"https://docs.archgw.com/\" rel=\"nofollow\">https://docs.archgw.com/</a><p>[7] <a href=\"https://huggingface.co/katanemo\" rel=\"nofollow\">https://huggingface.co/<em>katanemo</em></a>"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: archgw: open-source, intelligent proxy for AI agents, built on Envoy"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["katanemo"],"value":"https://github.com/<em>katanemo</em>/archgw"}},"_tags":["story","author_adilhafeez","story_42187132","show_hn"],"author":"adilhafeez","children":[42187214,42187219,42187352,42187646,42188134,42188203,42189040,42189705],"created_at":"2024-11-19T19:26:28Z","created_at_i":1732044388,"num_comments":14,"objectID":"42187132","points":20,"story_id":42187132,"story_text":"Hi HN! This is Adil, Salman, Co and Shuguang and we&#x27;re excited to introduce archgw [1], an open source intelligent proxy for agents built on Envoy [2]. Arch moves the critical but crufty work around safety, observability, and routing of prompts outside business logic. Arch is a uniquely intelligent infrastructure primitive, engineered with purpose-built fast LLMs [3] for tasks like intent detection over multi-turn, parameter identification and extraction, triggering single&#x2F;multiple function calls, and offers convenience features to auto dispatch LLM calls for summarization based on data from your APIs via system prompts configured in archgw.<p>Today, the approach to build a smart production-ready agent is weaving together a large set of mono-functional opinionated libraries, adding extra layers like LLM-based preprocessing to determine things like relevance and safety of the user&#x27;s prompt (e.g. applying governance and guardrails). Once past that stage, developers must extract relevant information from the user prompt to determine intent, extract parameters as necessary, package relevant tools calls to an LLM to trigger a backend API to execute particular domain-specific task. etc. After all that is done then only are developers ready to trigger an LLM call for summarization and must manage upstream error handling and retry logic themselves. Not to mention, if they want to experiment with multiple LLMs or move between LLM versions, they have to write crufty undifferentiated code. This entire experience is slow, error prone, cumbersome, and not specifically unique.<p>Prior to building archgw, the team spent time building Envoy [2] at Lyft, API Gateway at AWS, specialized search and intent models at Microsoft Research and worked on safety at Meta. archgw was born out of the belief that several rules based mono-functional tools should be converged into a multi-functional infrastructure primitive designed for prompts and agents. We built archgw on the highly popular, battle-tested open source proxy Envoy and re-imagined it for prompts and agents. For this we had to build blazing fast LLMs [3] that can handle crufty, ahead-in-the-request-path type of work in handling and processing prompts that are sent to an agent, so that developers can focus on what matters most: building fast personalized agents without the unnecessary prompt engineering and systems integration work needed to get there.<p>Here are some additional details about the open source project. arghw is written in rust, and the request path has three main parts:<p>* Listener subsystem which handles downstream (ingress) and upstream (egress) request processing.<p>* Prompt handler subsystem. This is where archgw makes decisions on the safety of the incoming request via its prompt_guard primitive and identifies where to forward the conversation to via its prompt_target primitive.<p>* Model serving subsystem is the interface that hosts all the lightweight LLMs engineered in archgw and offers a framework for things like hallucination detection of our these models<p>We loved building this open source project, and our belief is that this infra primitive would help developers build faster, safer and more personalized agents without all the manual prompt engineering and systems integration work needed to get there. We hope to invite other developers to use and improve Arch. Please give it a shot and leave feedback here, or at our discord channel [4]<p>Also here is a quick demo of the project in action [5]. You can check out our public docs here at [6]. Our models are also available here [7].<p>[1] <a href=\"https:&#x2F;&#x2F;github.com&#x2F;katanemo&#x2F;archgw\">https:&#x2F;&#x2F;github.com&#x2F;katanemo&#x2F;archgw</a><p>[2] <a href=\"https:&#x2F;&#x2F;www.envoyproxy.io&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;www.envoyproxy.io&#x2F;</a><p>[3] <a href=\"https:&#x2F;&#x2F;huggingface.co&#x2F;collections&#x2F;katanemo&#x2F;arch-function-66f209a693ea8df14317ad68\" rel=\"nofollow\">https:&#x2F;&#x2F;huggingface.co&#x2F;collections&#x2F;katanemo&#x2F;arch-function-66...</a><p>[4] <a href=\"https:&#x2F;&#x2F;discord.com&#x2F;channels&#x2F;1292630766827737088&#x2F;1292630768283029638\" rel=\"nofollow\">https:&#x2F;&#x2F;discord.com&#x2F;channels&#x2F;1292630766827737088&#x2F;12926307682...</a><p>[5] <a href=\"https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=I4Lbhr-NNXk\" rel=\"nofollow\">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=I4Lbhr-NNXk</a><p>[6] <a href=\"https:&#x2F;&#x2F;docs.archgw.com&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;docs.archgw.com&#x2F;</a><p>[7] <a href=\"https:&#x2F;&#x2F;huggingface.co&#x2F;katanemo\" rel=\"nofollow\">https:&#x2F;&#x2F;huggingface.co&#x2F;katanemo</a>","title":"Show HN: archgw: open-source, intelligent proxy for AI agents, built on Envoy","updated_at":"2024-11-20T20:08:11Z","url":"https://github.com/katanemo/archgw"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"adilhafeez"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["katanemo"],"value":"Hey HN \u2014 I\u2019m Adil from <em>Katanemo</em> (with Salman, Shuguang, and Meiyu)<p>We previously shared an early version of this project as ArchGW. Based on customer feedback, the scope expanded from \u201cLLM routing and model access\u201d into something broader: delivery infrastructure for agentic applications. We renamed it to Plano and reworked the architecture accordingly.<p>The problem<p>On-the-ground AI practitioners will tell you that calling an LLM is not the hard part. The really hard part is delivering agentic applications to production quickly and reliably, then iterating without rewriting system code every time. In practice, teams keep rebuilding the same concerns that sit outside any single agent\u2019s core logic:<p>This includes model agility \u2014 the ability to pull from a large set of LLMs and swap providers without refactoring prompts or streaming handlers. They need to learn from production by collecting signals and traces that tell them what to fix. They need consistent policy enforcement for moderation and jailbreak protection, rather than sprinkling hooks across codebases. And they need multi-agent patterns like handoff and specialization without turning their app into orchestration glue.<p>These concerns get rebuilt and maintained inside fast-changing frameworks and application code, coupling product logic to infrastructure decisions. It\u2019s brittle, and pulls teams away from core product work into plumbing they shouldn\u2019t have to own.<p>What Plano does<p>Plano moves core delivery concerns out of process into a modular proxy and dataplane designed for agents. It supports inbound listeners (agent orchestration, safety and moderation hooks), outbound listeners (hosted or API-based LLM routing), or both together.<p>Plano provides the following capabilities via a unified, protocol-native, framework-friendly dataplane:<p>- Orchestration: Low-latency routing and handoff between agents. Add or change agents without modifying app code, and evolve strategies centrally instead of duplicating logic across services.<p>- Guardrails &amp; Memory Hooks: Apply jailbreak protection, content policies, and context workflows (rewriting, retrieval, redaction) once via filter chains. This centralizes governance and ensures consistent behavior across your stack.<p>- Model Agility: Route by model name, semantic alias, or preference-based policies. Swap or add models without refactoring prompts, tool calls, or streaming handlers.<p>- Agentic Signals\u2122: Zero-code capture of behavior signals, traces, and metrics across every agent, surfacing traces, token usage, and learning signals in one place.<p>The goal is to keep application code focused on product logic while Plano owns delivery mechanics.<p>More on Architecture<p>Plano has two main parts:<p>Envoy-based data plane. Uses Envoy\u2019s HTTP connection management to talk to model APIs, services, and tool backends. We didn\u2019t build a separate model server\u2014Envoy already handles streaming, retries, timeouts, and connection pooling. Some of us are core Envoy contributors at <em>Katanemo</em>.<p>Brightstaff, a lightweight controller written in Rust. It inspects prompts and conversation state, decides which upstreams to call and in what order, and coordinates routing and fallback. It uses small LLMs (1\u20134B parameters) trained for constrained routing and orchestration. These models do not generate responses and fall back to static policies on failure. The models are open sourced here: <a href=\"https://huggingface.co/katanemo\" rel=\"nofollow\">https://huggingface.co/<em>katanemo</em></a><p>Plano runs alongside your app servers (cloud, on-prem, or local dev), doesn\u2019t require a GPU, and leaves GPUs where your models are hosted.<p>Repo <a href=\"https://github.com/katanemo/plano\" rel=\"nofollow\">https://github.com/<em>katanemo</em>/plano</a> + docs <a href=\"https://docs.planoai.dev/\" rel=\"nofollow\">https://docs.planoai.dev/</a>"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Plano \u2013 Edge and service proxy with orchestration for AI agents"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["katanemo"],"value":"https://github.com/<em>katanemo</em>/plano"}},"_tags":["story","author_adilhafeez","story_46517177","show_hn"],"author":"adilhafeez","children":[46517254,46563685],"created_at":"2026-01-06T19:20:42Z","created_at_i":1767727242,"num_comments":2,"objectID":"46517177","points":8,"story_id":46517177,"story_text":"Hey HN \u2014 I\u2019m Adil from Katanemo (with Salman, Shuguang, and Meiyu)<p>We previously shared an early version of this project as ArchGW. Based on customer feedback, the scope expanded from \u201cLLM routing and model access\u201d into something broader: delivery infrastructure for agentic applications. We renamed it to Plano and reworked the architecture accordingly.<p>The problem<p>On-the-ground AI practitioners will tell you that calling an LLM is not the hard part. The really hard part is delivering agentic applications to production quickly and reliably, then iterating without rewriting system code every time. In practice, teams keep rebuilding the same concerns that sit outside any single agent\u2019s core logic:<p>This includes model agility \u2014 the ability to pull from a large set of LLMs and swap providers without refactoring prompts or streaming handlers. They need to learn from production by collecting signals and traces that tell them what to fix. They need consistent policy enforcement for moderation and jailbreak protection, rather than sprinkling hooks across codebases. And they need multi-agent patterns like handoff and specialization without turning their app into orchestration glue.<p>These concerns get rebuilt and maintained inside fast-changing frameworks and application code, coupling product logic to infrastructure decisions. It\u2019s brittle, and pulls teams away from core product work into plumbing they shouldn\u2019t have to own.<p>What Plano does<p>Plano moves core delivery concerns out of process into a modular proxy and dataplane designed for agents. It supports inbound listeners (agent orchestration, safety and moderation hooks), outbound listeners (hosted or API-based LLM routing), or both together.<p>Plano provides the following capabilities via a unified, protocol-native, framework-friendly dataplane:<p>- Orchestration: Low-latency routing and handoff between agents. Add or change agents without modifying app code, and evolve strategies centrally instead of duplicating logic across services.<p>- Guardrails &amp; Memory Hooks: Apply jailbreak protection, content policies, and context workflows (rewriting, retrieval, redaction) once via filter chains. This centralizes governance and ensures consistent behavior across your stack.<p>- Model Agility: Route by model name, semantic alias, or preference-based policies. Swap or add models without refactoring prompts, tool calls, or streaming handlers.<p>- Agentic Signals\u2122: Zero-code capture of behavior signals, traces, and metrics across every agent, surfacing traces, token usage, and learning signals in one place.<p>The goal is to keep application code focused on product logic while Plano owns delivery mechanics.<p>More on Architecture<p>Plano has two main parts:<p>Envoy-based data plane. Uses Envoy\u2019s HTTP connection management to talk to model APIs, services, and tool backends. We didn\u2019t build a separate model server\u2014Envoy already handles streaming, retries, timeouts, and connection pooling. Some of us are core Envoy contributors at Katanemo.<p>Brightstaff, a lightweight controller written in Rust. It inspects prompts and conversation state, decides which upstreams to call and in what order, and coordinates routing and fallback. It uses small LLMs (1\u20134B parameters) trained for constrained routing and orchestration. These models do not generate responses and fall back to static policies on failure. The models are open sourced here: <a href=\"https:&#x2F;&#x2F;huggingface.co&#x2F;katanemo\" rel=\"nofollow\">https:&#x2F;&#x2F;huggingface.co&#x2F;katanemo</a><p>Plano runs alongside your app servers (cloud, on-prem, or local dev), doesn\u2019t require a GPU, and leaves GPUs where your models are hosted.<p>Repo <a href=\"https:&#x2F;&#x2F;github.com&#x2F;katanemo&#x2F;plano\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;katanemo&#x2F;plano</a> + docs <a href=\"https:&#x2F;&#x2F;docs.planoai.dev&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;docs.planoai.dev&#x2F;</a>","title":"Show HN: Plano \u2013 Edge and service proxy with orchestration for AI agents","updated_at":"2026-03-05T23:17:38Z","url":"https://github.com/katanemo/plano"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"sparacha"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["katanemo"],"value":"Hi HN<p>My name is Salman and I work on Arch GW - the intelligent gateway designed to protect, observe, and personalize LLM applications with your APIs. <a href=\"https://github.com/katanemo/arch\">https://github.com/<em>katanemo</em>/arch</a><p>Our team built Envoy Proxy at Lyft, and re-imagined it with the belief that: Prompts are nuanced and opaque user requests, which require the same capabilities as traditional HTTP requests including secure handling, intelligent routing, robust observability, and integration with backend (API) systems for personalization \u2013 all outside business logic.<p>Engineered with purpose-built LLMs, Arch handles the critical but undifferentiated tasks related to the handling and processing of prompts, including detecting and rejecting jailbreak attempts, intelligently calling &quot;backend&quot; APIs to fulfill the user's request represented in a prompt, routing to and offering disaster recovery between upstream LLMs, and managing the observability of prompts and LLM interactions in a centralized way.<p>Core Features:<p>* Built on Envoy: Arch runs alongside application servers, and builds on top of Envoy's proven HTTP management and scalability features to handle ingress and egress traffic related to prompts and LLMs.<p>* Function Calling for fast Agentic and RAG apps. Engineered with purpose-built LLMs to handle fast, cost-effective, and accurate prompt-based tasks like function/API calling, and parameter extraction from prompts.<p>* Prompt Guard: Arch centralizes prompt guardrails to prevent jailbreak attempts and ensure safe user interactions without writing a single line of code.<p>* Traffic Management: Arch manages LLM calls, offering smart retries, automatic cutover, and resilient upstream connections for continuous availability.<p>* Standards-based Observability: Arch uses the W3C Trace Context standard to enable complete request tracing across applications, ensuring compatibility with observability tools, and provides metrics to monitor latency, token usage, and error rates.<p>We are just getting started, and would love feedback and contribution from the community"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Arch GW \u2013 Distributed gateway for agents, engineered with small LLMs"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://docs.archgw.com/"}},"_tags":["story","author_sparacha","story_42094897","show_hn"],"author":"sparacha","created_at":"2024-11-09T15:25:07Z","created_at_i":1731165907,"num_comments":0,"objectID":"42094897","points":7,"story_id":42094897,"story_text":"Hi HN<p>My name is Salman and I work on Arch GW - the intelligent gateway designed to protect, observe, and personalize LLM applications with your APIs. <a href=\"https:&#x2F;&#x2F;github.com&#x2F;katanemo&#x2F;arch\">https:&#x2F;&#x2F;github.com&#x2F;katanemo&#x2F;arch</a><p>Our team built Envoy Proxy at Lyft, and re-imagined it with the belief that: Prompts are nuanced and opaque user requests, which require the same capabilities as traditional HTTP requests including secure handling, intelligent routing, robust observability, and integration with backend (API) systems for personalization \u2013 all outside business logic.<p>Engineered with purpose-built LLMs, Arch handles the critical but undifferentiated tasks related to the handling and processing of prompts, including detecting and rejecting jailbreak attempts, intelligently calling &quot;backend&quot; APIs to fulfill the user&#x27;s request represented in a prompt, routing to and offering disaster recovery between upstream LLMs, and managing the observability of prompts and LLM interactions in a centralized way.<p>Core Features:<p>* Built on Envoy: Arch runs alongside application servers, and builds on top of Envoy&#x27;s proven HTTP management and scalability features to handle ingress and egress traffic related to prompts and LLMs.<p>* Function Calling for fast Agentic and RAG apps. Engineered with purpose-built LLMs to handle fast, cost-effective, and accurate prompt-based tasks like function&#x2F;API calling, and parameter extraction from prompts.<p>* Prompt Guard: Arch centralizes prompt guardrails to prevent jailbreak attempts and ensure safe user interactions without writing a single line of code.<p>* Traffic Management: Arch manages LLM calls, offering smart retries, automatic cutover, and resilient upstream connections for continuous availability.<p>* Standards-based Observability: Arch uses the W3C Trace Context standard to enable complete request tracing across applications, ensuring compatibility with observability tools, and provides metrics to monitor latency, token usage, and error rates.<p>We are just getting started, and would love feedback and contribution from the community","title":"Show HN: Arch GW \u2013 Distributed gateway for agents, engineered with small LLMs","updated_at":"2024-11-09T17:29:57Z","url":"https://docs.archgw.com/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"honorable_judge"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["katanemo"],"value":"I've been building agentic apps for some large Fortune 500 companies (T-Mobile, Twilio, etc.) and developed a mental model that serves as a practical guide in building agentic apps: separate the high-level agent specific logic from low-level platform capabilities. I call it the L-MM: the Logical Mental Model for LLM applications.<p>This mental model has not only been tremendously helpful in building agents but also helping customers think about the development process - so when I am done with a consulting engagement they can move faster across the stack and enable engineers and platform teams to work concurrently without interference, boosting productivity.<p>So what is the high-level logic vs. the low-level platform work?<p>High-Level Logic (Agent &amp; Task Specific)<p>Tools and Environment - These are specific integrations and capabilities that allow agents to interact with external systems or APIs to perform real-world tasks. Examples include:<p><pre><code>    Booking a table via OpenTable API\n    Scheduling calendar events via Google Calendar or Microsoft Outlook\n    Retrieving and updating data from CRM platforms like Salesforce\n    Utilizing payment gateways to complete transactions\n</code></pre>\nRole and Instructions - Clearly defining an agent's persona, responsibilities, and explicit instructions is essential for predictable and coherent behavior. This includes:<p><pre><code>    The &quot;personality&quot; of the agent (e.g., professional assistant)\n    Explicit boundaries around task completion (&quot;done criteria&quot;)\n    Behavioral guidelines for handling unexpected inputs or situations\n</code></pre>\nLow-Level Logic (Common Platform Capabilities)<p>Routing - Efficiently coordinating tasks between multiple specialized agents, ensuring seamless hand-offs and effective delegation:<p><pre><code>    Implementing intelligent load balancing and dynamic agent selection based on task context\n    Supporting retries, failover strategies, and fallback mechanisms\n</code></pre>\nGuardrails - Centralized mechanisms to safeguard interactions and ensure reliability and safety:<p><pre><code>    Filtering or moderating sensitive or harmful content\n    Real-time compliance checks for industry-specific regulations (e.g., GDPR, HIPAA)\n    Threshold-based alerts and automated corrective actions to prevent misuse\n</code></pre>\nAccess to LLMs - Providing robust and centralized access to multiple LLMs ensures high availability and scalability:<p><pre><code>    Implementing smart retry logic with exponential backoff\n    Centralized rate limiting and quota management to optimize usage\n    Handling diverse LLM backends transparently (OpenAI, Cohere, local open-source models, etc.)\n</code></pre>\nObservability - Comprehensive visibility into system performance and interactions using industry-standard practices:\n    W3C Trace Context compatible distributed tracing for clear visibility across requests\n    Detailed logging and metrics collection (latency, throughput, error rates, token usage)\n    Easy integration with popular observability platforms like Grafana, Prometheus, Datadog, and OpenTelemetry<p>Why This Matters<p>By adopting this structured mental model, teams can achieve clear separation of concerns, improving collaboration, reducing complexity, and accelerating the development of scalable, reliable, and safe agentic applications.<p>I'm actively working on addressing challenges in this domain. If you're navigating similar problems or have insights to share, let's discuss further - i'll leave some links about the stack too if folks want it.<p>High-level framework - <a href=\"https://openai.github.io/openai-agents-python/\" rel=\"nofollow\">https://openai.github.io/openai-agents-python/</a>\nLow-level infrastructure - <a href=\"https://github.com/katanemo/archgw\">https://github.com/<em>katanemo</em>/archgw</a>"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: LMM for LLMs \u2013 A mental model for building LLM apps"}},"_tags":["story","author_honorable_judge","story_43765467","show_hn"],"author":"honorable_judge","children":[43765529],"created_at":"2025-04-22T19:32:26Z","created_at_i":1745350346,"num_comments":2,"objectID":"43765467","points":6,"story_id":43765467,"story_text":"I&#x27;ve been building agentic apps for some large Fortune 500 companies (T-Mobile, Twilio, etc.) and developed a mental model that serves as a practical guide in building agentic apps: separate the high-level agent specific logic from low-level platform capabilities. I call it the L-MM: the Logical Mental Model for LLM applications.<p>This mental model has not only been tremendously helpful in building agents but also helping customers think about the development process - so when I am done with a consulting engagement they can move faster across the stack and enable engineers and platform teams to work concurrently without interference, boosting productivity.<p>So what is the high-level logic vs. the low-level platform work?<p>High-Level Logic (Agent &amp; Task Specific)<p>Tools and Environment - These are specific integrations and capabilities that allow agents to interact with external systems or APIs to perform real-world tasks. Examples include:<p><pre><code>    Booking a table via OpenTable API\n    Scheduling calendar events via Google Calendar or Microsoft Outlook\n    Retrieving and updating data from CRM platforms like Salesforce\n    Utilizing payment gateways to complete transactions\n</code></pre>\nRole and Instructions - Clearly defining an agent&#x27;s persona, responsibilities, and explicit instructions is essential for predictable and coherent behavior. This includes:<p><pre><code>    The &quot;personality&quot; of the agent (e.g., professional assistant)\n    Explicit boundaries around task completion (&quot;done criteria&quot;)\n    Behavioral guidelines for handling unexpected inputs or situations\n</code></pre>\nLow-Level Logic (Common Platform Capabilities)<p>Routing - Efficiently coordinating tasks between multiple specialized agents, ensuring seamless hand-offs and effective delegation:<p><pre><code>    Implementing intelligent load balancing and dynamic agent selection based on task context\n    Supporting retries, failover strategies, and fallback mechanisms\n</code></pre>\nGuardrails - Centralized mechanisms to safeguard interactions and ensure reliability and safety:<p><pre><code>    Filtering or moderating sensitive or harmful content\n    Real-time compliance checks for industry-specific regulations (e.g., GDPR, HIPAA)\n    Threshold-based alerts and automated corrective actions to prevent misuse\n</code></pre>\nAccess to LLMs - Providing robust and centralized access to multiple LLMs ensures high availability and scalability:<p><pre><code>    Implementing smart retry logic with exponential backoff\n    Centralized rate limiting and quota management to optimize usage\n    Handling diverse LLM backends transparently (OpenAI, Cohere, local open-source models, etc.)\n</code></pre>\nObservability - Comprehensive visibility into system performance and interactions using industry-standard practices:\n    W3C Trace Context compatible distributed tracing for clear visibility across requests\n    Detailed logging and metrics collection (latency, throughput, error rates, token usage)\n    Easy integration with popular observability platforms like Grafana, Prometheus, Datadog, and OpenTelemetry<p>Why This Matters<p>By adopting this structured mental model, teams can achieve clear separation of concerns, improving collaboration, reducing complexity, and accelerating the development of scalable, reliable, and safe agentic applications.<p>I&#x27;m actively working on addressing challenges in this domain. If you&#x27;re navigating similar problems or have insights to share, let&#x27;s discuss further - i&#x27;ll leave some links about the stack too if folks want it.<p>High-level framework - <a href=\"https:&#x2F;&#x2F;openai.github.io&#x2F;openai-agents-python&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;openai.github.io&#x2F;openai-agents-python&#x2F;</a>\nLow-level infrastructure - <a href=\"https:&#x2F;&#x2F;github.com&#x2F;katanemo&#x2F;archgw\">https:&#x2F;&#x2F;github.com&#x2F;katanemo&#x2F;archgw</a>","title":"Show HN: LMM for LLMs \u2013 A mental model for building LLM apps","updated_at":"2025-05-06T23:59:09Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"sparacha"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["katanemo"],"value":"Hi HN!<p>My name is Salman Paracha. I aam the the Founder/CEO of <em>Katanemo</em> - the organization behind the open source Arch GW (an intelligent gateway for prompts - <a href=\"https://github.com/katanemo/arch\">https://github.com/<em>katanemo</em>/arch</a>). Today, we are making the (SOTA) LLMs engineered in Arch GW for function calling scenarios available under an OSS license that borrows from Llama's community license.<p>What is function calling? Function calling helps developers personalize apps by calling application-specific operations via user prompts. This involves any predefined functions or APIs you want to expose to perform tasks, gather information, or manipulate data - via prompts. With function calling, you get to support agentic workflows tailored to domain-specific use cases - from updating insurance claims to creating ad campaigns. Arch-Function analyzes prompts, extracts critical information from prompts, engages in lightweight conversations with the user to gather any missing parameters and makes API calls so that you can focus on writing business logic.<p>Arch-Function is an auto-regressive model that if run on the NVIDIA A100 GPUs using vLLM offers throughput of ~1900/output tokens per second, and a output token price of $0.10/M token. This is ~12x faster and 44x cheaper than GPT-4o."},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Arch-Function: 3B parameter LLM that beats GPT-4o on function calling"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["katanemo"],"value":"https://huggingface.co/<em>katanemo</em>/Arch-Function-3B"}},"_tags":["story","author_sparacha","story_41858608","show_hn"],"author":"sparacha","created_at":"2024-10-16T13:02:22Z","created_at_i":1729083742,"num_comments":0,"objectID":"41858608","points":5,"story_id":41858608,"story_text":"Hi HN!<p>My name is Salman Paracha. I aam the the Founder&#x2F;CEO of Katanemo - the organization behind the open source Arch GW (an intelligent gateway for prompts - <a href=\"https:&#x2F;&#x2F;github.com&#x2F;katanemo&#x2F;arch\">https:&#x2F;&#x2F;github.com&#x2F;katanemo&#x2F;arch</a>). Today, we are making the (SOTA) LLMs engineered in Arch GW for function calling scenarios available under an OSS license that borrows from Llama&#x27;s community license.<p>What is function calling? Function calling helps developers personalize apps by calling application-specific operations via user prompts. This involves any predefined functions or APIs you want to expose to perform tasks, gather information, or manipulate data - via prompts. With function calling, you get to support agentic workflows tailored to domain-specific use cases - from updating insurance claims to creating ad campaigns. Arch-Function analyzes prompts, extracts critical information from prompts, engages in lightweight conversations with the user to gather any missing parameters and makes API calls so that you can focus on writing business logic.<p>Arch-Function is an auto-regressive model that if run on the NVIDIA A100 GPUs using vLLM offers throughput of ~1900&#x2F;output tokens per second, and a output token price of $0.10&#x2F;M token. This is ~12x faster and 44x cheaper than GPT-4o.","title":"Show HN: Arch-Function: 3B parameter LLM that beats GPT-4o on function calling","updated_at":"2024-10-17T16:36:14Z","url":"https://huggingface.co/katanemo/Arch-Function-3B"}],"hitsPerPage":10,"nbHits":181,"nbPages":19,"page":0,"params":"query=katanemo&hitsPerPage=10&advancedSyntax=true&analyticsTags=backend","processingTimeMS":20,"processingTimingsMS":{"_request":{"queue":4,"roundTrip":20},"afterFetch":{"format":{"highlighting":2,"total":2}},"fetch":{"query":19,"total":20},"total":20},"query":"katanemo","serverTimeMS":28}
