{"author":"ij23","children":[{"author":"breckenedge","children":[{"author":"detente18","children":[{"author":"d4rkp4ttern","children":[{"author":"detente18","children":[],"created_at":"2023-08-12T14:53:38.000Z","created_at_i":1691852018,"id":37100787,"options":[],"parent_id":37099074,"points":null,"story_id":37095542,"text":"Hi @d4rkp4ttern - yes it&#x27;s standardized to the openai format - <a href=\"https:&#x2F;&#x2F;litellm.readthedocs.io&#x2F;en&#x2F;latest&#x2F;stream&#x2F;\" rel=\"nofollow noreferrer\">https:&#x2F;&#x2F;litellm.readthedocs.io&#x2F;en&#x2F;latest&#x2F;stream&#x2F;</a>","title":null,"type":"comment","url":null}],"created_at":"2023-08-12T11:27:45.000Z","created_at_i":1691839665,"id":37099074,"options":[],"parent_id":37096722,"points":null,"story_id":37095542,"text":"What about streaming output \u2014 is there a uniform interface for all models for streaming output?","title":null,"type":"comment","url":null}],"created_at":"2023-08-12T03:27:15.000Z","created_at_i":1691810835,"id":37096722,"options":[],"parent_id":37096039,"points":null,"story_id":37095542,"text":"Hey @breckenedge yes it does! Exactly the same way as the openai-python sdk. Here&#x27;s our code tutorial for it  <a href=\"https:&#x2F;&#x2F;github.com&#x2F;BerriAI&#x2F;litellm&#x2F;blob&#x2F;main&#x2F;cookbook&#x2F;liteLLM_function_calling.ipynb\">https:&#x2F;&#x2F;github.com&#x2F;BerriAI&#x2F;litellm&#x2F;blob&#x2F;main&#x2F;cookbook&#x2F;liteLL...</a>","title":null,"type":"comment","url":null}],"created_at":"2023-08-12T01:30:38.000Z","created_at_i":1691803838,"id":37096039,"options":[],"parent_id":37095542,"points":null,"story_id":37095542,"text":"Does it support the function-calling API? <a href=\"https:&#x2F;&#x2F;openai.com&#x2F;blog&#x2F;function-calling-and-other-api-updates\" rel=\"nofollow noreferrer\">https:&#x2F;&#x2F;openai.com&#x2F;blog&#x2F;function-calling-and-other-api-updat...</a>","title":null,"type":"comment","url":null},{"author":"jmorgan","children":[{"author":"detente18","children":[{"author":"jmorgan","children":[],"created_at":"2023-08-12T04:13:34.000Z","created_at_i":1691813614,"id":37096979,"options":[],"parent_id":37096748,"points":null,"story_id":37095542,"text":"Great. Let&#x27;s chat!","title":null,"type":"comment","url":null}],"created_at":"2023-08-12T03:31:58.000Z","created_at_i":1691811118,"id":37096748,"options":[],"parent_id":37096625,"points":null,"story_id":37095542,"text":"Hey @jmorgan - we love ollama!<p>Re: LocalLLM&#x27;s like Llama2 - yes we support self-deployed models, through Huggingface, Replicate, TogetherAI integrations.<p>We&#x27;re missing support for locally deployed models - and would love the help!<p>Re: ollama - we spent a couple hours trying to integrate ollama. We had a couple issues though, would love to try and support it. Got time to chat sometime this&#x2F;next week? I think this would be an awesome addition.","title":null,"type":"comment","url":null}],"created_at":"2023-08-12T03:09:41.000Z","created_at_i":1691809781,"id":37096625,"options":[],"parent_id":37095542,"points":null,"story_id":37095542,"text":"The idea of an LLM proxy is super compelling. There&#x27;s a lot of powerful ideas baked into the proxy form factor \u2013 I think you&#x27;ve listed out quite a few of them. It reminds me a bit of what Cloudflare did for the web: both making it faster and safer&#x2F;easier. Have you considered local LLMs at all for Llama 2? A few people and I have been working on <a href=\"https:&#x2F;&#x2F;github.com&#x2F;jmorganca&#x2F;ollama&#x2F;\">https:&#x2F;&#x2F;github.com&#x2F;jmorganca&#x2F;ollama&#x2F;</a> and was thinking it would be helpful to be able to augment it with a proxy layer like this. Not only that, but it might help folks dynamically choose to run locally (vs against a cloud LLM) for certain prompts.","title":null,"type":"comment","url":null},{"author":"romanzubenko","children":[{"author":"detente18","children":[],"created_at":"2023-08-12T03:27:42.000Z","created_at_i":1691810862,"id":37096725,"options":[],"parent_id":37096662,"points":null,"story_id":37095542,"text":"Thanks - that&#x27;s very kind.","title":null,"type":"comment","url":null},{"author":"jmorgan","children":[],"created_at":"2023-08-12T04:16:28.000Z","created_at_i":1691813788,"id":37096998,"options":[],"parent_id":37096662,"points":null,"story_id":37095542,"text":"I do think this is an apt analogy. I&#x27;ve heard a counterpoint that there won&#x27;t be enough &quot;destinations&quot; for this to work, but then it&#x27;s not hard to imagine a single order of magnitude more LLM &quot;destinations&quot; than today (all of which were launched in the last 12 months).<p>There&#x27;s also the fact that the data being sent to these LLM &quot;destinations&quot; could be significantly more valuable (or contain significantly more sensitive information) than the average segment identity or track objects.","title":null,"type":"comment","url":null}],"created_at":"2023-08-12T03:15:31.000Z","created_at_i":1691810131,"id":37096662,"options":[],"parent_id":37095542,"points":null,"story_id":37095542,"text":"Reminds me of analytics.js which later turned into Segment and 3b acquisition","title":null,"type":"comment","url":null},{"author":"downvotetruth","children":[{"author":"detente18","children":[],"created_at":"2023-08-12T03:26:00.000Z","created_at_i":1691810760,"id":37096715,"options":[],"parent_id":37096686,"points":null,"story_id":37095542,"text":"Hey @downvotetruth, the issues&#x2F;PR&#x27;s are actually ours (the creators). We&#x27;re actively working + using this repo, and use these as ways to keep track of ongoing work. We&#x27;re super open to contributions and help though!","title":null,"type":"comment","url":null}],"created_at":"2023-08-12T03:21:25.000Z","created_at_i":1691810485,"id":37096686,"options":[],"parent_id":37095542,"points":null,"story_id":37095542,"text":"There are unaddressed github issues like #29 added by whoiskatrin who is not listed as a contributor yet telephone #s and emails are listed on the readme- motive for that is ambiguous and is not likely to result in an optimal outcome.","title":null,"type":"comment","url":null},{"author":"deet","children":[{"author":"detente18","children":[],"created_at":"2023-08-12T04:01:57.000Z","created_at_i":1691812917,"id":37096925,"options":[],"parent_id":37096881,"points":null,"story_id":37095542,"text":"Hi @deet - Yes it is! This automatically stores the cost per query to the Supabase table - here&#x27;s how: <a href=\"https:&#x2F;&#x2F;github.com&#x2F;BerriAI&#x2F;litellm&#x2F;blob&#x2F;80d77fed7123af22201158528feadb83ab9a95bc&#x2F;litellm&#x2F;integrations&#x2F;supabase.py#L74\">https:&#x2F;&#x2F;github.com&#x2F;BerriAI&#x2F;litellm&#x2F;blob&#x2F;80d77fed7123af222011...</a><p>If you have ideas for improvement - we&#x27;d love a ticket&#x2F;PR!","title":null,"type":"comment","url":null}],"created_at":"2023-08-12T03:55:40.000Z","created_at_i":1691812540,"id":37096881,"options":[],"parent_id":37095542,"points":null,"story_id":37095542,"text":"Is user auth (and tracking token spend) within scope of this or is that better handled at a layer in front of this?","title":null,"type":"comment","url":null},{"author":"ada1981","children":[{"author":"ij23","children":[],"created_at":"2023-08-12T05:19:28.000Z","created_at_i":1691817568,"id":37097251,"options":[],"parent_id":37097078,"points":null,"story_id":37095542,"text":"Yes, you use your own API keys. You can set them as env variables. \nEither set them as os.environ[&#x27;OPENAI_API_KEY&#x27;] or set them in .env files:\n<a href=\"https:&#x2F;&#x2F;litellm.readthedocs.io&#x2F;en&#x2F;latest&#x2F;supported&#x2F;\" rel=\"nofollow noreferrer\">https:&#x2F;&#x2F;litellm.readthedocs.io&#x2F;en&#x2F;latest&#x2F;supported&#x2F;</a>","title":null,"type":"comment","url":null}],"created_at":"2023-08-12T04:32:43.000Z","created_at_i":1691814763,"id":37097078,"options":[],"parent_id":37095542,"points":null,"story_id":37095542,"text":"So I use my own accounts to access this correct? I didn\u2019t see in the documentation where I provide my credentials. I\u2019ll look again\u2026","title":null,"type":"comment","url":null},{"author":"kiratp","children":[{"author":"ij23","children":[{"author":"kiratp","children":[{"author":"bredren","children":[{"author":"detente18","children":[],"created_at":"2023-08-12T14:56:05.000Z","created_at_i":1691852165,"id":37100813,"options":[],"parent_id":37099410,"points":null,"story_id":37095542,"text":"Hey bredren - our supported list (if that&#x27;s helpful) is here <a href=\"https:&#x2F;&#x2F;litellm.readthedocs.io&#x2F;en&#x2F;latest&#x2F;supported&#x2F;\" rel=\"nofollow noreferrer\">https:&#x2F;&#x2F;litellm.readthedocs.io&#x2F;en&#x2F;latest&#x2F;supported&#x2F;</a><p>We&#x27;re adding new integrations every day, so if there&#x27;s any specific one you&#x27;d like to add feel free to let us know (discord&#x2F;ticket&#x2F;email&#x2F;etc.) - here&#x27;s my email: krrish@berri.ai","title":null,"type":"comment","url":null}],"created_at":"2023-08-12T12:15:04.000Z","created_at_i":1691842504,"id":37099410,"options":[],"parent_id":37097385,"points":null,"story_id":37095542,"text":"If it makes sense to expand scope to provide a particular model server and the group can easily be the best st it, I say go for it. But do it as a separate (but perhaps connected) project to this.<p>But in general I\u2019m in  agreement that this sounds like a separate concept than any given model server.<p>That said, where is a list of model servers for the most commonly wanted LLMs at this point?<p>Perhaps maintaining a list of those that do and don\u2019t work with the proxy would be helpful.","title":null,"type":"comment","url":null}],"created_at":"2023-08-12T05:47:03.000Z","created_at_i":1691819223,"id":37097385,"options":[],"parent_id":37097261,"points":null,"story_id":37095542,"text":"IMO don\u2019t try to be the one stop shop to host models. There are too many players with all sorts of advancements (eg: stopping grammar, continuous batching, novel quantization etc.) and you won\u2019t be able to keep up.<p>There is a ton of boilerplate around the actual model server that\u2019s just busy work , but if done wrong can be a huge performance suck. Solve that.<p>Build the proxy that works with the most model servers out there. Do it in a way that once you have mindshare, the model server makers will be find it easy to put up a PR so that they can claim your proxy supports their server.<p>Don\u2019t take a hard dependency on non-OSS stuff - being able to build an \u201con-prem\u201d solution (read \u201cdeployed into customer\u2019s VPC\u201d) is table stakes for anyone to use your offering for a lot of enterprise use cases.<p>Edit: another unsolved problem - different models need slightly different prompts to solve the same problem well\u2026","title":null,"type":"comment","url":null}],"created_at":"2023-08-12T05:21:15.000Z","created_at_i":1691817675,"id":37097261,"options":[],"parent_id":37097106,"points":null,"story_id":37095542,"text":"What local&#x2F;in-K8-cluster models servers would you recommend adding ?<p>Should we add support for llama.cpp and vllm.ai in the proxy server ? Or should we assume you can host them on your own infra and the proxy server requests your hosted model ?","title":null,"type":"comment","url":null}],"created_at":"2023-08-12T04:41:15.000Z","created_at_i":1691815275,"id":37097106,"options":[],"parent_id":37095542,"points":null,"story_id":37095542,"text":"This would be super useful if it supported local&#x2F;in-K8-cluster models.<p>Most production use cases probably need some sort of fall back of small -&gt; medium -&gt; large -&gt; GPT4.<p>Given the costs + low quota limits, I would be surprised if any significant portion of the market is falling back from one expensive proprietary* API to another.<p>With the rapid improvements in model servers like llama.cpp and vllm.ai, providing an abstraction layer for \u201cfastest model server of the month\u201d would be useful.","title":null,"type":"comment","url":null},{"author":"danielbln","children":[],"created_at":"2023-08-12T08:06:59.000Z","created_at_i":1691827619,"id":37098103,"options":[],"parent_id":37095542,"points":null,"story_id":37095542,"text":"Can you say more about the semantic caching, I can&#x27;t seem to find much in the docs.<p>edit: Found the cache notebook[1] and the calls to the cache and distance eval in the code[2]<p>[1] <a href=\"https:&#x2F;&#x2F;github.com&#x2F;BerriAI&#x2F;litellm&#x2F;blob&#x2F;main&#x2F;cookbook&#x2F;liteLLM_ChromaDB_Cache.ipynb\">https:&#x2F;&#x2F;github.com&#x2F;BerriAI&#x2F;litellm&#x2F;blob&#x2F;main&#x2F;cookbook&#x2F;liteLL...</a><p>[2] <a href=\"https:&#x2F;&#x2F;github.com&#x2F;BerriAI&#x2F;litellm&#x2F;blob&#x2F;80d77fed7123af22201158528feadb83ab9a95bc&#x2F;cookbook&#x2F;proxy-server&#x2F;main.py#L124\">https:&#x2F;&#x2F;github.com&#x2F;BerriAI&#x2F;litellm&#x2F;blob&#x2F;80d77fed7123af222011...</a>","title":null,"type":"comment","url":null},{"author":"bredren","children":[{"author":"detente18","children":[],"created_at":"2023-08-12T15:00:49.000Z","created_at_i":1691852449,"id":37100857,"options":[],"parent_id":37099449,"points":null,"story_id":37095542,"text":"Those are really cool use-cases. I wrote a bit about our initial motivation for LiteLLM here - <a href=\"https:&#x2F;&#x2F;hackernoon.com&#x2F;litellm-call-every-llm-api-like-its-openai\" rel=\"nofollow noreferrer\">https:&#x2F;&#x2F;hackernoon.com&#x2F;litellm-call-every-llm-api-like-its-o...</a>.<p>tldr; reliable model switching involved multiple 100-line if&#x2F;else statements, which made our code messy, and debugging in prod pretty hard.","title":null,"type":"comment","url":null},{"author":"detente18","children":[],"created_at":"2023-08-12T15:01:38.000Z","created_at_i":1691852498,"id":37100867,"options":[],"parent_id":37099449,"points":null,"story_id":37095542,"text":"If you&#x27;d like to add any of these ideas as notebooks to our cookbook - <a href=\"https:&#x2F;&#x2F;github.com&#x2F;BerriAI&#x2F;litellm&#x2F;tree&#x2F;main&#x2F;cookbook\">https:&#x2F;&#x2F;github.com&#x2F;BerriAI&#x2F;litellm&#x2F;tree&#x2F;main&#x2F;cookbook</a>, we&#x27;d love the contribution!","title":null,"type":"comment","url":null}],"created_at":"2023-08-12T12:20:05.000Z","created_at_i":1691842805,"id":37099449,"options":[],"parent_id":37095542,"points":null,"story_id":37095542,"text":"This makes sense to me. I\u2019d previously mentioned here [1]:<p>&gt; \u2026can\u2019t people just build their product with openAI or other and plan to move away based on the cost and fit for their circumstances?<p>&gt; Couldn\u2019t someone say prototype the entire product on some lower-quality LLM and occasionally pass requests to GPT4 to validate behavior?<p>This seems to reduce the switching costs to almost nothing.<p>[1] <a href=\"https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=36626943\">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=36626943</a>","title":null,"type":"comment","url":null},{"author":"marcopicentini","children":[{"author":"detente18","children":[{"author":"marcopicentini","children":[{"author":"reustle","children":[{"author":"marcopicentini","children":[{"author":"Menatombo","children":[],"created_at":"2023-08-17T02:15:21.000Z","created_at_i":1692238521,"id":37156350,"options":[],"parent_id":37104360,"points":null,"story_id":37095542,"text":"You have to build it if you want it for Ubuntu, Windows, or anything else. Just build Go on your machine and have at it.","title":null,"type":"comment","url":null}],"created_at":"2023-08-12T21:24:15.000Z","created_at_i":1691875455,"id":37104360,"options":[],"parent_id":37103503,"points":null,"story_id":37095542,"text":"It\u2019s only for MacOSX. I expect to load the model on a Ubuntu server, not on my local dev machine.","title":null,"type":"comment","url":null}],"created_at":"2023-08-12T19:36:51.000Z","created_at_i":1691869011,"id":37103503,"options":[],"parent_id":37102586,"points":null,"story_id":37095542,"text":"You can definitely start by checking out ollama, it was super helpful for me","title":null,"type":"comment","url":null},{"author":"detente18","children":[],"created_at":"2023-08-12T21:52:48.000Z","created_at_i":1691877168,"id":37104573,"options":[],"parent_id":37102586,"points":null,"story_id":37095542,"text":"Any reason you&#x27;re doing that vs. using Lambda Labs &#x2F; Replicate &#x2F; together.ai &#x2F; Banana.dev, etc.<p>There&#x27;s a lot of good model deployment platforms that would make it easy to call your model behind a hosted endpoint<p>--\nIf you do want to self-host - there&#x27;s some great libraries like <a href=\"https:&#x2F;&#x2F;github.com&#x2F;lm-sys&#x2F;FastChat\">https:&#x2F;&#x2F;github.com&#x2F;lm-sys&#x2F;FastChat</a> and <a href=\"https:&#x2F;&#x2F;github.com&#x2F;ggerganov&#x2F;llama.cpp\">https:&#x2F;&#x2F;github.com&#x2F;ggerganov&#x2F;llama.cpp</a> that might be helpful<p>If none of these really solve your issue - feel free to email me and I&#x27;m happy to help you figure something out - krrish@berri.ai","title":null,"type":"comment","url":null}],"created_at":"2023-08-12T17:53:54.000Z","created_at_i":1691862834,"id":37102586,"options":[],"parent_id":37100782,"points":null,"story_id":37095542,"text":"No idea yet. Which you recommend to start? \nIt will be hosted on a Ubuntu server (Digital Ocean, Linode etc..)","title":null,"type":"comment","url":null}],"created_at":"2023-08-12T14:53:06.000Z","created_at_i":1691851986,"id":37100782,"options":[],"parent_id":37100300,"points":null,"story_id":37095542,"text":"How do you plan on deploying llama 2? Is that via ollama&#x2F;fastchat&#x2F;etc.? We&#x27;re actively building out integrations to different providers, so if you have any preferred tooling - let me know!","title":null,"type":"comment","url":null}],"created_at":"2023-08-12T14:04:55.000Z","created_at_i":1691849095,"id":37100300,"options":[],"parent_id":37095542,"points":null,"story_id":37095542,"text":"Very cool! \nCould you add in the readme how to install llama 2 on the same machine ?","title":null,"type":"comment","url":null},{"author":"detente18","children":[],"created_at":"2023-08-12T23:48:14.000Z","created_at_i":1691884094,"id":37105236,"options":[],"parent_id":37095542,"points":null,"story_id":37095542,"text":"*Update*: for those asking how to run self-hosted llama2, we just added the ollama integration.<p>Here&#x27;s the tutorial - <a href=\"https:&#x2F;&#x2F;github.com&#x2F;BerriAI&#x2F;litellm&#x2F;blob&#x2F;main&#x2F;cookbook&#x2F;liteLLM_Ollama.ipynb\">https:&#x2F;&#x2F;github.com&#x2F;BerriAI&#x2F;litellm&#x2F;blob&#x2F;main&#x2F;cookbook&#x2F;liteLL...</a>","title":null,"type":"comment","url":null},{"author":"pkraft","children":[],"created_at":"2023-08-15T01:10:38.000Z","created_at_i":1692061838,"id":37128926,"options":[],"parent_id":37095542,"points":null,"story_id":37095542,"text":"Could you integrate with something like TextGen and then automatically have support for every locally running model they support?   Am I missing something?","title":null,"type":"comment","url":null}],"created_at":"2023-08-12T00:08:13.000Z","created_at_i":1691798893,"id":37095542,"options":[],"parent_id":null,"points":140,"story_id":37095542,"text":"Hello hacker news,<p>I\u2019m the maintainer of liteLLM() - package to simplify input&#x2F;output to OpenAI, Azure, Cohere, Anthropic, Hugging face API Endpoints: <a href=\"https:&#x2F;&#x2F;github.com&#x2F;BerriAI&#x2F;litellm&#x2F;\">https:&#x2F;&#x2F;github.com&#x2F;BerriAI&#x2F;litellm&#x2F;</a><p>We\u2019re open sourcing our implementation of liteLLM proxy: <a href=\"https:&#x2F;&#x2F;github.com&#x2F;BerriAI&#x2F;litellm&#x2F;blob&#x2F;main&#x2F;cookbook&#x2F;proxy-server&#x2F;readme.md\">https:&#x2F;&#x2F;github.com&#x2F;BerriAI&#x2F;litellm&#x2F;blob&#x2F;main&#x2F;cookbook&#x2F;proxy-...</a><p>TLDR: It has one API endpoint &#x2F;chat&#x2F;completions and standardizes input&#x2F;output for 50+ LLM models + handles logging, error tracking, caching, streaming<p>What can liteLLM proxy do?\n- It\u2019s a central place to manage all LLM provider integrations<p>- Consistent Input&#x2F;Output Format\n    - Call all models using the OpenAI format: completion(model, messages)\n    - Text responses will always be available at [&#x27;choices&#x27;][0][&#x27;message&#x27;][&#x27;content&#x27;]<p>- Error Handling Using Model Fallbacks (if GPT-4 fails, try llama2)<p>- Logging - Log Requests, Responses and Errors to Supabase, Posthog, Mixpanel, Sentry, Helicone<p>- Token Usage &amp; Spend - Track Input + Completion tokens used + Spend&#x2F;model<p>- Caching - Implementation of Semantic Caching<p>- Streaming &amp; Async Support - Return generators to stream text responses<p>You can deploy liteLLM to your own infrastructure using Railway, GCP, AWS, Azure<p>Happy completion() !","title":"Show HN: liteLLM Proxy Server: 50+ LLM Models, Error Handling, Caching","type":"story","url":"https://github.com/BerriAI/litellm/blob/main/cookbook/proxy-server/readme.md"}
