{"exhaustive":{"nbHits":false,"typo":false},"exhaustiveNbHits":false,"exhaustiveTypo":false,"hits":[{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"jacky2wong"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"<em>DeepEval</em> \u2013 Unit Testing for LLMs"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"https://github.com/mr-gpt/<em>deepeval</em>"}},"_tags":["story","author_jacky2wong","story_37157323"],"author":"jacky2wong","children":[37157644,37158122,37158426,37159322,37159837],"created_at":"2023-08-17T04:42:38Z","created_at_i":1692247358,"num_comments":31,"objectID":"37157323","points":79,"story_id":37157323,"title":"DeepEval \u2013 Unit Testing for LLMs","updated_at":"2025-10-11T18:55:08Z","url":"https://github.com/mr-gpt/deepeval"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"jacky2wong"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"Unit Test LlamaIndex with <em>DeepEval</em>"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://docs.confident-ai.com/docs/tutorials/evaluating-llamaindex"}},"_tags":["story","author_jacky2wong","story_37295408"],"author":"jacky2wong","children":[37298669],"created_at":"2023-08-28T15:14:21Z","created_at_i":1693235661,"num_comments":3,"objectID":"37295408","points":35,"story_id":37295408,"title":"Unit Test LlamaIndex with DeepEval","updated_at":"2024-09-20T15:05:02Z","url":"https://docs.confident-ai.com/docs/tutorials/evaluating-llamaindex"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"jacky2wong"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"Show HN: <em>DeepEval</em> \u2013 Evaluation and Unit Testing for LLMs"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"https://github.com/confident-ai/<em>deepeval</em>"}},"_tags":["story","author_jacky2wong","story_37649856","show_hn"],"author":"jacky2wong","children":[37649857,37650029,37650060,37650162],"created_at":"2023-09-25T20:08:36Z","created_at_i":1695672516,"num_comments":8,"objectID":"37649856","points":18,"story_id":37649856,"title":"Show HN: DeepEval \u2013 Evaluation and Unit Testing for LLMs","updated_at":"2024-09-20T15:13:00Z","url":"https://github.com/confident-ai/deepeval"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"jacky2wong"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"Show HN: <em>DeepEval</em> \u2013 Unit Testing for LLMs (Open Science)"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"https://github.com/confident-ai/<em>deepeval</em>"}},"_tags":["story","author_jacky2wong","story_37787328","show_hn"],"author":"jacky2wong","created_at":"2023-10-06T05:21:14Z","created_at_i":1696569674,"num_comments":0,"objectID":"37787328","points":6,"story_id":37787328,"title":"Show HN: DeepEval \u2013 Unit Testing for LLMs (Open Science)","updated_at":"2024-09-20T15:20:03Z","url":"https://github.com/confident-ai/deepeval"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"jacky2wong"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"Auto-Evaluation of LLMs with <em>DeepEval</em>"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://docs.confident-ai.com/docs/quickstart/synthetic-data-creation"}},"_tags":["story","author_jacky2wong","story_37348645"],"author":"jacky2wong","children":[37348646],"created_at":"2023-09-01T09:42:00Z","created_at_i":1693561320,"num_comments":1,"objectID":"37348645","points":2,"story_id":37348645,"title":"Auto-Evaluation of LLMs with DeepEval","updated_at":"2024-09-20T15:01:37Z","url":"https://docs.confident-ai.com/docs/quickstart/synthetic-data-creation"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"jacky2wong"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"<em>DeepEval</em> \u2013 PyTest for LLMs"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"https://github.com/mr-gpt/<em>deepeval</em>"}},"_tags":["story","author_jacky2wong","story_37142518"],"author":"jacky2wong","children":[37142519],"created_at":"2023-08-16T03:23:34Z","created_at_i":1692156214,"num_comments":1,"objectID":"37142518","points":2,"story_id":37142518,"title":"DeepEval \u2013 PyTest for LLMs","updated_at":"2024-09-20T14:49:40Z","url":"https://github.com/mr-gpt/deepeval"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"jacky2wong"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"<em>DeepEval</em> GuardRails \u2013 AI Alignment"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://docs.confident-ai.com/docs/tutorials/guardrails"}},"_tags":["story","author_jacky2wong","story_37713260"],"author":"jacky2wong","created_at":"2023-09-30T06:43:26Z","created_at_i":1696056206,"num_comments":0,"objectID":"37713260","points":2,"story_id":37713260,"title":"DeepEval GuardRails \u2013 AI Alignment","updated_at":"2024-09-20T15:11:24Z","url":"https://docs.confident-ai.com/docs/tutorials/guardrails"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"jacky2wong"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"<em>DeepEval</em> \u2013 Synthetic Data, Bulk Review, Custom Metric Logging"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"https://colabdoge.medium.com/<em>deepeval</em>-synthetic-data-bulk-review-custom-metric-logging-and-29b677946766"}},"_tags":["story","author_jacky2wong","story_37496563"],"author":"jacky2wong","created_at":"2023-09-13T13:37:20Z","created_at_i":1694612240,"num_comments":0,"objectID":"37496563","points":2,"story_id":37496563,"title":"DeepEval \u2013 Synthetic Data, Bulk Review, Custom Metric Logging","updated_at":"2024-09-20T15:08:25Z","url":"https://colabdoge.medium.com/deepeval-synthetic-data-bulk-review-custom-metric-logging-and-29b677946766"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"jacky2wong"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"<em>DeepEval</em> \u2013 Neural Framework for Testing LLMs"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"https://github.com/confident-ai/<em>deepeval</em>"}},"_tags":["story","author_jacky2wong","story_37310402"],"author":"jacky2wong","created_at":"2023-08-29T16:37:36Z","created_at_i":1693327056,"num_comments":0,"objectID":"37310402","points":2,"story_id":37310402,"title":"DeepEval \u2013 Neural Framework for Testing LLMs","updated_at":"2024-09-20T14:57:21Z","url":"https://github.com/confident-ai/deepeval"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"jacky2wong"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"<em>DeepEval</em> CLI"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"https://github.com/confident-ai/<em>deepeval</em>"}},"_tags":["story","author_jacky2wong","story_37284032"],"author":"jacky2wong","created_at":"2023-08-27T16:13:13Z","created_at_i":1693152793,"num_comments":0,"objectID":"37284032","points":2,"story_id":37284032,"title":"DeepEval CLI","updated_at":"2024-09-20T15:03:53Z","url":"https://github.com/confident-ai/deepeval"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"willmarquis"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"Has anyone ever used the Python framework \"<em>Deepeval</em>\"?"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"https://github.com/confident-ai/<em>deepeval</em>"}},"_tags":["story","author_willmarquis","story_44354992"],"author":"willmarquis","created_at":"2025-06-23T12:15:03Z","created_at_i":1750680903,"num_comments":0,"objectID":"44354992","points":1,"story_id":44354992,"title":"Has anyone ever used the Python framework \"Deepeval\"?","updated_at":"2025-06-23T12:19:20Z","url":"https://github.com/confident-ai/deepeval"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"jacky2wong"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"Testing for Image Similarity with <em>DeepEval</em>"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://docs.confident-ai.com/docs/image/image_similarity"}},"_tags":["story","author_jacky2wong","story_37742573"],"author":"jacky2wong","created_at":"2023-10-02T18:41:36Z","created_at_i":1696272096,"num_comments":0,"objectID":"37742573","points":1,"story_id":37742573,"title":"Testing for Image Similarity with DeepEval","updated_at":"2024-09-20T15:14:51Z","url":"https://docs.confident-ai.com/docs/image/image_similarity"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"jacky2wong"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"PDB Support for <em>DeepEval</em>"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://docs.confident-ai.com/docs/quickstart/"}},"_tags":["story","author_jacky2wong","story_37414366"],"author":"jacky2wong","created_at":"2023-09-07T03:19:51Z","created_at_i":1694056791,"num_comments":0,"objectID":"37414366","points":1,"story_id":37414366,"title":"PDB Support for DeepEval","updated_at":"2024-09-20T14:58:43Z","url":"https://docs.confident-ai.com/docs/quickstart/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"jeffreyip"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"Hi HN - we're Jeffrey and Kritin, and we're building Confident AI (<a href=\"https://confident-ai.com\">https://confident-ai.com</a>). This is the cloud platform for <em>DeepEval</em> (<a href=\"https://github.com/confident-ai/deepeval\">https://github.com/confident-ai/<em>deepeval</em></a>), our open-source package that helps engineers evaluate and unit-test LLM applications. Think Pytest for LLMs.<p>We spent the past year building <em>DeepEval</em> with the goal of providing the best LLM evaluation developer experience, growing it to run over 600K evaluations daily in CI/CD pipelines of enterprises like BCG, AstraZeneca, AXA, and Capgemini. But the fact that <em>DeepEval</em> simply runs, and does nothing with the data afterward, isn\u2019t the best experience. If you want to inspect failing test cases, identify regressions, or even pick the best model/prompt combination, you need more than just <em>DeepEval</em>. That\u2019s why we built a platform around it.<p>Here\u2019s a quick demo video of how everything works: <a href=\"https://youtu.be/PB3ngq7x4ko\" rel=\"nofollow\">https://youtu.be/PB3ngq7x4ko</a><p>Confident AI is great for RAG pipelines, agents, and chatbots. Typical use cases involve allowing companies to switch the underlying LLM, rewrite prompts for newer (and possibly cheaper) models, and keep test sets in sync with the codebase where <em>DeepEval</em> tests are run.<p>Our platform features a &quot;dataset editor,&quot; a &quot;regression catcher,&quot; and &quot;iteration insights&quot;. The datasets editor in Confident AI allows domain experts to edit datasets while keeping them in sync with your codebase for evaluation. We\u2019ll then generate sharable LLM testing/benchmark reports once <em>DeepEval</em> has finished running evaluations on these datasets that are pulled from the cloud. The regression catcher then identifies any regressions in your new implementation, and we use these evaluation results to determine the best iteration based on your metric scores.<p>Our goal is to make benchmarking LLM applications so reliable that picking the best implementation is as simple as reading the metric values off the dashboard. To achieve this, the quality of curated datasets and the accuracy and reliability of metrics must be the highest possible.<p>This brings us to our current limitations. Right now, <em>DeepEval</em>\u2019s primary evaluation method is LLM-as-a-judge. We use techniques such as GEval and question-answer generation to improve reliability, but these methods can still be inconsistent. Even with high-quality datasets curated by domain experts, our evaluation metrics remain the biggest blocker to our goal.<p>To address this, we recently released a DAG (Directed Acyclic Graph) metric in <em>DeepEval</em>. It is a decision-tree-based, LLM-as-a-judge metric that provides deterministic results by breaking a test case into finer atomic units. Each edge represents a decision, each node represents an LLM evaluation step, and each leaf node returns a score. It works best in scenarios where success criteria are clearly defined, such as text summarization.<p>The DAG metric is still in its early stages, but our hope is that by moving towards better, code-driven, open-source metrics, Confident AI can deliver deterministic LLM benchmarks that anyone can blindly trust.<p>We hope you\u2019ll give Confident AI a try. Quickstart here: <a href=\"https://docs.confident-ai.com/confident-ai/confident-ai-introduction\">https://docs.confident-ai.com/confident-ai/confident-ai-intr...</a><p>The platform runs on a freemium tier, and we've dropped the need to signup with a work email for the next four days.<p>Looking forward to your thoughts!"},"title":{"matchLevel":"none","matchedWords":[],"value":"Launch HN: Confident AI (YC W25) \u2013 Open-source evaluation framework for LLM apps"}},"_tags":["story","author_jeffreyip","story_43116633","launch_hn"],"author":"jeffreyip","children":[43118409,43118881,43118969,43119072,43119323,43120211,43121129,43121178,43124903,43126359,43131559,43172556],"created_at":"2025-02-20T16:23:56Z","created_at_i":1740068636,"num_comments":27,"objectID":"43116633","points":117,"story_id":43116633,"story_text":"Hi HN - we&#x27;re Jeffrey and Kritin, and we&#x27;re building Confident AI (<a href=\"https:&#x2F;&#x2F;confident-ai.com\">https:&#x2F;&#x2F;confident-ai.com</a>). This is the cloud platform for DeepEval (<a href=\"https:&#x2F;&#x2F;github.com&#x2F;confident-ai&#x2F;deepeval\">https:&#x2F;&#x2F;github.com&#x2F;confident-ai&#x2F;deepeval</a>), our open-source package that helps engineers evaluate and unit-test LLM applications. Think Pytest for LLMs.<p>We spent the past year building DeepEval with the goal of providing the best LLM evaluation developer experience, growing it to run over 600K evaluations daily in CI&#x2F;CD pipelines of enterprises like BCG, AstraZeneca, AXA, and Capgemini. But the fact that DeepEval simply runs, and does nothing with the data afterward, isn\u2019t the best experience. If you want to inspect failing test cases, identify regressions, or even pick the best model&#x2F;prompt combination, you need more than just DeepEval. That\u2019s why we built a platform around it.<p>Here\u2019s a quick demo video of how everything works: <a href=\"https:&#x2F;&#x2F;youtu.be&#x2F;PB3ngq7x4ko\" rel=\"nofollow\">https:&#x2F;&#x2F;youtu.be&#x2F;PB3ngq7x4ko</a><p>Confident AI is great for RAG pipelines, agents, and chatbots. Typical use cases involve allowing companies to switch the underlying LLM, rewrite prompts for newer (and possibly cheaper) models, and keep test sets in sync with the codebase where DeepEval tests are run.<p>Our platform features a &quot;dataset editor,&quot; a &quot;regression catcher,&quot; and &quot;iteration insights&quot;. The datasets editor in Confident AI allows domain experts to edit datasets while keeping them in sync with your codebase for evaluation. We\u2019ll then generate sharable LLM testing&#x2F;benchmark reports once DeepEval has finished running evaluations on these datasets that are pulled from the cloud. The regression catcher then identifies any regressions in your new implementation, and we use these evaluation results to determine the best iteration based on your metric scores.<p>Our goal is to make benchmarking LLM applications so reliable that picking the best implementation is as simple as reading the metric values off the dashboard. To achieve this, the quality of curated datasets and the accuracy and reliability of metrics must be the highest possible.<p>This brings us to our current limitations. Right now, DeepEval\u2019s primary evaluation method is LLM-as-a-judge. We use techniques such as GEval and question-answer generation to improve reliability, but these methods can still be inconsistent. Even with high-quality datasets curated by domain experts, our evaluation metrics remain the biggest blocker to our goal.<p>To address this, we recently released a DAG (Directed Acyclic Graph) metric in DeepEval. It is a decision-tree-based, LLM-as-a-judge metric that provides deterministic results by breaking a test case into finer atomic units. Each edge represents a decision, each node represents an LLM evaluation step, and each leaf node returns a score. It works best in scenarios where success criteria are clearly defined, such as text summarization.<p>The DAG metric is still in its early stages, but our hope is that by moving towards better, code-driven, open-source metrics, Confident AI can deliver deterministic LLM benchmarks that anyone can blindly trust.<p>We hope you\u2019ll give Confident AI a try. Quickstart here: <a href=\"https:&#x2F;&#x2F;docs.confident-ai.com&#x2F;confident-ai&#x2F;confident-ai-introduction\">https:&#x2F;&#x2F;docs.confident-ai.com&#x2F;confident-ai&#x2F;confident-ai-intr...</a><p>The platform runs on a freemium tier, and we&#x27;ve dropped the need to signup with a work email for the next four days.<p>Looking forward to your thoughts!","title":"Launch HN: Confident AI (YC W25) \u2013 Open-source evaluation framework for LLM apps","updated_at":"2026-06-14T02:35:39Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"sidmurali23"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"Hi HN, I\u2019m part of the Confident AI team and we\u2019re excited to share DeepTeam, an open-source framework that makes it trivial to penetration-test your LLM applications for 40+ security and safety risks. It has gained 400 on GitHub over the last month, and we\u2019d love your feedback!<p>Quick Introduction<p>- Detect vulnerabilities such as bias, misinformation, PII leakage, over-reliance on context, harmful content, and more\n- Simulate adversarial attacks with 10+ methods (jailbreaks, prompt injection, ROT13, automated evasion, data extraction, etc.)\n- Customize assessments to OWASP Top 10 for LLMs, NIST AI Risk Management, or your own security guidelines\n- Leverage <em>DeepEval</em> under the hood for robust metric evaluation, so you can run both regular and adversarial tests<p>Getting Started<p><pre><code>    # Install DeepTeam\n    pip install -U deepteam\n\n    # Clone the repo and run the example\n    git clone https://github.com/confident-ai/deepteam.git\n    cd deepteam\n    python3 -m venv venv &amp;&amp; source venv/bin/activate\n    python examples/red_teaming_example.py\n</code></pre>\nIn seconds you\u2019ll see a pass/fail breakdown for each vulnerability along with detailed test-case output. You can convert the results to a pandas DataFrame or save them for downstream analysis.<p>Why DeepTeam?<p>Most LLM safety tooling focuses on known benchmarks\u2014DeepTeam dynamically simulates attacks at runtime, so you catch novel, real-world threats and can track improvements over time via reusable attack suites.<p>We\u2019d love to hear:<p>- Which vulnerabilities you worry about most\n- How you integrate red-teaming into your CI/CD pipelines\n- Feature requests, contributions, and your toughest jailbreak stories<p>&lt;<a href=\"https://github.com/confident-ai/deepteam\">https://github.com/confident-ai/deepteam</a>&gt;"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: DeepTeam \u2013 Open-Source Red-Teaming Framework for LLM Security"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/confident-ai/deepteam"}},"_tags":["story","author_sidmurali23","story_44306963","show_hn"],"author":"sidmurali23","created_at":"2025-06-18T05:45:39Z","created_at_i":1750225539,"num_comments":0,"objectID":"44306963","points":4,"story_id":44306963,"story_text":"Hi HN, I\u2019m part of the Confident AI team and we\u2019re excited to share DeepTeam, an open-source framework that makes it trivial to penetration-test your LLM applications for 40+ security and safety risks. It has gained 400 on GitHub over the last month, and we\u2019d love your feedback!<p>Quick Introduction<p>- Detect vulnerabilities such as bias, misinformation, PII leakage, over-reliance on context, harmful content, and more\n- Simulate adversarial attacks with 10+ methods (jailbreaks, prompt injection, ROT13, automated evasion, data extraction, etc.)\n- Customize assessments to OWASP Top 10 for LLMs, NIST AI Risk Management, or your own security guidelines\n- Leverage DeepEval under the hood for robust metric evaluation, so you can run both regular and adversarial tests<p>Getting Started<p><pre><code>    # Install DeepTeam\n    pip install -U deepteam\n\n    # Clone the repo and run the example\n    git clone https:&#x2F;&#x2F;github.com&#x2F;confident-ai&#x2F;deepteam.git\n    cd deepteam\n    python3 -m venv venv &amp;&amp; source venv&#x2F;bin&#x2F;activate\n    python examples&#x2F;red_teaming_example.py\n</code></pre>\nIn seconds you\u2019ll see a pass&#x2F;fail breakdown for each vulnerability along with detailed test-case output. You can convert the results to a pandas DataFrame or save them for downstream analysis.<p>Why DeepTeam?<p>Most LLM safety tooling focuses on known benchmarks\u2014DeepTeam dynamically simulates attacks at runtime, so you catch novel, real-world threats and can track improvements over time via reusable attack suites.<p>We\u2019d love to hear:<p>- Which vulnerabilities you worry about most\n- How you integrate red-teaming into your CI&#x2F;CD pipelines\n- Feature requests, contributions, and your toughest jailbreak stories<p>&lt;<a href=\"https:&#x2F;&#x2F;github.com&#x2F;confident-ai&#x2F;deepteam\">https:&#x2F;&#x2F;github.com&#x2F;confident-ai&#x2F;deepteam</a>&gt;","title":"Show HN: DeepTeam \u2013 Open-Source Red-Teaming Framework for LLM Security","updated_at":"2025-06-18T08:03:41Z","url":"https://github.com/confident-ai/deepteam"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"nicolaib"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"Hi HN, I'm Nicolai. I'm working with a small team in Germany on Rhesis, an open-source platform for testing conversational LLM applications and agents. We\u2019re sharing an early community preview today.<p>Why we built this:\nWe saw teams repeatedly struggle with testing, e.g. scattered test cases, unclear or inconsistent metrics, and a lot of manual effort that still missed obvious failures before production. Most tools assume a single developer runs evals alone; in practice, testing tends to involve PMs, domain experts, QA, and engineers. We built Rhesis to make that collaboration straightforward.<p>What it does:\nRhesis is a self-hostable platform (with UI) where teams can create, run, and review tests for conversational AI systems.<p>A few core ideas:<p>- Test generation: Create and run tests for single-turns or full conversations; the platform can also assist with generating both single- and multi-turn scenarios using your domain context.<p>- Domain context / knowledge: Provide background material to guide test creation so you\u2019re not starting from an empty prompt.<p>- Collaboration tools: Non-technical teammates can write test cases, leave comments, and review results; developers can dig into failures with detailed traces and outputs.<p>- Unified metrics: Bring in eval metrics from <em>DeepEval</em>, RAGAS, and similar OSS frameworks without re-implementing them.<p>Current state:\nStill early. We shipped v0.4.2 last week with a zero-config Docker setup. Core flows work, but there are rough edges. Everything is MIT-licensed; an enterprise edition will come later, but the OSS core will remain free. We\u2019re currently focused on conversational applications because that\u2019s where we saw the biggest pain in evaluation and QA workflows.<p>Links:\nApp: app.rhesis.ai<p>GitHub: github.com/rhesis-ai/rhesis<p>Docs: docs.rhesis.ai<p>Happy to hear your thoughts and any answer questions about platform design, the architecture, or our thinking on collaborative testing workflows."},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Rhesis \u2013 Open-source platform for collaborative LLM application testing"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/rhesis-ai/rhesis"}},"_tags":["story","author_nicolaib","story_45970388","show_hn"],"author":"nicolaib","children":[45977066],"created_at":"2025-11-18T18:54:21Z","created_at_i":1763492061,"num_comments":2,"objectID":"45970388","points":3,"story_id":45970388,"story_text":"Hi HN, I&#x27;m Nicolai. I&#x27;m working with a small team in Germany on Rhesis, an open-source platform for testing conversational LLM applications and agents. We\u2019re sharing an early community preview today.<p>Why we built this:\nWe saw teams repeatedly struggle with testing, e.g. scattered test cases, unclear or inconsistent metrics, and a lot of manual effort that still missed obvious failures before production. Most tools assume a single developer runs evals alone; in practice, testing tends to involve PMs, domain experts, QA, and engineers. We built Rhesis to make that collaboration straightforward.<p>What it does:\nRhesis is a self-hostable platform (with UI) where teams can create, run, and review tests for conversational AI systems.<p>A few core ideas:<p>- Test generation: Create and run tests for single-turns or full conversations; the platform can also assist with generating both single- and multi-turn scenarios using your domain context.<p>- Domain context &#x2F; knowledge: Provide background material to guide test creation so you\u2019re not starting from an empty prompt.<p>- Collaboration tools: Non-technical teammates can write test cases, leave comments, and review results; developers can dig into failures with detailed traces and outputs.<p>- Unified metrics: Bring in eval metrics from DeepEval, RAGAS, and similar OSS frameworks without re-implementing them.<p>Current state:\nStill early. We shipped v0.4.2 last week with a zero-config Docker setup. Core flows work, but there are rough edges. Everything is MIT-licensed; an enterprise edition will come later, but the OSS core will remain free. We\u2019re currently focused on conversational applications because that\u2019s where we saw the biggest pain in evaluation and QA workflows.<p>Links:\nApp: app.rhesis.ai<p>GitHub: github.com&#x2F;rhesis-ai&#x2F;rhesis<p>Docs: docs.rhesis.ai<p>Happy to hear your thoughts and any answer questions about platform design, the architecture, or our thinking on collaborative testing workflows.","title":"Show HN: Rhesis \u2013 Open-source platform for collaborative LLM application testing","updated_at":"2026-03-05T23:02:21Z","url":"https://github.com/rhesis-ai/rhesis"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"jeffreyip"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"Hi HN, we\u2019re Jeffrey and Kritin, and we\u2019re building DeepTeam (<a href=\"https://trydeepteam.com\" rel=\"nofollow\">https://trydeepteam.com</a>), an open-source Python library to scan LLM apps for security vulnerabilities. You can start \u201cpenetration testing\u201d by defining a Python callback to your LLM app (e.g. `def model_callback(input: str)`), and DeepTeam will attempt to probe it with prompts designed to elicit unsafe or unintended behavior.<p>Note that the penetration testing process treats your LLM app as a black-box - which means that DeepTeam will not know whether PII leakage has occurred in a certain tool call or incorporated in the training data of your fine-tuned LLM, but rather just detect that it is present. Internally, we call this process \u201cend-to-end\u201d testing.<p>Before DeepTeam, we worked on <em>DeepEval</em>, an open-source framework to unit-test LLMs. Some of you might be thinking, well isn\u2019t this kind of similar to unit-testing?<p>Sort of, but not really. While LLM unit-testing focuses on 1) accurate eval metrics, 2) comprehensive eval datasets, penetration testing focuses on the haphazard simulation of attacks, and the orchestration of it. To users, this was a big and confusing paradigm shift, because it went from  \u201cDid this pass?\u201d to \u201cHow can this break?\u201d.<p>So we thought to ourselves, why not just release a new package to orchestrate the simulation of adversarial attacks for this new set of users and teams working specifically on AI safety, and borrow <em>DeepEval</em>\u2019s evals and ecosystem in the process?<p>Quickstart here: <a href=\"https://www.trydeepteam.com/docs/getting-started#detect-your-first-llm-vulnerability\" rel=\"nofollow\">https://www.trydeepteam.com/docs/getting-started#detect-your...</a><p>The first thing we did was offer as many attack methods as possible - simple encoding ones like ROT13, leetspeak, to prompt injections, roleplay, and jailbreaking. We then heard folks weren\u2019t happy because the attacks didn\u2019t persist across tests and hence they \u201clost\u201d their progress every time they tested, and so we added an option to `reuse_simulated_attacks`.<p>We abstracted everything away to make it as modular as possible - every vulnerability, attack, can be imported in Python as `Bias(type=[\u201crace\u201d])`, `LinearJailbreaking()`, etc. with methods such as `.enhance()` for teams to plug-and-play, build their own test suite, and even to add a few more rounds of attack enhancements to increase the likelihood of breaking your system.<p>Notably, there are a few limitations. Users might run into compliance errors when attempting to simulate attacks (especially for AzureOpenAI), and so we recommend setting `ignore_errors` to `True` in case that happens. You might also run into bottlenecks where DeepTeam does not cover your custom vulnerability type, and so we shipped a `CustomVulnerability` class as a \u201ccatch-all\u201d solution (still in beta).<p>You might be aware that some packages already exist that do a similar thing, often known as \u201cvulnerability scanning\u201d or \u201cred teaming\u201d. The difference is that DeepTeam is modular, lightweight, and code friendly. Take Nvidia Garak for example, although comprehensive, has so many CLI rules, environments to set up, it is definitely not the easiest to get started, let alone pick the library apart to build your own penetration testing pipeline. In DeepTeam, define a class, wrap it around your own implementations if necessary, and you\u2019re good to go.<p>We adopted a Apache 2.0 license (for now, and probably in the foreseeable future too), so if you want to get started, `pip install deepteam`, use any LLM for simulation, and you\u2019ll get a full penetration report within 1 minute (assuming you\u2019re running things asynchronously). GitHub: <a href=\"https://github.com/confident-ai/deepteam\">https://github.com/confident-ai/deepteam</a><p>Excited to share DeepTeam with everyone here \u2013 let us know what you think!"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: DeepTeam \u2013 Penetration Testing for LLMs"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/confident-ai/deepteam"}},"_tags":["story","author_jeffreyip","story_44117323","show_hn"],"author":"jeffreyip","created_at":"2025-05-28T15:49:43Z","created_at_i":1748447383,"num_comments":0,"objectID":"44117323","points":3,"story_id":44117323,"story_text":"Hi HN, we\u2019re Jeffrey and Kritin, and we\u2019re building DeepTeam (<a href=\"https:&#x2F;&#x2F;trydeepteam.com\" rel=\"nofollow\">https:&#x2F;&#x2F;trydeepteam.com</a>), an open-source Python library to scan LLM apps for security vulnerabilities. You can start \u201cpenetration testing\u201d by defining a Python callback to your LLM app (e.g. `def model_callback(input: str)`), and DeepTeam will attempt to probe it with prompts designed to elicit unsafe or unintended behavior.<p>Note that the penetration testing process treats your LLM app as a black-box - which means that DeepTeam will not know whether PII leakage has occurred in a certain tool call or incorporated in the training data of your fine-tuned LLM, but rather just detect that it is present. Internally, we call this process \u201cend-to-end\u201d testing.<p>Before DeepTeam, we worked on DeepEval, an open-source framework to unit-test LLMs. Some of you might be thinking, well isn\u2019t this kind of similar to unit-testing?<p>Sort of, but not really. While LLM unit-testing focuses on 1) accurate eval metrics, 2) comprehensive eval datasets, penetration testing focuses on the haphazard simulation of attacks, and the orchestration of it. To users, this was a big and confusing paradigm shift, because it went from  \u201cDid this pass?\u201d to \u201cHow can this break?\u201d.<p>So we thought to ourselves, why not just release a new package to orchestrate the simulation of adversarial attacks for this new set of users and teams working specifically on AI safety, and borrow DeepEval\u2019s evals and ecosystem in the process?<p>Quickstart here: <a href=\"https:&#x2F;&#x2F;www.trydeepteam.com&#x2F;docs&#x2F;getting-started#detect-your-first-llm-vulnerability\" rel=\"nofollow\">https:&#x2F;&#x2F;www.trydeepteam.com&#x2F;docs&#x2F;getting-started#detect-your...</a><p>The first thing we did was offer as many attack methods as possible - simple encoding ones like ROT13, leetspeak, to prompt injections, roleplay, and jailbreaking. We then heard folks weren\u2019t happy because the attacks didn\u2019t persist across tests and hence they \u201clost\u201d their progress every time they tested, and so we added an option to `reuse_simulated_attacks`.<p>We abstracted everything away to make it as modular as possible - every vulnerability, attack, can be imported in Python as `Bias(type=[\u201crace\u201d])`, `LinearJailbreaking()`, etc. with methods such as `.enhance()` for teams to plug-and-play, build their own test suite, and even to add a few more rounds of attack enhancements to increase the likelihood of breaking your system.<p>Notably, there are a few limitations. Users might run into compliance errors when attempting to simulate attacks (especially for AzureOpenAI), and so we recommend setting `ignore_errors` to `True` in case that happens. You might also run into bottlenecks where DeepTeam does not cover your custom vulnerability type, and so we shipped a `CustomVulnerability` class as a \u201ccatch-all\u201d solution (still in beta).<p>You might be aware that some packages already exist that do a similar thing, often known as \u201cvulnerability scanning\u201d or \u201cred teaming\u201d. The difference is that DeepTeam is modular, lightweight, and code friendly. Take Nvidia Garak for example, although comprehensive, has so many CLI rules, environments to set up, it is definitely not the easiest to get started, let alone pick the library apart to build your own penetration testing pipeline. In DeepTeam, define a class, wrap it around your own implementations if necessary, and you\u2019re good to go.<p>We adopted a Apache 2.0 license (for now, and probably in the foreseeable future too), so if you want to get started, `pip install deepteam`, use any LLM for simulation, and you\u2019ll get a full penetration report within 1 minute (assuming you\u2019re running things asynchronously). GitHub: <a href=\"https:&#x2F;&#x2F;github.com&#x2F;confident-ai&#x2F;deepteam\">https:&#x2F;&#x2F;github.com&#x2F;confident-ai&#x2F;deepteam</a><p>Excited to share DeepTeam with everyone here \u2013 let us know what you think!","title":"Show HN: DeepTeam \u2013 Penetration Testing for LLMs","updated_at":"2025-06-02T02:41:09Z","url":"https://github.com/confident-ai/deepteam"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"jeffreyip"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"Hi HN, we\u2019re Jeffrey and Kritin, and we\u2019re building DeepTeam (https://github.com/confident-ai/deepteam), an open-source Python library to scan LLM apps for security vulnerabilities. You can start \u201cpenetration testing\u201d by defining a Python callback to your LLM app (e.g. `def model_callback(input: str)`), and DeepTeam will attempt to probe it with prompts designed to elicit unsafe or unintended behavior.<p>Note that the penetration testing process treats your LLM app as a black-box - which means that DeepTeam will not know whether PII leakage has occurred in a certain tool call or incorporated in the training data of your fine-tuned LLM, but rather just detect that it is present. Internally, we call this process \u201cend-to-end\u201d testing.<p>Before DeepTeam, we worked on <em>DeepEval</em>, an open-source framework to unit-test LLMs. Some of you might be thinking, well isn\u2019t this kind of similar to unit-testing?<p>Sort of, but not really. While LLM unit-testing focuses on 1) accurate eval metrics, 2) comprehensive eval datasets, penetration testing focuses on the haphazard simulation of attacks, and the orchestration of it. To users, this was a big and confusing paradigm shift, because it went from \u201cDid this pass?\u201d to \u201cHow can this break?\u201d.<p>So we thought to ourselves, why not just release a new package to orchestrate the simulation of adversarial attacks for this new set of users and teams working specifically on AI safety, and borrow <em>DeepEval</em>\u2019s evals and ecosystem in the process?<p>Quickstart here: https://www.trydeepteam.com/docs/getting-started#detect-your-first-llm-vulnerability<p>The first thing we did was offer as many attack methods as possible - simple encoding ones like ROT13, leetspeak, to prompt injections, roleplay, and jailbreaking. We then heard folks weren\u2019t happy because the attacks didn\u2019t persist across tests and hence they \u201clost\u201d their progress every time they tested, and so we added an option to `reuse_simulated_attacks`.<p>We abstracted everything away to make it as modular as possible - every vulnerability, attack, can be imported in Python as `Bias(type=[\u201crace\u201d])`, `LinearJailbreaking()`, etc. with methods such as `.enhance()` for teams to plug-and-play, build their own test suite, and even to add a few more rounds of attack enhancements to increase the likelihood of breaking your system.<p>Notably, there are a few limitations. Users might run into compliance errors when attempting to simulate attacks (especially for AzureOpenAI), and so we recommend setting `ignore_errors` to `True` in case that happens. You might also run into bottlenecks where DeepTeam does not cover your custom vulnerability type, and so we shipped a `CustomVulnerability` class as a \u201ccatch-all\u201d solution (still in beta).<p>You might be aware that some packages already exist that do a similar thing, often known as \u201cvulnerability scanning\u201d or \u201cred teaming\u201d. The difference is that DeepTeam is modular, lightweight, and code friendly. Take Nvidia Garak for example, although comprehensive, has so many CLI rules, environments to set up, it is definitely not the easiest to get started, let alone pick the library apart to build your own penetration testing pipeline. In DeepTeam, define a class, wrap it around your own implementations if necessary, and you\u2019re good to go.<p>We adopted a Apache 2.0 license (for now, and probably in the foreseeable future too), so if you want to get started, `pip install deepteam`, use any LLM for simulation, and you\u2019ll get a full penetration report within 1 minute (assuming you\u2019re running things asynchronously). GitHub: https://github.com/confident-ai/deepteam<p>Excited to share DeepTeam with everyone here \u2013 let us know what you think!"},"title":{"matchLevel":"none","matchedWords":[],"value":"DeepTeam: Penetration Testing for LLMs"}},"_tags":["story","author_jeffreyip","story_44128270","ask_hn"],"author":"jeffreyip","created_at":"2025-05-29T17:35:12Z","created_at_i":1748540112,"num_comments":0,"objectID":"44128270","points":2,"story_id":44128270,"story_text":"Hi HN, we\u2019re Jeffrey and Kritin, and we\u2019re building DeepTeam (https:&#x2F;&#x2F;github.com&#x2F;confident-ai&#x2F;deepteam), an open-source Python library to scan LLM apps for security vulnerabilities. You can start \u201cpenetration testing\u201d by defining a Python callback to your LLM app (e.g. `def model_callback(input: str)`), and DeepTeam will attempt to probe it with prompts designed to elicit unsafe or unintended behavior.<p>Note that the penetration testing process treats your LLM app as a black-box - which means that DeepTeam will not know whether PII leakage has occurred in a certain tool call or incorporated in the training data of your fine-tuned LLM, but rather just detect that it is present. Internally, we call this process \u201cend-to-end\u201d testing.<p>Before DeepTeam, we worked on DeepEval, an open-source framework to unit-test LLMs. Some of you might be thinking, well isn\u2019t this kind of similar to unit-testing?<p>Sort of, but not really. While LLM unit-testing focuses on 1) accurate eval metrics, 2) comprehensive eval datasets, penetration testing focuses on the haphazard simulation of attacks, and the orchestration of it. To users, this was a big and confusing paradigm shift, because it went from \u201cDid this pass?\u201d to \u201cHow can this break?\u201d.<p>So we thought to ourselves, why not just release a new package to orchestrate the simulation of adversarial attacks for this new set of users and teams working specifically on AI safety, and borrow DeepEval\u2019s evals and ecosystem in the process?<p>Quickstart here: https:&#x2F;&#x2F;www.trydeepteam.com&#x2F;docs&#x2F;getting-started#detect-your-first-llm-vulnerability<p>The first thing we did was offer as many attack methods as possible - simple encoding ones like ROT13, leetspeak, to prompt injections, roleplay, and jailbreaking. We then heard folks weren\u2019t happy because the attacks didn\u2019t persist across tests and hence they \u201clost\u201d their progress every time they tested, and so we added an option to `reuse_simulated_attacks`.<p>We abstracted everything away to make it as modular as possible - every vulnerability, attack, can be imported in Python as `Bias(type=[\u201crace\u201d])`, `LinearJailbreaking()`, etc. with methods such as `.enhance()` for teams to plug-and-play, build their own test suite, and even to add a few more rounds of attack enhancements to increase the likelihood of breaking your system.<p>Notably, there are a few limitations. Users might run into compliance errors when attempting to simulate attacks (especially for AzureOpenAI), and so we recommend setting `ignore_errors` to `True` in case that happens. You might also run into bottlenecks where DeepTeam does not cover your custom vulnerability type, and so we shipped a `CustomVulnerability` class as a \u201ccatch-all\u201d solution (still in beta).<p>You might be aware that some packages already exist that do a similar thing, often known as \u201cvulnerability scanning\u201d or \u201cred teaming\u201d. The difference is that DeepTeam is modular, lightweight, and code friendly. Take Nvidia Garak for example, although comprehensive, has so many CLI rules, environments to set up, it is definitely not the easiest to get started, let alone pick the library apart to build your own penetration testing pipeline. In DeepTeam, define a class, wrap it around your own implementations if necessary, and you\u2019re good to go.<p>We adopted a Apache 2.0 license (for now, and probably in the foreseeable future too), so if you want to get started, `pip install deepteam`, use any LLM for simulation, and you\u2019ll get a full penetration report within 1 minute (assuming you\u2019re running things asynchronously). GitHub: https:&#x2F;&#x2F;github.com&#x2F;confident-ai&#x2F;deepteam<p>Excited to share DeepTeam with everyone here \u2013 let us know what you think!","title":"DeepTeam: Penetration Testing for LLMs","updated_at":"2025-05-29T23:18:42Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"dustfinger"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"I built a tag-driven changelog generator for an OSS project I maintain and used it to backfill our 2025 docs changelog.<p>It walks git tags, finds merged PRs between releases, buckets them into categories, and renders MDX (Docusaurus in our case) organized by year -&gt; month -&gt; category -&gt; version. There is an optional LLM mode that produces structured JSON via a Pydantic schema for PR entry and monthly summaries.<p>Example:<p>python .scripts/changelog/generate.py --year 2025 --github --ai --ai-model gpt-5.2\n(or --help)<p>Gotcha: if you use --github, set GITHUB_TOKEN or you will most likely hit GitHub rate limits.<p>Repo script: <a href=\"https://github.com/confident-ai/deepeval/blob/main/.scripts/changelog/generate.py\" rel=\"nofollow\">https://github.com/confident-ai/<em>deepeval</em>/blob/main/.scripts/...</a><p>Disclosure: I maintain <em>DeepEval</em>. Happy to answer questions and take feedback."},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Tag driven changelog generator (MDX) with optional LLM summaries"}},"_tags":["story","author_dustfinger","story_46569948","show_hn"],"author":"dustfinger","children":[46578861],"created_at":"2026-01-10T21:10:51Z","created_at_i":1768079451,"num_comments":1,"objectID":"46569948","points":1,"story_id":46569948,"story_text":"I built a tag-driven changelog generator for an OSS project I maintain and used it to backfill our 2025 docs changelog.<p>It walks git tags, finds merged PRs between releases, buckets them into categories, and renders MDX (Docusaurus in our case) organized by year -&gt; month -&gt; category -&gt; version. There is an optional LLM mode that produces structured JSON via a Pydantic schema for PR entry and monthly summaries.<p>Example:<p>python .scripts&#x2F;changelog&#x2F;generate.py --year 2025 --github --ai --ai-model gpt-5.2\n(or --help)<p>Gotcha: if you use --github, set GITHUB_TOKEN or you will most likely hit GitHub rate limits.<p>Repo script: <a href=\"https:&#x2F;&#x2F;github.com&#x2F;confident-ai&#x2F;deepeval&#x2F;blob&#x2F;main&#x2F;.scripts&#x2F;changelog&#x2F;generate.py\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;confident-ai&#x2F;deepeval&#x2F;blob&#x2F;main&#x2F;.scripts&#x2F;...</a><p>Disclosure: I maintain DeepEval. Happy to answer questions and take feedback.","title":"Show HN: Tag driven changelog generator (MDX) with optional LLM summaries","updated_at":"2026-03-05T23:20:53Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"jeffreyip"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["deepeval"],"value":"Hi HN, we\u2019re Jeffrey and Kritin, and we\u2019re building DeepTeam (https://github.com/confident-ai/deepteam), an open-source Python library to scan LLM apps for security vulnerabilities. You can start \u201cpenetration testing\u201d by defining a Python callback to your LLM app (e.g. `def model_callback(input: str)`), and DeepTeam will attempt to probe it with prompts designed to elicit unsafe or unintended behavior.\nNote that the penetration testing process treats your LLM app as a black-box - which means that DeepTeam will not know whether PII leakage has occurred in a certain tool call or incorporated in the training data of your fine-tuned LLM, but rather just detect that it is present. Internally, we call this process \u201cend-to-end\u201d testing.<p>Before DeepTeam, we worked on <em>DeepEval</em>, an open-source framework to unit-test LLMs. Some of you might be thinking, well isn\u2019t this kind of similar to unit-testing?<p>Sort of, but not really. While LLM unit-testing focuses on 1) accurate eval metrics, 2) comprehensive eval datasets, penetration testing focuses on the haphazard simulation of attacks, and the orchestration of it. To users, this was a big and confusing paradigm shift, because it went from \u201cDid this pass?\u201d to \u201cHow can this break?\u201d.<p>So we thought to ourselves, why not just release a new package to orchestrate the simulation of adversarial attacks for this new set of users and teams working specifically on AI safety, and borrow <em>DeepEval</em>\u2019s evals and ecosystem in the process?<p>Quickstart here: https://www.trydeepteam.com/docs/getting-started#detect-your-first-llm-vulnerability<p>The first thing we did was offer as many attack methods as possible - simple encoding ones like ROT13, leetspeak, to prompt injections, roleplay, and jailbreaking. We then heard folks weren\u2019t happy because the attacks didn\u2019t persist across tests and hence they \u201clost\u201d their progress every time they tested, and so we added an option to `reuse_simulated_attacks`.<p>We abstracted everything away to make it as modular as possible - every vulnerability, attack, can be imported in Python as `Bias(type=[\u201crace\u201d])`, `LinearJailbreaking()`, etc. with methods such as `.enhance()` for teams to plug-and-play, build their own test suite, and even to add a few more rounds of attack enhancements to increase the likelihood of breaking your system.<p>Notably, there are a few limitations. Users might run into compliance errors when attempting to simulate attacks (especially for AzureOpenAI), and so we recommend setting `ignore_errors` to `True` in case that happens. You might also run into bottlenecks where DeepTeam does not cover your custom vulnerability type, and so we shipped a `CustomVulnerability` class as a \u201ccatch-all\u201d solution (still in beta).<p>You might be aware that some packages already exist that do a similar thing, often known as \u201cvulnerability scanning\u201d or \u201cred teaming\u201d. The difference is that DeepTeam is modular, lightweight, and code friendly. Take Nvidia Garak for example, although comprehensive, has so many CLI rules, environments to set up, it is definitely not the easiest to get started, let alone pick the library apart to build your own penetration testing pipeline. In DeepTeam, define a class, wrap it around your own implementations if necessary, and you\u2019re good to go.<p>We adopted a Apache 2.0 license (for now, and probably in the foreseeable future too), so if you want to get started, `pip install deepteam`, use any LLM for simulation, and you\u2019ll get a full penetration report within 1 minute (assuming you\u2019re running things asynchronously). GitHub: https://github.com/confident-ai/deepteam<p>Excited to share DeepTeam with everyone here \u2013 let us know what you think!"},"title":{"matchLevel":"none","matchedWords":[],"value":"DeepTeam: Open-Source Pennetration Testing for LLMs"}},"_tags":["story","author_jeffreyip","story_44124610","ask_hn"],"author":"jeffreyip","created_at":"2025-05-29T10:21:59Z","created_at_i":1748514119,"num_comments":0,"objectID":"44124610","points":1,"story_id":44124610,"story_text":"Hi HN, we\u2019re Jeffrey and Kritin, and we\u2019re building DeepTeam (https:&#x2F;&#x2F;github.com&#x2F;confident-ai&#x2F;deepteam), an open-source Python library to scan LLM apps for security vulnerabilities. You can start \u201cpenetration testing\u201d by defining a Python callback to your LLM app (e.g. `def model_callback(input: str)`), and DeepTeam will attempt to probe it with prompts designed to elicit unsafe or unintended behavior.\nNote that the penetration testing process treats your LLM app as a black-box - which means that DeepTeam will not know whether PII leakage has occurred in a certain tool call or incorporated in the training data of your fine-tuned LLM, but rather just detect that it is present. Internally, we call this process \u201cend-to-end\u201d testing.<p>Before DeepTeam, we worked on DeepEval, an open-source framework to unit-test LLMs. Some of you might be thinking, well isn\u2019t this kind of similar to unit-testing?<p>Sort of, but not really. While LLM unit-testing focuses on 1) accurate eval metrics, 2) comprehensive eval datasets, penetration testing focuses on the haphazard simulation of attacks, and the orchestration of it. To users, this was a big and confusing paradigm shift, because it went from \u201cDid this pass?\u201d to \u201cHow can this break?\u201d.<p>So we thought to ourselves, why not just release a new package to orchestrate the simulation of adversarial attacks for this new set of users and teams working specifically on AI safety, and borrow DeepEval\u2019s evals and ecosystem in the process?<p>Quickstart here: https:&#x2F;&#x2F;www.trydeepteam.com&#x2F;docs&#x2F;getting-started#detect-your-first-llm-vulnerability<p>The first thing we did was offer as many attack methods as possible - simple encoding ones like ROT13, leetspeak, to prompt injections, roleplay, and jailbreaking. We then heard folks weren\u2019t happy because the attacks didn\u2019t persist across tests and hence they \u201clost\u201d their progress every time they tested, and so we added an option to `reuse_simulated_attacks`.<p>We abstracted everything away to make it as modular as possible - every vulnerability, attack, can be imported in Python as `Bias(type=[\u201crace\u201d])`, `LinearJailbreaking()`, etc. with methods such as `.enhance()` for teams to plug-and-play, build their own test suite, and even to add a few more rounds of attack enhancements to increase the likelihood of breaking your system.<p>Notably, there are a few limitations. Users might run into compliance errors when attempting to simulate attacks (especially for AzureOpenAI), and so we recommend setting `ignore_errors` to `True` in case that happens. You might also run into bottlenecks where DeepTeam does not cover your custom vulnerability type, and so we shipped a `CustomVulnerability` class as a \u201ccatch-all\u201d solution (still in beta).<p>You might be aware that some packages already exist that do a similar thing, often known as \u201cvulnerability scanning\u201d or \u201cred teaming\u201d. The difference is that DeepTeam is modular, lightweight, and code friendly. Take Nvidia Garak for example, although comprehensive, has so many CLI rules, environments to set up, it is definitely not the easiest to get started, let alone pick the library apart to build your own penetration testing pipeline. In DeepTeam, define a class, wrap it around your own implementations if necessary, and you\u2019re good to go.<p>We adopted a Apache 2.0 license (for now, and probably in the foreseeable future too), so if you want to get started, `pip install deepteam`, use any LLM for simulation, and you\u2019ll get a full penetration report within 1 minute (assuming you\u2019re running things asynchronously). GitHub: https:&#x2F;&#x2F;github.com&#x2F;confident-ai&#x2F;deepteam<p>Excited to share DeepTeam with everyone here \u2013 let us know what you think!","title":"DeepTeam: Open-Source Pennetration Testing for LLMs","updated_at":"2025-05-29T10:24:39Z"}],"hitsPerPage":20,"nbHits":341,"nbPages":18,"page":0,"params":"query=deepeval&advancedSyntax=true&analyticsTags=backend","processingTimeMS":9,"processingTimingsMS":{"_request":{"queue":100,"roundTrip":22},"afterFetch":{"format":{"highlighting":1,"total":1}},"fetch":{"query":7,"total":8},"total":9},"query":"deepeval","serverTimeMS":112}
