{"exhaustive":{"nbHits":false,"typo":false},"exhaustiveNbHits":false,"exhaustiveTypo":false,"hits":[{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"simonw"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["evals","faq","hamel"],"value":"This is <i>so hard</i>! I don't yet have a great solution for this myself, but I've been collecting notes about this on my &quot;<em>evals</em>&quot; tag for a while: <a href=\"https://simonwillison.net/tags/evals/\" rel=\"nofollow\">https://simonwillison.net/tags/<em>evals</em>/</a><p>The best writing I've seen about this is from <em>Hamel</em> Husain - <a href=\"https://hamel.dev/blog/posts/llm-judge/\" rel=\"nofollow\">https://<em>hamel</em>.dev/blog/posts/llm-judge/</a> and <a href=\"https://hamel.dev/blog/posts/evals-faq/\" rel=\"nofollow\">https://<em>hamel</em>.dev/blog/posts/<em>evals</em>-<em>faq</em>/</a> are both excellent."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Context engineering"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://chrisloy.dev/post/2025/08/03/context-engineering"}},"_tags":["comment","author_simonw","story_45788842"],"author":"simonw","comment_text":"This is <i>so hard</i>! I don&#x27;t yet have a great solution for this myself, but I&#x27;ve been collecting notes about this on my &quot;evals&quot; tag for a while: <a href=\"https:&#x2F;&#x2F;simonwillison.net&#x2F;tags&#x2F;evals&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;simonwillison.net&#x2F;tags&#x2F;evals&#x2F;</a><p>The best writing I&#x27;ve seen about this is from Hamel Husain - <a href=\"https:&#x2F;&#x2F;hamel.dev&#x2F;blog&#x2F;posts&#x2F;llm-judge&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;hamel.dev&#x2F;blog&#x2F;posts&#x2F;llm-judge&#x2F;</a> and <a href=\"https:&#x2F;&#x2F;hamel.dev&#x2F;blog&#x2F;posts&#x2F;evals-faq&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;hamel.dev&#x2F;blog&#x2F;posts&#x2F;evals-faq&#x2F;</a> are both excellent.","created_at":"2025-11-02T19:44:35Z","created_at_i":1762112675,"objectID":"45792830","parent_id":45792518,"story_id":45788842,"story_title":"Context engineering","story_url":"https://chrisloy.dev/post/2025/08/03/context-engineering","updated_at":"2026-03-05T23:00:37Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"TheIronYuppie"},"title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["evals"],"value":"About AI <em>Evals</em>"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["evals","faq","hamel"],"value":"https://<em>hamel</em>.dev/blog/posts/<em>evals</em>-<em>faq</em>/"}},"_tags":["story","author_TheIronYuppie","story_44430117"],"author":"TheIronYuppie","children":[44454094,44454273,44454889,44455227,44455971,44456935,44457032,44457557,44458524,44458811,44459047,44460878,44474563],"created_at":"2025-07-01T02:48:16Z","created_at_i":1751338096,"num_comments":43,"objectID":"44430117","points":189,"story_id":44430117,"title":"About AI Evals","updated_at":"2026-08-17T07:59:00Z","url":"https://hamel.dev/blog/posts/evals-faq/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"hamelsmu"},"title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["evals"],"value":"Frequently Asked Questions About AI <em>Evals</em>"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["evals","faq","hamel"],"value":"https://<em>hamel</em>.dev/blog/posts/<em>evals</em>-<em>faq</em>/"}},"_tags":["story","author_hamelsmu","story_44423661"],"author":"hamelsmu","children":[44423697,44423808],"created_at":"2025-06-30T14:21:43Z","created_at_i":1751293303,"num_comments":1,"objectID":"44423661","points":8,"story_id":44423661,"title":"Frequently Asked Questions About AI Evals","updated_at":"2025-07-01T23:53:49Z","url":"https://hamel.dev/blog/posts/evals-faq/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"tosh"},"title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["evals"],"value":"AI <em>Evals</em>: Everything You Need to Know"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["evals","faq","hamel"],"value":"https://<em>hamel</em>.dev/blog/posts/<em>evals</em>-<em>faq</em>/"}},"_tags":["story","author_tosh","story_49832346"],"author":"tosh","created_at":"2026-09-24T15:46:51Z","created_at_i":1790264811,"num_comments":0,"objectID":"49832346","points":2,"story_id":49832346,"title":"AI Evals: Everything You Need to Know","updated_at":"2026-09-24T16:10:24Z","url":"https://hamel.dev/blog/posts/evals-faq/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"pamelafox"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["evals","faq","hamel"],"value":"Fantastic <em>FAQ</em>, thank you <em>Hamel</em> for writing it up. We had an open space on AI <em>Evals</em> at Pycon this year, and had lots of discussion around similar questions. I only wrote down the questions, however:<p># Evaluation Metrics &amp; Methodology<p>* What metrics do you use (e.g., BERTScore, ROUGE, F1)? Are similarity metrics still useful?<p>* Do you use step-by-step evaluations or evaluate full responses?<p>* How do you evaluate VLM (vision-language model) summarization? Do you sample outputs or extract named entities?<p>* How do you approach offline (ground truth) vs. online evaluation?<p>* How do you handle uncertainty or &quot;don\u2019t know&quot; cases? (Temperature settings?)<p>* How do you evaluate multi-turn conversations?<p>* A/B comparisons and discrete labels (e.g., good/bad) are easier to interpret.<p>* It\u2019s important to counteract bias toward your own favorite eval questions\u2014ensure a diverse dataset.<p>## Prompting &amp; Models<p>* Do you modify prompts based on the specific app being evaluated?<p>* Where do you store prompts\u2014text files, Prompty, database, or in code?<p>* Do you have domain experts edit or review prompts?<p>* How do you choose which model to use?<p>## Evaluation Infrastructure<p>* How do you choose an evaluation framework?<p>* What platforms do you use to gather domain expert feedback or labels?<p>* Do domain experts label outputs or also help with prompt design?<p>## User Feedback &amp; Observability<p>* Do you collect thumbs up / thumbs down feedback?<p>* How does observability help identify failure modes?<p>* Do models tend to favor their own outputs? (There's research on this.)<p>I personally work on adding evaluation to our most popular Azure RAG samples, and put a Textual CLI interface in this repo that I've found helpful for reviewing the eval results:\n<a href=\"https://github.com/Azure-Samples/ai-rag-chat-evaluator\">https://github.com/Azure-Samples/ai-rag-chat-evaluator</a>"},"story_title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["evals"],"value":"About AI <em>Evals</em>"},"story_url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["evals","faq","hamel"],"value":"https://<em>hamel</em>.dev/blog/posts/<em>evals</em>-<em>faq</em>/"}},"_tags":["comment","author_pamelafox","story_44430117"],"author":"pamelafox","children":[44457577,44458600],"comment_text":"Fantastic FAQ, thank you Hamel for writing it up. We had an open space on AI Evals at Pycon this year, and had lots of discussion around similar questions. I only wrote down the questions, however:<p># Evaluation Metrics &amp; Methodology<p>* What metrics do you use (e.g., BERTScore, ROUGE, F1)? Are similarity metrics still useful?<p>* Do you use step-by-step evaluations or evaluate full responses?<p>* How do you evaluate VLM (vision-language model) summarization? Do you sample outputs or extract named entities?<p>* How do you approach offline (ground truth) vs. online evaluation?<p>* How do you handle uncertainty or &quot;don\u2019t know&quot; cases? (Temperature settings?)<p>* How do you evaluate multi-turn conversations?<p>* A&#x2F;B comparisons and discrete labels (e.g., good&#x2F;bad) are easier to interpret.<p>* It\u2019s important to counteract bias toward your own favorite eval questions\u2014ensure a diverse dataset.<p>## Prompting &amp; Models<p>* Do you modify prompts based on the specific app being evaluated?<p>* Where do you store prompts\u2014text files, Prompty, database, or in code?<p>* Do you have domain experts edit or review prompts?<p>* How do you choose which model to use?<p>## Evaluation Infrastructure<p>* How do you choose an evaluation framework?<p>* What platforms do you use to gather domain expert feedback or labels?<p>* Do domain experts label outputs or also help with prompt design?<p>## User Feedback &amp; Observability<p>* Do you collect thumbs up &#x2F; thumbs down feedback?<p>* How does observability help identify failure modes?<p>* Do models tend to favor their own outputs? (There&#x27;s research on this.)<p>I personally work on adding evaluation to our most popular Azure RAG samples, and put a Textual CLI interface in this repo that I&#x27;ve found helpful for reviewing the eval results:\n<a href=\"https:&#x2F;&#x2F;github.com&#x2F;Azure-Samples&#x2F;ai-rag-chat-evaluator\">https:&#x2F;&#x2F;github.com&#x2F;Azure-Samples&#x2F;ai-rag-chat-evaluator</a>","created_at":"2025-07-03T16:46:51Z","created_at_i":1751561211,"objectID":"44456935","parent_id":44430117,"story_id":44430117,"story_title":"About AI Evals","story_url":"https://hamel.dev/blog/posts/evals-faq/","updated_at":"2025-07-20T03:52:29Z"}],"hitsPerPage":20,"nbHits":5,"nbPages":1,"page":0,"params":"query=evals+faq+hamel&advancedSyntax=true&analyticsTags=backend","processingTimeMS":22,"processingTimingsMS":{"_request":{"queue":132,"roundTrip":15},"afterFetch":{"merge":{"total":1},"total":1},"fetch":{"query":13,"scanning":6,"total":20},"total":22},"query":"evals faq hamel","serverTimeMS":155}
