{"exhaustive":{"nbHits":true,"typo":true},"exhaustiveNbHits":true,"exhaustiveTypo":true,"hits":[{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"vardalab"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["tool","eval","bench"],"value":"<a href=\"https://github.com/SeraphimSerapis/tool-eval-bench\" rel=\"nofollow\">https://github.com/SeraphimSerapis/<em>tool-eval-bench</em></a>\nTrying to run this stuff really triggers it. Freaking frustrating.\nI have it set up some local inference and then I'm struggling to get the MTP working and it just refuses to work on evaluations."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://andonlabs.com/blog/fable5-vending-bench"}},"_tags":["comment","author_vardalab","story_48803762"],"author":"vardalab","comment_text":"<a href=\"https:&#x2F;&#x2F;github.com&#x2F;SeraphimSerapis&#x2F;tool-eval-bench\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;SeraphimSerapis&#x2F;tool-eval-bench</a>\nTrying to run this stuff really triggers it. Freaking frustrating.\nI have it set up some local inference and then I&#x27;m struggling to get the MTP working and it just refuses to work on evaluations.","created_at":"2026-07-07T01:42:30Z","created_at_i":1783388550,"objectID":"48812751","parent_id":48806201,"story_id":48803762,"story_title":"Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability","story_url":"https://andonlabs.com/blog/fable5-vending-bench","updated_at":"2026-07-07T01:47:14Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"redrove"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["tool","eval","bench"],"value":"&gt; My dream would be a local model that can do, say, 80% of the day to day tasks I need; &quot;how does X Handler connect to Y storage?&quot;, &quot;commit that feature, but leave out the bits that relate to billing&quot; etc.<p>Qwen 3.6 27B can do that today, but setup properly and in a good quant, I run an autoround [0] with weights in int8 and attention heads in f16 on a single RTX 6000 Pro Blackwell Max-Q via vllm with mtp=2 and full context, --max-num-seqs 3, KV in f16, mamba f32.<p>&gt;It would have 99% reliable tool calling<p>I managed to score 93/100 in <em>tool-eval-bench</em> [1]. For me this is very good already, at least in the pi coding harness I've never had an issue that wasn't auto-fixed in the next turn(s).<p>&gt;the ability to go &quot;this task is beyond my skills&quot; and refer to a Big Boy Online Model in a gigantic datacenter somewhere<p>This is heavy on the harness engineering side I think, but also quite contrary to the nature of LLMs today. If you figure this out I'd love to know.<p>[0] <a href=\"https://huggingface.co/Minachist/Qwen3.6-27B-INT8-AutoRound/tree/W8A16-GS128\" rel=\"nofollow\">https://huggingface.co/Minachist/Qwen3.6-27B-INT8-AutoRound/...</a><p>[1] <a href=\"https://github.com/SeraphimSerapis/tool-eval-bench\" rel=\"nofollow\">https://github.com/SeraphimSerapis/<em>tool-eval-bench</em></a>"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Local Qwen isn't a worse Opus, it's a different tool"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://blog.alexellis.io/local-ai-is-not-opus/"}},"_tags":["comment","author_redrove","story_48580209"],"author":"redrove","children":[48594632,48594821],"comment_text":"&gt; My dream would be a local model that can do, say, 80% of the day to day tasks I need; &quot;how does X Handler connect to Y storage?&quot;, &quot;commit that feature, but leave out the bits that relate to billing&quot; etc.<p>Qwen 3.6 27B can do that today, but setup properly and in a good quant, I run an autoround [0] with weights in int8 and attention heads in f16 on a single RTX 6000 Pro Blackwell Max-Q via vllm with mtp=2 and full context, --max-num-seqs 3, KV in f16, mamba f32.<p>&gt;It would have 99% reliable tool calling<p>I managed to score 93&#x2F;100 in tool-eval-bench [1]. For me this is very good already, at least in the pi coding harness I&#x27;ve never had an issue that wasn&#x27;t auto-fixed in the next turn(s).<p>&gt;the ability to go &quot;this task is beyond my skills&quot; and refer to a Big Boy Online Model in a gigantic datacenter somewhere<p>This is heavy on the harness engineering side I think, but also quite contrary to the nature of LLMs today. If you figure this out I&#x27;d love to know.<p>[0] <a href=\"https:&#x2F;&#x2F;huggingface.co&#x2F;Minachist&#x2F;Qwen3.6-27B-INT8-AutoRound&#x2F;tree&#x2F;W8A16-GS128\" rel=\"nofollow\">https:&#x2F;&#x2F;huggingface.co&#x2F;Minachist&#x2F;Qwen3.6-27B-INT8-AutoRound&#x2F;...</a><p>[1] <a href=\"https:&#x2F;&#x2F;github.com&#x2F;SeraphimSerapis&#x2F;tool-eval-bench\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;SeraphimSerapis&#x2F;tool-eval-bench</a>","created_at":"2026-06-18T07:50:03Z","created_at_i":1781769003,"objectID":"48582176","parent_id":48581623,"story_id":48580209,"story_title":"Local Qwen isn't a worse Opus, it's a different tool","story_url":"https://blog.alexellis.io/local-ai-is-not-opus/","updated_at":"2026-06-19T04:39:28Z"}],"hitsPerPage":20,"nbHits":2,"nbPages":1,"page":0,"params":"query=%22tool-eval-bench%22&advancedSyntax=true&analyticsTags=backend","processingTimeMS":5,"processingTimingsMS":{"_request":{"queue":477,"roundTrip":17},"fetch":{"total":1},"getIdx":{"load":{"gens":2,"total":3},"total":3},"total":5},"query":"\"tool-eval-bench\"","serverTimeMS":492}
