{"author":"pomarie","children":[{"author":"bumbledraven","children":[],"created_at":"2025-06-26T13:48:58.000Z","created_at_i":1750945738,"id":44387403,"options":[],"parent_id":44386887,"points":null,"story_id":44386887,"text":"What model were they using?","title":null,"type":"comment","url":null},{"author":"jangletown","children":[],"created_at":"2025-06-26T13:53:20.000Z","created_at_i":1750946000,"id":44387437,"options":[],"parent_id":44386887,"points":null,"story_id":44386887,"text":"&quot;51% fewer false positives&quot;, how were you measuring? is this an internal or benchmarking dataset?","title":null,"type":"comment","url":null},{"author":"N_Lens","children":[{"author":"flippyhead","children":[],"created_at":"2025-06-26T14:11:29.000Z","created_at_i":1750947089,"id":44387615,"options":[],"parent_id":44387467,"points":null,"story_id":44386887,"text":"I found it useful.","title":null,"type":"comment","url":null},{"author":"weego","children":[],"created_at":"2025-06-26T14:19:42.000Z","created_at_i":1750947582,"id":44387691,"options":[],"parent_id":44387467,"points":null,"story_id":44386887,"text":"It&#x27;s recreating the monolith vs micro-service argument by proxy for a new generation to plan conference talks around.","title":null,"type":"comment","url":null}],"created_at":"2025-06-26T13:55:56.000Z","created_at_i":1750946156,"id":44387467,"options":[],"parent_id":44386887,"points":null,"story_id":44386887,"text":"Very vague post light on details, and as usual, feels more like a marketing pitch for the website.","title":null,"type":"comment","url":null},{"author":"vinnymac","children":[],"created_at":"2025-06-26T13:58:01.000Z","created_at_i":1750946281,"id":44387491,"options":[],"parent_id":44386887,"points":null,"story_id":44386887,"text":"I\u2019ve been testing this for the last few months, and it is now much quieter than before, and even more useful.","title":null,"type":"comment","url":null},{"author":"kurtis_reed","children":[],"created_at":"2025-06-26T14:01:22.000Z","created_at_i":1750946482,"id":44387524,"options":[],"parent_id":44386887,"points":null,"story_id":44386887,"text":"There was a blog post from another AI code review tool: &quot;How to Make LLMs Shut Up&quot;<p><a href=\"https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=42451968\">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=42451968</a>","title":null,"type":"comment","url":null},{"author":"h1fra","children":[{"author":"bwfan123","children":[],"created_at":"2025-06-26T14:13:47.000Z","created_at_i":1750947227,"id":44387643,"options":[],"parent_id":44387583,"points":null,"story_id":44386887,"text":"code-reviews are not a good use-case for LLMs. here&#x27;s why: LLMs shine in usecases when their output is not evaluated on accuracy - for example, recommendations, semantic-search, sample snippets, images of people riding horses etc. code-reviews require accuracy.<p>What is a useful agent in the context of code-reviews in a large codebase is a semantic search agent which adds a comment containing related issues or PRs from the past for more context to human reviewers. This is a recommendation and is not rated on accuracy.","title":null,"type":"comment","url":null},{"author":"asdev","children":[{"author":"theonething","children":[],"created_at":"2025-06-28T02:58:53.000Z","created_at_i":1751079533,"id":44402040,"options":[],"parent_id":44388280,"points":null,"story_id":44386887,"text":"Isn&#x27;t it possible to feed that knowledge and context to it?  Have it scan your product website and docs, code documentation, git history, etc?","title":null,"type":"comment","url":null}],"created_at":"2025-06-26T15:21:06.000Z","created_at_i":1750951266,"id":44388280,"options":[],"parent_id":44387583,"points":null,"story_id":44386887,"text":"the code reviews can&#x27;t be effective because the LLM does not have the tribal knowledge and product context of the change. it&#x27;s just reading the code at face value","title":null,"type":"comment","url":null}],"created_at":"2025-06-26T14:08:08.000Z","created_at_i":1750946888,"id":44387583,"options":[],"parent_id":44386887,"points":null,"story_id":44386887,"text":"what I saw using 5-6 tools like this:<p>- PR description is never useful they barely summarize the file changes<p>- 90% of comments are wrong or irrelevant wether it&#x27;s because it&#x27;s missing context, missing tribal knowledge, missing code quality rules or wrongly interpret the code change<p>- 5-10% of the time it actually spots something<p>Not entirely sure it&#x27;s worth the noise","title":null,"type":"comment","url":null},{"author":"mosura","children":[{"author":"chanux","children":[{"author":"flippyhead","children":[],"created_at":"2025-06-26T14:12:45.000Z","created_at_i":1750947165,"id":44387631,"options":[],"parent_id":44387613,"points":null,"story_id":44386887,"text":"This is LITERALLY mind blowing.","title":null,"type":"comment","url":null},{"author":"criddell","children":[],"created_at":"2025-06-26T14:47:57.000Z","created_at_i":1750949277,"id":44387996,"options":[],"parent_id":44387613,"points":null,"story_id":44386887,"text":"I don&#x27;t like the word <i>learnings</i> either, but you write for your audience and this article was probably written with the hope that it would be shared on LinkedIn.<p><i>Learnings</i> might be the right choice here.<p>I wouldn&#x27;t complain if the HN headline mutator were to replace &quot;Learnings&quot; with &quot;lessons&quot;.","title":null,"type":"comment","url":null}],"created_at":"2025-06-26T14:11:27.000Z","created_at_i":1750947087,"id":44387613,"options":[],"parent_id":44387588,"points":null,"story_id":44386887,"text":"<a href=\"https:&#x2F;&#x2F;nolearnings.com&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;nolearnings.com&#x2F;</a>","title":null,"type":"comment","url":null}],"created_at":"2025-06-26T14:08:23.000Z","created_at_i":1750946903,"id":44387588,"options":[],"parent_id":44386887,"points":null,"story_id":44386887,"text":"Lessons.","title":null,"type":"comment","url":null},{"author":"curiousgal","children":[{"author":"nico","children":[{"author":"disgruntledphd2","children":[],"created_at":"2025-06-26T15:55:27.000Z","created_at_i":1750953327,"id":44388620,"options":[],"parent_id":44387840,"points":null,"story_id":44386887,"text":"Humans have a <i>lot</i> more introspection capabilities than any current LLM.","title":null,"type":"comment","url":null},{"author":"alganet","children":[],"created_at":"2025-06-26T16:12:04.000Z","created_at_i":1750954324,"id":44388768,"options":[],"parent_id":44387840,"points":null,"story_id":44386887,"text":"&gt; Several studies<p>Please, cite those studies. I want to read them.","title":null,"type":"comment","url":null}],"created_at":"2025-06-26T14:33:28.000Z","created_at_i":1750948408,"id":44387840,"options":[],"parent_id":44387666,"points":null,"story_id":44386887,"text":"Humans work pretty much the same way<p>Several studies have shown that we first make the decision and then we reason about it to justify it<p>In that sense, we are not much more rational than an LLM","title":null,"type":"comment","url":null},{"author":"elzbardico","children":[],"created_at":"2025-06-26T16:07:36.000Z","created_at_i":1750954056,"id":44388726,"options":[],"parent_id":44387666,"points":null,"story_id":44386887,"text":"The &quot;confidence&quot; field in the structured output was what really baffled me.","title":null,"type":"comment","url":null}],"created_at":"2025-06-26T14:16:40.000Z","created_at_i":1750947400,"id":44387666,"options":[],"parent_id":44386887,"points":null,"story_id":44386887,"text":"&gt; <i>Encouraged structured thinking by forcing the AI to justify its findings first, significantly reducing arbitrary conclusions.</i><p>Ah yes, because we know very well that the current generation of AI models reasons and draws conclusions based on logic and understanding... This is the true face palm.","title":null,"type":"comment","url":null},{"author":"nzach","children":[{"author":"exitb","children":[],"created_at":"2025-06-26T16:37:44.000Z","created_at_i":1750955864,"id":44388997,"options":[],"parent_id":44387704,"points":null,"story_id":44386887,"text":"This seems like really bad news for the \u201eAI will soon replace all software developers\u201d crowd.","title":null,"type":"comment","url":null}],"created_at":"2025-06-26T14:21:15.000Z","created_at_i":1750947675,"id":44387704,"options":[],"parent_id":44386887,"points":null,"story_id":44386887,"text":"I agree with the sentiment of this post. I my personal experience the usefulness of a LLM positively correlated with your ability to constrain the problem it should solve.<p>Prompts like &#x27;Update this regex to match this new pattern&#x27; generally give better results than &#x27;Fix this routing error in my server&#x27;.<p>Although this pattern seems true empirically,  I&#x27;ve never seen any hard data to confirm this property(?). And this post is interesting but seems like a missed opportunity to back this idea with some numbers.","title":null,"type":"comment","url":null},{"author":"singron","children":[{"author":"willsmith72","children":[],"created_at":"2025-06-26T14:41:32.000Z","created_at_i":1750948892,"id":44387929,"options":[],"parent_id":44387726,"points":null,"story_id":44386887,"text":"yep completely agreed, how can that be the best example they chose to use?<p>If I reviewed that PR, absolutely I&#x27;d question why you&#x27;re commenting that out. There better be a very good reason, or even a link to a ticket with a clear deadline of when it can be cleaned up&#x2F;reverted","title":null,"type":"comment","url":null}],"created_at":"2025-06-26T14:23:24.000Z","created_at_i":1750947804,"id":44387726,"options":[],"parent_id":44386887,"points":null,"story_id":44386887,"text":"I think they skipped over a non-obvious motivating example too fast. On first glance, commenting out your CI test suite would be very bad to sneak into a random PR, and that review note might be justified.<p>I could imagine the situation might actually be more nuanced (e.g. adding new tests and some of them are commented out), but there isn&#x27;t enough context to really determine that, and even in that case, it can be worth asking about commented out code in case the author left it that way by accident.<p>Aren&#x27;t there plenty of more obvious nitpicks to highlight? A great nitpick example would be one where the model will also ask to reverse the resolution. E.g.<p><pre><code>    final var items = List.copyOf(...);\n    &lt;-- Consider using an explicit type for the variable.\n\n    final List items = List.copyOf(...);\n    &lt;-- Consider using var to avoid redundant type name.\n</code></pre>\nThis is clearly aggravating since it will always make review comments.","title":null,"type":"comment","url":null},{"author":"mattas","children":[{"author":"s1mplicissimus","children":[{"author":"snapcaster","children":[],"created_at":"2025-06-26T14:55:19.000Z","created_at_i":1750949719,"id":44388062,"options":[],"parent_id":44387823,"points":null,"story_id":44386887,"text":"You&#x27;re saying alchemy is better than the scientific method?","title":null,"type":"comment","url":null}],"created_at":"2025-06-26T14:31:59.000Z","created_at_i":1750948319,"id":44387823,"options":[],"parent_id":44387757,"points":null,"story_id":44386887,"text":"Afaik alchemists had a more reliable method than ... whatever this state of affairs is ^^","title":null,"type":"comment","url":null},{"author":"AndrewKemendo","children":[{"author":"wrs","children":[{"author":"AndrewKemendo","children":[],"created_at":"2025-06-27T17:45:28.000Z","created_at_i":1751046328,"id":44398764,"options":[],"parent_id":44388838,"points":null,"story_id":44386887,"text":"All that means is that you verified the null hypothesis which should be that it doesn\u2019t work<p>If you create hypothesis tests that are not written in or specific enough then you\u2019re right you\u2019re not gonna be able to do science<p>Incidentally 99.9% of people I know have no instinct for how to actually do science or have rigor or focus to actually do it in a way that is usable","title":null,"type":"comment","url":null}],"created_at":"2025-06-26T16:19:23.000Z","created_at_i":1750954763,"id":44388838,"options":[],"parent_id":44387886,"points":null,"story_id":44386887,"text":"For one thing, what you learned can stop working when you switch to a new model, or just a newer version of the \u201csame\u201d model.","title":null,"type":"comment","url":null}],"created_at":"2025-06-26T14:38:30.000Z","created_at_i":1750948710,"id":44387886,"options":[],"parent_id":44387757,"points":null,"story_id":44386887,"text":"Otherwise known as science<p>1:Observation\n2:Hypothesis\n3:test\n4:GOTO:1<p>This is every thing ever built ever<p>What is the problem exactly?","title":null,"type":"comment","url":null},{"author":"neuronic","children":[],"created_at":"2025-06-26T19:10:15.000Z","created_at_i":1750965015,"id":44390361,"options":[],"parent_id":44387757,"points":null,"story_id":44386887,"text":"That&#x27;s because there is no intelligence or understanding involved. They are just trying to brute force a tool for a different purpose into their use case because marketing can&#x27;t stop overselling AI.","title":null,"type":"comment","url":null}],"created_at":"2025-06-26T14:26:01.000Z","created_at_i":1750947961,"id":44387757,"options":[],"parent_id":44386887,"points":null,"story_id":44386887,"text":"&quot;After extensive trial-and-error...&quot;<p>IMO, this is the difference between building deterministic software and non-deterministic software (like an AI agent). It often boils down to randomly making tweaks and evaluating the outcome of those tweaks.","title":null,"type":"comment","url":null},{"author":"nico","children":[],"created_at":"2025-06-26T14:26:59.000Z","created_at_i":1750948019,"id":44387772,"options":[],"parent_id":44386887,"points":null,"story_id":44386887,"text":"&gt; 2.3 Specialized Micro-Agents Over Generalized Rules\nInitially, our instinct was to continuously add more rules into a single large prompt to handle edge cases<p>This has been my experience as well. However, it seems like the platforms like Cursor&#x2F;Lovable&#x2F;v0&#x2F;et al are doing things differently<p>For example, this is Lovable\u2019s leaked system prompt, 1550 lines: <a href=\"https:&#x2F;&#x2F;github.com&#x2F;x1xhlol&#x2F;system-prompts-and-models-of-ai-tools&#x2F;blob&#x2F;main&#x2F;Lovable&#x2F;Prompt.txt\">https:&#x2F;&#x2F;github.com&#x2F;x1xhlol&#x2F;system-prompts-and-models-of-ai-t...</a><p>Is there a trick to making gigantic system prompts work well?","title":null,"type":"comment","url":null},{"author":"shenberg","children":[],"created_at":"2025-06-26T14:45:28.000Z","created_at_i":1750949128,"id":44387971,"options":[],"parent_id":44386887,"points":null,"story_id":44386887,"text":"When I read &quot;51% fewer false positives&quot; followed immediately by &quot;Median comments per pull request cut by half&quot; it makes me wonder how many true positives they find. That&#x27;s maybe unfair as my reference is automated tooling in the security world, where the true-positive&#x2F;false-positive ratio is so bad that a 50% reduction in false positives is a drop in the bucket","title":null,"type":"comment","url":null},{"author":"Oras","children":[{"author":"SparkyMcUnicorn","children":[{"author":"bjorgen","children":[{"author":"SparkyMcUnicorn","children":[],"created_at":"2025-06-27T19:57:42.000Z","created_at_i":1751054262,"id":44399771,"options":[],"parent_id":44390885,"points":null,"story_id":44386887,"text":"A (hopefully) clear and probably oversimplified example:<p>Query -&gt; Person Lookup -&gt; Result-&gt; Structured Output `{ firstName: &quot;&quot;, lastName: &quot;&quot; }`<p>When result doesn&#x27;t have relevant information, structured output will basically always output a name, whether it found the correct person or not, because it wants to output something, even if the fields are optional. With this example, prompting can help turn the names into &quot;Unknown&quot;, but the prompt usually ends up being excessive and&#x2F;or time consuming to get correct and fix edge cases for. Weaker models might struggle more on details or relevance with this prompt-only approach.<p>`{ found: boolean, missingFields: [], missingReason: &quot;&quot;, firstName, lastName }`<p>Including one or more of these additional text output properties has an almost magical affect sometimes, reducing the required prompting and hallucinated&#x2F;incorrect outputs.","title":null,"type":"comment","url":null}],"created_at":"2025-06-26T20:19:17.000Z","created_at_i":1750969157,"id":44390885,"options":[],"parent_id":44388797,"points":null,"story_id":44386887,"text":"Do you have an example of this in practice? I&#x27;m having a hard understanding this and have a very similar problem of the agent wanting to give a response on optional fields.","title":null,"type":"comment","url":null}],"created_at":"2025-06-26T16:14:09.000Z","created_at_i":1750954449,"id":44388797,"options":[],"parent_id":44388357,"points":null,"story_id":44386887,"text":"I&#x27;ve found that giving agents an &quot;opt out&quot; works pretty well.<p>For structured outputs, making fields optional isn&#x27;t usually enough. Providing an additional field for it to dump some output, along with a description for how&#x2F;when it should be used, covers several issues around this problem.<p>I&#x27;m not claiming this would solve the specific issues discussed in the post. Just a potentially helpful tip for others out there.","title":null,"type":"comment","url":null},{"author":"ffsm8","children":[],"created_at":"2025-06-26T16:34:56.000Z","created_at_i":1750955696,"id":44388975,"options":[],"parent_id":44388357,"points":null,"story_id":44386887,"text":"Likely because it&#x27;s temporary?<p>It takes less effort to re-enable if it&#x27;s just commented out and its more visible that there is something funky going on that <i>someone</i> should fix.<p>But yeah, even if it&#x27;s temporary, it really should have the rationale for commenting it out added... It takes like 5s and provides important context for reviewers and people looking through the file history in the future.","title":null,"type":"comment","url":null},{"author":"pancsta","children":[],"created_at":"2025-06-27T00:04:17.000Z","created_at_i":1750982657,"id":44392652,"options":[],"parent_id":44388357,"points":null,"story_id":44386887,"text":"By splitting prompts into smaller chunks you effectively get \u201cbias free\u201d opinions, especially when cross-checked. You can then turn them into local reasoning, which is different from \u201csending an email to the LLM\u201d which seems to be the case here. Remember, LLM is Rainman.","title":null,"type":"comment","url":null}],"created_at":"2025-06-26T15:28:30.000Z","created_at_i":1750951710,"id":44388357,"options":[],"parent_id":44386887,"points":null,"story_id":44386887,"text":"The problem is that, regardless of how you try to use &quot;micro-agents &quot; as a marketing term, LLMs are instructed to return a result.<p>They will always try to come up with something.<p>The example provided was a poor one. The comment from LLM was solid. Why would you comment out a step in the pipeline instead of just deleting it? I would comment the same in a PR.","title":null,"type":"comment","url":null},{"author":"elzbardico","children":[{"author":"sharkjacobs","children":[{"author":"ramity","children":[{"author":"bckr","children":[{"author":"baby","children":[{"author":"bckr","children":[],"created_at":"2025-06-27T16:06:20.000Z","created_at_i":1751040380,"id":44397886,"options":[],"parent_id":44391421,"points":null,"story_id":44386887,"text":"Thanks. I was talking about the confidence measure.","title":null,"type":"comment","url":null}],"created_at":"2025-06-26T21:11:04.000Z","created_at_i":1750972264,"id":44391421,"options":[],"parent_id":44390550,"points":null,"story_id":44386887,"text":"this trick is being used by many apps (including Github copilot reviews). The way I see it, is that if the agent has an eager-to-please problem, then you give it a way out","title":null,"type":"comment","url":null}],"created_at":"2025-06-26T19:38:51.000Z","created_at_i":1750966731,"id":44390550,"options":[],"parent_id":44389054,"points":null,"story_id":44386887,"text":"Is there research solid knowledge on this?","title":null,"type":"comment","url":null}],"created_at":"2025-06-26T16:44:51.000Z","created_at_i":1750956291,"id":44389054,"options":[],"parent_id":44388954,"points":null,"story_id":44386887,"text":"elzbardico is pointing out how the author is having the confidence value generated in the output of the response rather than it being the confidence of the output.","title":null,"type":"comment","url":null}],"created_at":"2025-06-26T16:33:21.000Z","created_at_i":1750955601,"id":44388954,"options":[],"parent_id":44388711,"points":null,"story_id":44386887,"text":"Do you mean that there is no correlation between confidence and false positives or other errors?","title":null,"type":"comment","url":null},{"author":"ramity","children":[],"created_at":"2025-06-26T16:34:01.000Z","created_at_i":1750955641,"id":44388965,"options":[],"parent_id":44388711,"points":null,"story_id":44386887,"text":"I too once fell into the trap of having an LLM generate a confidence value in a response. This is a very genuine concern to raise.","title":null,"type":"comment","url":null},{"author":"munificent","children":[{"author":"zengid","children":[{"author":"GardenLetter27","children":[{"author":"lgas","children":[],"created_at":"2025-06-26T20:32:15.000Z","created_at_i":1750969935,"id":44391024,"options":[],"parent_id":44390027,"points":null,"story_id":44386887,"text":"True in many situations in life.","title":null,"type":"comment","url":null}],"created_at":"2025-06-26T18:33:25.000Z","created_at_i":1750962805,"id":44390027,"options":[],"parent_id":44389649,"points":null,"story_id":44386887,"text":"Confidence is all you need.","title":null,"type":"comment","url":null}],"created_at":"2025-06-26T17:51:20.000Z","created_at_i":1750960280,"id":44389649,"options":[],"parent_id":44389225,"points":null,"story_id":44386887,"text":"confidence all the way down","title":null,"type":"comment","url":null},{"author":"paisawalla","children":[],"created_at":"2025-06-27T17:57:59.000Z","created_at_i":1751047079,"id":44398858,"options":[],"parent_id":44389225,"points":null,"story_id":44386887,"text":"Wasteful. `confidence`&#x27;s type should be Array&lt;number&gt;, wherein confidence[N] gives the Nth derivative confidence rating.","title":null,"type":"comment","url":null}],"created_at":"2025-06-26T17:05:54.000Z","created_at_i":1750957554,"id":44389225,"options":[],"parent_id":44388711,"points":null,"story_id":44386887,"text":"Easy fix, just have the LLM generate:<p><pre><code>    {\n      &quot;reasoning&quot;: &quot;`cfg` can be nil on line 42; dereferenced without check on line 47&quot;,\n      &quot;finding&quot;: &quot;Possible nil\u2011pointer dereference&quot;,\n      &quot;confidence&quot;: 0.81,\n      &quot;confidence_in_confidence_rating&quot;: 0.54,\n      &quot;confidence_in_confidence_rating_in_confidence_rating&quot;: 0.12,\n      &quot;confidence_in_confidence_rating_in_confidence_rating_in_confidence_rating&quot;: 0.98,\n      &#x2F;&#x2F; Etc...\n    }</code></pre>","title":null,"type":"comment","url":null},{"author":"volkk","children":[],"created_at":"2025-06-26T18:10:50.000Z","created_at_i":1750961450,"id":44389813,"options":[],"parent_id":44388711,"points":null,"story_id":44386887,"text":"i immediately noticed the same thing, but to be fair, we don&#x27;t know if it&#x27;s enriched by a separate service that checks the response and uses some heuristics to compute that value. If not, yeah, that is an entirely made up and useless value","title":null,"type":"comment","url":null},{"author":"MattSayar","children":[],"created_at":"2025-06-26T19:50:57.000Z","created_at_i":1750967457,"id":44390646,"options":[],"parent_id":44388711,"points":null,"story_id":44386887,"text":"Could you have a higher-order reasoning LLM generate a better confidence rating? That&#x27;s how eval frameworks generally work today","title":null,"type":"comment","url":null},{"author":"skipants","children":[],"created_at":"2025-06-26T20:23:44.000Z","created_at_i":1750969424,"id":44390925,"options":[],"parent_id":44388711,"points":null,"story_id":44386887,"text":"When I was younger and more into music, when I went to a concert I would often judge if a drummer was &quot;good&quot; based on if they were better than me or not. I knew enough about drumming to tell how good someone was at the different parts of having that skill but also knew enough to know that I was not even close to having what it took to be a professional drummer.<p>This is what I feel like with this blogpost. I&#x27;ve barely scratched the surface of the innards of LLMs but even I know it should be completely obvious to anyone that has a product built around it that these confidence levels are completely made up.<p>I&#x27;ve never heard or used cubic before today but that part of the blog post, along with the obvious LLM generated quality of it, gives a terrible first impression.","title":null,"type":"comment","url":null},{"author":"baby","children":[],"created_at":"2025-06-26T21:10:03.000Z","created_at_i":1750972203,"id":44391411,"options":[],"parent_id":44388711,"points":null,"story_id":44386887,"text":"you know everything is made up right? And yet it just works. I too use a confidence score in an bug finder app, Github seems to use them in copilot reviews, people will use them until it is shown not to work anymore.<p>on the other hand this post <a href=\"https:&#x2F;&#x2F;www.greptile.com&#x2F;blog&#x2F;make-llms-shut-up\">https:&#x2F;&#x2F;www.greptile.com&#x2F;blog&#x2F;make-llms-shut-up</a> says that it didn&#x27;t work in their case:<p>&gt; Sadly, this also failed. The LLMs judgment of its own output was nearly random. This also made the bot extremely slow because there was now a whole new inference call in the workflow.","title":null,"type":"comment","url":null}],"created_at":"2025-06-26T16:05:42.000Z","created_at_i":1750953942,"id":44388711,"options":[],"parent_id":44386887,"points":null,"story_id":44386887,"text":"Funny thing is the structured output in the last example.<p>```\n{\n  &quot;reasoning&quot;: &quot;`cfg` can be nil on line 42; dereferenced without check on line 47&quot;,\n  &quot;finding&quot;: &quot;Possible nil\u2011pointer dereference&quot;,\n  &quot;confidence&quot;: 0.81\n}\n```<p>You know the confidence value is completely bogus, don&#x27;t you?","title":null,"type":"comment","url":null},{"author":"jstummbillig","children":[{"author":"brabel","children":[],"created_at":"2025-06-26T16:49:16.000Z","created_at_i":1750956556,"id":44389092,"options":[],"parent_id":44388936,"points":null,"story_id":44386887,"text":"People creating products need to do what gives results right now. And I can attest that breaking up jobs into small steps seems to work better for most scenarios. When that becomes unnecessary, creating products that are useful will become much easier for sure, but I wouldn\u2019t hold my breath.","title":null,"type":"comment","url":null},{"author":"bckr","children":[],"created_at":"2025-06-26T19:41:54.000Z","created_at_i":1750966914,"id":44390580,"options":[],"parent_id":44388936,"points":null,"story_id":44386887,"text":"I\u2019m not being sarcastic when I say that I think supervisor agents and agent swarms in general are the way forward here","title":null,"type":"comment","url":null}],"created_at":"2025-06-26T16:31:42.000Z","created_at_i":1750955502,"id":44388936,"options":[],"parent_id":44386887,"points":null,"story_id":44386887,"text":"The multi agent thing with different roles is so obviously not a great concept, that I am very hesitant to build towards it, even thought it seems to win out right now. We want a AI that internally does what it needs to do to solve a problem, given a good enough problem description, tools and context. I really do not want to have to worry about breaking up tasks into chunks that are smaller than what I could handle myself, and I really hope that that in the near future this will go away.","title":null,"type":"comment","url":null},{"author":"EnPissant","children":[],"created_at":"2025-06-26T17:16:50.000Z","created_at_i":1750958210,"id":44389314,"options":[],"parent_id":44386887,"points":null,"story_id":44386887,"text":"&gt; Explicit reasoning improves clarity. Require your AI to clearly explain its rationale first\u2014this boosts accuracy and simplifies debugging.<p>I wonder what models they are using because reasoning models do this by default, even if they don&#x27;t give you that output.<p>This post reads more like a marketing blog post than any real world advice.","title":null,"type":"comment","url":null},{"author":"iandanforth","children":[{"author":"bckr","children":[],"created_at":"2025-06-26T19:40:13.000Z","created_at_i":1750966813,"id":44390565,"options":[],"parent_id":44390029,"points":null,"story_id":44386887,"text":"I think that article is talking about finding a previously unknown exploit. A known and well documented vulnerability should be much easier to identify","title":null,"type":"comment","url":null}],"created_at":"2025-06-26T18:33:50.000Z","created_at_i":1750962830,"id":44390029,"options":[],"parent_id":44386887,"points":null,"story_id":44386887,"text":"I learned from a recent post (<a href=\"https:&#x2F;&#x2F;sean.heelan.io&#x2F;2025&#x2F;05&#x2F;22&#x2F;how-i-used-o3-to-find-cve-2025-37899-a-remote-zeroday-vulnerability-in-the-linux-kernels-smb-implementation&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;sean.heelan.io&#x2F;2025&#x2F;05&#x2F;22&#x2F;how-i-used-o3-to-find-cve-...</a>) that finding security issues can take 100+ calls to an LLM to get good signal. So I wonder about agent implementers who are trying to get good signal out of single calls, even if they are specialized ones.","title":null,"type":"comment","url":null},{"author":"OnionBlender","children":[],"created_at":"2025-06-26T19:52:52.000Z","created_at_i":1750967572,"id":44390662,"options":[],"parent_id":44386887,"points":null,"story_id":44386887,"text":"What&#x27;s funny about the bullet points in section 3 is that it only compares to the previous noisy agent, rather than having no agent. 51% fewer false positives, median comments per pull request cut by half, spending less time managing irrelevant comments? Turn it off and you could get a 100% reduction in false positives and spend zero time on irrevant AI generated comments.","title":null,"type":"comment","url":null},{"author":"hbogert","children":[],"created_at":"2025-06-27T06:43:33.000Z","created_at_i":1751006613,"id":44394269,"options":[],"parent_id":44386887,"points":null,"story_id":44386887,"text":"ah the joy of non-determinism. Have fun tweaking till you die. Also I wish youa lot of fun giving your customers buttons to disable&#x2F;enable options.","title":null,"type":"comment","url":null}],"created_at":"2025-06-26T12:45:04.000Z","created_at_i":1750941904,"id":44386887,"options":[],"parent_id":null,"points":172,"story_id":44386887,"text":null,"title":"Learnings from building AI agents","type":"story","url":"https://www.cubic.dev/blog/learnings-from-building-ai-agents"}
