{"author":"ibobev","children":[{"author":"simonw","children":[{"author":"embedding-shape","children":[],"created_at":"2026-07-20T21:21:28.000Z","created_at_i":1784582488,"id":48985025,"options":[],"parent_id":48982698,"points":null,"story_id":48979475,"text":"&gt; My favorite trick for controlling the reasoning level is the hack<p>First described here I think, specifically for Qwen models: <a href=\"https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2501.19393\" rel=\"nofollow\">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2501.19393</a>, although it seems unclear how much it actually improves the final response unless you create a new fine-tuned model with this in mind, and also if it&#x27;s possible to replicate for other models. Still, got inspired and gonna give it a try with DiffusionGemma although slightly differently","title":null,"type":"comment","url":null}],"created_at":"2026-07-20T18:17:13.000Z","created_at_i":1784571433,"id":48982698,"options":[],"parent_id":48979475,"points":null,"story_id":48979475,"text":"I&#x27;m amused by how the whole reasoning model thing feels like a formalization of the old &quot;think step by step&quot; prompting hack, which was discovered against GPT-3 two years after that model was first released.<p>My favorite trick for controlling the reasoning level is the hack where you look at the output token stream and spot the token for &quot;the model has concluded reasoning&quot;... and then replace that with the tokens for &quot;wait, but&quot; and force it to keep going!","title":null,"type":"comment","url":null},{"author":"sva_","children":[],"created_at":"2026-07-20T19:00:27.000Z","created_at_i":1784574027,"id":48983333,"options":[],"parent_id":48979475,"points":null,"story_id":48979475,"text":"I recommend his book, &quot;Build a Reasoning Model (From Scratch)&quot;, which is also linked in the article.<p><a href=\"https:&#x2F;&#x2F;sebastianraschka.com&#x2F;books&#x2F;#build-a-reasoning-model-from-scratch\" rel=\"nofollow\">https:&#x2F;&#x2F;sebastianraschka.com&#x2F;books&#x2F;#build-a-reasoning-model-...</a>","title":null,"type":"comment","url":null},{"author":"razorbeamz","children":[{"author":"tenuousemphasis","children":[],"created_at":"2026-07-21T04:54:49.000Z","created_at_i":1784609689,"id":48988188,"options":[],"parent_id":48986685,"points":null,"story_id":48979475,"text":"So do most people.","title":null,"type":"comment","url":null},{"author":"solumunus","children":[],"created_at":"2026-07-21T17:45:47.000Z","created_at_i":1784655947,"id":48995642,"options":[],"parent_id":48986685,"points":null,"story_id":48979475,"text":"It has the same effect as reasoning. Its reasoning in a practical sense if not a philosophical one, and that\u2019s what most people care about.","title":null,"type":"comment","url":null}],"created_at":"2026-07-21T00:24:29.000Z","created_at_i":1784593469,"id":48986685,"options":[],"parent_id":48979475,"points":null,"story_id":48979475,"text":"LLMs don&#x27;t actually reason, they just create the illusion of reasoning by repeating things over and over.","title":null,"type":"comment","url":null},{"author":"cyanydeez","children":[],"created_at":"2026-07-21T01:50:48.000Z","created_at_i":1784598648,"id":48987155,"options":[],"parent_id":48979475,"points":null,"story_id":48979475,"text":"mmm, practically, llamacpp solved this with reasoning budget for tokens and a customized message.<p>ive tailored my local models to match agent with a message that pushes to use subagents and context compression.<p>it works pretty smoothly if theres proper vector and scope to the project. theres probably a way to also get it to record memories but ive not seen a memory system that includes spatial type reasoning which would create memories associated with file paths down to AST graphlike leaves.<p>then we could include remembering details about the path.<p>llamacpp also has a header for setting these that an intelligent harness could tailor per round message and budget. ideally you start with a small budget and expand based on some complexity criteria.<p>alas, i want to build things not related to AI.","title":null,"type":"comment","url":null},{"author":"nxtfari","children":[],"created_at":"2026-07-21T15:39:48.000Z","created_at_i":1784648388,"id":48993784,"options":[],"parent_id":48979475,"points":null,"story_id":48979475,"text":"Very high quality and info-dense article. Cheers.","title":null,"type":"comment","url":null}],"created_at":"2026-07-20T14:35:19.000Z","created_at_i":1784558119,"id":48979475,"options":[],"parent_id":null,"points":84,"story_id":48979475,"text":null,"title":"Controlling Reasoning Effort in LLMs","type":"story","url":"https://magazine.sebastianraschka.com/p/controlling-reasoning-effort-in-llms"}
