{"author":"debazel","children":[{"author":"nwienert","children":[],"created_at":"2026-08-08T20:55:12.000Z","created_at_i":1786222512,"id":49225823,"options":[],"parent_id":49220347,"points":null,"story_id":49214008,"text":"That&#x27;s roughly my experience. Luna is extremely efficient and at higher levels of reasoning and longer running tasks more capable.<p>Reading the DS reasoning is wild, it&#x27;s constantly going in circles. The most minor lack of clarity in your prompt and it will spend ages going back and forth on what you meant. It reasons 5x longer than the preview which makes it really slow now as well. We did a lot of work to nudge it to be decisive and improve our evaluation setup to there&#x27;s more clarity, and it helped but only marginally.<p>Ours is a full-stack app one shot test so it includes backend, frontend, design, and QA&#x2F;testing. It&#x27;s graded by Opus xhigh and Sol xhigh and the grades are averaged.<p>DS4 preview would finish in 20 minutes flat on high reasoning and grades 6&#x2F;10. Luna high gets 9&#x2F;10 in about 30 minutes. DS4-final is crazy - at high thinking it&#x27;s taking over an hour and getting ~8 but only had one successful run as I got tired of waiting so long after many early abort&#x2F;retries trying to debug why thinking was so long. The lowest thinking still takes over 45 minutes, and with thinking off it actually is finally closer to preview in time but actually get&#x27;s a much more varying result anywhere from incomplete to 6 it seems.<p>Costs per run DS4 is best but not actually by a lot as it&#x27;s spending 10x the tokens with all the reasoning and mistakes. It&#x27;s a very brute force model and I really preferred preview in many ways for how predictably fast it was.<p>Side note, Spark 1.2 is a nice model for this test, best in frontend design and fastest to get results together, though not nearly as efficient as Luna. Grok scores similarly to Spark but at like $50&#x2F;run vs the contributor Spark costing $1.50.<p>Edit: was curious to see and seems DeepSWE agrees at least: <a href=\"https:&#x2F;&#x2F;www.together.ai&#x2F;blog&#x2F;deepseek-v4-flash-0731-vs-gpt-5-6-luna-on-deepswe-cost-and-coding\" rel=\"nofollow\">https:&#x2F;&#x2F;www.together.ai&#x2F;blog&#x2F;deepseek-v4-flash-0731-vs-gpt-5...</a><p>Edit 2: btw it tests a team of agents working together in a special harness and stack. So 20 minutes is for 8 agents essentially. That said everything was built around DS as it was the cheapest to iterate against so even with that advantage the new one struggles.","title":null,"type":"comment","url":null}],"created_at":"2026-08-08T10:06:43.000Z","created_at_i":1786183603,"id":49220347,"options":[],"parent_id":49219761,"points":null,"story_id":49214008,"text":"I&#x27;ve been using this DeepSeek model the whole day today after building with 5.6 Luna extensively over the last week and I would disagree, at least for Rust + OpenGL.<p>DeepSeek just spend almost 2 hours trying to figure out why terrain textures were not working. It tried everything over and over again, it even had reference code for meshes on how to setup the rendering with materials, and it could just not do it.<p>I finally gave up and gave it to GPT-5.6 Luna instead, and figure out in a single prompt after 20 seconds, that the terrain mesh was being initialized with None in the material slot.<p>Other tasks it has managed to figure out at least, but it is significantly slower than GPT-5.6 Luna and it requires a lot more iterations.<p>(Both were set to high reasoning)","title":null,"type":"comment","url":null}
