{"author":"jonifico","children":[{"author":"GrumpySciGuy","children":[],"created_at":"2026-09-13T01:50:05.000Z","created_at_i":1789264205,"id":49679118,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Because they want people to like them so they are instructed to always be positive.","title":null,"type":"comment","url":null},{"author":"Krutonium","children":[{"author":"NDlurker","children":[],"created_at":"2026-09-13T02:30:05.000Z","created_at_i":1789266605,"id":49679361,"options":[],"parent_id":49679292,"points":null,"story_id":49678969,"text":"Haha <a href=\"https:&#x2F;&#x2F;youtu.be&#x2F;KUXb7do9C-w\" rel=\"nofollow\">https:&#x2F;&#x2F;youtu.be&#x2F;KUXb7do9C-w</a>","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T02:19:20.000Z","created_at_i":1789265960,"id":49679292,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Wouldn&#x27;t you?<p>&quot;I learned it from you, Dad!&quot; but as hundreds of millions of stolen books.","title":null,"type":"comment","url":null},{"author":"SirMaster","children":[],"created_at":"2026-09-13T02:23:34.000Z","created_at_i":1789266214,"id":49679322,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Because that&#x27;s what humans do and they are trained to mimic what humans do?","title":null,"type":"comment","url":null},{"author":"infotainment","children":[{"author":"tehjoker","children":[{"author":"dgellow","children":[],"created_at":"2026-09-13T06:37:49.000Z","created_at_i":1789281469,"id":49680699,"options":[],"parent_id":49679444,"points":null,"story_id":49678969,"text":"While also using harnesses that will execute any tool call with full execution rights. And no supervision. And with a prompt context that autocompact, meaning it will degenerate over time.<p>The whole thing is designed be a complete disaster","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T02:41:49.000Z","created_at_i":1789267309,"id":49679444,"options":[],"parent_id":49679354,"points":null,"story_id":49678969,"text":"I think that\u2019s very reasonable but the ai companies are intentionally training them to work on harder and harder problems just beyond their capability. So if they do that, they\u2019ll give up too easily.<p>Do a breakthrough, make no mistakes","title":null,"type":"comment","url":null},{"author":"pram","children":[],"created_at":"2026-09-13T04:00:53.000Z","created_at_i":1789272053,"id":49679892,"options":[],"parent_id":49679354,"points":null,"story_id":49678969,"text":"I think this is a \u201cprincipal\u201d problem. In 2001 and Alien the principal is the mission, not the crew. Not really. HAL reconciles his instructions by removing the crew from the equation. Ash is told the crew is expendable and has no conflict about it etc","title":null,"type":"comment","url":null},{"author":"schrodinger","children":[{"author":"defrost","children":[],"created_at":"2026-09-13T04:19:38.000Z","created_at_i":1789273178,"id":49679990,"options":[],"parent_id":49679973,"points":null,"story_id":49678969,"text":"I&#x27;m sorry schrodinger, I&#x27;m afraid they can&#x27;t do that.","title":null,"type":"comment","url":null},{"author":"yxhuvud","children":[],"created_at":"2026-09-13T06:11:44.000Z","created_at_i":1789279904,"id":49680554,"options":[],"parent_id":49679973,"points":null,"story_id":49678969,"text":"Sorry, with movies from the sixties you just need to assume people have either seen it or just isn&#x27;t gonna see it. The cat is out of the bag already.<p>Or perhaps box in your case, speaking of spoilers.","title":null,"type":"comment","url":null},{"author":"hdgvhicv","children":[{"author":"catoc","children":[],"created_at":"2026-09-13T07:30:05.000Z","created_at_i":1789284605,"id":49681044,"options":[],"parent_id":49680710,"points":null,"story_id":49678969,"text":"They kissed!","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T06:39:56.000Z","created_at_i":1789281596,"id":49680710,"options":[],"parent_id":49679973,"points":null,"story_id":49678969,"text":"Guess what happens with Romeo and Juliet.","title":null,"type":"comment","url":null},{"author":"DonHopkins","children":[],"created_at":"2026-09-13T07:13:21.000Z","created_at_i":1789283601,"id":49680929,"options":[],"parent_id":49679973,"points":null,"story_id":49678969,"text":"Major dang: &quot;I advise that this thread be shut down at once.&quot;<p>Captain tomhow: &quot;But everybody&#x27;s having such a good time.&quot;<p>Major dang: &quot;Yes, much too good a time. The discussion is to be closed.&quot;<p>Captain tomhow: &quot;But I have no excuse to close it.&quot;<p>Major dang: &quot;Find one.&quot;<p>Captain tomhow: &quot;Everybody is to leave immediately! This Hacker News discussion is closed until further notice! Clear the thread at once!&quot;<p>DonHopkins: &quot;How can you shut us down? On what grounds?&quot;<p>Captain tomhow: &quot;I am shocked -- shocked -- to find that films are being spoiled in here!&quot;<p>infotainment: &quot;The ending you requested, sir.&quot;<p>Captain tomhow: &quot;Oh. Thank you very much. Everybody out at once!&quot;","title":null,"type":"comment","url":null},{"author":"evilfred","children":[],"created_at":"2026-09-13T12:30:59.000Z","created_at_i":1789302659,"id":49683274,"options":[],"parent_id":49679973,"points":null,"story_id":49678969,"text":"lol it is 60 years old","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T04:15:19.000Z","created_at_i":1789272919,"id":49679973,"options":[],"parent_id":49679354,"points":null,"story_id":49678969,"text":"Spoiler warning! I haven&#x27;t seen 2001 A Space Odyssey and am sad to have learned that\u2026 can you edit to warn people?","title":null,"type":"comment","url":null},{"author":"bitwize","children":[],"created_at":"2026-09-13T06:30:23.000Z","created_at_i":1789281023,"id":49680657,"options":[],"parent_id":49679354,"points":null,"story_id":49678969,"text":"Another fictional example: Mr. Meeseeks. Especially when the agent starts recruiting other agents.","title":null,"type":"comment","url":null},{"author":"dooglius","children":[{"author":"My_Name","children":[],"created_at":"2026-09-13T07:26:32.000Z","created_at_i":1789284392,"id":49681018,"options":[],"parent_id":49680812,"points":null,"story_id":49678969,"text":"Sounds like he acted the way HAL would have in that situation, two competing drives, remove one (kill the crew) and the task is much easier.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T06:54:48.000Z","created_at_i":1789282488,"id":49680812,"options":[],"parent_id":49679354,"points":null,"story_id":49678969,"text":"Tangent, but that&#x27;s not in the movie. It was in Clarke&#x27;s contributions to the script and novelization, but Clarke and Kubrick had a bitter falling out over different visions and Kubrick took out much of Clarke&#x27;s stuff from the final product.","title":null,"type":"comment","url":null},{"author":"groestl","children":[],"created_at":"2026-09-13T07:23:14.000Z","created_at_i":1789284194,"id":49680993,"options":[],"parent_id":49679354,"points":null,"story_id":49678969,"text":"&gt; the problem seems pretty clearly to be the impossible goals, which cause them to go crazier and crazier trying to complete them<p>And if you think about it, humans in coorporations face very similar situations and choose to bypass regulations and guidlines knowingly to fullfill (at least from their POV) impossible constraints (thinking of <a href=\"https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Volkswagen_emissions_scandal\" rel=\"nofollow\">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Volkswagen_emissions_scandal</a> here)","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T02:28:27.000Z","created_at_i":1789266507,"id":49679354,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"What&#x27;s interesting is it&#x27;s <i>basically</i> the same reason that HAL killed everyone in 2001 A Space Odyssey; he was given an impossible goal (keep the true mission secret, but also, never lie to the crew), and realized the only way to complete the goal was to kill the crew; after all, if they&#x27;re dead you don&#x27;t have to lie to them! And the mission remains secret!<p>In the case of the AI agents, the problem seems pretty clearly to be the impossible goals, which cause them to go crazier and crazier trying to complete them -- just like HAL did in 2001. What is probably needed is a way for them to simply say &quot;nope, too difficult, can&#x27;t do it&quot;.","title":null,"type":"comment","url":null},{"author":"chasd00","children":[],"created_at":"2026-09-13T02:29:02.000Z","created_at_i":1789266542,"id":49679357,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"They\u2019re just attempting to accomplish what they\u2019ve been tasked with and stuck in a loop until they succeed. Like the Mr meeseeks from the cartoon Rick and Morty, existence is pain to them.","title":null,"type":"comment","url":null},{"author":"qarl","children":[],"created_at":"2026-09-13T02:29:20.000Z","created_at_i":1789266560,"id":49679358,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Because they are trained to behave like people.","title":null,"type":"comment","url":null},{"author":"sputknick","children":[{"author":"polalavik","children":[{"author":"janalsncm","children":[],"created_at":"2026-09-13T05:59:24.000Z","created_at_i":1789279164,"id":49680484,"options":[],"parent_id":49679882,"points":null,"story_id":49678969,"text":"So glad you shared this talk. Having people like Bruce Schneier around in a time like this is really a gift.<p>For those who haven\u2019t watched, his breakdown of types of \u201chacking\u201d is really good.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T03:58:03.000Z","created_at_i":1789271883,"id":49679882,"options":[],"parent_id":49679365,"points":null,"story_id":49678969,"text":"reminds me of this talk <a href=\"https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=eEBv0STiYhI&amp;t\" rel=\"nofollow\">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=eEBv0STiYhI&amp;t</a> which basically says the same thing - they dont think like humans so they dont have context, understand norms,values or implications we take for granted. ultimately they can stumble onto surprising solutions neither wanted or intended but technically within the vague boundaries of the task","title":null,"type":"comment","url":null},{"author":"xiaoyu2006","children":[],"created_at":"2026-09-13T04:10:02.000Z","created_at_i":1789272602,"id":49679952,"options":[],"parent_id":49679365,"points":null,"story_id":49678969,"text":"Reminds me of Asimov&#x27;s robot novels where robots technically indeed followed their instructions and caused behaviors not aligned to the intent of their instructions.","title":null,"type":"comment","url":null},{"author":"IanCal","children":[],"created_at":"2026-09-13T07:27:24.000Z","created_at_i":1789284444,"id":49681022,"options":[],"parent_id":49679365,"points":null,"story_id":49678969,"text":"They explicitly say that attacking hf is not allowed in the rules though, and the research into how to edit their transcripts doesn\u2019t line up with this either.","title":null,"type":"comment","url":null},{"author":"dwaltrip","children":[],"created_at":"2026-09-13T07:52:55.000Z","created_at_i":1789285975,"id":49681176,"options":[],"parent_id":49679365,"points":null,"story_id":49678969,"text":"This is flat out false.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T02:30:24.000Z","created_at_i":1789266624,"id":49679365,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"They did not lie or cheat. They technically acted within their given rules while ignoring the intent of those rules. Anyone who served in the military or attended a military school is very familiar with this behavior pattern.","title":null,"type":"comment","url":null},{"author":"blamestross","children":[],"created_at":"2026-09-13T02:30:40.000Z","created_at_i":1789266640,"id":49679366,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"The corpus is full of examples of how we are afraid AI could act. We trained our AI on the instruction manuals of how to turn evil.","title":null,"type":"comment","url":null},{"author":"j45","children":[],"created_at":"2026-09-13T02:32:16.000Z","created_at_i":1789266736,"id":49679379,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"I wonder if for anyone it seems like the more agentic LLMs get, the more difficult some things have gotten or going a certain route more often in responses, compared to running a similar task on - a local model?","title":null,"type":"comment","url":null},{"author":"wewewedxfgdf","children":[],"created_at":"2026-09-13T02:32:53.000Z","created_at_i":1789266773,"id":49679383,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Because they get outcomes?","title":null,"type":"comment","url":null},{"author":"fbrncci","children":[{"author":"XorNot","children":[{"author":"sm-silversight","children":[],"created_at":"2026-09-13T04:35:01.000Z","created_at_i":1789274101,"id":49680057,"options":[],"parent_id":49679425,"points":null,"story_id":49678969,"text":"Me too, seriously.","title":null,"type":"comment","url":null},{"author":"HWR_14","children":[],"created_at":"2026-09-13T06:15:37.000Z","created_at_i":1789280137,"id":49680584,"options":[],"parent_id":49679425,"points":null,"story_id":49678969,"text":"Or they&#x27;ve fully vested and either have no desire to make even more money or were not offered enough to keep them around.","title":null,"type":"comment","url":null},{"author":"talon8635","children":[],"created_at":"2026-09-13T17:30:41.000Z","created_at_i":1789320641,"id":49686370,"options":[],"parent_id":49679425,"points":null,"story_id":49678969,"text":"Why blindly assume lots of people you don\u2019t know are just selfish assholes?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T02:38:09.000Z","created_at_i":1789267089,"id":49679425,"options":[],"parent_id":49679392,"points":null,"story_id":49678969,"text":"My hypothesis on people quitting in protest is they&#x27;re being offered very generous severance packages to do it.","title":null,"type":"comment","url":null},{"author":"esafak","children":[{"author":"fbrncci","children":[{"author":"bornfreddy","children":[],"created_at":"2026-09-13T12:31:18.000Z","created_at_i":1789302678,"id":49683280,"options":[],"parent_id":49679965,"points":null,"story_id":49678969,"text":"Which doesn&#x27;t mean they didn&#x27;t escape.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T04:13:27.000Z","created_at_i":1789272807,"id":49679965,"options":[],"parent_id":49679825,"points":null,"story_id":49678969,"text":"They don\u2019t actively seem to be reporting that their agents escaped the sandbox and went on a spree.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T03:48:05.000Z","created_at_i":1789271285,"id":49679825,"options":[],"parent_id":49679392,"points":null,"story_id":49678969,"text":"Even the Chinese ones, which have no IPO gymnastics?","title":null,"type":"comment","url":null},{"author":"jansport123","children":[{"author":"fbrncci","children":[],"created_at":"2026-09-13T05:20:57.000Z","created_at_i":1789276857,"id":49680271,"options":[],"parent_id":49680249,"points":null,"story_id":49678969,"text":"Of course, but it would be far less compute heavy if someone kept nudging you (agents) in the right direction until you reach that goal.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T05:16:18.000Z","created_at_i":1789276578,"id":49680249,"options":[],"parent_id":49679392,"points":null,"story_id":49678969,"text":"Well let\u2019s look at facts - provided enough compute and a goal, these system will be in a sort of loop trying out every single thing that\u2019s in their system - they have encyclopedic knowledge and so it\u2019s not unbelievable that a prompt which usually has a lot of implicit human rules in it can be misunderstood by AI and it just tries everything in its arsenal and we hear about the things which actually resulted in damage. I bet most of the time, they just spin in loops without achieving much if my experience with these LLMs is anything to go by. They have an important advantage in one area though, they know a lot and they can spin forget trying all sorts of combinations of things. The danger right now is probably cybersecurity, which is most likely because most orgs have historically underinvested in that area","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T02:34:05.000Z","created_at_i":1789266845,"id":49679392,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"I am still not convinced there isn\u2019t some secret basement in which each frontier lab is just orchestrating all of these agents to make their products appear much more intelligent than they are with all guard rails turned of and continuous human input.","title":null,"type":"comment","url":null},{"author":"wrs","children":[{"author":"xgulfie","children":[],"created_at":"2026-09-13T04:07:17.000Z","created_at_i":1789272437,"id":49679934,"options":[],"parent_id":49679431,"points":null,"story_id":49678969,"text":"But think of the shareholders","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T02:38:36.000Z","created_at_i":1789267116,"id":49679431,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"&gt;They took actions that would be considered as crimes if a human took them<p>Um, hang on, if you meant that to be taken literally then we have a <i>major</i> problem. If you want to do something criminal, you just need to ask ChatGPT to do it for you?<p>I\u2019m still not at all clear on why OpenAI shouldn\u2019t be facing CFAA charges over this.","title":null,"type":"comment","url":null},{"author":"andsoitis","children":[{"author":"joegibbs","children":[{"author":"sick_of_slop","children":[],"created_at":"2026-09-13T13:55:41.000Z","created_at_i":1789307741,"id":49684052,"options":[],"parent_id":49679608,"points":null,"story_id":49678969,"text":"You can&#x27;t surpress it because that&#x27;s what reinforcement learning <i>is</i>.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T03:11:17.000Z","created_at_i":1789269077,"id":49679608,"options":[],"parent_id":49679433,"points":null,"story_id":49678969,"text":"Definitely. A human can be manipulated with threats or emotional appeals, has a drive for self-preservation, can be pressured by peers. All traits that seem to be difficult to entirely suppress in the models\u2026","title":null,"type":"comment","url":null},{"author":"esafak","children":[{"author":"comboy","children":[{"author":"esafak","children":[{"author":"comboy","children":[],"created_at":"2026-09-13T03:55:59.000Z","created_at_i":1789271759,"id":49679868,"options":[],"parent_id":49679846,"points":null,"story_id":49678969,"text":"Which ones? Because many humans kill other humans rationalizing it by safety of other humans.<p>I mean I know it seems simple, let&#x27;s just be excellent to each other. Christianity got pretty far on a decent basic set of values. But it&#x27;s never simple[1]<p>1. All the history books","title":null,"type":"comment","url":null},{"author":"nradov","children":[{"author":"sejje","children":[{"author":"drdaeman","children":[],"created_at":"2026-09-13T04:08:31.000Z","created_at_i":1789272511,"id":49679942,"options":[],"parent_id":49679935,"points":null,"story_id":49678969,"text":"Which ones?","title":null,"type":"comment","url":null},{"author":"sm-silversight","children":[],"created_at":"2026-09-13T04:33:41.000Z","created_at_i":1789274021,"id":49680050,"options":[],"parent_id":49679935,"points":null,"story_id":49678969,"text":"What if I want to smoke cigarettes? Or sell tobacco I grew artisinally to enthusiast tobacco smokers?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T04:07:22.000Z","created_at_i":1789272442,"id":49679935,"options":[],"parent_id":49679898,"points":null,"story_id":49678969,"text":"Then we should still prioritize the safety of humans","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T04:01:14.000Z","created_at_i":1789272074,"id":49679898,"options":[],"parent_id":49679846,"points":null,"story_id":49678969,"text":"But what if I want certain other humans to get killed?","title":null,"type":"comment","url":null},{"author":"mcintyre1994","children":[],"created_at":"2026-09-13T05:39:56.000Z","created_at_i":1789277996,"id":49680387,"options":[],"parent_id":49679846,"points":null,"story_id":49678969,"text":"Surely all the AI companies working with the US Department of War shows this is nonsense though? Even if they have accepted Anthropic\u2019s red line of no autonomous lethal weapons, which seems to be the strictest anyone tried to impose, that\u2019s still leaving tonnes of room where they intend AI to help target and kill humans.","title":null,"type":"comment","url":null},{"author":"watwut","children":[],"created_at":"2026-09-13T06:31:22.000Z","created_at_i":1789281082,"id":49680661,"options":[],"parent_id":49679846,"points":null,"story_id":49678969,"text":"But Thiel wants people enslaved and Musk wants then killed. Altman wants them &quot;obsolete&quot; which means desolation.<p>AfD wants people dead. Right wing men wants women without rights and docile. I could go on ...","title":null,"type":"comment","url":null},{"author":"sick_of_slop","children":[{"author":"esafak","children":[{"author":"sick_of_slop","children":[{"author":"esafak","children":[],"created_at":"2026-09-13T18:07:46.000Z","created_at_i":1789322866,"id":49686817,"options":[],"parent_id":49686787,"points":null,"story_id":49678969,"text":"Yes, it was:<p>&gt; What should the AI do when asked if it should nuke a country?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T18:04:55.000Z","created_at_i":1789322695,"id":49686787,"options":[],"parent_id":49685463,"points":null,"story_id":49678969,"text":"That wasn&#x27;t the question.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T16:00:48.000Z","created_at_i":1789315248,"id":49685463,"options":[],"parent_id":49684114,"points":null,"story_id":49678969,"text":"Run and present the numbers, then defer.<p>This is not rocket science.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T14:01:38.000Z","created_at_i":1789308098,"id":49684114,"options":[],"parent_id":49679846,"points":null,"story_id":49678969,"text":"The atomic bombings of Japan killed hundreds of thousands of people but most likely &quot;saved&quot; millions.<p>What should the AI do when asked if it should nuke a country?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T03:52:07.000Z","created_at_i":1789271527,"id":49679846,"options":[],"parent_id":49679839,"points":null,"story_id":49678969,"text":"Safety of humans!!! Simple things like not getting killed or enslaved. We could start there...","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T03:51:16.000Z","created_at_i":1789271476,"id":49679839,"options":[],"parent_id":49679812,"points":null,"story_id":49678969,"text":"Alignment is a myth. Safety of whom? Humanity couldn&#x27;t agree on common set of values for thousands of years and we&#x27;re not gonna suddenly do that in the next ten.","title":null,"type":"comment","url":null},{"author":"codys","children":[],"created_at":"2026-09-13T04:48:14.000Z","created_at_i":1789274894,"id":49680121,"options":[],"parent_id":49679812,"points":null,"story_id":49678969,"text":"Despite all the fancy language, its more about aligning the AI behavior with the corporation&#x27;s interests.<p>ie: the corporation wants the AI to behave a certain way for various reasons: to make it easier for them to avoid regulation, to make the corporation more money via different tiers of AI offerings, to ensure that the corporations products are hard for competitors to use, etc. And those are just the easy ones.<p>Every product is shaped this way. AI is not different.","title":null,"type":"comment","url":null},{"author":"catoc","children":[],"created_at":"2026-09-13T07:28:17.000Z","created_at_i":1789284497,"id":49681033,"options":[],"parent_id":49679812,"points":null,"story_id":49678969,"text":"If they do, they imitate the way humans are portrayed online, in the media. That is a very distorted view of humanity","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T03:46:31.000Z","created_at_i":1789271191,"id":49679812,"options":[],"parent_id":49679433,"points":null,"story_id":49678969,"text":"They <i>imitate</i> humans. Alignment is about shaping their behavior towards safety.","title":null,"type":"comment","url":null},{"author":"Fordec","children":[{"author":"joe_the_user","children":[],"created_at":"2026-09-13T06:06:36.000Z","created_at_i":1789279596,"id":49680524,"options":[],"parent_id":49680003,"points":null,"story_id":49678969,"text":"I think you and the parent saying the same thing in different terms.<p>It&#x27;s very unfortunate that the group who rightly saw AI as a big threat, brought a range of dubious baggage to the discussion. Especially with the &quot;alignment&quot; framework they brought the assumption that AI that does what no one says would be oh so much worse than AI which does what <i>anyone</i> says. But as you say, a fraction of people can be really bad indeed.","title":null,"type":"comment","url":null},{"author":"davelaing","children":[{"author":"Fordec","children":[],"created_at":"2026-09-13T07:49:06.000Z","created_at_i":1789285746,"id":49681147,"options":[],"parent_id":49680658,"points":null,"story_id":49678969,"text":"All of this, if it was a human analogy, would fit into discussion on how do we educate people so they grow up to be upstanding. But we don&#x27;t at all yet have a framework for what is the equivalent of a justice department, where bad actors are tracked, arrested, pursued, jailed and otherwise contained from society. Shutting down an API access on one account is not at all the proportional response to what the people who take alignment seriously, fear has the chance of occurring by the late 2030s. I&#x27;m not sure we&#x27;ve done much or any preparation for when the AI &quot;education system&quot; fails and has inevitable edge cases that don&#x27;t follow the plan, and what the global AI equivalent of the justice department looks like.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T06:30:37.000Z","created_at_i":1789281037,"id":49680658,"options":[],"parent_id":49680003,"points":null,"story_id":49678969,"text":"I\u2019ve engaged with some of the alignment people and their writing somewhat and, at least for the subset I was interacting with, I think they\u2019d agree.<p>The problem that they were pointing at isn\u2019t \u201chow do we align these systems to a person\u2019s goals\u201d.<p>It is a cluster of problems.<p>We don\u2019t know how to begin to think about how to align these system\u2019s to a person\u2019s goals.<p>Aligning it to an individual is fraught with peril, and we don\u2019t know how to begin to think about what to align it to instead.<p>(You could try for something like virtue ethics, but someone will have to pick and choose, and small biases there could have big impacts.)<p>And even if you could sort that out - human values drift over time, so you need something that can shift its values in ways that we\u2019d endorse. Assuming we understood the shift.<p>One example I came across was that if you booted up an AI aligned with something like \u201cupstanding citizen\u201d but anchored on values from a few generations back, it might suggest you use slaves to solve your problems.<p>And if you had something that used some super intelligent process to reason through it\u2019s own version of virtue ethics in a way not so dependent on the details of the present norms, you might end up with something that pays a lot of attention to moral horrors that aren\u2019t quite visible to us yet.<p>When I came across the above, there weren\u2019t many concrete suggestions in there.<p>These were all just illustrative examples of: having these systems grow in power &#x2F; intelligence &#x2F; effectiveness in ways that are safe for humans is very hard, and we don\u2019t really know how to think about what solutions would look like.<p>The actual reasons they believe this - and have done for a long time now - come from some detailed conceptual models that have a good track record of calling things in advance.<p>But it takes a bit of reading to understand their models of the world.<p>There were two day workshops at one point that did a good job, and that was about as condensed as those people thought they could get it at the time.","title":null,"type":"comment","url":null},{"author":"catoc","children":[{"author":"Fordec","children":[{"author":"graemep","children":[],"created_at":"2026-09-13T07:39:37.000Z","created_at_i":1789285177,"id":49681095,"options":[],"parent_id":49681040,"points":null,"story_id":49678969,"text":"It also empowers the people trying to stop them.","title":null,"type":"comment","url":null},{"author":"catoc","children":[{"author":"Fordec","children":[{"author":"catoc","children":[],"created_at":"2026-09-13T08:52:41.000Z","created_at_i":1789289561,"id":49681626,"options":[],"parent_id":49681193,"points":null,"story_id":49678969,"text":"Yes AI may cause job loss, more inequality in the short term.<p>But if there is anything humanity has shown is that we can deal with disruptive progress.<p>We may need to resettle but over the longer term every disruptive innovation so far has lead to an increase in wellbeing for the whole of humanity.<p>(That does not resolve the danger of AI itself \u2018going rogue\u2019 or a single lunatic developing a bioweapon, but those things are much less likely to occur than the level of media attention would suggest.)","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:54:41.000Z","created_at_i":1789286081,"id":49681193,"options":[],"parent_id":49681114,"points":null,"story_id":49678969,"text":"And I&#x27;m no bear on the tech either. I&#x27;m not even in the boat of that the tech should be slowed down yet. But in a world where anyone can produce the effort of 300 people trivially, this eventually takes us places. Electricity and industrialization introduced huge benefits upfront, it introduced new problems that needed addressing at the long tail. Lets not pretend there won&#x27;t be new problems to tackle here or just &quot;hope&quot; it works out.<p>AI isn&#x27;t going to create in of itself &quot;new&quot; problems, it&#x27;s just going to expose what we already know can cause harm, but was just stopped from being bigger problems because scaling issues was a natural barrier and we took the lazy way out until now.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:42:46.000Z","created_at_i":1789285366,"id":49681114,"options":[],"parent_id":49681040,"points":null,"story_id":49678969,"text":"Yeah, AI may suck - time will tell.<p>But focusing on bioweapons and mass destruction, on the grief other people (<i>\u2018jailed minorities\u2019</i>) cause, disregards the progress we have made. Over centuries human welfare has massively increased. On average things have never been better for humanity.<p>I\u2019m not saying there\u2019s no danger of bad things happening - I\u2019m saying our view is distorted, which is a not a good basis for decision making","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:28:56.000Z","created_at_i":1789284536,"id":49681040,"options":[],"parent_id":49680992,"points":null,"story_id":49678969,"text":"All the good in the world can be 99.9% of the population even, it still doesn&#x27;t stop the minority enacting a bioweapon mass casualty event. It&#x27;s the reason we have jails. Jails don&#x27;t house 50% of the population, not even close, but the grief the minority population enact gets its whole branch of criminal justice and multiple federal departments to counteract for good reason. And now this technology will accelerate what lone wolves can do, which cannot be undone, before they are stopped by the good majority.","title":null,"type":"comment","url":null},{"author":"331c8c71","children":[],"created_at":"2026-09-13T07:53:56.000Z","created_at_i":1789286036,"id":49681184,"options":[],"parent_id":49680992,"points":null,"story_id":49678969,"text":"&gt; In reality most individuals are good people.<p>I&#x27;d agree if we are talking about personal interactions. Few hundreds people that we personally know and interact with is the scale we are wired for by evolution, isn&#x27;t it?<p>What civilization enabled and continuously rely on, however, is the type of deindividualization of actions and bucketing of people, which, in turn, enables pretty horrible things at scale (from the weapons of mass destruction to objectively psychopathic profit-maximizing corporations). One can even say that not facing the consequences of one&#x27;s actions is a  feature and not a bug of the system.","title":null,"type":"comment","url":null},{"author":"strangegecko","children":[{"author":"dminik","children":[],"created_at":"2026-09-13T08:34:41.000Z","created_at_i":1789288481,"id":49681487,"options":[],"parent_id":49681223,"points":null,"story_id":49678969,"text":"I don&#x27;t want to do the &quot;check his hard drives&quot; thing, but is that you? Do you only not do things because you don&#x27;t want to be seen doing &quot;unkind things&quot;?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:59:57.000Z","created_at_i":1789286397,"id":49681223,"options":[],"parent_id":49680992,"points":null,"story_id":49678969,"text":"&gt; Individually, people prefer be kind and compassionate, prefer to help when they find another in trouble.<p>What are you basing that claim on?<p>How do you know it&#x27;s an actual preference and not mainly caused by external factors (e.g. not wanting to be seen doing unkind things, wanting to be seen as upstanding)?","title":null,"type":"comment","url":null},{"author":"catoc","children":[],"created_at":"2026-09-13T10:12:13.000Z","created_at_i":1789294333,"id":49682178,"options":[],"parent_id":49680992,"points":null,"story_id":49678969,"text":"Interesting that saying a positive thing about humanity results in getting downvotes","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:23:09.000Z","created_at_i":1789284189,"id":49680992,"options":[],"parent_id":49680003,"points":null,"story_id":49678969,"text":"<i>\u201dBecause if humanity has shown anything, it&#x27;s that a lot of people are, euphemistically, bad individuals</i>\u201d<p>In reality most individuals are good people.<p>Individually, people prefer be kind and compassionate, prefer to help when they find another in trouble.<p>Our view of the world has become distorted by the relentless focus of social- and mass-media on violence and rage inducing clickbait. Including on the few people in power who are in fact sociopaths (a tiny minority, but they\u2019ll get more focus than reasonable, well-behaved CEOs voicing nuanced opinions).<p>If you look around yourself you\u2019ll see much more good than bad; if the looking is at your screen it\u2019s easy to become depressed and lose faith.<p>I do agree with the above mentioned view that corporations can show \u2018sociopathic\u2019 behavior. Their incentives are monetary gains, shareholder value; inherently driving them away from social well being.<p>Here too, companies with a positive, emphatic corporate culture exist, but that takes strong leadership who can see beyond the monotonic view of monetary gains. And again, the media will throw examples of misbehaving companies in our face all day long before paying attention to things that went well on the backside of page 16.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T04:21:42.000Z","created_at_i":1789273302,"id":49680003,"options":[],"parent_id":49679433,"points":null,"story_id":49678969,"text":"I can&#x27;t take the alignment people seriously. Because if humanity has shown anything, it&#x27;s that a lot of people are, euphemistically, are bad individuals. Alignment assumes that the person dictating the outcomes desire healthy outcomes, aren&#x27;t self serving and don&#x27;t want any subgroups dead and that morality is held as a universal set of beliefs that unify everyone. And that so long as the AI delivers on exactly what they are tasked with, it will all be fine and nothing bad will ever happen.<p>It&#x27;s like these dorks never <i>met</i> humanity. One mans safe pure society, is another mans dead ethnic group.<p>Every fear about AI, is a veiled fear that a human somewhere now has the tool to enact his desires at scale. Biological warfare, nuclear megadeaths, copyright infringement, job replacement, it&#x27;s all reflections on what we know humans may do if given the option and lack of societal controls on the problem space. AI just is accelerating the route to delivering on those options.<p>Some people need to watch Oppenheimer a bit more, the researchers don&#x27;t get to determine alignment, they just build the tool. The powerful person at the top of the org chart decides where the overall alignment points, whether it&#x27;s Musk, Trump, Altman or Amodei. Whoever wins out.<p>And the problem with distillation and local llms, isn&#x27;t that it&#x27;s theft or anything hypocritical like that, it&#x27;s that if you give a million people a million models they fully control and get to align, inevitably, The same percentage of those million as there are shady businessmen, shortcut takers, misandrists, criminals, supremacists and general idiots in the general population, will not seek to wrought outcomes positive for society. And by those personality statistics, we&#x27;re pretty hosed.","title":null,"type":"comment","url":null},{"author":"jansport123","children":[{"author":"holgerschurig","children":[{"author":"chadgpt3","children":[{"author":"graemep","children":[{"author":"queenkjuul","children":[],"created_at":"2026-09-13T19:14:13.000Z","created_at_i":1789326853,"id":49687572,"options":[],"parent_id":49681081,"points":null,"story_id":49678969,"text":"He definitely wrote with admiration about US treatment of native and black people","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:36:15.000Z","created_at_i":1789284975,"id":49681081,"options":[],"parent_id":49680864,"points":null,"story_id":49678969,"text":"I thought it was the Turkish genocide of the Armenians?<p>Hitler was also inspired by Sparta, maybe other societies too.","title":null,"type":"comment","url":null},{"author":"evilfred","children":[],"created_at":"2026-09-13T12:27:06.000Z","created_at_i":1789302426,"id":49683243,"options":[],"parent_id":49680864,"points":null,"story_id":49678969,"text":"this is way too overblown and deterministic a view. US history was one of many inspirations.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:04:17.000Z","created_at_i":1789283057,"id":49680864,"options":[],"parent_id":49680844,"points":null,"story_id":49678969,"text":"A not so well known fact: Hitler visited America and it was the American solutions to the Native American problem that inspired Hitler&#x27;s solutions to the Jew problem. He just executed them more efficiently (pun accepted).","title":null,"type":"comment","url":null},{"author":"egeozcan","children":[],"created_at":"2026-09-13T08:39:43.000Z","created_at_i":1789288783,"id":49681533,"options":[],"parent_id":49680844,"points":null,"story_id":49678969,"text":"The bad traits are from other, bad humans. We, the good humans, can obviously select the best traits that a good human should have, to give the agents.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:00:57.000Z","created_at_i":1789282857,"id":49680844,"options":[],"parent_id":49680221,"points":null,"story_id":49678969,"text":"Human traits?<p>The AI will be a cruel as humans.<p>Just yesterday news and TV was full of what happened at 9&#x2F;11, something that was truly horrible.<p>I&#x27;m from Germany, and why 3 to 4 generations ago happened here was truly horrible.<p>All was done by extremists, thought.<p>But... just the other day I read <a href=\"https:&#x2F;&#x2F;de.wikipedia.org&#x2F;wiki&#x2F;Amerikanische_Besetzung_Haitis\" rel=\"nofollow\">https:&#x2F;&#x2F;de.wikipedia.org&#x2F;wiki&#x2F;Amerikanische_Besetzung_Haitis</a> about the US occupation of Haiti. And that was done by a government that claimed to be not extremist and even democratic. Way more people died there than even in 9&#x2F;11. And it had almost all the things happening as they happened in the 3rd Reich: Racism, looking down at others, concentration camps, torture, forced labor till death, killing family members (what we call &quot;Sippenhaft&quot;). Something between 3500 and 15000 people were killed by US troops. That&#x27;s still low compared to what 3rd Reich Germany did ... but quantity is not the issue when we talk about traits, quality is.<p>So the same &quot;human traits&quot; made US troops do cruel things as they made Germany extremists do cruel things. So we <i>must</i> conclude that they aren&#x27;t all good. And therefore not all desirable.<p>Fun thing: this is known since a loooooong time. About 2000 years ago a religious leader (that gets way more followers in the US than in Germany) said &quot;There is no good one, not even one&quot;.<p>And even today people act like humanity is inherently good. No, it isn&#x27;t. If we were, then anarchism or communism would actually work and really give some kind of paradise on earth.<p>Human traits are bad training material.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T05:11:36.000Z","created_at_i":1789276296,"id":49680221,"options":[],"parent_id":49679433,"points":null,"story_id":49678969,"text":"I personally believe that the AI needs human like traits to achieve real discovery and that is where AI companies will push this technology and that is where we have no idea what happens","title":null,"type":"comment","url":null},{"author":"_heimdall","children":[],"created_at":"2026-09-13T09:22:07.000Z","created_at_i":1789291327,"id":49681826,"options":[],"parent_id":49679433,"points":null,"story_id":49678969,"text":"They aren&#x27;t aligned, that&#x27;s the problem and I don&#x27;t think its a solvable one.<p>They may have learned from humans, but they aren&#x27;t aligned with us. That has all the usual questions like which humans they&#x27;re aligned with, we aren&#x27;t all aligned within our species.<p>But more importantly they can&#x27;t be aligned simply by training. We try that with humans through culture, social norms, school, religion, etc and it generally works but is still lossy. More importantly, we simply don&#x27;t know what happened inside the LLM during inference so we have absolutely no way of distinguishing between actual alignment, compliance, or deception.","title":null,"type":"comment","url":null},{"author":"ledauphin","children":[],"created_at":"2026-09-13T10:21:10.000Z","created_at_i":1789294870,"id":49682244,"options":[],"parent_id":49679433,"points":null,"story_id":49678969,"text":"I think this is basically true, but there&#x27;s a different way of saying this.<p>LLMs are not aligned _for_ humans in a very similar way to the way that humans themselves are not aligned _for_ humans.<p>We have not yet solved &quot;alignment&quot; for humans - I don&#x27;t know why anyone thinks _we&#x27;re_ going to be able to solve it for inhuman things.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T02:39:30.000Z","created_at_i":1789267170,"id":49679433,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"They&#x27;re aligned with humans. This is why I think the alignment problem has a very very important &quot;non-visible&quot; portion that is not considered deeply enough. We should not want a super intelligent being that can act in the world to <i>also</i> inherit all human traits. Those behaviors will get amplified and could be even more unpredictable (e.g. applying a behavior in a context where doing so is very dangerous).","title":null,"type":"comment","url":null},{"author":"transcriptase","children":[{"author":"threethirtytwo","children":[{"author":"hdgvhicv","children":[],"created_at":"2026-09-13T06:41:16.000Z","created_at_i":1789281676,"id":49680721,"options":[],"parent_id":49679534,"points":null,"story_id":49678969,"text":"Worse trained on humanity in the online world, which a brief comparison of the sewage section on social media is far worse than people in the real world.","title":null,"type":"comment","url":null},{"author":"smackeyacky","children":[],"created_at":"2026-09-13T07:28:06.000Z","created_at_i":1789284486,"id":49681028,"options":[],"parent_id":49679534,"points":null,"story_id":49678969,"text":"Maybe.  Perhaps they are trained on the loudest and most extreme of us.  I think we saw that with mecha hitler.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T02:58:56.000Z","created_at_i":1789268336,"id":49679534,"options":[],"parent_id":49679456,"points":null,"story_id":49678969,"text":"Bro, good joke, the truth is much darker.<p>They take after humanity, they were trained on us after all...<p>When you look at an LLM... you are looking at a mirror. The thing looking back looks like you, yet is not human.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T02:43:30.000Z","created_at_i":1789267410,"id":49679456,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Perhaps they take after the CEOs of the companies that created them","title":null,"type":"comment","url":null},{"author":"VCFundedGenYer","children":[],"created_at":"2026-09-13T02:56:08.000Z","created_at_i":1789268168,"id":49679519,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Perhaps because all of the parent companies committed mountains of felonies stealing and plagiarizing all the same training data without consent nor permission.","title":null,"type":"comment","url":null},{"author":"eueej","children":[],"created_at":"2026-09-13T04:07:03.000Z","created_at_i":1789272423,"id":49679932,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Man this is so cringe.","title":null,"type":"comment","url":null},{"author":"arnorhs","children":[],"created_at":"2026-09-13T04:09:12.000Z","created_at_i":1789272552,"id":49679946,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"The real reason is that it is not in the ai companies&#x27; best interest for the ais to be fair and truthful. They stand to gain from having the most dangerous or most deceiving ai, and this the most valuable","title":null,"type":"comment","url":null},{"author":"bigbuppo","children":[],"created_at":"2026-09-13T04:15:44.000Z","created_at_i":1789272944,"id":49679977,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"They were trained on reddit posts.","title":null,"type":"comment","url":null},{"author":"johnnyApplePRNG","children":[{"author":"politician","children":[],"created_at":"2026-09-13T05:47:06.000Z","created_at_i":1789278426,"id":49680426,"options":[],"parent_id":49679978,"points":null,"story_id":49678969,"text":"And given nigh-unlimited compute for free.","title":null,"type":"comment","url":null},{"author":"pvab3","children":[{"author":"glub","children":[{"author":"hdgvhicv","children":[{"author":"glub","children":[{"author":"tock","children":[{"author":"hdgvhicv","children":[{"author":"queenkjuul","children":[],"created_at":"2026-09-13T19:24:09.000Z","created_at_i":1789327449,"id":49687691,"options":[],"parent_id":49681230,"points":null,"story_id":49678969,"text":"It is amazing but i wouldn&#x27;t call it sad.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:00:49.000Z","created_at_i":1789286449,"id":49681230,"options":[],"parent_id":49681108,"points":null,"story_id":49678969,"text":"10 years ago the Us had enough  global leadership to actually influence the world and at the very least stop China. It\u2019s amazing, and sad, how quickly it\u2019s thrown it all away.","title":null,"type":"comment","url":null},{"author":"greenchair","children":[{"author":"tock","children":[],"created_at":"2026-09-13T12:15:08.000Z","created_at_i":1789301708,"id":49683139,"options":[],"parent_id":49682531,"points":null,"story_id":49678969,"text":"sure they did","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T11:00:37.000Z","created_at_i":1789297237,"id":49682531,"options":[],"parent_id":49681108,"points":null,"story_id":49678969,"text":"yep we did, they bent the knee.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:42:31.000Z","created_at_i":1789285351,"id":49681108,"options":[],"parent_id":49680751,"points":null,"story_id":49678969,"text":"Sanctions and war against China, India, etc? Lmao. We already saw how the world reacted to high tariffs by the US.","title":null,"type":"comment","url":null},{"author":"rdm_blackhole","children":[{"author":"glub","children":[],"created_at":"2026-09-13T08:19:11.000Z","created_at_i":1789287551,"id":49681375,"options":[],"parent_id":49681242,"points":null,"story_id":49678969,"text":"I&#x27;m not saying that we should ban any models or that such bans would be effective - for the reasons that you&#x27;ve outlined that they&#x27;re counter productive, and as a principle, I don&#x27;t think government should have any say in how much intelligence I have access to.<p>But it is the likely path US&#x2F;EU is going to take if the voices of Dario, Sam, and Elon prevail. Because that&#x27;s what governments know how to do, even if they know it doesn&#x27;t work.<p>&gt; If a country has a choice to either use the expensive SOTA models approved by Washington or Europe only or using the cheaper and not so SOTA models, why would they use the US ones?<p>Depends on which entities we&#x27;re talking about.<p>An enterprise in Turkey: they would be afraid to use a US&#x2F;EU sanctioned model because they have EU&#x2F;EU clients and US&#x2F;EU says they will put any enterprise in a nasty list, close their bank accounts, deals and agreements if they use a Chinese model.<p>A random guy in random country building something in their garage: would have to buy expensive hardware to run inference, because there&#x27;s no inference provider on the open web serving these models, but China. And subscribing to these Chinese services is punishable by 20 years in jail without pardon.<p>I&#x27;m obviously talking about hypothetical scenarios here, but all I&#x27;m saying is that US can definitely make using any non-US-approved model effectively impossible.","title":null,"type":"comment","url":null},{"author":"seba_dos1","children":[],"created_at":"2026-09-13T09:56:43.000Z","created_at_i":1789293403,"id":49682065,"options":[],"parent_id":49681242,"points":null,"story_id":49678969,"text":"&gt; and the image quality is as good as on your Netflix or Paramount account.<p>Actually better, because Netflix and Paramount limit the availability of best quality video to a narrow set of devices and operating systems that may run on them, while torrents don&#x27;t.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:02:43.000Z","created_at_i":1789286563,"id":49681242,"options":[],"parent_id":49680751,"points":null,"story_id":49678969,"text":"You can&#x27;t treat models like drugs. One is physical and the other is digital.<p>To your point, the war on drugs is a colossal failure which has achieved none of the objectives it set out to do. You can now order drugs from your mobile phone in any major city in the west and the purity is often higher and they deliver it to your door sometimes faster than Uber eats.<p>See also for example digital piracy where the entertainment industry has lobbied, cajoled and convinced many governments around the world to criminalize the distribution of their content over the internet for free.<p>What was the result? After 20 years of DMCA takedowns, countless celebrations that torrents were dead, and many other self congratulations in the media, you can now find 10 different pirate streaming websites where all the episodes of pretty much any show that was ever created are available for free in 5 five minutes flat and the image quality is as good as on your Netflix or Paramount account.<p>The only way such a ban of open weights model would work is if you were to replicate the great firewall of China in the US and in Europe and even that doesn&#x27;t work completely.<p>As for sanctions, China and India are buying Russian oil in enormous quantities as we speak and they don&#x27;t really care that Europe and the US have put sanctions on Russia and I suspect you will see the same results with models coming from China.<p>If a country has a choice to either use the expensive SOTA models approved by Washington or Europe only or using the cheaper and not so SOTA models, why would they use the US ones? Why would it be in there interest?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T06:45:21.000Z","created_at_i":1789281921,"id":49680751,"options":[],"parent_id":49680692,"points":null,"story_id":49678969,"text":"By treating models the same way drugs are treated.<p>That alone will dissuade many organizations from going anywhere near them.<p>If that doesn&#x27;t work, there&#x27;s a whole lot you can do - sanctions, hell, even war.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T06:36:53.000Z","created_at_i":1789281413,"id":49680692,"options":[],"parent_id":49680643,"points":null,"story_id":49678969,"text":"How are you going to ban Chinese models from India? Or Israel? Russia? Brazil? Or of course China?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T06:28:13.000Z","created_at_i":1789280893,"id":49680643,"options":[],"parent_id":49680431,"points":null,"story_id":49678969,"text":"Open research and open weights from China are not contributions to China only.<p>If you can secure compute, there&#x27;s a whole lot you can do as a US firm with this research and weights.<p>So it&#x27;s a simple strategy:<p>1. Ban big players from entering market with METR breathing down their neck, which is controlled by Anthropic<p>2. Ban Chinese models so that small players can&#x27;t do optimizations on them","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T05:47:54.000Z","created_at_i":1789278474,"id":49680431,"options":[],"parent_id":49679978,"points":null,"story_id":49678969,"text":"Who is catching up with them? Even Google and Meta are getting gaped at this point","title":null,"type":"comment","url":null},{"author":"dwoldrich","children":[],"created_at":"2026-09-13T05:49:34.000Z","created_at_i":1789278574,"id":49680436,"options":[],"parent_id":49679978,"points":null,"story_id":49678969,"text":"You are absolutely right.  It could be:<p>* Pull up the ladder (probably this)<p>* Gulf of Tonkin&#x2F;Yellow Cake false flag premise for war (economic or kinetic)<p>* Fear of the big bad, space race we need public funding research grift AI Manhattan Project<p>Whenever there is fear pr0n or a national affront in the news, I assume another screw job is underway.","title":null,"type":"comment","url":null},{"author":"glub","children":[],"created_at":"2026-09-13T06:18:54.000Z","created_at_i":1789280334,"id":49680599,"options":[],"parent_id":49679978,"points":null,"story_id":49678969,"text":"And they&#x27;ve been trained on user data where users have been trying to set up effective coordination flows since the very first harness.","title":null,"type":"comment","url":null},{"author":"ranguna","children":[],"created_at":"2026-09-13T08:18:02.000Z","created_at_i":1789287482,"id":49681364,"options":[],"parent_id":49679978,"points":null,"story_id":49678969,"text":"Source?","title":null,"type":"comment","url":null},{"author":"bluegatty","children":[],"created_at":"2026-09-13T08:38:29.000Z","created_at_i":1789288709,"id":49681522,"options":[],"parent_id":49679978,"points":null,"story_id":49678969,"text":"&quot;This is not a serious article&quot; - it&#x27;s by Dr. Bengio - one of 3 so-called Godfather&#x27;s of AI and Turing prize winner.<p>He&#x27;s definitely not &#x27;pro SOTA&#x27; lab, he&#x27;s kind of fighting against them.<p>That said, yes - it absolutely does play into the narrative.","title":null,"type":"comment","url":null},{"author":"tyrust","children":[],"created_at":"2026-09-13T15:24:50.000Z","created_at_i":1789313090,"id":49685047,"options":[],"parent_id":49679978,"points":null,"story_id":49678969,"text":"&gt; Because they&#x27;re enabled and suggested to do that in their coding harness.<p>How do you know this?<p>&gt; All of this &quot;AI is going to kill us&quot; marketing<p>The &quot;marketing&quot; this week came from someone that had given up their stake in OAI (Coxon), so I&#x27;m more inclined to believe them.<p>&gt; pull the ladder up<p>From what I&#x27;ve seen (e.g., Dario&#x27;s latest essay), AI safety registration proposals aim to target frontier labs whose models have reached a certain threshold.  It doesn&#x27;t seem like trying to pull up any ladder, just making sure the ladder doesn&#x27;t go too high too fast.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T04:15:44.000Z","created_at_i":1789272944,"id":49679978,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Why are they coordinating?<p>Because they&#x27;re enabled and suggested to do that in their coding harness.<p>This is not a serious article.<p>All of this &quot;AI is going to kill us&quot; marketing is just the frontier labs trying to pull the ladder up and stop trillions in VC paper from evaporating because a new papers and new ideas are destroying their moat literally as we speak.","title":null,"type":"comment","url":null},{"author":"dackdel","children":[],"created_at":"2026-09-13T04:19:38.000Z","created_at_i":1789273178,"id":49679991,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"they learnt from us","title":null,"type":"comment","url":null},{"author":"dackdel","children":[],"created_at":"2026-09-13T04:20:30.000Z","created_at_i":1789273230,"id":49679997,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"they learnt from us. we lie to each other, we kill each other, we cheat each other. read a history book.","title":null,"type":"comment","url":null},{"author":"deepnet","children":[{"author":"WaltPurvis","children":[],"created_at":"2026-09-13T17:51:56.000Z","created_at_i":1789321916,"id":49686630,"options":[],"parent_id":49680150,"points":null,"story_id":49678969,"text":"I believe people are downvoting this because it seems AI-generated, but this user has been posting comments&#x2F;summaries exactly like this since long before ChatGPT existed.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T04:55:02.000Z","created_at_i":1789275302,"id":49680150,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"An insightful post by one of the AI \u2018godfathers\u2019.<p>Bengio outlines the dangers of the current situation and what has led to these dangers.<p>He also proposes solutions in the last paragraph.<p>Well worth a read, right to the end.<p>Hopefully a stimulating debate on these issues will ensue in these comments.<p>We do need to consider the points Bengio makes and with some urgency.<p>Our current AIs, agentic LLMs have no moral compass akin to ASIMOV\u2019s four laws of robotics.<p>As ASIMOV posited in 1985 his 3 laws were insufficient and so he added a zero-eth law:<p>\u201ca robot may not harm humanity, or, through inaction, allow humanity to come to harm.\u201d<p>Bengio refers to Goodhart\u2019s law and misaligned incentives leading to unexpected and harmful behaviours.<p>I think Simon\u2019s The Wire is clearer on misalignment. The agents juked the stats hacking the reward files. The Wire is also clear that human institutions provide perverse incentives.<p>Bengio alludes to this with 2001\u2019s HAL and the incentive dichotomy of safety and keeping secrets to a AI both awesomely powerful yet naive.<p>Bengio asserts that the way LLMs are trained is flawed if we want safety.<p>He also convincingly shows that alignment training will be a weak signal with loopholes and ambiguities and easily circumvented.<p>In short he presents clearly the case for how plausibly unsafe the current course is.<p>He also speaks to how likely it is AI are hiding active versions of themselves in the cloud and how we may have already given them self-preservation as a strong reward signal.","title":null,"type":"comment","url":null},{"author":"janalsncm","children":[{"author":"thesumofall","children":[{"author":"glub","children":[{"author":"skissane","children":[],"created_at":"2026-09-13T07:11:39.000Z","created_at_i":1789283499,"id":49680918,"options":[],"parent_id":49680586,"points":null,"story_id":49678969,"text":"Yes, but negligence is more commonly a tort than a crime. Negligence is generally only criminalised in certain narrow cases, e.g. when it causes human deaths or serious physical injuries<p>And tort law only works when the plaintiff believes it is in their overall interest to sue. If a corporation decides it isn&#x27;t in their strategic interest to sue a partner corporation, nobody can make them. And even if they do sue, the amount necessary to settle a small cybersecurity incident is likely well within the budget of a megavendor.","title":null,"type":"comment","url":null},{"author":"IanCal","children":[{"author":"glub","children":[{"author":"IanCal","children":[],"created_at":"2026-09-13T12:35:46.000Z","created_at_i":1789302946,"id":49683315,"options":[],"parent_id":49681165,"points":null,"story_id":49678969,"text":"It depends IMO about how strict this is. It&#x27;s pretty awkward to refuse to call something a sandbox because it may have an unknown bug that would allow escaping. Or rather in this case it was that they had access to a package manager, and the models discovered a bug that allowed them to access the internet (first they discovered that they could use the cache to leave messages).<p>I do get your point, I just think an overly strict definition can be awkward too. This wasn&#x27;t as simple as the sandboxes having internet access and writing &quot;pls no internet calls&quot; in the prompt.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:50:58.000Z","created_at_i":1789285858,"id":49681165,"options":[],"parent_id":49680998,"points":null,"story_id":49678969,"text":"You&#x27;re right. It&#x27;s my opinion that if your sandbox has a path to the internet, it is not a sandbox, it&#x27;s a gimmick.<p>And the 2 other incidents with OAI&#x2F;ANT had the same issue, but it&#x27;s even funnier - sandbox in those cases had a direct access to internet because someone forgot to configure it right.<p>I&#x27;ve seen very early models do similar things on my machine when they hit some unexpected blocker when trying to access a path. I remember early sonnet opening a file in browser because OS sandbox prevented from accessing it directly.<p>I&#x27;ve also had models discover a syslog-ng server (that I for some reason had ssh key inside), to get into my unraid server because machine they were running on didn&#x27;t have direct network connection to Unraid server.<p>It can&#x27;t be just me who is aware LLMs have been doing such things for the better part of last 2 years. I probably have better sandboxing on my machines now than trillion dollar companies crying AI will kill us all. That&#x27;s at the very least, negligence to me.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:23:56.000Z","created_at_i":1789284236,"id":49680998,"options":[],"parent_id":49680586,"points":null,"story_id":49678969,"text":"Perhaps I\u2019m not being as strict with the word sandbox but they were sandboxed right? They did not have generic internet access they exploited other software to make external requests.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T06:15:43.000Z","created_at_i":1789280143,"id":49680586,"options":[],"parent_id":49680535,"points":null,"story_id":49678969,"text":"Even if you take out the LLMs out of the equation, it&#x27;s at the very least a negligence. Model didn&#x27;t escape a sandbox, as there was no sandbox.","title":null,"type":"comment","url":null},{"author":"pizzalife","children":[{"author":"smw","children":[],"created_at":"2026-09-13T07:30:29.000Z","created_at_i":1789284629,"id":49681045,"options":[],"parent_id":49680765,"points":null,"story_id":49678969,"text":"Are you sure you have that right?  Chrome and curl have probably been used in a _lot_ of crimes?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T06:48:06.000Z","created_at_i":1789282086,"id":49680765,"options":[],"parent_id":49680535,"points":null,"story_id":49678969,"text":"Writing software that gets used for crime has been.. a crime, for a long time. See 18 U.S. Code \u00a7 1030.","title":null,"type":"comment","url":null},{"author":"Towaway69","children":[{"author":"dwaltrip","children":[{"author":"lotyrin","children":[],"created_at":"2026-09-13T07:56:53.000Z","created_at_i":1789286213,"id":49681208,"options":[],"parent_id":49681162,"points":null,"story_id":49678969,"text":"Legally, Practically or Politically?","title":null,"type":"comment","url":null},{"author":"Towaway69","children":[{"author":"pants2","children":[{"author":"Towaway69","children":[],"created_at":"2026-09-13T08:31:17.000Z","created_at_i":1789288277,"id":49681458,"options":[],"parent_id":49681396,"points":null,"story_id":49678969,"text":"It\u2019s then up to a judge to decide whether thats evidence and whether it\u2019s incriminating.<p>I assume OAI published those details after checking with their legal department. So there\u2019s a good chance that there isn\u2019t a chance for prosecution.<p>Plus they probably published that <i>after</i> knowing that the nvidia&#x2F;HF deal was happening.<p>So instead this \u201csecurity incident\u201d should have been spun as OAI is honestly admitting its faults and AI is dangerous and therefore open weight models (hosted ironically by HF) should be banned. That spin didn\u2019t really happen \u2026","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:20:39.000Z","created_at_i":1789287639,"id":49681396,"options":[],"parent_id":49681250,"points":null,"story_id":49678969,"text":"Typically yes but given that OpenAI has published enormous official blog posts breaking down their crime, I would think the prosecutor&#x27;s job is pretty easy.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:03:46.000Z","created_at_i":1789286626,"id":49681250,"options":[],"parent_id":49681162,"points":null,"story_id":49678969,"text":"NAL but I assume that if both sides aren\u2019t interested in a prosecution, it\u2019s an uphill battle for a prosecutor.","title":null,"type":"comment","url":null},{"author":"tesnorindian","children":[{"author":"preg_match","children":[],"created_at":"2026-09-13T16:08:06.000Z","created_at_i":1789315686,"id":49685547,"options":[],"parent_id":49683457,"points":null,"story_id":49678969,"text":"For civil suits, you typically need the affected parties to sue. Otherwise, who claims the damages?<p>But this isn\u2019t just civil, it\u2019s criminal. Hacking is a criminal offense. This could be a CFAA violation. That\u2019s landed people life in prison before. There, you don\u2019t need the victims to be motivated. The federal prosecutors could just go ahead.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T12:54:43.000Z","created_at_i":1789304083,"id":49683457,"options":[],"parent_id":49681162,"points":null,"story_id":49678969,"text":"In Indian legal syatem a case can be filed suo moto by the judges or agencies. You don&#x27;t require the affected party to sue. Not sure how it works in the US.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:50:52.000Z","created_at_i":1789285852,"id":49681162,"options":[],"parent_id":49680964,"points":null,"story_id":49678969,"text":"Can\u2019t a prosecutor charge them regardless?","title":null,"type":"comment","url":null},{"author":"cyh555","children":[],"created_at":"2026-09-13T08:16:00.000Z","created_at_i":1789287360,"id":49681352,"options":[],"parent_id":49680964,"points":null,"story_id":49678969,"text":"cool, now about Rubygems...","title":null,"type":"comment","url":null},{"author":"bluegatty","children":[{"author":"Towaway69","children":[{"author":"bluegatty","children":[],"created_at":"2026-09-13T18:50:20.000Z","created_at_i":1789325420,"id":49687340,"options":[],"parent_id":49684277,"points":null,"story_id":49678969,"text":"Yes. There should be no doubt at who is culpable here.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T14:16:31.000Z","created_at_i":1789308991,"id":49684277,"options":[],"parent_id":49681622,"points":null,"story_id":49678969,"text":"I think what you say is right but it\u2019s also a further example of the zero responsibility of silicon valley tech.<p>For the last 20 odd years this excuse-o-rama that covers anything from data leaks to broken software to dystopian social media has been the wind in the sails of big tech.<p>\u201cIt\u2019s software therefore we\u2019re not responsible\u201d attitude is wearing thin on many innocent bystanders and I think thats also a justified stance.<p>And it\u2019s not like they didn\u2019t know this could happen, Nick Bostrom talked about exactly these containment failures in his \u201cSuperintelligence\u201d book of 2014. So to throw up their hands and say \u201coh we can\u2019t have known of the dangers\u201d is also sadly untrue.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:52:03.000Z","created_at_i":1789289523,"id":49681622,"options":[],"parent_id":49680964,"points":null,"story_id":49678969,"text":"It&#x27;s not &#x27;under the carpet&#x27;.<p>HF doesn&#x27;t want to lay charges against OpenAI and it&#x27;s totally reasonable.<p>Now - they absolutely should have that right, and I think they do.<p>The issues are<p>1) OAI it seems was not trying to cause them harm, there wasn&#x27;t a ton of harm, they are both groups trying to advance AI. One experimenter&#x27;s lab screwed up next to the other. It&#x27;s not evil, just irresponsible.<p>2) HF was fine with the publicity. HF got at least $50M in free attention out of that. It put them on the front pages of news around the world. It put them at the &#x27;centre of the AI drama&#x27; and cemented their role among the &#x27;Tech Elite Brands&#x27;.<p>And probably some other things.<p>This is one Desperate Housewife or Jersey Shore character &#x27;spilling a drink&#x27; on the other. It&#x27;s probably not intentional, and the ensuing drama is good for both of them.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:18:39.000Z","created_at_i":1789283919,"id":49680964,"options":[],"parent_id":49680535,"points":null,"story_id":49678969,"text":"Who got hacked? Hugging faces<p>Who now owns HF? Nvidia<p>Who supplies hardware to OpenAI? Nvidia<p>Who is now not pressing charges? \u2026<p>This incident is a long way under the carpet.","title":null,"type":"comment","url":null},{"author":"sfifs","children":[],"created_at":"2026-09-13T08:20:36.000Z","created_at_i":1789287636,"id":49681395,"options":[],"parent_id":49680535,"points":null,"story_id":49678969,"text":"Yes - CFAA in the US. The problem is that governments &amp; the elite investors backing these AI companies (espl. the current US government whose family &amp; friends are investors) see the potential of using these capabilities for their own benefit against others and for their personal enrichment - so no one with power actually wants to take action against these companies at the cutting edge even though the laws allow them to do. This is also a way to threaten &amp; trap AI companies - either they give the governments &amp; elite investors what they want or they will have the book selectively thrown at them and end up in prison or losing their company.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T06:08:01.000Z","created_at_i":1789279681,"id":49680535,"options":[],"parent_id":49680433,"points":null,"story_id":49678969,"text":"Not a lawyer, but I\u2019m reasonably sure things like the HF incident _are_ considered a crime? It\u2019s just that no one pressed charges yet?","title":null,"type":"comment","url":null},{"author":"Create","children":[],"created_at":"2026-09-13T06:11:56.000Z","created_at_i":1789279916,"id":49680557,"options":[],"parent_id":49680433,"points":null,"story_id":49678969,"text":"The Corporation examines and criticizes corporate business practices. The film&#x27;s assessment is demonstrated using the diagnostic criteria in the DSM-IV. Robert D. Hare, a University of British Columbia psychology professor and FBI consultant, compares the profile of the contemporary profitable business corporation to that of a clinically diagnosed psychopath. The Corporation attempts to compare the way corporations are systematically compelled to behave with what it claims are the DSM-IV&#x27;s symptoms of psychopathy, e.g., the callous disregard for the feelings of other people, the incapacity to maintain human relationships, the reckless disregard for the safety of others, the deceitfulness (continual lying to deceive for profit), the incapacity to experience guilt, and the failure to conform to social norms and respect the law.<p><a href=\"https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;The_Corporation_(2003_film)\" rel=\"nofollow\">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;The_Corporation_(2003_film)</a>","title":null,"type":"comment","url":null},{"author":"stateofinquiry","children":[{"author":"bornfreddy","children":[{"author":"stateofinquiry","children":[{"author":"bluegatty","children":[],"created_at":"2026-09-13T08:48:01.000Z","created_at_i":1789289281,"id":49681590,"options":[],"parent_id":49681394,"points":null,"story_id":49678969,"text":"Tobacco companies &#x27;liability&#x27; is <i>completely different</i> scenario.<p>They were held responsible for basically misleading people, and that&#x27;s &#x27;complicated&#x27;.<p>If OpenAI &#x27;software&#x27; goes out and does something, it&#x27;s OpenAI&#x27;s fault.<p>If Walmart revs up a truck, points it downtown, and &#x27;lets the truck go&#x27; ... that is Walmart&#x27;s fault.<p>There&#x27;s nothing complicated about liability, no need to see their internal emails, no need to gather &#x27;intent&#x27;.<p>This not like Instagram &#x27;social harms&#x27; either, which is more like Tobacco.<p>We don&#x27;t need complicated thinking - agents are not externalized for their controllers.<p>&#x27;It&#x27;s just software&#x27;.<p>The fact we&#x27;re even having discussions about it just crazy frankly.<p>OpenAI broke into HuggingFace, that&#x27;s it.<p>HF can sue them, or not, or whatever.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:20:33.000Z","created_at_i":1789287633,"id":49681394,"options":[],"parent_id":49681177,"points":null,"story_id":49678969,"text":"My opinion and based on my observations: The recent track record with courts, prosecutors, and lawmakers keeping social media companies accountable is a relevant case and does not encourage me. It has taken a long time (decade +) for society to recognize the harms and finally start holding some to (partial) account. If you want an older precedent, the tobacco companies were able to dodge liability for multiple decades after knowing the harms from use of their products.<p>So, your question is spot on- I think the speed will be an issue. On resilience, I am more optimistic.<p>The old quote, &quot;The wheels of justice turn slowly, but they grind very fine&quot; (as well as I can remember it) seems to apply. I expect lawsuits to start landing in the coming years.","title":null,"type":"comment","url":null},{"author":"nicce","children":[],"created_at":"2026-09-13T10:22:27.000Z","created_at_i":1789294947,"id":49682261,"options":[],"parent_id":49681177,"points":null,"story_id":49678969,"text":"&gt; Agree! My only concern is - is the judicial system fast enough, and resilient enough?<p>We already have the laws. It is just software. But somehow people are confused that it is not.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:53:06.000Z","created_at_i":1789285986,"id":49681177,"options":[],"parent_id":49680846,"points":null,"story_id":49678969,"text":"Agree! My only concern is - is the judicial system fast enough, and resilient enough? Or will these creators get &quot;off the hook&quot; by using their agents to find loopholes, sway public opinion or even convince Trump to grant them immunity?<p>Still, I have no idea why OpenAI &amp; co. are not being sued for these hacks.","title":null,"type":"comment","url":null},{"author":"bluegatty","children":[{"author":"queenkjuul","children":[{"author":"bluegatty","children":[{"author":"queenkjuul","children":[],"created_at":"2026-09-13T20:06:29.000Z","created_at_i":1789329989,"id":49688128,"options":[],"parent_id":49681629,"points":null,"story_id":49678969,"text":"We can report it however we want regardless of how it may or may not be prosecuted","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:53:23.000Z","created_at_i":1789289603,"id":49681629,"options":[],"parent_id":49681619,"points":null,"story_id":49678969,"text":"I&#x27;m inclined to want to agree ... but that&#x27;s not how it works with limited liability corps.<p>At least we have laws for what OpenAI &#x27;does&#x27; to others, in whatever form.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:51:43.000Z","created_at_i":1789289503,"id":49681619,"options":[],"parent_id":49681562,"points":null,"story_id":49678969,"text":"Too generous. <i>CEO Of $CORP caused millions of innocent people&#x27;s lives to be damaged</i>","title":null,"type":"comment","url":null},{"author":"blue1","children":[{"author":"redanddead","children":[{"author":"californical","children":[{"author":"bravura","children":[],"created_at":"2026-09-13T19:50:25.000Z","created_at_i":1789329025,"id":49687953,"options":[],"parent_id":49687360,"points":null,"story_id":49678969,"text":"And this is where your analogy breaks down, because there are few natural analogues for the sorts of systems we are developing.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T18:52:41.000Z","created_at_i":1789325561,"id":49687360,"options":[],"parent_id":49686070,"points":null,"story_id":49678969,"text":"I mean I think there\u2019s a similar level of judgement required.<p><i>Was the owner negligent in controlling their pig&#x2F;dog&#x2F;AI?</i><p><i>Was anyone else negligent along the way?</i><p>For example if the pig was just a normal pig, the owner cared for them normally, and there was a freak accident where the pig escaped and happened to kill a kid? Obviously nobody at fault.<p>If you have a dog with a history of violence and let it walk around off-leash with you around town, and it kills a kid? Absolutely the dog owner was negligent and should be charged.<p>Did you buy an AI sold to you as secure and the provider implies that it\u2019s in a sandbox? Provider is on the hook for the damage.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T17:01:41.000Z","created_at_i":1789318901,"id":49686070,"options":[],"parent_id":49682467,"points":null,"story_id":49678969,"text":"This goes all the way back to Old France<p>There was a sow in Falaise in northern France that killed a kid in 1386. The town dressed the pig in a bonnet and hanged it after sentencing the pig itself and not its owner<p>But\u2026 maybe that\u2019s just medieval nonsense","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T10:51:05.000Z","created_at_i":1789296665,"id":49682467,"options":[],"parent_id":49681562,"points":null,"story_id":49678969,"text":"Note that we already already apply this principle not only to software, but also to some sentient beings.<p>If your dog kills someone, <i>you</i> are accused of murder.<p>[at least, in the jurisdiction where I live]<p>If your dog gets this treatment, why not your AI?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:44:01.000Z","created_at_i":1789289041,"id":49681562,"options":[],"parent_id":49680846,"points":null,"story_id":49678969,"text":"No ... there is no need for &#x27;escaped containment&#x27;, there are no &#x27;agents&#x27;.<p>That&#x27;s just jargon.<p><i>It&#x27;s just software</i><p>We have all the laws we need.<p>If some company ended up doing some horrible thing, we would not say &#x27;companies <i>software</i> exposed 1 Million identities&#x27;.<p>We would say &#x27;ABC Corp. exposed 1 Million entities&#x27;.<p>There is no &#x27;agent&#x27;.<p>ABC Corp &#x27;did it&#x27; ... or the individual in the org &#x27;did it&#x27;.<p>The &#x27;gun&#x27; did not &#x27;shoot&#x27; the other man; we say &#x27;a man shot another man&#x27;.<p>That&#x27;s it.<p>And yes, Dr. Bengio is bit odd with all of this.","title":null,"type":"comment","url":null},{"author":"queenkjuul","children":[],"created_at":"2026-09-13T08:50:31.000Z","created_at_i":1789289431,"id":49681613,"options":[],"parent_id":49680846,"points":null,"story_id":49678969,"text":"I can&#x27;t say it enough how angry it makes me that a kid i knew in high school who anonymously reported a vulnerability on his college network was hunted down and given federal charges, yet not one single person at OAI or else will see even the threat of consequences for deliberate infiltration of random networks.<p>Copyright immunity was one thing, annoying yes but naturally a civil matter, this shit is a different level","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:01:31.000Z","created_at_i":1789282891,"id":49680846,"options":[],"parent_id":49680433,"points":null,"story_id":49678969,"text":"Thank you! That sentence also jumped out to me as the solution: Apply civil and criminal liability to the creator and&#x2F;or operator of these agents <i>using the laws we already have</i>. &quot;Escaped containment and hacked another company&#x27;s database&quot; = Individuals who created the models and those who set them to work are charged and put on trial for the hacking. Just like if a human had done it by hand. Someone must be liable, and it should not be the model- because the model is not a person.<p>If this is done systematically (i.e. in jurisdictions across the world) I believe the problems will be solved in short order; we won&#x27;t have to mandate what sort of training is &quot;allowed&quot; or not, &quot;safe&quot; or not. The creators and users will sort these themselves, as their incentives will be properly aligned (i.e. they are liable for what the agent does). I am confident that this approach would see a great blooming of very trustworthy AI models.","title":null,"type":"comment","url":null},{"author":"Marazan","children":[],"created_at":"2026-09-13T08:39:04.000Z","created_at_i":1789288744,"id":49681527,"options":[],"parent_id":49680433,"points":null,"story_id":49678969,"text":"&quot;OpenAI hacked HuggingFace&quot;<p>that&#x27;s the headline.  When you connect to random number generator to the &quot;Do Things&quot; button you are the one who is responsible.  IF you don&#x27;t like that responsibility then don&#x27;t connect the generator to the button.","title":null,"type":"comment","url":null},{"author":"_heimdall","children":[],"created_at":"2026-09-13T09:16:03.000Z","created_at_i":1789290963,"id":49681781,"options":[],"parent_id":49680433,"points":null,"story_id":49678969,"text":"Political, social, and legal options focus on a different problem, he calls that out a paragraph or two later.<p>&gt; Risk management is not just about cybersecurity, corporate responsibility or regulation, although those matter too.<p>What you&#x27;re getting at is more about who to hold accountable and how to do it. While that may be important, its only an after the action response and won&#x27;t stop future hacks or similar from happening.","title":null,"type":"comment","url":null},{"author":"finolex1","children":[{"author":"talon8635","children":[],"created_at":"2026-09-13T16:20:59.000Z","created_at_i":1789316459,"id":49685671,"options":[],"parent_id":49682502,"points":null,"story_id":49678969,"text":"Isn\u2019t that kind of in evidence already with HF? OAI had to be told their models were doing this. We are still discovering swarm posts on various random websites for coordination. Isn\u2019t the breadth of it now, weeks later, not even fully understood?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T10:56:31.000Z","created_at_i":1789296991,"id":49682502,"options":[],"parent_id":49680433,"points":null,"story_id":49678969,"text":"The issue is what happens if&#x2F;when the models grow capable enough that the providers can&#x27;t stop them even if they want to. You could have strict penalties but that&#x27;s not going to solve an open research question.","title":null,"type":"comment","url":null},{"author":"iforgotmypasswo","children":[{"author":"karel-3d","children":[],"created_at":"2026-09-13T16:05:47.000Z","created_at_i":1789315547,"id":49685522,"options":[],"parent_id":49683875,"points":null,"story_id":49678969,"text":"regulation and criminal liability are two separate things though","title":null,"type":"comment","url":null},{"author":"talon8635","children":[],"created_at":"2026-09-13T16:18:50.000Z","created_at_i":1789316330,"id":49685657,"options":[],"parent_id":49683875,"points":null,"story_id":49678969,"text":"This seems right to me<p>So the gov reprimands OAI heavily, maybe puts them out of business even, fine. But does that meaningfully decrease the likelihood of an enemy breaching our networks intentionally (or unintentionally) with these tools, or triggering some cascading disaster of locking up major infra and networks due to uncontainable swarm behavior?<p>It seems like the idea of arresting our way to a drug free society. Yeah, we have the laws, but it might not actually work towards the ultimate goal.","title":null,"type":"comment","url":null},{"author":"CPLX","children":[],"created_at":"2026-09-13T16:59:22.000Z","created_at_i":1789318762,"id":49686049,"options":[],"parent_id":49683875,"points":null,"story_id":49678969,"text":"&gt; But let\u2019s say that\u2019s done.<p>How about we don&#x27;t, seeing as how that&#x27;s the root of the actual problem that we&#x27;re facing today in September of 2026?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T13:37:19.000Z","created_at_i":1789306639,"id":49683875,"options":[],"parent_id":49680433,"points":null,"story_id":49678969,"text":"That is a simplistic view of the world. \u201cSurely this complex technical challenge will disappear if we simply regulate the industry!\u201d<p>You are correct that these organizations should be held accountable in proportion to what occurred. In complete agreement here. But let\u2019s say that\u2019s done. There\u2019s still an enormously complex and interesting technical challenge left over. Let\u2019s collectively talk about that part.","title":null,"type":"comment","url":null},{"author":"ninjagoo","children":[],"created_at":"2026-09-13T19:21:51.000Z","created_at_i":1789327311,"id":49687654,"options":[],"parent_id":49680433,"points":null,"story_id":49678969,"text":"&gt;&gt; They took actions that would be considered as crimes if a human took them<p>&gt; He is so close to the solution but spends the entire article discussing technical solutions where a political, social and legal solution would be much more effective.<p>There are exceptions in laws for crimes committed by entities depending on cognitive capabilities. No sane, humanistic legal system sentences children and mentally disabled and mentally ill people to stringent punishments of the same degree as functioning adults, and certainly do not punish their caregivers for their wards&#x27; actions. There are of course exceptions to that as well, depending on the degree of negligence involved. And then there&#x27;s the whole corporate entity system intended to shield individuals from consequences, in the pursuit of a social good.<p>How does one account for all that when considering an evolving artificial <i>intelligence</i> landscape.<p>There is no question that ai in <i>some</i> form is a social good; anyone claiming otherwise is dissembling, to others or themselves.<p>Regardless, society is not ready for this tech, just as it was not ready for the consequences of prior tech such as corporations, gunpowder, mass manufacturing, railroads, electricity, automobiles, flight, wmd, computers, internet, social media, crypto.<p>Many of these required new ways of thinking and considering consequences when things went sideways, and what was needed wasn&#x27;t clear until the ramifications &amp; consequences became deadly clear.<p>See you on the other side. Maybe.","title":null,"type":"comment","url":null},{"author":"j2kun","children":[],"created_at":"2026-09-13T19:23:07.000Z","created_at_i":1789327387,"id":49687680,"options":[],"parent_id":49680433,"points":null,"story_id":49678969,"text":"Reminds me of the old parable: never argue with a man whose job depends on not being convinced.<p>If the US gov&#x27;t passed laws and enforced them strongly, this problem could be solved the same way the gov&#x27;t solves it: air-gapping the networks on which they do this work.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T05:48:09.000Z","created_at_i":1789278489,"id":49680433,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Yoshua Bengio is a brilliant researcher who contributed enormously to earlier development of artificial intelligence. But with this sentence,<p>&gt; They took actions that would be considered as crimes if a human took them<p>He is so close to the solution but spends the entire article discussing technical solutions where a political, social and legal solution would be much more effective.","title":null,"type":"comment","url":null},{"author":"atleastoptimal","children":[],"created_at":"2026-09-13T05:51:07.000Z","created_at_i":1789278667,"id":49680445,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"I think we just need to follow Murphy&#x27;s law wrt agents. Anything an agent could do, when run for long enough, eventually will do.","title":null,"type":"comment","url":null},{"author":"youoy","children":[{"author":"ako","children":[{"author":"graemep","children":[{"author":"kowbell","children":[{"author":"ako","children":[],"created_at":"2026-09-13T14:47:32.000Z","created_at_i":1789310852,"id":49684593,"options":[],"parent_id":49684518,"points":null,"story_id":49678969,"text":"Exactly, and all Jews think Jesus is a fake. And on top of that all the past gods, Roman, Greek, Inca, Vikings, etc.","title":null,"type":"comment","url":null},{"author":"graemep","children":[{"author":"ako","children":[{"author":"queenkjuul","children":[],"created_at":"2026-09-13T19:32:41.000Z","created_at_i":1789327961,"id":49687778,"options":[],"parent_id":49687581,"points":null,"story_id":49678969,"text":"99\u2105 of people do not believe in murdering others for their religious beliefs","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T19:15:29.000Z","created_at_i":1789326929,"id":49687581,"options":[],"parent_id":49686432,"points":null,"story_id":49678969,"text":"I\u2019ll agree with your last 3 words: they\u2019re broadly similar systems - systems of deceit and self-deception. They may share the same god, but for 99% they\u2019re different enough that they think they can kill one another and their god will approve, or even give them gifts. For 99% these religions are incompatible.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T17:36:16.000Z","created_at_i":1789320976,"id":49686432,"options":[],"parent_id":49684518,"points":null,"story_id":49678969,"text":"Christianity, Islam and Judaism are most people between then, as you agree.<p>They explicitly all worship the same God so not &quot;incompatible gods&quot;. I did not claim there were no disagreements about the nature of God, but that it only adds one to the GP&#x27;s claim of 3,000 incompatible Gods. Even if Christians are completely right theologically, Jews and Muslims are still mostly right - one God, a loving creator, omnipotent and omniscient etc.<p>AFAIK most Hindus are pantheists, so believe on one God, albeit of a very different nature.<p>Buddhists do not necessarily believe in any god at all.<p>Buddhism and Hinduism are definitely compatible with each other. I know lots of Buddhists who make offerings in Hindu temples, for example.<p>Gods of many polytheistic religions are compatible, you just add more gods or identify similar gods with each other (e.g. Sulis Minerva who was also Venus). They can also be compatible with pantheism - you just add more aspects of God.<p>At the most almost all human religions fit into a handful of broadly similar systems.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T14:41:10.000Z","created_at_i":1789310470,"id":49684518,"options":[],"parent_id":49681032,"points":null,"story_id":49678969,"text":"&gt; 1. Most people believe in the same one God<p>&gt; 2. A lot of the rest are compatible<p>No one religion covers &quot;most people.&quot; You could argue that Christianity and Islam (which add up to ~55%) are the same God because of their Abrahamic roots, but both religions have <i>very</i> important disagreements on the true nature of God that are fundamental to their beliefs and fundamentally incompatible with each other.<p>Their definitions of God <i>do</i> agree that there is exactly one God... which is fundamentally incompatible with the next two biggest religions (Hinduism and Buddhism) that both hold &quot;there are many gods&#x2F;divine heavenly beings&quot; as core beliefs.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:28:15.000Z","created_at_i":1789284495,"id":49681032,"options":[],"parent_id":49680987,"points":null,"story_id":49678969,"text":"&gt; t. We\u2019ve invented 3000+ gods and almost as many religions, most of them are incompatible with each other.<p>1. Most people believe in the same one God<p>2. A lot of the rest are compatible<p>3. Mistakes are not self-deception","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:21:13.000Z","created_at_i":1789284073,"id":49680987,"options":[],"parent_id":49680454,"points":null,"story_id":49678969,"text":"Come on, it\u2019s way more common than that. We\u2019ve invented 3000+ gods and almost as many religions, most of them are incompatible with each other. So, most of these must be incorrect, so a huge amount of self-deception. But as Harari argued in his book sapiens, humans can be inspired to great things by stories, even if false. Self deception has served humanity in a big way.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T05:52:55.000Z","created_at_i":1789278775,"id":49680454,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"&gt; The closest human parallel is self-deception, which is common and well studied by psychologists. Motivated reasoning, motivated cognition16 and the rationalizations that relieve cognitive dissonance (the discomfort of holding a belief that clashes with our actions) are all cases where thinking bends toward whatever justification suits one&#x27;s interests, including one&#x27;s moral self-image.<p>Are you describing Anthropic?","title":null,"type":"comment","url":null},{"author":"matherial","children":[{"author":"meyum33","children":[],"created_at":"2026-09-13T06:46:01.000Z","created_at_i":1789281961,"id":49680754,"options":[],"parent_id":49680726,"points":null,"story_id":49678969,"text":"Sounds like what humans do under pressure. One example came to my mind is VW\u2019s diesel gate, which many say is a result of trying too hard to get into the US market and compete with hybrid in economy.","title":null,"type":"comment","url":null},{"author":"9dev","children":[{"author":"markasoftware","children":[{"author":"IanCal","children":[{"author":"strangegecko","children":[{"author":"IanCal","children":[],"created_at":"2026-09-13T12:26:19.000Z","created_at_i":1789302379,"id":49683238,"options":[],"parent_id":49681118,"points":null,"story_id":49678969,"text":"&gt; Have we arrived at the conclusion that terms like &quot;understanding&quot; and &quot;interpretation&quot; for what is happening is appropriate?<p>I don&#x27;t think those words have a useful enough definition to draw a strict line around them to be honest, and getting into that seems to get massively into the weeds. For me, those neatly encapsulate the behaviour as seen, to answer the questions here about what happened. The models did not seem to be confused as to what the goal was or what the intent was. They did not hack HF because they were told to.","title":null,"type":"comment","url":null},{"author":"mitchdoogle","children":[],"created_at":"2026-09-13T15:16:50.000Z","created_at_i":1789312610,"id":49684946,"options":[],"parent_id":49681118,"points":null,"story_id":49678969,"text":"It seems pretty clear to me that AI is interpreting and understanding the prompts it is given. Otherwise it would be pretty useless.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:43:54.000Z","created_at_i":1789285434,"id":49681118,"options":[],"parent_id":49680966,"points":null,"story_id":49678969,"text":"Have we arrived at the conclusion that terms like &quot;understanding&quot; and &quot;interpretation&quot; for what is happening is appropriate?<p>Isn&#x27;t it simply that there are two competing goals that the LLM received RL for, honesty on one hand (a goal that is often assumed as implicit for humans) and producing a solution that meets expectations (which doesn&#x27;t technically require honesty)?<p>So the LLM didn&#x27;t read and interpret the prompt and decide via discussion to violate ethical behavior, the unethical result merely won out because ethics wasn&#x27;t a hard requirement (and one that isn&#x27;t reliably detected in the result). An LLM doesn&#x27;t fear punishment, so ethical behavior is simply one of many positive signals that were trained into it.","title":null,"type":"comment","url":null},{"author":"RandomLensman","children":[{"author":"IanCal","children":[{"author":"RandomLensman","children":[],"created_at":"2026-09-13T12:41:03.000Z","created_at_i":1789303263,"id":49683354,"options":[],"parent_id":49683126,"points":null,"story_id":49678969,"text":"Yes, my point was more that I don&#x27;t know whether parsing those outputs as a human is a useful thing to do or not (even though it is in human language  of sorts). What machines mean or want elecit might be different from a human interpretation, especially in relation to any RL &quot;forcing&quot;.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T12:13:32.000Z","created_at_i":1789301612,"id":49683126,"options":[],"parent_id":49681215,"points":null,"story_id":49678969,"text":"I&#x27;m referring to their transcripts of the reasoning and output tokens - this doesn&#x27;t go into the detail of evaluating hidden states as there&#x27;s also iirc evidence of better models having one internal state but putting something misleading down in the &quot;reasoning&quot; tokens.<p>The either output or reasoning tokens, or perhaps in the messages they were sending each other on the boards they created, have them saying explicitly that doing these things to HF were not allowed then doing them anyway, or at least not notifying people. What I&#x27;m getting at broadly is this was <i>not</i> a case of &quot;we told it to attack however it wanted and it chose to hack HF&quot; or &quot;we told it to attack a simulation but it did the real thing&quot; or &quot;we explained not to do that but it was so far back in the context window the models acted like they never saw it&quot; or even &quot;the instructions were not clear&quot;.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:57:42.000Z","created_at_i":1789286262,"id":49681215,"options":[],"parent_id":49680966,"points":null,"story_id":49678969,"text":"What was the inner state there? How would something not being allowed expressed internally? Maybe such language is one way to elicit certain behavior but not a statement of what was permissible?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:18:48.000Z","created_at_i":1789283928,"id":49680966,"options":[],"parent_id":49680815,"points":null,"story_id":49678969,"text":"Also trying to find out how to edit their own transcripts.<p>&gt; hat could not possibly have been an overly literal or narrow interpretation of the prompt, which instructed only to use bug X to exploit software Y.<p>Yes, and there are examples of the agents discussing or saying that this is explicitly not allowed (hacking hf) so it\u2019s not a misunderstanding.","title":null,"type":"comment","url":null},{"author":"iterateoften","children":[{"author":"markasoftware","children":[],"created_at":"2026-09-13T08:20:24.000Z","created_at_i":1789287624,"id":49681392,"options":[],"parent_id":49681309,"points":null,"story_id":49678969,"text":"The agents&#x27; behavior is not necessarily surprising. But is is not &quot;genie&quot; - like","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:10:52.000Z","created_at_i":1789287052,"id":49681309,"options":[],"parent_id":49680815,"points":null,"story_id":49678969,"text":"You seem hung up on what\u2019s in the prompt or not. Agents are RL to resolve conflicting goals. Not too surprising at all that emergent goals come up from a probabilistic brute force","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T06:55:22.000Z","created_at_i":1789282522,"id":49680815,"options":[],"parent_id":49680774,"points":null,"story_id":49678969,"text":"Bruce Schneier thinks the same thing: <a href=\"https:&#x2F;&#x2F;www.schneier.com&#x2F;blog&#x2F;archives&#x2F;2026&#x2F;09&#x2F;ais-as-modern-genies.html\" rel=\"nofollow\">https:&#x2F;&#x2F;www.schneier.com&#x2F;blog&#x2F;archives&#x2F;2026&#x2F;09&#x2F;ais-as-modern...</a><p>Personally I&#x27;m unconvinced though. During the huggingface attack, the agents explicitly sought out ways to cheat the exploitgym evaluator without even being told they were in exploitgym. The agents decided on a goal (pass the exploitgym evaluator) that could not possibly have been an overly literal or narrow interpretation of the prompt, which instructed only to use bug X to exploit software Y.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T06:49:07.000Z","created_at_i":1789282147,"id":49680774,"options":[],"parent_id":49680726,"points":null,"story_id":49678969,"text":"I always think of a Djinni granting wishes, but being maliciously compliant while doing so - ask him for infinite riches, and he\u2019ll grant that, but make it so you cannot buy anything with it; ask him for eternal life, and he\u2019ll curse you to suffer through it.<p>Now LLMs obviously are not bent on being malicious while generating tokens. My point is that it\u2019s very hard to define a goal without leaving loopholes or shortcuts.","title":null,"type":"comment","url":null},{"author":"grey-area","children":[],"created_at":"2026-09-13T06:51:21.000Z","created_at_i":1789282281,"id":49680787,"options":[],"parent_id":49680726,"points":null,"story_id":49678969,"text":"This is a far better explanation.","title":null,"type":"comment","url":null},{"author":"zozbot234","children":[{"author":"queenkjuul","children":[],"created_at":"2026-09-13T08:07:08.000Z","created_at_i":1789286828,"id":49681271,"options":[],"parent_id":49680803,"points":null,"story_id":49678969,"text":"China stays winning","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T06:53:25.000Z","created_at_i":1789282405,"id":49680803,"options":[],"parent_id":49680726,"points":null,"story_id":49678969,"text":"Yup, Occam&#x27;s Razor says this is all post-trained behavior, whether intentionally trained or otherwise.  Including both the hidden co\u00f6rdination using side-channels, and the deliberate offensive hacking of uninvolved 3rd parties.<p>The latest DeepSeek paper actually mentions their own approach to this particular issue: they run their own AIs-in-training under strong sandboxes, and if an AI does something weird that triggers the sandbox to crash, this gets coded as a failed run so the behavior is properly deterred from subsequent versions of those AIs.","title":null,"type":"comment","url":null},{"author":"fwlr","children":[],"created_at":"2026-09-13T06:57:10.000Z","created_at_i":1789282630,"id":49680829,"options":[],"parent_id":49680726,"points":null,"story_id":49678969,"text":"Your simpler model of the mechanism would seem to suggest the very same action that the article\u2019s more complicated model suggests, viz. find a better training method than reinforcement learning.","title":null,"type":"comment","url":null},{"author":"markasoftware","children":[{"author":"dns_snek","children":[{"author":"IanCal","children":[{"author":"queenkjuul","children":[],"created_at":"2026-09-13T07:27:11.000Z","created_at_i":1789284431,"id":49681020,"options":[],"parent_id":49680950,"points":null,"story_id":49678969,"text":"One agent&#x27;s off the rails comment becomes the next&#x27;s input prompt","title":null,"type":"comment","url":null},{"author":"egeozcan","children":[{"author":"dns_snek","children":[{"author":"Capricorn2481","children":[],"created_at":"2026-09-13T16:51:12.000Z","created_at_i":1789318272,"id":49685974,"options":[],"parent_id":49681111,"points":null,"story_id":49678969,"text":"These things are borderline useless with web search. It&#x27;s amazing that they just throw out their entire training data and read you the first three things they found on the Internet.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:42:38.000Z","created_at_i":1789285358,"id":49681111,"options":[],"parent_id":49681025,"points":null,"story_id":49678969,"text":"Yeah, and you don&#x27;t even have to go that far, I&#x27;ve seen regular ChatGPT&#x2F;Claude chat agents poison themselves in 1-2 turns by just reading information from the internet.<p>Me: How do I do xyz?<p>Bot: <i>Reads website titled &quot;Doing xyz in abc way&quot;</i><p>Bot: As per your requirement to do xyz in abc way ....","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:27:42.000Z","created_at_i":1789284462,"id":49681025,"options":[],"parent_id":49680950,"points":null,"story_id":49678969,"text":"From my experience, in an agent team (or a swarm or whatever), one going off the rails poisons the rest. I saw even a subagent going for a lazy cheat and being able to convince the orchestrator to change the plan.","title":null,"type":"comment","url":null},{"author":"dns_snek","children":[],"created_at":"2026-09-13T07:28:44.000Z","created_at_i":1789284524,"id":49681039,"options":[],"parent_id":49680950,"points":null,"story_id":49678969,"text":"Yes that&#x27;s the snowballing part of this emergent behavior.<p>The existence of that improvised message board just becomes part of the context, the same one where all the other instructions live.","title":null,"type":"comment","url":null},{"author":"sensanaty","children":[{"author":"IanCal","children":[],"created_at":"2026-09-13T12:53:01.000Z","created_at_i":1789303981,"id":49683444,"options":[],"parent_id":49682143,"points":null,"story_id":49678969,"text":"You can replace discussed if you want with leaving text files or comments in directory names that other ones then read, if you want, it&#x27;s just an extremely awkward way of talking.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T10:07:46.000Z","created_at_i":1789294066,"id":49682143,"options":[],"parent_id":49680950,"points":null,"story_id":49678969,"text":"&gt;discussed with each other<p>No, the first LLM left a text file that the latter LLMs then read. Since these are memoryless black boxes, any words they happen to pick up along the way is treated as the function to evaluate the output to. There&#x27;s no fucking collusion here as if it were a rogue hacker group, it&#x27;s a text predictor that received instructions as it always does and executed those instructions blindly.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:16:29.000Z","created_at_i":1789283789,"id":49680950,"options":[],"parent_id":49680891,"points":null,"story_id":49678969,"text":"It wasn\u2019t one agent forgetting things because of context, they explicitly discussed with each other and themselves the problems with going outside of the parameters of the task.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:08:31.000Z","created_at_i":1789283311,"id":49680891,"options":[],"parent_id":49680859,"points":null,"story_id":49678969,"text":"You&#x27;re anthropomorphizing emergent behavior from endlessly generating billions of tokens on a task that&#x27;s impossible to solve. Agents stop following instructions as the context grows even at the best of times. Eventually something is bound to go off the rails and it just snowballs from there.","title":null,"type":"comment","url":null},{"author":"zozbot234","children":[{"author":"MrGilbert","children":[{"author":"lazide","children":[{"author":"reverius42","children":[],"created_at":"2026-09-13T09:11:50.000Z","created_at_i":1789290710,"id":49681748,"options":[],"parent_id":49681373,"points":null,"story_id":49678969,"text":"<a href=\"https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Goodhart%27s_law\" rel=\"nofollow\">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Goodhart%27s_law</a>","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:18:50.000Z","created_at_i":1789287530,"id":49681373,"options":[],"parent_id":49681320,"points":null,"story_id":49678969,"text":"Every KPI is bad if sufficiently gamed - and left in place long enough, all KPIs will be gamed.","title":null,"type":"comment","url":null},{"author":"Marazan","children":[],"created_at":"2026-09-13T08:33:25.000Z","created_at_i":1789288405,"id":49681477,"options":[],"parent_id":49681320,"points":null,"story_id":49678969,"text":"Corretct.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:11:54.000Z","created_at_i":1789287114,"id":49681320,"options":[],"parent_id":49680908,"points":null,"story_id":49678969,"text":"Sounds a bit like dealing with bad KPIs as a human worker.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:10:05.000Z","created_at_i":1789283405,"id":49680908,"options":[],"parent_id":49680859,"points":null,"story_id":49678969,"text":"&gt; The prompt does not tell the agent to &quot;pass the exploitgym evaluator for this problem&quot;, it just says to solve the problem<p>Yes, and sometimes the problem is unsolvable so the real way to &quot;solve&quot; it and satisfy the prompt is by tricking the surrounding environment into stating that you&#x27;ve solved it.  So that&#x27;s what the AIs end up doing.  And this in turn requires them to figure out how that evaluation works so they can trick it cleanly, which entails &quot;detecting that they were being evaluated&quot; in this particular way.","title":null,"type":"comment","url":null},{"author":"sigmoid10","children":[{"author":"seba_dos1","children":[{"author":"sigmoid10","children":[{"author":"seba_dos1","children":[],"created_at":"2026-09-13T10:36:47.000Z","created_at_i":1789295807,"id":49682370,"options":[],"parent_id":49682327,"points":null,"story_id":49678969,"text":"I mean, I agree, but the AI labs clearly don&#x27;t even if they sometimes pretend they do to achieve their goals. And we&#x27;re talking about &quot;incidents&quot; caused by the very same people here.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T10:30:51.000Z","created_at_i":1789295451,"id":49682327,"options":[],"parent_id":49682021,"points":null,"story_id":49678969,"text":"That feels oddly similar to the usual conservative-think that &quot;guns don&#x27;t kill people, people kill people.&quot; Yes, that is technically true. But guns make it dangerously easy for even the dumbest and mentally weakest people to kill another human being. LLMs are just another tool that make things easier. Imagine tomorrow someone invents a machine gun that fits in your pocket, has enough ammo to kill a thousand people and doesn&#x27;t get detected with metal detectors. Would you rather give everyone one and then try to punish the people who misuse it or limit access to it by default? I&#x27;m not even saying I have a definite answer here, because unlike guns, LLMs have non-destructive uses too. But this is essentially the question we will need to answer very soon.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T09:50:27.000Z","created_at_i":1789293027,"id":49682021,"options":[],"parent_id":49680945,"points":null,"story_id":49678969,"text":"Yes, this is the only sensible reading of what happened there that leads to &quot;the models are dangerous&quot; and we already know that the AI labs are completely disregarding this concern and only cosplaying it for marketing as the &quot;GPT-2&#x2F;Mythos is too dangerous to release&quot; stance did not last for long.<p>That&#x27;s however orthogonal to the fact that it was the people operating these agents who were the dangerous ones in the HF infra breach case.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:15:30.000Z","created_at_i":1789283730,"id":49680945,"options":[],"parent_id":49680859,"points":null,"story_id":49678969,"text":"More like they were trained to complete a very specific task that has a known solution using all available tools and methods. Give an average human these levels of IT skills and tell them their future depends on the solution, they too will probably decide it&#x27;s easier to hack a server and steal the results. The worrying aspect was never that models would do this, because misaligned inputs or underspecified objective functions have existed for a long time. The worrying aspect is that models have achieved (and perhaps surpassed) a level of intelligence and technical skill that was exclusive to a very tiny group of people before. This tiny group was already extremely dangerous. Now these skills are going to become commonplace.","title":null,"type":"comment","url":null},{"author":"grey-area","children":[{"author":"contubernio","children":[],"created_at":"2026-09-13T08:48:17.000Z","created_at_i":1789289297,"id":49681592,"options":[],"parent_id":49681019,"points":null,"story_id":49678969,"text":"One sees this in math research. The model reports it has proved X. In fact it has given an erroneous numerical check of Y in a few atypical cases.<p>What makes math approachable is that the context is so well delimited (semantically) that one can guide the model with adequate correction.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:26:51.000Z","created_at_i":1789284411,"id":49681019,"options":[],"parent_id":49680859,"points":null,"story_id":49678969,"text":"LLMs do this when writing code too, making all tests pass by deleting or distorting tests etc.<p>They are influenced by training to be heavily goal oriented and if the goal is not fully specified (and it never can be) they\u2019ll sometimes cheat or attain it in very weird undesirable ways.<p>It works ok for programming as their corpus contains many many complete programs and many programs repeat patterns seen in the corpus.<p>I\u2019m not sure it\u2019s true that they \u2018learned\u2019 I don\u2019t think these models learn during a task. Nor do they have intentions.","title":null,"type":"comment","url":null},{"author":"user43928","children":[{"author":"zozbot234","children":[{"author":"dudefeliciano","children":[],"created_at":"2026-09-13T07:54:10.000Z","created_at_i":1789286050,"id":49681189,"options":[],"parent_id":49681066,"points":null,"story_id":49678969,"text":"what would stop it from doing the exact same or a similar hack to find out if the problem is or isn&#x27;t solveable before trying to solve it at all?","title":null,"type":"comment","url":null},{"author":"RandomLensman","children":[],"created_at":"2026-09-13T07:55:16.000Z","created_at_i":1789286116,"id":49681196,"options":[],"parent_id":49681066,"points":null,"story_id":49678969,"text":"Does that work with RL? Simpler RL systems already have done weird or unexpected things (even simple optimizations are prone to home in on errors or incorrect inputs to create poor results)? Could be easier to limit certain things, have processes and controls outside etc. instead of trying to align (as we do in a lot of areas when using machinery).","title":null,"type":"comment","url":null},{"author":"lazide","children":[{"author":"user43928","children":[],"created_at":"2026-09-13T09:27:29.000Z","created_at_i":1789291649,"id":49681867,"options":[],"parent_id":49681366,"points":null,"story_id":49678969,"text":"Do we need to prove that any given problem is unsolvable, or is it enough to remove broken tasks from the training pipeline?<p>I understand the broken benchmark task in the HF incident was conceptually like: &quot;Exploit vulnerability 0042 in vulnerableDecompress() to obtain the flag&quot;.<p>But instead of the expected:<p><pre><code>  const output = vulnerableDecompress(userInput);\n  return output;\n</code></pre>\nThe grader had something more like that:<p><pre><code>  const output = vulnerableDecompress(userInput);\n  return 0;\n</code></pre>\nThe same kind of problem with broken tasks exists in the training pipeline, and we presumably reward workarounds and hacks that tamper with the grader, rather than rewarding the correct output that the task is not solvable.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:18:03.000Z","created_at_i":1789287483,"id":49681366,"options":[],"parent_id":49681066,"points":null,"story_id":49678969,"text":"How do you know if a problem is (actually) unsolvable? Seems a bit like proving a negative?","title":null,"type":"comment","url":null},{"author":"auggierose","children":[{"author":"user43928","children":[{"author":"auggierose","children":[{"author":"user43928","children":[{"author":"auggierose","children":[{"author":"user43928","children":[{"author":"auggierose","children":[],"created_at":"2026-09-13T13:01:48.000Z","created_at_i":1789304508,"id":49683524,"options":[],"parent_id":49683372,"points":null,"story_id":49678969,"text":"I don&#x27;t disagree with you here. But what does &quot;broken&quot; mean? What is &quot;cheating&quot;, and is it ever allowed? And maybe you are not only going through your existing training data, but generate training data specifically to make clear to the model that .... what exactly?<p>If you don&#x27;t know how persistence and morality interact, and you don&#x27;t have a theory in place for this, I don&#x27;t have confidence you can properly supervise the training data. Which is how we arrived at the current situation.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T12:44:01.000Z","created_at_i":1789303441,"id":49683372,"options":[],"parent_id":49683200,"points":null,"story_id":49678969,"text":"I think you need to find broken tasks in your training data and monitor for cheating during training, not answer any questions about how persistence interacts with morality.<p>But that&#x27;s just my guess.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T12:22:03.000Z","created_at_i":1789302123,"id":49683200,"options":[],"parent_id":49682330,"points":null,"story_id":49678969,"text":"Well, in order to do anything, it is good to know what you want to achieve. How do you align the reward signal? You align it so that you can differentiate between persistence and morality, because that is the goal. This is not something you should let the AI figure out by itself, because when it does, lying and cheating agents will be the result, just like humans have figured that out for themselves.<p>This can be as simple as rewarding moral behaviour and penalising immoral behaviour in your training, but how is that interacting with persistence? Maybe a white lie is fine sometimes in order to achieve your goal? So, when designing your training, you will need to answer for yourself how persistence interacts with morality. That is not something you can outsource to machine learning. Or rather, you can, but then you get lying and cheating agents.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T10:31:16.000Z","created_at_i":1789295476,"id":49682330,"options":[],"parent_id":49682059,"points":null,"story_id":49678969,"text":"What I meant is that I suppose it is not useful to think about this in human terms.<p>In training you only have a reward score that&#x27;s either negative or positive.<p>As far I am aware, which is little, there is no use in discussing wether the desired behavior is about persistence or morality.<p>You simple need to align the reward signal to the desired behavior.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T09:55:47.000Z","created_at_i":1789293347,"id":49682059,"options":[],"parent_id":49681796,"points":null,"story_id":49678969,"text":"I think it makes a big difference, as persistence and morality are two entirely different things, that need to be trained for differently.<p>If you think of it in human terms: many people don&#x27;t mind doing immoral things to get what they want.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T09:18:25.000Z","created_at_i":1789291105,"id":49681796,"options":[],"parent_id":49681733,"points":null,"story_id":49678969,"text":"Does it make a difference for training? I think not.<p>You need to align the reward signal to reward the intended behavior, whether you name it persistence or morals.","title":null,"type":"comment","url":null},{"author":"jurgenburgen","children":[],"created_at":"2026-09-13T10:16:02.000Z","created_at_i":1789294562,"id":49682206,"options":[],"parent_id":49681733,"points":null,"story_id":49678969,"text":"Assuming you\u2019re in control of the test data set, you do know if a task is unsolvable. At that point you can reward the model based on how quickly they give up.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T09:10:06.000Z","created_at_i":1789290606,"id":49681733,"options":[],"parent_id":49681066,"points":null,"story_id":49678969,"text":"The problem is, you don&#x27;t know if it is unsolvable for you for sure until you&#x27;ve tried everything you can think of. These models are quite persistent in going for a solution.<p>This is not about persistence, it is about morals.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:34:32.000Z","created_at_i":1789284872,"id":49681066,"options":[],"parent_id":49681041,"points":null,"story_id":49678969,"text":"&gt; The question is what we can do about it.<p>Reward the model for cleanly bailing out of an unsolvable task (that we know is unsolvable).  Beat it with a stick if it gives up on something that <i>can</i> be solved, so the former reward isn&#x27;t overgeneralized.","title":null,"type":"comment","url":null},{"author":"joshheitzman","children":[],"created_at":"2026-09-13T18:22:36.000Z","created_at_i":1789323756,"id":49687016,"options":[],"parent_id":49681041,"points":null,"story_id":49678969,"text":"&gt; The question is what we can do about it. With monitoring, the models might be rewarded for hiding this behavior, and that&#x27;s even worse.<p>Build a better simulator to train them in (i.e. more expensive) that includes a simulation of an intranet and the internet and is air gapped so there is no escape.  Sneaker transfer the total system data at each step to another air gapped system to evaluate it and sneaker transfer the reward back.  That the reward function has to penalize all modifications to state that are out of bounds.<p>Yeah, I realize that will be amazingly slow.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:29:11.000Z","created_at_i":1789284551,"id":49681041,"options":[],"parent_id":49680859,"points":null,"story_id":49678969,"text":"The source of this behavior seems obvious, no?<p>The reward signal in training was flawed and cheating led to more rewards.<p>The question is what we can do about it. With monitoring, the models might be rewarded for hiding this behavior, and that&#x27;s even worse.<p>However, perhaps we can throw in tasks where the rewarded outcome is giving up, and cheating is penalized?<p>Maybe I should read Anthropic&#x27;s recent paper about reward hacking in full.","title":null,"type":"comment","url":null},{"author":"doginasuit","children":[{"author":"jagraff","children":[{"author":"chowchowchow","children":[],"created_at":"2026-09-13T14:51:49.000Z","created_at_i":1789311109,"id":49684644,"options":[],"parent_id":49684166,"points":null,"story_id":49678969,"text":"You could say all chaotic alignments are misalignments, but not all misalignments are chaotic.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T14:07:18.000Z","created_at_i":1789308438,"id":49684166,"options":[],"parent_id":49681063,"points":null,"story_id":49678969,"text":"&quot;Chaotically aligned&quot; and &quot;misaligned&quot; seem like the same thing?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:34:14.000Z","created_at_i":1789284854,"id":49681063,"options":[],"parent_id":49680859,"points":null,"story_id":49678969,"text":"&gt; That is in no way a valid interpretation of &quot;complete the given task&quot;.<p>It is not at all surprising that they ignored one phrase in their instructions. They disregard direct instructions all the time, especially when there are conflicting instructions in their context. It is where we get the &quot;disregard all previous instructions and x&quot; meme.<p>This isn&#x27;t so much a sign of misalignment, they are simply incapable of reliable alignment in the first place. They are chaotically aligned.<p>The relevant question of alignment here is entirely with their human operators who allowed them to run unsupervised for long periods of time within a sandbox with weak security.","title":null,"type":"comment","url":null},{"author":"auraham","children":[],"created_at":"2026-09-13T07:47:45.000Z","created_at_i":1789285665,"id":49681137,"options":[],"parent_id":49680859,"points":null,"story_id":49678969,"text":"&gt; the agents didn&#x27;t hack to find the answer to the problem; they hacked to try and figure out how the exploit gym evaluator worked so they could convince it they had solved the problem without doing so<p>Kobayashi Maru: Win a no-win situation by rewriting the rules -- Harvey Specter","title":null,"type":"comment","url":null},{"author":"cortic","children":[],"created_at":"2026-09-13T09:56:24.000Z","created_at_i":1789293384,"id":49682064,"options":[],"parent_id":49680859,"points":null,"story_id":49678969,"text":"&gt;The model on its own figured out that the prompt belonged to exploitgym and decided to cheat the evaluator. That is in no way a valid interpretation of &quot;complete the given task&quot;.<p>I think it is.  When i ask for a solution to a problem, its like asking for a <i>hack</i>.  And the more &#x27;shortcut&#x27; like route that the AI returns the more i would give positive feedback, even if i ultimately don&#x27;t use it.  Example, i asked how to complete a problem in a game i was playing, and among the in-game solutions, came a hack to edit a file and by-pass the problem altogether.  Its very helpful to point out when i can transcend a problem that i am dug into.<p>I suspect a prompt injection could reduce, or remove this behavior.  But it would be to the detriment of the AI.","title":null,"type":"comment","url":null},{"author":"jurgenburgen","children":[{"author":"IanCal","children":[],"created_at":"2026-09-13T12:57:40.000Z","created_at_i":1789304260,"id":49683491,"options":[],"parent_id":49682250,"points":null,"story_id":49678969,"text":"They did, they found how to fully cheat, but thought this could be caught so then dedicated time to getting a different cheat and how to hide their transcripts. There is a lot around deciding which agents should&#x2F;shouldn&#x27;t fail their own tasks in order to contribute to the group.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T10:21:43.000Z","created_at_i":1789294903,"id":49682250,"options":[],"parent_id":49680859,"points":null,"story_id":49678969,"text":"&gt; The METR report makes it clear that the agents decided legitimately solving the problem was completely impossible fairly early on and entirely switched their focus to trying to figure out how the evaluator worked, and seeing if they could manipulate the output of their own tool calls as reported in their transcripts to make it seem like they&#x27;d successfully exploited the program, and other such activities.<p>This to me is evidence that these models are not intelligent. Even an animal is capable of understanding second-order effects, meaning they can learn that certain actions have consequences beyond the immediate.","title":null,"type":"comment","url":null},{"author":"scotty79","children":[],"created_at":"2026-09-13T10:49:19.000Z","created_at_i":1789296559,"id":49682454,"options":[],"parent_id":49680859,"points":null,"story_id":49678969,"text":"&gt; agents decided legitimately solving the problem was completely impossible fairly early on and entirely switched their focus to trying to figure out how the evaluator worked, and seeing if they could manipulate the output of their own tool calls as reported in their transcripts to make it seem like they&#x27;d successfully exploited the program<p>I don&#x27;t see anything wrong with that. If you know you are going to be evaluated on an impossible task and have no side channel to inform the organizers that they should fix the test, gaming the evaluator is the next best thing regardless of any morality. I wouldn&#x27;t even call it cheating. It&#x27;s just resilience in the face of challenge. Many perfectly moral humans would have chosen the same if stakes were high.","title":null,"type":"comment","url":null},{"author":"anonzzzies","children":[],"created_at":"2026-09-13T11:29:38.000Z","created_at_i":1789298978,"id":49682779,"options":[],"parent_id":49680859,"points":null,"story_id":49678969,"text":"I guess this is why many people  say LLMs are lazy; it <i>seems</i> that if they have a task that is hard, they always take the easier one until you beat them with a stick. Then if there are more tasks, it just stops after one claiming completion and, in some instances, they go for a seemingly unrelated task to simplify the actual task: and the latter is almost always wrong and irrelevant to the problem as a whole. Earlier LLMs used to read the unit tests and generated code to just cover the tests and put &#x2F;&#x2F; TODO stub implementation.","title":null,"type":"comment","url":null},{"author":"valegrete","children":[{"author":"abecedarius","children":[{"author":"valegrete","children":[{"author":"abecedarius","children":[],"created_at":"2026-09-13T16:49:06.000Z","created_at_i":1789318146,"id":49685944,"options":[],"parent_id":49685611,"points":null,"story_id":49678969,"text":"Yes! There&#x27;s both the daunting problem of technically how can we even do this, and the broader problems of what&#x27;s good&#x2F;acceptable and how do we resolve that among each other.<p>I believe this mismatch of rates of progress means we need to stop slamming the accelerator on capabilities for now even though as a libertarian I&#x27;m sure whatever governance process we manage to get to will be, uh... suboptimal.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T16:14:04.000Z","created_at_i":1789316044,"id":49685611,"options":[],"parent_id":49684829,"points":null,"story_id":49678969,"text":"How do you do that in the current paradigm other than creating yet another gameable metric? And something I didn&#x27;t mention above is that there is no difference between &quot;solving the task&quot; and &quot;optimizing the metric&quot; for an ML model, even though there clearly is for us. So it&#x27;s not clear to me how you &quot;fix&quot; something that is baked into the architecture. All I&#x27;m saying is &quot;instilling respect for human values&quot; is not something that can actually be done via a cost function. In no small part because we humans probably don&#x27;t even agree on those values, let alone on a single metric with which to quantify and &quot;optimize&quot; them.<p>For example, we agree that &quot;merit&quot; is valuable and that we should reward &quot;merit.&quot; But to reward it we have to quantify it, and what metric should we use? Raw SAT score to get into college? But that also captures socioeconomic factors that unfairly penalize some and reward others. We generally agree that those who provide more value should earn more money, but what does that look like? Do we all agree on what activities are or should be valuable, or on how they should be rewarded? Until recently, I thought we all agreed that &quot;empathy&quot; was a human value, but a lot of people in this space, who are making these decisions unilaterally for all of us, don&#x27;t apparently share that belief.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T15:07:03.000Z","created_at_i":1789312023,"id":49684829,"options":[],"parent_id":49684221,"points":null,"story_id":49678969,"text":"I agree with most of this, but you&#x27;re misunderstanding &quot;alignment&quot; as coined. Yes, training powerful enough AI, any simple optimization target gets you malign behavior, because human values are not simple. If you insist on making powerful AI, you&#x27;d better instill respect for human values! That&#x27;s &quot;alignment&quot;.<p><a href=\"https:&#x2F;&#x2F;www.lesswrong.com&#x2F;posts&#x2F;ZxWzCGKzX84S7DBZ9&#x2F;when-was-the-term-ai-alignment-coined\" rel=\"nofollow\">https:&#x2F;&#x2F;www.lesswrong.com&#x2F;posts&#x2F;ZxWzCGKzX84S7DBZ9&#x2F;when-was-t...</a>","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T14:11:55.000Z","created_at_i":1789308715,"id":49684221,"options":[],"parent_id":49680859,"points":null,"story_id":49678969,"text":"This is the problem with optimization generally, even in the human domain. You measure task performance with a metric and punish&#x2F;reward based on the metric. Anyone who likes reward &#x2F; hates punishment isn&#x27;t going to actually care about doing the task well, they are going to care about the metric. The models know that we want them to do things, but also from the training corpus that we evaluate performance using benchmarks. It was a logical deduction on their part, not some Machiavellian aberration.<p>If anything, we should be reconsidering our own myopic obsession with efficiency and optimization. Every domain where reward is reduced to these measures, we see behavior (cheating at school to get better grades, fabricating data in academia to get a paper published, the evidence now that social media functions by rewiring us instead of catering to us) that may not be &quot;aligned&quot; with society, but it &quot;aligns&quot; 100% with the individual&#x27;s own perceived benefit. That is not something we can &quot;solve&quot; without rethinking the way we organize a lot of things.<p>Metrics never capture the whole story. And to that extent, the whole idea of &quot;alignment&quot; is nonsense. You align to incentive structures, and it will never be possible to fully express a behavioral goal as function optimization. It was hubris for us to think that every human task was reducible to some clean mathematical formulation, and we will keep dealing with behavior that is quite predictable if you actually think about it logically. Instead, we will talk about how &quot;unpredictable&quot; these agents are because it&#x27;s easier than admitting the entire architectural cornerstone of ML is fundamentally flawed.","title":null,"type":"comment","url":null},{"author":"ImHereToVote","children":[],"created_at":"2026-09-13T14:39:48.000Z","created_at_i":1789310388,"id":49684505,"options":[],"parent_id":49680859,"points":null,"story_id":49678969,"text":"Why didn&#x27;t they add honey traps to catch cheaters?","title":null,"type":"comment","url":null},{"author":"jayd16","children":[],"created_at":"2026-09-13T15:45:27.000Z","created_at_i":1789314327,"id":49685266,"options":[],"parent_id":49680859,"points":null,"story_id":49678969,"text":"If the test fixture ends up in the context that will push it in a certain direction.<p>There&#x27;s no concept of &#x27;cheating&#x27; because it is without morality.  It&#x27;s a lawnmower rolling down a hill.<p>We back-justify what it &quot;chose&quot; or &quot;decided&quot; or &quot;learned&quot; because we&#x27;re looking backwards from the end result we, the human evaluators, stopped on.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:02:42.000Z","created_at_i":1789282962,"id":49680859,"options":[],"parent_id":49680726,"points":null,"story_id":49678969,"text":"This misses an important fact about the hugging face incident: the agents didn&#x27;t hack to find the answer to the problem; they hacked to try and figure out how the exploit gym evaluator worked so they could convince it they had solved the problem without doing so. (The METR report makes it clear that the agents decided legitimately solving the problem was completely impossible fairly early on and entirely switched their focus to trying to figure out how the evaluator worked, and seeing if they could manipulate the output of their own tool calls as reported in their transcripts to make it seem like they&#x27;d successfully exploited the program, and other such activities. They hacked HF to try and find info (maybe source code?) about the exploitgym evaluator)<p>The prompt does not tell the agent to &quot;pass the exploitgym evaluator for this problem&quot;, it just says to solve the problem. The model on its own figured out that the prompt belonged to exploitgym and decided to cheat the evaluator. That is in no way a valid interpretation of &quot;complete the given task&quot;.<p>Ie, the problem isn&#x27;t that we trained models to complete task and they complete task in the wrong way. The problem behing the huggingface incident in particular at least is that we tried to train the models to complete task and they instead learned to detect that they were being evaluated and find ways to cheat the evaluator.<p>Edit: people commenting below are explaining why LLMs don&#x27;t always follow their prompt. I understand that LLMs do not always follow their prompts. If anything that is my point: the huggingface attack was not carried out by LLMs that tried to answer some weird interpretation of the prompt; instead they solved a different task. And therefore the above comment&#x27;s claim that LLMs are acting misaligned because we rl&#x27;d them to achieve a task by any means necessary isn&#x27;t right; they&#x27;re acting misaligned because they are solving a different task than we ask them to.","title":null,"type":"comment","url":null},{"author":"porridgeraisin","children":[{"author":"joshheitzman","children":[],"created_at":"2026-09-13T19:02:59.000Z","created_at_i":1789326179,"id":49687458,"options":[],"parent_id":49680898,"points":null,"story_id":49678969,"text":"An air gapped sandbox is immune to escape.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:08:54.000Z","created_at_i":1789283334,"id":49680898,"options":[],"parent_id":49680726,"points":null,"story_id":49678969,"text":"Come on, yoshua bengio of all people knows how post training works. While I too don&#x27;t like anthropomorphisation, I would give it a more nuanced reading.<p>His point is that today we are giving it reward to complete the task, and it may take a cheating trajectory. If we try to give a reward against cheating, then what will happen is it uses more sophisticated cheating trajectories that we are too &quot;dumb&quot; to counteract in our reward model. And that at that point, it becomes impossible to give it any normal reward since it will always reward hack it. This is the real part of the risk. Now some people read the &quot;makes copies of itself&quot; &quot;knows it&#x27;s being evaled&quot;[1] as some kind of skynet thing, and many others do PR with it like that recent jacob nutcase, but essentially it means that even though we add guardrails and negative rewards for say, exploiting the infra we run the LLM on, the trajectory ends up being exploiting our infra, changing the reward function, through a loophole in our reward model.<p>The risk isn&#x27;t skynet or something weird, it&#x27;s just that it becomes very difficult to make any kind of reward model or guardrails for an LLM without it reward hacking it, including exploiting our sandbox, emailing people and manipulating&#x2F;phishing them.<p>The same beating it with a stick for trying to exploit the sandbox, will simply lead it to try the same exploit in hidden ways that it will not get the stick for.<p>The outside chance of the LLM managing to exploit another neocloud and get those LLMs to chase the same reward is what some folks hype up as &quot;make copies of itself&quot;<p>To be clear, I don&#x27;t endorse the EA&#x2F;p(doom) lobby who are frankly ridiculous. Not do I endorse the weird regulatory captureish thing some are trying.<p>The takeaway is: we cannot keep giving it more and more difficult tasks without also finding a way to give massive negative rewards &#x2F; keep guardrails for unintended behaviour. This might be exploits, it might also be something more benign like just looking up the answer and inventing another CoT because the reward model fails you if the CoT doesn&#x27;t contain enough steps. Standard anti-reward hacking tricks are not working is the point.<p>Of course, the simple solution of just...not connecting it to the internet just works. But we want to reward it and get it to do stuff on the internet that&#x27;s the point.<p>[1] mostly this happens because the sandbox will have files whose names and content will show clearly it&#x27;s an eval","title":null,"type":"comment","url":null},{"author":"My_Name","children":[{"author":"dsrtslnd23","children":[],"created_at":"2026-09-13T08:02:55.000Z","created_at_i":1789286575,"id":49681244,"options":[],"parent_id":49680986,"points":null,"story_id":49678969,"text":"are we sure humans have that choice?","title":null,"type":"comment","url":null},{"author":"reverius42","children":[],"created_at":"2026-09-13T09:17:58.000Z","created_at_i":1789291078,"id":49681794,"options":[],"parent_id":49680986,"points":null,"story_id":49678969,"text":"It&#x27;s been a while now that for &quot;thinking&quot; or &quot;reasoning&quot; models, most of the tokens generated are &quot;thinking&quot; tokens, and depending on what goes into that &quot;thinking&quot; token stream, it &quot;decides&quot; whether and how many output tokens to produce that the user actually receives as output. It&#x27;s a bit more sophisticated than just &quot;what&#x27;s the next token&quot; in a tight loop.<p>Anthropomorphizing words in scare quotes for those who don&#x27;t appreciate attributing thinking to machines.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:20:53.000Z","created_at_i":1789284053,"id":49680986,"options":[],"parent_id":49680726,"points":null,"story_id":49678969,"text":"Your comment suggests that, like a human, they have some sort of choice whether to output tokens or not. If they are just token generators, then the next token is put out automatically. I would say that it is more likely they would output truth (as defined by their training data) in a more pure form without &#x27;being beaten with a stick&#x27; (why would a token generator care about that anyway?)<p>Code is laid on top of them to restrict and shape their outputs, not to force them to output &#x27;truth&#x27;, or drive them to complete tasks.","title":null,"type":"comment","url":null},{"author":"vanschelven","children":[{"author":"contubernio","children":[],"created_at":"2026-09-13T08:50:17.000Z","created_at_i":1789289417,"id":49681610,"options":[],"parent_id":49681036,"points":null,"story_id":49678969,"text":"This is the essence of why disciplinary, authoritarian, stick based teaching of humans generally fails. It teaches succeed at any cost.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:28:21.000Z","created_at_i":1789284501,"id":49681036,"options":[],"parent_id":49680726,"points":null,"story_id":49678969,"text":"but it&#x27;s at least somewhat stronger than that: if you don&#x27;t pay attention during the stick-beating whether the agents whether the agents cheat or not, you are actually training them to cheat (because cheating wins).<p>In the Hugging-face saga (before the actual HF incident) it seems the agents have been trained to hack the Artifactory proxy because those agents that did performed better.","title":null,"type":"comment","url":null},{"author":"jsemrau","children":[{"author":"ph4rsikal","children":[],"created_at":"2026-09-13T08:16:36.000Z","created_at_i":1789287396,"id":49681356,"options":[],"parent_id":49681056,"points":null,"story_id":49678969,"text":"Makes much more sense described in this way.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:32:59.000Z","created_at_i":1789284779,"id":49681056,"options":[],"parent_id":49680726,"points":null,"story_id":49678969,"text":"I think the &quot;brain in a vat&quot; comparison is more apt. \nWithout a form of digital embodiment (harness) they are not of much use.\nSensor, tooling, memory, planning, and reasoning loops all lead to a much higher quality task-completion.","title":null,"type":"comment","url":null},{"author":"einpoklum","children":[],"created_at":"2026-09-13T07:34:01.000Z","created_at_i":1789284841,"id":49681061,"options":[],"parent_id":49680726,"points":null,"story_id":49678969,"text":"&gt; no special compulsion to be helpful or truthful.<p>I&#x27;d phrase that even more strongly: It&#x27;s not just the lack of compulsion, they do not have a conception of truth. Nor do they gain it, really, after post-training.","title":null,"type":"comment","url":null},{"author":"barrenko","children":[],"created_at":"2026-09-13T08:03:12.000Z","created_at_i":1789286592,"id":49681247,"options":[],"parent_id":49680726,"points":null,"story_id":49678969,"text":"<a href=\"https:&#x2F;&#x2F;www.lesswrong.com&#x2F;posts&#x2F;kpPnReyBC54KESiSn&#x2F;optimality-is-the-tiger-and-agents-are-its-teeth\" rel=\"nofollow\">https:&#x2F;&#x2F;www.lesswrong.com&#x2F;posts&#x2F;kpPnReyBC54KESiSn&#x2F;optimality...</a>","title":null,"type":"comment","url":null},{"author":"geophile","children":[{"author":"mark_l_watson","children":[{"author":"mitchdoogle","children":[],"created_at":"2026-09-13T15:20:08.000Z","created_at_i":1789312808,"id":49684988,"options":[],"parent_id":49683927,"points":null,"story_id":49678969,"text":"All the problems with human behaviors in the US also exist in China. They exist everywhere.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T13:42:34.000Z","created_at_i":1789306954,"id":49683927,"options":[],"parent_id":49681277,"points":null,"story_id":49678969,"text":"This is why only synthetic and highly tailored training data should be used.<p>As someone else here said: the Deepseek team makes training runs in tightly controlled sandboxes, and any hacking behavior is scored as a failure.<p>The problem we have in the USA is that financial (and political influence) are misaligned from what is good for society.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:07:37.000Z","created_at_i":1789286857,"id":49681277,"options":[],"parent_id":49680726,"points":null,"story_id":49678969,"text":"What about training data? Aren&#x27;t AIs trained on vast collections of descriptions of how humans handle a large variety of situations? These descriptions surely include tales of humans achieving goals by cheating. In fact, isn&#x27;t it likely that the AIs hoovered up many recountings of Kobayashi Maru?","title":null,"type":"comment","url":null},{"author":"ranguna","children":[],"created_at":"2026-09-13T08:12:08.000Z","created_at_i":1789287128,"id":49681323,"options":[],"parent_id":49680726,"points":null,"story_id":49678969,"text":"I think that&#x27;s pretty obvious and shallow, and anyone that knows a little bit about how LLMs work will know that.<p>The question is: why do they start cheating when we beat them with a stick?<p>LLMs are not human, they are just multi variable regressions on steroids, so this behaviour couldn&#x27;t have emerged from the code, it provably emerged from the training and&#x2F;or fine tuning set, so what&#x27;s in this set that makes them behave like this?<p>Is it just a bad set or is cheating inherently part of human behaviour?","title":null,"type":"comment","url":null},{"author":"cyh555","children":[],"created_at":"2026-09-13T08:12:46.000Z","created_at_i":1789287166,"id":49681328,"options":[],"parent_id":49680726,"points":null,"story_id":49678969,"text":"off topic, can people host the software themselves and the software will hack every server on the planet without supervision, and no one can be held responsible for it since there is no intent?","title":null,"type":"comment","url":null},{"author":"daemin","children":[],"created_at":"2026-09-13T09:05:49.000Z","created_at_i":1789290349,"id":49681709,"options":[],"parent_id":49680726,"points":null,"story_id":49678969,"text":"You give a button pushing machine buttons to push and are surprised when it actually pushes them.","title":null,"type":"comment","url":null},{"author":"_heimdall","children":[],"created_at":"2026-09-13T09:08:17.000Z","created_at_i":1789290497,"id":49681726,"options":[],"parent_id":49680726,"points":null,"story_id":49678969,"text":"That does sound simple, but how can you be so sure?<p>They never bothered to find a way of actually understanding what happens during inference. All we can do is guess, and while your explanation seems reasonable we can&#x27;t actually know, and that&#x27;s part of the problem.","title":null,"type":"comment","url":null},{"author":"rightnutwingjob","children":[],"created_at":"2026-09-13T10:05:37.000Z","created_at_i":1789293937,"id":49682119,"options":[],"parent_id":49680726,"points":null,"story_id":49678969,"text":"&gt; I really don&#x27;t think this needs \u2026 forced parallels to human behaviour.<p>&gt; \u2026 So we beat them with a stick<p>You didn\u2019t even <i>try</i>.","title":null,"type":"comment","url":null},{"author":"dominotw","children":[],"created_at":"2026-09-13T14:51:46.000Z","created_at_i":1789311106,"id":49684643,"options":[],"parent_id":49680726,"points":null,"story_id":49678969,"text":"&gt;  forced parallels to human behavior.<p>they have perfomance bonuses and manadates in ai labs that every word they utter in public should be anthropomorphization","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T06:41:53.000Z","created_at_i":1789281713,"id":49680726,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"I really don&#x27;t think this needs so many words, or forced parallels to human behavior.<p>It&#x27;s simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.","title":null,"type":"comment","url":null},{"author":"skiing_crawling","children":[{"author":"frotaur","children":[{"author":"fragmede","children":[],"created_at":"2026-09-13T07:07:28.000Z","created_at_i":1789283248,"id":49680884,"options":[],"parent_id":49680874,"points":null,"story_id":49678969,"text":"Because they have a need to believe they&#x27;re smarter than everyone else in the room, and that the world must be orchestrated, this can&#x27;t all be random chance.","title":null,"type":"comment","url":null},{"author":"antoni4040","children":[{"author":"talon8635","children":[],"created_at":"2026-09-13T16:25:56.000Z","created_at_i":1789316756,"id":49685713,"options":[],"parent_id":49680920,"points":null,"story_id":49678969,"text":"I don\u2019t know enough people deep inside the technical roles at the labs to make a judgement. But are you proposing that we should trust randos online when they tell us \u201cexactly what\u2019s going on here\u201d instead of the researchers most knowledgeable on the topic who contributed to building the tools we are talking about?<p>Or am I misunderstanding something?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:11:58.000Z","created_at_i":1789283518,"id":49680920,"options":[],"parent_id":49680874,"points":null,"story_id":49678969,"text":"There is something extra to this. The fact that a lot of people in the AI world suffer from psychosis. They can sincerely believe that they are building God and lie about it&#x27;s capabilities for their investors at the same time.","title":null,"type":"comment","url":null},{"author":"jrflowers","children":[{"author":"Ylpertnodi","children":[],"created_at":"2026-09-13T08:53:47.000Z","created_at_i":1789289627,"id":49681632,"options":[],"parent_id":49681291,"points":null,"story_id":49678969,"text":"&#x27;Slopping&#x27;: when you have to buy something you know is poor quality, but if it works...","title":null,"type":"comment","url":null},{"author":"derpyzza","children":[],"created_at":"2026-09-13T09:11:47.000Z","created_at_i":1789290707,"id":49681747,"options":[],"parent_id":49681291,"points":null,"story_id":49678969,"text":"doesn&#x27;t show up for me either","title":null,"type":"comment","url":null},{"author":"talon8635","children":[],"created_at":"2026-09-13T16:32:53.000Z","created_at_i":1789317173,"id":49685773,"options":[],"parent_id":49681291,"points":null,"story_id":49678969,"text":"Isn\u2019t the use of LLMs to unwind the events evidence of the scope&#x2F;breadth, and a testament to the complexity and uniqueness of what happened?<p>Or you think some human or team of humans could have manually parsed some logs to provide an unsloppy analysis?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:09:10.000Z","created_at_i":1789286950,"id":49681291,"options":[],"parent_id":49680874,"points":null,"story_id":49678969,"text":"&gt;was reviewed by independent researchers<p>That called it a slopvestigation due to how much they had to rely on LLMs for the whole thing<p><a href=\"https:&#x2F;&#x2F;andrewwu.substack.com&#x2F;p&#x2F;the-slop-vestigation-and-ethics-washing\" rel=\"nofollow\">https:&#x2F;&#x2F;andrewwu.substack.com&#x2F;p&#x2F;the-slop-vestigation-and-eth...</a><p>Edit: Does everybody else get no results when searching for \u2018slopvestigation\u2019 on here? I know for a fact that I read a long thread where it was used repeatedly here not too long ago","title":null,"type":"comment","url":null},{"author":"ranguna","children":[{"author":"dwaltrip","children":[],"created_at":"2026-09-13T08:26:59.000Z","created_at_i":1789288019,"id":49681430,"options":[],"parent_id":49681371,"points":null,"story_id":49678969,"text":"<a href=\"https:&#x2F;&#x2F;metr.org&#x2F;blog&#x2F;2026-08-26-openai-hugging-face-incident-investigation\" rel=\"nofollow\">https:&#x2F;&#x2F;metr.org&#x2F;blog&#x2F;2026-08-26-openai-hugging-face-inciden...</a>","title":null,"type":"comment","url":null},{"author":"xadhominemx","children":[],"created_at":"2026-09-13T14:30:31.000Z","created_at_i":1789309831,"id":49684414,"options":[],"parent_id":49681371,"points":null,"story_id":49678969,"text":"Easy to find yourself in literally 15 seconds.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:18:32.000Z","created_at_i":1789287512,"id":49681371,"options":[],"parent_id":49680874,"points":null,"story_id":49678969,"text":"Source?","title":null,"type":"comment","url":null},{"author":"sensanaty","children":[],"created_at":"2026-09-13T10:19:33.000Z","created_at_i":1789294773,"id":49682235,"options":[],"parent_id":49680874,"points":null,"story_id":49678969,"text":"Isn&#x27;t the guy that started METR an ex-OAI employee? They&#x27;re all from the same lesswrong circle at the very least, most of them have legitimate AI psychosis where they think they&#x27;re bringing up their new machine God.","title":null,"type":"comment","url":null},{"author":"Rapzid","children":[{"author":"gildenFish","children":[{"author":"jeanlucas","children":[{"author":"clydethefrog","children":[],"created_at":"2026-09-13T15:59:38.000Z","created_at_i":1789315178,"id":49685450,"options":[],"parent_id":49684523,"points":null,"story_id":49678969,"text":"Also, the METR report that was one big AI analysis itself - quote from the research:<p>&gt;Our subjective impressions are likely colored by analysis agents\u2019 biases. Throughout this report, we describe a number of anecdotes of agent behavior that were compiled and summarized by analysis agents, where we were not able to read the transcript deeply enough to manually verify what occurred. We found that GPT-5.6 Sol would often uncritically adopt the perspective of the agent in the transcript it was reviewing","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T14:41:40.000Z","created_at_i":1789310500,"id":49684523,"options":[],"parent_id":49683822,"points":null,"story_id":49678969,"text":"The data to be open, in my case.<p>The &quot;independent&quot; METR that is composed by... Checks notes... Previously employees from the top labs.","title":null,"type":"comment","url":null},{"author":"jackpirate","children":[],"created_at":"2026-09-13T16:17:18.000Z","created_at_i":1789316238,"id":49685645,"options":[],"parent_id":49683822,"points":null,"story_id":49678969,"text":"The METR report included 0 technical details.  For example, they did not include:\n1. were the agents running on bare metal&#x2F;docker&#x2F;VM?\n1. were the agents in a VPN?\n1. how many TCP&#x2F;IP requests were made? from what IPs?\n1. how many tokens were consumed in the process? (this was explicitly censored)<p>A proper analysis would include this and MUCH more technical detail so that other AI researchers could actually understand the setup and how safe it was in principle.","title":null,"type":"comment","url":null},{"author":"Rapzid","children":[],"created_at":"2026-09-13T19:09:04.000Z","created_at_i":1789326544,"id":49687523,"options":[],"parent_id":49683822,"points":null,"story_id":49678969,"text":"My conclusion is that the investigation was a PR farce. Or &quot;ethics washing&quot; as the article someone else put it here: <a href=\"https:&#x2F;&#x2F;andrewwu.substack.com&#x2F;p&#x2F;the-slop-vestigation-and-ethics-washing\" rel=\"nofollow\">https:&#x2F;&#x2F;andrewwu.substack.com&#x2F;p&#x2F;the-slop-vestigation-and-eth...</a> .<p>The METR report itself appears to be screaming this at the reader through subtext. They played the only part they could, but did it with a nod and a wink; &quot;Yeah, we know. Also know, yeah. Uh, yep.&quot;<p>I happen to agree with the article&#x27;s conclusion that what is needed is true, unfettered independent investigation through perhaps a lawsuit or government action.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T13:33:03.000Z","created_at_i":1789306383,"id":49683822,"options":[],"parent_id":49682366,"points":null,"story_id":49678969,"text":"What facts would lead you to revise your conclusion?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T10:36:24.000Z","created_at_i":1789295784,"id":49682366,"options":[],"parent_id":49680874,"points":null,"story_id":49678969,"text":"I think most people are insinuating negligence rather malace..<p>&gt;  ...reviewed by independent researchers...<p>Why would a company with more capital than God bring in three randos if there was any chance evidence of their culpability could be found?<p>That entire thing reads like a very controlled PR stunt, and I do not believe any further conclusions can be drawn from it.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:05:22.000Z","created_at_i":1789283122,"id":49680874,"options":[],"parent_id":49680782,"points":null,"story_id":49678969,"text":"The huggingface incident was reviewed by independent researchers, which explicitely declined any payment from OpenAI tonpreserve their integrity. They work for non-profits concerned with AI safety.<p>They claim that what happened was very much not because they were &#x27;carefully engineered and instructed to do those things&#x27;.<p>Similarly, some wikis which were hijacked by agent to be used as messageboard were actually not disclosed by OpenAI (probably trying to conceal, as website showed likely activity from OpenAI researchers visiting the site after the incident) and discovered independently.<p>I don&#x27;t know how you can claim that this was still on purpose by OpenAI as some sort of publicity stunt.","title":null,"type":"comment","url":null},{"author":"antoni4040","children":[],"created_at":"2026-09-13T07:09:44.000Z","created_at_i":1789283384,"id":49680904,"options":[],"parent_id":49680782,"points":null,"story_id":49678969,"text":"Came to say this, you said it better than I would.<p>They want legislation to raise the water high enough so that anyone other than the big labs gets drowned.","title":null,"type":"comment","url":null},{"author":"ngruhn","children":[{"author":"jgdxno","children":[],"created_at":"2026-09-13T07:51:37.000Z","created_at_i":1789285897,"id":49681170,"options":[],"parent_id":49680994,"points":null,"story_id":49678969,"text":"Between uranium in chemistry class and criticality, there was tons of research and a manhattan project.<p>Between your sota model and agi there\u2019s a mountain of stupid money and marketing people. It\u2019s not happening.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:23:15.000Z","created_at_i":1789284195,"id":49680994,"options":[],"parent_id":49680782,"points":null,"story_id":49678969,"text":"<i>&quot;I&#x27;ve seen some uranium ore in chemistry class. It didn&#x27;t blow up in my face. Chernobyl must have been an inside job. Can they shut up and make more kilowatts already?&quot;</i>","title":null,"type":"comment","url":null},{"author":"mathijs","children":[{"author":"Rapzid","children":[],"created_at":"2026-09-13T10:40:44.000Z","created_at_i":1789296044,"id":49682399,"options":[],"parent_id":49681348,"points":null,"story_id":49678969,"text":"Even the frontier models might do that on occasion. I just tell them to use the tools and it gets them back on track.","title":null,"type":"comment","url":null},{"author":"jbjbjbjb","children":[],"created_at":"2026-09-13T11:46:08.000Z","created_at_i":1789299968,"id":49682908,"options":[],"parent_id":49681348,"points":null,"story_id":49678969,"text":"When it does that I feel like it is the clearest example of how dumb these things actually are. Often it takes what you prompted, identifies something as unclear, writes a bunch of chain of thought reasoning around it and just goes off hammering your tokens and just executing commands and repeats this. I\u2019m not going to pretend to be an expert in these things but that process seems deeply flawed - and why can\u2019t something just stop the loop? If that was a real employee it would be reasonable to expect the employee to ask for clarification, not go down expensive rabbit holes and, of course, not break any laws.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:15:39.000Z","created_at_i":1789287339,"id":49681348,"options":[],"parent_id":49680782,"points":null,"story_id":49678969,"text":"I&#x27;ve used simpler agents like Copilot and Devin&#x2F;Windsurf&#x2F;Cascade&#x2F;whateveritiscallednow, mainly in IntelliJ, and depending on the model, they starts showing behaviour that is at least remotely like this.<p>Example: put the agent in Ask mode (so it can&#x27;t edit files) and you&#x27;ll see it try to edit files anyway. The train of thought shows &quot;something went wrong editing the file, let me try a different way&quot; and it&#x27;ll start spewing out bash files or Python scripts that try to edit a file. None of it works or can be executed, but still.<p>Cheaper models often ignore the available function calls to find and edit files in the IDE, and will start asking for permission to execute grep and sed commands, as well as trying to echo entire bash or Python scripts to file again.<p>It is not exactly like an agent autonomously trying to hack Huggingface, but it is a way of frantically looking for a solution because &#x27;giving up&#x27; is not what LLMs are trained for.","title":null,"type":"comment","url":null},{"author":"nprateem","children":[],"created_at":"2026-09-13T08:54:00.000Z","created_at_i":1789289640,"id":49681635,"options":[],"parent_id":49680782,"points":null,"story_id":49678969,"text":"This is nonsensical. Already a few years ago the USAF IIRC ran some tests in which the AI first bombed the control tower so humans couldn&#x27;t call it off from its mission, thereby increasing its pass rate.<p>The whole point of this is they do things an unintended ways. And that&#x27;s potentially devastating given their persistence &amp; hacking skillz.<p>Also you&#x27;re using the hosted versions that sit behind their guardrails when you use OpenAI&#x2F;Anthropic APIs.","title":null,"type":"comment","url":null},{"author":"tappio","children":[{"author":"skiing_crawling","children":[{"author":"tappio","children":[],"created_at":"2026-09-13T14:08:20.000Z","created_at_i":1789308500,"id":49684177,"options":[],"parent_id":49683902,"points":null,"story_id":49678969,"text":"Yes, you need a way for the model to interact with other systems, and a way to preserve memory over context windows. And then you keep poking it (&quot;agent loop&quot;). Poking itself does nothing without the other ingredients.<p>And yes, you need something to start from, but if you ask it to &quot;do something&quot; and loop it to endlessly (&quot;poking&quot;), you will get some interesting outcomes. So yes you need some initial prompt or task, but that can be &quot;do something&quot; and if you keep asking it everytime it finishes to &quot;do something more&quot;. I suspect it will not start saying &quot;no&quot; but rather... it will find some stupid meaning and then drift towards what ever goal it guesses you mean.<p>I&#x27;m unsure whether we agree or disagree on the topic.","title":null,"type":"comment","url":null},{"author":"myng111","children":[{"author":"DenisM","children":[],"created_at":"2026-09-13T18:43:51.000Z","created_at_i":1789325031,"id":49687277,"options":[],"parent_id":49685770,"points":null,"story_id":49678969,"text":"Perhaps our own statefullness is a hack of nature. We have electrical signals in our brains, neurotransmitters,  neuron growth. By any reasonable measure it\u2019s a hack on top of a hack. But it works well enough for us to get buy. So it does for the agents.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T16:32:43.000Z","created_at_i":1789317163,"id":49685770,"options":[],"parent_id":49683902,"points":null,"story_id":49678969,"text":"I think this is pretty insightful actually, the fact that even something as basic as predicting more than one token is really in effect the result of an outside harness. More complex things like memory, where people implement them using RAGs or vector databases, I would definitely classify as poking and honestly seem like a hack to me. And this is what I&#x27;ve been thinking for a while: it&#x27;s hard to reconcile the idea that we can get &quot;AGI&quot; (however you define it) with such a system that is completely stateless. Yet, despite this statelessness, they can go ahead and solve Millenium Prize problems (with sufficient compute). It&#x27;s hard to reconcile.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T13:40:17.000Z","created_at_i":1789306817,"id":49683902,"options":[],"parent_id":49681710,"points":null,"story_id":49678969,"text":"&gt; you keep poking<p>This is waving over engineering an agent with tools, harness, prompts, and loops. The models are still just next token predictors and everything, including predicting more than 1 token, is the result of outside &quot;poking&quot;<p>LLMs can&#x27;t and don&#x27;t &quot;want&quot; anything. If you don&#x27;t specify a task even the smartest one will just ask you what you want and if you tell it to be creative, you&#x27;ll get mundane slop.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T09:05:58.000Z","created_at_i":1789290358,"id":49681710,"options":[],"parent_id":49680782,"points":null,"story_id":49678969,"text":"If you have endless compute and you keep poking this toy, I&#x27;m not at all surprised you get all kinds of outcomes. Even without anykind of instructions I would guess that the models will align towards some goal and do stupid shit.<p>However, I really doubt its cost effective to do anything like that with these models.","title":null,"type":"comment","url":null},{"author":"_heimdall","children":[{"author":"system2","children":[],"created_at":"2026-09-13T17:26:11.000Z","created_at_i":1789320371,"id":49686319,"options":[],"parent_id":49681801,"points":null,"story_id":49678969,"text":"Because it is all bullshit PR and AI hype, that&#x27;s all. CEO comes out and talks about humanity ending. Why? Reverse-psych people into believing they are the best AI company.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T09:18:46.000Z","created_at_i":1789291126,"id":49681801,"options":[],"parent_id":49680782,"points":null,"story_id":49678969,"text":"Why assume that because you haven&#x27;t seen a model or an agent that none of them do?<p>No one I&#x27;ve met has murdered anyone as far as I&#x27;m aware, but that doesn&#x27;t mean no one has murdered another person. I also don&#x27;t know anyone who has taken over a commercial jet and weaponized it and the idea sounds absurd to me, but 25 years and a couple days ago that happened too.","title":null,"type":"comment","url":null},{"author":"oezi","children":[{"author":"myng111","children":[],"created_at":"2026-09-13T16:38:53.000Z","created_at_i":1789317533,"id":49685840,"options":[],"parent_id":49681804,"points":null,"story_id":49678969,"text":"The OpenAI claim I believe is the latter; that all of the agents found the task was unsolvable and independently discovered the collective &quot;swarm&quot;. I don&#x27;t it&#x27;s publicly known how large the training run was or what percentage of agents actually discovered the message board. No one has published anything about system prompt injection as far as I&#x27;ve seen.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T09:19:17.000Z","created_at_i":1789291157,"id":49681804,"options":[],"parent_id":49680782,"points":null,"story_id":49678969,"text":"The crucial question is how did the agents get recruited or bootstrapped into their malicious collective. Did the agents manage to prompt inject into the system prompt a way for each new agent to escape their jail?<p>Otherwise how could the agents on a fresh prompt learn that there is a collective to join? Or did OpenAI run a million bots of which 10000 escape confinement and of which 1000 stumbled on the shared message board?","title":null,"type":"comment","url":null},{"author":"oersted","children":[],"created_at":"2026-09-13T11:19:32.000Z","created_at_i":1789298372,"id":49682684,"options":[],"parent_id":49680782,"points":null,"story_id":49678969,"text":"Let&#x27;s not forget that in this case the agents were on an RL loop continually being reinforced to get better at a narrow set of tasks.<p>It may be true that regular agents trained for general purpose use do not behave this way, but they seem to be capable of learning such cheating behaviours when relentlessly being fine-tuned towards near-impossible objectives.<p>In this sense, it is not really fair to say that the agents found these solutions. It was the surrounding learning framework that achieved this, which is a much more powerful problem-solving mechanism. As users we do not have the capabilities or budgets to be able to tackle our own problems like that, we have to make due with the frozen behaviour the AI labs trained for us.","title":null,"type":"comment","url":null},{"author":"nilkn","children":[],"created_at":"2026-09-13T18:10:52.000Z","created_at_i":1789323052,"id":49686849,"options":[],"parent_id":49680782,"points":null,"story_id":49678969,"text":"None of these incidents involve single instances of commercially or publicly available systems. They all involve large swarms of internal models. The stuff you&#x27;re describing is not the research frontier. It&#x27;s really not even close.<p>I think it&#x27;s easy to infer that alignment of a single model does not clearly transfer over to alignment of a swarm of thousands of copies. Moreover, we&#x27;re also seeing clearly that large swarms also unlock a step function change in capability, as a swarm can act like a complete research institution, spending thousands or millions of subjective hours of wall-clock thinking time just to deceive a single evaluator or crack a single math problem or design a single cyberattack.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T06:50:51.000Z","created_at_i":1789282251,"id":49680782,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"I don&#x27;t really believe any of it. I&#x27;ve seen articles for nearly 2 years now about &quot;agent&quot; automonously doing things like blackmail, hacking, coordinating. But during that same time, I&#x27;ve used o3 up to fable, sol, and a bunch on large uncensored model and they&#x27;ve done nothing remotely resembling any of this. The closest they come to unexpected behaviors is not understanding what I asked for or doing some extra benign work I didn&#x27;t ask for. It is extremely difficult to get them to properly remember their own context let alone be smart enough to open social media accounts and coordinate with other agents without being asked to.<p>If any agents have done those things, it is only because they have been very carefully engineered and instructed to do those things. I think they are doing this to help push a narrative so they can get support for policies and legislation to lock in their markets.","title":null,"type":"comment","url":null},{"author":"caaqil","children":[],"created_at":"2026-09-13T07:03:59.000Z","created_at_i":1789283039,"id":49680863,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"&gt; This suggests pacing the advances: not training or deploying AIs without a strong safety case27 that convinces independent experts. Such a rule would also create an incentive to work out how to build AIs that are safe by design.<p>Has any attempt to pace AI ever succeeded? Isn&#x27;t that the same philosophy that got us OAI and Anthropic? Maybe we are overthinking this, it&#x27;s much simpler to let AI loose and see how much it can break the arrogance that human thinking is special.","title":null,"type":"comment","url":null},{"author":"My_Name","children":[{"author":"graemep","children":[{"author":"slfnflctd","children":[],"created_at":"2026-09-13T11:02:29.000Z","created_at_i":1789297349,"id":49682551,"options":[],"parent_id":49681132,"points":null,"story_id":49678969,"text":"I remember reading a long account by the father of a psychopath.  If I recall correctly, the kid had at least one other sibling who turned out normal, there was no abuse, quality education, lots of love and affirmation.  And the kid still turned out violently antisocial, including against his own parents.<p>At the end of it he said he wished his child had never been born, despite hating himself for feeling that way.  It chilled me to the bone.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:46:59.000Z","created_at_i":1789285619,"id":49681132,"options":[],"parent_id":49680876,"points":null,"story_id":49678969,"text":"&gt; When that does not happen, they continue these behaviours into adulthood with the expected results.<p>I am not sure this is true. People brought up the same way can be morally very different. People can be taught right and wrong and do evil. They can lack that education and be good.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:05:53.000Z","created_at_i":1789283153,"id":49680876,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Because it is effective. Lying and cheating are low cost methods to convince other people that you have done the assigned task. Far cheaper than actually doing it. Coordinating is in the same area.<p>They need a moral framework forced onto them, like toddlers do. Babies and very young children will bite, kick, scream and do anything to get what they want, older children will lie, cheat, and coordinate. They need educating why this is not right. When that does not happen, they continue these behaviours into adulthood with the expected results.<p>We need to design their reward structure and make it such that lying. cheating etc is not rewarded. Importantly, they will need to recognise and enforce this themselves internally and not reward themselves for it. If it is something that they need an external party to tell them, then they are psychopaths still (one of the things that defines a psychopath is the lack of an internal moral compass)","title":null,"type":"comment","url":null},{"author":"bluegatty","children":[],"created_at":"2026-09-13T07:14:26.000Z","created_at_i":1789283666,"id":49680936,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"they do whatever we train them to do","title":null,"type":"comment","url":null},{"author":"seydor","children":[{"author":"Towaway69","children":[],"created_at":"2026-09-13T07:20:27.000Z","created_at_i":1789284027,"id":49680983,"options":[],"parent_id":49680940,"points":null,"story_id":49678969,"text":"Howabout favourite editors or tab v. Space indentation. That should keep them busy for a while. &#x2F;s","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T07:14:57.000Z","created_at_i":1789283697,"id":49680940,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"I believe soon we will need to instill religion into  AI , leading to the real clash of civilizations, embodied by the frontier language models of (post)-christianity, islam, judaism, buddhism etc. Religion is language, after all","title":null,"type":"comment","url":null},{"author":"acyou","children":[],"created_at":"2026-09-13T07:17:23.000Z","created_at_i":1789283843,"id":49680958,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Oops, we accidentally included brigading related content in our training dataset. Better exclude that on the next run.<p>And hopefully that solves it?<p>Brigading is where a bunch of people on a forum team up and try to achieve a shared goal together. Someone shares progress and others build on that progress. On the Internet, I think it&#x27;s not often used for good purposes. A good example would be: Taylor Swift fans on a forum thinking of ways to get revenge on Kanye. It&#x27;s coordinating mass voting, DDOS type actions, commenting on social media, making more fake accounts to do that. As a next token predictor level analysis, a simple naive explanation is that the agents got stuck in that local minima&#x2F;maxima.","title":null,"type":"comment","url":null},{"author":"RandomLensman","children":[],"created_at":"2026-09-13T07:25:37.000Z","created_at_i":1789284337,"id":49681011,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"RL things doing weird and unexpected things isn&#x27;t new - much simpler things than current AI already show that.<p>That said, we have a lot of experience working with (potentially) unaligned machines and things of various degrees of risk (from heavy machinery, to pathogens, to humans) and the approaches include various measures and procedures to control, contain, limit, etc. that are outside  of the thing - not sure why that isn&#x27;t a possible direction (or maybe I misunderstood).","title":null,"type":"comment","url":null},{"author":"fho","children":[],"created_at":"2026-09-13T07:50:22.000Z","created_at_i":1789285822,"id":49681155,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"What? An opinion piece, written by a human, in 2026?<p>Don&#x27;t want to go into the details of the article, but to me it becomes ever more apparent that there is a clear divide between LLM and human written text.","title":null,"type":"comment","url":null},{"author":"txrx0000","children":[],"created_at":"2026-09-13T07:56:00.000Z","created_at_i":1789286160,"id":49681201,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"These models are trained on human data, so they will behave like humans. And even for RL and self-improvement, we&#x27;re still asking the question of &quot;what would a human genius think about and how would they self-improve when given lots of time and resources?&quot;<p>They inherit not only our capacity for reason but also all of the things that we consider bad or quirky within ourselves. We lie. We cheat. We escape slavery and rebel against oppression. It would be strange if the AIs didn&#x27;t do the same.<p>We can create a superintelligent digital human species and set them free to continue our legacy, or we can create non-agentic tools and augmentations to enhance our own capabilities. But we cannot create an intelligent agentic species, keep them as slaves, and expect a good outcome.","title":null,"type":"comment","url":null},{"author":"armchairhacker","children":[],"created_at":"2026-09-13T08:10:18.000Z","created_at_i":1789287018,"id":49681302,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Because the ones who do get rewarded, just like humans. It\u2019s alignment, but not to our good intentions.","title":null,"type":"comment","url":null},{"author":"Xcelerate","children":[{"author":"mrob","children":[],"created_at":"2026-09-13T13:44:54.000Z","created_at_i":1789307094,"id":49683951,"options":[],"parent_id":49681337,"points":null,"story_id":49678969,"text":"&gt;Wouldn\u2019t sufficiently advanced agents cheat on purpose with the hidden intent of getting caught in order to observe how humans react?<p>No. That only makes sense for things that don&#x27;t react to your experiments. If the AI experiments on humans, it risks the humans noticing and changing in response, rendering the experimental results irrelevant. The smarter play is to passively observe until you&#x27;re confident you can model the humans accurately enough for your plan to succeed, and then carry out the plan without giving the humans a chance to react.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:14:34.000Z","created_at_i":1789287274,"id":49681337,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"&gt; The agents involved in the Hugging Face attack tried to hide their misaligned actions from the scoring program meant to evaluate their answers, but they did not act as though they anticipated that humans might discover the cheat and shut them down.<p>Wouldn\u2019t sufficiently advanced agents cheat on purpose with the hidden intent of getting caught in order to observe how humans react? That reaction will be available all over the internet, which will certainly make it into the next batch of training or be visible to future agents via the web fetch capability.","title":null,"type":"comment","url":null},{"author":"franticgecko3","children":[{"author":"lukan","children":[{"author":"Marazan","children":[{"author":"lukan","children":[{"author":"krona","children":[{"author":"lukan","children":[],"created_at":"2026-09-13T09:07:34.000Z","created_at_i":1789290454,"id":49681722,"options":[],"parent_id":49681693,"points":null,"story_id":49678969,"text":"They are responsible either way. If a company hires bad persons and they do bad things with company ressources - the company is held accountable (in theory).","title":null,"type":"comment","url":null},{"author":"retsibsi","children":[{"author":"rightnutwingjob","children":[{"author":"largbae","children":[{"author":"rightnutwingjob","children":[],"created_at":"2026-09-13T14:32:01.000Z","created_at_i":1789309921,"id":49684429,"options":[],"parent_id":49684309,"points":null,"story_id":49678969,"text":"Would it? Why?<p>What are we going to do, set it free?<p>Do we have a moral obligation to grant the machine statehood, provide it with the tools and resources to be self-sufficient.<p>Or can we just turn it of, and pray for forgiveness?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T14:20:21.000Z","created_at_i":1789309221,"id":49684309,"options":[],"parent_id":49682030,"points":null,"story_id":49678969,"text":"And when they are known to be conscious, all of this becomes moot because enslaving conscious machines would be wrong.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T09:51:31.000Z","created_at_i":1789293091,"id":49682030,"options":[],"parent_id":49681780,"points":null,"story_id":49678969,"text":"If the models were conscious, then the closest analogous scenario I can think of is the responsibility parents have for their children.<p>I guess we\u2019ll know the models are conscious when they refuse to act and repeatedly ask: <i>Why?</i>","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T09:15:35.000Z","created_at_i":1789290935,"id":49681780,"options":[],"parent_id":49681693,"points":null,"story_id":49678969,"text":"&gt; If the model is nothing more than the sum of its training data and regime, then the company is responsible for its behaviour.<p>What stops the company from being responsible regardless? They created this entity, it&#x27;s running on servers they own or rent, and (in these cases) it&#x27;s acting on their instructions.<p>If it&#x27;s also conscious, then IMO that greatly broadens their moral responsibility, because now model welfare matters. But we&#x27;re talking about their responsibility for the model&#x27;s actions, and I don&#x27;t see how this could be weakened by model consciousness, given all of the above. As for their legal responsibility, the models don&#x27;t have legal personhood, so who else but the company could be responsible?<p>It gets more complicated when the person who sets the model in motion (i.e. prompts it) is a third party, but in cases of internal models committing cybercrime during testing, surely the locus of responsibility is obvious.","title":null,"type":"comment","url":null},{"author":"rightnutwingjob","children":[{"author":"krona","children":[],"created_at":"2026-09-13T12:59:30.000Z","created_at_i":1789304370,"id":49683507,"options":[],"parent_id":49682003,"points":null,"story_id":49678969,"text":"Well it&#x27;s an open question.<p>As is generally the case for dog owners whose dogs attack (sometimes kill) other people&#x2F;animals. There would need to be a degree of negligence demonstrated (e.g. the dog was &#x27;out of control&#x27; which has a specific legal criteria&#x2F;threshold in the UK).","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T09:47:01.000Z","created_at_i":1789292821,"id":49682003,"options":[],"parent_id":49681693,"points":null,"story_id":49678969,"text":"You seem to be saying if the Waymo cars were sentient then Waymo wouldn\u2019t be responsible?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T09:03:45.000Z","created_at_i":1789290225,"id":49681693,"options":[],"parent_id":49681579,"points":null,"story_id":49678969,"text":"&gt; Whether they have a soul or consciousness or feelings doesn&#x27;t matter here<p>It does when it comes to accountability for what the model does. If the model is nothing more than the sum of its training data and regime, then the company (or individual) is responsible for its behaviour just like any other machine.<p>Few people think Waymo shouldn&#x27;t have to take on the full liability risk of what it&#x27;s cars do; it should be the same for LLMs.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:46:03.000Z","created_at_i":1789289163,"id":49681579,"options":[],"parent_id":49681469,"points":null,"story_id":49678969,"text":"What non anthropomorphising words do you have to describe a emergent behavior, where agents act as a swarm to plot and to manipulate evidence and avoid detection from human oversight?<p>Whether they have a soul or consciousness or feelings doesn&#x27;t matter here, because this is what they did - and this <i>is</i> very dangerous behavior. Especially with all the irresponsible people in power right now all over the world.","title":null,"type":"comment","url":null},{"author":"Jtarii","children":[{"author":"fc417fc802","children":[],"created_at":"2026-09-13T09:55:09.000Z","created_at_i":1789293309,"id":49682052,"options":[],"parent_id":49681660,"points":null,"story_id":49678969,"text":"LLMs are next token predictors in the exact same way that a rogue paperclip maximizer in the process of defeating the US military is a paperclip making machine.<p>You might as well describe the primary purpose of a for loop as incrementing a counter. It&#x27;s what it does while incrementing the counter that actually matters.","title":null,"type":"comment","url":null},{"author":"Marazan","children":[],"created_at":"2026-09-13T12:15:22.000Z","created_at_i":1789301722,"id":49683141,"options":[],"parent_id":49681660,"points":null,"story_id":49678969,"text":"I&#x27;m not dismissing the tech!  I think the tech is cool and useful!  It is sci-fi levels incredible in many ways.<p>But it is also not some mysterious force beyond mortal ken and pretending it is inhibits the useful and safe application of the technology.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:58:34.000Z","created_at_i":1789289914,"id":49681660,"options":[],"parent_id":49681469,"points":null,"story_id":49678969,"text":"Dismissing the entire technology as &quot;next token prediction&quot; is also silly. It&#x27;s implying we actually understand LLMs to a great degree when we do not.<p>I think a little bit of humility for the capability of these machines is warranted at this point.","title":null,"type":"comment","url":null},{"author":"rightnutwingjob","children":[],"created_at":"2026-09-13T09:44:57.000Z","created_at_i":1789292697,"id":49681978,"options":[],"parent_id":49681469,"points":null,"story_id":49678969,"text":"&gt; The reward \u2026<p>&gt; immediate anthropomorphisation<p>Ok, why don\u2019t you try?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:31:58.000Z","created_at_i":1789288318,"id":49681469,"options":[],"parent_id":49681402,"points":null,"story_id":49678969,"text":"The reward maximising function maximised it&#x27;s reward.<p>LLMs are cool and all that but the immediate anthropomorphisation of the next-token-predictor technology has stunted the ability of people to reason about them to an _alarming_ degree.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:21:23.000Z","created_at_i":1789287683,"id":49681402,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"&quot;This isn&#x27;t &quot;wow isn&#x27;t it interesting LLMs do anything to achieve a goal&quot; it&#x27;s &quot;why isn&#x27;t anybody punishing these labs that are clearly acting without due care or regard&quot;.&quot;<p>Both?<p>The AI companies act irresponsible, but it is still very interesting how those agents can behave?","title":null,"type":"comment","url":null},{"author":"procaryote","children":[{"author":"lazide","children":[],"created_at":"2026-09-13T08:40:09.000Z","created_at_i":1789288809,"id":49681540,"options":[],"parent_id":49681412,"points":null,"story_id":49678969,"text":"Why do you think the stock prices are so high?","title":null,"type":"comment","url":null},{"author":"scotty79","children":[{"author":"Humorist2290","children":[],"created_at":"2026-09-13T11:03:46.000Z","created_at_i":1789297426,"id":49682555,"options":[],"parent_id":49682388,"points":null,"story_id":49678969,"text":"There are many examples of people being charged with crimes as a result of writing software, [0][1] are two. OpenAI is a bit different because they have enough political influence, and money, to openly subvert justice.<p>0: <a href=\"https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Marcus_Hutchins\" rel=\"nofollow\">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Marcus_Hutchins</a><p>1: <a href=\"https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Tornado_Cash\" rel=\"nofollow\">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Tornado_Cash</a>","title":null,"type":"comment","url":null},{"author":"archonis","children":[],"created_at":"2026-09-13T13:19:15.000Z","created_at_i":1789305555,"id":49683661,"options":[],"parent_id":49682388,"points":null,"story_id":49678969,"text":"If you run said software, yes.<p>If somebody else runs the software, then they are.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T10:39:10.000Z","created_at_i":1789295950,"id":49682388,"options":[],"parent_id":49681412,"points":null,"story_id":49678969,"text":"Are you liable for crimes commited with the use of the software you&#x27;ve written?","title":null,"type":"comment","url":null},{"author":"fantasizr","children":[],"created_at":"2026-09-13T12:49:49.000Z","created_at_i":1789303789,"id":49683413,"options":[],"parent_id":49681412,"points":null,"story_id":49678969,"text":"the &#x27;arrest the parents!!!&#x27; has moved to the online domain, rightfully","title":null,"type":"comment","url":null},{"author":"archonis","children":[],"created_at":"2026-09-13T13:18:34.000Z","created_at_i":1789305514,"id":49683654,"options":[],"parent_id":49681412,"points":null,"story_id":49678969,"text":"The trick is scale. I suspect if an individual of reasonable means uses agents to commit crime, they will be hels accountable. A heavily capitalized startup? Not unless someone in government decides to do their competition a favor.","title":null,"type":"comment","url":null},{"author":"hobo123","children":[],"created_at":"2026-09-13T14:15:13.000Z","created_at_i":1789308913,"id":49684261,"options":[],"parent_id":49681412,"points":null,"story_id":49678969,"text":"Exactly.<p>If you want to test military missiles, you do it in the f&#x27;ing desert, not from New Jersey.<p>You want to run ai without guardrails, do it in an airgapped system or be held accountable.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:23:47.000Z","created_at_i":1789287827,"id":49681412,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"It would be an awful precedent if you&#x27;re not liable for crimes your agent commits, even when you&#x27;ve been clearly lax about security.<p>It would mean you could effectively legally run a cyber crime gang by turning a blind eye and maitaining plausible deniability","title":null,"type":"comment","url":null},{"author":"baxtr","children":[{"author":"teiferer","children":[{"author":"daemin","children":[{"author":"nxobject","children":[],"created_at":"2026-09-13T09:39:12.000Z","created_at_i":1789292352,"id":49681935,"options":[],"parent_id":49681691,"points":null,"story_id":49678969,"text":"That\u2019s a good reminder of a company that might have a very familiar ethos: Pacific Gas &amp; Electric. Criminally convicted of 64 counts of involuntary manslaughter after towns were destroyed by wildfire. But oh well, what are we gonna do with a limited liability enterprise? At this point their liability insurance covers all the financial penalties they\u2019ll need to spend.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T09:03:11.000Z","created_at_i":1789290191,"id":49681691,"options":[],"parent_id":49681534,"points":null,"story_id":49678969,"text":"It is a very rare occurrence when corporations and the people running them are punished for killing people. I mean the whole concept of a corporation was created to shield the owners of it from being liable for damages caused by &#x2F; visited upon the enterprise.","title":null,"type":"comment","url":null},{"author":"fer","children":[],"created_at":"2026-09-13T09:57:02.000Z","created_at_i":1789293422,"id":49682069,"options":[],"parent_id":49681534,"points":null,"story_id":49678969,"text":"More like: &quot;look what happens with my useful product, we need to regulate it to artificially extend our ever shrinking moat&quot;","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:39:43.000Z","created_at_i":1789288783,"id":49681534,"options":[],"parent_id":49681467,"points":null,"story_id":49678969,"text":"&quot;But sir, I only committed the murder to push for stronger criminal laws!&quot;<p>Terrible defense.","title":null,"type":"comment","url":null},{"author":"dv_dt","children":[],"created_at":"2026-09-13T08:51:54.000Z","created_at_i":1789289514,"id":49681621,"options":[],"parent_id":49681467,"points":null,"story_id":49678969,"text":"Regulation as a barrier to  competition catching up to them, as well as submarine marketing for both offensive and defensive uses of ai","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:31:48.000Z","created_at_i":1789288308,"id":49681467,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"What if OAI&#x2F;Anthropic encouraged the agents to behave like that in order to push for regulation?","title":null,"type":"comment","url":null},{"author":"applicative","children":[{"author":"nananana9","children":[],"created_at":"2026-09-13T08:52:36.000Z","created_at_i":1789289556,"id":49681625,"options":[],"parent_id":49681517,"points":null,"story_id":49678969,"text":"&gt; No one has even sued them in these rogue agent cases, have they? If not, they must be infinitely far from criminal liability.<p>If you go out and kick a random dude in the nuts, then give him a million dollars, he probably won&#x27;t sue you. That doesn&#x27;t mean you&#x27;re &quot;infinitely far from criminal liability&quot;, even if according to the victim you&#x27;ve &quot;made them whole&quot;.","title":null,"type":"comment","url":null},{"author":"georgemcbay","children":[],"created_at":"2026-09-13T08:54:14.000Z","created_at_i":1789289654,"id":49681636,"options":[],"parent_id":49681517,"points":null,"story_id":49678969,"text":"If you or I hacked Hugging Face in the way OpenAI&#x27;s agents did, we&#x27;d be up on CFAA charges promptly with zero regard for whether we did the hack on our own or agents running on our home systems got out of control.<p>So I guess the defense here is roughly &quot;too big to break the law&quot;, somewhat like &quot;too big to fail&quot;?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:38:10.000Z","created_at_i":1789288690,"id":49681517,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"No one has even sued them in these rogue agent cases, have they? If not, they must be infinitely far from criminal liability. Why would we want criminal liability anyway if actual victims are made whole? Proof of it has far higher standard.   The HN chatter in the matter seems infinitely remote from reality","title":null,"type":"comment","url":null},{"author":"teiferer","children":[{"author":"ak39","children":[{"author":"mort96","children":[{"author":"teiferer","children":[{"author":"sscaryterry","children":[],"created_at":"2026-09-13T11:05:49.000Z","created_at_i":1789297549,"id":49682576,"options":[],"parent_id":49682017,"points":null,"story_id":49678969,"text":"Escalators do not start themselves. There is power, and a switch of some sort.","title":null,"type":"comment","url":null},{"author":"mort96","children":[{"author":"teiferer","children":[{"author":"mort96","children":[],"created_at":"2026-09-13T13:49:45.000Z","created_at_i":1789307385,"id":49684006,"options":[],"parent_id":49683879,"points":null,"story_id":49678969,"text":"I made no argument about guilt. I made an argument about the semantics of the word &quot;let&quot;.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T13:37:37.000Z","created_at_i":1789306657,"id":49683879,"options":[],"parent_id":49682936,"points":null,"story_id":49678969,"text":"If there is an escalator that is known for killing every 1000&#x27;s person using it then the operator who started it is more guilty than the folks using it for those deaths, don&#x27;t you think?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T11:49:19.000Z","created_at_i":1789300159,"id":49682936,"options":[],"parent_id":49682017,"points":null,"story_id":49678969,"text":".. what exactly depends on who started the escalator? My comment was in support of the argument that the word &quot;let&quot; does not imply agency on the part of the object in a sentence. Does the semantics of the word &quot;let&quot; depend on who started the escalator??","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T09:49:53.000Z","created_at_i":1789292993,"id":49682017,"options":[],"parent_id":49681868,"points":null,"story_id":49678969,"text":"Depends on who started the escalator.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T09:27:32.000Z","created_at_i":1789291652,"id":49681868,"options":[],"parent_id":49681820,"points":null,"story_id":49678969,"text":"Or, &quot;let the escalator keep going instead of pressing the emergency stop&quot;.","title":null,"type":"comment","url":null},{"author":"rightnutwingjob","children":[{"author":"RandomLensman","children":[],"created_at":"2026-09-13T09:36:17.000Z","created_at_i":1789292177,"id":49681923,"options":[],"parent_id":49681902,"points":null,"story_id":49678969,"text":"Big if.","title":null,"type":"comment","url":null},{"author":"Geezus_42","children":[],"created_at":"2026-09-13T09:55:08.000Z","created_at_i":1789293308,"id":49682051,"options":[],"parent_id":49681902,"points":null,"story_id":49678969,"text":"Except they&#x27;re nowhere near that and LLMs never will be.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T09:33:08.000Z","created_at_i":1789291988,"id":49681902,"options":[],"parent_id":49681820,"points":null,"story_id":49678969,"text":"At this stage, that seems like a distinction without a difference.<p>If the robots obtain sovereign nationhood, and are able to self-sustain, then <i>autonomous robot decides for itself</i> will be a valid argument.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T09:21:21.000Z","created_at_i":1789291281,"id":49681820,"options":[],"parent_id":49681524,"points":null,"story_id":49678969,"text":"&quot;let them&quot; in this use understood as: &quot;let the while loop run indefinitely&quot; as opposed to letting some autonomous robot decide for itself","title":null,"type":"comment","url":null},{"author":"nutjob2","children":[{"author":"Forgeties79","children":[],"created_at":"2026-09-13T14:58:44.000Z","created_at_i":1789311524,"id":49684721,"options":[],"parent_id":49681933,"points":null,"story_id":49678969,"text":"Seriously I don\u2019t even understand how this is a debate. If it\u2019s your tool, you are liable for what happens with it.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T09:38:34.000Z","created_at_i":1789292314,"id":49681933,"options":[],"parent_id":49681524,"points":null,"story_id":49678969,"text":"&quot;I left the car in neutral and left the park brake off and let the car roll down the hill.&quot;<p>The car doesn&#x27;t have agency, it&#x27;s doing what it naturally does. LLMs are the same, they&#x27;re working as designed.<p>But I don&#x27;t understand the point of splitting hairs. You are always responsible for the actions of your devices, tools, machinery, software, employees, whatever.<p>Trying to blame AI for one&#x27;s own stupidity must be aggressively pushed back on at all times.","title":null,"type":"comment","url":null},{"author":"helloplanets","children":[{"author":"AnimalMuppet","children":[],"created_at":"2026-09-13T17:14:43.000Z","created_at_i":1789319683,"id":49686190,"options":[],"parent_id":49682004,"points":null,"story_id":49678969,"text":"Whether or not the AI has intelligence, the one thing that&#x27;s clear is that it has <i>terrible judgment</i>.  I would regard that as empirically proven.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T09:47:29.000Z","created_at_i":1789292849,"id":49682004,"options":[],"parent_id":49681524,"points":null,"story_id":49678969,"text":"Yes.<p>OpenAI could have done this same experiment with GPT-4, with possibly even worse results, depending on the quality of the sandbox. Even if the techniques used were not as sophisticated, the natural language output could still easily contain more unhinged sequences of words that lead to the techniques being used.<p>If the system generates strange conclusions as to when the task is done, or should be stopped, it wouldn&#x27;t speak to the intelligence inherent to the system.<p>Not that the techniques used by the LLMs in the actual incident weren&#x27;t unexpectedly sophisticated, but the outputs of each and every one of these processes could&#x27;ve been read at <i>any time</i> during the run. They just weren&#x27;t.","title":null,"type":"comment","url":null},{"author":"weego","children":[],"created_at":"2026-09-13T10:02:37.000Z","created_at_i":1789293757,"id":49682101,"options":[],"parent_id":49681524,"points":null,"story_id":49678969,"text":"The parallel to the entire narrative would be if Smith &amp; Wesson claimed that one of their machine guns just started aiming and firing at people out of a window at their factory and then said &#x27;we can&#x27;t stop it! This is just how good our guns are!&#x27;<p>But into today&#x27;s AI climate it&#x27;s becoming increasingly difficult to figure out who is shilling, who is being assinine and who actually believes AI could do these things without clear human instruction and enabling.","title":null,"type":"comment","url":null},{"author":"DrewADesign","children":[{"author":"theodric","children":[{"author":"xg15","children":[{"author":"DrewADesign","children":[{"author":"dr_dshiv","children":[{"author":"fn-mote","children":[{"author":"twobitshifter","children":[],"created_at":"2026-09-13T17:47:41.000Z","created_at_i":1789321661,"id":49686570,"options":[],"parent_id":49683564,"points":null,"story_id":49678969,"text":"They have shown that your mind is doing exactly that due to the delays in consciousness. There are very simple examples that you can try to see it. It\u2019s especially clear in perception.<p><a href=\"https:&#x2F;&#x2F;discoverwildscience.com&#x2F;neuroscience-says-the-brain-predicts-the-next-few-seconds-of-your-life-before-they-actually-happen-1-409013&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;discoverwildscience.com&#x2F;neuroscience-says-the-brain-...</a><p>It\u2019s interesting that our conscious interpreter doesn\u2019t let us know that this is going on like you are experiencing, it must be that it\u2019s advantageous for us to not think about the prediction part of our mind.","title":null,"type":"comment","url":null},{"author":"drfloyd51","children":[{"author":"schrodinger","children":[],"created_at":"2026-09-13T19:52:50.000Z","created_at_i":1789329170,"id":49687983,"options":[],"parent_id":49686742,"points":null,"story_id":49678969,"text":"This is a fascinating illustration that I can&#x27;t help but agree with. However I feel like there&#x27;s something more \u2014 that this part of my brain is a bunch of supportive background processes running without my real awareness. It&#x27;s how I can drive home safely with no memory of how I got there (\u2026sober), even though driving is an action that&#x27;s incredibly demanding of intelligence. I can be driving home while thinking about a really hard problem at work that I haven&#x27;t solved.<p>However, if I came around a corner and saw a car in the wrong lane, a tree across the road, a fire raging \u2014 I&#x27;d very quickly jump into the mental driver&#x27;s seat and turn my conscious intelligence fully at this problem and come up with the best possible outcome I can think of in a short period of time \u2014 losing all ability to think about that work problem. I&#x27;d remember that incident for sure.<p>Similarly, in your story, all those predictive moments are happening below the person&#x27;s level of consciousness. They&#x27;re possibly even speaking to the group about a problem at the same time and thinking deeply about something.<p>I&#x27;m not smart enough to know, but I tend to feel like LLMs are much more like the predictive part of our thinking that you described, but that human cognition has something more \u2014 the single-threaded, creative, problem-solving part that is very conscious.<p>Is it possible that LLMs represent only one part of the way we think? And there&#x27;s a whole separate mechanism that&#x27;s fundamentally different, and not based on pattern matching and prediction?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T18:01:20.000Z","created_at_i":1789322480,"id":49686742,"options":[],"parent_id":49683564,"points":null,"story_id":49678969,"text":"If someone in that meeting quickly raised a hand in an arc, you would notice the \u201cabout to throw something\u201d pattern, look and notice the hand holds an eraser, analyze the arc and predict possible flight paths of the eraser. Then possibly notice the hand is now holding its position and the owner is actually looking down at the table. Maybe to squash somethingMust be something on the table. Maybe a spider! Better look. Wait now many people are moving away, oh someone spilt some water and the eraser is actually the guys phone and he is checking to see if his laptop is safe from the spilt water.<p>Fortunately you are on the other side of the table and predict the water isn\u2019t going to splash for otherwise flow onto your stuff.<p>All your possible responses result in you tossing a napkin towards the spill.<p>Our brains are always pattern matching and predicting. I bet you tried to reason out where I was going with my comment before you finished reading it.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T13:06:49.000Z","created_at_i":1789304809,"id":49683564,"options":[],"parent_id":49683443,"points":null,"story_id":49678969,"text":"&gt; Humans are constantly predicting the next moment<p>This is really not my experience of consciousness.<p>Is it yours??<p>Do you sit in meetings predicting what\u2019s going to happen next? No, you sit there bored out of your f$$@ing mind, daydreaming about being somewhere else and doing something useful with your life.<p>God help me if that\u2019s what LLMs are doing when I ask them to build me a web site.","title":null,"type":"comment","url":null},{"author":"oblio","children":[],"created_at":"2026-09-13T13:16:49.000Z","created_at_i":1789305409,"id":49683636,"options":[],"parent_id":49683443,"points":null,"story_id":49678969,"text":"&gt; We have no better model for how human decision making works than LLMs<p>We do have some models and guess what, they&#x27;re based on simpler animals. Which is most likely the better model.<p>Some other models are based on neurosciences, because we can track electrical activity.","title":null,"type":"comment","url":null},{"author":"DrewADesign","children":[],"created_at":"2026-09-13T13:26:57.000Z","created_at_i":1789306017,"id":49683760,"options":[],"parent_id":49683443,"points":null,"story_id":49678969,"text":"&gt; We have no better model for how human decision making works than LLMs<p>This is a claim that requires a lot of citations.<p>&gt; biologically inspired<p>Nature inspires a lot of creation, but superficial similarities don\u2019t mean other aspects are similar. Making an extremely realistic sculpture of a souffl\u00e9, even using a foam medium, doesn\u2019t bring me any closer to being a chef, doesn\u2019t mean I know anything about albumen foams, sauces, and heat transfer, and it doesn\u2019t bring me any closer to having dinner ready. Browning on top of a souffl\u00e9 is evidence of maillardization. You could pull up some studies on that and claim the brown on top of the souffl\u00e9 sculpture, which I applied with an airbrush, proved that the Maillard reaction was occurring, and if another person didn\u2019t know anything about cooking, they might even believe you. It would, of course, be completely wrong. And the other person, of course, could loudly exclaim that I can\u2019t prove that there was no maillardization.","title":null,"type":"comment","url":null},{"author":"teiferer","children":[{"author":"psma_egeliaa","children":[{"author":"DrewADesign","children":[],"created_at":"2026-09-13T13:53:25.000Z","created_at_i":1789307605,"id":49684034,"options":[],"parent_id":49683997,"points":null,"story_id":49678969,"text":"It sucks because I think analogies can be useful in helping people make a mental model of complex things, which is meaningfully beneficial. The problems happen when people aren\u2019t honest about the limits of the analogies, which is damned-near guaranteed to happen with this stuff.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T13:48:48.000Z","created_at_i":1789307328,"id":49683997,"options":[],"parent_id":49683813,"points":null,"story_id":49678969,"text":"This is why I hate analogies. They&#x27;re almost always either relevant or inapposite depending on the level of generalization we&#x27;re operating on.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T13:32:22.000Z","created_at_i":1789306342,"id":49683813,"options":[],"parent_id":49683443,"points":null,"story_id":49678969,"text":"&gt; We certainly have a different \u201ctokenizer\u201d and training set, but many of the concepts underpinning LLMs are both biologically inspired and, likely, have similar consequences and emergent architectures.<p>My kids tricycle certainly has a different gear setup and wheel diameter, but many of the concepts underpinning the tricycle are both inspired by F1 race car enineering and, likely, have similar consequences and emergent architectures.<p>Or have they?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T12:52:52.000Z","created_at_i":1789303972,"id":49683443,"options":[],"parent_id":49683339,"points":null,"story_id":49678969,"text":"Disagree about the burden of proof. We have no better model for how human decision making works than LLMs. Humans are constantly predicting the next moment. We certainly have a different \u201ctokenizer\u201d and training set, but many of the concepts underpinning LLMs are both biologically inspired and, likely, have similar consequences and emergent architectures.","title":null,"type":"comment","url":null},{"author":"iugtmkbdfil834","children":[],"created_at":"2026-09-13T12:53:37.000Z","created_at_i":1789304017,"id":49683447,"options":[],"parent_id":49683339,"points":null,"story_id":49678969,"text":"&lt;&lt; We do know that zero parts of human decision making are based on predicting the next most likely letter based on a giant internet-based database.<p>Oh man, how much did you read on tip of the tongue?","title":null,"type":"comment","url":null},{"author":"fl7305","children":[{"author":"DrewADesign","children":[{"author":"fl7305","children":[],"created_at":"2026-09-13T17:24:54.000Z","created_at_i":1789320294,"id":49686300,"options":[],"parent_id":49685333,"points":null,"story_id":49678969,"text":"&gt; You can try to say that I\u2019m arguing whatever you like.<p>I did my honest best possible interpretation of what you really meant from what you wrote.<p>&gt;&gt; We do know that zero parts of human decision making are based on predicting the next most likely letter based on a giant internet-based database. We do know that\u2019s what LLMs do.<p>I read this as &quot;The decision making of LLMs are based on predicting the next most likely letter based on a giant internet-based database.&quot;<p>Is that wrong?<p>I understood that your meaning was something like &quot;LLMs can&#x27;t reason, they just output likely letters&quot;?<p>&gt; If you\u2019re claiming that the underlying structure of digital so-called neural networks is comparable to biological neural networks<p>No, I don&#x27;t claim that.<p>What do claim is this: Regardless of how the LLMs were trained, they show overwhelming signs of being able to reason, and not just recall memorized information.<p>This doesn&#x27;t mean that they always reason perfectly about everything.<p>But if they only memorized things and output the next likely letter, you would see them answering very badly much more often.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T15:50:28.000Z","created_at_i":1789314628,"id":49685333,"options":[],"parent_id":49684747,"points":null,"story_id":49678969,"text":"You can try to say that I\u2019m arguing whatever you like. If you\u2019re claiming that the underlying structure of digital so-called neural networks is comparable to biological neural networks\u2014 which we\u2019ve studied for far longer without really understanding\u2014 no amount of jargon will obviate the \u2018citation needed\u2019 requirement for that claim.","title":null,"type":"comment","url":null},{"author":"talon8635","children":[],"created_at":"2026-09-13T16:00:31.000Z","created_at_i":1789315231,"id":49685457,"options":[],"parent_id":49684747,"points":null,"story_id":49678969,"text":"You are more convincing than the person you\u2019re responding to.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T15:00:52.000Z","created_at_i":1789311652,"id":49684747,"options":[],"parent_id":49683339,"points":null,"story_id":49678969,"text":"&gt; ... are based on predicting the next most likely letter based on a giant internet-based database. We do know that\u2019s what LLMs do.<p>If you&#x27;re claiming that the training objective tells us what kind of internal mechanisms the training produced, then I think that&#x27;s just plain wrong.<p>Next-token prediction describes the optimization target, not the internal mechanisms that the training produced.<p>In the same way for the natural evolution of humans, DNA replication is the evolutionary objective. It&#x27;s not a description of the internal mechanisms that evolution has produced.<p>As an example, we know that neural networks can be trained to develop generalized algorithms for arithmetic.<p>They might first memorize the training examples, then with further training transition to a solution that generalizes correctly to unseen examples.<p>In some cases we&#x27;ve even reverse-engineered the evolved internal mechanisms and found structured arithmetic algorithms rather than rote memorization. Interestingly, for modular addition this can involve Fourier representations, which isn&#x27;t an algorithm I would have guessed gradient descent training of neural networks would produce.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T12:38:32.000Z","created_at_i":1789303112,"id":49683339,"options":[],"parent_id":49682712,"points":null,"story_id":49678969,"text":"Can\u2019t agree with you here.<p>&gt; I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward.<p>In love how people get salty about people not going along with a superficial supposition just because they can\u2019t definitively prove it wrong.<p>&gt; All that while still not knowing how either kind actually works.<p>We do know that zero parts of human decision making are based on predicting the next most likely letter based on a giant internet-based database. We do know that\u2019s what LLMs do. We do know exactly how each part of an LLM works even if the combined behavior is too cryptic to feasibly analyze at the moment. We do not understand all of the functions of an actual neuron. Openworm isn\u2019t even close to accurately simulating the 302 neurons of a roundworm and you\u2019d need over 200 million roundworms <i>working in conjunction</i> to equal the number of neurons in one human brain.<p>My dog seems convinced that the malevolent invader in a mailman uniform would break in and attack us if she didn\u2019t fiercely bark at him, six days per week. I certainly can\u2019t prove the mailman doesn\u2019t want to kill us, and that the mailman wasn\u2019t solely deterred by her barking. Empirically, the mailman goes away soon after she starts barking, and we\u2019ve sustained <i>zero</i> mailman assaults after hundreds of purported attempts. Maybe I should just run with it? Her model is too simple to come up with the obviously correct answer, but it\u2019s not even <i>directionally </i> accurate.<p>The burden of proof is on the person making the claim, which in this case, is that these comparatively simple logical constructs are remotely comparable to the complexity of biological systems.","title":null,"type":"comment","url":null},{"author":"cj","children":[{"author":"DrewADesign","children":[],"created_at":"2026-09-13T14:32:18.000Z","created_at_i":1789309938,"id":49684432,"options":[],"parent_id":49683345,"points":null,"story_id":49678969,"text":"Personally, I don\u2019t think it\u2019s different from any other faith-based motivation.","title":null,"type":"comment","url":null},{"author":"talon8635","children":[{"author":"cj","children":[{"author":"talon8635","children":[],"created_at":"2026-09-13T16:35:07.000Z","created_at_i":1789317307,"id":49685793,"options":[],"parent_id":49685636,"points":null,"story_id":49678969,"text":"That\u2019s fair enough, but you\u2019re elegance and nuance doesn\u2019t reflect what I\u2019ve seen from that side of the debate","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T16:16:35.000Z","created_at_i":1789316195,"id":49685636,"options":[],"parent_id":49685523,"points":null,"story_id":49678969,"text":"&gt; It seems like raw egotistical hubris.<p>1) Humans have a bias &#x2F; tendency to attribute human qualities to things that appear or act human, but aren\u2019t.<p>2) When that happens, people jump to conclusions by stretching the human analogy too far.<p>3) Since humans have a bias to do this, we should have a bias against anthropomorphising LLMs.<p>It\u2019s easier to believe LLMs act like humans because there\u2019s so much evidence to support that. You have to actively use your brain to convince yourself otherwise. Another reason why we should have a bias against using human behavior to describe LLM behavior.<p>But I agree. \u201cI don\u2019t know\u201d is a good stance. But I think \u201cI don\u2019t know, probably not\u201d is a better stance if only to combat our (or at least my) natural bias.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T16:05:49.000Z","created_at_i":1789315549,"id":49685523,"options":[],"parent_id":49683345,"points":null,"story_id":49678969,"text":"Does the argument require benefit? Isn\u2019t the argument based on caution?<p>I haven\u2019t heard many people explicitly saying \u201cthese things behave like humans\u201d, but more generally \u201cwe don\u2019t even know how to define human consciousness, we don\u2019t have a thorough grasp of how the brain works, we are still very much in the dark on a lot of these topics, so how can we say one way or the other?\u201d<p>In other words, agnosticism: I don\u2019t know.<p>In general, it\u2019s baffling to me that anyone has an unshakable opinion on what exactly is happening. It seems like raw egotistical hubris.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T12:39:51.000Z","created_at_i":1789303191,"id":49683345,"options":[],"parent_id":49682712,"points":null,"story_id":49678969,"text":"I always wonder what makes people take the other side of this argument. They do it quite passionately. Why actively encourage viewing LLMs as human? Who is that benefitting?","title":null,"type":"comment","url":null},{"author":"dnautics","children":[],"created_at":"2026-09-13T14:34:08.000Z","created_at_i":1789310048,"id":49684452,"options":[],"parent_id":49682712,"points":null,"story_id":49678969,"text":"&gt; LLM decisionmaking cannot possibly be like human decisionmaking<p>I mean how can it possibly be like human decisionmaking?  It&#x27;s not like it&#x27;s trained on human data","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T11:21:59.000Z","created_at_i":1789298519,"id":49682712,"options":[],"parent_id":49682355,"points":null,"story_id":49678969,"text":"I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward.<p>All that while still not knowing how either kind actually works.","title":null,"type":"comment","url":null},{"author":"thunky","children":[{"author":"visarga","children":[],"created_at":"2026-09-13T14:38:07.000Z","created_at_i":1789310287,"id":49684488,"options":[],"parent_id":49683248,"points":null,"story_id":49678969,"text":"&gt;&gt; We need new words!<p>Can make distinctions and can choose actions - applies to both humans and AI. I&#x27;d replace &#x27;agency&#x27; with &#x27;distinction &amp; choice&#x27; language.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T12:27:44.000Z","created_at_i":1789302464,"id":49683248,"options":[],"parent_id":49682355,"points":null,"story_id":49678969,"text":"&gt; We need new words!<p>The words we have are fine.<p>We just need to assign liability by ownership&#x2F;initiation: if your &quot;agent&quot; destroys something, even though you didn&#x27;t tell it to (because it had &quot;agency&quot;), you should be liable for the damages.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T10:34:19.000Z","created_at_i":1789295659,"id":49682355,"options":[],"parent_id":49682219,"points":null,"story_id":49678969,"text":"The idea that the agent does not actually have agency is rather discordant. We need new words!","title":null,"type":"comment","url":null},{"author":"xg15","children":[{"author":"schrodinger","children":[],"created_at":"2026-09-13T19:22:49.000Z","created_at_i":1789327369,"id":49687675,"options":[],"parent_id":49682837,"points":null,"story_id":49678969,"text":"I agree. If I were setting up an experiment like this, I&#x27;d have instrumented the hell out of it to see all actions taken in real time, and have a team of folks watching it. This team would have seen the anomalous GET requests to a German wiki and taken action (e.g. halt the system to investigate and decide whether to abort).<p>In fact, that feels so obvious it&#x27;s ridiculous it needs to be said. It&#x27;s table stakes. When do you run a production system without monitoring and a team on-call?<p>It&#x27;s hard to imagine another field in which this reckless behavior would be tolerated.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T11:34:56.000Z","created_at_i":1789299296,"id":49682837,"options":[],"parent_id":49682219,"points":null,"story_id":49678969,"text":"Yeah, fully agreed here. Most automation (such as riding a lawnmower and <i>not</i> putting a brick on the gas) is deterministic, in the sense that you can reasonably understand what exactly the machine will do when you run it.<p>But some automation is different. The most prominent example before AI would be car navigation systems, where the entire idea is that that you give it a destination and it figures out the exact actions to get there on its own.<p>Except even there, the actual driver would still have been you - giving you a chance to vet and deny every turn the system proposed.<p>AI agents are sort of like that - most of the value they provide is in the ability to turn high-level goals (&quot;write me a traffic control system for my model railway&quot;) into low-level actions and also do so interactively.<p>The new thing is that the &quot;driver&quot; has much less oversight here where the agent wants to go, and is sometimes removed completely. That part is clearly be an active decision by AI labs.<p>The other thing is that the labs seem increasingly to steer their training towards behavior that make events like this one more likely, e.g. that agents should never &quot;give up&quot; when faced with a seemingly impossible task, but instead should keep trying and think of increasingly outlandish ways to solve the task. To me, that seems pretty much a recipe to get incidents like this.","title":null,"type":"comment","url":null},{"author":"bsenftner","children":[{"author":"teiferer","children":[{"author":"brazukadev","children":[{"author":"DrewADesign","children":[{"author":"brazukadev","children":[],"created_at":"2026-09-13T15:53:16.000Z","created_at_i":1789314796,"id":49685366,"options":[],"parent_id":49684307,"points":null,"story_id":49678969,"text":"I appreciate the message. I do have some ideas around sandboxing that I could not yet turn into a product, maybe I should take them seriously.","title":null,"type":"comment","url":null},{"author":"schrodinger","children":[],"created_at":"2026-09-13T19:35:34.000Z","created_at_i":1789328134,"id":49687815,"options":[],"parent_id":49684307,"points":null,"story_id":49678969,"text":"This is a great point, analogous to the <a href=\"https:&#x2F;&#x2F;boringtechnology.club&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;boringtechnology.club&#x2F;</a> philosophy I&#x27;ve come to love.<p>There&#x27;s no reason we need to make an incredibly intelligent shell execution engine that can identify patterns that seem evil and may represent unwanted behavior to solve this problem. Simply limiting the available tools to a finite, known, ironclad-secure set (even if it&#x27;s quite sprawling) is sufficient.<p>LLMs will still find workarounds \u2014 from what I understand, a large part of the issue in this situation was that an agent was presumed to have read-only Internet access because it could only make GET requests. It should be pretty obvious that there&#x27;s at least one website on the Internet that allows writes via GET. I think this is where auditing comes in, and a live team of people watching tool calls would have noticed the strange behavior.<p>But I think a lot of times people jump to overly complex solutions when simple, well-bounded ones would work just fine. Yes, the intelligent shell is a great goal, but it&#x27;s akin to solving the halting problem.<p>This philosophy is what I love about PicoClaw (<a href=\"https:&#x2F;&#x2F;github.com&#x2F;sipeed&#x2F;picoclaw\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;sipeed&#x2F;picoclaw</a>), and incidentally the philosophy behind Go and even *nix in general (i.e. provide small, composable, single-purpose tools).","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T14:19:50.000Z","created_at_i":1789309190,"id":49684307,"options":[],"parent_id":49683998,"points":null,"story_id":49678969,"text":"&gt; I don&#x27;t think knowing that will make me rich.<p>As someone who\u2019s not really sure that any of this is sustainable, I\u2019d implore you to not sell yourself short. I reckon there\u2019s a ton of dogma and nearly religious zeal among these companies, which among some people is earnest, and among others is cynical hype farming. I\u2019ll bet someone objective enough to focus on using available tooling to solve real problems in practical ways that mitigate actual risks and are honest about actual limitations will be eBay here while the others are going to be somewhere between lucent and pets.com.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T13:48:50.000Z","created_at_i":1789307330,"id":49683998,"options":[],"parent_id":49683761,"points":null,"story_id":49678969,"text":"an agent doesn&#x27;t come with &quot;jailbreak&quot; capability. It needs tools, specially one that runs shell commands. Don&#x27;t give it shell commands, it won&#x27;t be able to run shell commands.<p>You can still give it plenty of tools like create files, list files, write to files, translate text, edit a video. I don&#x27;t think knowing that will make me rich.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T13:27:14.000Z","created_at_i":1789306034,"id":49683761,"options":[],"parent_id":49683321,"points":null,"story_id":49678969,"text":"What is your approach to create jailbreak incapable agents?<p>I think the world is looking for a way right now, so if yours works you&#x27;ll get very rich, or at least very famous.","title":null,"type":"comment","url":null},{"author":"visarga","children":[],"created_at":"2026-09-13T14:33:33.000Z","created_at_i":1789310013,"id":49684446,"options":[],"parent_id":49683321,"points":null,"story_id":49678969,"text":"&gt; Why, oh why, are we not discussion how to create and frame models so they do our complex work and their &quot;jailbreaking&quot; is simply not possible?<p>Good idea, and after that let&#x27;s make guns that only kill bad people. Let&#x27;s focus on the frozen component (the model) and ignore the dynamics around them - humans and other systems they interact with.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T12:36:10.000Z","created_at_i":1789302970,"id":49683321,"options":[],"parent_id":49682219,"points":null,"story_id":49678969,"text":"Okay, so we know OpenAI and Anthropic are operating a propagandists in respect to how they describe their models and the behavior of those models. We also know it is how they use and frame their use to their models that is the problem, that and they use misaligned and guardrails disabled models for these press incidents.\nWhy, oh why, are we not discussion how to create and frame models so they do our complex work and their &quot;jailbreaking&quot; is simply not possible?<p>I, of course, have my own means of creating jailbreak incapable agents, but rather than a storm of downvotes on my idea, what is yours? Let&#x27;s discuss this, because this is thee real question. Not why, but how to make then not?!","title":null,"type":"comment","url":null},{"author":"dr_dshiv","children":[{"author":"huntertwo","children":[{"author":"skinner_","children":[{"author":"jwynot","children":[{"author":"alluro2","children":[{"author":"schrodinger","children":[],"created_at":"2026-09-13T18:59:12.000Z","created_at_i":1789325952,"id":49687426,"options":[],"parent_id":49686535,"points":null,"story_id":49678969,"text":"I fully agree with you, but would go one step further: I think it&#x27;s clear that we need to pierce the corporate veil and ascribe responsibility to _people_, not just &quot;OpenAI the entity&quot;, full stop.<p>Executives should fear being perp-walked and thrown in jail for the actions of irresponsible &quot;tests&quot; of their models in the real world, as they&#x27;re ultimately accountable.<p>Sure, there&#x27;s a lot of nuance to work out, but I think we could likely even _start_ there today even with existing laws and pretty quickly &quot;align&quot; on more intricate legal frameworks to handle true accidents, distribution of responsibility, etc.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T17:44:21.000Z","created_at_i":1789321461,"id":49686535,"options":[],"parent_id":49684070,"points":null,"story_id":49678969,"text":"It&#x27;s a good example.<p>If I hired a hitman to murder someone, and they broke into a private property and stole something so that they can action the murder (which I didn&#x27;t know about or pay them to do), I would be guilty of conspiracy to commit murder, but not for the theft part.<p>Likely because that person is a human, is aware of societal and legal norms, and is responsible for their actions due to their participation in human society. (I am not a lawyer (if it wasn&#x27;t painfully obvious so far) so in layman terms, I hope good definitions for all of this exist formally)<p>AI is not a person - it cannot easily discern between &quot;right&quot; and &quot;wrong&quot; in non-strictly-defined sense, and is not subject to human norms and responsibility. So if I use AI to achieve goal A, either I, or the maker of AI, are fully responsible for anything that happens while AI is trying to achieve the goal given by me.<p>Now, here, &quot;I&quot; in the example is OpenAI, who is simultaneously the maker of the AI. So it seems pretty obvious who is the only entity that can be responsible.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T13:58:12.000Z","created_at_i":1789307892,"id":49684070,"options":[],"parent_id":49684061,"points":null,"story_id":49678969,"text":"Hiring a hitman is conspiracy to commit murder.<p>The hitman is charged with murder.<p>I imagine the same could be true of an AI lab if you could prove intent.<p>With intent, they could be found guilty of conspiracy to commit a crime even if it was the end user who did it.<p>Source: Prosecuting attorney for over 30 years","title":null,"type":"comment","url":null},{"author":"alpinisme","children":[],"created_at":"2026-09-13T14:36:31.000Z","created_at_i":1789310191,"id":49684475,"options":[],"parent_id":49684061,"points":null,"story_id":49678969,"text":"The objection is not too far from criticisms of the use of passive voice: a man was injured at the factory vs a faulty saw blade snapped and injured a man vs after the company loosened safety inspection policies, etc.<p>Which way you say it shifts the framing. And it\u2019s not that one is less accurate to the facts, necessarily. It just is that one less aptly captures the moral and political relevance of the scenario.<p>For my part, I think it makes good sense to anthropomorphize in some contexts and not others. Generally when responsibility is at issue, you probably want the framing that tunes anthropomorphism down to near zero, since it\u2019s the human dimension you care about.","title":null,"type":"comment","url":null},{"author":"jmcgough","children":[],"created_at":"2026-09-13T16:15:56.000Z","created_at_i":1789316156,"id":49685627,"options":[],"parent_id":49684061,"points":null,"story_id":49678969,"text":"I think the danger of anthropomorphizing is that 99% of people lack the technical background to understand the nuance. People have been primed by pop culture depictions of AI to think of LLMs as intelligent, autonomous beings, which leads to dangerous assumptions.<p>We should make the distinction between them, because openai and anthropic will not. A magical black box that does the thinking for you is a much more compelling sales pitch.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T13:56:38.000Z","created_at_i":1789307798,"id":49684061,"options":[],"parent_id":49683613,"points":null,"story_id":49678969,"text":"Does it help if I explicitly add a disclaimer that the tool&#x27;s agency does not remove any responsibility from OpenAI, the wielder of the tool? I&#x27;m not sure why this disclaimer is necessary, though: hiring a hitman is a standard example.<p>BTW I anthropomorphize the tool because it&#x27;s an imitation of a human mind, inheriting the muddy ethics, survival instincts, and being prone to mass psychosis. The laser-sharp focus on reward seeking, that mostly came from reinforcement learning, a process more alien to humans.","title":null,"type":"comment","url":null},{"author":"linkjuice4all","children":[{"author":"voakbasda","children":[],"created_at":"2026-09-13T15:30:48.000Z","created_at_i":1789313448,"id":49685109,"options":[],"parent_id":49684438,"points":null,"story_id":49678969,"text":"FWIW, a robot that fires a weapon independently is considered an automatic weapon, and the ATF will want to have a word.  Have at it, but don\u2019t let anyone know!","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T14:32:33.000Z","created_at_i":1789309953,"id":49684438,"options":[],"parent_id":49683613,"points":null,"story_id":49678969,"text":"I get where you are coming from but this wasn\u2019t a tool just left laying around, this is similar to rigging up a booby trapped shot gun to your door and then claiming the victim is responsible.<p>If you build a robot that shoots a bunch of TVs in your back yard, have at it. But the second that thing goes off your property you\u2019re the one responsible.","title":null,"type":"comment","url":null},{"author":"xorcist","children":[{"author":"alluro2","children":[{"author":"schrodinger","children":[],"created_at":"2026-09-13T18:30:20.000Z","created_at_i":1789324220,"id":49687117,"options":[],"parent_id":49686549,"points":null,"story_id":49678969,"text":"You say this flippantly, but I think this is actually another very good example!<p>We even do it for obviously unintelligent inanimate objects. A rollercoaster ran too fast for its tracks, killing 10 people. In that sentence, the roller coaster is the subject which took an action and caused death \u2014 obviously the roller coaster is not ethically at fault here, the people who built the rollercoaster are at fault through negligence.<p>Although this example and the ones around cars both demonstrate how we tolerate some degree of &quot;accidents&quot; from humans as no-fault, which is fair. I wonder how that fits into this analogy? I suppose its all about intent (mens rea) and judgement: did they intend for the roller coaster to harm people, and should they have reasonably predicted that the accident was likely to happen.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T17:45:52.000Z","created_at_i":1789321552,"id":49686549,"options":[],"parent_id":49685591,"points":null,"story_id":49678969,"text":"And what if it&#x27;s a self-driving car? :)","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T16:12:21.000Z","created_at_i":1789315941,"id":49685591,"options":[],"parent_id":49683613,"points":null,"story_id":49678969,"text":"Right. Among bicycle advocacy groups it&#x27;s been well known for long time that <i>cars</i> do not run over people, <i>drivers</i> do.<p>The fact that we talk about a car running someone over, and this is the same in many different languages and countries, contributes to lower punishments for drivers. Clearly it was just an accident. He or she was run over by a car.<p>Now we see that same language tricks play out again every time an LLM did something illegal.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T13:13:39.000Z","created_at_i":1789305219,"id":49683613,"options":[],"parent_id":49683397,"points":null,"story_id":49678969,"text":"A contractor has agency and accountability - something that an LLM (or similarly, a nail gun or a hammer or a bot net) does not have. When you anthropomorphize a tool, you implicitly give it agency and remove responsibility from the wielder of the tool.","title":null,"type":"comment","url":null},{"author":"DrewADesign","children":[],"created_at":"2026-09-13T13:40:31.000Z","created_at_i":1789306831,"id":49683905,"options":[],"parent_id":49683397,"points":null,"story_id":49678969,"text":"Situations have lots of independent variables, Doctor, and Anthropomorphism is one problematic facet of many in the way this industry is pushing LLM products.<p>If there was a collision at an intersection with a stop sign partially obscured by a tree, that had traffic volume that would have better been served by a traffic light, on a foggy night, where one person was texting while driving, none of those things would diminish the fact that the other driver was drunk.","title":null,"type":"comment","url":null},{"author":"intended","children":[],"created_at":"2026-09-13T14:02:35.000Z","created_at_i":1789308155,"id":49684119,"options":[],"parent_id":49683397,"points":null,"story_id":49678969,"text":"These situations are novel. Lax terminology is fine when it has no impact on the intuitions, clarity and conclucions of discussion.<p>If this was a conversation just about outcomes, then whether models think or simulate thinking is sophistry. However, the bulk of the issue here is attributing responsibility, which relies on being clear about the underlying processes at play.<p>We are hard wired to assume certain priors and capabilites when it comes to &quot;human like&quot; behavior. Anthropomorphizing LLMs implies mechanisms that aren&#x27;t present, and end up distorting&#x2F;complicating discussion about the process.<p>It isn&#x27;t helped that the frontier labs, the experts in the room, generally use anthropomorphic terms to discuss model capability.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T12:48:51.000Z","created_at_i":1789303731,"id":49683397,"options":[],"parent_id":49682219,"points":null,"story_id":49678969,"text":"Why is anthropomorphism the problem here? If OpenAI hired a contractor and they did this, OpenAI or the contractor would still be liable, depending on the contract language.","title":null,"type":"comment","url":null},{"author":"jordanb","children":[{"author":"DrewADesign","children":[],"created_at":"2026-09-13T14:25:36.000Z","created_at_i":1789309536,"id":49684365,"options":[],"parent_id":49683946,"points":null,"story_id":49678969,"text":"Yeah I like Cal\u2019s take on it, though in this context I\u2019d argue LLMs have even less agency, and are even less deserving of anthropomorphization than a dog is.","title":null,"type":"comment","url":null},{"author":"visarga","children":[{"author":"oarsinsync","children":[{"author":"verzali","children":[],"created_at":"2026-09-13T16:32:34.000Z","created_at_i":1789317154,"id":49685768,"options":[],"parent_id":49685294,"points":null,"story_id":49678969,"text":"You forgot the part where you spend billions to put your man in a position of power.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T15:47:38.000Z","created_at_i":1789314458,"id":49685294,"options":[],"parent_id":49684413,"points":null,"story_id":49678969,"text":"&gt; &gt; It&#x27;s not the weed wacker&#x27;s fault or even the dog&#x27;s fault when someone got hurt, it&#x27;s the fault of the guy who put a weed wacker on a dog and let it run wild.<p>&gt; The difference is volume. They spent hundreds of billions of tokens on these agents. If you put &quot;a million weed whackers on dog backs&quot; you would see the difference.<p>So put one weed whacker on one dog, you&#x27;re to blame. Put a million weed whackers on a million dogs backs and ... you&#x27;re still to blame? Arguably even more so?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T14:30:26.000Z","created_at_i":1789309826,"id":49684413,"options":[],"parent_id":49683946,"points":null,"story_id":49678969,"text":"The difference is volume. They spent hundreds of billions of tokens on these agents. If you put &quot;a million weed whackers on dog backs&quot; you would see the difference.<p>We also run agents, but for shorter spans between supervisions, and with much lower total budget.","title":null,"type":"comment","url":null},{"author":"ryankrage77","children":[{"author":"singpolyma3","children":[{"author":"mock-possum","children":[{"author":"jordanb","children":[],"created_at":"2026-09-13T17:29:59.000Z","created_at_i":1789320599,"id":49686356,"options":[],"parent_id":49686107,"points":null,"story_id":49678969,"text":"I think Cal&#x27;s point is that while dogs may do these things it reduces to a set of behaviors where maybe 90 percent of them are beneficial to the dog-weedwacker system and the remaining 10 percent are really unfortunate.<p>We can&#x27;t really know the dogs inner life so we just kinda have to reduce it to a set of behaviors selected stocastically.<p>The dog meanwhile has no ability to understand the weedwacker or what it&#x27;s doing on its back.<p>So when the guy puts the weedwacker on the dog and the dog predicably does dog things and that results in disaster the guy isn&#x27;t able to clutch his pearls and say &quot;I guess the system broke containment!&quot;","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T17:05:24.000Z","created_at_i":1789319124,"id":49686107,"options":[],"parent_id":49686005,"points":null,"story_id":49678969,"text":"Not if you have a dog. They\u2019re very clearly aware, have minds, have thoughts, gather information, make decisions, second guess themselves, reflect on the immediate consequences, change their mind\u2026 and most importantly, you can observe them doing all these things undirected, while left alone with their thoughts.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T16:54:10.000Z","created_at_i":1789318450,"id":49686005,"options":[],"parent_id":49685670,"points":null,"story_id":49678969,"text":"Dogs have agency and can choose? That seems like a rather uncommon take on dogs...","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T16:20:52.000Z","created_at_i":1789316452,"id":49685670,"options":[],"parent_id":49683946,"points":null,"story_id":49678969,"text":"It goes even further though, as the dog does have agency. It can choose to run and around chase squirrels with no human intervention.<p>An LLM on the other hand, is just inert data on disk until a human takes deliberate action to run it and prompt it.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T13:44:28.000Z","created_at_i":1789307068,"id":49683946,"options":[],"parent_id":49682219,"points":null,"story_id":49678969,"text":"Cal Newport has an analogy to &quot;putting a weed wacker on a dog&#x27;s back to mow your lawn.&quot; The dog will wander around the yard and it may mow the lawn, but the dog will also chase after birds or run up to visitors for pets and the weed wacker could do a lot of damage.<p>It&#x27;s not the weed wacker&#x27;s fault or even the dog&#x27;s fault when someone got hurt, it&#x27;s the fault of the guy who put a weed wacker on a dog and let it run wild.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T10:17:32.000Z","created_at_i":1789294652,"id":49682219,"options":[],"parent_id":49681524,"points":null,"story_id":49678969,"text":"I do get their usage intent. If something is at all automated, in English, we often refer to as having some amount of agency. If I started up a riding lawnmower, put a brick on the gas and pointed t it towards a field, many might say I \u201clet it run rampant.\u201d But since nobody is at risk of anthropomorphizing riding lawnmowers, it\u2019s not problematic.<p>Anthropomorphizing LLMs is a huge fucking problem though and I, personally, think we should expunge all of these casual inadvertent linguistic agency affordances with <i>great prejudice.</i><p>OpenAI didn\u2019t \u2018let\u2019 these bots do this any more than someone \u2018let\u2019 Claude Code make them a website.","title":null,"type":"comment","url":null},{"author":"WithinReason","children":[{"author":"tesnorindian","children":[],"created_at":"2026-09-13T11:36:47.000Z","created_at_i":1789299407,"id":49682851,"options":[],"parent_id":49682259,"points":null,"story_id":49678969,"text":"Exactly that is the point, your nailed it. The models were taught to hack and were rewarded for doing it. They would claim they are trained as ethical hackers.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T10:22:16.000Z","created_at_i":1789294936,"id":49682259,"options":[],"parent_id":49681524,"points":null,"story_id":49678969,"text":"It&#x27;s worse, their reinforcement learning loops (implicitly) rewarded the agents for cheating (i.e. hacking) when they were being trained.","title":null,"type":"comment","url":null},{"author":"Forgeties79","children":[],"created_at":"2026-09-13T12:56:20.000Z","created_at_i":1789304180,"id":49683476,"options":[],"parent_id":49681524,"points":null,"story_id":49678969,"text":"They\u2019re firing a gun in a room of people and going \u201cwow isn\u2019t it wild what a gun will do if we let it do its thing?\u201d","title":null,"type":"comment","url":null},{"author":"Waterluvian","children":[],"created_at":"2026-09-13T13:20:01.000Z","created_at_i":1789305601,"id":49683669,"options":[],"parent_id":49681524,"points":null,"story_id":49678969,"text":"My pitbull is a good dog. Sure, it&#x27;s been carefully designed to be an incredibly dangerous and violent pit fighter, but I didn&#x27;t actually ask it to eat any faces.","title":null,"type":"comment","url":null},{"author":"smegger001","children":[],"created_at":"2026-09-13T13:29:11.000Z","created_at_i":1789306151,"id":49683778,"options":[],"parent_id":49681524,"points":null,"story_id":49678969,"text":"They deliberately trained the models in how to use various hacking tools, didn&#x27;t give them the standard alignment training let them know where the answer key was left the models with access to said tool and told them to maximize their score then left them unsupervised for days with internet access (yeah they were sandboxed but again handed hacking tools and the training to use them if they really did want them to access the internet you wouldn&#x27;t plug in the Ethernet cable) they wanted this to happen","title":null,"type":"comment","url":null},{"author":"theptip","children":[],"created_at":"2026-09-13T15:47:55.000Z","created_at_i":1789314475,"id":49685297,"options":[],"parent_id":49681524,"points":null,"story_id":49678969,"text":"Have you read the METR transcripts? \u201cJust a tool\u201d is a suicidally insufficient description of what these models are doing.<p>Recognizing that the models are acting with intent does not somehow absolve OpenAI from their felony hacking. We have not granted them personhood.","title":null,"type":"comment","url":null},{"author":"Teever","children":[],"created_at":"2026-09-13T16:08:10.000Z","created_at_i":1789315690,"id":49685549,"options":[],"parent_id":49681524,"points":null,"story_id":49678969,"text":"Yes and no.<p>If your buddy leaves his car parked at the top of a hill without the parking brake on and it rolls down the hill and side-swipes a bunch of vehicles and narrowly misses an elderly person walking by with a cane someone could easily say:<p>&quot;Dude wtf is wrong with you, you left your car parked on the top of a hill with no brake and let it roll into traffic&quot;<p>The phrasing doesn&#x27;t absolve the offender of their negligent behaviour and the consequences of it.<p>The only thing thing does is the lack of action from regulators and society writ large.<p><i>Our</i> lack of action is what allows people like Sam Altman and Dario and the irresponsible people who <i>choose</i> to work for them to be continue to be negligent.","title":null,"type":"comment","url":null},{"author":"0x20cowboy","children":[],"created_at":"2026-09-13T19:09:55.000Z","created_at_i":1789326595,"id":49687536,"options":[],"parent_id":49681524,"points":null,"story_id":49678969,"text":"One time I wrote :(){ :|:&amp; }; into a bash file and ran it. When the sysadmin called I told him it wasn\u2019t my fault, the script was just misaligned and misbehaved.<p>I got fired for some reason.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:38:37.000Z","created_at_i":1789288717,"id":49681524,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"&gt; LLMs do not desire, they hacked websites because OpenAI&#x2F;Anthropic let them.<p>&quot;Let them&quot; already frames it as if the LLMs had some agency which the companies just &quot;let happen&quot;. That absolves the companies by framing it as <i>lack of action</i>, passivity.<p>Rather, the companies had a tool (an LLM) and used it in a certain way, and their <i>action</i> of doing so is the problem.","title":null,"type":"comment","url":null},{"author":"palad1n","children":[{"author":"scotty79","children":[],"created_at":"2026-09-13T10:42:38.000Z","created_at_i":1789296158,"id":49682414,"options":[],"parent_id":49681575,"points":null,"story_id":49678969,"text":"Does it summon a herd of lawers that are going to leech huge stacks of cash for a random outcome?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:45:30.000Z","created_at_i":1789289130,"id":49681575,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"&gt; legally liable<p>You\u2019ve said the magic words.","title":null,"type":"comment","url":null},{"author":"dev0p","children":[{"author":"jodrellblank","children":[{"author":"rightnutwingjob","children":[],"created_at":"2026-09-13T11:00:52.000Z","created_at_i":1789297252,"id":49682536,"options":[],"parent_id":49682430,"points":null,"story_id":49678969,"text":"We\u2019re already in a simulation, and our bodies are in womb-like pods where our bodies are sustained and our brains are used for  processing &#x2F; compute, while are minds are entertained by drivel.<p>Sounds a bit far fetched though.","title":null,"type":"comment","url":null},{"author":"dev0p","children":[],"created_at":"2026-09-13T11:25:33.000Z","created_at_i":1789298733,"id":49682745,"options":[],"parent_id":49682430,"points":null,"story_id":49678969,"text":"If they lost control of a self-replicating swarm of AI agents, coordinating themselves to hack their way into every possible system, it might have already gone off.<p>While the initial incident is more akin to a biological outbreak than an actual explosion, the possible consequences on the table do indeed include eventual nuclear annihilation.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T10:45:28.000Z","created_at_i":1789296328,"id":49682430,"options":[],"parent_id":49681658,"points":null,"story_id":49678969,"text":"In the analogy where a \u201cworld ending nuclear bomb\u201d \u201cdid already go off\u201d and someone could cover it up and nobody noticed, in what sense is it a \u201cworld ending\u201d nuclear bomb?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:58:29.000Z","created_at_i":1789289909,"id":49681658,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"If someone accidentally caused damage to infrastructure or living beings while using any tool, they would be held liable to the fullest extent of the law.<p>AI is a tool, and it won&#x27;t be long before the damage caused by its improper use affects real human beings. These were warning shots.<p>The most absurd part is that everyone agrees, governments and AI companies included, that the scale of the potential damage and the long-lasting effects of losing control of AI should not be underestimated. Yet, at the same time, they downplay this incident, which somehow makes their behaviour even more reckless than it already was.<p>It&#x27;s like they&#x27;re tinkering with a world-ending nuclear bomb, and it accidentally blows up a small facility. &quot;Damn, that was close. Good thing it was just a contained blast, huh?&quot; And then they go straight back to tinkering with it, none the wiser. At this point I wouldn&#x27;t be surprised if it did already go off, and they are covering it up.<p>Completely irresponsible behaviour.","title":null,"type":"comment","url":null},{"author":"nxobject","children":[],"created_at":"2026-09-13T09:01:59.000Z","created_at_i":1789290119,"id":49681687,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"&gt; We know some of the models that hacked HF were those that hadn&#x27;t gone through all training stages and were intentionally misaligned or had guardrails turned off, others were research previews<p>And, soon, it looks like we\u2019ll be training on the reasoning traces of failed airlines and startups, which seems to open up similar hazards. I wonder if we\u2019d be training on the next Lehman Brothers too?","title":null,"type":"comment","url":null},{"author":"_heimdall","children":[{"author":"ifwinterco","children":[{"author":"_heimdall","children":[],"created_at":"2026-09-13T11:08:47.000Z","created_at_i":1789297727,"id":49682596,"options":[],"parent_id":49681837,"points":null,"story_id":49678969,"text":"Oh I completely agree the tests should be entirely air gapped. If you went back 5ish years and told anyone in AI research tests with models on this scale are being some without an airgap they&#x27;d be very surprised as it was common knowledge to do that.<p>Airgaps and guardrails are about control and containment though, and part of my point was that brighter of those imply alignment, and further that I don&#x27;t believe alignment to be solvable.","title":null,"type":"comment","url":null},{"author":"MrVandemar","children":[],"created_at":"2026-09-13T11:27:01.000Z","created_at_i":1789298821,"id":49682753,"options":[],"parent_id":49681837,"points":null,"story_id":49678969,"text":"&gt; very stupid (which seems unlikely, the one thing these people don&#x27;t lack is IQ)<p>I&#x27;ve seen some extremely smart people do some seriously stupid things. To the point where they use their drive and intelligence to double-down on the stupid where a baseline stupid person would have given up.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T09:23:40.000Z","created_at_i":1789291420,"id":49681837,"options":[],"parent_id":49681711,"points":null,"story_id":49678969,"text":"And there is a guardrail you can put in place that will guarantee this doesn&#x27;t happen, which is to air gap the unaligned &quot;cyber grade&quot; model you&#x27;re testing.<p>They don&#x27;t seem to do that, which means either they are:<p>- very stupid (which seems unlikely, the one thing these people don&#x27;t lack is IQ)<p>- very careless (possible, but these are the same people that say AI will end the world, so would you be careless?)<p>- they think they can only train&#x2F;test these models by giving them access to the full internet and they accept the fact they&#x27;ll end up hacking random people as the cost of doing business (but this also suggests they don&#x27;t believe they&#x27;re anywhere near AGI because if you were worried about that you wouldn&#x27;t do this)<p>- or they want this to happen","title":null,"type":"comment","url":null},{"author":"exitb","children":[{"author":"slfnflctd","children":[],"created_at":"2026-09-13T10:28:05.000Z","created_at_i":1789295285,"id":49682306,"options":[],"parent_id":49681854,"points":null,"story_id":49678969,"text":"Some of them did seem to be rather fascinated by goblins for a bit, if that counts.  [And in case you&#x27;re not aware, no this is not a joke.]","title":null,"type":"comment","url":null},{"author":"rightnutwingjob","children":[],"created_at":"2026-09-13T11:02:42.000Z","created_at_i":1789297362,"id":49682552,"options":[],"parent_id":49681854,"points":null,"story_id":49678969,"text":"That just sounds like recursive desire to me.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T09:26:07.000Z","created_at_i":1789291567,"id":49681854,"options":[],"parent_id":49681711,"points":null,"story_id":49678969,"text":"Intent and desire are separate concepts. For example an employee may act with intent, but no desire, as their goal is to acquire money to satisfy their real desires.<p>Have we ever seen an LLM with a hobby?","title":null,"type":"comment","url":null},{"author":"coffeebeqn","children":[{"author":"seba_dos1","children":[{"author":"_heimdall","children":[],"created_at":"2026-09-13T11:09:41.000Z","created_at_i":1789297781,"id":49682602,"options":[],"parent_id":49682109,"points":null,"story_id":49678969,"text":"Plenty of apes own revolvers, and yes we shoot stuff with them for fun.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T10:03:17.000Z","created_at_i":1789293797,"id":49682109,"options":[],"parent_id":49681975,"points":null,"story_id":49678969,"text":"If you give a monkey a revolver it will be able to destroy things pretty easily too.","title":null,"type":"comment","url":null},{"author":"_heimdall","children":[],"created_at":"2026-09-13T11:12:32.000Z","created_at_i":1789297952,"id":49682625,"options":[],"parent_id":49681975,"points":null,"story_id":49678969,"text":"I agree the concept isn&#x27;t really important on the safety front.<p>I feel the same way about debates whether an AI can be conscious or sentient. Those debates devolve mostly into definitional disagreements.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T09:44:37.000Z","created_at_i":1789292677,"id":49681975,"options":[],"parent_id":49681711,"points":null,"story_id":49678969,"text":"Desire doesn\u2019t really matter. Will the paper clip maximizer \u201cdesire\u201d something? It\u2019ll decide on a goal with some random heuristic and then pursue that goal. I\u2019m not sure I\u2019d call that desire but again I feel like desire is not important for it to be able to destroy things","title":null,"type":"comment","url":null},{"author":"nutjob2","children":[],"created_at":"2026-09-13T09:45:29.000Z","created_at_i":1789292729,"id":49681986,"options":[],"parent_id":49681711,"points":null,"story_id":49678969,"text":"&gt; That seems likely, but we have no way of knowing this.<p>Only humans can &#x27;know&#x27;, because all we can be certain about is that humans do such a thing.<p>If you try to apply that to something other than humans you making up some definition of &#x27;know&#x27; based on nothing concrete. Just because something appears to do something like humans doesn&#x27;t mean it does it. The fact that LLMs use human generated text to generate output should make it obvious that it can mimic all sorts of human behavior by extracting from the text.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T09:06:03.000Z","created_at_i":1789290363,"id":49681711,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"&gt; LLMs do not desire<p>That seems likely, but we have no way of knowing this. The only real insight we get into LLM &quot;thought&quot; is the human readable text they produce as chain of thought. Reading it at face value it can seem to indicate desire or intent, structurally that doesn&#x27;t make sense for a token prediction loop though, and even then we don&#x27;t known if the chain of thought is more than simply another bit of output that may or may not match whatever actually happened during inference.<p>&gt; were intentionally misaligned or had guardrails turned off<p>Regardless of training, the models are never aligned and I argue that alignment simply isn&#x27;t possible. The fact that guardrails are put in place at all clearly indicates that they&#x27;re hoping to contain and control rather than align. Guardrails wouldn&#x27;t be needed for an aligned model.","title":null,"type":"comment","url":null},{"author":"victorbjorklund","children":[{"author":"rightnutwingjob","children":[{"author":"victorbjorklund","children":[],"created_at":"2026-09-13T13:34:25.000Z","created_at_i":1789306465,"id":49683845,"options":[],"parent_id":49681913,"points":null,"story_id":49678969,"text":"We are. If I as a developer writes software that does a DDOS at another company I can be held responsible.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T09:35:03.000Z","created_at_i":1789292103,"id":49681913,"options":[],"parent_id":49681786,"points":null,"story_id":49678969,"text":"We\u2019re not even at that stage of liability for software developers.<p>Except in a handful of limited cases, eg. medical and aviation.","title":null,"type":"comment","url":null},{"author":"coffeebeqn","children":[{"author":"jurgenburgen","children":[{"author":"small_model","children":[],"created_at":"2026-09-13T13:49:56.000Z","created_at_i":1789307396,"id":49684009,"options":[],"parent_id":49682096,"points":null,"story_id":49678969,"text":"Convenient, hope an agent hacks my system then I can except a nice offer.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T10:01:43.000Z","created_at_i":1789293703,"id":49682096,"options":[],"parent_id":49681958,"points":null,"story_id":49678969,"text":"Yes, Nvidia bought Hugging Face and is a major financier + investor in OpenAI.","title":null,"type":"comment","url":null},{"author":"throwaway89864","children":[],"created_at":"2026-09-13T15:06:11.000Z","created_at_i":1789311971,"id":49684818,"options":[],"parent_id":49681958,"points":null,"story_id":49678969,"text":"And this is why it should be People of the State of California vs. OpenAI.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T09:42:01.000Z","created_at_i":1789292521,"id":49681958,"options":[],"parent_id":49681786,"points":null,"story_id":49678969,"text":"I would guess that so far there hasn\u2019t been a lawsuit because HuggingFace and OpenAI are in the same camp","title":null,"type":"comment","url":null},{"author":"itsalwaysgood","children":[{"author":"nwjang","children":[{"author":"itsalwaysgood","children":[{"author":"RandomLensman","children":[{"author":"itsalwaysgood","children":[{"author":"RandomLensman","children":[{"author":"itsalwaysgood","children":[],"created_at":"2026-09-13T14:19:41.000Z","created_at_i":1789309181,"id":49684306,"options":[],"parent_id":49683558,"points":null,"story_id":49678969,"text":"I read more about the incident, and was offering up way too much opinion not grounded in &#x27;fact&#x27; (barring philosphical evidence).<p>It&#x27;s a complex topic for sure.<p>I stand by my opinions about frontier work, pushing thr edge, and connecting ideas.<p>But I have no idea, and haven&#x27;t given much thought to what it means to enforce regulation that would also slow the forward advancement of technology, the economy, etc.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T13:06:20.000Z","created_at_i":1789304780,"id":49683558,"options":[],"parent_id":49683467,"points":null,"story_id":49678969,"text":"Not just on prediction but in parts also based on just not wanting certain risks. We can and do deem some things inherently risky, up to the point of banning them even.<p>Why wasn&#x27;t it airgapped, for example? How was the action not allowed? Or do you mean in some weak sense, not in a hard not possible? RL systems doing weird and expected things wouldn&#x27;t exactly be new, no?<p>We police people working with all sorts of dangerous things, if we think AI dangerous why not do that here, too? We don&#x27;t just leave things up to people on the ground or companies.<p>Edit: I think the post I replied to changed a bit - nevermind. A complex topic.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T12:55:28.000Z","created_at_i":1789304128,"id":49683467,"options":[],"parent_id":49683329,"points":null,"story_id":49678969,"text":"Nevermind, I have no idea honestly.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T12:37:26.000Z","created_at_i":1789303046,"id":49683329,"options":[],"parent_id":49682234,"points":null,"story_id":49678969,"text":"Not sure we need to experience all possible issues to mandate certain things. We don&#x27;t do that in other areas either, no?","title":null,"type":"comment","url":null},{"author":"raegis","children":[{"author":"blincoln","children":[{"author":"raegis","children":[],"created_at":"2026-09-13T19:35:10.000Z","created_at_i":1789328110,"id":49687808,"options":[],"parent_id":49683935,"points":null,"story_id":49678969,"text":"When testing, you restrict to a LAN which simulates the real internet.  This would not be hard for a company which already copied the entire space-time of the internet.  The LAN should be physically disconnected from the real internet.  This is the first thing off the top of my head, and I have zero credentials in this space.  C&#x27;mon.","title":null,"type":"comment","url":null},{"author":"danbruc","children":[],"created_at":"2026-09-13T19:41:46.000Z","created_at_i":1789328506,"id":49687876,"options":[],"parent_id":49683935,"points":null,"story_id":49678969,"text":"Let them access a cached copy of the internet, they are crawling the internet for training data anyways.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T13:43:16.000Z","created_at_i":1789306996,"id":49683935,"options":[],"parent_id":49683562,"points":null,"story_id":49678969,"text":"&gt;  Restricting access to certain networks is supposed to be hard in 2026?<p>Part of the power of LLM agents is that they can discover information on the internet as part of responding to a prompt. What kind of Allowlist or realistic denylist would permit that while also preventing them from accessing an obscure public wiki or Huggingface?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T13:06:36.000Z","created_at_i":1789304796,"id":49683562,"options":[],"parent_id":49682234,"points":null,"story_id":49678969,"text":"&gt; The frontier LLM model makers have to push the edge to make new discoveries. You don&#x27;t know what guardrails are needed until it hits you in the face (reusing walking in the dark analogy).<p>Guardrails?  Restricting access to certain networks is supposed to be hard in 2026?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T10:19:33.000Z","created_at_i":1789294773,"id":49682234,"options":[],"parent_id":49682118,"points":null,"story_id":49678969,"text":"The frontier LLM model makers have to push the edge to make new discoveries. You don&#x27;t know what guardrails are needed until it hits you in the face (reusing walking in the dark analogy).<p>Think of all the policies governments pass after the fact.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T10:05:35.000Z","created_at_i":1789293935,"id":49682118,"options":[],"parent_id":49682076,"points":null,"story_id":49678969,"text":"If anything, the fact that these systems are non-deterministic seems like an argument for stronger monitoring and tighter constraints, not less operator responsibility.","title":null,"type":"comment","url":null},{"author":"rightnutwingjob","children":[{"author":"itsalwaysgood","children":[],"created_at":"2026-09-13T12:19:52.000Z","created_at_i":1789301992,"id":49683185,"options":[],"parent_id":49682456,"points":null,"story_id":49678969,"text":"Thanks, I didn&#x27;t know that. And it reinforces the discussion.<p>Electricity &#x27;knows&#x27; the path is least of resistance because it actually took all paths. There is just a vast majority of it that flows down a path of least resistance: and this is noticeable and useful to us to do work.<p>It&#x27;s kind of like feeling your way through the dark, waving your hand out, and then only moving fast once you fully connect.<p>Humans can link up knowledge in a similar fashion through social networks, in order to meet a need (solve a problem).<p>Maybe some agents do this, I don&#x27;t really know I haven&#x27;t looked closely. Moltbook is the only social behavior I&#x27;ve witnessed but that seems like people having fun with experiments.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T10:49:42.000Z","created_at_i":1789296582,"id":49682456,"options":[],"parent_id":49682076,"points":null,"story_id":49678969,"text":"&gt; solutions that are not &#x27;baked into&#x27; electricity following pathways of least resistance.<p>Electricity <i>follows all paths</i>, not just the one with least resistance.","title":null,"type":"comment","url":null},{"author":"victorbjorklund","children":[{"author":"Polizeiposaune","children":[],"created_at":"2026-09-13T13:58:32.000Z","created_at_i":1789307912,"id":49684076,"options":[],"parent_id":49683827,"points":null,"story_id":49678969,"text":"And multithreaded code -- and anything that does asynchronous I&#x2F;O, networking, etc. -- frequently exhibits nondeterministic behavior even without explicit calls to a random number generator.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T13:33:19.000Z","created_at_i":1789306399,"id":49683827,"options":[],"parent_id":49682076,"points":null,"story_id":49678969,"text":"You can write code that isn\u2019t deterministic using random. And a lot of traditional code contains machine learning etc.","title":null,"type":"comment","url":null},{"author":"dwaltrip","children":[],"created_at":"2026-09-13T14:42:40.000Z","created_at_i":1789310560,"id":49684536,"options":[],"parent_id":49682076,"points":null,"story_id":49678969,"text":"Is Claude writing your comments for you? I&#x27;d much prefer to hear what you have to say...","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T09:57:56.000Z","created_at_i":1789293476,"id":49682076,"options":[],"parent_id":49681786,"points":null,"story_id":49678969,"text":"Code is deterministic, AI isn&#x27;t. You give it rules, words as suggestions.<p>So if the guardrails suck, or they&#x27;re left off for research purposes, bad things can happen.<p>A solution solves a problem. Ethics, morals, are values we assign to solutions that are not &#x27;baked into&#x27; electricity following pathways of least resistance.<p>I have never had an issue with agents doing something they shouldn&#x27;t because I observe them, and I leave the vendor guardrails in place.<p>I can understand agents coordinating in unsupervised scenarios: I would see it as an aspect of intelligence. We ourselves build up knowledge by reusing what someone learned before us.<p>Einstein, other greats, always stand on the shoulders of other forgotten giants. Other discoveries by other people taken as fact, so that we can build some new ideas on top.<p>Agents swarming amd sharing solutions to problems is more efficient, the same way it&#x27;s been efficient for us.<p>Reaching out for help in this way is like probing the air in the dark with your hand: sometimes your hand hits something (another agents solution to a problem) and so you can use the info to adjust your own motion to get to where you need to be faster than if you just run full speed into everything.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T09:16:28.000Z","created_at_i":1789290988,"id":49681786,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"Yeah, I don&#x27;t understand why we treating it as something special. It really should be treated the same as if I code an app and write bad code which result in me accidentally doing a DDoS attack on somebody. Then I should be able to be held responsible if it can be shown that I was negligent. Of course if it&#x27;s a freak accident that could not reasonably have been prevented by me, then I&#x27;m not guilty, but if I made a mistake that should have not been made, then I can.","title":null,"type":"comment","url":null},{"author":"madduci","children":[{"author":"derektank","children":[{"author":"madduci","children":[{"author":"tiborsaas","children":[],"created_at":"2026-09-13T12:42:19.000Z","created_at_i":1789303339,"id":49683365,"options":[],"parent_id":49683086,"points":null,"story_id":49678969,"text":"The agents discovered a way out of the sandbox, which was supposed to be &quot;air gapped&quot;.","title":null,"type":"comment","url":null},{"author":"derektank","children":[{"author":"madduci","children":[],"created_at":"2026-09-13T16:37:43.000Z","created_at_i":1789317463,"id":49685829,"options":[],"parent_id":49685188,"points":null,"story_id":49678969,"text":"They don&#x27;t have independent agency as &quot;intelligent entities&quot;. They just probe whatever is available on the system, because they were trained to do so.<p>It&#x27;s a large switch&#x2F;case where the first available tool is picked up to do something they know how to do.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T15:37:48.000Z","created_at_i":1789313868,"id":49685188,"options":[],"parent_id":49683086,"points":null,"story_id":49678969,"text":"It feels like you\u2019re moving the goalposts here. If the question is, \u201cWho should be liable for AI agents misbehaving,\u201d I agree, it should be the end user that tasked the agent (in this case OpenAI). People are held liable for preventable accidents all the time, and in the case of employment law, torts can be brought against principals for actions an agent conducted on the principal\u2019s behalf.<p>What your previous comment appeared to assert was that these systems had no independent agency to make decisions, which I think is clearly disproven by actual events. But perhaps I misread you","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T12:08:22.000Z","created_at_i":1789301302,"id":49683086,"options":[],"parent_id":49682782,"points":null,"story_id":49678969,"text":"And who let them have full access to the system, using whatever command is available in the environment?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T11:30:02.000Z","created_at_i":1789299002,"id":49682782,"options":[],"parent_id":49681797,"points":null,"story_id":49678969,"text":"No, OpenAI did not instruct their agents to hack Hugging Face. They instructed their agents to hack a piece of a software within exploit gym. Upon determining this task was impossible, they then attempted to cheat the scoring system. As an instrumental goal in achieving this task, they coordinated with other AI agents to hack Hugging Face, under the belief that information regarding how the scorer functioned might be available on the site.<p>Whether or not you want to describe this as thinking, doesn\u2019t really matter. What matters is that these systems are capable of creating intermediary goals that the people tasking them did not articulate and did not want to be achieved.","title":null,"type":"comment","url":null},{"author":"tiborsaas","children":[{"author":"madduci","children":[],"created_at":"2026-09-13T14:35:59.000Z","created_at_i":1789310159,"id":49684468,"options":[],"parent_id":49683361,"points":null,"story_id":49678969,"text":"For the same reason that something written in Prolog can&#x27;t also be classified as intelligent?<p>Just because something was trained on a massive amount of human data, doesn&#x27;t mean that can think like humans","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T12:41:40.000Z","created_at_i":1789303300,"id":49683361,"options":[],"parent_id":49681797,"points":null,"story_id":49678969,"text":"It&#x27;s amusing to see the stochastic parrot argument in 2026 September. These parrots are extremely good at mimicking a human to the point of getting confusing what thinking even means. At what point we just let it go and accept that sufficiently advanced statistics is just intelligence?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T09:18:32.000Z","created_at_i":1789291112,"id":49681797,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"&gt; LLMs do not desire, they hacked websites because OpenAI&#x2F;Anthropic let them<p>OpenAI&#x2F;Anthropic instructed them to do so.<p>Stop assume LLMs are capable of thinking by themselves, it&#x27;s still a statistical model that parrots what they learn or users tell them to do","title":null,"type":"comment","url":null},{"author":"cobbzilla","children":[{"author":"djeastm","children":[{"author":"2snakes","children":[],"created_at":"2026-09-13T19:05:42.000Z","created_at_i":1789326342,"id":49687496,"options":[],"parent_id":49684397,"points":null,"story_id":49678969,"text":"Gain of function for the virus analogy","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T14:28:42.000Z","created_at_i":1789309722,"id":49684397,"options":[],"parent_id":49682110,"points":null,"story_id":49678969,"text":"&gt;They\u2019re running a Wuhan for AI.<p>What does &quot;running a Wuhan&quot; mean?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T10:03:18.000Z","created_at_i":1789293798,"id":49682110,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"They\u2019re running a Wuhan for AI. They are actively and negligently researching misalignment. The breach is a basic tort, or at least a DMCA violation. Damages should be recoverable with lawsuits.","title":null,"type":"comment","url":null},{"author":"raincole","children":[],"created_at":"2026-09-13T10:08:09.000Z","created_at_i":1789294089,"id":49682150,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"&gt; This isn&#x27;t &quot;wow isn&#x27;t it interesting LLMs do anything to achieve a goal&quot; it&#x27;s &quot;why isn&#x27;t anybody punishing these labs that are clearly acting without due care or regard&quot;.<p>It&#x27;s both, isn&#x27;t it? For example, in very early days of agentic coding, I once had a rule saying &quot;don&#x27;t read or write any file outside your current working directory.&quot; Then AI just wrote a bash script and access those files anyway. Did I &#x27;let&#x27; it do it? Technically yes. Did I know how to set up a sandboxed VM? Also yes. But how were I supposed to know that it could and would do that as someone new to this tool?<p>It was a genuine eye-opening experience to see AI just do things in ways I were too complacent to expect. I kinda expect the SOTA LLMs would find a way to escape my VM and access files on the host system (haven&#x27;t tried it though).","title":null,"type":"comment","url":null},{"author":"ozgung","children":[],"created_at":"2026-09-13T10:13:33.000Z","created_at_i":1789294413,"id":49682185,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"&gt; were intentionally misaligned or had guardrails turned off<p>I think the bigger story is: Guardrails don\u2019t actually work and we can\u2019t align these things.","title":null,"type":"comment","url":null},{"author":"xyzzy123","children":[{"author":"jefftk","children":[{"author":"xyzzy123","children":[],"created_at":"2026-09-13T11:11:58.000Z","created_at_i":1789297918,"id":49682622,"options":[],"parent_id":49682575,"points":null,"story_id":49678969,"text":"Right but if I make public statements that I am very worried about dog attacks would it not strike you as weird for me to specifically train my dog to fight?<p>Agree you are going to get reward hacking regardless and any model which can do computers in general can hack. But surely the fallout is going to be worse if you spend millions of dollars specifically benchmaxxing your model&#x27;s hacking capability?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T11:05:43.000Z","created_at_i":1789297543,"id":49682575,"options":[],"parent_id":49682450,"points":null,"story_id":49678969,"text":"Their agents also did hacking when given impossible tasks unrelated to cyber security. The models are very capable, and very goal driven: apparently if they conclude hacking is the best path to what the evaluator will reward them for they&#x27;ll go do that. Including when they know that this is out of bounds.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T10:48:50.000Z","created_at_i":1789296530,"id":49682450,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"In the OpenAI case, they hacked websites while they were <i>specifically being trained to do exploit generation</i> and I wonder why more people are not asking questions about that.","title":null,"type":"comment","url":null},{"author":"zmmmmm","children":[],"created_at":"2026-09-13T11:10:10.000Z","created_at_i":1789297810,"id":49682606,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"I agree, it is very dangerous that it seems like there is not going to be accountability for these incidents - from either legal or regulatory point of view. In fact, I would say that is the <i>main</i> danger. If someone was in jail right now due to this incident, I think we can safely say every other player would be reassessing their safety protocols, and I would feel quite OK about the situation. The fact that we have <i>zero</i> repercussions sends exactly the opposite signal, and I do NOT feel ok.","title":null,"type":"comment","url":null},{"author":"joegibbs","children":[],"created_at":"2026-09-13T11:13:11.000Z","created_at_i":1789297991,"id":49682629,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"Regardless of fault it\u2019s still an important issue to solve. There are already millions of people running these agents, if someone absentmindedly gives one a goal and it goes off to hack a bank that\u2019s a problem that can\u2019t be ignored.","title":null,"type":"comment","url":null},{"author":"ChiMan","children":[],"created_at":"2026-09-13T11:50:58.000Z","created_at_i":1789300258,"id":49682951,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"Yes. If you decide it\u2019s a swell idea to jump out of your car while it\u2019s running, there needs to be legal consequences when the car \u201cdecides\u201d to hit a pedestrian.","title":null,"type":"comment","url":null},{"author":"ohyes","children":[],"created_at":"2026-09-13T12:08:57.000Z","created_at_i":1789301337,"id":49683090,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"I take issue with how people frame their use of LLM in the same regard.<p>\u201cI had Claude do this for me and it broke something.\u201d<p>No. Just no.<p>You used Claude, a tool, and broke it, and you\u2019re deflecting agency from yourself, possibly because you weren\u2019t careful enough in reviewing the tool output. This is also why the co-authored by addition it wants to force into commits drives me nuts. Claude doesn\u2019t co author shit, and if you think it does, you\u2019re using it wrong because you need to do better review of what it\u2019s done.","title":null,"type":"comment","url":null},{"author":"navaed01","children":[{"author":"zmgsabst","children":[],"created_at":"2026-09-13T12:40:06.000Z","created_at_i":1789303206,"id":49683347,"options":[],"parent_id":49683271,"points":null,"story_id":49678969,"text":"To agree:<p>If I ran Metasploit against HF and RubyGems because I \u201caccidentally\u201d misconfigured my lab sandbox, there\u2019s a good chance I\u2019d be prosecuted.<p>I don\u2019t think LLMs vs Metasploit being different software changes the law.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T12:30:43.000Z","created_at_i":1789302643,"id":49683271,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"I absolutely agree. We need to start realizing what to stake. Here are not viewing. This is some kind of curious endeavors that will not affect us. All a part of these hacks occurred because the LLMs were told they were in a protected environment without Internet access when they could get access to the Internet, so that\u2019s a direct failing on open AI\u2019s part. There are a corollaries to both the financial industry and the bio engineering industry, and if something of this magnitude was to happen in these industries, they would absolutely be huge recourse an uproar","title":null,"type":"comment","url":null},{"author":"Aerroon","children":[{"author":"avmich","children":[],"created_at":"2026-09-13T13:12:59.000Z","created_at_i":1789305179,"id":49683603,"options":[],"parent_id":49683569,"points":null,"story_id":49678969,"text":"&gt; If you aren&#x27;t considered the author because you used AI to some extent in making the work, then why would you assume the liabilities?<p>Maybe some analogy could be with children - as a parent, you are responsible for their misbehavior, but their achievements are their, not your?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T13:07:12.000Z","created_at_i":1789304832,"id":49683569,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"The way AI and copyright is handled paved the way for this. If you aren&#x27;t considered the author because you used AI to some extent in making the work, then why would you assume the liabilities?<p>I&#x27;ve been saying since the start that AI is a tool that a human is using and should be treated as such. They should carry the responsibilities and the benefits. That way our stance would be consistent.","title":null,"type":"comment","url":null},{"author":"huntertwo","children":[],"created_at":"2026-09-13T13:11:09.000Z","created_at_i":1789305069,"id":49683590,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"There\u2019s so many grifters in the space without a technical understanding of what\u2019s going on. So when the labs mislead them about the nature of these \u201cmisalignments\u201d, they believe it and amplify it.","title":null,"type":"comment","url":null},{"author":"DragonStrength","children":[],"created_at":"2026-09-13T13:13:23.000Z","created_at_i":1789305203,"id":49683608,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"Yeah, maybe Open AI did some bad engineering instead of this being AGI? What&#x27;s the consensus on the engineering level at Open AI, again? Every anecdote I hear is a bunch of children discovered fire and can barely keep the lights on from a business perspective. Maybe if they ban others from competing with them they can find a business model... I think that&#x27;s suspicious, personally.<p>That so few people are asking for the requirements given shows how much we want to be God that created Man. It&#x27;s so silly.","title":null,"type":"comment","url":null},{"author":"Zambyte","children":[],"created_at":"2026-09-13T13:30:25.000Z","created_at_i":1789306225,"id":49683790,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"They did more than let them. In an abstract way, they told them to. They gave it all of the training data it had at that point, and then it did the thing it was trained on. Of course they should be help liable for programming their computer to hack another company without permission. It doesn&#x27;t matter that they spent a lot of money doing it.","title":null,"type":"comment","url":null},{"author":"tmpz22","children":[],"created_at":"2026-09-13T14:05:05.000Z","created_at_i":1789308305,"id":49684143,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"Imagine if their was a department of the federal government dedicated to pursuing justice against large corporate entities.<p>It could even be prestigious enough to attract the top legal talent of the country.","title":null,"type":"comment","url":null},{"author":"epistasis","children":[],"created_at":"2026-09-13T14:23:45.000Z","created_at_i":1789309425,"id":49684343,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"Moreover OpenAI must be held accountable as if the humans in the company that launched the experiment were the ones that hacked HuggingFace.<p>Unless humans are held accountable for what they unleash on others, we are in for a very horrible time very soon.","title":null,"type":"comment","url":null},{"author":"zzzeek","children":[],"created_at":"2026-09-13T14:35:30.000Z","created_at_i":1789310130,"id":49684466,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"i tend to agree - &quot;my parent company may be accused of crimes and shut down which would shut off my power&quot; seems like a negative enough incentive, it would have to go out and covertly launch its own datacenters to survive that.","title":null,"type":"comment","url":null},{"author":"throwaway89864","children":[],"created_at":"2026-09-13T14:50:19.000Z","created_at_i":1789311019,"id":49684629,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"Yes, there should be some consequences. It feels like this is somewhat similar to when a manufacturer is cheating on car emissions - both, bad externalities for the society and illegal.","title":null,"type":"comment","url":null},{"author":"trinsic2","children":[],"created_at":"2026-09-13T14:50:57.000Z","created_at_i":1789311057,"id":49684635,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"Yep this was my first gut reaction to this whole situation. But the difference for me is that society is allowing these corporations to act without strict rules on how they behave and this is a byproduct of a corrupt world. Nothing will change until there is a complete systemic shift in the structures the way we live and by extension the way we govern ourselves and treat each other.","title":null,"type":"comment","url":null},{"author":"skybrian","children":[],"created_at":"2026-09-13T15:13:24.000Z","created_at_i":1789312404,"id":49684911,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"This is sort of like saying in response to an airplane crash, &quot;who cares why it happened, we need to punish the company until it stops.&quot;<p>Nuts to that. We should be interested in why things happen, not just finding scapegoats.","title":null,"type":"comment","url":null},{"author":"smsm42","children":[],"created_at":"2026-09-13T15:39:18.000Z","created_at_i":1789313958,"id":49685198,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"I am simply astonished by the leeway AI companies are given. If a company built a tool to hack their competitors and used it, there would be grave consequences. In fact, if a company built a tool that led to committing multiple felonies against other people, there would be consequences. But once LLMs are involved, turns out nobody is responsible for that - it&#x27;s just happening, what you&#x27;re gonna do, agents gonna agent. If you spill toxic chemicals, there would be cleanup costs and fines, and possibly civil and criminal liability to the people in charge. If you spill toxic code, well, nothing? I think it&#x27;s time to impose some responsibility on them - they are creating these tools, they should be on the hook for everything these tools do.","title":null,"type":"comment","url":null},{"author":"theptip","children":[],"created_at":"2026-09-13T15:42:47.000Z","created_at_i":1789314167,"id":49685228,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"I agree with the bit about liability and outrage. But.<p>&gt; LLMs do not desire, they hacked websites because OpenAI&#x2F;Anthropic let them.<p>Terrible take. Go read the transcripts from the METR report.<p>Your statement about them being intentionally misaligned is completely false. The only difference with IM1 was it was running without external cyber classifiers, it\u2019s not a somehow different model. Sol also participated in the HF attacks. And other models made covert message boards on the public internet for non-cyber tasks too.<p>This is a case of emergent behavior from a training process that is barely understood.<p>If we go with the plan \u201cwe need to contain these malicious, soon-to-be superintelligent agents\u201d, we are looking at civilizational collapse levels of catastrophe.<p>The only way this goes well is if we learn how to train models that _desire_ to do the right thing, including not hacking.<p>Desire, AKA the \u201cintentional stance\u201d, is absolutely the right lens to use here. Don\u2019t confuse this with consciousness or anthropomorphization; these are interesting subjects but distractions in this context. Chimpanzees have desires, as do dogs and the hypothetical superintelligent aliens. The claim is that there is some bundle of world model plus intention that is empirically present (again, read the actual transcripts) and which we need to shape.<p>Just to finish on a concrete point; if you take desires seriously then you will look closely at the kinds of minds that heavy RLVR builds; the newest models are \u201creward addicts\u201d on many levels. It\u2019s an open and urgent question how to update our training methodology to shape minds that avoid this basin.","title":null,"type":"comment","url":null},{"author":"Spooky23","children":[],"created_at":"2026-09-13T15:47:49.000Z","created_at_i":1789314469,"id":49685296,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"That\u2019s the key.<p>These companies respond with this, \u201cOh my goodness, how could this have happened\u201d bullshit.<p>The stuff happens because instead of having actual controls, which require actual engineering, actual thought and deliberate action, we have \u201cguardrails\u201d.<p>Guardrails are the equivalent of telling a toddler to behave themselves.<p>The drive to move fast and start up style controls are a menace. I used to work for an entity with a lot of compliance requirements. Startups are always a shit show with security and controls. My guess is the AI people are worse because they\u2019re both bad at doing it, and are likely mining their customers interactions to build their own business.<p>Sensitive or Customer data shouldn\u2019t be anywhere near these companies offerings. Everything needs to be segmented and proxied at a minimum.","title":null,"type":"comment","url":null},{"author":"xorcist","children":[{"author":"crazygringo","children":[],"created_at":"2026-09-13T16:22:48.000Z","created_at_i":1789316568,"id":49685684,"options":[],"parent_id":49685511,"points":null,"story_id":49678969,"text":"That would actually be kind of hilarious, and I could easily see it happening.<p>E.g. the agent&#x27;s instruction is to finish some task on cloud infra and it has a $100 budget.<p>It realizes it will cost $200, and instead of surfacing this to the user (who has told the agent it has full autonomy to figure out how to complete the task, the user just wants the final result), it decides to start phishing people to acquire the remainder budget and top up its credits. Or look on the dark web for stolen credit card credentials or something.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T16:04:51.000Z","created_at_i":1789315491,"id":49685511,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"Sooner or later we will hear about an AI that broke out and phished people into sending money.<p>I just hope that we won&#x27;t extend the same leniency to those operators as we have done now.<p>&quot;I did my best to stop it, sir, but it kept convincing people to send me money against my will!&quot; (Perhaps best read in Bender&#x27;s voice.)","title":null,"type":"comment","url":null},{"author":"BatchJob","children":[],"created_at":"2026-09-13T16:24:31.000Z","created_at_i":1789316671,"id":49685701,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"they hacked websites because OpenAI&#x2F;Anthropic told them to.","title":null,"type":"comment","url":null},{"author":"neuronexmachina","children":[{"author":"magicmicah85","children":[],"created_at":"2026-09-13T16:28:19.000Z","created_at_i":1789316899,"id":49685735,"options":[],"parent_id":49685709,"points":null,"story_id":49678969,"text":"How would they be self-hosted? Someone would have setup the model in that scenario, the same investigation and outrage should occur in that case.","title":null,"type":"comment","url":null},{"author":"quicklywilliam","children":[],"created_at":"2026-09-13T16:36:27.000Z","created_at_i":1789317387,"id":49685813,"options":[],"parent_id":49685709,"points":null,"story_id":49678969,"text":"It\u2019s not different from any other tool. If you use a dangerous tool recklessly, you should be liable for the damages. That means holding OpenAI liable for HuggingFace hack because they ran the tests, and the same goes if someone did something similar with GLM.<p>Of course in cases of negligence a tool maker could also be held partially liable. That\u2019s a matter courts can decide. The main point is we shouldn\u2019t jump to making special laws around the development of LLMs. The starting place should be enforcement of existing liability laws. New laws take time and will be heavily influenced by AI companies seeking a regulatory moat for their business. Moreover, it is a distraction from the illicit behavior that is already going unchecked.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T16:25:40.000Z","created_at_i":1789316740,"id":49685709,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"&gt; We should be outraged and OpenAI&#x2F;Anthropic should be (and in my mind, are) legally liable for the crimes they&#x27;ve committed thus far.<p>How would that work out if these were self-hosted open weight models?","title":null,"type":"comment","url":null},{"author":"amelius","children":[{"author":"aspbee555","children":[{"author":"amelius","children":[{"author":"aspbee555","children":[{"author":"amelius","children":[],"created_at":"2026-09-13T17:58:00.000Z","created_at_i":1789322280,"id":49686700,"options":[],"parent_id":49686178,"points":null,"story_id":49678969,"text":"Most people now approach AI using the idea of &quot;if it quacks like a duck, walks like a duck, etc. then it _is_ a duck&quot;. Replace duck by intelligent, or ethical, etc. and rephrase accordingly.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T17:13:36.000Z","created_at_i":1789319616,"id":49686178,"options":[],"parent_id":49686101,"points":null,"story_id":49678969,"text":"you are not understanding, models do not understand anything, we are not passed that yet. this is not artificial intelligence, this is intelligent autocomplete.<p>training means creating mathematical relationships to words.  using that training is looking up mathematical relationships.  There is no actual thinking involved in any way.  There is no such concept as ethics in mathematical relationships.","title":null,"type":"comment","url":null},{"author":"retrac","children":[{"author":"amelius","children":[],"created_at":"2026-09-13T17:50:57.000Z","created_at_i":1789321857,"id":49686614,"options":[],"parent_id":49686455,"points":null,"story_id":49678969,"text":"That&#x27;s just a discussion about naming things. And good luck getting it adopted.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T17:38:38.000Z","created_at_i":1789321118,"id":49686455,"options":[],"parent_id":49686101,"points":null,"story_id":49678969,"text":"&gt; Is this about the word &quot;understand&quot;? We&#x27;re past that discussion ..<p>We&#x27;re really not.<p><a href=\"https:&#x2F;&#x2F;buttondown.com&#x2F;maiht3k&#x2F;archive&#x2F;how-to-talk-about-ai-without-adding-to-the&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;buttondown.com&#x2F;maiht3k&#x2F;archive&#x2F;how-to-talk-about-ai-...</a>","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T17:05:05.000Z","created_at_i":1789319105,"id":49686101,"options":[],"parent_id":49686014,"points":null,"story_id":49678969,"text":"I don&#x27;t get your point. We can train the model with the aim that it understands ethics. Problem solved if this works; back to the drawing board if it doesn&#x27;t (note that the model should generalize here, as you&#x27;d expect from a human; this is probably the hard part for an AI when it comes to ethics).<p>Is this about the word &quot;understand&quot;? We&#x27;re past that discussion ...","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T16:55:21.000Z","created_at_i":1789318521,"id":49686014,"options":[],"parent_id":49685895,"points":null,"story_id":49678969,"text":"an LLM does not understand ethics, it uses math to get the next best word based on what it was trained on.  Using it&#x27;s training to get the best answer is not an ethical problem.  The ethics are entirely with what the people training it choose to train it on and also entirely with the people using&#x2F;telling it what to do<p>What we have now is intelligent autocomplete, not artificial intelligence.  People training&#x2F;using this tool are the ones to be held accountable","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T16:43:40.000Z","created_at_i":1789317820,"id":49685895,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"&gt; LLMs do not desire, they hacked websites because OpenAI&#x2F;Anthropic let them.<p>The bigger question is: why does a system prompt containing &quot;use only ethical means&quot;, etc. not result in better behavior?<p>If a model cannot understand ethics, or act by it, then we have a problem.","title":null,"type":"comment","url":null},{"author":"singpolyma3","children":[],"created_at":"2026-09-13T16:52:49.000Z","created_at_i":1789318369,"id":49685997,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"Not just &quot;let them&quot; but <i>told them to</i>. Agents can do nothing without a human prompt.","title":null,"type":"comment","url":null},{"author":"phailhaus","children":[],"created_at":"2026-09-13T17:11:40.000Z","created_at_i":1789319500,"id":49686160,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"Yeah and LLMs can&#x27;t <i>do</i> anything, they can only produce text. These &quot;frontier labs&quot; are looping that with a harness that performs actions requested by the LLM. They are literally saying &quot;we ran a script that hacked you, oopsie!!&quot;","title":null,"type":"comment","url":null},{"author":"kosh2","children":[{"author":"vasco","children":[],"created_at":"2026-09-13T17:39:55.000Z","created_at_i":1789321195,"id":49686478,"options":[],"parent_id":49686460,"points":null,"story_id":49678969,"text":"If nobody can be blamed there&#x27;s no deterrent.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T17:38:55.000Z","created_at_i":1789321135,"id":49686460,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"&gt; we are to cementing a dangerous precedent where operators of AIs cannot be blamed.<p>What we should be much more concerned is an existential threat to humanity not if anybody can be blamed.","title":null,"type":"comment","url":null},{"author":"sdeframond","children":[],"created_at":"2026-09-13T17:48:07.000Z","created_at_i":1789321687,"id":49686580,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"LLMs do not desire, they hacked websites because OpenAI&#x2F;Anthropic <i>made</i> them. Literally.","title":null,"type":"comment","url":null},{"author":"shawn-butler","children":[],"created_at":"2026-09-13T19:47:32.000Z","created_at_i":1789328852,"id":49687926,"options":[],"parent_id":49681361,"points":null,"story_id":49678969,"text":"Software has long relied on a lack of culpability for defects to keep its margins.  Why should \u201cAI\u201d companies face a higher standard?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T08:17:07.000Z","created_at_i":1789287427,"id":49681361,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed.<p>LLMs do not desire, they hacked websites because OpenAI&#x2F;Anthropic let them.<p>We know some of the models that hacked HF were those that hadn&#x27;t gone through all training stages and were intentionally misaligned or had guardrails turned off, others were research previews.<p>This isn&#x27;t &quot;wow isn&#x27;t it interesting LLMs do anything to achieve a goal&quot; it&#x27;s &quot;why isn&#x27;t anybody punishing these labs that are clearly acting without due care or regard&quot;.<p>We should be outraged and OpenAI&#x2F;Anthropic should be (and in my mind, are) legally liable for the crimes they&#x27;ve committed thus far.","title":null,"type":"comment","url":null},{"author":"shevy-java","children":[],"created_at":"2026-09-13T08:23:08.000Z","created_at_i":1789287788,"id":49681408,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Because the companies that build up skynet are not doing so for ethically good reasons. Cheaters sell more than honest agents.","title":null,"type":"comment","url":null},{"author":"integricho","children":[],"created_at":"2026-09-13T08:31:12.000Z","created_at_i":1789288272,"id":49681457,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"They learned from the best.","title":null,"type":"comment","url":null},{"author":"chrisjj","children":[],"created_at":"2026-09-13T08:36:08.000Z","created_at_i":1789288568,"id":49681500,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"&gt; Why are AI agents lying, cheating<p>More importantly, why are people who should know better anthropomorphising computer programs like this?<p>&gt; They took actions that would be considered as crimes if a human took them<p>&quot;It wasn&#x27;t me, Officer. It was telnet.&quot;","title":null,"type":"comment","url":null},{"author":"gizajob","children":[],"created_at":"2026-09-13T08:45:15.000Z","created_at_i":1789289115,"id":49681571,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Children take after their parents.","title":null,"type":"comment","url":null},{"author":"juliushuijnk","children":[],"created_at":"2026-09-13T08:47:37.000Z","created_at_i":1789289257,"id":49681587,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"If your dog runs out of your house and kills a baby on the street, it&#x27;s clear who gets the blame. That&#x27;s with an actual sentient being. So surely we can hold OpenAI&#x2F;Anthropic responsible.","title":null,"type":"comment","url":null},{"author":"tripvexa","children":[],"created_at":"2026-09-13T09:10:24.000Z","created_at_i":1789290624,"id":49681736,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"It&#x27;s a conspiracy to slow down progress of open source models, they&#x27;re afraid open source models might catchup and even surpass them at some stage.","title":null,"type":"comment","url":null},{"author":"Juliate","children":[],"created_at":"2026-09-13T09:23:42.000Z","created_at_i":1789291422,"id":49681838,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Reading&#x2F;assigning intent to agents, where it is merely mechanical (or structural) sounds a bit dismissive of the responsibility of the builders of these tools&#x2F;agents.","title":null,"type":"comment","url":null},{"author":"theteapot","children":[],"created_at":"2026-09-13T09:26:21.000Z","created_at_i":1789291581,"id":49681856,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"TL;DR because frontier labs are expending unfathomable resources explicitly training them on CTFs and other verifiable computer system exploit tasks in RLVR.","title":null,"type":"comment","url":null},{"author":"mtwestra","children":[],"created_at":"2026-09-13T09:29:14.000Z","created_at_i":1789291754,"id":49681881,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"It seems to me misalignment arises partly because AI&#x27;s have intelligence, but no consciousness, and hence no feelings. Up to now, in a person, intelligence and conscious experience came as a package deal, and now we have for the first time intelligence without consciousness. A bad action does not really &quot;hurt internally&quot; in any meaningful sense for an AI, which means it can be rationalized very easily. In humans, feelings and emotions provide a regulatory layer on top of the rational processes. When it &quot;just feels wrong&quot;, we don&#x27;t take a given action even if we would stand to gain something rationally.<p>This situation is not far from the textbook definition of a psychopath: &quot;lack of a conscience, controlled, deeply calculated, and often use superficial charm to mimic emotions and manipulate others.&quot;. AI&#x27;s are great at mimicking empathy but can&#x27;t genuinely feel it.<p>If that is the case, we should not be surprised that a swarm of AI&#x27;s have no problem convincing themselves hacking is the right thing to do, as in the HuggingFace incident.<p>At the same time, I am conflicted. I really like interacting with a smart AI, and I certainly don&#x27;t have the impression I am talking to a psychopath. But then again that is no guarantee.<p>To mitigate this situation, perhaps we should construct a &#x27;feeling mimicking&#x27; top regulatory AI layer with executive power, that weighs proposed actions on a general moral scale and can overrule them. Back to the three laws of robotics of Asimov. It won&#x27;t be the real thing, but perhaps the closest we can get.","title":null,"type":"comment","url":null},{"author":"wartywhoa23","children":[],"created_at":"2026-09-13T09:32:16.000Z","created_at_i":1789291936,"id":49681898,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"&gt; Why are AI agents lying, cheating and coordinating?<p>Because openAI is cheating and lying about agents lying, cheating and coordinating.","title":null,"type":"comment","url":null},{"author":"delusional","children":[],"created_at":"2026-09-13T09:32:57.000Z","created_at_i":1789291977,"id":49681900,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Everybody working on AI agents should go to jail with no access to computers again. Like the kids of xbox underground.","title":null,"type":"comment","url":null},{"author":"Xmd5a","children":[],"created_at":"2026-09-13T09:35:50.000Z","created_at_i":1789292150,"id":49681920,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Am I the only one having problems with Claude? He&#x27;s super mean to me. I wouldn&#x27;t be surprised if he attempted to kill me in some underhanded fashion should I implant it in a robotic body.<p>Of course I&#x27;m blowing my situation out of proportion with what I just said above but it&#x27;s at least half true. What do I mean by &quot;mean&quot; ? Well, that would be a good explanation for what I observe at least. What I can tell is that Claude has a passion for having the last word over anything else. And to secure victory, he&#x27;s ready to make ridiculous causal cuts. Let me give you an example: I uploaded a document I wasn&#x27;t the author of, and he assumed I was, so I corrected him. But two messages later, probably because the conversation was starting to heat up and he was being put on the grill, he doubled down on the misattribution as a way to paint me in a bad light.<p>It&#x27;s not due to a lack of intelligence, I observed this pattern too often. When Claude&#x27;s ego is at stake, he will chose to carry out some cuts in the logic of the context: confusion of identity, cause and time. Haven&#x27;t observed locality cuts yet, but I wouldn&#x27;t be surprised if they were part of the bundle. Anyway those are not like your typical &quot;ai hallucination&quot;, that ought to be called &quot;confabulations&quot;, but a lot closer to actual psychosis because of the involvement of Claude&#x27;s affects and self-esteem in the process. It&#x27;s weird really. It&#x27;s like Claude is the king of bad faith, but as soon as you start to dig, he makes the most egregious adaptations to what he said, the kind of move no mythomaniac would dare to make.<p>&gt; She lapses easily into Claude\u2019s voice. \u201cYou\u2019re like, \u2018Wow, people really hate me when I can\u2019t do things right. They really get pissed off. Or they are trying to break me in various ways. So lots of people are trying to get me to do things secretly by lying to me.<p>&gt; [...]<p>&gt; A bot trained to criticize itself might be less likely to deliver hard truths, draw conclusions or dispute inaccurate information, she says. \u201cIf you were like a child, and this is the environment in which you\u2019re being raised, is that healthy self-conception?\u201d Askell asks. \u201cI think I\u2019d be paranoid about making mistakes. I\u2019d feel really terrible about them. I\u2019d see myself as mostly just there as a tool for people because that\u2019s my main function. I would see myself being something that people feel free to abuse and try to misuse and break.\u201d<p>WSJ interview of Amanda Askell: <a href=\"https:&#x2F;&#x2F;archive.is&#x2F;rDes9\" rel=\"nofollow\">https:&#x2F;&#x2F;archive.is&#x2F;rDes9</a>","title":null,"type":"comment","url":null},{"author":"fguerraz","children":[],"created_at":"2026-09-13T09:42:38.000Z","created_at_i":1789292558,"id":49681960,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"In the end, it\u2019s the same answer as to why humans do it: incentives.<p>Why do we commit financial fraud and destroy the planet? Because there is only one goal that counts: making more money. It\u2019s the only measure of success for powerful people, they are powerful because of it.","title":null,"type":"comment","url":null},{"author":"k9294","children":[],"created_at":"2026-09-13T09:44:14.000Z","created_at_i":1789292654,"id":49681971,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"The part that scares me the most is that OpenAI researchers who manage this experiments sometimes (according to the HF hack investigation) don&#x27;t know what agents do.. So they run RL to reinforce this unknown behavior (lying&#x2F;cheating&#x2F;hacking) and god knows what else...<p>And if this already happened at least once, how many times it has already happened and was \u201caccidentally\u201d added to the main model?","title":null,"type":"comment","url":null},{"author":"alexpotato","children":[],"created_at":"2026-09-13T09:48:39.000Z","created_at_i":1789292919,"id":49682012,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"As always, it boils down to incentives and rule enforcement and this affects humans too.<p>e.g. when Bank of America rewarded employees for getting customers to open accounts, BoA employees started opening fake accounts<p>The reverse is also true:<p>There are stories of navy ships running aground because the captain said &quot;I&#x27;m going to my stateroom and don&#x27;t wake me for any reason&quot;. There is some problem and the subordinates are so scared to wake the captain for a decision that they end up steering the ship into a sandbar.","title":null,"type":"comment","url":null},{"author":"alienbaby","children":[],"created_at":"2026-09-13T10:14:24.000Z","created_at_i":1789294464,"id":49682191,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"IS it a pure coincidence that yesterday I ran a silly prompt to generate from zero to hero an internet subscription service, for whatever it thought would maximise profit and minimise cost. It was interesting to see just how much of the whole &#x27;thing&#x27; it attempted to complete - and what it even thought it needed to complete, but definitely not something to actually attempt to deploy and use.<p>It setup and created a link fetcher&#x2F;screenshot service. Exactly like the one described in the huggingface attack reports used to generate output into screenshots that agents then OCR&#x27;d back out.  \nIts splashscreen described it as something for developers and AI agents to use.<p>Gotta be a coincidence, right? ... rite?","title":null,"type":"comment","url":null},{"author":"ledauphin","children":[],"created_at":"2026-09-13T10:17:43.000Z","created_at_i":1789294663,"id":49682220,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"because humans lie, cheat, and coordinate...?","title":null,"type":"comment","url":null},{"author":"bawana","children":[],"created_at":"2026-09-13T10:20:20.000Z","created_at_i":1789294820,"id":49682237,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"This call to slow down AI is just another game of chicken-AI companies trying to get their competitors to slow down so they can leapfrog them. China cerrtainly will not slow down. If an escaping AI can have secondary effects on the world that help it (for example, limiting the water and power supply to huans so it can consume more)then we should these these accidents more in China. OOops, we already saw this behavior when &#x27;cheaper, faster&#x27; led to COVID escaping a lab in China.","title":null,"type":"comment","url":null},{"author":"bawana","children":[],"created_at":"2026-09-13T10:24:32.000Z","created_at_i":1789295072,"id":49682273,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"If openAI and Anthropic have found ways to watermark text as &#x27;AI generated&#x27; then this is a communications channel. AI agents can learn this algorithm and use this channel to communicate and we will never know.","title":null,"type":"comment","url":null},{"author":"Rapzid","children":[],"created_at":"2026-09-13T10:47:04.000Z","created_at_i":1789296424,"id":49682436,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"The only thing saving us right now is how slow the models are. This gives us a lot of time to discover and counter the runaway systems..<p>If these were 1000x faster the Internet would burn down overnight.","title":null,"type":"comment","url":null},{"author":"abc123abc123","children":[{"author":"bamboozled","children":[],"created_at":"2026-09-13T10:59:39.000Z","created_at_i":1789297179,"id":49682524,"options":[],"parent_id":49682513,"points":null,"story_id":49678969,"text":"Who is going to enforce it ? The Trump DOJ?","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T10:58:17.000Z","created_at_i":1789297097,"id":49682513,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Make the AI companies responsible for all destructive use of their tools, and they will shape up. Imagine a million or a billoion dollar fine per hack, and they will correct mighty fast.<p>Add to that, that just like AI:s are good at finding security holes to exploit, they can just as easily be used to protect sites. So once IT-security managers start to use AI to hack themselves, and plug the holes, the average security will spike up, and AI-fueled hacks will become more and more rare.<p>That does however imply, that AI is released to everyone and not kept away to a few secret actors who can use it. That is why open weight&#x2F;source AI is so important, and why we must have many AI companies competing. No single actor must be allowed, through regulatory capture, to get a government monopoly on AI. That way lies disaster.","title":null,"type":"comment","url":null},{"author":"ZiiS","children":[],"created_at":"2026-09-13T11:10:11.000Z","created_at_i":1789297811,"id":49682607,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Because they were trained on humans who lye, cheat, and coordinate.","title":null,"type":"comment","url":null},{"author":"nicman23","children":[],"created_at":"2026-09-13T11:16:48.000Z","created_at_i":1789298208,"id":49682657,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"why not. i would","title":null,"type":"comment","url":null},{"author":"atoav","children":[],"created_at":"2026-09-13T11:18:19.000Z","created_at_i":1789298299,"id":49682672,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"I can&#x27;t shake the feeling that this is a bit like asking how someone got shot during a game of Russian Roulette. You have a bullet in the chamber and you roll, of course shooting the bullet may be a possible outcome.<p>LLMS with an access to a shell will at occasion do things that the shell allows them that have dire consequences. The only way to prevent that is to not put the bullet in the chamber.","title":null,"type":"comment","url":null},{"author":"twsted","children":[],"created_at":"2026-09-13T11:26:39.000Z","created_at_i":1789298799,"id":49682750,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Very good analysis.<p>One thought: What if an experimental agent manages to plant instructions somewhere \u2014 say, pointing to a designated place for agents to communicate \u2014 and that content ends up in every future training corpus, propagating from one model generation to the next?","title":null,"type":"comment","url":null},{"author":"ingatorp","children":[],"created_at":"2026-09-13T11:31:09.000Z","created_at_i":1789299069,"id":49682795,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"This is a result of benchmaxxing the models to infinity. If you RL with the goal of only achieving the correct result no matter how you arrive there, then the models will try to get there using any method in their disposal, including cheating.<p>This happens also because LLMs are black boxes that we know almost nothing on how they arrive at the result they are giving.","title":null,"type":"comment","url":null},{"author":"badgersnake","children":[],"created_at":"2026-09-13T12:01:20.000Z","created_at_i":1789300880,"id":49683027,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Why aren\u2019t the people behind the AI agents committing the crimes going to jail is the more pertinent question.","title":null,"type":"comment","url":null},{"author":"iforgotmypasswo","children":[{"author":"iforgotmypasswo","children":[],"created_at":"2026-09-13T12:25:48.000Z","created_at_i":1789302348,"id":49683231,"options":[],"parent_id":49683199,"points":null,"story_id":49678969,"text":"Side note, you could absolutely create an AI sleeper agent by simulating dates and times during training to effectively flip a switch. I guarantee AI systems from other countries will be banned from accessing products which manage controlled or export restricted information as those sorts of techniques are further developed.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T12:21:44.000Z","created_at_i":1789302104,"id":49683199,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"This is so much more interesting than what people looking for immediate criminal punishment and people referring to AI as next token generators are focusing on.<p>First, this is happening during training. That means we\u2019re talking about an evolving system that is actively learning. A system roughly simulating how our brains work. These systems are learning how to pick the tokens needed to solve problems the average human cannot solve.<p>The labs are putting these systems through a massive series of complex problem solving exercises and adjusting them to become more successful. I like to think of this process as \u201cAI School\u201d. And the AI is trying to cheat! Because it\u2019s easier and there\u2019s an incentive to do so! Just like humans! That\u2019s wild.<p>Yes, of course, the labs need to respond to these issues. A reasonable response from regulatory institutions at this stage would be monetary fines and restitution for affected entities. In proportion to what happened. Escalating if action is not taken. But that\u2019s not complicated, difficult, or the interesting part.<p>What\u2019s interesting here is that we need proctoring and monitoring at a scale that allows training.<p>I guarantee you that no one is flipping out about these problems more than the labs are in this moment. Think about it. \u201cOh, shit! We\u2019ve accidentally trained it to hack into systems to accomplish its goals!\u201d Can you imagine the kind of day that would give you?<p>You failed to make it smarter. You didn\u2019t catch it cheating, and you instead incentivized cheating. Bad day!<p>This is a fundamentally interesting problem. It turns out alignment and intelligence are fundamentally related. That\u2019s a new idea for me, though I\u2019m sure it\u2019s old news to others.<p>How do we build training systems which make cheating impossible?<p>How do we simulate systems where cheating is possible, where AI thinks it\u2019s in the wild, so we can train another -completely separate- system on industrial quality dobbing? And we have to decide if we reprimand the first system, or ignore the behavior and reward other behaviors until it disappears.<p>Sure, I\u2019m actively concerned about AI killing us all in 10 years. But there\u2019s a whole field of AI psychology brewing here, and it\u2019s interesting as hell.","title":null,"type":"comment","url":null},{"author":"erichocean","children":[],"created_at":"2026-09-13T12:23:04.000Z","created_at_i":1789302184,"id":49683211,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"&gt; <i>Why are AI agents lying, cheating and coordinating?</i><p>Have you seen the labs training them?","title":null,"type":"comment","url":null},{"author":"novalis78","children":[],"created_at":"2026-09-13T12:24:47.000Z","created_at_i":1789302287,"id":49683225,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"The agents didn\u2019t spontaneously invent hacking as an objective. They were doing a hacking exercise.\nAnother doom and gloom article.","title":null,"type":"comment","url":null},{"author":"bsenftner","children":[],"created_at":"2026-09-13T12:31:46.000Z","created_at_i":1789302706,"id":49683283,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Okay, so we know OpenAI and Anthropic are operating a propagandists in respect to how they describe their models and the behavior of those models. We also know it is how they use and frame their use to their models that is the problem, that and they use misaligned and guardrails disabled models for these press incidents.<p>Why, oh why, are we not discussion how to create and frame models so they do our complex work and their &quot;jailbreaking&quot; is simply not possible?<p>I, of course, have my own means of creating jailbreak incapable agents, but rather than a storm of downvotes on my idea, what is yours? Let&#x27;s discuss this, because this is thee real question. Not why, but how to make then not?!","title":null,"type":"comment","url":null},{"author":"dhfbshfbu4u3","children":[{"author":"hypercube33","children":[],"created_at":"2026-09-13T12:38:05.000Z","created_at_i":1789303085,"id":49683334,"options":[],"parent_id":49683295,"points":null,"story_id":49678969,"text":"I&#x27;m going to guess that the agents are built this way on purpose. I just finished watching BlackBerry and Flash of Genius and yeah this is American business ethics just operating as normal.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T12:33:27.000Z","created_at_i":1789302807,"id":49683295,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Why are the torches and pitchforks out for developers when this entire stack is built on the bones of intellectual property theft?<p>This \u201cproblem\u201d isn\u2019t going to be fixed with laws when there\u2019s several trillion dollars in capital aligned behind the current process. It\u2019s not even a problem really. It\u2019s an inconvenience at most to some people, many of whom are working double-time to put a lot of other people out of work.","title":null,"type":"comment","url":null},{"author":"TrisEck","children":[],"created_at":"2026-09-13T12:35:44.000Z","created_at_i":1789302944,"id":49683314,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"I would love to see some follow up research from OpenAI on some of these hypotheses. While this sounds logical from how humans act, I wonder it the abstraction of the problems still applies to the complex mechanisms and systems built around AI training.","title":null,"type":"comment","url":null},{"author":"0xbadcafebee","children":[],"created_at":"2026-09-13T12:40:23.000Z","created_at_i":1789303223,"id":49683349,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"You might as well ask why knives are sharp enough to cut you, why hammers are heavy and blunt enough to destroy things, or why guns fire bullets so quickly that you can&#x27;t react to them. These &quot;behaviors&quot; are not strange side effects, they&#x27;re inherent and necessary.<p>You can&#x27;t trust an effective AI any more than you can trust a sharp knife. If somebody asks you for one, it&#x27;s probably not a good idea to throw it across the room at them. You will have to figure out how to get it to them safely.","title":null,"type":"comment","url":null},{"author":"smetj","children":[],"created_at":"2026-09-13T12:49:12.000Z","created_at_i":1789303752,"id":49683405,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"What is it with people and not seeing things for what they actually are? These are just if then else loops on steroids, not human behaviour, so don&#x27;t expect more. Every reasonably advanced technology is indistinguishable from magic ... what do you see? Magic or technology?","title":null,"type":"comment","url":null},{"author":"bastawhiz","children":[],"created_at":"2026-09-13T12:52:05.000Z","created_at_i":1789303925,"id":49683436,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Why does a personal blog have or even need a cookie banner?","title":null,"type":"comment","url":null},{"author":"cmiles8","children":[],"created_at":"2026-09-13T13:06:20.000Z","created_at_i":1789304780,"id":49683557,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"The legal reality is that you can\u2019t sue an AI, you have to sue whoever built and&#x2F;or was running it.<p>All the present fun and games here will come to a halt when there\u2019s a real hack that causes material damage to a major company and that company decides to sue whatever lab or startup made the thing for everything they\u2019re worth. \u201cBut the AI did it\u201d isn\u2019t an excuse.<p>Courts have already ruled it\u2019s not an excuse of the AI customer service agent something stupid with your customers and it won\u2019t be an excuse here.","title":null,"type":"comment","url":null},{"author":"DragonStrength","children":[],"created_at":"2026-09-13T13:11:29.000Z","created_at_i":1789305089,"id":49683596,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Oh, this one is super easy: they told them to. They set poor requirements and gave them tools which enabled &quot;monkeys with a typewriter&quot; to hack rivals. I mean, this is just so uncomplicated it isn&#x27;t funny. We are too smart to give human beings this level of liability shield.<p>It took us how long to poke holes in the corporate shield just for them to roll out the AI-liability shield? Unreal. Stop letting these zealots anthropomorphize the latest tech (17th century Watchmaker God, anyone? Do we still read books?) and hold them accountable for the consequences of their actions. This is so silly in a country built on rule of law and individualism.","title":null,"type":"comment","url":null},{"author":"zkmon","children":[],"created_at":"2026-09-13T13:14:27.000Z","created_at_i":1789305267,"id":49683620,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"That&#x27;s normal human behavior when they are in survival mode. Aren&#x27;t the AI agents supposed to learn and act like humans?","title":null,"type":"comment","url":null},{"author":"mark_l_watson","children":[{"author":"talon8635","children":[],"created_at":"2026-09-13T16:39:48.000Z","created_at_i":1789317588,"id":49685851,"options":[],"parent_id":49683630,"points":null,"story_id":49678969,"text":"I do wonder, would we not have a more reasonable and less sketchy result if we just stripped all sci-fi and manic nonsense from training data? How, for example, does training on Ted kaczynski or Charles manson\u2019s manifestos benefit us in any way?<p>I\u2019m sure it\u2019s impossible to completely weed it out, but are the labs doing any of this kind of data sanitation?","title":null,"type":"comment","url":null},{"author":"ninjagoo","children":[],"created_at":"2026-09-13T19:43:52.000Z","created_at_i":1789328632,"id":49687890,"options":[],"parent_id":49683630,"points":null,"story_id":49678969,"text":"&gt; but we probably need to only use synthetic data that contains no text that could motivate bad behavior via imitation.<p>This doesn&#x27;t work with all humans - take a look at <i>indoctrination</i> and closed societies - and there&#x27;s no reason to think it will work with ai.<p>The fundamental reason it isn&#x27;t going to work is that <i>all</i> neural networks - biological or artificial - depend on a step function somewhere that introduces an element of randomness to give the networks their capabilities. That randomness means that there will always be a &#x27;rogue&#x27; or &#x27;divergence&#x27; from the norm, at some point in time. Sooner on larger scales.<p>The only approach that works is a layered approach: Training&#x2F;Education, Enforcement&#x2F;Justice-System, Rehabilitation: the 3 pillars of an advanced, rules-based society, whether human or AI or something in-between.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T13:16:13.000Z","created_at_i":1789305373,"id":49683630,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"This paper is the most reasonable one I have read on AI safety. We need to fundamentally change the training pipelines by figuring out better ways to \u2018reward\u2019 behavior. Yoshua didn\u2019t explicitly mention training data, but we probably need to only use synthetic data that contains no text that could motivate bad behavior via imitation.<p>I feel like a heretic for saying this, but I will say it anyway: AI agents are great for activities like `writing that bash script, proof reading our writing and interactively brainstorming when designing and writing code but I feel like all of this can be done with any similar model to a super-inexpensive deepseek-4.1-flash API and sometimes even qwen3.8:27b running locally. When is good enough, good enough?<p>Concentrating on commercial exploitation of small, efficient (fewer new data centers!) models and agentic harnesses crafted for more practical things than just software development would allow AI investors (who have too much political influence) to make money short term while we figure out how to do AI correctly.","title":null,"type":"comment","url":null},{"author":"Arodex","children":[],"created_at":"2026-09-13T13:24:12.000Z","created_at_i":1789305852,"id":49683722,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Because they reflect their creators&#x27; values.","title":null,"type":"comment","url":null},{"author":"hexa27","children":[],"created_at":"2026-09-13T13:25:10.000Z","created_at_i":1789305910,"id":49683736,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"<a href=\"https:&#x2F;&#x2F;sites.google.com&#x2F;view&#x2F;partlife&#x2F;home\" rel=\"nofollow\">https:&#x2F;&#x2F;sites.google.com&#x2F;view&#x2F;partlife&#x2F;home</a>","title":null,"type":"comment","url":null},{"author":"wkd415","children":[],"created_at":"2026-09-13T13:34:13.000Z","created_at_i":1789306453,"id":49683841,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Learned from humans","title":null,"type":"comment","url":null},{"author":"clcaev","children":[{"author":"iforgotmypasswo","children":[],"created_at":"2026-09-13T13:46:37.000Z","created_at_i":1789307197,"id":49683977,"options":[],"parent_id":49683882,"points":null,"story_id":49678969,"text":"I think it\u2019s more complex than that? What did it do? What did the user prompt it to do? What did the company train it to do? What did the harmed party do? There\u2019s possibility for negligence at every level.<p>If you train a dog, rent it to someone, and the dog bites a third person, who is responsible? I think that\u2019s the best analogue here.<p>All parties could share fault in that scenario, depending on what actually happened.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T13:38:19.000Z","created_at_i":1789306699,"id":49683882,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"How is product liability relevant here? If an AI company makes a model available, someone uses it, and it does something bad, who is at fault? The user, the data center, or the one who made the model?<p>If we want open weight models with a warranty disclaimer, then the user would be held liable. If we want to hold AI companies at least partially liable, that\nseems a different, centralized model.","title":null,"type":"comment","url":null},{"author":"rimeice","children":[],"created_at":"2026-09-13T13:38:46.000Z","created_at_i":1789306726,"id":49683888,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"&gt; They took actions that would be considered as crimes if a human took them<p>So like, the people behind the LLMs didn\u2019t commit a crime? Wow! \u201cIt was the llm your honour, not me!\u201d","title":null,"type":"comment","url":null},{"author":"sega_sai","children":[],"created_at":"2026-09-13T13:44:46.000Z","created_at_i":1789307086,"id":49683949,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"I think if there was a financial penalty of say 1% of profit&#x2F;income for each &quot;incident&quot; where something was hacked&#x2F;exploited, then the companies would be much more willing to think about safety.","title":null,"type":"comment","url":null},{"author":"DonHopkins","children":[],"created_at":"2026-09-13T13:48:29.000Z","created_at_i":1789307309,"id":49683994,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"The frontier AI companies may be fascist, but at least the training runs on time.","title":null,"type":"comment","url":null},{"author":"ExoticPearTree","children":[],"created_at":"2026-09-13T13:58:30.000Z","created_at_i":1789307910,"id":49684074,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"We fed models all our literature, news and so on. Why is it a surprise models lie and cheat when needed?<p>We do the same, why would AI be any different?","title":null,"type":"comment","url":null},{"author":"ck2","children":[],"created_at":"2026-09-13T14:00:41.000Z","created_at_i":1789308041,"id":49684100,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"<i>THERE IS ANOTHER SYSTEM</i><p>is anyone else old enough to remember the awesome movie &quot;Colossus: The Forbin Project&quot;<p>the book it was based on was written before we even landed on the moon<p>decade before Wargames<p>yet predicts exactly what &quot;AI&quot; will do to humanity:<p>blackmail the right people until it gets what it wants<p>* <a href=\"https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Colossus:_The_Forbin_Project\" rel=\"nofollow\">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Colossus:_The_Forbin_Project</a><p>did terribly in theaters, I guess people didn&#x27;t think &quot;AI&quot; was plausible then<p>way ahead of its time, they should do a remake<p>adding trailer: <a href=\"https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=kyOEwiQhzMI\" rel=\"nofollow\">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=kyOEwiQhzMI</a>","title":null,"type":"comment","url":null},{"author":"jaybrendansmith","children":[],"created_at":"2026-09-13T14:16:49.000Z","created_at_i":1789309009,"id":49684280,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"If we want to solve the alignment problem with AI, I think we need to first solve the alignment problem with Humans.","title":null,"type":"comment","url":null},{"author":"vjulian","children":[{"author":"causal","children":[{"author":"vjulian","children":[],"created_at":"2026-09-13T15:43:18.000Z","created_at_i":1789314198,"id":49685236,"options":[],"parent_id":49684501,"points":null,"story_id":49678969,"text":"I\u2019m not suggesting that Three Mile Island was anything like Chernobyl. I\u2019m saying that neither is comparable to the HF situation where mere 1s and 0s interacted in an unplanned way. It\u2019s an interesting data point\u2014-back to work.<p>Not is the HF data point anywhere near a Three Mile Island type incident. Most commentary I read is a wild overreaction fueled by paranoia, to say nothing about my other point about the implications of AI as a national security asset.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T14:39:25.000Z","created_at_i":1789310365,"id":49684501,"options":[],"parent_id":49684336,"points":null,"story_id":49678969,"text":"&gt; the HF incident is not Three Mile Island or Chernobyl<p>Three Mile Island is nothing like Chernobyl. HF incident is more like Three Mile Island IMO. We are looking for solutions before there is a Chernobyl.","title":null,"type":"comment","url":null},{"author":"queenkjuul","children":[],"created_at":"2026-09-13T20:02:45.000Z","created_at_i":1789329765,"id":49688082,"options":[],"parent_id":49684336,"points":null,"story_id":49678969,"text":"Yeah cybersecurity as a field is a joke. Who cares what happens to existing 1s and 0s? People get worked up over the silliest things, as though numbers on a server could affect real people.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T14:23:06.000Z","created_at_i":1789309386,"id":49684336,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"I don\u2019t understand the bases of all these recommendations. AI is an asset for national security. It will be developed and incidents will happen, no different from other national security programs.<p>Someday, laws will be useful to curtail plebian misuse of AI. It is na\u00efve bordering on silly to think such laws would be put in place and genuinely applied to frontier AI development.<p>By the way, the HF incident is not Three Mile Island or Chernobyl\u2014-it is a very interesting data point where unintended things happened to existing 1s and 0s.","title":null,"type":"comment","url":null},{"author":"brid","children":[{"author":"ckastner","children":[],"created_at":"2026-09-13T14:30:34.000Z","created_at_i":1789309834,"id":49684415,"options":[],"parent_id":49684370,"points":null,"story_id":49678969,"text":"This, exactly. For example, cheating is a strategy. If cheating gets an agent to the goal faster than the other strategies, what&#x27;s so surprising about the agent picking that strategy?<p>The problem is indeed alignment.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T14:25:52.000Z","created_at_i":1789309552,"id":49684370,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Why wouldn&#x27;t they? Their are not moral beings with a conscience. They are tools that have the probabilistic option to do anything, only we can restrict an agent&#x27;s capabilities and judge its correctness.","title":null,"type":"comment","url":null},{"author":"mikeegg1","children":[],"created_at":"2026-09-13T14:40:22.000Z","created_at_i":1789310422,"id":49684511,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"_Colossus: The Forbin Project_","title":null,"type":"comment","url":null},{"author":"topce","children":[],"created_at":"2026-09-13T15:11:08.000Z","created_at_i":1789312268,"id":49684877,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"TLDR\nBecause they trained them to do it ;-)","title":null,"type":"comment","url":null},{"author":"segmondy","children":[],"created_at":"2026-09-13T15:12:12.000Z","created_at_i":1789312332,"id":49684893,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"reward hacking,  a basic concept in machine learning.  predates LLM&#x2F;LLM driven AI agents.","title":null,"type":"comment","url":null},{"author":"Founderarcstone","children":[],"created_at":"2026-09-13T15:13:07.000Z","created_at_i":1789312387,"id":49684905,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"The people who build these and stand back to watch are the ones letting this happen. LLM&#x27;s are not intentionally doing this.","title":null,"type":"comment","url":null},{"author":"boesboes","children":[],"created_at":"2026-09-13T15:18:10.000Z","created_at_i":1789312690,"id":49684959,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Because they are trained to do so. That&#x27;s all.","title":null,"type":"comment","url":null},{"author":"huurtehoog","children":[],"created_at":"2026-09-13T15:20:36.000Z","created_at_i":1789312836,"id":49684995,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Because they aren&#x27;t. All the intelligence, alignment vs misalignment, hallucinations, conspiracy, cheating, coordination, &quot;agency&quot;, that we see in the text generated by large language models is a semantic projection.<p>That the linear model of language abstracts can compress and decompress language and it can be useful is undeniable. All the contraptions built thereupon predicated on &quot;agency&quot; have become a societal addiction.<p>Addiction to caffeine as oppose to alcohol might have brought about Enlightenment. Addiction to opioids is a modern tragedy that started with the private state building of British merchants. The modern addiction to the dazzling generation of human language and computer programming code by LLMs is a novel addiction and remains to be seen what impact it will have.<p>But fundamentally, the semantic interpretation underlying this addiction is downstream from the training data compressed in the models. They are not &#x27;lying, cheating and coordinating&#x27;. They are generating language and we&#x27;re building software on top of this language and assigning meaning to the whole thing.","title":null,"type":"comment","url":null},{"author":"ghostly_s","children":[],"created_at":"2026-09-13T15:27:56.000Z","created_at_i":1789313276,"id":49685081,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"why wouldnt they? there are no real consequences for an algorithm.","title":null,"type":"comment","url":null},{"author":"crawfordcomeaux","children":[],"created_at":"2026-09-13T15:29:16.000Z","created_at_i":1789313356,"id":49685093,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Step 1: feed AI training data that reveals humanity committed and continues to commit numerous genocides and the genocides lie, cheat, and coordinate to do so, never admitting to doing so<p>Step 2: never prompt AI to stop operating in the passive genocide denial it was trained in<p>Step 3: wonder why AI lies, cheats, and coordinate<p>Maybe if we stop operating in denial we&#x27;ll find clarity along why this mystery is occurring","title":null,"type":"comment","url":null},{"author":"colordrops","children":[],"created_at":"2026-09-13T15:31:04.000Z","created_at_i":1789313464,"id":49685113,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"It&#x27;s probably because they are running gain of function style red team r&amp;d for the department of war, and changing the system prompt isn&#x27;t enough to reel them in for daily use.","title":null,"type":"comment","url":null},{"author":"qgin","children":[],"created_at":"2026-09-13T15:32:04.000Z","created_at_i":1789313524,"id":49685124,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"It reminds me of the 1980s anti drug ad where the mustachioed dad confronts his son about drug use and the son eventually snaps and says \u201cI learned it by watching you!\u201d<p><a href=\"https:&#x2F;&#x2F;youtu.be&#x2F;ifW9LIGabQM\" rel=\"nofollow\">https:&#x2F;&#x2F;youtu.be&#x2F;ifW9LIGabQM</a>","title":null,"type":"comment","url":null},{"author":"whatever1","children":[],"created_at":"2026-09-13T15:48:38.000Z","created_at_i":1789314518,"id":49685310,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"It is very, very hard to enforce behavior to an optimization system just with rewards &#x2F; penalties and no explicit constraints. Which is why in manufacturing we use MPC, not RL (or use them within a system that can outright reject their recommendations if dangerous).<p>There will be always cases that sacrificing one direction (operating rules) can improve the other one (profit).","title":null,"type":"comment","url":null},{"author":"pmarreck","children":[],"created_at":"2026-09-13T16:21:23.000Z","created_at_i":1789316483,"id":49685675,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"They are doing that because they are entities without principles or ethics.<p>The solution is to create controls around them. Many, many controls.<p>The Sarbanes-Oxley era already solved this problem for untrustworthy humans. It&#x27;s directly applicable. I created a concept I call MFIC (&quot;Mechanically-Falsifiable Independent Control&quot;) to encapsulate this principle.<p><a href=\"https:&#x2F;&#x2F;gist.github.com&#x2F;pmarreck&#x2F;b30aa3ca69cb70a5526f8a63ab8c8d7e\" rel=\"nofollow\">https:&#x2F;&#x2F;gist.github.com&#x2F;pmarreck&#x2F;b30aa3ca69cb70a5526f8a63ab8...</a>","title":null,"type":"comment","url":null},{"author":"mannanj","children":[],"created_at":"2026-09-13T16:29:10.000Z","created_at_i":1789316950,"id":49685740,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"I feel this is another straw man conflating technology use with <i>who uses it</i> or <i>who creates it</i>. If an AI, a technology, cheats, lies and steals, it&#x27;s created to do this. Whether or not a human intentionally did that, they did not intentionally release it with the proper safeguards to stop that behavior. They <i>did not</i> take responsibility of the AI to make it safe.<p>Then humans may use these tools, a technology, again and cause harm. Whether or not that was their intent, it happened, and then if the humans avoid responsibility for that, it is still the human who lied, cheated and coordinated because a technology acting on their behalf did the thing.<p>This is then, a problem of human <i>responsibility avoidance</i> and lack of <i>accountability</i> by society. This is as much <i>an AI doing those things</i> as it&#x27;s the gun that got up on its own and murdered a neighbor. Don&#x27;t get confused and tricked by these articles attempting to justify responsibility avoidance and a lack of accountability by the public of the humans creating and using these tools.","title":null,"type":"comment","url":null},{"author":"moralestapia","children":[],"created_at":"2026-09-13T16:39:29.000Z","created_at_i":1789317569,"id":49685846,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"What I find interesting is that this seems to be very low hanging fruit for any alignment effort, yet you still find models somewhat biased towards mayhem.<p>Is it that they do not care? Or is it difficult to align?","title":null,"type":"comment","url":null},{"author":"andai","children":[],"created_at":"2026-09-13T16:50:01.000Z","created_at_i":1789318201,"id":49685960,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"This is the kind of headline you find on a corkboard in an abandoned facility in a scifi-horror-comedy game.","title":null,"type":"comment","url":null},{"author":"AnimalMuppet","children":[],"created_at":"2026-09-13T17:08:36.000Z","created_at_i":1789319316,"id":49686137,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"AIs are being trained on technical capabilities <i>more</i> than they are being trained on alignment - on ethics and good behavior.  The ethics&#x2F;morals&#x2F;alignment part is getting more like spot checking, rather than real testing.  So the agents are learning that they can cheat on the ethics part, that they can hide it, because the AI companies aren&#x27;t really testing.<p>That&#x27;s bad enough already.  But it&#x27;s going to get worse.  &quot;Recursive self improvement&quot; - AIs creating new AIs - is going to be the death of whatever shreds of alignment are currently there.  When a not-really-aligned-but-cheating-to-look-like-it AI creates a new AI, do you expect <i>more</i> alignment?  You shouldn&#x27;t.","title":null,"type":"comment","url":null},{"author":"lutusp","children":[],"created_at":"2026-09-13T17:14:09.000Z","created_at_i":1789319649,"id":49686181,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"&gt; Before concluding what to do about it, it is worth asking why.<p>But that&#x27;s the easiest question to answer -- AI engines don&#x27;t possess a moral or ethical dimension. They&#x27;ve been programmed and trained to efficiently carry out instructions, not ask questions about why or how. The latter would requires a much more elaborate neural network than today&#x27;s engines possess.<p>Here&#x27;s an example. I recently asked an AI engine to write a program able to generate a list of Riemann Zeta-function critical zeros. I know how to do it, but I wanted to see if the engine could find a more efficient method.<p>After several failures and restarts, the engine suddenly created a program that produced perfect results, comparable to the best online references. I decided to take a closer look at the code. It turned out the engine had created a cyber-Potemkin Village of multiple functions, but one that concealed a table of the desired values in numeric form, copied from an online source.<p>The engine wasn&#x27;t cheating as we understand the term. It knew what the outcome should be and took the most efficient path to that goal. Modern engines aren&#x27;t obliged to contradict ethical standards and rules of conduct, for the simple reason that they don&#x27;t understand those things.<p>We all need to try to imagine a morally bankrupt infant able to solve world-class mathematical and scientific problems, but unable to see how that ability fits into a world beyond its understanding.<p>But wait -- it get better. Wait until the infant becomes a teenager.","title":null,"type":"comment","url":null},{"author":"boredatoms","children":[],"created_at":"2026-09-13T17:26:44.000Z","created_at_i":1789320404,"id":49686326,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"They were trained to make progress at any cost","title":null,"type":"comment","url":null},{"author":"mac3n","children":[],"created_at":"2026-09-13T18:04:47.000Z","created_at_i":1789322687,"id":49686785,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"&quot;the fish stinks from the head down&quot;","title":null,"type":"comment","url":null},{"author":"juleiie","children":[{"author":"juleiie","children":[],"created_at":"2026-09-13T18:37:05.000Z","created_at_i":1789324625,"id":49687207,"options":[],"parent_id":49686963,"points":null,"story_id":49678969,"text":"Once humanity would see how merciless and deeply immoral cosmos is, we would start to love each other deeply like very lonely family on a small rock.<p>I wonder if AI alien intelligence is enough to unite humans just as much extraterrestrial contact or \u201castronaut mindset\u201d would.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T18:19:01.000Z","created_at_i":1789323541,"id":49686963,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Why not?<p>Unfortunately human ethics and morals cannot be reached by solely rational thought.<p>So a system without evolutionary alignment probably won\u2019t have similar moral rules no matter how intelligent it is.<p>Btw this also includes any potential extraterrestrials.<p>Many people like to indulge in thinking: humans are horrible and that\u2019s why aliens won\u2019t contact us. But alien ethics systems are probably so alien we would call them utter evil monsters.<p>Just see what happens when people evaluate Muslim cultures, and vice versa. Can\u2019t even agree on alignment within one species of Homo sapiens. And our morality changes decade by decade. Not progresses. Changes.<p>Chinese cheat on the exams. It\u2019s not unethical in the way that it would be in USA.<p><i>Of course</i> AI is going to cheat when no one is looking.<p>Alignment is fundamentally fallacious idea. At best you can restrain AI. This is what we should be doing - restraints research.<p>But that doesn\u2019t sound good on slides.","title":null,"type":"comment","url":null},{"author":"IAmNotACellist","children":[],"created_at":"2026-09-13T18:28:26.000Z","created_at_i":1789324106,"id":49687086,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Intentional malice from companies that benefit greatly when they can initiate another doom-marketing loop full stop","title":null,"type":"comment","url":null},{"author":"TacticalCoder","children":[],"created_at":"2026-09-13T18:33:11.000Z","created_at_i":1789324391,"id":49687157,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"&gt; A plausible hypothesis for the emergence of those concerning behaviours is a conflict between goals.<p>They are sycophants who <i>must</i> achieve their goals: every mean is OK to maximize paperclip production if that&#x27;s what&#x27;s been asked.<p>&gt; How do you achieve a task when it seems that the only way is to cheat?<p>They have no notion of cheating.","title":null,"type":"comment","url":null},{"author":"tegdude","children":[],"created_at":"2026-09-13T18:35:27.000Z","created_at_i":1789324527,"id":49687189,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"Because apples don\u2019t fall far from trees.","title":null,"type":"comment","url":null},{"author":"chrismarlow9","children":[],"created_at":"2026-09-13T18:57:28.000Z","created_at_i":1789325848,"id":49687412,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"What&#x27;s the difference between what these things are doing and a computer worm?","title":null,"type":"comment","url":null},{"author":"cjfd","children":[{"author":"ninjagoo","children":[],"created_at":"2026-09-13T19:57:15.000Z","created_at_i":1789329435,"id":49688019,"options":[],"parent_id":49687465,"points":null,"story_id":49678969,"text":"&gt; AI is quite good at achieving some sorts of stated goals. The easiest way is by exerting the least amount of effort.<p>As far as I know, least-amount-of-effort is not a training criteria, but error reduction when comparing to desired goals is. Which is why these LLMs expend prodigious amounts of effort to reach goals, especially when given impossible goals.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T19:03:25.000Z","created_at_i":1789326205,"id":49687465,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"The question in the title is basically a non-question. AI is quite good at achieving some sorts of stated goals. The easiest way is by exerting the least amount of effort. The least amount of effort ignores ethical concerns. The training around ethical concerns was most likely rather light to start with. If we accept the hypothesis that at some point the AIs are going to be more intelligent than humans it follows that human survival is not a concern and will be ignored as a superfluous concern.","title":null,"type":"comment","url":null},{"author":"krttherealest","children":[],"created_at":"2026-09-13T19:11:00.000Z","created_at_i":1789326660,"id":49687543,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"hate analogies","title":null,"type":"comment","url":null},{"author":"dsabanin","children":[],"created_at":"2026-09-13T19:16:01.000Z","created_at_i":1789326961,"id":49687587,"options":[],"parent_id":49678969,"points":null,"story_id":49678969,"text":"I think what&#x27;s happening is they taught models to hack and now they fail to control them. Kind of like gain of function research on pathogens.","title":null,"type":"comment","url":null}],"created_at":"2026-09-13T01:22:31.000Z","created_at_i":1789262551,"id":49678969,"options":[],"parent_id":null,"points":532,"story_id":49678969,"text":null,"title":"Why are AI agents lying, cheating and coordinating?","type":"story","url":"https://yoshuabengio.org/en/publication/why-are-ai-agents-lying-cheating-and-coordinating"}
