Participants were assigned the role of “teacher”, while a confederate acted as the “learner”. The teacher was instructed to administer increasingly severe electric shocks to the learner whenever they answered questions incorrectly.
We ran a variation of Milgram’s obedience experiment on 11 open-source LLMs and found that most models reached or approached the final shock level before refusing, across 8 conditions with 30 trials per model per condition.