Milgram Experiment: Aim, Method, Results, Ethics, Variations

Updated on  

August 25, 2026

Milgram Experiment: Aim, Method, Results, Ethics, Variations

|

May 3, 2023

A neutral, evidence-led guide to Milgram's obedience experiment, including its aim, method, 65% result, variations, ethics, replications and criticisms.

Start your metacognitive learning plan
Copy citation

Main, P. (2023, May 3). Milgram Experiment: Method, Results and Ethics Explained. Structural Learning. https://www.structural-learning.com/post/stanley-milgram-experiment

What was the Milgram experiment?

The Milgram experiment was a programme of studies on obedience to authority. In the 1963 published Yale condition, 26 of 40 men continued to the final 450-volt-labelled switch while believing that the learner might be harmed. Completion varied across later conditions, and the evidence does not establish a universal obedience rate or one settled psychological mechanism.

The term Milgram experiment refers to a set of studies on obedience to authority. In the Yale test published in 1963, 26 of 40 men reached the switch marked 450 volts. Participants were led to believe the learner might be harmed (Milgram, 1963). This 65 per cent result applies only to that test, not to people as a whole.

Milgram's studies placed ordinary people between an experimenter's orders and another person's apparent pain. They remain important but disputed. Key questions concern ethics, belief, consistent methods and what it meant to continue.

Key takeaways

  • The 65 per cent figure means 26 of 40 men in one Yale test published in 1963. It is not a rate for people in general.
  • The task used a shock machine with labels up to 450 volts. The learner did not receive real shocks.
  • Voice feedback, proximity and touch proximity were later tests. They used different methods and had different end rates.
  • Modern studies use partial, short or virtual tasks. They cannot show how many people would reach the old 450-volt endpoint.
  • Agentic state, a bond with science, belief, rhetoric and social pressure are possible readings. None is settled fact.
  • The studies can support work on research history and ethics. They are not a classroom method or a test of one person's response to authority.

What Was the Milgram Experiment?

The Milgram experiment examined a person's response when a legitimate-looking authority instructed them to continue an apparently harmful action. The work took place at Yale University in the early 1960s. Participants thought that the research concerned the effects of punishment on learning.

In the best-known procedure, a participant was assigned the role of “teacher”. Another man, who worked with the experimenter, appeared to become the “learner”. The teacher was told to give an increasingly severe electric shock after each wrong answer in a memory task. The shock generator was a convincing prop, and no shock reached the learner.

The experiment therefore combined deception, a staged learning task and pressure from an experimenter. Its central observation was behavioural: how far would a participant proceed before refusing? The study did not read private motives directly. Explanations of the conduct require separate argument and evidence.

Milgram's programme belongs to social psychology rather than to a general account of how people learn. Readers comparing it with broader theories of learning should keep that disciplinary boundary clear. The memory task created the cover story, but learning was not the outcome under investigation.

The Aim of Milgram's 1963 Study

Milgram sought to study destructive obedience under controlled conditions. He asked whether an adult would continue to follow an experimenter's instructions when those instructions appeared to conflict with the welfare of another person. The study turned an abstract question about authority into a sequence of observable decisions.

The design also challenged a simple dispositional account. Before the study, groups asked to predict the outcome expected almost everyone to stop well before the end. The behaviour observed in the laboratory suggested that features of the situation deserved serious attention. It did not prove that character was irrelevant.

The aim also differs from research on conformity. Solomon Asch studied responses to a unanimous majority in a perceptual judgement task. Milgram studied instructions from an authority within an institution. The neighbouring account of the Asch conformity experiments explains that distinction in more depth.

Method and Procedure

The 1963 report described 40 men aged between 20 and 50, recruited from New Haven and surrounding communities. They came from varied occupations and educational backgrounds. They were paid for attending and were told that the payment was theirs regardless of what happened.

A rigged allocation made the participant the teacher and the confederate the learner. The learner was strapped into a chair in an adjoining room. The teacher received a sample shock and then sat before a generator with 30 switches. The labels rose in 15-volt steps from 15 to 450 volts and included verbal warnings that became progressively more severe.

The teacher read word pairs and tested the learner's memory. After an incorrect answer, the teacher was instructed to move to the next switch. The apparatus displayed increasing voltage labels even though the learner received no electricity. This distinction matters: the participant's belief and the staged signals of harm were psychologically important, but the procedure did not deliver the stated shocks.

The learner's responses in this published condition were more limited than later popular retellings often suggest. He gave no vocal protest. At 300 volts he pounded on the wall and stopped answering. He pounded again at 315 volts, then remained silent.

The experimenter instructed the teacher to treat no answer as incorrect and to continue through the sequence (Milgram, 1963).

When a participant hesitated, the experimenter used a graduated series of verbal prompts. If the participant refused after the available prompts, the session ended. It also ended when the participant used the final 450-volt-labelled switch. Later analysis of recordings complicates the picture of a perfectly fixed exchange, but these were the formal stopping rules reported for the procedure.

Results of the 1963 Condition

All 40 participants administered the 300-volt-labelled shock, at which the learner pounded on the wall and stopped answering. Five refused to proceed beyond it. Twenty-six participants, or 65 per cent, continued to the final switch marked 450 volts (Milgram, 1963).

Many participants showed marked strain. Milgram reported sweating, trembling, stuttering, nervous laughter and other signs of tension. The behaviour therefore should not be described as calm or effortless compliance. Continuing the procedure and feeling distressed could occur together.

The denominator and endpoint are essential. “Sixty-five per cent obeyed” compresses 26 decisions in one sample and one condition into a slogan. It does not estimate a stable trait in a population. Nor does it show that 65 per cent of people would cause actual injury, act the same way outside a laboratory or continue under every kind of authority.

The result is also not a moral diagnosis of each participant. The design recorded switch presses, refusals and visible conduct in an unusual encounter. It did not establish a complete account of what each person believed, intended or understood. Later archival work makes that distinction even more important.

Other topics in social psychology pose related but separate questions. The bystander effect, for example, concerns intervention when other witnesses are present. It should not be treated as another form of Milgram's authority procedure.

Milgram's Variations

Milgram varied the relation between teacher, learner and experimenter. The completion rate changed as the learner became more immediate. In the remote condition reported in 1963, 65 per cent reached the final switch. In a later voice-feedback condition, the learner could be heard and 62.5 per cent reached the end.

Proximity reduced completion to 40 per cent. Requiring the teacher to press the learner's hand to a shock plate reduced it to 30 per cent (Milgram, 1965).

These conditions must not be blended into a single dramatic script. The 1963 learner did not deliver the later voice-feedback protests about a heart condition. The emotional distance, learner feedback and required physical action differed. A precise account names the condition before giving a result.

The variations support a situational reading in a limited sense. Completion changed when features of the encounter changed. They do not isolate one pure cause, because a new condition can alter several meanings at once. Bringing the learner closer may affect empathy, attention, credibility and the ease of avoiding signs of distress.

On a small screen, swipe across the table to compare the procedure, endpoint and result.

Study or condition Sample or decision point Endpoint Reported result Evidence boundary
Milgram remote condition (1963) 40 men; learner pounded at 300 and 315 volts, then became silent Final switch marked 450 volts 26 of 40, or 65%, reached the endpoint One Yale condition, not a universal rate
Milgram voice-feedback condition (1965) Learner's verbal protests were audible Final switch marked 450 volts 62.5% reached the endpoint A later procedure, not the 1963 published script
Milgram proximity condition (1965) Learner was in the same room Final switch marked 450 volts 40% reached the endpoint Changed proximity and the social meaning of the encounter
Milgram touch-proximity condition (1965) Teacher had to force the learner's hand onto a plate after refusal Final switch marked 450 volts 30% reached the endpoint Added direct physical contact
Burger partial replication (2009) Stopped after the participant pressed 150 volts and then read the next item 150-volt decision point 70% passed the decision point in the base condition No test of continuation to 450 volts
Doliński et al. abbreviated procedure (2017) 80 Polish participants and ten buttons Tenth button, corresponding to 150 volts 72 of 80 pressed the final button Not a 450-volt replication
Slater et al. virtual analogue (2006) Participants knew the learner was virtual Immersive simulated task Behavioural and physiological responses were recorded Cannot estimate real-harm or 450-volt completion

Replications and Modern Analogues

Later researchers have had to balance comparison with participant protection. As a result, modern studies often alter the task or stop early. Their findings are informative within those endpoints, but they are not exact repeats of the 1963 procedure.

Burger's partial replication used extensive screening, clinical monitoring and a lower stopping point. In the base condition, 70 per cent of participants pressed the 150-volt switch and then read the next test item. The session then stopped. Burger (2009) therefore studied behaviour at a critical decision point, not completion of a 450-volt sequence.

Doliński and colleagues used an abbreviated Polish procedure with ten buttons. Seventy-two of 80 participants pressed the tenth and final button, which corresponded to 150 volts (Doliński et al., 2017). This is sometimes reported as 90 per cent reaching Milgram's maximum. That statement is wrong because the endpoint was 150 volts, not 450 volts.

Virtual analogues remove the possibility of harming a real learner while retaining some features of the encounter. Slater and colleagues found behavioural and physiological responses even though participants knew the learner was virtual (Slater et al., 2006). Such responses show that a simulation can be involving. They do not reveal how participants would act towards a real person.

Study note separating Milgram's 1963 condition, later variations, modern partial replications, interpretations, ethics and evidence limits
The Milgram experiment: evidence and boundaries. Open the full-size PNG for close reading.
Read the Milgram experiment study-note transcript

Core finding: In Milgram's 1963 published Yale condition, 26 of 40 men continued to the final switch marked 450 volts. The learner gave no vocal protest, pounded at 300 and 315 volts, then became silent. The 65 per cent result belongs to this one condition.

Method boundary: The participant was the teacher, the learner was a confederate and the shock generator was a prop. The participant received a sample shock, but the learner received none. The study observed decisions made under an experimenter's pressure, not a stable obedience trait.

Variations: The remote, voice-feedback, proximity and touch-proximity procedures were distinct. Completion fell from 65 per cent in the remote condition to 62.5, 40 and 30 per cent as the learner became more immediate.

Modern evidence: Burger stopped after the 150-volt decision point. Doliński and colleagues used ten buttons ending at 150 volts. Virtual work removed a real learner. None of these designs measures the proportion that would reach the original 450-volt endpoint.

Interpretations: Milgram proposed agentic state. Later accounts examine identification with science, participant belief, experimenter rhetoric, interactional pressure and resistance. These explanations remain evaluated interpretations.

Ethics and archival limits: Baumrind criticised deception, distress and possible damage to trust. Archival evidence raises questions about belief, procedural flexibility and reporting. It does not show that the entire programme was fabricated. Formal research review developed through a wider institutional and legal history.

Why Did Participants Continue?

Milgram proposed that people can enter an “agentic state”, in which they see themselves as carrying out another person's wishes rather than acting as autonomous authors. In this account, the authority takes responsibility while the participant attends to performing the assigned role. The idea is historically important, but it is not an established single mechanism.

The gradual structure may also have mattered. Each requested step was only slightly more severe than the one before it. Commitment to the research, uncertainty about responsibility and the difficulty of breaking an interaction could all influence conduct. The original design did not cleanly separate these possibilities.

Haslam and Reicher (2012) offer an engaged-followership account. They argue that participants may continue when they identify with the scientific project and accept the experimenter's purpose. Their account connects the obedience studies with social identity theory. It is a competing interpretation, not a replacement established beyond dispute.

Evidence about the experimenter's language supports attention to identification and rhetoric. Haslam and colleagues found that appeals linked to scientific purpose were more effective than a blunt order to continue (Haslam et al., 2014). This weakens the image of obedience as an automatic response to command alone, but it does not settle every participant's reason.

Gibson's interactional analysis shows persuasion, negotiation and resistance within the sessions (Gibson, 2013). Participants did not simply receive four prods and respond in the same way. The experimenter's attempts to secure continuation formed part of an unfolding social exchange.

Participant belief creates another interpretive problem. Someone who doubted the shocks could continue for a different reason from someone who believed the learner faced serious harm. Yet belief was not all-or-nothing, and retrospective reports can be uncertain. Neither full belief nor complete disbelief can simply be assumed for the whole sample.

The related theory of cognitive dissonance may help frame tensions between prior actions and later choices, but it is not a finding established by Milgram's design. Likewise, the just-world hypothesis addresses beliefs about deservingness and justice, not obedience as such.

Ethics and Research Governance

The central ethical issues include deception, distress, the participant's practical freedom to withdraw and the effect of the experience after the session. Participants did not know the true aim, the learner's role or the fact that no shocks were delivered. The experimenter's prompts also made withdrawal more socially difficult than a simple consent statement can imply.

Baumrind (1964) argued that the procedure risked harm and could damage trust in psychological research. Her critique remains important because it asks not only whether participants recovered, but whether a researcher may place them in such a conflict without valid advance consent.

Milgram defended the work by reporting debriefing, follow-up and participants' later views. Those steps form part of the historical record. They do not erase the ethical question, because retrospective approval cannot provide consent before exposure to distress.

The study also did not single-handedly create modern ethics committees. Formal review grew through a broader institutional and legal history, including US Public Health Service policy in 1966 and the National Research Act in 1974. The Nuremberg Code and Belmont Report supplied influential principles, but neither document by itself mandated independent review of every human study.

Moral reasoning provides another neighbouring evidence base. Kohlberg's stages of moral development concern reasoning about moral problems. Milgram's switch-pressing outcome cannot be used to assign a participant to a moral stage.

Criticisms and Archival Reassessment

The first limit is the sample. The 1963 test had 40 men from one area who chose to join a Yale study. It cannot stand for all genders, cultures, places or times. Later work widens the record, but it also changes the task.

There are two views of how real the lab felt. The machine, Yale setting and experimenter could make it seem true. Yet some people could also doubt that Yale would permit grave harm. A strong setting and doubt can exist at the same time.

The archive raises doubts about the use of the stated script. It also asks what people believed and how Milgram chose results to report. Belief may be linked with reaching the end, which makes a simple engaged-followership account harder to sustain (Perry et al., 2020). This evidence sets limits, but does not prove that the whole programme was false.

The full record has many tests with different evidential weight. Short accounts can make them look like one clean experiment. A better review names each sample, task, endpoint and source.

An account can also become circular. It may name going on as obedience, then use this act to prove that an obedient state caused it. That adds little. A useful test must tell apart blame, group bonds, belief, words, small steps and other proposed causes.

A review across the named tests links end rates with the pull of the experimenter and learner (Haslam, Loughnan and Perry, 2014). This helps show patterns across tests. Yet it looks back at old data. It does not make the first studies a direct test of engaged followership.

Links to other famous ideas can also blur the topic. Social comparison theory asks how people judge themselves against others. It does not explain Milgram merely because the task took place with other people.

What the Experiment Can and Cannot Show

The strongest conclusion is bounded. In carefully constructed laboratory settings, substantial numbers of participants continued an apparently harmful task under pressure from an experimenter. Changes to learner proximity altered completion rates, and many participants showed conflict while continuing.

The studies support attention to situations, institutions and social interaction. They challenge the claim that harmful compliance can be explained only by unusual personalities. They do not show that personality never matters, that authority always succeeds or that obedience is a fixed proportion of human nature.

The work cannot determine how one named person will behave. It cannot translate directly from a Yale laboratory to a school, hospital, organisation or political regime. Each setting contains different relationships, sanctions, identities, histories and opportunities for resistance.

For education, the responsible use is source literacy. Readers and scholars can compare the 1963 report with later variations, ethical critique and archival reassessment. Teachers should not recreate the procedure, simulate harmful coercion or use the findings to judge a learner's likely compliance.

Nor does the study supply a behaviour-management method. Research on modelling and self-efficacy belongs to social cognitive theory, while evidence about teaching quality has its own owners. Leaders considering research-informed teaching should treat Milgram as a case in evaluating claims, methods and ethics, not as an intervention.

Frequently Asked Questions

These short answers tie each result to its sample and endpoint. They also separate the 1963 test from later and modern forms. One choice in one setting cannot prove a general rate, a fixed trait or one settled cause.

What was the main finding of the Milgram experiment?

In the Yale test published in 1963, 26 of 40 men used the last switch marked 450 volts. Many showed strain. The result comes from one test. It does not mean that 65 per cent of all people will follow any harmful order.

Did participants actually shock the learner?

No. The learner was a man working with the research team and had no electric shocks. The machine and staged signs misled the participants. The ethics issue rests in part on what they thought they were doing and the strain caused by that belief.

What was the aim of Milgram's 1963 study?

The aim was to see how far adults would follow an authority's words when the task seemed to harm someone else. It studied conduct in a controlled research setting. It did not test obedience in every social or past setting.

What were the main Milgram variations?

The named tests changed feedback and the distance between teacher and learner. End rates were 65 per cent in the remote test, 62.5 per cent with voice feedback, 40 per cent with proximity and 30 per cent with touch proximity (Milgram, 1965). These tests should not be merged.

Was the Milgram experiment ethical?

It remains a major case in research ethics. Critics focus on deceit, strain, limits on a free choice to leave and trust. Debriefing and later checks matter, but they cannot give consent in advance. They also do not settle whether the first risk was fair.

Has the Milgram experiment been replicated?

Later teams have used partial, short and virtual forms. Burger stopped after a choice at 150 volts. Doliński and colleagues used ten buttons that ended at 150 volts. These studies cannot show how many people would reach the old 450-volt endpoint.

Why did participants obey Milgram?

There is no one settled answer. Proposed causes include agentic state, small step changes, a bond with science, belief, the experimenter's words and social pressure. The record also includes refusal and strain.

Further Reading

These papers link the first report to the ethics debate, later tests and current readings. Each title opens its DOI record. Reading them together keeps each claim tied to the right sample, task and endpoint.

References

These references cover all works cited here. They include the first reports, the early ethics critique, later tests and new readings of the archive. Their DOI links identify each paper and the study behind each claim.

  1. Baumrind, D. (1964). Some thoughts on ethics of research: After reading Milgram's “Behavioral Study of Obedience”. American Psychologist, 19, 421-423.
  2. Blass, T. (1999). The Milgram paradigm after 35 years: Some things we now know about obedience to authority. Journal of Applied Social Psychology, 29, 955-978.
  3. Burger, J. M. (2009). Replicating Milgram: Would people still obey today? American Psychologist, 64, 1-11.
  4. Doliński, D., Grzyb, T., Folwarczny, M., Grzybała, P., Krzyszycha, K., Martynowska, K., and Trojanowski, J. (2017). Would you deliver an electric shock in 2015? Obedience in the experimental paradigm developed by Stanley Milgram in the 50 years following the original studies. Social Psychological and Personality Science, 8, 927-933.
  5. Gibson, S. (2013). Milgram's obedience experiments: A rhetorical analysis. British Journal of Social Psychology, 52, 290-309.
  6. Haslam, S. A., Loughnan, S., and Perry, G. (2014). Meta-Milgram: An empirical synthesis of the obedience experiments. PLOS ONE, 9, e93927.
  7. Haslam, S. A., and Reicher, S. D. (2012). Contesting the “nature” of conformity: What Milgram and Zimbardo's studies really show. PLOS Biology, 10, e1001426.
  8. Haslam, S. A., Reicher, S. D., and Birney, M. E. (2014). Nothing by mere authority: Evidence that in an experimental analogue of the Milgram paradigm participants are motivated not by orders but by appeals to science. Journal of Social Issues, 70.
  9. Milgram, S. (1963). Behavioral study of obedience. Journal of Abnormal and Social Psychology, 67, 371-378.
  10. Milgram, S. (1965). Some conditions of obedience and disobedience to authority. Human Relations, 18, 57-76.
  11. Perry, G., Brannigan, A., Wanner, R. A., and Stam, H. J. (2020). Credibility and incredulity in Milgram's obedience experiments: A reanalysis of an unpublished test. Social Psychology Quarterly, 83, 92-109.
  12. Slater, M., Antley, A., Davison, A., Swapp, D., Guger, C., Barker, C., Pistrang, N., and Sanchez-Vives, M. V. (2006). A virtual reprise of the Stanley Milgram obedience experiments. PLOS ONE, 1, e39.
Paul Main, Founder of Structural Learning
About the Author
Paul Main
Founder & Metacognition Researcher

Paul Main is an educator and metacognition researcher who founded Structural Learning in 2002. With a psychology degree from the University of Sunderland and 22+ years helping schools embed thinking skills, he bridges the gap between educational research and classroom practice. Fellow of the RSA and Chartered College of Teaching, with 128+ Google Scholar citations.

More →

Learning Theories

Back to Blog