Learning, quality, and reliability
Nancy Leveson and Clark Turner's reconstruction documents six known Therac-25 overdose accidents from 1985 to 1987, with serious injuries and deaths. Reports, interface evidence, and corrective action traveled unevenly across patients, hospitals, the manufacturer, users, and regulators; later quality, learning, reliability, and software practices are comparisons rather than one documented lineage.
Governing questionWhat lets a disturbing local event become a system defect before the same design harms someone elsewhere?
Six known overdoses bound the case
Nancy Leveson and Clark Turner report six known Therac-25 accidents between June 1985 and January 1987 in which patients received massive radiation overdoses, with resulting serious injuries and deaths. Their 1993 IEEE Computer article reconstructs the events from public legal, hospital, manufacturer, user, and government material and says that the authors sought corroboration for important facts. It also says that detailed evidence about the manufacturer's software development, management, and quality control was unavailable and that some conclusions had to be inferred.1
“Six known” is therefore a boundary on the surviving reconstruction, not proof that every overdose was detected, reported, preserved, or counted. The record supports six documented accidents and their consequences; it does not support a complete incidence rate, a single fatality count with uniform causal adjudication, or a claim that unreported events did or did not occur.1
Atomic Energy of Canada Limited designed the Therac-25 so that software carried more safety responsibility than in the Therac-20 and did not duplicate all of the earlier machine's hardware safety mechanisms. The 1983 safety analysis described by Leveson and Turner excluded residual software errors and assigned computer-failure probabilities without stated justification. The authors treat software flaws, hardware architecture, interface design, incident follow-up, risk assessment, engineering practice, management, and regulation as interacting contributors rather than naming software or an operator as the sole cause.1
Local observations did not become prompt cross-site containment
The first reconstructed accident at Kennestone involved severe radiation burns, lasting disability, and eventual removal of the patient's breast. The article reports uncertainty about when AECL learned of the event, no prompt investigation, delayed FDA reporting, and other Therac-25 users remaining unaware until later accidents. At Hamilton, AECL could not reproduce the suspected failure, changed a microswitch design, and did not then install the independent turntable-position interlock recommended by Canadian officials and an independent consultant.2
After a later Yakima patient developed an unusual injury, hospital staff wrote to AECL. AECL replied in February 1986 that neither machine malfunction nor operator error could have caused the damage and said there had apparently been no similar instance. The hospital initially recorded the cause as unknown; only after another Yakima overdose did it revisit the earlier event. The reconstructed sequence supports a failure to join observations across sites, while leaving open exactly which incidents AECL and each regulator knew at each point.2
That distinction matters. Patients reported heat, pain, or electric-shock-like sensations; operators reported abnormal behavior; physicists investigated; hospitals contacted the manufacturer; the manufacturer tested and replied; and regulators eventually acted. The evidence does not support a story in which nobody spoke. It supports uneven reporting, incomplete propagation, initial nonreproduction, misleading interface information, and confidence claims that made each local report harder to interpret as a common design hazard.234
Operator expertise made Malfunction 54 reproducible
At the East Texas Cancer Center in March 1986, an experienced operator quickly corrected an X-ray-mode entry to electron mode. The console displayed “Malfunction 54,” described the condition as a low-priority treatment pause, and showed an apparent underdose; the local documentation did not explain that the message could accompany a dangerous overdose. AECL engineers tested the machine after the injury but did not reproduce the condition.3
Three weeks later, the same operator saw the same message during another treatment and immediately involved the hospital physicist, Fritz Hager. Working from her remembered sequence, they found that rapid editing was necessary to reproduce the failure. AECL could reproduce it after Hager explained the timing condition. The article's software analysis identifies a race condition in the shared state governing edited treatment data.3
The first Tyler patient died from complications of the overdose five months after his treatment; the second died three weeks after his overdose. Dose figures in the record are reconstructions and simulations, and the authors warn that exact delivered doses cannot be established because results varied across machines and simulations.3 The supported learning claim is not that operator speed caused the catastrophe. Operator familiarity exposed a timing-dependent unsafe state that the interface obscured and the safety architecture failed to contain.
Corrective action eventually changed the safety architecture
After the Tyler events, FDA correspondence reproduced in the article criticized AECL's first user notice for failing to describe the defect or hazard, required a corrective action plan, and requested detailed software procedures, documentation, testing, and interlock information. Therac users formed a group and complained about incomplete information propagation. Later meetings brought users, AECL, U.S. and Canadian regulators, and technical representatives together to review all six known accidents and proposed modifications.4
In February 1987, U.S. regulators concluded that software alone could not assure safe operation, required additional hardware interlocking, and recommended that routine patient treatment stop until an amended plan was completed. The final 1987 plan described in the reconstruction included independent hardware and software single-pulse shutdown, turntable monitoring and interlocks, restrictions on editing and restart behavior, meaningful malfunction messages, manual changes, and fixes for the known Tyler and Yakima faults.4
Those changes support a bounded inference: the response moved beyond repairing one input sequence and changed independent containment, operator information, and the premise that software correctness could carry the safety case alone. They do not prove that every remaining hazard was eliminated, that the response was timely, or that the final system produced a measured long-run safety effect.4
Quality traditions pose different questions
The NIST/SEMATECH engineering handbook describes guild apprenticeship and demonstrated mastery as historical quality practices, then dates effective statistical application to quality control to the 1920s. It identifies Walter Shewhart's 1924 control-chart memorandum, his 1931 Economic Control of Quality of Manufactured Product, and sampling work by Harold Dodge and Harry Romig. Its process-control section separates monitoring against control limits from acceptance sampling of finished lots.5
A 1956 National Bureau of Standards report says World War II military procurement expanded sampling inspection, intensive statistical-quality-control training, and wartime statistical research. That contemporaneous retrospective supports the bounded contribution of I11-N03; it does not show uniform adoption, one standard operating model, or the effect of those methods across the entire mobilization.6
W. Edwards Deming presents Out of the Crisis as a transformation of management, while the Juran Institute describes Joseph Juran's quality trilogy as planning, control, and improvement. The publisher and institute records are authoritative statements of those frameworks, not comparative trials or evidence that either author investigated Therac-25.7
Toyota describes jidoka as stopping a machine automatically or allowing an operator to stop the line when an abnormality appears, then detecting the abnormality clearly and preventing recurrence. That is Toyota's own current account of its production system, not independent evidence about who can safely exercise stop authority in practice or proof that an assembly-line mechanism transfers to radiation therapy.8
Statistical process control, management responsibility, quality planning, and stop-on-abnormality can each illuminate a different part of the Therac record. None decides how many catastrophic injuries form an acceptable sample, replaces fail-safe architecture, or protects patients at another site unless information and containment cross hospital, manufacturer, and regulatory boundaries.
Learning, reliability, voice, and software remain comparisons
Chris Argyris defines organizational learning as detecting and correcting error and distinguishes correction within existing policies from double-loop learning that challenges underlying policies and objectives. His 1977 article is a primary conceptual statement; it is not a Therac investigation or a test of whether adding an interlock constitutes double-loop learning.9
Organizational Learning II is the associated Argyris and Donald Schön work route. The comparison asks whether a correction changes only one faulty sequence or also the assumption that software can stand alone as a catastrophic-harm barrier. It does not assert that AECL, users, or regulators applied that theory.
Karl Weick and Kathleen Sutcliffe organize Managing the Unexpected around preoccupation with failure, reluctance to simplify, sensitivity to operations, commitment to resilience, and deference to expertise. Amy C. Edmondson studied 51 teams in one manufacturing firm and found team psychological safety associated with learning behavior. The book record states a framework; the field study reports a bounded association. They do not form one proven mechanism, establish a causal effect across settings, or analyze Therac-25.10
Google's site-reliability chapters describe its use of written, reviewed, widely shared postmortems and an outage tracker that groups and analyzes alerts across services. The authors also note that postmortems centered on large events can miss small but frequent problems and cross-service opportunities. These are useful primary practitioner records for I11-N08, but they are not independent evaluations, do not support every practice labeled “modern software delivery,” and do not establish transfer to medical-device safety.11
A blameless review can protect candor without erasing differentiated accountability. The relevant test is whether the review changes unsafe design, authority, resources, or incentives while protecting patients and reporters—not whether it avoids naming every accountable decision.
A joined learning record is a design proposal
The Therac reconstruction suggests fields that a safety learning record may need to join: expected safe behavior, patient outcome, machine and software version, operator-visible state, prior reports, investigation, decision authority, containment, corrective action, and later verification.14 Turning that suggestion into a product requirement remains a design choice, not an evaluated result of the case.
Joining records is insufficient if the manufacturer alone controls meaning, patients cannot inspect or contest the account, reporters can be punished, or a closure metric substitutes for verified removal of the repeated path to harm. A learning system needs minimization, access control, appeal, independent review, and bounded stop authority alongside traceability.
Four product hypotheses remain unvalidated
- I11-P01 proposes joining event, artifact, review, decision, and outcome in one trace. Evaluation must compare reconstruction quality, correction, repeat failures, surveillance burden, blame, and information overload against less integrated records.
- I11-P02 proposes recording expected results before action. Evaluation must test forecast calibration and later reasoning while checking gaming, risk-avoidance, punishment for honest error, and delay of necessary action.
- I11-P03 proposes classifying preventable deviation, complex-system failure, intelligent experiment, and governance breach. Evaluation must test reviewer agreement, response fit, recurrence, and whether labels excuse governance, blame frontline workers, or relabel foreseeable harm.
- I11-P04 proposes a bounded evidence-backed stop or escalation available to any coworker with protection for people who raise risk. Evaluation must track use, response, substantiation, false positives, remediation, power differences, and retaliation.
The Therac sources motivate these questions but validate none of the four mechanisms. Prospective comparison, abuse review, affected-party evidence, and evidence of later outcomes are required before promotion.
Eight nodes and seven comparative edges bound the map
The eight nodes are not a continuous historical ladder. I11-N01 groups apprenticeship and craft inspection; I11-N02 marks statistical quality control; I11-N03 marks World War II training and operations research; I11-N04 groups Deming and Juran; I11-N05 marks Toyota; I11-N06 groups Argyris and Schön; I11-N07 groups Weick, high-reliability research, and Edmondson; and I11-N08 marks modern software delivery.
The NIST historical overview supports the broad I11-N01 and I11-N02 labels but does not establish that the Roman military, Venetian Arsenal, and British Royal Navy transmitted one method to Shewhart.5 The National Bureau of Standards report bounds I11-N03.6 Framework records bound I11-N04 through I11-N08 within the self-description, setting, and transfer limits stated above.7891011
The edge records expose rather than conceal those limits:
- I11-E01, grade C and comparative, uses source IDs
roman-military,venetian-arsenal-and-republic,british-royal-navy, andeconomic-control-of-quality-of-manufactured-product. The paths establish separate craft and statistical endpoints, not continuous descent. - I11-E02, grade C and comparative, uses source IDs
economic-control-of-quality-of-manufactured-productandallied-and-us-world-war-ii-mobilization. External history supports wartime expansion, not its uniformity or effect across mobilization. - I11-E03, grade C and comparative, uses source IDs
allied-and-us-world-war-ii-mobilization,w-edwards-deming,out-of-the-crisis, andjuran-on-planning-for-quality. It places historical scaling beside later management frameworks without asserting a single transfer path; Juran on Planning for Quality is the local Juran work route. - I11-E04, grade C and comparative, uses source IDs
out-of-the-crisis,juran-on-planning-for-quality,toyota, andjapanese-quality-control-circles. Toyota's self-description supports its stated endpoint, not independent implementation effects. - I11-E05, grade C and comparative, uses source IDs
toyota,chris-argyris, andorganizational-learning-ii. It compares operating response with a conceptual distinction and makes no influence claim. - I11-E06, grade C and comparative, uses source IDs
organizational-learning-ii,karl-e-weick,managing-the-unexpected, andamy-c-edmondson. These sources differ in unit, method, and evidence base and do not validate a combined mechanism. - I11-E07, grade C and comparative, uses source IDs
managing-the-unexpected,amy-c-edmondson, andgoogle-alphabet. Google's practitioner record establishes reported software practices, not theoretical descent, independent effects, or medical-device transfer; Google / Alphabet is the corporate source path.
Evidence grades describe support for each bounded relation, not importance, generalizability, moral worth, or effect size. These relations are editorial comparisons, not evidence of causal influence or institutional descent.
Seven related records are routes, not prerequisites
- Economic Control of Quality of Manufactured Product supplies the Shewhart work route.
- Out of the Crisis supplies the Deming work route.
- Organizational Learning II supplies the Argyris and Schön work route.
- Managing the Unexpected supplies the Weick and Sutcliffe work route.
- Amy C. Edmondson supplies the psychological-safety research route.
- organizational ignorance supplies an adjacent route for missing, fragmented, or institutionally produced unknowns.
- benefit for all life supplies the normative route for asking who or what is protected by reliability.
No structured reading dependency is recorded. These routes can be read in any order, and relatedness does not establish agreement, historical influence, shared findings, or equivalent moral stakes.
Established institutions are comparison paths, not one lineage
Military, industrial, production, and technical comparison paths:
- Allied and U.S. World War II mobilization, Toyota, NASA Apollo program, U.S. Navy nuclear propulsion and submarine command, 3M, Boeing, Intel, and Zeiss.
Health, education, public-service, infrastructure, and crisis comparison paths:
- Aravind Eye Care System, Bali's Subak irrigation governance, BRAC, Brazil's Family Health Strategy in SUS, Cuban National Literacy Campaign, Eskom under state capture, Ethiopian Airlines, Fiji Locally Managed Marine Area Network, Ghana's Community-based Health Planning and Services, and Haiti's 2010 humanitarian cluster response.
Community, ecosystem, learning-network, and accountability comparison paths:
- ITRI–TSMC semiconductor ecosystem, Japanese quality-control circles, Karuk cultural fire institutions, Te Kōhanga Reo, MINUSTAH and Haiti's cholera response, Niger's farmer-managed natural regeneration, Nigeria's Ebola Emergency Operations Center, South Africa's Treatment Action Campaign, Ushahidi, and Vale and the Brumadinho dam disaster.
Their local records carry their own evidence. Grouping them here creates questions about inspection, feedback, stop authority, recurrence, adaptation, public accountability, and who bears failure. It does not assert a shared practice, theory adoption, performance level, moral status, or descent from the Therac response.
The affected-party path is explicit and impact remains unclassified
The documented path runs from patients and families through operators, physicists, clinicians, hospitals, AECL engineers and managers, other Therac users, and U.S. and Canadian regulators. These positions had different exposure, information, technical access, authority, and ability to contain use. Treating one position as a proxy for all of them would erase the distribution of both knowledge and harm.234
No structured impact record is assigned. The two Tyler patients died from the documented overdoses, and the six known accidents included severe burns, necrosis, pain, disability, major surgery, paralysis, and death.23 Those consequences are known rather than neutral. A single net impact classification remains unassigned because the reviewed sources do not produce a commensurable account across patients, families, care workers, hospital staff, manufacturer workers, regulators, later patients, communities, animals, ecosystems, or future generations.
The reviewed record does not measure material or ecological effects of the machines, their repair, or their replacement. Missing ecological evidence is not evidence of ecological safety or harm. The absence of a structured impact record means impact is unclassified, not absent.
Evidence still needed
- Contemporaneous hospital, AECL, Canadian, and FDA files in a documented chain of custody, including original incident reports, correspondence, corrective action revisions, meeting minutes, tests, and final verification records.
- A patient-by-patient clinical and legal record that distinguishes injury, later death, stated cause of death, settlement, and uncertainty without converting people into a single casualty statistic.
- A cross-site chronology of who received which report, when, through what channel, with what authority and response; the reconstruction often reports conflicting accounts or incomplete knowledge.
- Original software source, specifications, version history, test plans, hazard analysis, audit trails, hardware-change records, and independent validation of the final safety architecture.
- Accounts from patients, families, operators, physicists, nurses, clinicians, service technicians, engineers, managers, user-group participants, and regulators rather than treating any one institutional voice as complete.
- Independent studies separating statistical control, wartime scaling, Deming and Juran, Toyota practice, double-loop learning, high reliability, psychological safety, and software incident learning instead of treating them as one cumulative method.
- Prospective tests of I11-P01 through I11-P04, including repeat-failure rates, false reassurance, reporting and retaliation, privacy, appeal, decision quality, and whether affected people can stop exposure.
- Human, labor, distributional, material, and ecological evidence sufficient to classify impacts across the full affected-party path.
Paths into deeper study
- Follow the six accidents through the organizational ignorance route and distinguish missing observation, failed propagation, disbelief, and absent authority to contain.
- Compare Economic Control of Quality of Manufactured Product, Out of the Crisis, Organizational Learning II, and Managing the Unexpected as different questions rather than stages of one method.
- Continue through Amy C. Edmondson for the bounded team-level evidence on interpersonal risk and learning behavior.
- Use benefit for all life to ask whose safety defines reliability, who may authorize exposure, and who can demand repair.
Additional reciprocal institution comparisons
Newly developed institutional records add these reciprocal comparison paths:
- Linux ecosystem — institution-comparison
- Singapore — institution-comparison
- Steward Health Care — institution-comparison
- W3C — institution-comparison
- Wikipedia — institution-comparison
Each path identifies a sourced case where this idea is a defining emphasis. The relation is editorial comparison, not evidence of direct influence, shared terminology, or equivalent outcomes.
Source notes
Nancy G. Leveson and Clark S. Turner, “An Investigation of the Therac-25 Accidents,” IEEE Computer 26, no. 7 (1993), pp. 18–41; bibliographic record and abstract, University of California eScholarship, and opening, “Genesis of the Therac-25,” safety-analysis assumptions, and “Accident history,” MIT reprint, part I. This is independent scholarly retrospective reconstruction based on public primary and legal material, with corroboration where possible. The authors explicitly report unavailable manufacturer development and management evidence, possible documentary error, and necessary inference; it cannot establish undetected events or every actor's knowledge.
↩ ↩ ↩ ↩Leveson and Turner, “Accident history,” Kennestone and Hamilton sections in MIT reprint, part I, and first Yakima section in MIT reprint, part II. The article reconstructs patient outcomes, reporting, nonreproduction, manufacturer replies, and proposed interlocks from later-accessible records. It is not a contemporaneous unified incident log, and it preserves disputes over notification and causation rather than resolving them by assumption.
↩ ↩ ↩ ↩ ↩Leveson and Turner, East Texas Cancer Center sections and dose limitation in MIT reprint, part II, and opening race-condition analysis in MIT reprint, part III. This is the strongest reviewed reconstruction of the operator and physicist's reproduction sequence and the two Tyler deaths. Much of the operational account derives from deposition, participant, simulation, and later technical evidence; exact delivered doses are not recoverable.
↩ ↩ ↩ ↩ ↩ ↩Leveson and Turner, FDA corrective-action correspondence, user-group information complaints, and plan review in MIT reprint, part III, and February shutdown history, multi-party review, final corrective-action features, and “Lessons learned” in MIT reprint, part IV. The scholarly article reproduces and summarizes primary regulatory and participant material; the original FDA files were not independently reviewed here, and completion of a plan does not itself establish long-run effect.
↩ ↩ ↩ ↩ ↩ ↩NIST/SEMATECH, “How Did Statistical Quality Control Begin?” sections on guild apprenticeship, twentieth-century statistics, Shewhart, Dodge, and Romig, National Institute of Standards and Technology, and “What Are Process Control Techniques?” sections on process monitoring and acceptance sampling, NIST. This official technical handbook is authoritative for the described methods and supplies a concise historical overview. It is not a primary craft archive, an institutional genealogy, or evidence that statistical methods are sufficient for rare catastrophic hazards.
↩ ↩National Bureau of Standards, Manual on Experimental Statistics for Ordnance Engineers: Progress Report for the Period Ending 30 June 1956, NBS Report 4817 (1956), introduction, pp. 1–2, on military sampling, intensive wartime courses, and the Statistical Research Group, National Bureau of Standards. This government technical report is a near-contemporaneous primary institutional retrospective. Its broad introduction does not measure adoption, effectiveness, coercive procurement effects, or worker experience across wartime production.
↩ ↩W. Edwards Deming, Out of the Crisis, publisher description of the 14 Points and transformation-of-management argument, MIT Press. Maureen Goldman, “The Juran Trilogy,” sections defining quality planning, control, and improvement, Juran Institute. These are authoritative destinations for the authors' stated frameworks. The publisher description is not the full Deming argument, and the Juran Institute is a successor practice organization; neither is an independent outcome evaluation, a Therac source, or proof of one shared framework.
↩ ↩Toyota Motor Corporation, “Toyota Production System,” “The Two Pillars of TPS” and “Jidoka,” Toyota. This is an authoritative corporate self-description of stopping on an abnormality, operator stop cords, building quality into a process, and recurrence prevention. It is not independent observation of daily practice, retaliation risk, comparative effect, or transfer to safety-critical care.
↩ ↩Chris Argyris, “Double Loop Learning in Organizations,” Harvard Business Review, September 1977, canonical article destination, Harvard Business Review, and official product description defining organizational, single-loop, and double-loop learning, Harvard Business Publishing. This is a primary conceptual statement of detecting and correcting error, single-loop correction, and challenge to underlying policy and objectives. It does not study Therac-25, validate the analogy, or by itself cover the later Argyris and Schön treatment of defensive routines.
↩ ↩Karl E. Weick and Kathleen M. Sutcliffe, Managing the Unexpected, 3rd ed., publisher description and chapter contents listing the five principles, Wiley. Amy C. Edmondson, “Psychological Safety and Learning Behavior in Work Teams,” Administrative Science Quarterly 44, no. 2 (1999), pp. 350–383, abstract and study description, SAGE. The Wiley record is an authoritative conceptual destination, not an independent test of all high-reliability claims. Edmondson's original multimethod field study covers 51 teams in one manufacturing firm and reports association; it does not establish universal causality, cross-site reporting, or Therac interpretation.
↩ ↩John Lunney and Sue Lueder, “Postmortem Culture: Learning from Failure,” sections defining postmortems, review, sharing, and aggregate analysis, Google Site Reliability Engineering. Gabe Krabbe, “Tracking Outages,” sections on grouping alerts and the limits of large-incident postmortems, Google Site Reliability Engineering. These are primary corporate practitioner accounts of Google's reported practices. They provide no independent causal estimate, do not establish all software organizations' practice, and cannot substitute for medical-device hazard analysis, regulation, or independent barriers.
↩ ↩
Research record
Evidence basis
Claim Cited. Material claims carry source locators; comparative interpretation may still evolve.
Open questions and affected lives
Benefit-to-life status: Seed
- Who is allowed to name an error, stop work, or challenge the rule that produced a failure?
- Who bears the human, community, and ecological costs of an experiment or reliability failure?
- How can learning records support repair without becoming instruments of blame or surveillance?
- Reliability for whose purpose, and at what cost to adaptability, dignity, and voice?
These questions remain open; absence from the record does not imply absence of benefit or harm.
Structured atlas record
Lineage nodes
- Apprenticeship and craft inspectionembed feedback in direct observation and mastery
- Statistical quality controluses variation and sampling to control processes
- World War II training and operations researchscales statistical quality-control training, acceptance sampling, and analytic work for wartime production
- Deming and Juranmake quality a management responsibility and improvement a system discipline
- Toyotabuilds quality at source, stop-the-line authority, standard work, and kaizen into operations
- Argyris and Schöndistinguish correcting action from revising governing variables
- Weick, high-reliability research, and Edmondsonconnect weak signals, mindful organizing, voice, and learning behavior
- Modern software deliveryadds shared postmortems, monitoring records, and cross-service incident aggregation to operational learning
Typed relationships
CApprenticeship and craft inspection → Statistical quality control
Enabling Technology: places craft inspection and demonstrated mastery beside statistical process and sampling methods as different ways to detect unacceptable work
The source paths establish separate endpoints; they do not establish one continuous line of descent from the named craft institutions to Shewhart's methods.CStatistical quality control → World War II training and operations research
Enabling Technology: places statistical quality-control methods beside their wartime expansion through military acceptance sampling, training, and analytic programs
The external historical record supports wartime use and expansion, while the local source paths establish the two endpoints; the edge does not quantify adoption or effectiveness across the mobilization.CWorld War II training and operations research → Deming and Juran
Managerial Generalization: places wartime statistical scaling beside Deming's management transformation and Juran's planning-control-improvement scheme
The source paths and external framework records establish distinct historical and conceptual endpoints, not a documented single transfer from wartime programs to both authors.CDeming and Juran → Toyota
Operationalization: places Deming's and Juran's management frameworks beside Toyota's stated jidoka, stop-on-abnormality, and recurrence-prevention practices
Toyota's corporate account establishes its stated practices but not independent implementation effects; no direct influence claim is made beyond the separate local source paths.CToyota → Argyris and Schön
Theoretical Codification: places Toyota's operating response to abnormalities beside Argyris and Schön's distinction between correcting action and revising governing variables
The comparison concerns different learning questions; the reviewed evidence does not establish Toyota as an empirical case or historical source for Argyris and Schön.CArgyris and Schön → Weick, high-reliability research, and Edmondson
Managerial Generalization: places double-loop learning beside high-reliability principles and team psychological safety as distinct accounts of noticing, questioning, and responding
The theories have different units and evidence bases; the edge neither combines them into one validated mechanism nor establishes a line of influence.CWeick, high-reliability research, and Edmondson → Modern software delivery
Operationalization: places reliability and voice concepts beside Google's stated use of shared postmortems, monitoring records, and incident aggregation
Google's practitioner chapters establish its reported software practices, not descent from the cited theories, independent causal effects, or transfer to medical-device safety.Provenance and sources
Online anchors
- https://escholarship.org/uc/item/5dr206s3
- https://web.mit.edu/6.033/2004/wwwdocs/papers/Therac_1.html
- https://web.mit.edu/6.033/2004/wwwdocs/papers/Therac_2.html
- https://web.mit.edu/6.033/2004/wwwdocs/papers/Therac_3.html
- https://web.mit.edu/6.033/2004/wwwdocs/papers/Therac_4.html
- https://www.itl.nist.gov/div898/handbook/pmc/section1/pmc11.htm
- https://www.itl.nist.gov/div898/handbook/pmc/section1/pmc12.htm
- https://nvlpubs.nist.gov/nistpubs/Legacy/RPT/nbsreport4817.pdf
- https://mitpress.mit.edu/9780262350037/out-of-the-crisis/
- https://www.juran.com/blog/the-juran-trilogy-/
- https://global.toyota/en/company/vision-and-philosophy/production-system/
- https://hbr.org/1977/09/double-loop-learning-in-organizations
- https://store.hbr.org/product/double-loop-learning-in-organizations/77502
- https://onlinelibrary.wiley.com/doi/book/10.1002/9781119175834
- https://journals.sagepub.com/doi/10.2307/2666999
- https://sre.google/sre-book/postmortem-culture/
- https://sre.google/sre-book/tracking-outages/