← Atlas
Work

Accelerate

Forsgren, Humble, and Kim built Accelerate from four years of voluntary, self-reported DevOps surveys, using statistical models to connect software-delivery measures with technical practice, culture, and organizational outcomes. The findings challenged a supposed tradeoff between speed and stability, but the study design establishes associations rather than a universal causal recipe, and its famous metrics have continued to change.

Working · Claim Cited

Most of the engineers' time was not going into new features

In 2008, Hewlett-Packard's LaserJet firmware organization began a transformation of a codebase with millions of lines and a twenty-five-year compatibility legacy. The participant account describes about 400 developers across four business units, four U.S. states, and three continents. It reports that only 5 percent of capacity initially went to new features.1

The organization invested in a new architecture, automated testing, and continuous-integration infrastructure. After four years, the same account put 40 percent of capacity into innovation and reported a 40 percent reduction in overall development cost, a 140 percent increase in programs under development, and a reduction in its measured cycle time from two months to one day.1 The executive, program manager, and architect who led the work wrote the case. It is unusually concrete participant evidence, not an independent comparison or controlled experiment.

The case supplies a plausible mechanism for the argument later made in Accelerate. When integration is rare and releases are large, each change accumulates more uncertainty and coordination. When tests, deployment, and feedback become routine, smaller changes can move without making every release a special event. That is a systems hypothesis, not a conclusion the HP case can establish by itself.

A survey program turned that proposition into measured associations

Nicole Forsgren, Jez Humble, and Gene Kim built Accelerate from four annual cross-sectional studies. The book reports more than 23,000 responses from more than 2,000 organizations, spanning startups, large enterprises, regulated industries, and nonprofit and government work.2

The first report makes the scale concrete. More than 9,200 people in 110 countries answered the December 2013 survey. Eighty-three percent identified as practitioners, 14 percent as managers, directors, or executives, and 3 percent as C-level executives.3 In later methodological disclosures, the researchers describe recruiting practitioners and leaders familiar with DevOps through email lists, social media, sponsors, and respondent referrals. This was snowball sampling from a defined target population, not a probability sample of all software organizations.4

Respondents estimated deployment frequency, change lead time, recovery time, and change failure and assessed practices, culture, job conditions, and organizational results. The researchers tested multi-question latent constructs, used hierarchical clustering to identify delivery profiles, and applied regression and partial-least-squares structural models to theory-based relationships.5

The program found that respondents who reported frequent deployment also tended to report short lead times and fast recovery. It associated stronger software delivery performance with respondents' ratings of organizational performance, and associated continuous delivery, loosely coupled architecture, a generative information culture, and lean product practice with delivery or organizational outcomes.6 These are observational, self-reported associations. The claim that particular capabilities produce the outcomes is a theory-guided interpretation, not a randomized or longitudinal estimate of their causal effect.5

Even the famous “four keys” were not one settled object. Change failure rate did not join the other three measures in the 2014 latent construct. DORA later renamed and narrowed recovery time to failed-deployment recovery time, and in 2024 added deployment rework rate, producing a five-metric model organized as throughput and instability.7 Accelerate records one influential stage of a continuing research program.

The metrics are outcomes, not a theory of the whole job

The book's four measures were deployment frequency, lead time from commit to production, time to restore service, and the percentage of changes requiring remediation.8 By definition, they begin around committed code and production delivery. They do not measure which work should be built, whether people need it, or what happens beyond delivery and immediate operational failure. A team can improve every delivery measure while producing the wrong thing efficiently.

The capability model tries to explain why the measures move. Its twenty-four capabilities include version control, deployment and test automation, continuous integration, trunk-based development, small batches, deployable architecture, customer feedback, work-in-process limits, learning, collaboration, and transformational leadership.9 These are managerial prescriptions informed by survey associations and practitioner experience; the research did not experimentally install each practice and observe the result. The model therefore belongs beside a broader inquiry into measurement, accounting, and control: the same measure can reveal a constraint or become a target that hides it.

Current DORA guidance warns against turning a deployment rate into a universal goal, comparing unlike applications, assigning isolated metrics to siloed teams, or making teams compete. It says those uses invite gaming, misleading comparisons, friction, and finger-pointing.10 A memorable framework does not make every use of it valid.

What the design can and cannot establish

The four annual samples are repeated cross-sections, not a panel following the same organizations through transformation. Recruitment was concentrated among people already familiar with DevOps. The same respondent could characterize technical practice, workplace climate, and organizational performance. The authors report tests for early-versus-late response bias and common-method variance and say those checks did not identify a problem, but internal consistency does not turn perceptions into production telemetry or audited financial results.11 Nor can a cross-sectional association by itself rule out reverse causation or an unmeasured factor affecting both practice and performance.

Independent studies have tested narrower pieces rather than independently replicating the book's whole model. A 2021 Swiss Post study automated the four metrics for one ten-person team. Six participants generally found them useful, but were less confident in the stability measures, expected both motivation and pressure, and said the metrics did not cover all quality or feedback work.12 A 2024 ACM case study collected telemetry for 37 microservices over four weeks. It found that an aggregate team score could conceal large differences among systems, while also cautioning that its single-team setting limits generalization.13 A separate 123-person international survey found many positive perceived effects from DevOps practices, but correlations varied by practice and some respondents reported less predictable delivery or less enjoyable work.14

The people inside the delivery system make this more than a measurement dispute. Accelerate explicitly studied deployment pain and burnout and associated technical and lean practices with lower reported levels of both; it did not experimentally establish that adopting those practices causes healthier work.15 Later implementation research likewise records the risk that measurement creates pressure, and current DORA guidance rejects using the metrics as competitive targets.12 Fast feedback can reduce rework and release fear, or become continuous pressure on developers, quality staff, security teams, and on-call responders.

The delivery measures also do not record accessibility, surveillance, energy use, content harm, or whether the product deserves to exist.8 Accelerate makes one part of the delivery system unusually visible. It leaves the institution responsible for deciding what that system should serve and whose evidence counts when speed produces costs elsewhere.

Structured reading paths and evidence limits

The dependency on Out of the Crisis is an editorial digital-operationalization comparison: modern delivery studies extend feedback, variation, and system thinking without establishing that Deming's program caused DevOps. The dependency on Amazon Shareholder Letters compares executive operating mechanisms with delivery research. It does not show that Amazon implemented the book's sampled capabilities or research design.

The path to Google/Alphabet identifies an institution associated with the continuing DORA program, not a claim that one company represents the survey population. Paths to learning, quality, and reliability, measurement, accounting, and control, and culture, informal organization, trust, and voice distinguish feedback, metrics, and work climate. Organizational intelligence connects those mechanisms, while benefit for all life asks whether faster, more stable delivery serves a justified purpose.

These are editorial reading paths. No structured impacts or typed relations are encoded, so the record supports no net-impact estimate or historical-influence claim. The evidence includes authorial studies, program reports, one participant case, and bounded independent implementations and surveys. It does not contain representative outcome evidence from on-call responders, contractors, moderators, users harmed by software, infrastructure communities, or ecologically affected life. Those gaps limit any inference from delivery performance to human or ecological benefit.16

Source notes

  1. Participant case history: Gary Gruver, Mike Young, and Pat Fulghum, A Practical Approach to Large-Scale Agile Development: How HP Transformed LaserJet FutureSmart Firmware (Addison-Wesley, 2013), foreword and preface, pp. xiv–xvii, and “Summary,” p. 177, official publisher sample. The authors led the transformation; the reported figures are HP's own activity and outcome measures.

  2. Primary synthesis: Nicole Forsgren, Jez Humble, and Gene Kim, Accelerate: The Science of Lean Software and DevOps (IT Revolution, 2018), preface, pp. xxi–xxiii; chapter 15, “The Data for the Project,” pp. 169–175; and appendix B, pp. 211–222, official publisher record.

  3. Program report: Puppet Labs, IT Revolution Press, and ThoughtWorks, 2014 State of DevOps Report, “Who Took the Survey,” pp. 7–8, report PDF.

  4. Program methodology: Alanna Brown et al., 2017 State of DevOps Report (Puppet and DORA), pp. 48–49, especially “Target population & sampling method” and “Study design,” report PDF; Forsgren, Humble, and Kim, Accelerate, chapter 15, pp. 169–175. The report describes the design as cross-sectional and the recruitment as snowball sampling, likely limited to teams familiar with DevOps.

  5. Authors' methodology: Forsgren, Humble, and Kim, Accelerate, chapter 12, pp. 129–143; chapter 15, pp. 169–175; and appendix C, pp. 223–230; Brown et al., 2017 State of DevOps Report, pp. 48–49, report PDF. The book classifies its analyses as descriptive, exploratory, and inferential-predictive rather than causal.

  6. Authors' findings: Forsgren, Humble, and Kim, Accelerate, chapter 2, pp. 13–27; chapters 3–8, pp. 28–88; and appendix A, pp. 201–210; Puppet Labs et al., 2014 State of DevOps Report, pp. 11–20, report PDF. Organizational performance was a respondent-rated construct covering profitability, market share, and productivity, not audited company data.

  7. DORA's retrospective account: Nathen Harvey, “A history of DORA's software delivery metrics,” January 2, 2026, sections “Origins,” “Refining definitions,” and “From four keys to five,” DORA, accessed July 14, 2026. This is the research program's account of its own changing model.

  8. Primary definitions: Forsgren, Humble, and Kim, Accelerate, chapter 2, especially pp. 13–18 and figure 2.1, official publisher record. The conclusion about what falls outside the measures follows from those defined boundaries.

  9. Primary synthesis: Forsgren, Humble, and Kim, Accelerate, “Quick Reference: Capabilities to Drive Improvement” and appendix A, pp. 201–210, official publisher record.

  10. Current program guidance: Nathen Harvey, “DORA's software delivery performance metrics,” sections “Context matters” and “Common pitfalls,” last updated January 5, 2026, DORA, accessed July 14, 2026.

  11. Authors' methods and reported checks: Forsgren, Humble, and Kim, Accelerate, chapter 14, pp. 153–168; chapter 15, pp. 169–175; and appendix C, pp. 223–230, especially “Tests for Bias,” official publisher record.

  12. Independent implementation study: Marc Sallin et al., “Measuring Software Delivery Performance Using the Four Key Metrics of DevOps,” in Agile Processes in Software Engineering and Extreme Programming (XP 2021), pp. 103–119, especially sections 3.3, 5.2–5.3, and 6.2–6.4, Springer, doi:10.1007/978-3-030-78098-2_7.

  13. Independent telemetry case: Janick Rüegger et al., “Fully Automated DORA Metrics Measurement for Continuous Improvement,” ICSSP '24, pp. 36–45, especially pp. 36–38 and 43–45, author-hosted paper, doi:10.1145/3666015.3666020.

  14. Independent survey: Tyron Offerman et al., “A Study of Adoption and Effects of DevOps Practices,” 2022 IEEE ICE/ITMC–IAMOT, especially sections V–VII, author manuscript, doi:10.1109/ICE/ITMC-IAMOT55089.2022.10033313. The 123-respondent study is itself cross-sectional and perception-based; it is corroboration and counterevidence, not a causal replication.

  15. Authors' worker-outcome analysis: Forsgren, Humble, and Kim, Accelerate, chapter 9, “Making Work Sustainable,” pp. 89–100, especially figure 9.1, official publisher record.

  16. Evidence-composition audit: the authorial survey design and its limits are documented in [^sampling-design], [^analysis], and [^cmv-limits]; independent bounded tests appear in [^sallin], [^ruegger], and [^offerman]. None is a representative affected-user, infrastructure- community, labor-chain, or ecological outcome study, so no such result is inferred.

Research record

Evidence basis

Claim Cited. Material claims carry source locators; comparative interpretation may still evolve.

Open questions and affected lives

Benefit-to-life status: Seed

  • When do delivery metrics help teams learn, and when do they become targets used to rank, surveil, or accelerate workers?
  • Whose reliability matters when fast delivery benefits customers but transfers interruption, moderation, security, or on-call costs to less powerful people?
  • What environmental and material costs of software infrastructure remain outside the book's organizational-performance measures?
  • How can leaders preserve the enabling capabilities behind the metrics instead of gaming the visible measures?

These questions remain open; absence from the record does not imply absence of benefit or harm.

Structured atlas record

Reading prerequisites

Provenance and sources

Online anchors