An engineering manager brings a number to planning: developers spend 13.5 hours a week on technical debt. Should the team reserve a third of its next sprint to pay debt down? The number sounds precise enough to set a budget. It came from a 2018 survey asking people to estimate someone else's average week, so it cannot tell this team what to schedule.
Technical-debt statistics become useful when the question behind each number stays attached. Hours spent on maintenance, effort spent managing debt, and a ranking of debt sources describe different things. None tells you how much time a particular intervention will return.
The weekly-hours estimate is a survey answer
In Stripe's 2018 Developer Coefficient report, produced with Harris Poll, developers estimated that the average developer at their company spent 13.5 hours a week addressing technical debt. A separate question produced a 17.3-hour estimate for maintenance, described as dealing with bad code or errors, debugging, refactoring, and modifying. The reported average workweek was 41.1 hours. The report surveyed more than 1,000 developers and more than 1,000 C-level executives. Its methodology names the US, UK, France, Germany, and Singapore, although the introduction says six countries.
These are respondents' estimates, not time logs. The report does not establish that the 13.5 hours sit outside the 17.3 maintenance hours. A refactor could satisfy both descriptions; adding the figures would count some work twice. Maintenance can also be ordinary work on a healthy system. A routine dependency update or an expected product change is not automatically repayment of an earlier shortcut. The survey's age, sponsorship, and wording matter when someone repeats its figures as though they describe today's team.
A share of effort has a different denominator
Martini, Besker, and Bosch's 2018 study asked practitioners how much overall development effort usually went to technical-debt management activities. Of 226 completed survey responses across 15 organizations within eight large companies, 215 answered that question. The researchers converted response ranges to their midpoints and calculated a 25.9% mean and 25% median. In a separate case-study phase, they examined three companies that had started tracking debt.
That percentage is an estimate of development effort in those organizations. It is not a measured share of each person's workweek, and it is not interchangeable with Stripe's weekly hours. The studied organizations were selected for a large-company setting; most worked on embedded software. The paper says its result should not be generalized to small organizations. Its 26% figure for respondents using a tool to track debt addresses a separate practice, not the portion of effort a tool saves.
A ranking is not a fraction of debt
In Ernst and colleagues' 2015 study, 1,831 people at three large organizations began a survey and 536 answered every question. Individual questions had different response counts. For one question, 544 respondents ranked 14 possible debt sources for their current or most recent project. Bad architecture choices appeared in the top three for 296 of them, or about 54%.
That means 54% of the people answering that ranking question placed bad architecture choices near the top of a supplied list. It does not mean architecture caused 54% of their debt or consumed 54% of their time. The study captures practitioners' judgments about debt sources. It does not observe all debt items, assign their costs, or test whether a diagram prevents any of them.
Use the studies to frame a local decision
Return to the fictional team planning its sprint. It maintains a billing service and wants to know whether to spend time untangling a shared data-access module. The module is used by several features. Engineers say that changing one billing rule often requires edits elsewhere, but they have not established how often that happens or what it costs. The manager cannot turn Stripe's hours into a sprint allocation, Martini's percentage into a savings target, or Ernst's ranking into proof that this module is the team's largest debt item.
The team can define the decision first: should it change the shared module now, or leave it in place while shipping the next billing feature? For recent work on that module, it can review tickets and ask engineers to identify the extra steps caused by the shared dependency. A ticket might include ordinary feature work, a necessary test update, and two days spent repairing unrelated callers. Only the last part is a candidate for the module's ongoing cost. The team should record why it attributes that work to the dependency, because memory and ticket labels alone will not make the classification reliable.
Suppose the review finds several such tickets. That still does not produce an automatic return on a rewrite. The team needs a rough change cost, the expected frequency of similar work, and the risk that a replacement breaks existing billing behavior. If the next feature barely touches the module, postponing the change may be reasonable. If most planned features must cross it, the same evidence may support a bounded refactor. The studies can justify looking carefully; the local work history and upcoming changes determine the choice.
For a planning note, write the measure beside the number: the time window, whose work was counted, what qualified as extra effort, and what was excluded. Keep estimates separate from observed ticket or time data. Mark disputed classifications rather than smoothing them into a precise total. If the team's disagreement is about which billing callers depend on the shared module, they can sketch that relationship in Lycana and revise it through spoken or typed instructions as they inspect the code. The drawing can make the proposed dependency explicit; it cannot measure past effort or prove a refactor will save time.
The strongest use of an industry statistic is to sharpen a question for your own system. Before putting a number on a roadmap slide, ask whether it describes reported hours, a share of effort, or people's ranking of causes. Then collect the evidence that the actual decision needs.


