How to Measure Technical Debt in a Codebase You Did Not Write
Last updated October 2026
To measure technical debt in a codebase you did not write, work from the two artifacts you always inherit: the version control history and the code itself. Score file-level health, rank hotspots where change frequency meets complexity, convert the remediation estimate into a ratio, then price that ratio at your own loaded engineering cost.
Why does the usual advice on how to measure technical debt fail on inherited code?
Almost every published metric list assumes four things an inheritor does not have. It assumes a ticket system wired to commits, so cycle time and defect density can be computed per story. It assumes a baseline from before the decay, so this quarter can be compared to something. It assumes engineers who remember why a workaround exists. And it assumes sprint history long enough for a trend line.
Inherit a codebase through an acquisition, a vendor handover, or a founder who built it alone, and you get none of that. The Jira project was never used, or it was used by a team that no longer exists. What you do get, always, is the repository: every commit, every author, every file as it stands today. That is enough, and it is more objective than the ticket data, because nobody was grooming it to look good.
The number matters because the decision it feeds is a budget decision. In McKinsey’s survey of 50 CIOs at companies above $1 billion in revenue, respondents put tech debt at 20 to 40 percent of the value of their entire technology estate before depreciation, and said 10 to 20 percent of the budget earmarked for new products gets diverted into debt-related work. Sixty percent said the problem had grown perceptibly over three years. An estimate you present to a board has to survive at that scale, which means it has to be traceable to artifacts, not impressions.
What can you measure with nothing but the repository?
Five measurements, all derivable from a clone of the repo and a static analyzer. None of them require a ticket, a former employee, or a dashboard subscription.
| Measurement | Where it comes from | Treat as a red flag when | Time to produce |
|---|---|---|---|
| File-level health or complexity score | Static analysis across the whole tree | The worst decile of files holds the code you ship from weekly | Half a day |
| Hotspots | Commit counts per file from the git log, crossed with complexity | The top ten files by change frequency are also the ten least healthy | Half a day |
| Duplication share | A clone detector over the tree | More than 15 percent of lines sit inside duplicated blocks | Two hours |
| Knowledge concentration | Author counts per file from the git log | More than half the hotspot files have a single author who has left | One hour |
| Remediation estimate and ratio | The analyzer’s rule-level effort totals | The ratio lands above 20 percent | One hour, once analysis has run |
Run them in that order. Health and hotspots tell you where to look, duplication and ownership tell you why the code resists change, and the ratio is the only output a non-engineer will read.
How do you measure technical debt as a ratio instead of an adjective?
The convention most tools implement comes from the SQALE method. Per SonarSource’s own metric definitions, technical debt is the sum of estimated remediation minutes for every maintainability issue found, and the debt ratio is that total divided by the cost of writing the code in the first place, where the default cost to develop one line is 30 minutes. Their worked example: 122,563 minutes of debt against 63,987 lines of code gives a ratio of 6.4 percent. The letter grade follows a fixed grid.
| Maintainability rating | Debt ratio | Reasonable reading |
|---|---|---|
| A | 0 to 5 percent | Normal maintenance, no programme needed |
| B | 5 to under 10 percent | Fix opportunistically inside feature work |
| C | 10 to under 20 percent | Budget a named remediation track |
| D | 20 to under 50 percent | Feature estimates are already unreliable |
| E | 50 percent and above | Decide between remediation and replacing a layer |
Be honest with yourself about what that number is. The 30 minutes per line is a configurable default, not a measurement of your team, and the remediation minutes come from rule authors who never saw your system. The ratio is useful as a comparator across modules of the same codebase and as a tracked line over time. It is close to useless as an absolute claim about one repository on one day. Calibrate the cost-per-line against what a feature of known size actually took your team, and say in the report which number you used.
Which files should you look at first?
The hotspot list, and only the hotspot list, because debt is not spread evenly. In the study Code Red, by Adam Tornhill and Markus Borg, 39 proprietary production codebases and 30,737 files were scored for code health and matched against issue data. Files in the worst category carried 3.70 defects on average against 0.25 for healthy files, roughly fifteen times more. Work in those files took 124 percent more time in development, and their worst-case cycle times ran about nine times longer than healthy files, which is the part that destroys forecasting rather than throughput.
Their 2024 follow-up across 79 projects and 46,211 files found the distribution brutally skewed: 87 percent of files healthy, 12 percent in warning, and 0.55 percent in the alert band. The practical implication for an inherited codebase is a relief. You are not remediating a repository. You are remediating a few dozen files, and your first job is to find out whether those files are the ones your roadmap touches. If they are not, the debt is real and can wait.
How do you turn the ratio into a number a CFO will accept?
- Take the remediation estimate in hours from the analyzer, for the hotspot files only, not the whole tree.
- Multiply by your fully loaded hourly cost, not base salary. Use the same loaded figure you use for headcount planning, as covered in our guide to budgeting an engineering team for the next 12 months.
- Add the carrying cost: the hours per month your team currently spends on those files, multiplied by the 1.24 time penalty from the Code Red data, to express what the debt costs you if you do nothing.
- Add the forecast penalty separately. Nine times worst-case cycle time is a missed-commitment risk, not a labour cost. Name it in the report rather than folding it into a dollar figure.
- Present three numbers: cost to fix, annual cost to carry, and the date by which the roadmap collides with the hotspots. That third number is what actually moves a decision.
Two of those lines are estimates built on published defaults, so label them that way. A finance team will accept a stated assumption far more readily than a confident round number with no derivation behind it.
What changes when the code was written by AI?
The signal order changes, because the failure shape is different. GitClear’s Maintainability Gap research, covering 623 million changed lines from 2023 through 2026, found duplicated code blocks up 81 percent over 2023 to 73.0 duplicated lines per million changed, while moved code, the fingerprint of actual refactoring, fell from 21 percent of changed lines in 2022 to 3.8 percent in 2026 so far. Copy and paste rose from 9.4 percent to 15.7 percent over the same window. Cross-file function connectivity dropped 35 percent, from 343 method calls per thousand changed lines to 223, and changes to code untouched for twelve months or more fell 74 percent, from 1.7 percent to 0.46 percent.
Read those together and the diagnosis is specific. An AI-heavy codebase is usually wide rather than tangled: the same logic repeated in many places, each copy unaware of the others, and almost nothing ever revisited. So duplication share and knowledge concentration are leading indicators here, while cyclomatic complexity per function can look deceptively fine. The other number worth watching is error-masking constructs, up 47 percent in the same dataset, because swallowed exceptions are how an inherited system hides its real defect count from you.
Do not expect an AI reviewer to produce this picture for you. As we set out in detail on why AI code review tools do not catch architecture problems, those tools read the diff, and inherited debt is a property of the system. For the mechanics of why AI-assisted work compounds faster in the first place, start with the pillar on why AI-generated code compounds technical debt faster than human code.
Book a Code Review
Everything above is runnable in about three days by an engineer with repository access, and doing it yourself is the right first move. Where an outside read earns its keep is the judgement call: which hotspots are load-bearing for your roadmap, and whether the schema under them is sound. Our $4,950 Rescue Audit runs that work to an executive readout in ten business days and is credited toward remediation, so the diagnosis is not a sunk cost. The method behind it is written up in how to audit an AI-generated codebase in 10 days. If the codebase is already live and you want the carrying cost tracked rather than estimated once, our PULSE telemetry reports more than 400 delivery data points a week against exactly these files. Book a code review and get the three numbers in writing.