Rewrite or Rescue? A Decision Framework for AI-Built Applications

Last updated September 2026

Rewrite vs refactor for an AI app turns on one question: is the data model worth keeping? If the schema and the validated product behavior are sound, rescue the code around them. If the schema itself encodes a misunderstanding of the domain, rewrite that layer. The age of the code is irrelevant. What it knows is not.

Why does the usual rewrite vs refactor advice fail for an AI app?

The standard advice against rewriting comes from a specific claim: old code is ugly because it accumulated fixes for real bugs, and throwing it away throws away years of hard-won knowledge that exists nowhere else. That argument is correct for a fifteen year old billing system. It is close to meaningless for an application a founder generated in six weeks with Cursor or Lovable, because there is no accumulated knowledge in it. There is one model’s guess at how the problem should be structured, repeated across forty files.

The structural evidence backs this up. GitClear’s 2026 Maintainability Gap research, covering 623 million changed lines from 2023 through 2026, found duplicated code blocks rose 81 percent over 2023 levels to 73.0 duplicated lines per million changed, while moved code, the signal that someone actually refactored, collapsed from 21 percent in 2022 to 3.8 percent so far in 2026. Cross-file function connectivity fell 35 percent, from 343 method calls per thousand changed lines to 223. That last number is the one that matters for this decision. Low connectivity means the code is not a web of dependencies you have to untangle. It means the same logic is sitting in fifteen places, unaware of the other fourteen.

A tangled legacy monolith resists rewriting because everything touches everything. An AI-generated codebase usually resists nothing. It is wide, shallow, and repetitive, which makes the cost of replacing any single slice far lower than the cost of replacing a slice of a mature system. The classic rule inverts.

What should you measure before choosing rewrite vs refactor for an AI app?

Seven signals, all of which you can measure in an afternoon with the repository and a staging environment. Score one point for each condition on the right that is true.

Signal How to measure it Scores a point if
Schema fit Read the migrations and name the entity each table represents Tables do not map to real domain entities, or core state lives in a JSON blob
Test coverage that means anything Delete a random line of business logic, run the suite The suite still passes
Auth and payment correctness Attempt access to another tenant’s record with a valid session It returns data
Duplication Run a clone detector across the repo Over 15 percent of lines are in duplicated blocks
Dependency reality Verify every package in the manifest exists and is the one intended Any package is unresolvable, abandoned, or typosquatted
Comprehension Ask the strongest engineer you have to explain the request path end to end They cannot do it in twenty minutes
Change cost Time the last three small feature changes from branch to production A one-line change takes more than a day

Two of those rows carry more weight than the rest. Tenant isolation and payment paths are the failures that end companies rather than sprints, and they are common. Veracode’s 2025 GenAI Code Security Report, which tested more than 100 large language models on real coding tasks, found 45 percent of generated samples introduced an OWASP Top 10 vulnerability, rising to 72 percent for Java, and that the models failed to defend against cross-site scripting in 86 percent of the relevant cases. Newer and larger models did not do better. If you want the full list of what turns up, we catalogued the nine security flaws we find most often in AI-generated code.

What does the score tell you to do?

  1. 0 to 2 points: rescue. The application is structurally sound and the work is remediation. Fix in place, add the tests that were never written, and keep shipping features while you do it.
  2. 3 to 4 points: strangle one layer. Something specific is wrong, usually the service layer or the data access code. Replace that layer behind its existing interface while the rest of the app keeps running. Do not stop feature work for it.
  3. 5 or more, including the schema row: rewrite the core, keep the edges. A wrong data model cannot be refactored into a right one cheaply, because every query, every API response, and every screen encodes the mistake. Rebuild the model and the logic above it. Keep the front end, the infrastructure, and anything that already passes a security test.
  4. Any score, if auth or payments scores a point: stop the decision and fix that first. It is a live incident, not an architecture question.

Notice what is missing from the list: a full ground-up rewrite. It almost never wins on the merits, because the failure is rarely everywhere at once. The front end usually works. The deployment pipeline usually works. The product decisions embedded in the UI were validated by real users, which is the one piece of accumulated knowledge these codebases genuinely do hold.

What should never go in the rewrite pile?

Four things, regardless of how bad the code around them looks.

  • Production data and its migration history. You can replace the schema. You cannot replace the six months of customer records shaped by it, and the migration path is the actual project.
  • Anything already passing a security or compliance test. If a payment integration survived a penetration test, it is an asset. Rewriting it restarts the clock on that evidence.
  • Validated product behavior. Every quirk users depend on is a requirement whether or not it is documented. Capture it as tests before anyone deletes the code that implements it.
  • The team’s understanding. If one engineer knows why a workaround exists, get it written down this week. That knowledge leaves with them, and in an AI-built codebase it is often the only documentation.

How long does each path take, and what does it cost?

These are the shapes of the three paths as we run them, with the published prices attached. A $4,950 Rescue Audit is credited toward remediation, so the diagnosis is not a sunk cost whichever way the decision goes.

Path Trigger Typical elapsed time What you keep Main risk
Audit first Any score, before committing 10 business days Everything, plus a ranked findings report Delay, if you already know the answer
Rescue in place 0 to 2 points 4 to 12 weeks alongside feature work The whole application Remediation stalls if nobody owns the test suite
Strangle one layer 3 to 4 points 6 to 16 weeks Front end, infrastructure, data The interface you strangle behind turns out to leak
Rewrite the core 5 or more, schema included 3 to 6 months Front end, infrastructure, product decisions Feature freeze, and the second system is worse

The feature freeze is what kills rewrites, and it is worth being honest about why. A rewrite competes with a running product that keeps shipping. If the rewrite takes six months and the old system ships for all six, the new system has to catch up to a moving target on its last day. That is the arithmetic behind every rewrite that quietly becomes a parallel codebase nobody launches.

Will AI make the rewrite fast enough to be worth it?

This is the assumption that does the most damage, because it is the reason people choose rewrite over refactor without scoring anything. The evidence says to discount it. METR’s randomized controlled trial of experienced open-source developers found they were 19 percent slower on real tasks in their own repositories when using AI tools, while estimating afterward that they had been 20 percent faster. In Stack Overflow’s 2025 Developer Survey of roughly 49,000 respondents, 66 percent named “AI solutions that are almost right, but not quite” as their top frustration, and 45 percent said debugging AI-generated code takes more time than writing it themselves.

Google’s 2025 DORA report, drawing on nearly 5,000 technology professionals, found the same pattern at the organizational level: AI adoption correlates positively with delivery throughput and negatively with delivery stability. AI makes the rewrite faster to type and not faster to trust. That gap is the entire project.

The practical consequence is that the rewrite only pays when a senior engineer who understands the domain is driving it. We staff these from a vetted pool of 15,643 developers averaging more than ten years of experience, because the failure mode of an AI-assisted rewrite is not slow typing. It is a second codebase with the same misunderstanding, produced faster.

Book a Code Review

Score the seven signals yourself first. If the answer is obvious, act on it and skip the meeting. If it is not, or if the schema row is the one you are arguing about, that argument is worth an outside read. A code review walks your repository with a senior engineer and returns the score, the ranked findings, and the path, in writing. For the full picture of what that diagnosis covers, start with our pillar on what an AI code rescue is and what it costs, or the ten day method in how to audit an AI-generated codebase. When you are ready, book a code review and get the decision settled before the quarter is.