Reverse Engineering Legacy Code Using AI

Most enterprises run software that was written before the engineers now responsible for it were hired. Banking systems built on COBOL in the 1980s. Healthcare platforms written in PL/I that predate modern compliance standards. Government infrastructure whose original developers retired a decade ago. The code still runs. The institutional knowledge behind it is gone.

Modernizing any of these systems starts with understanding what the code actually does, function by function, dependency by dependency. That work is called legacy code reverse engineering, and it has long been the slowest and most expensive phase of any modernization program. Teams spent months reading undocumented code, interviewing whoever still remembered how it worked, and assembling a picture of systems nobody had documented properly in decades. AI changes both how long that takes and what it produces.

Why Legacy Codebases Resist Understanding

Reverse engineering a codebase means reconstructing its logic, business rules, and data flows from the code itself, without the documentation that usually no longer exists. For a system built and patched over 30 or 40 years by engineers who have since left, the work is enormous and the risk of missing something is real.

The talent challenge makes modernization even harder. Despite decades of digital transformation, 43% of banking systems still run on COBOL, according to Reuters research, leaving many institutions dependent on aging platforms and an increasingly scarce pool of specialized talent needed to maintain and modernize them.

The engineers who understand these systems are aging into retirement faster than new ones are learning the language. When the last person who knew a system leaves, the knowledge leaves with them, and modernization stalls at discovery because nobody can agree on what the current system does before deciding how to replace it.

That stall is expensive. Forrester found that organizations dedicate an average of 20% of their IT budget to managing technical debt. Every deferred migration raises maintenance cost, widens security exposure, and makes the next integration harder. The discovery phase, which should be the foundation that makes everything else faster, becomes the phase where timelines and budgets collapse.

What Changes When AI Reads the Code

The shift is one of scale. A senior engineer reading legacy code by hand can process a limited volume per week, and much of that work ends up in their head rather than in any document. AI models read structure and relationships at a different order of magnitude: comments, naming patterns, control flows, dependencies between modules, and call chains that span thousands of lines. They turn that analysis into structured, shareable output.

The practical result is functional documentation generated directly from the code: plain-language descriptions of what each module does, the business rules it enforces, the data it depends on, and unit tests that capture expected behavior as a validation baseline before any modernization begins. Engineers work from those specifications instead of spending weeks producing them by hand.

The market is moving decisively in this direction. Gartner predicts that by 2027, generative AI tools will cut legacy modernization costs by 70 percent by explaining legacy applications and generating their replacements. For most institutions, the question is no longer whether AI belongs in reverse engineering. It is whether they can run it with the rigor a regulated environment demands.

The Three Outputs That Make Modernization Possible

AI-assisted reverse engineering produces three artifacts that determine whether the program that follows succeeds.
The first is a functional specification: a readable account of what the system does, the business rules embedded in the code, and the data flows it depends on. For many enterprises this documentation has never existed in written form. AAVA™ generates it directly from the codebase, giving engineering teams a foundation to work from before they write a line of replacement code.

The second is an architecture map. AI traces dependencies between modules, identifies tightly coupled components, and flags where technical debt is concentrated. That map lets engineering and program leadership sequence the work deliberately, deciding what to replatform, what to refactor, and what is stable enough to leave until a later phase.

The third addresses security and compliance. Legacy systems accumulate unpatched vulnerabilities over decades, and banks and healthcare providers often have limited visibility into where those vulnerabilities sit. AI-assisted static analysis surfaces known vulnerability patterns at speed, giving compliance teams a prioritized list to address before or during modernization, and turning a risk audit that would take months into output available in weeks.

What This Looks Like in a Production Engagement

The clearest evidence comes from regulated industries, where legacy systems are oldest and the cost of error is highest.

For a digital-first banking pioneer running more than 900,000 lines of 1980s code, AAVA completed the reverse engineering in three weeks. The modernization that followed finished in half the time and a third of the cost of a traditional approach, with 23 go-to-market capabilities defined directly from the reverse-engineered codebase. The reverse engineering phase made those outcomes possible: it produced a complete picture of a 40-year-old system in the time a manual process would have spent reading the first quarter of it.

For a 200-year-old UK bank, AAVA agents mapped the architecture in weeks rather than months, following a £50 million transformation attempt that had failed. That analysis gave engineers the foundation to protect 5.2 million customers and reach a 50 to 75 percent velocity gain on the rebuild. Architecture mapped in weeks was what separated this attempt from the one that spent £50 million and shipped nothing.

Engineers carried the judgment in both engagements. They defined what a good reverse engineering output looked like, validated the specifications AAVA produced, and made the calls on sequencing and priority. AAVA handled the volume: reading millions of lines, identifying patterns, and generating documentation at a rate that gives a senior engineer full context instead of forcing them to build it by hand.

Why Orchestration Determines Whether AI Actually Ships

A persistent problem in enterprise AI is that individual tools produce outputs that do not connect. A code-analysis tool generates findings an engineer feeds by hand into a documentation generator, which produces artifacts an engineer then interprets for the testing team. None of them share context, and the integration work between them consumes much of the time AI was supposed to save.

Agentic platforms close that gap by running coordinated agents inside a single governed workflow. AAVA orchestrates agents across the full software development lifecycle, from code analysis through documentation, test generation, and deployment, with each agent passing structured output to the next. In reverse engineering, the architecture map feeds directly into test generation, and the business-rule documentation becomes the specification that guides the rewrite. The integration work is already done.

In regulated industries, the governance layer is not optional. Banks and healthcare organizations need audit trails, consistent guidelines, and human review checkpoints around agent activity. AAVA provides those controls by design. That is the difference between an enterprise-grade deployment and a pilot that works in a demo and fails in production.

What Determines the Quality of the Output

AI produces better reverse engineering when given richer context. A model reading raw COBOL in isolation returns weaker results than one given the same code alongside related test files, database schemas, and deployment configuration. Engineers who know the system improve output quality by supplying that context.

Validation is the other factor. AI-generated specifications have to be checked against actual system behavior, which means running the generated test suite against the existing system in a controlled environment. A multi-pass approach, where inaccurate specifications trigger targeted re-analysis of the affected modules, preserves the time savings while catching errors before they reach modernization work.

Organizations that treat reverse engineering as a one-time discovery phase get less from it than those that keep it as a living artifact. As modernization proceeds, current specifications and architecture maps give teams ongoing visibility into what has migrated, what remains, and where new dependencies have appeared. That visibility is what separates programs that finish from programs that stall midway.

How Ascendion Delivers Legacy Code Reverse Engineering

Ascendion’s Engineering to the Power of AI™ model runs AAVA and its engineers as one delivery system rather than separate workstreams. AAVA reads millions of lines of undocumented code and produces specifications, architecture maps, and test suites at a pace no manual process can match. Ascendion’s engineers carry the judgment AI cannot: validating outputs, sequencing the work, and making the decisions that determine whether the program delivers on its business case.

The client outcomes follow from that combination, and they hold up where delivery is hardest: a 40-year-old banking platform modernized at a third of the cost, a 200-year-old bank’s architecture mapped in weeks after a £50 million failure. Understanding the existing system accurately is what makes everything downstream faster and cheaper.

For engineering and technology leaders weighing a legacy modernization program, the discovery phase is where the economics of the whole effort are set. It is the right place to start.

See AAVA in production. Explore the AAVA platform →

Ascendion is the AI-native disruptor reinventing how global enterprises build software for impact. Its engineering teams, powered by AAVA, the company’s proprietary agentic AI platform, deliver measurable business outcomes: accelerating growth, unlocking capital, and de-risking transformation. With 11,000+ engineering professionals and 10,000+ AI agents working across 12 countries, Ascendion delivers the promise of AI to more than a third of the Fortune 500. Learn more at https://www.ascendion.com.

Engineering to the Power of AI™, AAVA™, EngineeringAI, Engineering to Elevate Life™, Enterprise PlatformsAI, Data & InsightsAI, ExperienceAI, GCCAI, OperationsAI, Platform EngineeringAI, ProductAI, and Quality EngineeringAI are trademarks or service marks of Ascendion®. AAVA™ is pending registration. Unauthorized use is strictly prohibited.