AI agents in legacy code: modernizing with risk under control

02. 10. 2026
AI agent upravující vymezenou část staršího systému s kontrolou před nasazením.

The proposition is appealing: give an AI agent an older application and receive a cleaner system that is easier to maintain. For business software, though, the new implementation needs to do more than look well structured. It must continue supporting people's work, connected applications, and less common operational scenarios.

Addy Osmani addresses this challenge in Brownfield Agentic Engineering. He emphasizes risk-based work boundaries, knowledge that lives outside the repository, verification of existing behavior, completed migrations, and careful introduction of parallel agent work. The workflow below applies those principles to an illustrative business application.

Start with a specific business constraint

A legacy system is an existing application whose development may be constrained by older technologies, tightly coupled components, missing tests, or incomplete documentation. Age alone does not determine the problem. What matters is whether the team can make and verify changes at an acceptable cost.

“Modernize our application with AI” is therefore an incomplete brief. A better starting point identifies a specific obstacle: creating a new export format takes too long, changing one screen affects several modules, or a dependency prevents further upgrades.

In custom software development, a technical change should have a verifiable outcome. For example, separating export formatting could allow the team to add another format without modifying order processing. That gives the agent a bounded assignment and gives reviewers a concrete basis for assessing the result.

This distinction also helps prioritize investment. A component might be untidy yet rarely change. Another might delay every customer request. A modernization plan should reflect the cost of those constraints, rather than treating every outdated part of the application as equally urgent.

Illustrative example: an order export

Consider an internal application that exports orders to another system. Its export module combines data selection, file formatting, and updates to order status. The proposed change separates formatting into its own component. This is an illustrative scenario, not a customer case study.

The team first defines which outcomes must remain unchanged. The internal structure may differ, but the agent cannot independently change column order, the meaning of an empty value, or the point at which an order becomes marked as exported.

This creates an assignment reviewers can evaluate. “Cleaner code” is an aspiration. “Equivalent export files and state changes, with formatting separated from order processing” describes observable results. It also makes disagreements easier to resolve before implementation begins.

Set agent authority according to the impact of a change

Green, yellow, and red zones can provide a practical decision aid. Someone who understands the application and its operation must set the boundaries. The presence of tests alone does not automatically make a component low risk.

Zone Example in the illustrative application Suggested working mode
Green An isolated preview formatter with meaningful tests The agent prepares a small change and runs the required checks.
Yellow An export with poorly understood historical exceptions The team establishes behavior and adds tests before allowing a scoped modification.
Red Export permissions or calculations that affect invoicing An accountable developer directly supervises each change.

Consider reach as well as local complexity. A tested utility might be used by exports, invoicing, and reporting. Changing it could have broader consequences than editing a single screen, even if the patch is short.

A useful agent brief identifies permitted files, interfaces that must remain compatible, and stop conditions. For example, if the agent discovers that separating formatting requires a database schema change, it reports that finding. The team then evaluates the expanded scope before proceeding.

The same principle applies to deployment authority. Permission to prepare code does not have to include permission to release it. Those are separate steps with different consequences and can have different owners.

Capture business rules that source code cannot explain

In the order export, an unusual date format might be required by an older accounting application. An empty value might mean something different from zero. Certain orders might be processed the following day because of an agreed operational procedure.

Record those constraints alongside the assignment and identify who confirmed them. A useful note could read: “Preserve an empty delivery date in the export. The receiving application treats it as a date awaiting confirmation. Confirmed by the process owner.”

If nobody knows the reason, record an open question. An assumption must not quietly become a business requirement. The team can investigate with an application user or an appropriately prepared sample of data.

Keep coding agents distinct from AI features running inside the application. They involve different decisions. AI implementation support might help automate data handling within a business system. Separating an export module does not itself require adding such a feature.

For a decision maker, this distinction matters when evaluating proposals. Ask what changes in the development workflow, what changes in the application, and which benefits depend on each. A proposal that combines the two should make both scopes explicit.

Establish today's behavior before changing the implementation

Characterization tests capture how existing software behaves for selected inputs. They help detect unintended differences during refactoring, but do not prove that the original behavior is correct for the business. The practical characterization testing guide explains this distinction.

For the illustrative export, the team prepares orders covering ordinary and unusual cases: a missing delivery date, several line items, and a value containing the file's column separator. Outputs from the original version become the comparison baseline.

The file is only part of the result. Reviewers also need to check changes in stored order state and the failure path. If the export fails, an order must not accidentally appear to have been successfully delivered to the receiving system.

Expected results should come from the original implementation and confirmed requirements. If the agent builds a replacement and then changes test expectations to match it, the team loses an independent comparison.

When a test reveals an existing defect, separate that correction from the structural refactor. First decide what the business requires, then update the test and implementation deliberately. Preserving a known defect should not become an automatic modernization objective.

A practical review therefore asks two different questions: did the restructuring preserve the agreed behavior, and are any behavior changes intentional? Keeping those questions separate helps the team explain what it is approving.

Define completion in operational terms

An incremental migration may temporarily run old and new components side by side. Microsoft's Strangler Fig pattern documentation describes this approach. The transition needs an explicit plan for retiring the replaced functionality.

In our export scenario, the team can compare outputs from both implementations in a test environment. In production, however, it must be clear which implementation actually sends data. Sending an order twice could create a duplicate in the receiving application.

Define completion before release: every agreed order type uses the new export, nothing still calls the old function, documentation reflects the current design, and operational checks show no unresolved differences.

Include a recovery procedure. Reverting code may be sufficient for some changes. If data has changed or an order has already been sent to another application, restoring the previous release will not necessarily restore that state. The plan must address those consequences separately.

This connects modernization to application support and maintenance. A release needs monitoring, useful failure records, and someone responsible for deciding what happens next.

Plan an observation period appropriate to the process. An export used once a month cannot be assessed solely from an uneventful afternoon after release. The relevant business cycle should inform when the team considers the change operationally verified.

Add parallel agents when the team can absorb their output

Concurrent work is useful for independent assignments. One agent might prepare the export change while another updates documentation for an unrelated module. If both modify a shared interface, they need coordination.

Begin with one bounded assignment and evaluate the whole delivery process. How much time went into preparation, verification, and corrections? Could a developer who did not prepare the patch understand it? Did the change leave unexplained exceptions behind?

Use a consistent handover format: purpose, affected components, checks performed, known limitations, and release procedure. Team capacity depends on the ability to assess and accept results as well as the number of agents generating them.

If review becomes the bottleneck, adding more implementation capacity can lengthen the queue. Investigate whether assignments are too broad, acceptance criteria are unclear, or the same missing knowledge repeatedly causes rework before expanding parallel execution.

Prepare five answers before the first assignment

Choose a small part of the system with a specific benefit for the first pilot. Before work begins, the team should answer these questions:

  1. What problem does the change solve, and how will we assess the benefit?
  2. Which behavior must remain unchanged, and who can confirm it?
  3. What may the agent modify, and what discovery requires it to stop?
  4. How will we verify the result before release and during operation?
  5. What counts as completion, and how will we handle a failed release?

If the answers are missing, gathering them can be the first assignment. That is a useful outcome in its own right: the team gains evidence for deciding the scope, cost, and approach to modernization.

Are you maintaining an older application and considering where AI could help its development? At 1. Web IT, we can assess the current system and help plan the next steps. Tell us what is holding back its development.

More articles