The proposition is appealing: give an AI agent an older application and receive a cleaner system that is easier to maintain. For business software, though, the new implementation needs to do more than look well structured. It must continue supporting people's work, connected applications, and less common operational scenarios.
Addy Osmani addresses this challenge in Brownfield Agentic Engineering. He emphasizes risk-based work boundaries, knowledge that lives outside the repository, verification of existing behavior, completed migrations, and careful introduction of parallel agent work. The workflow below applies those principles to an illustrative business application.
A legacy system is an existing application whose development may be constrained by older technologies, tightly coupled components, missing tests, or incomplete documentation. Age alone does not determine the problem. What matters is whether the team can make and verify changes at an acceptable cost.
“Modernize our application with AI” is therefore an incomplete brief. A better starting point identifies a specific obstacle: creating a new export format takes too long, changing one screen affects several modules, or a dependency prevents further upgrades.
In custom software development, a technical change should have a verifiable outcome. For example, separating export formatting could allow the team to add another format without modifying order processing. That gives the agent a bounded assignment and gives reviewers a concrete basis for assessing the result.
This distinction also helps prioritize investment. A component might be untidy yet rarely change. Another might delay every customer request. A modernization plan should reflect the cost of those constraints, rather than treating every outdated part of the application as equally urgent.
Consider an internal application that exports orders to another system. Its export module combines data selection, file formatting, and updates to order status. The proposed change separates formatting into its own component. This is an illustrative scenario, not a customer case study.
The team first defines which outcomes must remain unchanged. The internal structure may differ, but the agent cannot independently change column order, the meaning of an empty value, or the point at which an order becomes marked as exported.
This creates an assignment reviewers can evaluate. “Cleaner code” is an aspiration. “Equivalent export files and state changes, with formatting separated from order processing” describes observable results. It also makes disagreements easier to resolve before implementation begins.
Green, yellow, and red zones can provide a practical decision aid. Someone who understands the application and its operation must set the boundaries. The presence of tests alone does not automatically make a component low risk.
| Zone | Example in the illustrative application | Suggested working mode |
|---|---|---|
| Green | An isolated preview formatter with meaningful tests | The agent prepares a small change and runs the required checks. |
| Yellow | An export with poorly understood historical exceptions | The team establishes behavior and adds tests before allowing a scoped modification. |
| Red | Export permissions or calculations that affect invoicing | An accountable developer directly supervises each change. |
Consider reach as well as local complexity. A tested utility might be used by exports, invoicing, and reporting. Changing it could have broader consequences than editing a single screen, even if the patch is short.
A useful agent brief identifies permitted files, interfaces that must remain compatible, and stop conditions. For example, if the agent discovers that separating formatting requires a database schema change, it reports that finding. The team then evaluates the expanded scope before proceeding.
The same principle applies to deployment authority. Permission to prepare code does not have to include permission to release it. Those are separate steps with different consequences and can have different owners.
In the order export, an unusual date format might be required by an older accounting application. An empty value might mean something different from zero. Certain orders might be processed the following day because of an agreed operational procedure.
Record those constraints alongside the assignment and identify who confirmed them. A useful note could read: “Preserve an empty delivery date in the export. The receiving application treats it as a date awaiting confirmation. Confirmed by the process owner.”
If nobody knows the reason, record an open question. An assumption must not quietly become a business requirement. The team can investigate with an application user or an appropriately prepared sample of data.
Keep coding agents distinct from AI features running inside the application. They involve different decisions. AI implementation support might help automate data handling within a business system. Separating an export module does not itself require adding such a feature.
For a decision maker, this distinction matters when evaluating proposals. Ask what changes in the development workflow, what changes in the application, and which benefits depend on each. A proposal that combines the two should make both scopes explicit.
Characterization tests capture how existing software behaves for selected inputs. They help detect unintended differences during refactoring, but do not prove that the original behavior is correct for the business. The practical characterization testing guide explains this distinction.
For the illustrative export, the team prepares orders covering ordinary and unusual cases: a missing delivery date, several line items, and a value containing the file's column separator. Outputs from the original version become the comparison baseline.
The file is only part of the result. Reviewers also need to check changes in stored order state and the failure path. If the export fails, an order must not accidentally appear to have been successfully delivered to the receiving system.
Expected results should come from the original implementation and confirmed requirements. If the agent builds a replacement and then changes test expectations to match it, the team loses an independent comparison.
When a test reveals an existing defect, separate that correction from the structural refactor. First decide what the business requires, then update the test and implementation deliberately. Preserving a known defect should not become an automatic modernization objective.
A practical review therefore asks two different questions: did the restructuring preserve the agreed behavior, and are any behavior changes intentional? Keeping those questions separate helps the team explain what it is approving.
An incremental migration may temporarily run old and new components side by side. Microsoft's Strangler Fig pattern documentation describes this approach. The transition needs an explicit plan for retiring the replaced functionality.
In our export scenario, the team can compare outputs from both implementations in a test environment. In production, however, it must be clear which implementation actually sends data. Sending an order twice could create a duplicate in the receiving application.
Define completion before release: every agreed order type uses the new export, nothing still calls the old function, documentation reflects the current design, and operational checks show no unresolved differences.
Include a recovery procedure. Reverting code may be sufficient for some changes. If data has changed or an order has already been sent to another application, restoring the previous release will not necessarily restore that state. The plan must address those consequences separately.
This connects modernization to application support and maintenance. A release needs monitoring, useful failure records, and someone responsible for deciding what happens next.
Plan an observation period appropriate to the process. An export used once a month cannot be assessed solely from an uneventful afternoon after release. The relevant business cycle should inform when the team considers the change operationally verified.
Concurrent work is useful for independent assignments. One agent might prepare the export change while another updates documentation for an unrelated module. If both modify a shared interface, they need coordination.
Begin with one bounded assignment and evaluate the whole delivery process. How much time went into preparation, verification, and corrections? Could a developer who did not prepare the patch understand it? Did the change leave unexplained exceptions behind?
Use a consistent handover format: purpose, affected components, checks performed, known limitations, and release procedure. Team capacity depends on the ability to assess and accept results as well as the number of agents generating them.
If review becomes the bottleneck, adding more implementation capacity can lengthen the queue. Investigate whether assignments are too broad, acceptance criteria are unclear, or the same missing knowledge repeatedly causes rework before expanding parallel execution.
Choose a small part of the system with a specific benefit for the first pilot. Before work begins, the team should answer these questions:
If the answers are missing, gathering them can be the first assignment. That is a useful outcome in its own right: the team gains evidence for deciding the scope, cost, and approach to modernization.
Are you maintaining an older application and considering where AI could help its development? At 1. Web IT, we can assess the current system and help plan the next steps. Tell us what is holding back its development.