Engineering
Engineering B.C. #3: A.C. Okay, Smart Guy. What About the Code I Already Have?
I've had several requests and have been working on many other things, but folks are right: we rarely get to work on greenfielding a project. We all have existing projects we'd like to move forward, now with AI in the mix. How do we recover intent, distinguish architecture from historical accident, document what we've learned, and then give AI useful context?
To be totally transparent
Retrofitting AI into an existing project is not as clear-cut as any other process you may undertake. As a matter of fact, the process can be downright murky. A couple of key decisions up front will make the journey smoother, but as the map makers of olde used to say, "Here, there be dragons!"
There has been much discussion, including published white papers taking deep dives on doing exactly what you're asking. I have found some of the information quite practical, while other bits leave me shaking my head.
What I hope to help you with is stripping away the cruft so you can arrive at a brownfield AI project which has many of the features and documentation allowing these projects to move forward in an AI world where you're establishing processes, guardrails, architectural decisions, and establishing a CONSTITUTION.
Eventually, doing this for your applications makes it easier for everyone.

Your Codebase Doesn't Know Why
Now, at the risk of sounding stupendously ridiculous, I want to point out your codebase doesn't know why it exists, and your documentation may not either. Well, your documentation should illustrate the why, but it may not be as clear-cut as you'd wish. Some documentation absolutely does contain why; our problem is we may not know which documentation, how authoritative it is, whether it's current, or what was never documented at all. Combining this uncertainty with the digital smorgasbord of choices you and your team made over time is central to the archaeology.
It's like turning to your senior principal engineer, Bob, and asking, "Can we add the tax rules for this region to our app so customers' transactions are taxed properly?" but you cannot see how he processes that in his head.
It's the same for your code.
Maybe even less so.
But, what if we could add Bob's (and yours, and others) thought processes and design decisions alongside the code so everyone doesn't have to reach out to Bob every time this section of the application needs changing?
Just to reiterate
Up until this point, I have been giving guidance for greenfield development:
Understanding → Architecture → Context → Code
For an existing system, we're almost doing:
A.C. Code + Evidence → Intent Discovery → Architecture Recovery → Understanding
Then we can deliberately construct the forward path again.
This is a really interesting engineering problem. And it really isn't retrofitting. It is retro-understanding.
Can't cover it all in one post
This is not “How to use AI with legacy code.” What we want is to know how do we recover enough intent from an existing system to give it the engineering artifacts we wish it had accumulated all along and then resume the forward Engineering B.C. process.
Much like I continue to write about greenfield projects, brownfield projects require several different areas of focus, so this will be a series of posts intended to guide you through the process and help you to make decisions which will work for you and your projects. Here is what we'll cover:
- A.C. #1 — Okay, Smart Guy. What About the Code I Already Have? We're laying the groundwork for how we plan to move our brownfield projects to a place where Engineering B.C. can enter the equation.
- A.C. #2 — Your Codebase Doesn't Know Why. We'll go looking for the evidence: code, tests, Git, infrastructure, README files, Confluence, Google Docs, Jira, diagrams, incident reports—and Bob. We'll also establish an evidence model: Observed / Inferred / Human-confirmed / Unknown.
- A.C. #3 — Don't Rewrite History. Recover It. Here, we'll turn the evidence into artifacts: recovered architecture, retrospective/recovered ADRs, constraints, invariants, candidate constitutional principles.
- A.C. #4 — Bootstrap It Backward. This is the experiment. Take a real brownfield repository and actually do the thing:
Evidence → Archaeology → Recovery → Human Validation → CONSTITUTION + ADRs + Context → Bootstrap → Forward Engineering
We aren't starting by changing anything.
We're not migrating frameworks. We're not fixing Bob's questionable service boundary. We're not generating documentation and declaring it truth. We're certainly not turning an agent loose with instructions to “modernize the application.” We're not even going to pick tools.
Before we change the system, we're going to understand the system.
The repository isn't the application.
While the repo may contain all of the code and a nice (hopefully updated) README and CHANGELOG, the rest of the evidence you need is probably scattered all over creation. Some is in the repository. Some is in Jira. Some is in Confluence. Some is in an ancient PowerPoint. Some is in an incident report. And some exists only in somebody's head.
That gives us the first practical principle of A.C.:
Start with what you already have.
The documentation isn't the application, either.
We're not going to do a six-month documentation project before AI can help. We're not demanding everybody clean up Confluence or SharePoint. What we will do is take the low-hanging fruit first and let AI help identify what is missing and which missing things actually matter.
It's a very important distinction. We're not asking humans, “Please document everything you remember.”
We're trying to get the machine far enough to be able to ask Bob:
“These two services solve apparently similar problems in completely different ways. Was that intentional?”
And Bob says:
“Oh. Yeah. Funny story…”
Now we have something worth documenting.
We may need to conduct some interviews
And by "we", I mean us and our AI collaborators. But we need to give our collaborators as much context as we can and we have to determine what threads we want our AI interviewers to pull on.
Because when Bob says, "Funny story..." there are likely some gems to be dug up which could be critical to the project. We'll set up a Markdown template which you can use to kick off our interview process.
We will not modify existing code
Not yet, at least.
In many, if not most, cases, I expect our first pull request against the brownfield application might contain no application code at all.
What does the research show?
Oh, boy.
We aren't starting from scratch. Architecture recovery/software archaeology existed long before LLMs. Researchers were already confronting missing documentation, lost design knowledge and the insufficiency of source code alone. Towards Immersive Software Archaeology in Legacy Systems Using Feature Modeling for Program Comprehension and Software Architecture Recovery — 2004 IEEE ECBS
AI changes what we can practically do with all of the evidence.
The closest thing I've found to what we're attempting is:
Many of the concepts here fit our plan well and have been used as part of the process:
Put on some coffee and have a look. There is a lot here. I am also including these two publications, because they support and even defend, what we're doing. Especially handy when someone asks us what we're doing or planning to do.
Additionally, Using Feature Modeling for Program Comprehension and Software Architecture Recovery — 2004 IEEE ECBS revealed that source alone isn't enough and domain knowledge matters.
Also, AI-Assisted Documentation Across the SDLC — U.S. DoD/AI4SDLC, probably our best practical neighbor supports AI-assisted recovery from existing artifacts while keeping humans responsible for validation. It also reinforces our idea of gathering evidence rather than treating generated documentation as authoritative.
But you should know this....
I still have questions
As I have read and practiced retro-understanding of existing apps, there were some things I didn't find in the processes that I want to validate and explore.
- Can we make uncertainty operationally useful rather than merely recording confidence? Can the unknowns themselves become a work queue?
- Can those uncertainty signals drive AI-assisted interviews with the people who actually hold the missing institutional knowledge? I'm looking at you, Bob.
- How do we preserve provenance once human recollection becomes another evidence source—and how do we preserve conflicting recollections rather than reconcile them prematurely? GADR becomes very relevant here because it demonstrates both the usefulness of extracting decisions from messy human conversation and the danger of enrichment introducing material which may not be faithful to the source.
- When does a recovered observation deserve to become a governing architectural decision? Just because we observe a pattern doesn't mean it was an architectural decision. How do we determine when recovered knowledge should become an ADR, constraint, principle, or other governing artifact?
- When do we know enough? Complete understanding isn't the goal. How do we know when we've recovered enough trustworthy understanding to stop digging and begin deliberate engineering again?
The plot twist?
Our goal is simple: get the existing application to the point where Engineering B.C. can begin.
Because A.C. isn't Engineering B.C. in reverse. A.C. is the recovery process required to reach Engineering B.C.'s starting line.
A.C. Code + Evidence → Intent Discovery → Architecture Recovery → Understanding
B.C. Understanding → Architecture → Context → Code
That shared Understanding node is the hinge.
And, well... A.C. comes before B.C.