KAMRAN QADRI.← All notes
Field note · AI-native engineering

A weekend with agents changed how I think about teams.

I've been the technical lead for Namazee for more than four years. It helps people find nearby masjids and their jamat times, and it runs on a dataset of over a thousand masjids that dozens of area volunteers keep up to date. In those years I've made the product and technical calls, written parts of the app myself, managed freelancers, and reviewed a lot of their work.

So I know how long things take on this product. I know where the time actually goes.

When we planned the next phase, the mobile app was the problem. It sits on an old React Native version that's hard to upgrade because everything depends on everything else. A few years ago I would have upgraded it piece by piece. This time I proposed rebuilding it in Flutter, because I suspected AI agents had made a rebuild cheaper than I was used to assuming. I wanted to find out if that was true, and I had a weekend.

What I was trying to find out

Most AI coding examples are small: a simple landing page, a simple app, a bug fix. I wanted to know something harder. Could agents move a real product with a live app, a NestJS backend, an admin panel and years of history across several codebases, without losing the product decisions, the architecture, or the documentation along the way?

I also set a limit. The current app is the baseline. Flutter gives us a better foundation; it doesn't give an agent permission to redesign whatever it likes.

Before starting, I estimated the same scope done the way I've always done it on this product: upgrading the backend, tests and security work, a new admin panel, mobile design, the Flutter app, and the planning around it. My rough number was more than three months for one engineer. It's a gut estimate from years on this codebase, not a benchmark, but it gave me something to compare against.

Four repositories, and one of them holds the product

Three repositories hold code: backend, admin, and the Flutter app. The fourth holds the product: PRDs, decisions, designs, plans, and a description of what the app actually does today.

I built that description first. I had agents read the existing codebases and write down what is really shipped, because the old PRDs describe what someone meant to build, and the code shows what users actually got. The two had drifted in places.

Then I worked in the product repository and asked for the next phase as three coordinated plans, one each for backend, admin and mobile. The plans reference each other. The backend plan says which of its pieces unblock the Flutter app. The admin plan says which screens have to wait for backend work.

After that, I opened each code repository and told its agent to read its plan and build against the real code. When an agent finds the plan is wrong, it doesn't quietly decide something different. It writes a handoff back to the product repository. The first round of that happened almost immediately: the Flutter agent reported that its own brief contradicted itself, and the backend agent came back with a few defects and steps to reproduce them. Both were checked and recorded in the product repository instead of being quietly worked around.

Plan centrally. Execute locally. Hand changes back. Conceptual. A product repository holding PRDs, shipped reality, designs, decisions and coordinated plans sends plans to three implementation repositories: backend, admin and Flutter, each with its own agent. Each repository hands discoveries back to the product repository. Backend work unblocks admin and mobile work. HOW I COORDINATE AGENTS ACROSS ONE PRODUCT Plan centrally. Execute locally. Hand changes back. PRODUCT REPOSITORY The source of intent PRDs Shipped reality Designs Decisions Coordinated plans PLAN HANDOFF PLAN HANDOFF PLAN HANDOFF Backend Admin Flutter app own agent · own code own agent · own code own agent · own code IMPLEMENTATION TRUTH IMPLEMENTATION TRUTH IMPLEMENTATION TRUTH BACKEND WORK UNBLOCKS ADMIN AND MOBILE Conceptual workflow · Namazee rebuild KAMRAN QADRI.
The product repository holds intent; each code repository holds what's really built. Plans go down, handoffs come back up.

This is the part I'd keep even if every tool changed. In my experience, docs and code slowly drift into two different products. The handoffs keep pulling them back together.

Designing before building

I didn't start the mobile work with “build me a Flutter home screen.” I designed every screen in Pen.dev first and reviewed the flows before any Flutter code existed. So the implementation agent had a design, a plan, and clear rules for what counts as done for each milestone.

Agents write code fast, and if the direction is vague they build the wrong thing fast. The cheaper the code gets, the more I want the design and the plan settled before building starts.

Where it stands

It isn't finished. On Sunday evening I still thought I'd have a working first version by the end of the day; I didn't.

So no, AI didn't build the product in a weekend. What surprised me was everything around the building.

The waiting disappeared

For years a Namazee change went like this. I'd define it and find a freelancer who had time. Then I'd explain the context, wait for the work, and review it. We'd go back and forth on revisions, and if the change touched the backend and the app, that meant another person and more waiting. I put in part-time hours every week, and a lot of those hours went to coordinating rather than building.

The code itself was rarely what held us up. Someone was busy that week. A decision needed correcting after it was already built. Feedback sat until the next round. A new person needed the whole context again, and when someone left, what they knew left with them.

A lot of that waiting went away. An agent picks up the next step as soon as the last one finishes. It reads the plan and the code, runs the tests, and updates the docs. Nobody has to schedule a call on Tuesday to explain why the backend changed on Saturday.

The engineering work was all still there. There was just very little waiting between the pieces.

What kept the agent on track

The work ran long enough that I filled an entire very large context window. The agent compacted it and kept going deep into a second one. After compaction it didn't act like someone joining the project that morning. It didn't forget the project.

I think that's because the important project knowledge didn't live in the conversation. It lived in KIS, the small project-memory method I built and use on my projects:

When a session starts, a hook loads the top of the state file. Before doing anything, the agent knows the branch situation, the task in progress, how it will verify that task, and the known defects. It reads deeper files only when it needs them.

It isn't perfect. Only the top of the state file loads automatically, so sometimes I had to point the agent at the rest. What I care about now is less the size of the context window and more whether an agent can recover its context when a conversation ends.

I'm not using VS Code anymore

My day-to-day workspace is now super.engineering on macOS and Orca on Linux. Both let me run several agents side by side, read their code and docs, and edit when I need to. Claude Code with Opus does most of the heavy implementation. When work gets split up, I let the lead agent decide whether a subtask needs Opus or whether Sonnet is enough.

The Linux side is new too. Around the same time I set up Omarchy. It's fast, and every part of it is configuration I can read and change: I added a Raycast-style launcher and a workspace overview like Mission Control. For the first time my operating system feels like open-source software I can configure however I want. When I want something to work differently, I can find where it's configured and change it. That deserves its own note.

My loop used to be: open the editor, change files, run the code. Now it's closer to: shape the problem, design it, plan it, hand it to agents, read what comes back, use the app, correct, update the state. I still edit files. That's just no longer where most of my time goes.

More agents didn't mean better

I tried adding reviewer sub-agents that check an implementation against its spec. It was expensive. A fresh reviewer starts knowing nothing, so it reads the spec, the state, the repository, the changed files and the tests before it can say anything useful. Most of its effort went into catching up on what the main agent already knew.

I still want independent review. But a reviewer should get a small package: the requirement, the decisions already made, the changed files and the test results. It shouldn't have to read the whole project. I haven't built that yet.

For now my review is simple. The agent runs the tests, I use the actual app, and sometimes I read the diff. Testing on a real device isn't optional. Agents make it very easy to produce code nobody has actually checked.

What it cost

The agent workspace reported hundreds of agents spawned, days of combined agent time, well over a billion tokens, and an estimated model cost of over a thousand dollars.

Those numbers need context. The cost is the tool's estimate, not an invoice. Almost all of the tokens were cached context being re-read, not new work. And I was pushing hard on purpose.

Still, the work wasn't free; the cost moved. I used to pay in waiting, availability and handoffs. This weekend I paid a lot of that in compute instead. Compared with a small side project, that's expensive. Compared with months of work across several people, it looks very different. One weekend doesn't settle it, but it's worth measuring properly.

What changed in how I think about teams

I already believed AI would make engineers more productive. What changed this weekend is how much I think one strong engineer can take responsibility for.

That doesn't mean one person replaces a team. The agents needed me for exactly the things I'd expect: deciding what to build, the architecture, whether a design is good, which trade-off is acceptable, turning a vague goal into a plan, and noticing when the plan had to change. I spent more of my time on those this weekend, not less.

What changed was how much building I could direct at once, and how little of it waited on anyone. That matters for how teams are shaped, how we estimate, and who we hire. Giving everyone an AI subscription won't get you there on its own. The plans, the state and the checks around the agents made the difference for me.

What I want to work on next

My setup is still more manual than I'd like. I decide what to build, approve the plan, start the work, watch what's blocked, check the result, and start the next plan.

The system already knows most of what that takes: the current state, which plans are done, what's blocked, and how the repositories depend on each other. So the next question for me is how much of that sequencing can run on its own. Could it notice a backend piece has landed and start the Flutter work that was waiting for it, then only come to me when a real decision is needed?

I don't want autonomy for its own sake. I want less waiting, without the agent making product decisions that should be mine.

I wrote earlier about orchestration becoming the bottleneck in a different multi-agent system. This weekend I ran into the same problem again, this time on a real product.

This is the first of a few notes on how I'm changing the way I build with agents. Next: the product-repository setup in more detail, KIS and long-running context, and what agent-heavy work really costs.