From prompting to specifying
The shift is not that engineers write less code. It is that when the output is wrong, the thing you edit stops being the prompt and becomes the spec or the agent definition.
The argument
A prompt is a conversation you have once. A spec and an agent definition are files you change, version and re-run, so the fix applies next time and to everybody.
Which creates a problem the old way never had. You cannot tell from a single run whether the change did what you wanted, so teams need several branches going at once.
The first way to work with an agent is by prompting it. You ask, it answers, the answer is not quite right, so you ask again with more detail. It works, and it is the right place to start. What it does not do is accumulate. The understanding you built up getting to a good answer lives in that session and goes when the session does, and the next person starts where you started.
The shift is that when the output is wrong you stop reaching for a better prompt and start changing one of two things: the specification the work was written against, or the definition of the agent that did the work. Both are files. Both are versioned and reviewed like anything else. Both apply on the next run, to everyone, without anybody having to be told.
That is a different skill from prompting, and it is closer to something engineers already know. A bad output is a defect, and a defect has a cause that lives somewhere you can edit. The question stops being how do I ask for this better and becomes which file was wrong, the one describing what I wanted or the one describing who does it.
The problem this creates
Prompting answers immediately. Changing a spec or an agent definition does not, because what you are moving is how it behaves rather than what it said once, and one run tells you little. You want the old version and the new one on the same work at the same time.
How you run them depends on what you changed, and it is a fork rather than a ladder.
Change a spec and two working directories of one clone is right: one repository, two directories, already comparing against a shared base.
Change an agent definition and that shape tells you nothing. Two working directories are one machine, so everything at the machine level, your environment, your tools, your authenticated connections, is shared between the arms and cannot be varied. A change whose effect depends on any of it reads as no effect. You need two environments, two machines or two workspaces in a remote one, so that tier becomes something you can move. The bill is that a shared constant is now a per-arm variable somebody has to equalize on purpose.
Two things to do before you believe the result
Freeze the task before you write the change. A task picked afterwards is one the change happens to help, and nothing in either run will show you that.
Run an identical pair first. Same task, same procedure, no change between the arms. If two identical arms differ about as much as two different arms did, the method has no resolution on that task and any verdict from it is noise. Cannot tell has to be an answer you are willing to record, or you will only ever get the answer you went looking for.
Two smaller ones, and neither is a control. Run the arms close together and note when each ran, because the model can move underneath a stable name between them, and that is the one difference nobody can find afterwards. Read the results without knowing which is which, if you can. The timing note bounds how much drift could explain rather than removing it, and where one person writes the change and reads the results, reading blind is not available at all.
The question to keep asking
Is it producing the outcome you wanted? Not is the code good, and not was the run green. Whether the change you made to the spec or to the agent definition moved the work in the direction you intended, and whether it will keep doing that when somebody else runs it.
None of that is a way of knowing it did. It is a way of finding out, and the difference matters: it ranks two arms and measures neither, it is one run each at unknown resolution, and it is manual, so it cannot run on every change. I have not seen a verdict produced this way, mine or anybody's. What it does prevent is changing a file, liking the next answer, and calling that evidence.
Teams that get here need time for it. Refining an agent team produces nothing shippable that week, so it is the first thing traded away when delivery is late. It is also the thing that decides what every following week produces.