Crinaro.AI · horizontal CRINARO.AI

Same request, same answer

How you scope a team used to be a question about coordination cost. It now decides whether you can hand that team a specification at all, and a team you can only prompt gives you nothing to test the answer against.

The argument

In an estate other teams already own, a change that crosses their systems cannot be specified by one team alone: the spec would describe things it does not own. A team scoped to something it owns can be handed a spec.

That is not a matter of taste. A prompt is a conversation, and the answer can differ every time it is run, with nothing to check the difference against. Systems a business depends on need the same request to land somewhere you can test.

For a long time the argument about how to shape delivery teams was an argument about coordination. Split people by feature and one team carries a change from end to end, paying for it in breadth. Split them by component and each team goes deep, paying for it in handoffs. Both sides had evidence, and both were arguing about the same thing: what it costs for one group of people to get something from another group of people.

That argument assumed a person on each end. The receiving team could be told roughly what was wanted and work the rest out, because working the rest out is what people do. It is why the handoff was expensive, and it is also why it was survivable.

What an agent team can be given

An agent team is scoped to what it can be pointed at, and the question is whether that scope has an edge. A team whose remit is everything a request might touch has nothing to write a specification against. Not because breaking the request into parts is impossible: people do it every day. Because there is no reason to. A team that owns the whole surface can always absorb the work, so the decomposition happens only when a boundary makes it happen. What gets handed over instead is a description of intent, refined by asking again.

This is not an argument about vertical against horizontal. A team aligned to a stream of work, owning one bounded context end to end, has an edge and passes the same test. The failing case is not a slice through the stack. It is a remit with no boundary anywhere in it.

A team scoped to something it owns has a request with an edge. It can be written down, against a contract, and handed over.

Scoped without a boundary the request has no edge Everything a request might touch Nothing to specify against Variation with nothing to test it Scoped to something it owns the request has an edge One contract it is answerable for You can write it a specification Variation inside something testable The difference is not speed. It is whether there is anything to check the second run against.

Why that is not a preference

Prompting is the right way to explore, and it is where the work starts. What it does not give you is a result you can check. Ask a second time and you get a second answer, as reasonable as the first and different from it, and nothing in the exchange tells you which one to keep.

Runs vary for reasons that have nothing to do with scope. The model moves under a stable name, sampling is not deterministic, and the same request can be served differently depending on what else the machine is doing. Drawing a boundary does not fix any of that, and a note on this site already says so: the method in the piece on specifying starts by running two identical arms, which is only worth doing because two identical arms differ.

So the claim here is not that a specification produces the same output twice. It does not. The claim is narrower and more useful: a bounded request constrains the space an answer can land in, and a constrained space is what makes a test mean something. A contract can be replayed against whatever came back. Intent cannot.

Variation is fine while you are exploring. It is not fine in a system a business runs on. Enterprise systems earn their keep by being predictable: the same input produces the same behavior, a release does what the last release did except where it was meant to change, and a customer relying on it can plan around it. That is what quality means at that scale, and it is built out of repeatability rather than out of any single good answer.

So the question of how you scope a team stopped being an organizational preference and became a question about whether your delivery is repeatable. Not how fast the team is. Whether the same request lands somewhere you can test, every time it is run.

The objection, and why it does not hold

The obvious answer is that a party who does not own something cannot write a good specification for it: it will miss the invariants, the failure modes and the couplings that only the owner knows, so the work satisfies the request and does not work. If that were disqualifying, contract-first development would never work, and it does.

But it works for a reason worth being precise about, because the reason is the answer to the objection. A contract is agreed by both parties and then verified against the real thing, every build. The requester writes what it needs; the owner says whether that is what the component does; the test replays it. Nobody in that practice believes a specification thrown over a wall is sufficient, and nothing here should be read as proposing one. A contract never fully describes behavior either. Somebody ends up depending on something it did not promise, which is an argument for the owner being in the loop rather than against writing it down.

When a specification is wrong, the cause can sit upstream of the team receiving it, in analysis that was not done. That was true before agents. What has changed is the cost of being wrong: the turnaround is hours, so a specification that missed something gets corrected and re-run rather than absorbed into a release nobody wants to reopen. The argument for keeping the work inside one team was that going outside it was slow. That is the part that moved, and it is worth saying that it moved for the writing rather than for the agreeing. Getting another team to care is the same problem it always was.

And the specification is not overhead you pay for the privilege of splitting the work. The team receiving it has to update what its component says about itself, and the specification is what that update is written from. A request that never got written down leaves nothing behind for the next one to read.

What to call the unit, which matters more than it sounds

Component, service, module, capability: organizations fill these words in differently, which is most of why arguments between delivery frameworks talk past each other. Two people can agree completely about features and disagree entirely about what a feature is.

So the unit here is defined by a test rather than by a word. It is whatever satisfies two conditions at once: one team can own what it says about itself, and a specification can be written against it without describing anything else. If those hold, the name does not matter. If they do not hold, no name will fix it.

What is not known. That a bounded request produces output which varies less, or varies more usefully, than an unbounded one is stated here as a mechanism and not as a measurement. Nothing here has run the comparison. It is a real experiment: the same request given to a broadly scoped team and to a narrowly scoped one, several times each, with the runs compared to each other rather than to a target. Until somebody does that, the honest version of this note is that the bounded request is the one you can write a test for, which is a claim about what is checkable rather than about what varies.

Next

From prompting to specifying

What changes for an engineer once the spec, not the prompt, is the thing being edited.

Written by John Kelly. If a piece of this does not hold in your architecture, at your size, that is the mail worth sending: john@crinaro.ai. Today I read it myself.
The rest are in the notes.