Ship or Learn

Every day ends in someone's hands, or with a named blocker

Everyone who has run delivery teams knows the feature that was "80% done" on Monday, "almost there" on Wednesday, and "just needs a few tweaks" on Friday. The status report said on track, because the work was still moving. Nobody could say whether it was right, because nobody had judged it against evidence. A hundred status meetings have been built on exactly that sentence.

The in-progress state is comfortable precisely because it avoids forcing a decision. As long as work is in progress, nobody has to declare it accepted and own the call, or rejected and name what is wrong. The ambiguity protects everyone from accountability, and it compounds daily until the Effort stalls.

Ship or Learn replaces the in-progress state with a daily forcing function. The cycle has two legitimate endings. Either something ships, or something is learned. There is no third option where work drifts overnight without a destination.

The unacceptable end to a day is "it is not quite there yet." The acceptable ends are: accepted and landed, not accepted with the failing criteria named, or not accepted because a decision is required that exceeds the Decision Owner's envelope.

Ship or Learn: three destinations from the daily cycle
Ship or Learn: three destinations from the daily cycle

Ship

The clearance multiplier: before and after clearance for a change class
The clearance multiplier: before and after clearance for a change class

The preferred destination is production, same day. Real users using real work in a real environment produce the highest-quality signal about whether the increment is right. Every intermediary, staging environments, review boards, release trains, dilutes that signal.

One clearance decision, worked. An internal servicing dashboard was classified as a change class in week one. Reversible, no external commitments, policy checks in the control layer, rollback rehearsed once in front of the Flow Council. Every increment on that Effort shipped same-day for the rest of the outcome, and nobody discussed deployment again. The payout calculation path on the same portfolio was deliberately not cleared. Those increments landed on the floor, in the hands-on environment, while their harness evidence accumulated into the case the compliance function needed for the next clearance conversation. Same model, two speeds, both of them chosen rather than suffered.

Getting there runs through a clearance test, applied once per outcome rather than per increment. Does the change class have an approved deployment path? Does the control layer of the harness cover the applicable policy constraints? Is rollback automatic and proven? Where all three hold, ship. Establishing clearance for a class of change is some of the highest-value Flow Council work there is, because one clearance decision converts weeks of future increments from queued releases into same-day ones.

The floor is a hands-on environment. Where same-day production is not available, the increment lands somewhere stakeholders can open it and use it themselves. Three requirements, none negotiable:

  • Realistic data. Synthetic data that looks nothing like the real thing produces synthetic feedback. The data must be close enough to production that a stakeholder can form a genuine judgment.
  • Reachable without a scheduled session. If a stakeholder has to book a demo slot, the environment is a presentation tool. The point is that they visit on their own schedule and form an independent opinion.
  • Current within one build day. An environment two weeks behind the work delivers feedback two weeks late, which means two weeks of increments were built without it.

And the floor rules some things out. Never a slide, never a status report, never a recorded walkthrough as a substitute. A slide describes the work. A report summarizes someone's interpretation of it. A video shows one path through one scenario chosen by the person who recorded it. None of them let a stakeholder use the thing, and using the thing is the only reliable way to surface a misunderstanding before it becomes expensive.


Or learn

The increment is not accepted, and that is not a failure. The word is chosen deliberately. The day produced something valuable: knowledge that changes the next cycle. What makes the learn branch legitimate rather than a euphemism is that it must end with three specific things.

The failing criteria, named. Not "it did not quite work." Which criteria, by number, and what evidence shows they failed. Precision here is what makes tomorrow's spec actionable.

The decision required, if one is. When the blocker is a decision above the Decision Owner's envelope, the exact question is written down. "Should we handle case X by method A or method B, given that B touches a policy constraint?" That goes to the Flow Council, which resolves it within hours or carries it to the Intent Council.

The redirect for tomorrow's spec. The learning changes the queue, as an amendment or a replacement, but always explicitly.

When the demo invalidates the plan and the new direction is not yet clear, do not force a build day. Run tomorrow as a spec-and-explore cycle: the fleet investigates, prototypes throwaway options, and gathers what the owner needs to decide. The day still ends with something in someone's hands. The artifact is a decision rather than an increment, and a cycle that produces a clear direction for the next five cycles is worth more than an increment built on a guess.

The Confirm, Amend, Replace rates from daily planning are the diagnostic here, and they are covered in the one-day cycle. The short version: if every demo overturns the plan, your slices are too big. A demo that adjusts tomorrow means you sliced correctly.

Two edge cases test the binary and neither breaks it. The multi-day increment, a data migration or an infrastructure cutover that cannot reach users in a day, is sliced by provable states instead of features. The day the rehearsal runs clean in the hands-on environment is a day that landed, because the evidence is in someone's hands even though the cutover is not. And the increment nobody can see, the backend contract with no screen, lands as something technical hands can use. A working endpoint a stakeholder's analyst can call, a dataset diff a domain expert can read. In someone's hands includes those hands.


The binary

It stops ambiguity from accumulating. Monday's "almost there" becomes Friday's "we need to re-scope," and a week of capacity is gone without anything a stakeholder could use. The binary forces clarity every day. There is no "almost."

It makes progress visible without reports. A Flow Council lead, a sponsoring officer, or a stakeholder can look at the daily record and see exactly what happened. Landed increments are in the hands-on environment. Named blockers are in the planning record. Status is read, not written.

It puts pressure on the spec. A vague spec produces a vague increment, a vague acceptance, and a day that ends in ambiguity. When every day must end in ship or learn, the spec has to be precise enough to produce that ending. The pressure travels upstream, which is where you want it.

Underneath all three is the same principle: activity is not value. A large volume of completed work is not evidence that an outcome changed. Something in a user's hands creates that evidence, and when landing is not possible, a named blocker or a resolved uncertainty creates evidence of a different kind. The learn branch has a strict condition, though. The learning must change the next decision. An interesting demo or another day of partial progress does not close the loop.


The failure modes

The carry-forward habit. Two or more consecutive days of reported progress without acceptance is the binary eroding. Split the increment. Land what is provable today, and carry the remainder into tomorrow's spec as a new increment.

The cosmetic red. Accepting on a red harness because the failures look cosmetic teaches the team the harness is optional, and that lesson propagates in weeks. Red harness, no landing. No exceptions.

The unescalated blocker. The team ends on a learn outcome but raises nothing to the Flow Council, so the blocker survives into tomorrow. Daily planning explicitly asks what the council must clear tonight. If nothing was raised and the team did not ship, something was missed.

The environment nobody visits. Landing in a hands-on environment only helps if stakeholders use it. If nobody has visited for a week, the feedback loop is broken even though the work landed. The correction is specific rather than exhortative. Each increment names the stakeholder whose week it touches, and the Decision Owner walks that person in once. After that they return on their own, because the environment is current and the work is theirs. Track engagement in the weekly calibration. See What to Measure.


Its place in the Model

In The AI-Native Operating Model, Ship or Learn sits at the landing stage of the daily cycle, the right side of the execution band, with three labeled destinations: Preferred (production), Floor (hands-on environment), and Learn (redirect). The green border reflects that landing is ultimately a domain-judgment decision by the Decision Owner.

Ship or Learn connects to:


Read the Ship or Learn white paper (PDF). Return to the Model to see how landing connects to the rest of the system.