An idea, a research report — or a screen recording. shipd turns a recording into a cited brief, grounded frame by frame: every claim carries a timestamp, a speaker, and the frame that was on screen when it was said.
Agents investigate your codebase first, climb the question ladder for what's left, and compile intent into three reviewable artifacts: a plan with its decisions, testable requirement deltas per capability, and a mechanical task list.
A context-sufficiency gate checks the plan against the real codebase before a line of code is written. Insufficient context parks the plan as rejected for human enrichment — the system never builds on a guess.
The orchestrator designs on the strongest model; execution agents one tier down claim tasks atomically. Then an independent validator tries to refute every scenario in the spec against the real, running code. Refuted goes back; only confirmed moves on.
CI and a semantic review that must be explicitly dispositioned gate the merge. On ship, the deltas merge into a versioned capability library and the change archives — the system always knows exactly what it can do.
Before shipd asks you anything, it climbs a ladder: the codebase first, then the workspace wiki and personal memory store, then an "ask-first" oracle holding your standing positions. Only a question none of them can answer reaches a human — and once you answer, the answer is recorded so it's never asked twice.
Every example on this page builds the same thing: Unkanny Banny, a kanban tool with boards, columns, cards and WIP limits. One repo, one team, one running product — so each command lands in the context of the one before it. First the full surface of what shipd can do, then how to state intent so a plan comes back right the first time, then the patterns for when something has already gone wrong.
Twenty-four commands, six jobs. Skills run inside your agent session as /s:<name>; the shipd CLI reads the same library from your terminal and speaks JSON when you ask it to.
Planning is the only step where your judgement is required. The gate, the task list and every scenario the validator later attacks are all derived from what you said here. Four habits separate a plan that comes back right from one that comes back with six questions.
Which tool you pick encodes a diagnosis. Reach for the wrong one and you will rewrite a spec that was right, or patch code that was only ever doing what the spec told it to. Nine things that go wrong, and what each one actually is.