How it works

From stated intent to a merged, verified pull request.

Every change flows through the same five stages, with a context gate before any code is written and an independent validation before anything merges. Gates are only ever skipped through explicit configuration.

01
Intent
prose · brief · recording

An idea, a research report — or a screen recording. shipd turns a recording into a cited brief, grounded frame by frame: every claim carries a timestamp, a speaker, and the frame that was on screen when it was said.

→ research/<slug>/report.md · video/<slug>/brief.md
02
Converge
codebase-first investigation

Agents investigate your codebase first, climb the question ladder for what's left, and compile intent into three reviewable artifacts: a plan with its decisions, testable requirement deltas per capability, and a mechanical task list.

→ plan.md · specs/<capability>/spec.md · tasks.md
03
Gate
deterministic · LLM-free engine

A context-sufficiency gate checks the plan against the real codebase before a line of code is written. Insufficient context parks the plan as rejected for human enrichment — the system never builds on a guess.

→ status: ready | rejected (with findings in plan.md)
04
Build & refute
orchestrator · executors · validator

The orchestrator designs on the strongest model; execution agents one tier down claim tasks atomically. Then an independent validator tries to refute every scenario in the spec against the real, running code. Refuted goes back; only confirmed moves on.

→ one change · one worktree · one branch · one PR
05
Ship & record
CI + semantic review · auto-merge

CI and a semantic review that must be explicitly dispositioned gate the merge. On ship, the deltas merge into a versioned capability library and the change archives — the system always knows exactly what it can do.

→ verified/<capability>/spec.md · completed/<date>-<change>/
the capability library compiles context for the next change↺
The question ladder

Most questions never reach you.

Before shipd asks you anything, it climbs a ladder: the codebase first, then the workspace wiki and personal memory store, then an "ask-first" oracle holding your standing positions. Only a question none of them can answer reaches a human — and once you answer, the answer is recorded so it's never asked twice.

01Codebase investigationanswers most▾
The plan reads before it asks: call sites, tests, config, and the verified capability specs from every change that already shipped. Most questions die here — the repository already answers them.
↓ escalates only if unanswered
The workspace wiki holds the job's recorded decisions; your personal memory store holds preferences that follow you across projects. Both are consulted silently. Every /s:teach adds coverage, so this rung catches more the longer you use it.
↓ escalates only if unanswered
An ask-first agent holding your standing positions, answering on your behalf with citations. Every consultation is recorded in the plan as a Q<n> ledger entry — and /s:teach <change> Q<n> replays one so you can correct a position it got wrong.
↓ escalates only if unanswered
Only a question no lower rung can answer reaches you. Your answer is written back into the wiki and the oracle's positions — so the same question never climbs the ladder again. The goal: a path you're needed on less and less.
every answer you give teaches a lower rung — over time, less escalates to you
Division of labor

Three roles, tiered by model strength.

ORCHESTRATOR
Plans and designs
Runs on the strongest model. Owns investigation, the spec, and the architecture of every change.
EXECUTION AGENTS
Claim and implement
One tier down. Claim tasks atomically from the task list and implement against the spec — never the transcript.
VALIDATOR
Tries to break it
Independent and adversarial. Attempts to refute every scenario in the spec against the real, running code before anything merges.
every change:worktree→branch→PR→CI→semantic review (explicitly dispositioned)→merged→capability library
Autopilot

Autopilot: drive a full epic.

One human approval — the epic — and the autopilot drives every member to a shipped PR, in risk-ascending order. Unattended doesn't mean unguarded; every member passes the same gates a hands-on change does:

the oracleopen decisions are answered from the wiki and memory; unanswered ones proceed on the oracle's recommendation and are parked as questions for you — the run never blocks.
/s:gatethe deterministic context gate still decides whether a member builds at all — one it finds lacking is parked rejected for you to enrich, and the run continues without it.
validatoran adversarial validator still tries to refute every scenario against the running code before a member can merge.
AI reviewCI plus the semantic review gate every PR, explicitly dispositioned on record — and a live board tracks the whole run.
the autopilot pipeline · per member, sequential, risk-ordered
1 · you approve the epic
the only human touchpoint on the happy path
2 · plan
3 · context gate
insufficient? side branch:
wiki → memory → oracle
still open → parked as rejected
you enrich → ↺ rejoins at plan
other members continue meanwhile
4 · build
execution agents claim tasks atomically
5 · refute
validator attacks every scenario
6 · ci + review
7 · merged ↺ next member
a parked member pauses only itself — the run continues with the rest and reports it
What can it do?

Everything it does, shown on one real product.

Every example on this page builds the same thing: Unkanny Banny, a kanban tool with boards, columns, cards and WIP limits. One repo, one team, one running product — so each command lands in the context of the one before it. First the full surface of what shipd can do, then how to state intent so a plan comes back right the first time, then the patterns for when something has already gone wrong.

the running example: unkanny-banny · a kanban tool · boards, columns, cards, WIP limits
The full surface

Everything it can do.

Twenty-four commands, six jobs. Skills run inside your agent session as /s:<name>; the shipd CLI reads the same library from your terminal and speaks JSON when you ask it to.

THE CORE LOOP
/s:plan <intent>
Plan a change
Reads your codebase first, climbs the question ladder for whatever is left, and compiles intent into reviewable spec artifacts. Stops for your review — it never builds from a guess.
→ plan.md · delta specs · tasks.md — status: draft
/s:build [change]
Build a change
Runs the context gate, checks the plan has not been superseded, then delegates every task to execution agents and lets a validator refute each scenario before CI and semantic review.
→ one worktree · one branch · one merged PR
/s:status [change]
Check a change
Reports a change's lifecycle position, validates its structure, or runs a guarded transition that asks before forcing past a guard.
→ draft → ready → active → complete → verified
/s:review
Review a diff
An AST-aware semantic review of local changes against a base ref: files grouped into cohorts, changed signatures chased to their call sites, findings rated high/medium/low.
→ findings by cohort · a ship-it or fix-required verdict
/s:fix <symptom>
Fix a bug
Retrieves the related specs, reproduces the failure, and repairs the code that drifted from documented behaviour — with a regression test. Never edits a spec, never opens a PR.
→ a reproduction · a regression test · a fix
CONVERGING ON INTENT
/s:research <question>
Research a question
Decomposes an open question into bounded parts, searches each, reads the strongest sources, and composes a report in which every claim carries a citation.
→ research/<slug>/report.md, fully cited
/s:video-ingest <file>
Plan from a recording
Transcribes a screen recording on-device, grounds each spoken claim against the frame that was on screen, and resolves contradictions by recency.
→ brief.md — timestamp · speaker · frame per claim
/s:epic <feature>
Decompose an epic
Breaks a feature too big for one PR into shared decisions, a design, and a table of member changes with complexity ratings — then stops. Members are planned one at a time.
→ epic.md + a risk-ordered member table
/s:initiative
Track an initiative
Authors a lint-clean outcome brief from a workspace-first interview, reports every initiative's requirement progress, or attaches one to an epic.
→ a brief · outcome progress · an epic tag
ASKING ONCE
/s:ask <decision>
Ask the oracle
An ask-first agent answers from your standing positions, with citations, before any person is interrupted. What it cannot answer is queued for a human — the run never blocks.
→ a cited recommendation, or a queued question
/s:teach
Teach the system
Distills this repo's spec artifacts and answered queue entries into the workspace wiki, interviewing you only on the gaps and contradictions the scan surfaced.
→ wiki pages + drained queue entries
/s:remember <preference>
Remember a preference
Captures a durable preference into the personal memory store through one confirmed write. Consulted by the question ladder on every future plan.
→ a memory page in the personal store
/s:memory · /s:forget
Browse or forget
List every preference the store holds, or remove one after a single confirmation — the read and delete counterparts to remembering.
→ the stored preferences, edited
WORKING AT SCALE
/s:autopilot <epic>
Deliver an epic unattended
One human approval — the epic — and every unplanned member is driven plan → gate → build → merged PR in risk-ascending order, one worktree and branch each. A parked member pauses only itself.
→ member PRs + a resume pointer per parked member
/s:workspace clone <url>
Stand up a workspace
Clones a portable workspace — manifest plus job wiki — and materializes its member repos on this machine via the sync ladder: worktree, reference-clone, or full clone.
→ a ready cross-repo job directory
VISIBILITY
shipd board
Watch delivery live
The full-screen delivery board: what is in flight, what is parked, what merged this week, and which agent is doing what right now.
→ the live board, or JSON for your own tooling
shipd list · shipd metrics
List and measure
In-flight changes across the repo root and every worktree, plus throughput metrics derived from the immutable archive.
→ human-readable or JSON-first output
shipd epic <slug>
Inspect an epic
An epic's status, metadata and per-member state — which members are stubs, which are planned, which already merged.
→ member states + the epic's derived status
SETUP AND GUARDRAILS
/s:onboard
Take the tour
A nine-step walkthrough you drive one step at a time. Steps 1–7 explain the artifacts over a throwaway example, step 8 builds it for real, step 9 hands you your first command.
→ a working example + your first command
/s:doctor
Diagnose the setup
A read-only preflight reporting ok / warn / fail per check. Proposes one remedy per remediable finding and runs only what you consent to, then re-runs and reports before and after.
→ a green preflight, or the exact blocker named
/s:gate
Install the review gate
Installs the managed files, requires the semantic-review check on the repository, and enables auto-merge — taking your consent before anything changes on GitHub.
→ PRs blocked until the review is dispositioned
shipd lint [change]
Validate the library
Structurally validates specs and change deltas: requirement grammar, scenario coverage, delta consistency against the capability library.
→ findings with file + line · an exit code for CI
Stating intent

How to ask for it.

Planning is the only step where your judgement is required. The gate, the task list and every scenario the validator later attacks are all derived from what you said here. Four habits separate a plan that comes back right from one that comes back with six questions.

01Describe the outcome, not the mechanism.
DON'T
/s:plan add drag and drop to the board
Nothing here says what dragging has to survive. The plan cannot write a single scenario from it, so it comes back with four questions before any work starts.
DO
/s:plan Let a user drag a card between columns on the Unkanny Banny board. The card's new column and its position inside that column must survive a page reload, and the existing keyboard reorder shortcuts must keep working.
Three testable outcomes, one stated constraint. The plan reads the board store, asks the oracle about persistence, and emits.
$ /s:plan Let a user drag a card between columns …
 
→ read app/board/column.tsx · app/board/card.tsx · lib/board-store.ts
→ read .shipd/verified/board-interaction/spec.md
→ oracle Q1 where card order persists → ANSWER (cited: wiki/board-store)
 
.shipd/planned/card-drag-between-columns/
plan.md status: draft
specs/board-interaction/spec.md +2 requirements · 5 scenarios
tasks.md 9 tasks
Why it works — Outcomes are testable; mechanisms are not. "Survives a reload" and "keyboard shortcuts keep working" each became a #### Scenario: the validator will later try to refute against the running board. "Drag and drop" would have become nothing at all.
02Say what must not change.
DON'T
/s:plan add WIP limits to columns
Silent on every board that has no limit. The safest reading is also the most disruptive one, and you will only find out at review.
DO
/s:plan Add a per-column WIP limit to Unkanny Banny. When a column is at its limit, dropping another card onto it is refused and the column header turns red. Limits are opt-in per column — boards with no limit set must behave exactly as they do today, and existing boards get none.
The non-goal becomes a requirement of its own, not an assumption.
$ /s:plan Add a per-column WIP limit to Unkanny Banny …
 
specs/board-interaction/spec.md
## ADDED Requirements
wip-limit-enforcement
#### Scenario: Dropping onto a full column is refused
#### Scenario: A full column header renders red
#### Scenario: Columns without a limit accept every drop ← the non-goal
Why it works — Non-goals are the cheapest thing you will ever write. Stated once, "boards with no limit behave exactly as today" becomes a scenario that fails loudly if a later change quietly breaks it. Left unstated, it is just something you hoped everyone assumed.
03Bring evidence, not adjectives.
DON'T
/s:plan the board feels slow
"Slow" is not a target, so there is nothing to verify against. The deterministic context gate parks this plan as rejected and writes the findings into it rather than letting an agent guess at a number.
DO
/s:plan Unkanny Banny's board takes 4.1s to first paint with 200 cards on a mid-tier laptop, and every card re-renders when any one card moves. Get first paint under 1s at 200 cards, and stop cards re-rendering when a sibling moves.
A measured before, a stated after, and a named mechanism. The gate passes.
$ /s:plan the board feels slow
→ gate spec_gate.py board-render-perf → REJECTED
· no measurable target — "slow" resolves to no scenario
· findings written to plan.md · status: rejected
 
$ /s:plan Unkanny Banny's board takes 4.1s to first paint with 200 cards …
→ gate spec_gate.py board-render-perf → PASS
#### Scenario: First paint under 1s with 200 cards
#### Scenario: Moving one card re-renders only that card
Why it works — A number is a scenario; a feeling is not. If all you have is the feeling, get the evidence first: /s:video-ingest turns a screen recording into a brief where every claim is pinned to the frame that was on screen, and /s:research turns an open question into a report where every claim is cited. Both flow straight into a plan.
04Split anything bigger than one PR.
DON'T
/s:plan add swimlanes, card aging and WIP limits
Three features, one branch, one unreviewable PR. Nothing stops you — but a failure in any one of them now blocks the other two.
DO
/s:epic Unkanny Banny board governance: swimlanes, per-column WIP limits, and card aging
Shared decisions recorded once, five members sized and risk-ordered, each planned and shipped on its own.
$ /s:epic Unkanny Banny board governance …
 
.shipd/epics/board-governance/epic.md
member complexity status
wip-limit-enforcement simple stub
column-header-warning simple stub
swimlane-grouping moderate stub
card-aging-indicator moderate stub
swimlane-wip-interaction complex stub
 
$ /s:autopilot board-governance
→ PR #341 wip-limit-enforcement merged
→ PR #342 column-header-warning merged
→ PR #343 swimlane-grouping merged
→ PR #344 card-aging-indicator merged
→ swimlane-wip-interaction parked · resume pointer written
Why it works — An epic records the decisions its members share exactly once, then hands each member its own worktree, branch and PR. The simple members merge while you are still thinking about the complex one — and the one that parks pauses only itself.
a good intent carries:the outcomewhat must not changeevidence, not adjectivesone PR's worth
When it goes wrong

Reach for the right tool.

Which tool you pick encodes a diagnosis. Reach for the wrong one and you will rewrite a spec that was right, or patch code that was only ever doing what the spec told it to. Nine things that go wrong, and what each one actually is.

▸ Cards land in the wrong column on touch devices.
REACH FOR/s:fix cards drop into the wrong column on touch
The behaviour is documented and the code stopped matching it — this is drift, not a design question. /s:fix retrieves the board-interaction spec, reproduces the drop on a touch viewport, writes the failing regression test first, then repairs the handler. It never edits a spec artifact and never opens a PR: it stops at the diagnosis, the fix and the evidence.
→ a reproduction · a regression test · a fix on your working tree
▸ A dropped card lands at the bottom of a column. It should land at the top.
REACH FOR/s:plan — not /s:fix
Nothing drifted here. The code does exactly what the spec says and the spec is what is wrong, so a patch would leave the library asserting the opposite of the shipped behaviour. /s:fix detects this case and hands off rather than quietly rewriting the contract. Plan it instead: the delta carries a MODIFIED requirement with a base: hash, and the validator refutes the new scenario before it merges.
→ a recorded correction, not a silent edit
▸ The gate parked my plan as rejected.
REACH FORenrich plan.md, then re-run /s:build
The context gate is deterministic and LLM-free: it compared the plan against your actual codebase, found something it could not resolve — an unreferenced file, an open decision, an unmeasurable target — and wrote the findings into the plan. Answer what it named and the change rejoins the pipeline at plan. Nothing was built on the guess.
→ status: rejected → ready, with the findings answered in place
▸ The validator refuted a scenario.
REACH FORthe fix loop, then re-validate
An independent agent attacked every scenario against the running code and one did not hold. A refuted scenario blocks verified outright — the finding routes back to the execution agents, they fix it, and the validator runs again from a clean context. If the validator was wrong, the scenario was ambiguous: sharpen the scenario, never the verdict.
→ nothing merges until every scenario is confirmed
▸ I answer the same question on every single plan.
REACH FOR/s:remember · /s:teach
Unkanny Banny keeps board order in lib/board-store.ts, and every plan asks about it. Record it once. /s:remember captures a durable preference into your personal memory store; /s:teach distills the repo's decisions and answered queue entries into the workspace wiki. Both rungs are read by the question ladder before anything is escalated to you.
→ a question answered in March does not come back in August
▸ A plan I wrote last week no longer matches main.
REACH FORnothing — /s:build checks this itself
Before a single execution agent is spawned, the build merges the base branch and runs the supersession check against your own masters. Content drift carries forward into the plan review and is reconciled there. A genuine supersession — a requirement id collision, or a merged change that already did this work — stops the build and asks whether to abandon or re-scope.
→ clean · content drift · superseded — only the third one stops
▸ A member of my epic parked overnight and I do not know why.
REACH FORshipd board · shipd epic board-governance
The autopilot parks a member rather than guessing at a resolution, and a parked member pauses only itself — the rest of the run carries on and reports it. The board shows every member's state alongside the resume pointer written for the parked one, so you pick it back up exactly where it stopped rather than restarting the epic.
→ the resume pointer, and the run's other PRs already merged
▸ The PR is open, CI is green, and nothing is merging.
REACH FOR/s:review · check the semantic-review status
The gate holds the merge until the semantic review is explicitly dispositioned — every finding either implemented or cleared with a stated reason. There is no rubber stamp to apply. Run /s:review locally against the base ref before you push and the gate has far less to say when the PR opens.
→ findings by cohort · a ship-it or fix-required verdict
▸ Something about my setup is off and I cannot tell what.
REACH FOR/s:doctor
A read-only preflight reports ok, warn or fail for each check — plugin version, tokens, the review gate, the workspace layout, the spec library's structure. It proposes one remedy per remediable finding, runs only what you consent to, then re-runs the preflight and shows you before and after.
→ a green preflight, or the exact blocker named

Try it on your next change.

GET STARTEDREAD THE FAQ