Who this is for
You build sites or small apps for clients and you already use Claude Code. You have seen the four-agent setups going around (a planner, a coder, a tester, a reviewer) and you want one you can put your name on when a client's money is involved.
What you get at the end
- Four roles with hard edges: the planner stops on open questions, the coder stops when blocked instead of guessing, the tester writes tests the coder is not allowed to touch, the reviewer can read but not write.
ship_gate.py, which says "ship" only when three things are true: a test file changed, no locked test was edited or deleted, and the proof command exits 0.- A delivery process you can describe to a client in one sentence: "nothing ships until the tests someone else wrote pass."
What it costs
- Your Claude plan. Each role is a Claude Code session or a
claude -pcall; no API keys. - Git and Python 3.
- On a small feature, expect four short sessions and up to two fix rounds.
Build it
- Copy the
no-fake-testsskill into~/.claude/skills/. - Make a job folder outside the repo, for example
~/jobs/12/. Handoffs live there:plan.md,changes.md,test-results.md,review.md. Never inside the repo, where a later job can read a stale one or commit it. - Run the planner. It writes acceptance lines a test can check, one proof command, and a list of open questions. If the list is not empty, answer it before anything else happens.
- Record the base with
git rev-parse HEAD, then run the coder againstplan.md. - Run the tester. Its only job is to break the coder's work. When it is done, lock its files:
python3 scripts/ship_gate.py lock --out ~/jobs/12/locked.json. - Failing tests go back to the coder, two rounds at most, with the locked tests unchanged.
- Run the gate:
python3 scripts/ship_gate.py check --base <ref> --proof "python3 -m unittest" --locked ~/jobs/12/locked.json. Exit 0 ships. Anything else goes back with the reasons. - Run the reviewer last. Its "ship" never overrides the gate's "needs work".
Prompts to copy
Planner:
You are the planner. Read the request below and the repo. Do not write code.
Write plan.md with: 1) acceptance lines, each one checkable by a test,
2) exactly one proof command that runs all tests, 3) OPEN QUESTIONS: anything
ambiguous that would change the plan. If OPEN QUESTIONS is not empty, stop there.
Request: {paste the client's request}
Tester:
You are the tester. Read plan.md and changes.md. Your only job is to break this.
Write real test files for edge cases, bad input, empty input, error paths and
anything in changes.md marked "attack first". Write test files only. Run the
proof command and write test-results.md with what failed and why.
Reviewer:
You are the reviewer. Read plan.md, the diff since {base}, the tests and
test-results.md. Flag at most 5 things that matter to the client, most serious
first. Verdict: ship, needs work, or blocked. You cannot edit files.
Sell it
You do not sell "agents". You sell a fixed-price build with a promise about quality you can prove. Our own price sheet: a one-page site at $1,250, a business site at $2,750, a cinematic build from $5,800, and care from $75 a month after launch. The four-agent pipeline is why those prices hold up on a one-person schedule.
Hi {name}, I build {kind of site/app} for {type of business} at a fixed price.
Every change goes through a separate tester that tries to break it, and nothing
ships until those tests pass. You get the code, the tests and a one-page report
of what was checked. Want to see one I did for a business like yours?
Where it breaks
- The gate proves tests exist, were not tampered with, and pass. It does not prove they test the right thing. Read the tester's file names against the acceptance lines yourself.
- Tests are found by file pattern. A project with tests in an unusual folder needs
--testsor the gate will say "no tests". - The proof command runs with your permissions. On code you did not write, run the coder and the proof in a sandbox or a throwaway clone.
- A planner with no open questions on a vague request is a red flag, not a green light.
Receipts
This is not a thought experiment. Our HQ's build department (four agents named Ada, Linus, Vera and Morgan) runs this pattern. Its log showed that 2 of 3 recent build proofs had failed, and one job had shipped with exit 127: the proof command was not found, and the old gate treated "nothing ran" as "nothing failed". That is why ship_gate.py treats 127 as a failure and has a test for it.
Prove it: python3 -m unittest discover -s tests inside the skill runs 10 tests on throwaway git repos, including the exit 127 case and edited and deleted locked tests.