All guides

Keyword: shipgate

A four-agent dev team that cannot fake its tests

Planner, coder, tester, reviewer. The popular version lets the agents decide when the work is done. Ours lets git and exit codes decide, and it caught our own team shipping a build whose test command did not exist.

Updated 2026-10-10 · paired skill no-fake-tests

Who this is for

You build sites or small apps for clients and you already use Claude Code. You have seen the four-agent setups going around (a planner, a coder, a tester, a reviewer) and you want one you can put your name on when a client's money is involved.

What you get at the end

What it costs

Build it

  1. Copy the no-fake-tests skill into ~/.claude/skills/.
  2. Make a job folder outside the repo, for example ~/jobs/12/. Handoffs live there: plan.md, changes.md, test-results.md, review.md. Never inside the repo, where a later job can read a stale one or commit it.
  3. Run the planner. It writes acceptance lines a test can check, one proof command, and a list of open questions. If the list is not empty, answer it before anything else happens.
  4. Record the base with git rev-parse HEAD, then run the coder against plan.md.
  5. Run the tester. Its only job is to break the coder's work. When it is done, lock its files: python3 scripts/ship_gate.py lock --out ~/jobs/12/locked.json.
  6. Failing tests go back to the coder, two rounds at most, with the locked tests unchanged.
  7. Run the gate: python3 scripts/ship_gate.py check --base <ref> --proof "python3 -m unittest" --locked ~/jobs/12/locked.json. Exit 0 ships. Anything else goes back with the reasons.
  8. Run the reviewer last. Its "ship" never overrides the gate's "needs work".

Prompts to copy

Planner:

You are the planner. Read the request below and the repo. Do not write code.
Write plan.md with: 1) acceptance lines, each one checkable by a test,
2) exactly one proof command that runs all tests, 3) OPEN QUESTIONS: anything
ambiguous that would change the plan. If OPEN QUESTIONS is not empty, stop there.
Request: {paste the client's request}

Tester:

You are the tester. Read plan.md and changes.md. Your only job is to break this.
Write real test files for edge cases, bad input, empty input, error paths and
anything in changes.md marked "attack first". Write test files only. Run the
proof command and write test-results.md with what failed and why.

Reviewer:

You are the reviewer. Read plan.md, the diff since {base}, the tests and
test-results.md. Flag at most 5 things that matter to the client, most serious
first. Verdict: ship, needs work, or blocked. You cannot edit files.

Sell it

You do not sell "agents". You sell a fixed-price build with a promise about quality you can prove. Our own price sheet: a one-page site at $1,250, a business site at $2,750, a cinematic build from $5,800, and care from $75 a month after launch. The four-agent pipeline is why those prices hold up on a one-person schedule.

Hi {name}, I build {kind of site/app} for {type of business} at a fixed price.
Every change goes through a separate tester that tries to break it, and nothing
ships until those tests pass. You get the code, the tests and a one-page report
of what was checked. Want to see one I did for a business like yours?

Where it breaks

Receipts

This is not a thought experiment. Our HQ's build department (four agents named Ada, Linus, Vera and Morgan) runs this pattern. Its log showed that 2 of 3 recent build proofs had failed, and one job had shipped with exit 127: the proof command was not found, and the old gate treated "nothing ran" as "nothing failed". That is why ship_gate.py treats 127 as a failure and has a test for it.

Prove it: python3 -m unittest discover -s tests inside the skill runs 10 tests on throwaway git repos, including the exit 127 case and edited and deleted locked tests.

Go further

free

Field Notes

Every guide on this site, free, no email wall. The weekly build note goes to the Substack list.

Read
$29

Skill Pack

The Claude Code skill behind one guide: SKILL.md, scripts, and the test suite that proves it works. Files you keep. Packed only if its tests pass.

Get it
$79

HQ Kit

Every skill pack plus the fenced crew starter: the guard hook, the role charters and the audit log layout HQ runs on.

Get it
from $3,500

HQ Install

Done for you: a fenced agent crew on your machine, wired to your tools, with one human gate. Same price ladder as Elevare's Ops Install.

Get it
from $1,250

Website Sprint

A fixed-price site for a local business, live in 5 business days, built with the same pipeline the guides teach.

Get it