A team can use Codex or Claude Code to produce the implementation, its tests, and the completion summary in one workflow. That workflow does not give the team independent evidence that the code is safe to merge.
The specific risk is correlated evidence. If the implementation, the tests, the documentation, and the summary all come from one workflow and share the same wrong assumption, they agree with one another while the checks stay green.
This article is for CTOs, architects, and engineering leaders moving coding agents from individual experiments into regular production delivery. It explains how the failure happens, then gives you repository rules, a pull-request template, an independent-review prompt, and isolation commands that you can apply in your own repository.
We developed these controls while building revsko, the sales follow-up product we build and run. Coding agents implement schema migrations, API handlers, tests, and many review comments in the repository. One founder runs the development loop, and there is no separate QA team: the controls in this article are how we own quality.
In one week in July, three failures exposed the limits of green checks and plausible summaries:
- A test continued to pass after we deleted the feature it claimed to protect.
- Two concurrent agent sessions changed each other's test results through a shared database.
- Independent review found confident coverage and completion claims that the repository did not support.
These failures looked like evidence of success: passing tests, professional documentation, and confident completion summaries. The six controls below make those signals easier to challenge before a change reaches production.
What to notice: none of the six controls makes the agent smarter. Each control makes an unsupported claim easier to detect. On a phone, scroll the diagram horizontally to read every control.
Set up a minimum verification system
Start with one release-critical behavior. Good candidates are authorization, billing, tenant isolation, consent, and irreversible state changes. Do not try to redesign the entire development process in one pull request.
The complete agentic coding verification kit contains the four files below. You can copy the files separately or use the inline versions in this section.
Step 1: add a completion contract to the repository
Paste this block into the repository's existing AGENTS.md, CLAUDE.md, or equivalent instruction file. Keep the repository's real build and test commands beside it.
## Completion contract
For every behavior-changing task:
1. Restate the acceptance criterion as an externally observable result before
editing code.
2. Identify the production entry point that must produce that result.
3. Identify or add a test that reaches the production entry point. A test that
writes the expected database state directly does not prove the application
produces that state.
4. Do not delete, skip, weaken, or replace an assertion only to make a check pass.
Report any required test-contract change explicitly.
5. Run the focused test, then run the repository's required local gate.
6. Record the exact commands and observed results. Do not replace command output
with “tests pass.”
7. Report gaps, skipped checks, shared-state dependencies, and unverified claims.
For authorization, billing, tenant isolation, consent, and irreversible state
changes, the implementation session cannot approve its own evidence. An independent
reviewer must observe the responsible test fail after a deliberate, reversible
break and pass after restoration.
This contract changes the definition of “done.” A completion summary is now an index to evidence, not evidence by itself. Download the repository rule block.
Step 2: require evidence in the pull request
Add these fields to .github/PULL_REQUEST_TEMPLATE.md:
## Agent-written change: acceptance and evidence
- Acceptance criterion:
- Production entry point:
- Test that reaches that entry point:
- Release risk if the behavior is wrong:
### Negative control
- Deliberate, reversible break:
- Focused test command:
- Expected failing assertion:
- Observed failing output:
- Restoration performed:
- Observed passing output after restoration:
### Full verification
- Repository gate command:
- Observed result:
- Independent reviewer:
- Unverified or skipped checks:
- Documentation changed or removed:
The fields force the author to connect a business behavior to production code, a responsible test, and an observed result. Download the full pull-request template, which includes an approval checklist.
Step 3: open a separate reviewer session
Do not ask the implementation session whether its own claim is true. Start a new Codex or Claude Code session and give it a verification task instead of a general code-review task. Run the first review without allowing edits.
You are the independent verifier for an agent-written change.
Inputs
- Mode: READ_ONLY
- Acceptance criterion: <one externally observable behavior>
- Pull request or revision: <PR URL, branch, or commit SHA>
- Focused test command: <the smallest test command for this behavior>
- Repository gate command: <the required local CI command>
Do not edit, create, delete, move, format, stage, commit, or restore files. Do not
trust the pull-request description, test name, or completion summary as evidence.
Name the production entry point, show whether the responsible test calls it, and
propose one small negative control. Do not perform the negative control in this
mode.
Return: VERDICT, CLAIM MAP, NEGATIVE CONTROL, RESTORATION, UNVERIFIED CLAIMS, and
RECOMMENDATION. Use VERIFIED, REFUTED, or INCONCLUSIVE for the verdict.
Use the full independent-review prompt. It contains READ_ONLY and NEGATIVE_CONTROL modes, mutation safety rules, and an exact response format. Review the proposed break before you run the second mode.
Step 4: give the reviewer a disposable environment
For a GitHub pull request, replace 123 with the pull-request number and run:
git fetch origin pull/123/head:refs/remotes/origin/pr/123
git worktree add --detach ../verify-pr-123 refs/remotes/origin/pr/123
cd ../verify-pr-123
git status --short
git status --short must produce no output before the reviewer changes a file. If the repository uses Docker Compose for tests, start an isolated Compose project:
docker compose -p verify-pr-123 up -d --wait
docker compose -p verify-pr-123 ps
The Compose project name isolates containers, networks, and named volumes. It does not isolate fixed host ports or an external database. Assign separate ports and a separate database before concurrent runs. The worktree verification guide includes cleanup commands and common test commands.
This minimum system is intentionally narrow. It verifies that one named test protects one named production behavior. It does not replace security review, threat modeling, load testing, or human approval.
Control 1: write specifications a machine can check
Every feature starts with a specification that contains numbered success criteria. Each criterion must describe an observable result that a test or review command can verify.
A machine-checkable criterion is only as strong as the test that claims to check it. One of our specifications required every advertised status value to have an application path that produced the value. The associated test wrote the expected value directly to the database, so it could pass even when the application path was absent.
The specification is also why a short prompt can produce a complete feature. The prompt can stay short because the engineering decisions already exist in the repository. Without those decisions, the coding agent must invent them during implementation.
Control 2: isolate each agent session
Parallel sessions increase delivery capacity, but they can also change each other's results. Git worktrees isolate source files. They do not isolate databases, ports, test data, or background processes.
Our test helpers reused fixed tenant names in one shared database. Because the tenant identifier was globally unique, a concurrent test or residual row could cause a unique-key collision. Some tests appeared reliable only because another test happened to remove the shared data first.
During the same week, one session left a test process connected to the shared database while another session ran the local CI gate. The second session reported an unrelated 401 response. Two intermediate failures had been suppressed, which made the shared database difficult to identify as the cause.
Use unique test data for each run. Make cleanup failures visible. Give concurrent sessions separate databases and processes when shared infrastructure can change the result.
Control 3: make the local verification gate truthful
Before a pull request, our local CI script runs the same test tiers as the remote gate. That local gate is useful only when its configuration matches the system it verifies.
Our developer documentation specified a ten-minute timeout for an integration suite that needed nearly fifteen minutes. A coding agent that followed the documented limit would terminate a passing run and report a failure.
When agents execute repository instructions, those instructions become part of the development system. We now record the source and capture date for measured values, and we use automated checks where configuration can drift.
Control 4: review completion claims separately from code diffs
A coding-agent change contains at least two types of evidence: the code and the description of what the code does. A diff review does not automatically verify the description.
In one review cycle, our review agents found a lock leak in the test tooling and several false statements about the code, including a coverage claim with no supporting test.
Our review system uses agent roles that are versioned in the repository. Some reviewers inspect one failure class, such as production paths, data integrity, or schema contracts. A separate refuter tries to disprove each review finding before a human acts on it. Ground-truth commands search the repository at a pinned revision instead of trusting the implementation notes.
One review found a test that continued to pass after we deleted the feature it claimed to protect. The reviewer challenged the coverage claim, a repository search showed that no test called the application path, and a controlled deletion confirmed the gap.
Control 5: store reproducible output, not only conclusions
Agent sessions end and context windows reset. A later session can use a lesson only when the repository contains enough evidence to reproduce it.
Each completed change therefore leaves an implementation note that records where the implementation diverged from the specification. Each review also records which agent role raised a finding and whether the team confirmed or refuted it.
Store commands and outputs with the conclusion. When we deliberately break a feature to verify a test, the pull request keeps both observations: the expected failure while the feature is broken and the passing result after the feature is restored.
We adopted this rule after an implementation note claimed that an application path was covered by the existing workflow tests. The statement was specific and false. A future agent could not verify the claim without rerunning the search and test.
Control 6: delete instructions that stopped being true
The other five controls generate specifications, implementation notes, review findings, and operating instructions. Those files become harmful when the underlying system changes but the text does not.
Our documentation system gives each session a small set of required files, stores each fact in one canonical location, and checks generated references for drift. When an instruction becomes false, we remove or replace it instead of adding another corrective note.
We adopted this rule after manually calculated repository statistics became stale. A specific number can look verified after the source has changed. When a number affects a decision, we record the command, commit, and capture date that produced it.
What the control system costs
A major specification can require substantial design work before an agent writes the first implementation line. Independent review requires additional runs, and the local gate adds time to each verification cycle. This approach is slower per feature than prompting an agent and merging the first plausible result.
The controls let us sustain a high change volume without treating a green test, a plausible note, or an agent's confidence as enough proof. They also reduce dependence on one session remembering what another session learned.
Update, September 2026. Here is one recent week, for scale. From 21 to 27 September 2026, AI coding agents working for me merged 266 pull requests into the revsko codebase. Almost half of the new code written that week was tests. The automated checks ran 325 times and stopped the work 66 times. A separate AI reviews every change, the tests must pass, and I decide what goes live. That week, before launch, an agent also checked every written rule of a feature against its tests, found one with no working proof, and wrote a test that exposed a real bug in what downgraded teams could still do in the app. We fixed it the same day, before any customer saw it. At that volume, the checks have to be able to say no. (The 266 changes were merged, not yet deployed.)
Adopt the controls in this order
Do not install all six controls at once. Start with the failure that already affects your team:
- Reproduce the failure and run an explicit local gate. Keep the failing output before changing the implementation. Do not weaken an assertion to make a broken behavior pass.
- Use unique test data. Add this control when concurrent or repeated runs can reuse the same records, ports, or identifiers.
- Review claims independently. Add a reviewer or refuter when pull-request volume exceeds the team's ability to verify every completion statement.
- Isolate session infrastructure. Give each concurrent session separate databases and processes when worktrees no longer provide enough isolation.
- Retain reproducible evidence. Store the commands and outputs that support release decisions.
- Remove stale instructions. Give each fact one canonical location and check values that can drift.
The smaller pull-request gate
If you cannot add the complete template yet, use this smaller checklist for agent-written changes that affect authorization, billing, tenant isolation, consent, or irreversible state:
Acceptance criterion:
Production entry point that implements it:
Test that exercises that entry point:
Deliberate failure or negative control:
Expected failing assertion:
Observed failing output:
Observed passing output after restoration:
Independent reviewer:
Documentation changed or removed:
The completed fields give a reviewer evidence to inspect. Empty fields show where the pull request still depends on an unverified statement.
If you want this implemented in your repository
Use the downloadable kit first. It is enough to test the method on one release-critical behavior without hiring CoEdify.
If your team is adopting coding agents in a production codebase and would rather have these controls set up with you than assembled from blog posts, this is CoEdify's agentic development service: we install the system in your repository, on your stack, and prove it on one real deliverable. CoEdify is the engineering company behind revsko, and this is the system we run on it.
The engagement starts with one scoped two-week phase ending in a reviewable working deliverable against the agreed scope. If that first phase does not satisfy you, you do not pay for it. Book a scoping call.