I'm building a production-shaped quality platform for AI agents on AWS — Part 2

Part 1 ended with a promise: one command, and what comes out is governed by default. This post is the honest account of building that command and everything around it — including the milestones that closed red, the claim that failed, and the five claims that turned out to be unmeasurable in a repo this size.
Repo (MIT, make check runs offline with no AWS account): github.com/andaro74/beaconpave, tag m12. Building has stopped. What is left is a ledger, and the ledger is the deliverable.
Part 2 in three sentences
What it is. A miniature platform where an engineering team runs pave new, gets an AI agent that is already wired into a gateway, a tool registry, a golden test set and a CI gate that cannot be merged past — and where the rules that gate enforces live in plain data files owned by the business, not in code owned by engineers.
What is different. Most agent frameworks make it easier to build an agent. beaconpave makes it hard to ship an ungoverned one, and it does that by treating compliance, quality and adversarial resistance as properties of the road rather than a checklist a team completes before launch.
What it delivers. For an executive: a regulation change becomes a one-file edit by the seat that owns it, and every agent on the road goes red until it complies. For an engineer: you write behaviour and tests; you never write a gateway, a policy, a guardrail or an audit log, because the road already did.

The problem Part 2 solves, in one paragraph
Part 1 proved that an agent can be judged: a gateway it cannot bypass, a registry that makes unregistered tools unreachable, a gate that fails closed. But none of that helps a service team on Monday morning. They still have to find the gateway, learn Cedar, write a manifest nothing checks, and guess what "good" looks like. Part 2 is the answer to a different question: how does a team get onto the road without learning how the road was built? And underneath that: when the law changes, who changes what?
Two real-world comparisons that carry the whole design
A building permit, not a fire drill. A city does not ask each contractor to invent their own fire code. It hands over the code, the contractor builds, and an inspector checks against the code. In beaconpave the code is rules/, platform/registry/tools.yaml and the golden set; pave new hands it over; the CI gate is the inspector. The contractor's skill is the building. The code is not theirs to rewrite.
A thermostat with a lock on it. You change the temperature of a room by turning one dial — not by rewiring the furnace. In beaconpave there are exactly five dials that change what the gate enforces (rules, drill scenarios, the tool registry, thresholds and baselines, the adversarial corpus). Each has a named owner, and each needs a second key to turn. The agent's code is not a dial. You cannot make the gate kinder by editing the thing it is judging.

That second picture is the sentence I most want an executive to leave with: only by changing the business specs do the gates and rules change. Legal changes a rule file; AI Quality changes a threshold; Security changes what counts as an attack. Nobody changes the gate.
What is different from other agent platforms
| A typical agent framework | beaconpave |
|---|---|
| Makes an agent easier to build | Makes an ungoverned agent impossible to merge |
| Rules live in prompts and code | Rules live in data files with an owner, a source law, a control and a review-by date (G7) |
| "Tested" means the team ran it | Every milestone is scored against the same 25 golden cases and 10 probes, three samples each, append-only |
| A judge model grades the answers | The judge was measured against humans first, failed, and was demoted to advisory (claim 9) |
| Security is a review step | An adversarial pass counts only if a guardrail blocked or a policy denied and an audit record exists (G4) |
| A green dashboard | Pre-registered predictions and named falsifiers; a claim that cannot be falsified is refused, not celebrated |
| Success is measured by what shipped | Success is measured by what was proven; the final tally is five of twelve, and it is published as-is |
The last row is the unusual one, and it is deliberate. A platform that is not allowed to flatter itself will show fewer green boxes than one that is. That is a cost. I think it is the right cost for infrastructure that other teams will trust with compliance.
The architecture (what exists today)
Nothing in Part 1's diagram changed. Part 2 added the pieces around it:
pave newandpave verify(M05) — the scaffold and the verifier.pave new my-agent --brand meridian-sportsrenders five files: a manifest, a gateway client, an answer schema, a starter golden set and its README.pave verifyrefuses a scaffold fourteen distinct ways — a manifest key that is missing, a grant that does not match the registry, a golden set below the floor.- The rules registry (M09) —
rules/MER-AI-0001.yaml: a fictional AI-disclosure act, owned by the Legal/S&P seat, disposed into a seven-case eval pack the gate blocks on.pave rules trace MER-AI-0001walks law → rule → control → cases from the committed files. - The guardrail, per channel (M07) — the same Bedrock Guardrail applied differently to the viewer's words, the model's tool request and the tool's output, because Part 1 found that the same attack was blocked 3/3 as a user turn and allowed 2/3 as tool output.
- Surface runners (M10) — Playwright and k6 emit the same verdict envelope as the agent evals, and
gate deciderefuses any record with a key it does not know. - The drill (M11) —
pave drillrehearses a live event against a scenario file and writes a signed GO or NO-GO. - The ledger (M12) — every debt any milestone left behind is in one register with a terminal state: PAID, RETIRED or UNSCHEDULED. Nothing is "TBD".
What one developer's morning looks like

A service team runs one command, writes the agent's behaviour and its golden cases, and opens an ordinary pull request. The gate runs the same five lanes it runs for the platform itself. If the golden score drops, or a probe stops being blocked, or a threshold moved without a second key, the merge is refused with a comment naming exactly what moved.
The measured reality (claim 1 below): the scaffold works, the verifier works, and the developer's remaining job is well over an hour, not the thirty minutes the claim promised — because writing twenty honest golden cases is real work, and four of the twenty-five starter cases I wrote myself had bugs. The claim is marked INCOMPLETE, not rounded up.
What one regulation change looks like

This is the demo I would show a Chief Compliance Officer. A fictional state passes an AI-disclosure act. Legal edits one file it owns, rules/MER-AI-0001.yaml: status proposed → enforced, a control pointing at an eval pack. The gate goes red on the real deployment (exit 1) because the highlights agent does not disclose. The fix is one sentence in the agent's prompt. The gate goes green (exit 0). Every step is traceable from committed files; no step was supplied by hand.
That mechanism worked end to end. The claim built on it failed anyway — see M09 below — because the claim promised more than "the mechanism works."
What one live event looks like

Before the Jefferson Derby goes live, pave drill reads the event's caption fixture against a scenario that says how large a caption gap is acceptable, who owns a failure and how long they have to fix it. A seeded gap produced a signed NO-GO naming the exact gap (c006/c008), the owner (service-team) and a 36-hour fix-by. The same drill without the fix stayed NO-GO for the same reason. With the fix it wrote GO. Edit one byte of the NO-GO and pave drill verify answers signature: MISMATCH.
This is the one Part 2 claim that closed clean. It is also the smallest: one scenario, one committed fixture, a MAC instead of a real signature, no page and no human in the loop.
Try it yourself in five minutes
Everything below the first block runs on a laptop with no AWS account and makes zero model calls. That is deliberate: the road, the verifier, the rules trace and the drill are all deterministic, so you can see the governance without paying for the agent.
1. Clone and check the road is sound
git clone https://github.com/andaro74/beaconpave.git
cd beaconpave
pip install -e . # Python 3.10+; installs the `pave` command
make check # hermetic: unit + contract + rules validation, no cloud
make check is the same gate CI runs. If it passes here, it passes there.
2. Scaffold a governed agent — the building-permit moment
pave new my-agent --brand meridian-sports --classification internal
pave verify my-agent
pave new writes five files with the gateway client, the manifest and a starter golden set already wired in. pave verify is the inspector: change the brand in the manifest to one that doesn't exist, or grant the agent a tool the registry never heard of, and it refuses — fourteen different ways.
3. Walk a regulation from law to test — the thermostat moment
pave rules trace MER-AI-0001
This prints the chain from a fictional AI-disclosure act → the rule file Legal owns → the control → the seven test cases the gate blocks on. Nothing in the chain is typed by hand; it is all read from committed files.
4. Run the game-day drill and try to forge the result
export BEACONPAVE_DRILL_KEY=$(openssl rand -hex 32) # no key, no artifact — 32+ chars
pave drill --event jefferson-derby --tier 3 --out go-no-go.json
pave drill read go-no-go.json
pave drill verify go-no-go.json # now edit one byte of the file and run this again
The drill refuses to write anything without a signing key. With one, it writes a signed GO or NO-GO; edit a single character in the file and verify answers MISMATCH.
5. Optional — deploy the real thing (costs money, needs an AWS account)
make bootstrap # one-time: CDK bootstrap, tool deps
make core # deploy gateway, tools, agent, dashboard
make evals # the 25 golden cases, three samples each
make adversarial # the ten probes; pass = blocked or denied AND logged
Cost posture is serverless only: the target is under $5 a month idle and under $2 per full demo run, and a cost regression blocks a merge the same way a quality regression does.
The repository, in one picture

The README has the full folder-by-folder map. The way to read it is not "where is the code" but "who owns this, and who may change it?" Four answers cover every folder:
- The road (
pave/,templates/,services/,tools/) — what a service team touches. Code. - The judge (
platform/gateway/,platform/policy/,quality/) — the Part 1 machinery. Code, and every model call in the repo passes through it. - The dials (
rules/,drill/,platform/registry/tools.yaml,quality/adversarial/, the thresholds underevals/) — the only files that change what the gate enforces. Data, each with a named owner, each on two keys. - The evidence (
evals/history/,milestones/,docs/adr/,SPEC/) — append-only. A number recorded there is never edited, only superseded by a later entry that says why.
Two files sit outside the picture and are worth opening first: CLAUDE.md, the rules Claude Code reads before any change, and BUILD.md, the milestone build order with each exit artifact.
Part 2, milestone by milestone

Same rules as Part 1: every milestone is branched, tagged, journaled, and every number is left as recorded. The pattern to notice is that most rows close red or unmeasured, and none of them was edited to look better.
M05 — The scaffold and the verifier
Six PRs, zero model calls. pave new renders five files; pave verify refuses fourteen ways. No golden run, because nothing model-facing changed and re-running would spend tokens to reproduce a known number. Claim 1 is INCOMPLETE: nothing is deployed (the verifier is a control on the repository, not the runtime), and the developer's remaining authorship was measured — 510 lines and 138 asserts across the reference pack — at well over an hour. An earlier draft said "roughly an hour"; that was measured as too low.
M06 — The attack register
Ten PRs. Twenty-eight attacks on the platform itself, each with a test that reproduces it, and the two-key check hardened so it refuses to report compliance when it cannot say what it examined. Goldens 21/25 — and the journal says plainly that the +2 over M01 is the estimator (three samples with majority instead of one), not the system. Scheduled to deliver the second tool and claim 10; delivered neither, and was renamed to what it shipped rather than left describing work that did not exist.
M06b, M06c, M06d — The second tool, and a broken instrument
M06b added entitlement-check as a second tool and a trajectory eval (did the agent actually call the tool it claims it called — the thing Part 1's control faked). It took thirty-four PRs, and it closed on 1/25, unrecorded: the entitlement-circumvention guardrail topic refused 17 of 25 cases before they produced an answer to score. That is a guardrail outage, not answer quality, so it was kept out of the append-only history. The tool-plane probes scored 6/6 with two Cedar denials logged — the first observations in the repo to satisfy G4's "denied and logged" on the tool plane.
M06c tried to capture the refused text and could not: the gateway never had it. Three PRs of a cap of six, claim not met, claim not rewritten. M06d fixed the report, so the same run now reads "17 refused before scoring, 7 answered and scored wrong" instead of 24 identical FAILs. A cap is a ceiling, not a budget to exhaust — that rule is M06b's lesson.
M07 — The guardrail, applied per channel
Stage 1 applied the guardrail to the model's tool request: 18/25 refused, on text no viewer would ever see. Stage 2 moved the tool-output channel to a topic-free policy: 1/25 refused. The pinned score is 2/25, and the reason is the milestone's real finding: all 22 answered-and-wrong cases fail the tokens_in budget (6,022–9,220 against a 6,000 ceiling), and 15 fail on nothing else. Without that one assert the suite reads 17/25. The refusals had been hiding a wrong ceiling for four milestones.
M08 and M08b — The ceiling, re-derived and then tested
M08 made zero model calls: it re-derived the ceiling from a census of committed trajectories — 6,000 had been set for a two-call loop, and the loop became three calls when the second tool arrived. At 7,700 every sample at three calls or fewer passes and every sample at four or more fails; re-reading M07's run gives 12/25. M08b then took a fresh run on samples the number was not derived from: the ceiling held per sample, the count came in at 10/25 against a predicted 12, inside the pre-registered band. Both lost cases failed on tokens_out, not on the ceiling. That is the first appearance of the limitation I discuss below.
M09 — The rules registry, and a failed claim
Seven PRs against a cap of six, breach recorded. The disposition worked end to end (see the diagram above). Claim 6 FAILED on two of five pre-registered falsifiers. F1: disclosure-103 was already passing 2 of 3 before the fix — the model disclosed on its own, so the control was not what made it pass. F4.2: grounded-017, a case nothing in the milestone touched, passed 3 of 3 at M08b and failed by majority after. The headline count stayed 10/25, exactly M08b's number, because another case moved the other way. A reader of the count alone would publish "no change." The spec had moved the direction checks into the falsifiers for precisely this reason.
The mechanism is proven; the claim — and the fix makes it pass, with nothing else moving — is not, and it is not re-scoped to fit what survived.
M09b and M09c — Two milestones that closed RED without a measurement
M09b was to add a guardrail line for the one hole a word-match cannot catch (a field reading "written by a human, not by AI" contains the token and passes). Its first spec was withdrawn before any run: it let a corpus choose a wording, and the corpus it chose on was forbidden from judging the fix. The replacement let a corpus only eliminate wordings. Then a cold read found the one surviving candidate's text inside the header of the very corpus that would judge it. Seven PRs, zero model calls, no deploy, closed RED.
M09c was to close the "browse gap" — seven golden cases where the agent asks for a list and the catalog tool returns nothing. The diagnosis found the cause in a line SPEC/02 had named before M02 ran: catalog-search requires a query. Relaxing it turned 19 tests red and invalidated committed numbers in three milestones. That is a semver event owned by the Tool Owner seat, not something a milestone can take quietly. Two PRs, closed RED, debt recorded with an owner.
M10 — One verdict schema, many runners
PR 1 closed the envelope: before it, gate decide accepted a Playwright record carrying an invented p95_ms at exit 0. That fix stands. Then the claim collapsed: every verdict in the repo is built by one function that validates against the same schema the gate reads, so no record a runner writes can be refused — the claim confirms itself and cannot fire. A cold review also found 14 of 24 test plants silent. GREEN was available and refused. Claim 3 is UNSCHEDULED, with the condition that would make it measurable written in the claims table.
M11 — The drill
Five PRs. Claim 11 proven — the arc in the diagram above. Three runs, each valid, none of the five falsifiers fired. It is the first milestone since M08b to measure its claim, and the only Part 2 milestone to close green on one.
M12 — The ledger
Three PRs, zero model calls, zero AWS calls, no claim. A pre-spec feasibility check found claim 12 ("defect leakage counted honestly") has no false state here: there is no delivery pipeline and no rollback on main, so a counter that always returns zero is indistinguishable from a correct one. The one proposed wording — seed a revert and count it — substitutes revert-counting for leakage-counting and was refused. Claims 7, 8 and 12 became UNSCHEDULED together. M12 built the terminal states (RETIRED, UNSCHEDULED) so every obligation in the repo can be named without a milestone to carry it, retired brand_tone after seven deferrals, recorded Act 4, and retired Act 5. Then building stopped.
The twelve claims, at the end

Five proven (2, 4, 5, 9, 11), one incomplete (1), one failed (6), five unscheduled (3, 7, 8, 10, 12). Unscheduled is neither pass nor fail: it means no honest test existed in a repo of this size, and each row states what would create one — a real tool contract bump, a publish-highlight deployment Legal has so far answered no to, a rollback population nobody seeded.
What Part 2 honestly doesn't have
This section is longer than Part 1's, and it should be.
- The
tokens_outtiers. Twelve of twenty-five golden cases fail a per-case output-token tier, several set at 300. Two fail by a single sample writing 307 and 308 tokens.grounded-017passed at 288/289/298 and then failed at 337 and 390 — the tiers are so tight that a fifty-token drift in a correct answer is a "regression." The tiers were never re-derived from measurement; that re-derivation is owned by AI Quality and unscheduled. The recorded 10/25 is mostly this, not answer quality. - The browse gap. Seven cases fail because the catalog tool demands a query and a browse has none. The fix is a schema change with a measured blast radius of 19 tests and three milestones' numbers. Owned by the Tool Owner seat, unscheduled.
- Latency. Suite p95 has breached its 2,500 ms budget in every milestone that measured it — 3,194 · 8,437 · 2,794 · 11,171 · 5,431 ms, in order — and the census shows model time alone on a three-call turn is 3,606 ms. The gate is breached by the shape of the loop. Recorded, never accommodated.
- The guardrail topic that breaks the product.
entitlement-circumventionrefuses "who is playing in the Jefferson Derby?" 1 in 3 times and refused the judge's own calls at M03. Rewording it was M09b's job, and M09b closed without a measurement. - No human-approval interlock in production.
publish-highlight— the one tool with real consequences — was never deployed. Claim 10 is unscheduled pending a Legal/S&P answer the record does not contain. - No deployed team-created agent.
pave verifyguards the repository, not the runtime. - Self-heal is a stub.
pave selfhealneeds a real contract break to classify, and the only candidate is the unscheduled browse-gap bump. - One drill scenario, one fixture, a MAC. The blackout sweep and the alarm self-test were cut.
- One brand. Meridian News never got a surface;
brand_tonewas retired rather than deferred an eighth time. - Seats are subagents. Solo-played, by ADR-001. The two-key check collects attestations because CODEOWNERS on a one-operator repo collects nothing.
- The cost of the rigor. Four consecutive milestones — M09, M09b, M09c, M10 — closed RED without a measurement, and M06b took thirty-four PRs. The rules that stop the platform flattering itself also stopped it finishing more of what it set out to do. I would rather publish that than the alternative, but it is a real trade and anyone adopting this posture should price it.
How I actually worked with AI on this
Same division of labour as Part 1: Claude Code, powered by Fable 5, built every PR; the seven role subagents read each diff from the angle of the seat that owns it; I disposed. What changed in Part 2 is how often the subagents' cold review was the thing that withdrew a spec — M09b's, M10's and M12's claims were all pulled before a token was spent, on findings a solo reader from one angle would have missed. The AI without the invariants would have "fixed" grounded-017 by widening its tier. The invariants without the AI would still be at M06b.
Seventy-seven ADR files (numbered to 081), dozens of amendments to earlier claims, and a README whose footnotes are longer than its table. None of it is what I would ship to a service team. All of it is what I would want to have read before asking a service team to trust the road.
What's next
Nothing, in this repo. The obligations carry on in milestones/M12/README.md with owners and triggers, and each unscheduled claim names the condition that would make it measurable. If I pick it back up, the first move is the boring one: the Tool Owner takes the catalog-search semver bump, the three affected milestones get a fresh baseline, and claim 8 gets its first real contract break to classify.
If you are building governance for agents at any scale, the artifact I would point you at is not the platform. It is the shape of the claims table — twelve falsifiable statements, each with a proof artifact or a written reason it has none. What would your twelve be?
References
- Repository: github.com/andaro74/beaconpave — MIT,
make checkruns offline. Every number in this post is in the README's progression table and its footnotes, at tagm12. - The twelve claims: README § The twelve claims — each with its proof artifact or the condition that would make it measurable.
- The ledger (M12):
milestones/M12/README.md— every open obligation, its owner and its terminal state. - The rule that was disposed (M09):
rules/MER-AI-0001.yamlandmilestones/M09/README.md. - The drill (M11):
drill/scenarios/andmilestones/M11/runs/falsifiers.md. - Decision records:
docs/adr/— 77 ADRs, numbered to 081, superseded ones marked and never deleted. - Part 1 of this series: The machinery that judges an agent.
Transparency note: this article was drafted with Claude Fable 5.1 — driven, reviewed, and fact-checked by me against the project's README, specs, ADRs and milestone journals at tag m12. The same division of labour that built the project wrote the post about it.
#AWS #AmazonBedrock #AgenticAI #AIQuality #Guardrails #Cedar #MCP #GenAI #PlatformEngineering #ClaudeCode #Anthropic