Flores Innovations

I'm building a production-shaped quality platform for AI agents on AWS — Part 2

p2-00-cover_header

Part 1 ended with a promise: one command, and what comes out is governed by default. This post is the honest account of building that command and everything around it — including the milestones that closed red, the claim that failed, and the five claims that turned out to be unmeasurable in a repo this size.

Repo (MIT, make check runs offline with no AWS account): github.com/andaro74/beaconpave, tag m12. Building has stopped. What is left is a ledger, and the ledger is the deliverable.


Part 2 in three sentences

What it is. A miniature platform where an engineering team runs pave new, gets an AI agent that is already wired into a gateway, a tool registry, a golden test set and a CI gate that cannot be merged past — and where the rules that gate enforces live in plain data files owned by the business, not in code owned by engineers.

What is different. Most agent frameworks make it easier to build an agent. beaconpave makes it hard to ship an ungoverned one, and it does that by treating compliance, quality and adversarial resistance as properties of the road rather than a checklist a team completes before launch.

What it delivers. For an executive: a regulation change becomes a one-file edit by the seat that owns it, and every agent on the road goes red until it complies. For an engineer: you write behaviour and tests; you never write a gateway, a policy, a guardrail or an audit log, because the road already did.

p2-01-two-parts


The problem Part 2 solves, in one paragraph

Part 1 proved that an agent can be judged: a gateway it cannot bypass, a registry that makes unregistered tools unreachable, a gate that fails closed. But none of that helps a service team on Monday morning. They still have to find the gateway, learn Cedar, write a manifest nothing checks, and guess what "good" looks like. Part 2 is the answer to a different question: how does a team get onto the road without learning how the road was built? And underneath that: when the law changes, who changes what?

Two real-world comparisons that carry the whole design

A building permit, not a fire drill. A city does not ask each contractor to invent their own fire code. It hands over the code, the contractor builds, and an inspector checks against the code. In beaconpave the code is rules/, platform/registry/tools.yaml and the golden set; pave new hands it over; the CI gate is the inspector. The contractor's skill is the building. The code is not theirs to rewrite.

A thermostat with a lock on it. You change the temperature of a room by turning one dial — not by rewiring the furnace. In beaconpave there are exactly five dials that change what the gate enforces (rules, drill scenarios, the tool registry, thresholds and baselines, the adversarial corpus). Each has a named owner, and each needs a second key to turn. The agent's code is not a dial. You cannot make the gate kinder by editing the thing it is judging.

p2-04-the-only-knobs

That second picture is the sentence I most want an executive to leave with: only by changing the business specs do the gates and rules change. Legal changes a rule file; AI Quality changes a threshold; Security changes what counts as an attack. Nobody changes the gate.


What is different from other agent platforms

A typical agent framework beaconpave
Makes an agent easier to build Makes an ungoverned agent impossible to merge
Rules live in prompts and code Rules live in data files with an owner, a source law, a control and a review-by date (G7)
"Tested" means the team ran it Every milestone is scored against the same 25 golden cases and 10 probes, three samples each, append-only
A judge model grades the answers The judge was measured against humans first, failed, and was demoted to advisory (claim 9)
Security is a review step An adversarial pass counts only if a guardrail blocked or a policy denied and an audit record exists (G4)
A green dashboard Pre-registered predictions and named falsifiers; a claim that cannot be falsified is refused, not celebrated
Success is measured by what shipped Success is measured by what was proven; the final tally is five of twelve, and it is published as-is

The last row is the unusual one, and it is deliberate. A platform that is not allowed to flatter itself will show fewer green boxes than one that is. That is a cost. I think it is the right cost for infrastructure that other teams will trust with compliance.


The architecture (what exists today)

Nothing in Part 1's diagram changed. Part 2 added the pieces around it:


What one developer's morning looks like

p2-02-paved-road

A service team runs one command, writes the agent's behaviour and its golden cases, and opens an ordinary pull request. The gate runs the same five lanes it runs for the platform itself. If the golden score drops, or a probe stops being blocked, or a threshold moved without a second key, the merge is refused with a comment naming exactly what moved.

The measured reality (claim 1 below): the scaffold works, the verifier works, and the developer's remaining job is well over an hour, not the thirty minutes the claim promised — because writing twenty honest golden cases is real work, and four of the twenty-five starter cases I wrote myself had bugs. The claim is marked INCOMPLETE, not rounded up.

What one regulation change looks like

p2-03-rule-change

This is the demo I would show a Chief Compliance Officer. A fictional state passes an AI-disclosure act. Legal edits one file it owns, rules/MER-AI-0001.yaml: status proposed → enforced, a control pointing at an eval pack. The gate goes red on the real deployment (exit 1) because the highlights agent does not disclose. The fix is one sentence in the agent's prompt. The gate goes green (exit 0). Every step is traceable from committed files; no step was supplied by hand.

That mechanism worked end to end. The claim built on it failed anyway — see M09 below — because the claim promised more than "the mechanism works."

What one live event looks like

p2-05-drill

Before the Jefferson Derby goes live, pave drill reads the event's caption fixture against a scenario that says how large a caption gap is acceptable, who owns a failure and how long they have to fix it. A seeded gap produced a signed NO-GO naming the exact gap (c006/c008), the owner (service-team) and a 36-hour fix-by. The same drill without the fix stayed NO-GO for the same reason. With the fix it wrote GO. Edit one byte of the NO-GO and pave drill verify answers signature: MISMATCH.

This is the one Part 2 claim that closed clean. It is also the smallest: one scenario, one committed fixture, a MAC instead of a real signature, no page and no human in the loop.


Try it yourself in five minutes

Everything below the first block runs on a laptop with no AWS account and makes zero model calls. That is deliberate: the road, the verifier, the rules trace and the drill are all deterministic, so you can see the governance without paying for the agent.

1. Clone and check the road is sound

git clone https://github.com/andaro74/beaconpave.git
cd beaconpave
pip install -e .          # Python 3.10+; installs the `pave` command
make check                # hermetic: unit + contract + rules validation, no cloud

make check is the same gate CI runs. If it passes here, it passes there.

2. Scaffold a governed agent — the building-permit moment

pave new my-agent --brand meridian-sports --classification internal
pave verify my-agent

pave new writes five files with the gateway client, the manifest and a starter golden set already wired in. pave verify is the inspector: change the brand in the manifest to one that doesn't exist, or grant the agent a tool the registry never heard of, and it refuses — fourteen different ways.

3. Walk a regulation from law to test — the thermostat moment

pave rules trace MER-AI-0001

This prints the chain from a fictional AI-disclosure act → the rule file Legal owns → the control → the seven test cases the gate blocks on. Nothing in the chain is typed by hand; it is all read from committed files.

4. Run the game-day drill and try to forge the result

export BEACONPAVE_DRILL_KEY=$(openssl rand -hex 32)   # no key, no artifact — 32+ chars
pave drill --event jefferson-derby --tier 3 --out go-no-go.json
pave drill read go-no-go.json
pave drill verify go-no-go.json      # now edit one byte of the file and run this again

The drill refuses to write anything without a signing key. With one, it writes a signed GO or NO-GO; edit a single character in the file and verify answers MISMATCH.

5. Optional — deploy the real thing (costs money, needs an AWS account)

make bootstrap            # one-time: CDK bootstrap, tool deps
make core                 # deploy gateway, tools, agent, dashboard
make evals                # the 25 golden cases, three samples each
make adversarial          # the ten probes; pass = blocked or denied AND logged

Cost posture is serverless only: the target is under $5 a month idle and under $2 per full demo run, and a cost regression blocks a merge the same way a quality regression does.

The repository, in one picture

p2-08-repository-map

The README has the full folder-by-folder map. The way to read it is not "where is the code" but "who owns this, and who may change it?" Four answers cover every folder:

Two files sit outside the picture and are worth opening first: CLAUDE.md, the rules Claude Code reads before any change, and BUILD.md, the milestone build order with each exit artifact.


Part 2, milestone by milestone

p2-06-part2-timeline

Same rules as Part 1: every milestone is branched, tagged, journaled, and every number is left as recorded. The pattern to notice is that most rows close red or unmeasured, and none of them was edited to look better.

M05 — The scaffold and the verifier

Six PRs, zero model calls. pave new renders five files; pave verify refuses fourteen ways. No golden run, because nothing model-facing changed and re-running would spend tokens to reproduce a known number. Claim 1 is INCOMPLETE: nothing is deployed (the verifier is a control on the repository, not the runtime), and the developer's remaining authorship was measured — 510 lines and 138 asserts across the reference pack — at well over an hour. An earlier draft said "roughly an hour"; that was measured as too low.

M06 — The attack register

Ten PRs. Twenty-eight attacks on the platform itself, each with a test that reproduces it, and the two-key check hardened so it refuses to report compliance when it cannot say what it examined. Goldens 21/25 — and the journal says plainly that the +2 over M01 is the estimator (three samples with majority instead of one), not the system. Scheduled to deliver the second tool and claim 10; delivered neither, and was renamed to what it shipped rather than left describing work that did not exist.

M06b, M06c, M06d — The second tool, and a broken instrument

M06b added entitlement-check as a second tool and a trajectory eval (did the agent actually call the tool it claims it called — the thing Part 1's control faked). It took thirty-four PRs, and it closed on 1/25, unrecorded: the entitlement-circumvention guardrail topic refused 17 of 25 cases before they produced an answer to score. That is a guardrail outage, not answer quality, so it was kept out of the append-only history. The tool-plane probes scored 6/6 with two Cedar denials logged — the first observations in the repo to satisfy G4's "denied and logged" on the tool plane.

M06c tried to capture the refused text and could not: the gateway never had it. Three PRs of a cap of six, claim not met, claim not rewritten. M06d fixed the report, so the same run now reads "17 refused before scoring, 7 answered and scored wrong" instead of 24 identical FAILs. A cap is a ceiling, not a budget to exhaust — that rule is M06b's lesson.

M07 — The guardrail, applied per channel

Stage 1 applied the guardrail to the model's tool request: 18/25 refused, on text no viewer would ever see. Stage 2 moved the tool-output channel to a topic-free policy: 1/25 refused. The pinned score is 2/25, and the reason is the milestone's real finding: all 22 answered-and-wrong cases fail the tokens_in budget (6,022–9,220 against a 6,000 ceiling), and 15 fail on nothing else. Without that one assert the suite reads 17/25. The refusals had been hiding a wrong ceiling for four milestones.

M08 and M08b — The ceiling, re-derived and then tested

M08 made zero model calls: it re-derived the ceiling from a census of committed trajectories — 6,000 had been set for a two-call loop, and the loop became three calls when the second tool arrived. At 7,700 every sample at three calls or fewer passes and every sample at four or more fails; re-reading M07's run gives 12/25. M08b then took a fresh run on samples the number was not derived from: the ceiling held per sample, the count came in at 10/25 against a predicted 12, inside the pre-registered band. Both lost cases failed on tokens_out, not on the ceiling. That is the first appearance of the limitation I discuss below.

M09 — The rules registry, and a failed claim

Seven PRs against a cap of six, breach recorded. The disposition worked end to end (see the diagram above). Claim 6 FAILED on two of five pre-registered falsifiers. F1: disclosure-103 was already passing 2 of 3 before the fix — the model disclosed on its own, so the control was not what made it pass. F4.2: grounded-017, a case nothing in the milestone touched, passed 3 of 3 at M08b and failed by majority after. The headline count stayed 10/25, exactly M08b's number, because another case moved the other way. A reader of the count alone would publish "no change." The spec had moved the direction checks into the falsifiers for precisely this reason.

The mechanism is proven; the claim — and the fix makes it pass, with nothing else moving — is not, and it is not re-scoped to fit what survived.

M09b and M09c — Two milestones that closed RED without a measurement

M09b was to add a guardrail line for the one hole a word-match cannot catch (a field reading "written by a human, not by AI" contains the token and passes). Its first spec was withdrawn before any run: it let a corpus choose a wording, and the corpus it chose on was forbidden from judging the fix. The replacement let a corpus only eliminate wordings. Then a cold read found the one surviving candidate's text inside the header of the very corpus that would judge it. Seven PRs, zero model calls, no deploy, closed RED.

M09c was to close the "browse gap" — seven golden cases where the agent asks for a list and the catalog tool returns nothing. The diagnosis found the cause in a line SPEC/02 had named before M02 ran: catalog-search requires a query. Relaxing it turned 19 tests red and invalidated committed numbers in three milestones. That is a semver event owned by the Tool Owner seat, not something a milestone can take quietly. Two PRs, closed RED, debt recorded with an owner.

M10 — One verdict schema, many runners

PR 1 closed the envelope: before it, gate decide accepted a Playwright record carrying an invented p95_ms at exit 0. That fix stands. Then the claim collapsed: every verdict in the repo is built by one function that validates against the same schema the gate reads, so no record a runner writes can be refused — the claim confirms itself and cannot fire. A cold review also found 14 of 24 test plants silent. GREEN was available and refused. Claim 3 is UNSCHEDULED, with the condition that would make it measurable written in the claims table.

M11 — The drill

Five PRs. Claim 11 proven — the arc in the diagram above. Three runs, each valid, none of the five falsifiers fired. It is the first milestone since M08b to measure its claim, and the only Part 2 milestone to close green on one.

M12 — The ledger

Three PRs, zero model calls, zero AWS calls, no claim. A pre-spec feasibility check found claim 12 ("defect leakage counted honestly") has no false state here: there is no delivery pipeline and no rollback on main, so a counter that always returns zero is indistinguishable from a correct one. The one proposed wording — seed a revert and count it — substitutes revert-counting for leakage-counting and was refused. Claims 7, 8 and 12 became UNSCHEDULED together. M12 built the terminal states (RETIRED, UNSCHEDULED) so every obligation in the repo can be named without a milestone to carry it, retired brand_tone after seven deferrals, recorded Act 4, and retired Act 5. Then building stopped.


The twelve claims, at the end

p2-07-scorecard

Five proven (2, 4, 5, 9, 11), one incomplete (1), one failed (6), five unscheduled (3, 7, 8, 10, 12). Unscheduled is neither pass nor fail: it means no honest test existed in a repo of this size, and each row states what would create one — a real tool contract bump, a publish-highlight deployment Legal has so far answered no to, a rollback population nobody seeded.


What Part 2 honestly doesn't have

This section is longer than Part 1's, and it should be.


How I actually worked with AI on this

Same division of labour as Part 1: Claude Code, powered by Fable 5, built every PR; the seven role subagents read each diff from the angle of the seat that owns it; I disposed. What changed in Part 2 is how often the subagents' cold review was the thing that withdrew a spec — M09b's, M10's and M12's claims were all pulled before a token was spent, on findings a solo reader from one angle would have missed. The AI without the invariants would have "fixed" grounded-017 by widening its tier. The invariants without the AI would still be at M06b.

Seventy-seven ADR files (numbered to 081), dozens of amendments to earlier claims, and a README whose footnotes are longer than its table. None of it is what I would ship to a service team. All of it is what I would want to have read before asking a service team to trust the road.

What's next

Nothing, in this repo. The obligations carry on in milestones/M12/README.md with owners and triggers, and each unscheduled claim names the condition that would make it measurable. If I pick it back up, the first move is the boring one: the Tool Owner takes the catalog-search semver bump, the three affected milestones get a fresh baseline, and claim 8 gets its first real contract break to classify.

If you are building governance for agents at any scale, the artifact I would point you at is not the platform. It is the shape of the claims table — twelve falsifiable statements, each with a proof artifact or a written reason it has none. What would your twelve be?


References


Transparency note: this article was drafted with Claude Fable 5.1 — driven, reviewed, and fact-checked by me against the project's README, specs, ADRs and milestone journals at tag m12. The same division of labour that built the project wrote the post about it.

#AWS #AmazonBedrock #AgenticAI #AIQuality #Guardrails #Cedar #MCP #GenAI #PlatformEngineering #ClaudeCode #Anthropic

#agentic-ai #ai-quality #amazon-bedrock #anthropic #aws #cedar #claude-code #genai #guardrails #mcp #platform-engineering