← The work

Case study · platform · 2026

email innovation hub: nine agents behind a quality gate.

AI agents scaffold, fix and review production email; a ten-point gate decides what ships. Built solo, around the hardest consistency problem I know.

My own live product, built outside this concept. Not work for Fin.

specialised agents
9
automated gate checks
10
API endpoints
261
tests
7,055

The context

Ten years in agency email teams showed the same day everywhere: table HTML rebuilt by hand, breaking in Outlook, dark mode inverting a logo, QA squinting at ninety screenshots at 6pm. The expensive part was holding one design consistent across dozens of rendering engines.

The bet

If the consistency rules can be written down, machines can hold them. Agents do the scaffolding and the Outlook fixes; the gate checks everything before a human reviews. People keep the brief and the brand.

The email hub workspace: agent crew panel on the left, live email build in the centre, pipeline activity and a human-in-the-loop approval prompt on the right The approval view: client verdicts, comments and an audit trail inside the platform The QA engine: deterministic checks with scores and verdicts, a repair pipeline of eight stages, and an adversarial pass summary
The workspace: agent crew left, live build centre, pipeline activity and a human-in-the-loop pause right. Tap a capture to open it full size.

The design decisions

The workspace is a single surface.

Code editor, live preview and agent chat sit together, because email work is a conversation between intent and rendering, and every tab switch between those was waste.

The gate is a UX, not a report.

Each of the ten checks returns a verdict a non-developer can act on, next to the thing it refers to, so a failed check becomes a task rather than a finding.

The design system is built from the tokens up.

Tiered tokens beneath a semantic contract, components above. Agents generated the densest surfaces against that one contract and they read as one product. The system re-skins in one file.

Design syncs from Figma.

The source of truth stays in Figma; the platform pulls it into coded, versioned components, so design and production cannot quietly diverge.

Approval is part of the product.

Clients review, comment and sign off inside the platform with an audit trail, because in agency life the approval loop is where weeks actually go.

Personas make QA honest.

Test profiles for device, email client and dark mode, so "does it work" is a checklist run by machines, not a hope.

The platform remembers.

A knowledge module answers from the team’s own record: past fixes, client rules, rendering quirks. Agents consult it before proposing, so judgement compounds instead of resetting with every brief.

What shipped

Deployed and live: FastAPI and Next.js, 261 endpoints, 7,055 tests, built solo with the IP mine. Several features were put forward for adoption at Dentsu before my role was made redundant.

What I learned

The gate outperforms the most careful reviewer at 6pm and never tires of Outlook. Quality held by machines is now the centre of my practice.