Built because: "Allowing an agent to take long horizon decisions bounded only by the availability of tools is a recipe for unreliable behaviour."
Where Fin did not stop, and what it cost
Here Fin does not stop. It answers a pricing question wrongly, the customer acts on the answer twice, and no rule fires. Nothing flags the conversation.
The bet: the contract has to work when no rule fires. The human who catches the error is the only source of the fix, so their correction gets captured.
A designed scenario. The wrong answer is authored, not a measurement of how often Fin is wrong.
Designed scenario, not a recording of Fin's product.
Fin is answering this conversation.
The replay did not start.
This page was opened from the file system, so the browser refuses to load its module script and the scenario cannot be fetched. Serve this directory over HTTP and reload.
d-025 · Scenario 2 is a wrong answer nobody flagged, and the discovery beat is the stop
- Decision
- Scenario 2 is a confident wrong answer that nothing flagged.
- Because
- A model that raises its hand is the easy half of the argument.
- Rejected
- Have the customer ask for a human
- Would measure
- Whether the missing guideline row reads as a bug.