Built because: "Allowing an agent to take long horizon decisions bounded only by the availability of tools is a recipe for unreliable behaviour."

Where Fin did not stop, and what it cost

Here Fin does not stop. It answers a pricing question wrongly, the customer acts on the answer twice, and no rule fires. Nothing flags the conversation.

The bet: the contract has to work when no rule fires. The human who catches the error is the only source of the fix, so their correction gets captured.

A designed scenario. The wrong answer is authored, not a measurement of how often Fin is wrong.

Designed scenario, not a recording of Fin's product.

    Fin is answering this conversation.

    d-025 · Scenario 2 is a wrong answer nobody flagged, and the discovery beat is the stop

    Decision
    Scenario 2 is a confident wrong answer that nothing flagged.
    Because
    A model that raises its hand is the easy half of the argument.
    Rejected
    Have the customer ask for a human
    Would measure
    Whether the missing guideline row reads as a bug.