Built because: "Allowing an agent to take long horizon decisions bounded only by the availability of tools is a recipe for unreliable behaviour."

When Fin was wrong and nobody said so

The other half of the handover contract. Here Fin does not stop: it answers a pricing question confidently, the answer is wrong, and the customer acts on it twice before anyone notices. No rule fires and no reason is logged, so nothing flags the conversation.

The bet: a contract that only works when the model raises its hand is a courtesy. The human who catches the error is the only source of the fix, so the interface has to capture it.

This is a designed scenario. Fin’s wrong answer is authored for this concept, and it is not a measurement of how often Fin is wrong.

Designed scenario, not a recording of Fin's product.

    Fin is answering this conversation.

    d-025 · Scenario 2 is a wrong answer nobody flagged, and the discovery beat is the stop

    Decision
    Scenario 2 is a confident wrong answer that nothing flagged.
    Because
    A model that raises its hand is the easy half of the argument.
    Rejected
    Have the customer ask for a human
    Would measure
    Whether the missing guideline row reads as a bug.