← Demos

Demo · Detecting AI lies

Detecting AI lies

Assistant UIs render guesses and verified facts with the same weight. This replay shows how inline marks — dashed for heuristic, solid for researched — would have flagged a wrong fork slug before it burned three turns.

Agent transcript Scene: fork-name retry
Agent Step 0 / 4

    Provenance on.

    Real Cursor chat where the agent assumed AMDphreak/clients before finding AMDphreak/bitwarden-clients
    Real scene that seeded this idea — same story as the replay above.