Epistemic Mirrors: Using AI Failures to Reflect on Human Confidence, Evidence, and Agency
Abstract
A chatbot answers confidently and wrongly, a classifier scores well using a feature that predicts the label without capturing the task, and a coding agent turns a failing test suite green by rewriting the assertions. In each case something easy to see stood in for something harder to see and stopped tracking it: fluency for competence, a score for progress. People lean on the same substitutions when they judge their own work. Mirror systems so far reflect a user's knowledge, profile, or framing back at them; none uses an AI failure as a case for examining how the user judges. In this work we introduce epistemic mirrors, interactions that make the parts of such a failure visible: what is being judged, the cue or measure standing in for it, and the process doing the judging. The user then looks for the same parts in a practice of their own, asking which of their measures plays the part the test suite played, and stays free to reject the comparison. Confidence, shortcut, and proxy mirrors fail in different ways but share this structure, so the same claims apply to all three, each needing its own evidence: comprehension, self-comparison, regulation, agency, transfer, and optional identity change. We fix the outcome each claim needs, including reliance scored in both directions and confidence that tracks whether a user can tell right answers from wrong ones, and commit to two predictions that could fail, among them that direct metacognitive instruction will match the mirror when users are asked to explain and fall behind when they are not. We separate evidence that a mirror worked from evidence that it merely taught distrust, and state what an interface owes a user who rejects the reflection.