Please note: This PhD defence will take place in DC 2314.
Marvin Pafla, PhD candidate
David R. Cheriton School of Computer Science
Supervisors: Professors Kate Larson, Mark Hancock
This thesis asks where explanation actually lives, and argues that explanation should build a person’s understanding in interaction rather than verify a model’s output. It begins by comparing human and machine explanations of correct and incorrect answers from an AI model, finding that the explanations people preferred did no better at helping them catch errors, because a persuasive explanation lends the same credibility to wrong answers as to right ones. Indeed, greater trust in the AI and higher-rated explanations correlated with lower accuracy. This dilemma of AI errors turns out to be ontological rather than merely technical, because understanding is not stored inside a model but built in interaction, so an explanation must reckon with what the explainee brings to it.
This motivates an account of embodied explainability, on which a system is explainable insofar as it offers affordances people can use to coordinate, check, and repair what it is doing, rather than by exposing its inner workings. A pre-registered study then shows that a system contrasting against what a learner already understands, through minimal edits to shared material, builds more understanding than selecting the samples that maximize an information-theoretic criterion. It thereby addresses the selection problem—the question of which foil to contrast the explanandum against—by contrasting against the explainee’s ever-changing understanding, and so locates explanation in the interaction itself.