Do Not Trust a Model Because It is Confident: Uncovering and Characterizing Unknown Unknowns to Student Success Predictors in Online-Based Learning

December 16, 2022 · Declared Dead · 🏛 International Conference on Learning Analytics and Knowledge

Authors Roberta Galici, Tanja Käser, Gianni Fenu, Mirko Marras arXiv ID 2212.08532 Category cs.LG: Machine Learning Cross-listed cs.CY Citations 7 Venue International Conference on Learning Analytics and Knowledge Repository https://github.com/epfl-ml4ed/unknown-unknowns Last Checked 2 months ago

Abstract

Student success models might be prone to develop weak spots, i.e., examples hard to accurately classify due to insufficient representation during model creation. This weakness is one of the main factors undermining users' trust, since model predictions could for instance lead an instructor to not intervene on a student in need. In this paper, we unveil the need of detecting and characterizing unknown unknowns in student success prediction in order to better understand when models may fail. Unknown unknowns include the students for which the model is highly confident in its predictions, but is actually wrong. Therefore, we cannot solely rely on the model's confidence when evaluating the predictions quality. We first introduce a framework for the identification and characterization of unknown unknowns. We then assess its informativeness on log data collected from flipped courses and online courses using quantitative analyses and interviews with instructors. Our results show that unknown unknowns are a critical issue in this domain and that our framework can be applied to support their detection. The source code is available at https://github.com/epfl-ml4ed/unknown-unknowns.