Publications
Eye gaze patterns can reveal user attention, reading fluency, corrective responding, and other reading processes, suggesting they can be used to develop automated, real-time assessments of comprehension. However, past work has focused on modeling factual comprehension, whereas we ask whether gaze patterns reflect deeper levels of comprehension where inferencing and elaboration are key. We trained linear regression and random forest models to predict the quality of users' open-ended self-explanations (SEs) collected both during and after reading and scored on a continuous scale by human raters. Our models use theoretically grounded eye tracking features (number and duration of fixations, saccade distance, proportion of regressive and horizontal saccades, spatial dispersion of fixations, and reading time) captured from a remote, head-free eye tracker (Tobii TX300) as adult users read a long expository text (6500 words) in two studies (N = 106 and 131; 247 total). Our models: (1) demonstrated convergence with human-scored SEs (r = .322 and .354), by capturing both within-user and between-user differences in comprehension; (2) were distinct from alternate models of mind-wandering and shallow comprehension; (3) predicted multiple-choice posttests of inference-level comprehension (r = .288, .354) measured immediately after reading and after a week-long delay beyond the comparison models; and (4) generalized across new users and datasets. Such models could be embedded in digital reading interfaces to improve comprehension outcomes by delivering interventions based on users' level of comprehension.
Psychological science can benefit from and contribute to emerging approaches from the computing and information sciences driven by the availability of real-world data and advances in sensing and computing. We focus on one such approach, machine-learned computational models (MLCMs)-computer programs learned from data, typically with human supervision. We introduce MLCMs and discuss how they contrast with traditional computational models and assessment in the psychological sciences. Examples of MLCMs from cognitive and affective science, neuroscience, education, organizational psychology, and personality and social psychology are provided. We consider the accuracy and generalizability of MLCM-based measures, cautioning researchers to consider the underlying context and intended use when interpreting their performance. We conclude that in addition to known data privacy and security concerns, the use of MLCMs entails a reconceptualization of fairness, bias, interpretability, and responsible use.
What can eye movements reveal about reading, a complex skill ubiquitous in everyday life? Research suggests that gaze can reflect short-term comprehension for facts, but it is unknown whether it can measure long-term, deep comprehension. We tracked gaze while 147 participants read long, connected, informative texts and completed assessments of rote (factual) and inference comprehension (connecting ideas) while reading a text, after reading a text, after reading five texts, and after a seven-day delay. Gaze-based student-independent computational models predicted both immediate and long-term rote and inference comprehension with moderate accuracies. Surprisingly, the models were most accurate for comprehension assessed after reading all texts and predicted comprehension even after a week-long delay. This shows that eye movements can provide a lens into the cognitive processes underlying reading comprehension, including inference formation, and the consolidation of information into long-term memory, which has implications for intelligent student interfaces that can automatically detect and repair comprehension in real-time. © 2022 Copyright is held by the author(s).
Can Computers Outperform Humans in Detecting User Zone-Outs? Implications for Intelligent Interfaces
The ability to identify whether a user is zoning out (mind wandering) from video has many HCI (e.g., distance learning, high-stakes vigilance tasks). However, it remains unknown how well humans can perform this task, how they compare to automatic computerized approaches, and how a fusion of the two might improve accuracy. We analyzed videos of users' faces and upper bodies recorded 10s prior to self-reported mind wandering (i.e., ground truth) while they engaged in a computerized reading task. We found that a state-of-the-art machine learning model had comparable accuracy to aggregated judgments of nine untrained human observers (area under receiver operating characteristic curve [AUC] =.598 versus .589). A fusion of the two (AUC = .644) outperformed each, presumably because each focused on complementary cues. Furthermore, adding more humans beyond 3-4 observers yielded diminishing returns. We discuss implications of human-computer fusion as a means to improve accuracy in complex tasks.
Eye movements provide a window into cognitive processes, but much of the research harnessing this data has been confined to the laboratory. We address whether eye gaze can be passively, reliably, and privately recorded in real-world environments across extended timeframes using commercial-off-the-shelf (COTS) sensors. We recorded eye gaze data from a COTS tracker embedded in participants (N=20) work environments at pseudorandom intervals across a two-week period. We found that valid samples were recorded approximately 30% of the time despite calibrating the eye tracker only once and without placing any other restrictions on participants. The number of valid samples decreased over days with the degree of decrease dependent on contextual variables (i.e., frequency of video conferencing) and individual difference attributes (e.g., sleep quality and multitasking ability). Participants reported that sensors did not change or impact their work. Our findings suggest the potential for the collection of eye-gaze in authentic environments. © 2022 ACM.
Eye tracking has been a research tool for decades, providing insights into interactions, usability, and, more recently, gaze-enabled interfaces. Recent work has utilized consumer-grade and webcam-based eye tracking, but is limited by the need to repeatedly calibrate the tracker, which becomes cumbersome for use outside the lab. To address this limitation, we developed an unsupervised algorithm that maps gaze vectors from a webcam to fxation features used for user modeling, bypassing the need for screen-based gaze coordinates, which require a calibration process. We evaluated our approach using three datasets (N=377) encompassing diferent UIs (computerized reading, an Intelligent Tutoring System), environments (laboratory or the classroom), and a traditional gaze tracker used for comparison. Our research shows that webcam-based gaze features correlate moderately with eye-tracker-based features and can model user engagement and comprehension as accurately as the latter. We discuss applications for research and gaze-enabled user interfaces for long-term use in the wild.
Student engagement is a key component of learning and teaching, resulting in a plethora of automated methods to measure it. Whereas most of the literature explores student engagement analysis using computer-based learning often in the lab, we focus on using classroom instruction in authentic learning environments. We collected audiovisual recordings of secondary school classes over a one and a half month period, acquired continuous engagement labeling per student (N=15) in repeated sessions, and explored computer vision methods to classify engagement from facial videos. We learned deep embeddings for attentional and affective features by training Attention-Net for head pose estimation and Affect-Net for facial expression recognition using previously-collected large-scale datasets. We used these representations to train engagement classifiers on our data, in individual and multiple channel settings, considering temporal dependencies. The best performing engagement classifiers achieved student-independent AUCs of .620 and .720 for grades 8 and 12, respectively, with attention-based features outperforming affective features. Score-level fusion either improved the engagement classifiers or was on par with the best performing modality. We also investigated the effect of personalization and found that only 60 seconds of person-specific data, selected by margin uncertainty of the base classifier, yielded an average AUC improvement of .084.
A large proportion of thoughts are internally generated. Of these, mind wandering-when attention shifts away from the current activity to an internal stream of thought-is frequent during reading and is negatively related to comprehension outcomes. Our goal is to review research on mind wandering during reading with an interdisciplinary and integrative lens that spans the cognitive, behavioural, computing and intervention sciences. We begin with theoretical developments on mind wandering, both in general and in the context of reading. Next, we discuss psychological research on how the text, context and reader interact to influence mind wandering and on associations between mind wandering and reading outcomes. We integrate the findings in a (working) theoretical account of mind wandering during reading. We then turn to computational models of mind wandering, including a short tutorial with examples on how to use machine learning to construct these models. Finally, we discuss emerging intervention research aimed at proactively reducing the occurrence of mind wandering or mitigating its effects. We conclude with open questions and directions for future research.


