Publications
To effectively process complex information within intelligent tutoring systems (ITSs), learners are required to engage in metacognitive monitoring micro-processes (content evaluations [CEs], judgments of learning [JOLs], feelings of knowing [FOKs], and monitoring progress towards goals [MPTGs]). Learners' average monitoring micro-process strategy frequencies were used to examine learning gains using a person-centered approach as they interacted with MetaTutor. Undergraduates (n = 94) engaged in self-initiated and system-facilitated self-regulated learning (SRL) strategies as they studied the human circulatory system with MetaTutor, a hypermedia-based ITS. Using hierarchical clustering, results showed a difference in learning between clusters differing in metacognitive monitoring process usage. Specifically, learners who used both CEs and FOKs for a greater proportion of monitoring strategy usage had significantly greater learning gains than learners who used MPTGs. Implications for monitoring strategy usage across different micro-processes and the development of ITSs to facilitate and scaffold learners' interactions with these micro-processes via prompting are discussed.
New ways to identify students in need of assistance are imperative to the evolution of online tutoring platforms. Currently implemented models to identify struggling students use costly and tedious classroom observation paired with student's platform usage, and are often suitable for only a subset of students. With the recent influx of new students to online tutoring platforms due to COVID-19, a simple method to quickly identify struggling students could help facilitate effective remote learning. To this end, we created an anomaly detection algorithm that models the normal behavior of students during remote learning and recognizes when students deviate from this behavior. We demonstrated how anomalous behavior revealed which students needed additional assistance and predicted student learning outcomes.
Emotions in Intelligent Tutoring Systems (ITS) are often modeled as single affective states, however there is evidence that emotions co-occur during learning, with implications for affect-aware ITS that need to have a comprehensive understanding of a student's affective state to react accordingly. In this paper we broaden the evidence that emotions co-occur in an educational context, and present a first attempt to predict these co-occurrences from data, using the MetaTutor ITS as a test-bed. We show that boredom+frustration, as well as curiosity+anxiety, frequently co-occur in MetaTutor, and that we can predict when these emotions co-occur significantly better than a baseline using eye-tracking and interaction data. These findings provide a first step toward building affect-aware ITS that can adapt to these complex co-occurring affective states.
Multi-modal approaches have increasingly shown promise in exploring the human side of engineering via the assessment of authentic responses to learning or working environments. This study explores the utility of non-invasive physiological wrist sensors in measuring the reactive and regulatory responses of a group of 161 engineering students taking an authentic engineering practice exam. The practice exam was categorized into Conceptual problems (e.g., rote memorization) and Analytical problems (e.g., requiring application of learned concepts through equations and free-body diagrams). Responses were measured through electrodermal activity and indicators of performance. Findings identified that the type of practice exam problem, even if designed to be within the moderate range of difficulty, influenced how students reacted to and regulated their performance to the problem (as seen by stronger positive correlations in the Analytical problems) and that these may occur via multi-componential processes.
One of the central goals of doctoral programs is to develop independent researchers and scholars who will lead the next generation of knowledge production. Despite extant evidence of inequalities in doctoral education, few studies have closely examined the experiences of first-generation college students who pursue a Ph.D. We examine how first-generation and continuing-generation doctoral students conceptualize the role of the faculty advisor/principal investigator (PI) in supporting their development as researchers. Our analysis of interviews from 111 first-year Ph.D. students in the biological sciences indicates that first-generation and continuing-generation students had similar overarching conceptions of PIs and the role of PIs in their development. However, the two groups ascribed different meanings to the same concepts. First-generation students expected more direct, skill-based guidance and assistance with learning to do research the right way. Conversely, continuing-generation students expected independence and support for their specific needs. We rely on Bourdieu's conceptualization of habitus to explain these differences and conclude by offering implications for advancing equity in doctoral education and supporting first-generation students, particularly regarding the alignment of student-advisor expectations.
This paper presents an analysis of how school mathematics credentials, particularly calculus, coupled with elite college degrees can allow for presumed proficiencies, resulting in acceptance to a selective alternative route program (SARP) for secondary mathematics. However, such credentials do not guarantee deep mathematics knowledge nor effective teaching, particularly in low-income schools serving a majority of Black and Latinx students. We present a portrait of two NYC Teaching Fellows secondary mathematics teachers, their mathematics preparation and credentials, and how their supervisors viewed them. We further present findings from two years of classroom observations and show that these SARP mathematics teachers resorted to the direct, procedural teaching style they had known themselves, held superficial mathematical understandings, regularly made math errors, and often struggled to coherently answer students' mathematical questions.
Interactions around unexpected, incorrect, or dis-preferred responses can be powerful sites of learning for both teachers and students. The information that teachers uncover through probing student thinking can then guide their pedagogical response. We report on a study of prospective teachers' skills and capabilities around a particular problem of practice: eliciting student thinking when a student has an incorrect answer. In this case, if the student's thinking is sufficiently probed, the student is able to recognize the mistake and revise their work. Focusing on prospective teachers at the beginning of a teacher preparation program, we illustrate how knowledge of this kind of eliciting skill can be gathered through the use of a live teaching simulation. Our findings reveal that these prospective teachers were more fluent with eliciting the student's process than the student's conceptual understanding. Further, they focused more on eliciting the revised method and/or solution than asking about why the mistake was made. We consider the findings in terms of the skills brought by prospective teachers that could be built upon in teacher education, skills that need to be learned, and skills that need to be unlearned.
A correction to this paper has been published: https://doi.org/10.1007/s10857-021-09491-7. © 2021, Springer Nature B.V.
We systematically compared two coding approaches to generate training datasets for machine learning (ML): (i) a holistic approach based on learning progression levels and (ii) a dichotomous, analytic approach of multiple concepts in student reasoning, deconstructed from holistic rubrics. We evaluated four constructed response assessment items for undergraduate physiology, each targeting five levels of a developing flux learning progression in an ion context. Human-coded datasets were used to train two ML models: (i) an 8-classification algorithm ensemble implemented in the Constructed Response Classifier (CRC), and (ii) a single classification algorithm implemented in LightSide Researcher's Workbench. Human coding agreement on approximately 700 student responses per item was high for both approaches with Cohen's kappas ranging from 0.75 to 0.87 on holistic scoring and from 0.78 to 0.89 on analytic composite scoring. ML model performance varied across items and rubric type. For two items, training sets from both coding approaches produced similarly accurate ML models, with differences in Cohen's kappa between machine and human scores of 0.002 and 0.041. For the other items, ML models trained with analytic coded responses and used for a composite score, achieved better performance as compared to using holistic scores for training, with increases in Cohen's kappa of 0.043 and 0.117. These items used a more complex scenario involving movement of two ions. It may be that analytic coding is beneficial to unpacking this additional complexity.
Machine learning (ML) has been increasingly employed in science assessment to facilitate automatic scoring efforts, although with varying degrees of success (i.e., magnitudes of machine-human score agreements [MHAs]). Little work has empirically examined the factors that impact MHA disparities in this growing field, thus constraining the improvement of machine scoring capacity and its wide applications in science education. We performed a meta-analysis of 110 studies of MHAs in order to identify the factors most strongly contributing to scoring success (i.e., high Cohen's kappa [kappa]). We empirically examined six factors proposed as contributors to MHA magnitudes: algorithm, subject domain, assessment format, construct, school level, and machine supervision type. Our analyses of 110 MHAs revealed substantial heterogeneity in kappa(mean=.64; range = .09-.97, taking weights into consideration). Using three-level random-effects modeling, MHA score heterogeneity was explained by the variability both within publications (i.e., the assessment task level: 82.6%) and between publications (i.e., the individual study level: 16.7%). Our results also suggest that all six factors have significant moderator effects on scoring success magnitudes. Among these, algorithm and subject domain had significantly larger effects than the other factors, suggesting that technical features and assessment external features might be primary targets for improving MHAs and ML-based science assessments.


