Publications
This study investigates the effects on different racial/ethnic groups of middle school students when learning with a digital learning game, Decimal Point, and a comparable computer tutor. Using data from three classroom studies with 835 students, we compared learning outcomes and engagement among students from racial/ethnic groups that are well-represented in STEM (white and Asian) to those that are underrepresented in STEM (Black, Hispanic/Latine, Indigenous, and multiracial). Relative to students from underrepresented groups, students from well-represented groups in STEM scored higher on all tests (pre, post, and delayed, despite similar learning gains from pre-to-post and pre-to-delayed) and showed more engagement and less anxiety. The game also enhanced the experience of mastery only among students from well-represented groups. At the same time, students from underrepresented groups learned from the intervention and matched students from well-represented groups in learning efficiency. In short, we found similar learning gains from the game and tutor interventions among students from well-represented and underrepresented racial/ethnic groups, despite the lower performance and lower engagement among students from underrepresented groups. These insights highlight how students from diverse backgrounds may engage differently with educational technology, guiding future efforts in making Decimal Point - as well as digital learning tools in general - more inclusive.
We conducted a 2 x 2 study comparing the digital learning game Decimal Point to a comparable non-game tutor with or without self-explanation prompting. We expected to replicate previous studies showing the game improved learning compared to the non-game tutor, and that self-explanation prompting would enhance learning across platforms. Additionally, prior research with Decimal Point suggested that self-explanation was driving gender differences in which girls learned more than boys. To better understand these effects, we manipulated the presence of self-explanation prompts and incorporated a multidimensional gender measure. We hypothesized that girls and students with stronger feminine-typed characteristics would learn more than boys and students with stronger masculine-typed characteristics in the game with self-explanation condition, but not in the game without self-explanation or in the non-game conditions. Results showed no advantage for the game over the non-game or for including self-explanation, but an analysis of hint usage indicated that students in the game conditions used (and abused) hints more than in the non-game conditions, which in turn was associated with worse learning outcomes. When we controlled for hint use, students in the game conditions learned more than students in the non-game tutor. We replicated a gender effect favoring boys and students with masculine-typed characteristics on the pretest, but there were no gender differences on the posttests. Finally, results indicated that the multidimensional framework explained variance in pretest performance better than a binary gender measure, adding further evidence that this framework may be a more effective, inclusive approach to understanding gender effects in game-based learning.
Dialogue Acts (DAs) can be used to explain what expert tutors do and what students know during the tutoring process. Most empirical studies adopt the random sampling method to obtain sentence samples for manual annotation of DAs, which are then used to train DA classifiers. However, these studies have paid little attention to sample informativeness, which can reflect the information quantity of the selected samples and inform the extent to which a classifier can learn patterns. Notably, the informativeness level may vary among the samples and the classifier might only need a small amount of low informative samples to learn the patterns. Random sampling may overlook sample informativeness, which consumes human labelling costs and contributes less to training the classifiers. As an alternative, researchers suggest employing statistical sampling methods of Active Learning (AL) to identify the informative samples for training the classifiers. However, the use of AL methods in educational DA classification tasks is under-explored. In this paper, we examine the informativeness of annotated sentence samples. Then, the study investigates how the AL methods can select informative samples to support DA classifiers in the AL sampling process. The results reveal that most annotated sentences present low informativeness in the training dataset and the patterns of these sentences can be easily captured by the DA classifier. We also demonstrate how AL methods can reduce the cost of manual annotation in the AL sampling process. © 2023, The Author(s), under exclusive license to Springer Nature Switzerland AG.
Mentoring promotes underserved students’ persistence in STEM but is difficult to scale up. Conversational virtual agents can help address this problem by conveying a mentor’s experiences to larger audiences. The present study examined college students’ (N= 138 ) utilization of CareerFair.ai, an online platform featuring virtual agent-mentors that were self-recorded by sixteen real-life mentors and built using principles from the earlier MentorPal framework. Participants completed a single-session study which included 30 min of active interaction with CareerFair.ai, sandwiched between pre-test and post-test surveys. Students’ user experience and learning gains were examined, both for the overall sample and with a lens of diversity and equity across different, potentially underserved demographic groups. Findings included positive pre/post changes in intent to pursue STEM coursework and high user acceptance ratings (e.g., expected benefit, ease of use), with under-represented minority (URM) students giving significantly higher ratings on average than non-URM students. Self-reported learning gains of interest, actual content viewed on the CareerFair.ai platform, and actual learning gains were associated with one another, suggesting that the platform may be a useful resource in meeting a wide range of career exploration needs. Overall, the CareerFair.ai platform shows promise in scaling up aspects of mentoring to serve the needs of diverse groups of college students. © 2023, The Author(s), under exclusive license to Springer Nature Switzerland AG.
Developing models and using mathematics are two key practices in internationally recognized science education standards, such as the Next Generation Science Standards (NGSS) [1]. However, students often struggle at the intersection of these practices, i.e., developing mathematical models about scientific phenomena. In this paper, we present the design and initial classroom test of AI-scaffolded virtual labs that help students practice these competencies. The labs automatically assess fine-grained sub-components of students’ mathematical modeling competencies based on the actions they take to build their mathematical models within the labs. We describe how we leveraged underlying machine-learned and knowledge-engineered algorithms to trigger scaffolds, delivered proactively by a pedagogical agent, that address students’ individual difficulties as they work. Results show that students who received automated scaffolds for a given practice on their first virtual lab improved on that practice for the next virtual lab on the same science topic in a different scenario (a near-transfer task). These findings suggest that real-time automated scaffolds based on fine-grained assessment data can help students improve on mathematical modeling. © 2023, The Author(s), under exclusive license to Springer Nature Switzerland AG.
In education, intelligent learning environments allow students to choose how to tackle open-ended tasks while monitoring performance and behavior, allowing for the creation of adaptive support to help students overcome challenges. Timely feedback is critical to aid students’ progression toward learning and improved problem-solving. Feedback on text-based student responses can be delayed when teachers are overloaded with work. Automated evaluation can provide quick student feedback while easing the manual evaluation burden for teachers in areas with a high teacher-to-student ratio. Current methods of evaluating student essay responses to questions have included transformer-based natural language processing models with varying degrees of success. One main challenge in training these models is the scarcity of data for student-generated data. Larger volumes of training data are needed to create models that perform at a sufficient level of accuracy. Some studies have vast data, but large quantities are difficult to obtain when educational studies involve student-generated text. To overcome this data scarcity issue, text augmentation techniques have been employed to balance and expand the data set so that models can be trained with higher accuracy, leading to more reliable evaluation and categorization of student answers to aid teachers in the student’s learning progression. This paper examines the text-generating AI model, GPT-3.5, to determine if prompt-based text-generation methods are viable for generating additional text to supplement small sets of student responses for machine learning model training. We augmented student responses across two domains using GPT-3.5 completions and used that data to train a multilingual BERT model. Our results show that text generation can improve model performance on small data sets over simple self-augmentation. © 2023, The Author(s), under exclusive license to Springer Nature Switzerland AG.
Reading comprehension is essential for both knowledge acquisition and memory reinforcement. Automated modeling of the comprehension process provides insights into the efficacy of specific texts as learning tools. This paper introduces an improved version of the Automated Model of Comprehension, version 3.0 (AMoC v3.0). AMoC v3.0 is based on two theoretical models of the comprehension process, namely the Construction-Integration and the Landscape models. In addition to the lessons learned from the previous versions, AMoC v3.0 uses Transformer-based contextualized embeddings to build and update the concept graph as a simulation of reading. Besides taking into account generative language models and presenting a visual walkthrough of how the model works, AMoC v3.0 surpasses the previous version in terms of the Spearman correlations between our activation scores and the values reported in the original Landscape Model for the presented use case. Moreover, features derived from AMoC significantly differentiate between high-low cohesion texts, thus arguing for the model’s capabilities to simulate different reading conditions. © 2023, The Author(s), under exclusive license to Springer Nature Switzerland AG.
In second language vocabulary learning, existing works have primarily focused on either the learning interface or scheduling personalized retrieval practices to maximize memory retention. However, the learning content, i.e., the information presented on flashcards, has mostly remained constant. Keyword mnemonic is a notable learning strategy that relates new vocabulary to existing knowledge by building an acoustic and imagery link using a keyword that sounds alike. Beyond that, producing verbal and visual cues associated with the keyword to facilitate building these links requires a manual process and is not scalable. In this paper, we explore an opportunity to use large language models to automatically generate verbal and visual cues for keyword mnemonics. Our approach, an end-to-end pipeline for auto-generating verbal and visual cues, can automatically generate highly memorable cues. We investigate the effectiveness of our approach via a human participant experiment by comparing it with manually generated cues. © 2023, The Author(s), under exclusive license to Springer Nature Switzerland AG.
Writing argumentative essays is a critical component of students’ learning. Previous works on automatic assessments on essay writing often focused on providing a holistic score for the input essay, which only summarized the essay’s overall quality. However, to provide more pedagogical value and equitable educational opportunities for all students, an automatized system needs to provide detailed feedback on students’ essays. To address this issue, we developed an essay argumentative structure feedback system to support educators and students. We employed natural language processing (NLP) and data mining techniques to explore the association between argumentative structure and essay scores. First, we proposed a cross-prompt, sentence-level ensemble model to classify the argumentative elements and extract the argumentative structures from the essay. The model worked across multiple datasets and achieved high performance. Second, after applying the classification model on the ACT writing tests, we performed a sequential mining process to extract representative argumentative structures. Our findings highlight the role of organizational argumentative structure in essay scoring. Furthermore, we found a common argumentative structure used by the high-scored essays. Finally, with the knowledge of argumentative elements and structures used in the previous essays, we proposed a feedback tool design to complement the current AES systems and help students improve their argument writing skill. © 2023, The Author(s), under exclusive license to Springer Nature Switzerland AG.
This work compares two approaches to provide metacognitive interventions and their impact on preparing students for future learning across Intelligent Tutoring Systems (ITSs). In two consecutive semesters, we conducted two classroom experiments: Exp. 1 used a classic artificial intelligence approach to classify students into different metacognitive groups and provide static interventions based on their classified groups. In Exp. 2, we leveraged Deep Reinforcement Learning (DRL) to provide adaptive interventions that consider the dynamic changes in the student's metacognitive levels. In both experiments, students received these interventions that taught how and when to use a backward-chaining (BC) strategy on a logic tutor that supports a default forward-chaining strategy. Six weeks later, we trained students on a probability tutor that only supports BC without interventions. Our results show that adaptive DRL-based interventions closed the metacognitive skills gap between students. In contrast, static classifier-based interventions only benefited a subset of students who knew how to use BC in advance. Additionally, our DRL agent prepared the experimental students for future learning by significantly surpassing their control peers on both ITSs.


