Publications
This paper introduces innovative software for efficient English language learning that incorporates machine learning and natural language processing techniques to personalize the vocabulary acquisition process for English language learners. The software was designed to enhance extensive reading by generating English language learning materials for learners based on their interests and proficiency levels. The software begins by administering a brief and straightforward vocabulary test to learners. The test is used to identify words that may be unknown to the learners from 12,000 words. The machine-learning algorithm identifies words that are most likely to be unknown to the learner based on their test performance. The software then prompts the learner to select a topic of interest such as science or music. Thereafter, the software generates personalized English language learning materials for the learner, which contain texts with specific vocabulary related to the selected topic. The material is generated using ChatGPT. The software highlights unknown words in the text, which the learner can check using a dictionary. The software then generates new material incorporating the unknown words that the learner has checked, ensuring that the learner is exposed to a wide vocabulary in his area of interest. This process is repeated multiple times with the software generating new materials and incorporating new words that the learner has checked, thereby facilitating the efficient acquisition of new vocabulary. Through this process, the learners can engage in extensive reading, enabling them to read more in English and develop their reading skills, while simultaneously acquiring new vocabulary related to their interests. The innovative approach of the software in English language learning offers a personalized, adaptive, and efficient approach to extensive reading. This can help learners improve their English proficiency by reading in a manner tailored to their interests and proficiency levels. © 2023, The Author(s), under exclusive license to Springer Nature Switzerland AG.
Classroom observation is an effective way for teachers to improve professional development, and the analysis of student-teacher interactions is critical and significant to classroom observation. However, the traditional methods of the classroom observation are mainly based on manual coding by domain experts. Although several studies have been conducted to automate the coding and analyzing process, they are either based on audio information or video information collected from the classroom, which fails to jointly utilize multimodal information like domain experts. We thus propose a student-teacher multimodal interaction analysis system that conducts the analysis using both video and audio information and accordingly generates the informative reports based on the analysis results. A preliminary evaluation of the system validates the effectiveness of the built system and the analysis results could be further used for the evidence-based teaching behavior evaluation. The current limitations and possible optimization on the built system are discussed as well. We are planning to keep improving the system and deploy it to 1000 schools located at the rural areas in three years. © 2023, The Author(s), under exclusive license to Springer Nature Switzerland AG.
We present a randomized field trial delivered in Carnegie Learning’s MATHia’s intelligent tutoring system to 12,374 learners intended to test whether rewriting content in “word problems” improves student mathematics performance within this content, especially among students who are emerging as English language readers. In addition to describing facets of word problems targeted for rewriting and the design of the experiment, we present an artificial intelligence-driven approach to evaluating the effectiveness of the rewrite intervention for emerging readers. Data about students’ reading ability is generally neither collected nor available to MATHia’s developers. Instead, we rely on a recently developed neural network predictive model that infers whether students will likely be in this target sub-population. We present the results of the intervention on a variety of performance metrics in MATHia and compare performance of the intervention group to the entire user base of MATHia, as well as by comparing likely emerging readers to those who are not inferred to be emerging readers. We conclude with areas for future work using more comprehensive models of learners. © 2023, The Author(s), under exclusive license to Springer Nature Switzerland AG.
Interactive simulations encourage students to practice skills essential to understanding and learning sciences. Alas, inquiry learning with interactive simulations is challenging. In this paper, we seek to identify inquiry patterns across topics and evaluate their stability with regard to common behaviors and student membership. Applying a clustering approach, we propose an encoding through which we can model students’ strategies in diverse environments. Specifically, we encode each sequence with three different levels of granularity which range from simulation-specific characteristics to simulation-agnostic features. Using this generalizable encoding, we find two clusters for each of two simulations. The formed groups exhibit similar learning patterns across environments. One systematically cycles through exploring and recording systematically over all variables. The other group explores the simulation more freely. This suggests that our feature encoding captures inherent quality of inquiry with simulations and can be used to characterize learners knowledge of productive exploration. © 2023, The Author(s), under exclusive license to Springer Nature Switzerland AG.
Item selection is the key process for computerized adaptive testing (CAT) to effectively assess examinees’ knowledge states. Existing item selection algorithms mainly rely on information metrics, suffering two issues: one is that the implicit cognitive information like relations between testing items as well as knowledge components cannot be captured by the information-based methods, and the other one is that the information-based algorithms computes item’s suitableness depending on examinees’ knowledge states which are estimated and imprecise inherently. To address these two issues, this work proposes to employ reinforcement learning technology to learn the item selection algorithm automatically in a data-driven manner. It is also able to properly capture the implicit cognitive relations between different testing items and avoid unnecessary item testing, and does not depend on examinees’ estimated knowledge states at all. © 2023, The Author(s), under exclusive license to Springer Nature Switzerland AG.
Improving social presence has been popularly researched for massive open online courses (MOOCs). Despite this, little consideration has been given to how students intend to utilize new technology from a diversity, equity, and inclusion (DEI) standpoint. In this study, we examined the role of social presence in diverse MOOC students’ behavioral intentions and sentiments toward using a learning assistant chatbot. Considering potentially disparate perceptions of the technology depending on the students’ demographic factors (age, gender, region, and native language), we investigated the relationships among their social presence, age, behavioral intentions (before and after use), sentiment levels, and numbers of turn-takings with the chatbot. The investigation showed positive and moderate correlations between social presence and post-behavioral intention and between social presence and pre-behavioral intention. Also, the sentiment and social presence toward the chatbot turned out to be different according to the students’ region factor based on 13 different geographical categorizations. The native language factor was also influential in creating different sentiment levels between native and non-native English user groups. The findings implicate the importance of promoting a better understanding of the role of social presence in students’ learning experiences with AI-based applications from a DEI perspective. © 2023, The Author(s), under exclusive license to Springer Nature Switzerland AG.
In this paper, we explore the use of data augmentation through generative adversarial networks (GANs) for improving the performance of machine learning models in detecting at-risk students in the context of e-learning institutions. It is well known that balancing datasets can have a positive effect on improving the performance of machine learning models, especially for deep neural networks. However, undersampling can potentially result in the loss of valuable data, so data augmentation seems to be more meaningful solution when the dataset is relatively small. One of the most popular data augmentation approaches is the use of GAN networks due to their ability to generate high-quality synthetic samples that belong to the distribution of the original dataset. On the other hand, detecting at-risk students is a hot topic in learning analytics, and ability to detect these students early with high accuracy enables e-learning institutions to take necessary steps to motivate and retain students during the course. We apply this approach to the OULA dataset, a commonly used dataset in learning analytics that includes labeled at-risk students. The OULA dataset is not highly-imbalanced, making it more challenging to improve model performance through these techniques. © 2023, The Author(s), under exclusive license to Springer Nature Switzerland AG.
Simulation-based experiential learning environments used in nurse training programs offer numerous advantages, including the opportunity for students to increase their self-confidence through deliberate repeated practice in a safe and controlled environment. However, measuring and monitoring students’ self-confidence is challenging due to its subjective nature. In this work, we show that students’ self-confidence can be predicted using multimodal data collected from the training environment. By extracting features from student eye gaze and speech patterns and combining them as inputs into a single regression model, we show that students’ self-rated confidence can be predicted with high accuracy. Such predictive models may be utilized as part of a larger assessment framework designed to give instructors additional tools to support and improve student learning and patient outcomes. © 2023, The Author(s), under exclusive license to Springer Nature Switzerland AG.
We develop models to classify desirable reasoning revisions in argumentative writing. We explore two approaches – multi-task learning and transfer learning – to take advantage of auxiliary sources of revision data for similar tasks. Results of intrinsic and extrinsic evaluations show that both approaches can indeed improve classifier performance over baselines. While multi-task learning shows that training on different sources of data at the same time may improve performance, transfer-learning better represents the relationship between the data. © 2023, The Author(s), under exclusive license to Springer Nature Switzerland AG.
Providing timely assistance to students in intelligent tutoring systems is a challenging research problem. In this study, we aim to address this problem by determining when to provide proactive help with autoencoder based feature learning and a deep reinforcement learning (DRL) model. To increase generalizability, we only use domain-independent features for the policy. The proposed pedagogical policy provides next-step proactive hints based on the prediction of the DRL model. We conduct a study to examine the effectiveness of the new policy in an intelligent logic tutor. Our findings provide insight into the use of DRL policies utilizing autoencoder based feature learning to determine when to provide proactive help to students. © 2023, The Author(s), under exclusive license to Springer Nature Switzerland AG.


