Publications
We consider the task of automatically generating math word problems (MWPs) of various difficulties that meet the needs of teachers in teaching and testing students in corresponding educational stages. Existing methods fail to produce high-quality problems while allowing the teacher control over the problem difficulty level. In this work, we introduce a controllable MWP generation pipeline that samples from an energy language model with various expert model components for realizing the target attributes. We control the difficulty of the resulting MWPs from mathematical and linguistic aspects by imposing constraints on equations, vocabulary, and topics. We also use other control attributes including fluency and distance to the conditioning sequence to manage language quality and creativity. Experiments and evaluation results demonstrate our approach improves upon the baselines in generating solvable, well-formed, and diverse MWPs of controlled difficulty levels. Lastly, we solicit feedback from various math educators who approve the effectiveness of our system for their MWP design processes. They suggest our outputs align with the expectations of problem designers showing a possibility of using such problem generators in real-life educational scenarios. Our code and data are available on request. © 2023, The Author(s), under exclusive license to Springer Nature Switzerland AG.
Solving mathematical problems is cognitively complex, involving strategy formulation, solution development, and the application of learned concepts. However, gaps in students' knowledge or weakly grasped concepts can lead to errors. Teachers play a crucial role in predicting and addressing these difficulties, which directly influence learning outcomes. However, preemptively identifying misconceptions leading to errors can be challenging. This study leverages historical data to assist teachers in recognizing common errors and addressing gaps in knowledge through feedback. We present a longitudinal analysis of incorrect answers from the 2015-2020 academic years on two curricula, Illustrative Math and EngageNY, for grades 6, 7, and 8. We find consistent errors across 5 years despite varying student and teacher populations. Based on these Common Wrong Answers (CWAs), we designed a crowdsourcing platform for teachers to provide Common Wrong Answer Feedback (CWAF). This paper reports on an in vivo randomized study testing the effectiveness of CWAFs in two scenarios: next-problem-correctness within-skill and next-problem-correctness within-assignment, regardless of the skill. We find that receiving CWAF leads to a significant increase in correctness for consecutive problems within-skill. However, the effect was not significant for all consecutive problems within-assignment, irrespective of the associated skill. This paper investigates the potential of scalable approaches in identifying Common Wrong Answers (CWAs) and how the use of crowdsourced CWAFs can enhance student learning through remediation.
While classroom video data are detailed sources for mining student learning insights, their complex and unstructured nature makes them less than straightforward for researchers to analyze. In this paper, we compared the differences between the processes of expert- informed manual feature engineering and automated feature engi- neering using positional data for predicting student group interac- tion in four middle school and high school mathematics classroom videos. Our results highlighted notable differences, including im- proved model accuracy for the combined (manual features + au- tomated features) models compared to the only-manual-features models (mean AUC = .778 vs. .706) at the cost of feature interpretabil- ity, increased number of features for automated feature engineering (1523 vs. 178), and engineering approach (domain-agnostic in au- tomated vs. domain-knowledge-informed in manual). We carried out feature importance analyses and discuss the implications of the results for potentially augmenting human perspectives about quali- tatively coding classroom video data by confirming and expanding views on which body areas and characteristics may be relevant to the target interaction behavior. Lastly, we discuss our study’s limitations and future work.
We examined associations between mind wandering - where attention shifts from the task at hand to task-unrelated thoughts - and learning outcomes. Our data consisted of 177 students who self-reported mind wandering while reading five long, connected texts on scientific research methods and completed learning assessments targeting multiple depths of processing (rote, inference, integration) at different timescales (during and after reading each text, after reading all texts, and after a week-long delay). We found that mind wandering negatively predicted measures of factual, text-based (explicit) information and global integration of information across multiple parts of the text, but not measures requiring a local inference on a single sentence. Further, mind wandering only predicted comprehension measures assessed during the reading session and not after a week-long delay. Our findings provide important nuances to the established negative link between mind wandering and learning outcomes, which has predominantly focused on rote comprehension assessed during the learning session itself. Implications for interventions to address mind wandering during learning are discussed. © 2023 Owner/Author.
As evidence grows supporting the importance of non-cognitive factors in learning, computer-assisted learning platforms increasingly incorporate non-academic interventions to influence student learning and learning related-behaviors. Non-cognitive interventions often attempt to influence students' mindset, motivation, or metacognitive reflection to impact learning behaviors and outcomes. In the current paper, we analyze data from five experiments, involving seven treatment conditions embedded in mastery-based learning activities hosted on a computer-assisted learning platform focused on middle school mathematics. Each treatment condition embodied a specific non-cognitive theoretical perspective. Over seven school years, 20,472 students participated in the experiments. We estimated the effects of each treatment condition on students' response time, hint usage, likelihood of mastering knowledge components, learning efficiency, and post-tests performance. Our analyses reveal a mix of both positive and negative treatment effects on student learning behaviors and performance. Few interventions impacted learning as assessed by the post-tests. These findings highlight the difficulty in positively influencing student learning behaviors and outcomes using non-cognitive interventions. [This paper was published in: "LAK23: 13th International Learning Analytics and Knowledge Conference Proceedings," March 13-17, 2023.]
There have been numerous efforts documenting the effects of open science in existing papers; however, these efforts typically only consider the author's analyses and supplemental materials from the papers. While understanding the current rate of open science adoption is important, it is also vital that we explore the factors that may encourage such adoption. One such factor may be publishing organizations setting open science requirements for submitted articles: encouraging researchers to adopt more rigorous reporting and research practices. For example, within the education technology discipline, the ACM Conference on Learning @ Scale (L@S) has been promoting open science practices since 2018 through a Call For Papers statement. The purpose of this study was to replicate previous papers within the proceedings of L@S and compare the degree of open science adoption and robust reproducibility practices to other conferences in education technology without a statement on open science. Specifically, we examined 93 papers and documented the open science practices used. We then attempted to reproduce the results with invitation from authors to bolster the chance of success. Finally, we compared the overall adoption rates to those from other conferences in education technology. Although the overall responses to the survey were low, our cursory review suggests that researchers at L@S might be more familiar with open science practices compared to the researchers who published in the International Conference on Artificial Intelligence in Education (AIED) and the International Conference on Educational Data Mining (EDM): 13 of 28 AIED and EDM responses were unfamiliar with preregistrations and 7 unfamiliar with preprints, while only 2 of 7 L@S responses were unfamiliar with preregistrations and 0 with preprints. The overall adoption of open science practices at L@S was much lower with only 1% of papers providing open data, 5% providing open materials, and no papers had a preregistration. All openly accessible work can be found in an Open Science Framework project(1).
Computer science learning in primary school classrooms has expanded, necessitating effective instructional strategies for this age group. GalleryWalks are a common activity to allow peers to share their work and give feedback on peers' work. Like other skills, providing effective feedback may require scaffolding and/or instruction for some students. Currently, there is little work exploring how CS students provide peer feedback without extensive instruction, especially at the primary school level (aged 5-11). We analyzed the feedback 4th grade students gave on a structured worksheet during a gallery walk. We found that students often provided both compliments and suggestions when prompted, but their feedback focused primarily on program aesthetics (e.g., characters, sounds, or storyline) rather than programmed elements, even when directed to focus on the program. In addition, bilingual and non-bilingual classes differed in how students followed the worksheet structure.
Motivation. Teachers can play a role in disrupting social inequities that are reflected in education, such as racial disparities in who succeeds in CS. Professional learning addressing inequities causes teachers to confront difficult topics, including how their own identities impact these problems. Understanding the differing ways teachers' identities surface can provide insights into designing better supports for their professional learning. Objectives. The goal of this paper is to examine the teaching and racial identities of two secondary CS teachers who participated in professional learning focused on combining CS content and equity pedagogy. The second goal of this paper is to demonstrate how discourse analytic methods can be used to examine interviews and other interactional data. Method. Teachers were interviewed individually about their teaching identity, racial identity, and professional learning. Drawing on Bucholtz and Hall's identity and interaction framework, interviews were examined for linguistic and discursive features reflecting positionality (i.e., how identity surfaces through the way individuals present themselves to and are perceived by others) and indexicality (i.e., various ways of referring to an identity). Results. Participants used personal deictics, quotative markers, code choice, and affective and epistemic stances when discussing and negotiating their identities with the interviewer. The data reflected ways teachers problematized questions about teaching identity, negotiated tensions in their disciplinary identities, found the topic of race difficult to address, and highlighted other aspects of their identities relevant to understanding and discussing race. Discussion.The study provides a demonstration of how discourse analytic methods can reveal nuances of teacher identity that may be overlooked with other qualitative approaches. Findings also revealed how teachers' ethnic identities might be used as a lever in helping teachers discuss the difficult topic of race in education. Discourse analytic methods are encouraged for future CS education research focused on interactional analyses.
Debugging is a distinct subject in programming that is both comprehensive and challenging for novice programmers. However, instructors have limited opportunities to gain insights into the difficulties students encountered in isolated debugging processes. While qualitative studies have identified debugging strategies that novice programmers use and how they relate to theoretical debugging frameworks, limited larger scale quantitative analyses have been conducted to investigate how students' debugging behaviors observed in log data align with the identified strategies and how they relate to successful debugging. In this study, we used submission log data to understand how the existing debugging strategies are employed by students in an introductory CS course when solving homework problems. We identified strategies from existing debugging literature that can be observed with trace data and extracted features to reveal how efficient debugging is associated with debugging strategy usage. Our findings both align with and contradict past assumptions from previous studies by suggesting that minor code edition can be a beneficial strategy and that width and depth aggregations of the same debugging behavior can reveal opposite effects on debugging efficiency. © 2023 ACM.
Many online learning platforms and MOOCs incorporate some amount of video-based content into their platform, but there are few randomized controlled experiments that evaluate the effectiveness of the different methods of video integration. Given the large amount of publicly available educational videos, an investigation into this content's impact on students could help lead to more effective and accessible video integration within learning platforms. In this work, a new feature was added into an existing online learning platform that allowed students to request skill-related videos while completing their online middle-school mathematics assignments. A total of 18,535 students participated in two large-scale randomized controlled experiments related to providing students with publicly available educational videos. The first experiment investigated the effect of providing students with the opportunity to request these videos, and the second experiment investigated the effect of using a multi-armed bandit algorithm to recommend relevant videos. Additionally, this work investigated which features of the videos were significantly predictive of students' performance and which features could be used to personalize students' learning. Ultimately, students were mostly disinterested in the skill-related videos, preferring instead to use the platforms existing problem-specific support, and there was no statistically significant findings in either experiment. Additionally, while no video features were significantly predictive of students' performance, two video features had significant qualitative interactions with students' prior knowledge, which showed that different content creators were more effective for different groups of students. These findings can be used to inform the design of future video-based features within online learning platforms and the creation of different educational videos specifically targeting higher or lower knowledge students. The data and code used in this work can be found at https://osf.io/cxkzf/.


