Publications
Automatic Short Answer Grading aims to automatically grade short answers authored by students. Recent work has shown that this task can be effectively reformulated as a Natural Language Inference problem. State-of-the-art is defined by the use of large pretrained language models fine-tuned in the domain dataset. But how to quantify the effectiveness of the models in small data regimes still remains an open issue. In this work we present a set of experiments to analyse the impact of different annotation strategies when not enough training examples for fine-tuning the model are available. We find that when annotating few examples, it is preferable to have more question variability than more answers per question. With this annotation strategy, our model outperforms state-of-the-art systems utilizing only 10% of the full-training set. Finally, experiments show that the use of out-of-domain annotated question-answer examples can be harmful when fine-tuning the models. © 2023, The Author(s), under exclusive license to Springer Nature Switzerland AG.
Automated feedback as students answer open-ended math questions has significant potential in improving learning outcomes at large scale. A key part of automated feedback systems is an error classification component, which identifies student errors and enables appropriate, predefined feedback to be deployed. Most existing approaches to error classification use a rule-based method, which has limited capacity to generalize. Existing data-driven methods avoid these limitations but specifically require mathematical expressions in student responses to be parsed into syntax trees. This requirement is itself a limitation, since student responses are not always syntactically valid and cannot be converted into trees. In this work, we introduce a flexible method for error classification using pre-trained large language models. We demonstrate that our method can outperform existing methods in algebra error classification, and is able to classify a larger set of student responses. Additionally, we analyze common classification errors made by our method and discuss limitations of automated error classification. © 2023, The Author(s), under exclusive license to Springer Nature Switzerland AG.
The Physiology Core Concept of flow down gradients is a major concept in physiology, as pressure gradients are the key driving force for the bulk flow of fluids in biology. However, students struggle to understand that this principle is foundational to the mechanisms governing bulk flow across diverse physiological systems (e.g., blood flow, phloem sap flow). Our objective was to investigate whether bulk flow items that differ in scenario context (i.e., taxa, amount of scientific terminology, living or nonliving system) or in which aspect of the pressure gradient is kept constant (i.e., starting pressure or pressure gradient) influence under-graduate students' reasoning. Item scenario context did not impact the type of reasoning students used. However, students were more likely to use the Physiology Core Concept of flow down [pressure] gradients when the pressure gradient was kept constant and less likely to use this concept when the starting pressure was kept constant. We also investigated whether item scenario context or which aspect of the pressure gradient is kept constant impacted how consistent students were in the type of reasoning they used across two bulk flow items on the same homework. Most students were consistent across item scenario contexts (76%) and aspects of the pressure gradient kept constant (70%). Students who reasoned using flow down gradients on the first item were the most consistent (86, 89%), whereas students using pressures indicate (but don't cause) flow were the least consistent (43, 34%). Students who are less consistent know that pressure is somehow involved or indicates fluid flow but do not have a firm grasp of the concept of a pressure gradient as the driving force for fluid flow. These findings are the first empirical evidence to support the claim that using Physiology Core Concept reasoning supports transfer of knowledge across different physiological systems.NEW & NOTEWORTHY These findings are the first empirical evidence to support the claim that using Physiology Core Concept reasoning supports transfer of knowledge across different physiological systems.
This study examines the relationship between alternatively certified mathematics teachers' stated reasons for entry and their odds of retention at the school level and at the district level. Study participants were members of the 2006 and 2007 cohorts of New York City Teaching Fellows who completed three surveys over a 9-year period. Administrative data sets from the New York City Department of Education (NYCDOE) provided employment history of cohort members as well as demographic information about the teachers and sites of employment in New York City (NYC) public schools. Drawing on retention and survey data, we found that, of the four reasons for entry factors, two were predictive of NYCTF mathematics teacher retention (i.e., job benefits and alternative certification) and two were not (i.e., altruism and meaningful job). Given the cost associated with recruiting and training alternatively certified teachers, information to improve the initial selection process and increase the rate of retention makes financial sense for districts that employ alternatively certified teachers.
Collaborative game-based learning environments offer the promise of combining the strengths of computer-supported collaborative learning and gamebased learning to enable students to work collectively towards achieving problemsolving goals in engaging storyworlds. Group chat plays an important role in such environments, enabling students to communicate with team members while exploring the learning environment and collaborating on problem solving. However, students may engage in chat behavior that negatively affects learning. To help address this problem, we introduce a multidimensional stealth assessment model for jointly predicting students' out-of-domain contributions to group chat as well as their learning outcomes with multi-task learning. Results from evaluating the model indicate that multi-task learning, which simultaneously performs the multidimensional stealth assessment, utilizing predictive features extracted from in-game actions and group chat data outperforms single-task variants and suggest that multi-task learning can effectively support stealth assessment in collaborative game-based learning environments.
We analyzed a population-based cohort (N = 10,922) to investigate the onset and stability of racial and ethnic disparities in advanced (i.e., above the 90(th) percentile) science and mathematics achievement during elementary school as well as the antecedent, opportunity, and propensity factors that explained these disparities. About 13% to 16% of White students versus 3% to 4% of Black or Hispanic students displayed advanced science or mathematics achievement during kindergarten. The antecedent factor of family socioeconomic status and the propensity factors of student science, mathematics, and reading achievement by kindergarten consistently explained whether students displayed advanced science or mathematics achievement during first, second, third, fourth, or fifth grade. These and additional factors substantially or fully explained initially observed disparities between Black or Hispanic and White students in advanced science or mathematics achievement during elementary school. Economic and educational policies designed to increase racial and ethnic representation in STEM course taking, degree completion, and workforce participation may need to begin by elementary school.
Prior work on teacher candidates in Washington State has shown that about two thirds of individuals who trained to become teachers between 2005 and 2015 and received a teaching credential did not enter the state's public teaching workforce immediately after graduation, while about one third never entered a public teaching job in the state at all. In this analysis, we link data on these teacher candidates to unemployment insurance data in the state to provide a descriptive portrait of the future earnings and wages of these individuals inside and outside of public schools. Candidates who initially became public school teachers earned considerably more, on average, than candidates who were initially employed either in other education positions or in other sectors of the state's workforce. These differences persisted ten years into the average career and across transitions into and out of teaching. There is therefore little evidence that teacher candidates who did not become teachers were lured into other professions by higher compensation. Instead, the patterns are consistent with demand-side constraints on teacher hiring during this time period that resulted in individuals who wanted to become teachers taking positions that offered lower wages but could lead to future teaching positions.
Mass balance (MB) reasoning offers a rich topic for examination of students' scientific thinking and skills, as it requires students to account for multiple inputs and outputs within a system and apply covariational reasoning. Using previously validated constructed response prompts for MB, we examined 1,920 student-constructed responses (CRs) aligned to an emerging learning progression to determine how student language changes from low (1) to high (4) covariational reasoning levels. As students' abilities and thinking change with Context, we used the same general prompt in six physiological contexts. We asked how Level and Context affect student language and what language is conserved across Contexts at higher reasoning Levels. Using diversity methods, we found student language becomes more similar as covariational reasoning level increases. Using text analysis, we found context-dependent words at each Level; however, the type of context words changed. Specifically, at Level 1, students used context words that are tangential to MB reasoning, while Level 4 responses used words that specify inputs and outputs for the given Item Context. Further, at Level 4, students shared 30% of language across the six contexts and leveraged context-independent words including rate, equal, and some form of slower/lower/smaller. Together, these data demonstrate that Context affects undergraduate MB language at all covariational reasoning levels, but that the language becomes more specific and similar as Level increases. These findings encourage instructors to foster context-independent, comparative, and summative language during instruction to functionally build MB and covariational reasoning skills across contexts.
The basis for mastering neurophysiology is understanding ion movement across cell membranes. The Electrochemical Gradients Assessment Device (EGAD) is a 17-item test assessing students' understanding of fundamental concepts of neurophysiology, e.g., electrochemical gradients and resistance, synaptic transmission, and stimulus strength. We collected responses to the EGAD from 534 students from seven institutions nationwide, before and after instruction. We determined the relative difficulty of neurophysiology topics and noted that students did better on what questions compared to how questions, particularly those integrating concentration gradient and electric forces to predict ion movement. We also found that, even after instruction, stu-dents selected one incorrect answer, at a rate greater than random chance for nine questions. We termed these incorrect answers attractive distractors. Most attractive distractors contained terms associated with concentration gradients, equilibrium, or anthropomorphic and teleological reasoning, and incorrect answers containing multiple terms were more attractive. We used v2 analysis and alluvial diagrams to investigate how individual students moved or did not move between answer choices on the pre-and posttest. Interestingly, students selecting the attractive distractor on the pretest were just as likely as other incorrect students to move to the correct answer on the posttest. In contrast, of students incorrect on both the pre-and posttest, students who selected the attractive distractor on the pretest were more likely to stick with this answer on the posttest than students choosing other incorrect answers. Combining the EGAD results with alluvial diagrams can inform neurophysiology instruction to address points of student confusion.NEW & NOTEWORTHY Investigating students' alternative reasoning in neurophysiology, this research is the first to investigate how analyzing the most common incorrect answer can shed light on the concepts students struggle with when reasoning about neurophysiological problems, especially those dealing with both chemical and electrical driving forces to predict ion movement across cell membranes.
Proof Blocks is a software tool that allows students to practice writing mathematical proofs by dragging and dropping lines instead of writing proofs from scratch. Proof Blocks offers the capability of assigning partial credit and providing solution quality feedback to students. This is done by computing the edit distance from a student’s submission to some predefined set of solutions. In this work, we propose an algorithm for the edit distance problem that significantly outperforms the baseline procedure of exhaustively enumerating over the entire search space. Our algorithm relies on a reduction to the minimum vertex cover problem. We benchmark our algorithm on thousands of student submissions from multiple courses, showing that the baseline algorithm is intractable, and that our proposed algorithm is critical to enable classroom deployment. Our new algorithm has also been used for problems in many other domains where the solution space can be modeled as a DAG, including but not limited to Parsons Problems for writing code, helping students understand packet ordering in networking protocols, and helping students sketch solution steps for physics problems. Integrated into multiple learning management systems, the algorithm serves thousands of students each year. © 2023, The Author(s), under exclusive license to Springer Nature Switzerland AG.


