Publications
Machine learning models were constructed to predict student performance in an introductory mechanics class at a large land-grant university in the United States using data from 2061 students. Students were classified as either being at risk of failing the course (earning a D or F) or not at risk (earning an A, B, or C). The models focused on variables available in the first few weeks of the class which could potentially allow for early interventions to help at-risk students. Multiple types of variables were used in the model: in-class variables (average homework and clicker quiz scores), institutional variables [college grade point average (GPA)], and noncognitive variables (self-efficacy). The substantial imbalance between the pass and fail rates of the course, with only about 10% of students failing, required modification to the machine learning algorithms. Decision threshold tuning and upsampling were successful in improving performance for at-risk students. Logistic regression combined with a decision threshold tuned to maximize balanced accuracy yielded the strongest classifier, with a DF accuracy of 83% and an ABC accuracy of 81%. Measures of variable importance involving changes in balanced accuracy identified homework grades, clicker grades, college GPA, and the fraction of college classes successfully completed as the most important variables in predicting success in introductory physics. Noncognitive variables added little predictive power to the models. Classification models with performance near the best-performing models using the full set of variables could be constructed with very few variables (homework average, clicker scores, and college GPA) using straightforward to implement algorithms, suggesting the application of these technologies may be fairly easy to include in many physics classes.
This study proposes methods of reporting results of physics conceptual evaluations that more fully characterize the range of outcomes experienced by students with differing levels of prior preparation, allowing for more meaningful comparison of the outcomes of educational interventions within and across institutions. Factors leading to variation in post-test scores on the Force and Motion Conceptual Evaluation (FMCE) across different instructors, semesters, and course models in a sample collected in introductory calculus-based mechanics at a large, eastern land-grant university were examined. The sample was collected over nine years and contains a total of N 1/4 4409 matched pretest and post-test records. The data showed a systematic semester-by-semester variation in both pretest scores and ACT or SAT mathematics percentile scores. Neither the normalized gain nor Cohen's d removed the semester-to-semester variation observed in post-test scores. The local average curve plotting post-test scores against pretest scores, which we call a conceptual growth curve, allowed for the characterization of outcomes for students with different pretest scores. Regression models were used to produce an approximation to this curve. By using either the full curve or a mathematical approximation developed through linear regression, the post-test score that would be observed if a class enrolled students with a given level of prior preparation measured by pretest scores can be predicted. This predicted post-test score can then be used to calculate the predicted normalized gain if desired. These methods rely on using the natural variation of incoming student preparation at one institution to predict how a class would perform if it enrolled students with different prior preparation. The study presents an example of converting the outcomes at an institution with a weakly prepared student population to the outcomes which would have been observed if the course enrolled a more prepared student population; converting the outcomes for a different student population dramatically changed the interpretation of how the class studied was functioning.
Curricular analytics (CA) is a quantitative method that analyzes the sequence of courses (curriculum) that students in an undergraduate academic program must complete to fulfill the requirements of the program. The main hypothesis of CA is that the less complex a curriculum is, the more likely it is that students complete the program. This study compares the curricular complexity of undergraduate physics programs at 60 institutions in the United States. The institutions were divided into three tiers based on national rankings of the physics graduate program, and the means of each tier were compared. No significant difference between the means of each tier was found, indicating that there is not a relationship between program curricular complexity and program ranking. Further analysis focused on the physics, chemistry, and mathematics courses, defined as the core courses of the curriculum. Significant differences in the number of required core courses and the complexity per core course were measured between the tiers; both were measured as large effects. Programs with the highest rankings required fewer core courses while having a higher complexity per core course. These institutions have more strict prerequisite requirements than lower ranking programs. This study also showed complexity was quantitatively related to curricular flexibility operationalized as the number of available eight-semester degree plans. The number of available degree plans exponentially decreased with increasing core complexity per course. Modifications to a curriculum at one institution were analyzed; a similar relationship between the number of available degree plans and increasing complexity per core course was found.
Self-regulated learning (SRL) is a cognitive and metacognitive process through which students develop the self-awareness necessary to direct their learning based on their needs to reach a desired outcome. Despite 40 years of literature, SRL has no singular definition, as it is often used in domain-specific research that is not always transferable to other fields. Regardless, much of the literature speaks to the importance of SRL regarding academic success. This paper details the development of an SRL instrument designed to identify key self-regulatory constructs in an undergraduate introductory physics classroom. Confirmatory factor analysis supported a four-factor model measuring Planning, Time & Environment Management, Comprehension Monitoring & Evaluation, and Peer Learning & Help-Seeking as unique facets of self-regulated learning. While most behaviors did not significantly evolve over one semester, students reported significantly lower scores on the Comprehension Monitoring & Evaluation factor between the beginning and end of the semester. Higher performing students, as measured by their average homework grades, scored significantly higher on the Time & Environment Management factor and the Peer Learning & Help-Seeking factor at both time points. Additionally, SRL behaviors were significantly predicted by personality facets from the Big Five Inventory, with Conscientiousness, Extraversion, and Openness being the most related to certain behaviors.
This study examines high school preparation measures [ACT/SAT scores, high school grade point average (HSGPA), and conceptual physics pretest scores], in-class behavior measures (homework submission rates and lecture attendance rates), and in-class achievement measures (homework and test averages) for the last two fully face-to-face prepandemic and the first two fully face-to-face postpandemic semesters of an introductory calculus-based electricity and magnetism class. This class was offered at a large eastern land grant university in the United States. The total number of students for the four semesters was 1033. While some significant differences were measured (higher postpandemic HSGPA, lower postpandemic conceptual pretest scores, higher postpandemic homework average (fall semesters only), and lower postpandemic lecture attendance (spring semesters only), none were larger than a small effect. As such, student achievement, attendance rates, and assignment completion rates were largely unchanged after the pandemic.
The Force Concept Inventory (FCI) is a popular multiple-choice instrument used to measure a student's conceptual understanding of Newtonian mechanics. Recently, a network analytic technique called module analysis has been used to identify responses to the FCI and other conceptual instruments that are preferentially selected together by students; these groups of responses are called communities. This study uses module analysis to explore the misconception structure of the FCI at five U.S. institutions with varying undergraduate populations (sample sizes of N 1/4 9606, 4360, 1496, 466, and 213). Students from these universities had a broad range of prior knowledge in physics and of general high school academic preparation, resulting in large differences in FCI normalized gain, pretest, and post-test scores. In the current work, modified module analysis partial was applied and communities of consistently selected responses within the FCI were identified at the five institutions studied. There was substantial similarity between the communities identified postinstruction; somewhat less similarity preinstruction. This suggests that consistently applied Newtonian misconceptions exist both before and after instruction at a wide range of institutions. The most frequently applied misconceptions were largest force determines motion, Newton's third law misconceptions, and motion implies active forces.These misconceptions were still consistently applied even after instruction by a substantial number of students at all but the highest performing of the five institutions.
Self-efficacy has emerged as one of the most important noncognitive variables explaining academic behavior. It has been shown to influence students' academic and career decisions as well as their academic performance. Multiple studies have reported differences in self-efficacy between men and women in science, technology, engineering, and mathematics classes. A student's personality, characterized by the five-factor model, is also related to academic performance; some personality facets are substantially different for men and women. This work examines the relations among the five-factor model of personality (agreeableness, conscientiousness, extraversion, neuroticism, and openness), self-efficacy toward physics and mathematics, and course outcomes in university physics and mathematics classes. Women reported significantly higher neuroticism in all classes, a medium to large effect size, and significantly higher conscientiousness in Calculus 1 and Physics 1, small effects. Men reported higher self-efficacy in two-semester Calculus 1, one-semester Calculus 1, Physics 1, and Physics 2, small effects. Conscientiousness and neuroticism had competing mediational effects on the relation of gender to self-efficacy. The path through neuroticism accounted for 25%-47% of the total effect of gender on self-efficacy (increasing self-efficacy for men) and the path through conscientiousness accounted for 12%-23% of the total effect (increasing self-efficacy for women). Self-efficacy mediated the relation of conscientiousness to course grade in all classes, accounting for 30%-45% of the total effect.
This study investigated factors influencing Force and Motion Conceptual Evaluation (FMCE) pretest and post-test scores for a sample (N 1/4 1116 students) collected in the introductory calculus-based mechanics class at a large eastern land-grant university. Several academic and noncognitive factors were examined using correlation analysis and linear regression analysis to understand their relation to students??? physics conceptual understanding. High school physics preparation was the most important factor in predicting FMCE pretest score. The kind of high school physics class (normal or Advanced Placement) and the student???s academic performance in that class also greatly affected pretest scores. The optimal linear regression model explained 28% of the variance of pretest scores. Controlling for pretest score, ACT or SAT verbal and mathematics scores, students??? grade expectation, and self-efficacy significantly predicted post-test score. The optimal linear regression model explained 54% of the variance of post-test scores. Pretest scores completely captured the effect of high school preparation on post-test scores; if pretest scores were included in a model predicting post-test scores, then high school physics preparation variables were not significant. Gender differences were observed on both the pretest and the post-test. These differences were not substantially mediated by either academic or noncognitive factors.
Visualizing and Predicting the Path to an Undergraduate Physics Degree at Two Different Institutions
This study examined physics major retention to degree at two institutions with substantially different admissions selectivity. Two modes of leaving the physics major were examined: leaving college and changing to another major while staying in college. The risk of leaving college while still enrolled as a physics major was highest in the spring freshman semester. The changing major risk was substantially different between the two institutions. For the less selective institution, the students changed major at the highest rate in the fall sophomore semester. For the more selective institution, the risk of changing major was high through the first two years of college with highest risk in the fall freshman semester and the fall junior semester. Different features were important in predicting the two modes of leaving; these also differed between institutions. For the less selective institution, math readiness (being academically prepared to enroll in Calculus 1 in the fall freshman semester) was the most predictive feature for leaving the physics major while staying in college; high school GPA was the most important feature for predicting both leaving college and graduating with a physics degree. For the more selective institution, ACT composite scores were the only significant predictor of retention. The role of math readiness was dramatic at the less selective institution with 41% of students not math ready upon enrolling in college as physics majors; 59% of these students failed to enroll in the first required physics class.
The Conceptual Survey of Electricity and Magnetism (CSEM) is a widely used multiple-choice instrument measuring a student's conceptual understanding of electricity and magnetism. This study applied modified module analysis (MMA) and modified module analysis-partial (MMA-P), network analytic methods that identify groups of correlated responses, to CSEM data from two institutions (N = 2538 and 3595). In module analysis, groups of correlated responses are called communities. As in previous applications of MMA and MMA-P to mechanics conceptual inventories, a number of communities related to physics concepts and some communities related to the structure of blocked items in the inventory were identified. An item block is a set of items all referring to each other or to a common stem. Many blocked communities involved responses where the response to the later item would be correct if the response to the earlier item was correct. This suggests a modified scoring rubric for the CSEM is needed to account for these connections between items. A modified scoring rubric is proposed; however, the modified overall average scores changed by less than 1%. The communities of incorrect responses to the CSEM related to physical concepts had varied explanations. These explanations ranged from seemingly straightforward errors (the electric field pointing to higher potential or reversing the right-hand rule), to misconceptions about Newton's 2nd and 3rd laws carried over from mechanics, to naive reasoning conflating general topics in electricity and magnetism. The identification of incorrect communities allowed the computation of misconception scores showing how prevalent the misconceptions were in the classes studied.


