Publications
The Conceptual Survey of Electricity and Magnetism (CSEM) is a widely used multiple-choice instrument measuring a student's conceptual understanding of electricity and magnetism. This study applied modified module analysis (MMA) and modified module analysis-partial (MMA-P), network analytic methods that identify groups of correlated responses, to CSEM data from two institutions (N = 2538 and 3595). In module analysis, groups of correlated responses are called communities. As in previous applications of MMA and MMA-P to mechanics conceptual inventories, a number of communities related to physics concepts and some communities related to the structure of blocked items in the inventory were identified. An item block is a set of items all referring to each other or to a common stem. Many blocked communities involved responses where the response to the later item would be correct if the response to the earlier item was correct. This suggests a modified scoring rubric for the CSEM is needed to account for these connections between items. A modified scoring rubric is proposed; however, the modified overall average scores changed by less than 1%. The communities of incorrect responses to the CSEM related to physical concepts had varied explanations. These explanations ranged from seemingly straightforward errors (the electric field pointing to higher potential or reversing the right-hand rule), to misconceptions about Newton's 2nd and 3rd laws carried over from mechanics, to naive reasoning conflating general topics in electricity and magnetism. The identification of incorrect communities allowed the computation of misconception scores showing how prevalent the misconceptions were in the classes studied.
According to Bandura's social cognitive theory, a student's self-efficacy influences his or her academic and career decisions, and his or her performance outcomes; as such, a student's self-efficacy changes with time in response to the student's experiences. Self-efficacy may also vary by academic domain. Differences in STEM self-efficacy have often been reported between men and women. The purpose of this study is to explore the evolution of domain-specific STEM self-efficacy in students in gateway physics and mathematics courses and how academic feedback influences the evolution of these differences with time. Further, this study explored whether gender differences in self-efficacy are consistent across STEM domains and how these differences change in response to academic feedback. Self-efficacy in multiple academic domains (current mathematics/science class, other STEM classes, and intended profession) was assessed at multiple time points with subscales adapted from the Motivated Strategies for Learning Questionnaire. Linear mixed effects modeling was used to understand how academic feedback provided by test scores influenced changes in self-efficacy. Students in all classes expressed different levels of self-efficacy toward different domains with the lowest self-efficacy toward their current class and the highest toward their intended profession. Only the current math/science class self-efficacy of men and women differed significantly, with women expressing lower self-efficacy. The differences in current class self-efficacy were evident very early in the class before substantive class feedback was received. The evolution of self-efficacy within the class and between classes was the same for men and women.
Machine learning algorithms have recently been used to predict students' performance in an introductory physics class. The prediction model classified students as those likely to receive an A or B or students likely to receive a grade of C, D, F or withdraw from the class. Early prediction could better allow the direction of educational interventions and the allocation of educational resources. However, the performance metrics used in that study become unreliable when used to classify whether a student would receive an A, B, or C (the ABC outcome) or if they would receive a D, F or withdraw (W) from the class (the DFW outcome) because the outcome is substantially unbalanced with between 10% to 20% of the students receiving a D, F, or W. This work presents techniques to adjust the prediction models and alternate model performance metrics more appropriate for unbalanced outcome variables. These techniques were applied to three samples drawn from introductory mechanics classes at two institutions (N = 7184, 1683, and 926). Applying the same methods as the earlier study produced a classifier that was very inaccurate, classifying only 16% of the DFW cases correctly; tuning the model increased the DFW classification accuracy to 43%. Using a combination of institutional and in-class data improved DFW accuracy to 53% by the second week of class. As in the prior study, demographic variables such as gender, underrepresented minority status, fast-generation college student status, and low socioeconomic status were not important variables in the final prediction models.
Investigating student learning and understanding of conceptual physics is a primary research area within physics education research. Multiple quantitative methods have been employed to analyze commonly used mechanics conceptual inventories: the Force Concept Inventory (FCI) and the Force and Motion Conceptual Evaluation (FMCE). Recently, researchers have applied network analytic techniques to explore the structure of the incorrect responses to the FCI identifying communities of incorrect responses which could be mapped on to common misconceptions. In this study, the method used to analyze the FCI, modified module analysis was applied to a large sample of FMCE pretest and post-test responses (N-pre = 3956, N-post = 3719). The communities of incorrect responses identified were consistent with the item groups described in previous works. As in the work with the FCI, the network was simplified by only retaining nodes selected by a substantial number of students. Retaining as nodes only those incorrect answer choices selected by at least 20% of the students produced communities associated with only four misconceptions. The incorrect response communities identified for men and women were substantially different, as was the change in these communities from pretest to post-test. The 20% threshold was far more restrictive than the 4% threshold applied to the FCI in the prior work that generated similar structures. Retaining nodes selected by 5% or 10% of students generated a large number of complex communities. The communities identified at the 10% threshold were generally associated with common misconceptions producing a far richer set of incorrect communities than the FCI; this may indicate that the FMCE is a superior instrument for characterizing the breadth of student misconceptions about Newtonian mechanics.
This study examines the correlation of physics conceptual inventory pretest scores with post-instruction achievement measures (post-test scores, test averages, and course grades). The correlation for demographic groups in the minority in the physics classes studied (women, underrepresented racial/enthic students, first generation college students, and rural students) were compared with their majority peers. Three conceptual inventories were examined: the Force and Motion Conceptual Evaluation (FMCE) (N = 2450), the Force Concept Inventory (FCI) (N = 2373) and the CSEM (N1 = 1796, N2 = 2537). While many of the correlations were similar, for some of the demographic groups, the correlations were substantially different. There was little consistency in the differences measured. In most cases where the correlations differed, the correlation for the group in the minority was the smaller. As such, pretest scores may not predict course performance for some minority demographic groups as accurately as they predict outcomes for majority students. The pattern of correlation differences did not appear to be related to the size of the pretest score. If pretest scores are used for instructional decisions that have academic consequences, instructors should be aware of these potential inaccuracies and ensure the pretest used is equally valid for all students.
This work applied exploratory factor analysis (EFA) and graded Item Response Theory (IRT) to a large sample ( N = 4522) of Colorado Learning Attitudes about Science Survey (CLASS) post-test scores. EFA failed to reproduce the factor structure suggested by the authors of the CLASS and strongly supported the alternate 3-factor model suggested by Douglas et al. Graded IRT allowed an examination of the progression from non-expert-like to expert-like beliefs. This progression was generally uniform with a linear relation between the difficulty of each step in the progression. Some items within the factors identified by Douglas et al. had difficulty and discrimination parameters substantially different from other items in the factor suggesting the subscale is not unidimensional. The expert-like latent ability trait estimated by IRT correlated more strongly with measures of physics performance than measures of general academic performance indicating that expert-like beliefs are not a general property of high performing students.
The Force and Motion Conceptual Evaluation is commonly used to measure the conceptual understanding of Newtonian mechanics. Several studies have reported a substantial difference in pretest scores between men and women. This study examines the contribution of several prior preparation factors to explain the variance in pretest score and whether these factors explain gender differences in the pretest score. The study examined a large sample (N = 1060) of students taking introductory calculus-based mechanics at the university level. Women outperformed men on most prior preparation and college achievement measures. No significant differences between men and women were found in high school physics taking patterns. Linear regression analysis showed only 23% of the variance in FMCE pretest score could be explained using a linear combination of prior preparation variables. Controlling for these variables failed to explain the gender difference in pretest scores; conversely, the gender difference increased controlling for prior preparation.
Many studies have examined the structure and properties of the Force Concept Inventory (FCI); however, far less research has investigated the Force and Motion Conceptual Evaluation (FMCE). This study applied Multidimensional Item Response Theory (MIRT) to a sample of N=4528 FMCE post-test responses. Exploratory factor analysis showed that 5, 9, and 10-factor models optimized some fit statistics. The FMCE uses extensive blocking of items into groups with a common stem; these blocks factored together in most models. A confirmatory analysis, which constrained the MIRT models to a theoretical model constructed from expert solutions, produced a model requiring only 8 principles, fundamental reasoning steps. This was substantially fewer than the 19 principles identified in the FCI by a previous study. Correlation analysis also demonstrated that the two instruments were very dissimilar. The reduced number of principles and the repetition of items using a single principle allowed the extraction of eight single-principle subscales, seven with Cronbach's alpha greater than the 0.7 required for acceptable internal consistency. The differences between the FCI and the FMCE suggest that the two instruments could provide complementary, but different, information about student understanding of Newton's laws with the FCI measuring an integrated Newtonian force concept and the FMCE measuring details of that force concept. © 2019 authors. Published by the American Physical Society. Published by the American Physical Society under the terms of the "https://creativecommons.org/licenses/by/4.0/" Creative Commons Attribution 4.0 International license. Further distribution of this work must maintain attribution to the author(s) and the published article's title, journal citation, and DOI.
The use of machine learning and data mining techniques across many disciplines has exploded in recent years with the field of educational data mining growing significantly in the past 15 years. In this study, random forest and logistic regression models were used to construct early warning models of student success in introductory calculus-based mechanics (Physics 1) and electricity and magnetism (Physics 2) courses at a large eastern land-grant university. By combining in-class variables such as homework grades with institutional variables such as cumulative GPA, we can predict if a student will receive less than a B in the course with 73% accuracy in Physics 1 and 81% accuracy in Physics 2 with only data available in the first week of class using logistic regression models. The institutional variables were critical for high accuracy in the first four weeks of the semester. In-class variables became more important only after the first in-semester examination was administered. The student's cumulative college GPA was consistently the most important institutional variable. Homework grade became the most important in-class variable after the first week and consistently increased in importance as the semester progressed; homework grade became more important than cumulative GPA after the first in-semester examination. Demographic variables including gender, race or ethnicity, and first generation status were not important variables for predicting course grade.
While many studies have examined the structure, validity, and reliability of the Force Concept Inventory, far less research has been performed on other conceptual instruments in widespread use in physics education research. This study performs a confirmatory analysis of the Conceptual Survey of Electricity and Magnetism (CSEM) guided by a theoretical model of expert understanding of electricity and magnetism. Multidimensional Item Response Theory (MIRT) with the discrimination matrix constrained to the theoretical model was used to investigate two large datasets (N-1 = 2014 and N-2 = 2657) from two research universities in the United States. The optimal model identified by MIRT was similar, but not identical, for the two datasets and had very good model fit with comparative fit indices of 0.975 and 0.984, respectively. The most parsimonious optimal model required 23 independent principles of electricity and magnetism and was significantly better fitting than a more general model dividing the CSEM into 6 general topics. The optimal models for the two samples were quite similar, sharing 22 of a possible 26 conceptual principles. Most of the overall item difficulties and discriminations were significantly different between the two samples; however, the rank order of the overall difficulty and discrimination were generally similar. There was much more similarity between the discrimination by item of the individual principles. Five items had a difficulty ranking that was substantially different between the two samples, indicating that while generally similar, relative difficulty does depend on the student population and instructional environment.


