Publications
Argumentation is fundamental to science education, both as a prominent feature of scientific reasoning and as an effective mode of learning-a perspective reflected in contemporary frameworks and standards. The successful implementation of argumentation in school science, however, requires a paradigm shift in science assessment from the measurement of knowledge and understanding to the measurement of performance and knowledge in use. Performance tasks requiring argumentation must capture the many ways students can construct and evaluate arguments in science, yet such tasks are both expensive and resource-intensive to score. In this study we explore how machine learning text classification techniques can be applied to develop efficient, valid, and accurate constructed-response measures of students' competency with written scientific argumentation that are aligned with a validated argumentation learning progression. Data come from 933 middle school students in the San Francisco Bay Area and are based on three sets of argumentation items in three different science contexts. The findings demonstrate that we have been able to develop computer scoring models that can achieve substantial to almost perfect agreement between human-assigned and computer-predicted scores. Model performance was slightly weaker for harder items targeting higher levels of the learning progression, largely due to the linguistic complexity of these responses and the sparsity of higher-level responses in the training data set. Comparing the efficacy of different scoring approaches revealed that breaking down students' arguments into multiple components (e.g., the presence of an accurate claim or providing sufficient evidence), developing computer models for each component, and combining scores from these analytic components into a holistic score produced better results than holistic scoring approaches. However, this analytical approach was found to be differentially biased when scoring responses from English learners (EL) students as compared to responses from non-EL students on some items. Differences in the severity between human and computer scores for EL between these approaches are explored, and potential sources of bias in automated scoring are discussed.
Argumentation, a key scientific practice presented in the Framework for K-12 Science Education, requires students to construct and critique arguments, but timely evaluation of arguments in large-scale classrooms is challenging. Recent work has shown the potential of automated scoring systems for open response assessments, leveraging machine learning (ML) and artificial intelligence (AI) to aid the scoring of written arguments in complex assessments. Moreover, research has amplified that the features (i.e., complexity, diversity, and structure) of assessment construct are critical to ML scoring accuracy, yet how the assessment construct may be associated with machine scoring accuracy remains unknown. This study investigated how the features associated with the assessment construct of a scientific argumentation assessment item affected machine scoring performance. Specifically, we conceptualized the construct in three dimensions: complexity, diversity, and structure. We employed human experts to code characteristics of the assessment tasks and score middle school student responses to 17 argumentation tasks aligned to three levels of a validated learning progression of scientific argumentation. We randomly selected 361 responses to use as training sets to build machine-learning scoring models for each item. The scoring models yielded a range of agreements with human consensus scores, measured by Cohen’s kappa (mean = 0.60; range 0.38 − 0.89), indicating good to almost perfect performance. We found that higher levels of Complexity and Diversity of the assessment task were associated with decreased model performance, similarly the relationship between levels of Structure and model performance showed a somewhat negative linear trend. These findings highlight the importance of considering these construct characteristics when developing ML models for scoring assessments, particularly for higher complexity items and multidimensional assessments. © The Author(s) 2023.
In this study, we developed machine learning algorithms to automatically score students' written arguments and then applied the cognitive diagnostic modeling (CDM) approach to examine students' cognitive patterns of scientific argumentation. We abstracted three types of skills (i.e., attributes) critical for successful argumentation practice: making claims, using evidence, and providing warrants. We developed 19 constructed response items, with each item requiring multiple cognitive skills. We collected responses from 932 students in Grades 5 to 8 and developed machine learning algorithmic models to automatically score their responses. We then applied CDM to analyze their cognitive patterns. Results indicate that machine scoring achieved the average machine-human agreements of Cohen's kappa = 0.73, SD= 0.09. We found that students were clustered in 21 groups based on their argumentation performance, each revealing a different cognitive pattern. Within each group, students showed different abilities regarding making claims, using evidence, and providing warrants to justify how the evidence supports a claim. The 9 most frequent groups accounted for more than 70% of the students in the study. Our in-depth analysis of individual students suggests that students with the same total ability score might vary in the specific cognitive skills required to accomplish argumentation. This result illustrates the advantage of CDM in assessing the fine-grained cognition of students during argumentation practices in science and other scientific practices.
IntroductionThe Framework for K-12 Science Education promotes supporting the development of knowledge application skills along previously validated learning progressions (LPs). Effective assessment of knowledge application requires LP-aligned constructed-response (CR) assessments. But these assessments are time-consuming and expensive to score and provide feedback for. As part of artificial intelligence, machine learning (ML) presents an invaluable tool for conducting validation studies and providing immediate feedback. To fully evaluate the validity of machine-based scores, it is important to investigate human-machine score consistency beyond observed scores. Importantly, no formal studies have explored the nature of disagreements between human and machine-assigned scores as related to LP levels. MethodsWe used quantitative and qualitative approaches to investigate the nature of disagreements among human and scores generated by two approaches to machine learning using a previously validated assessment instrument aligned to LP for scientific argumentation. ResultsWe applied quantitative approaches, including agreement measures, confirmatory factor analysis, and generalizability studies, to identify items that represent threats to validity for different machine scoring approaches. This analysis allowed us to determine specific elements of argumentation practice at each level of the LP that are associated with a higher percentage of misscores by each of the scoring approaches. We further used qualitative analysis of the items identified by quantitative methods to examine the consistency between the misscores, the scoring rubrics, and student responses. We found that rubrics that require interpretation by human coders and items which target more sophisticated argumentation practice present the greatest threats to the validity of machine scores. DiscussionWe use this information to construct a fine-grained validity argument for machine scores, which is an important piece because it provides insights for improving the design of LP-aligned assessments and artificial intelligence-enabled scoring of those assessments.
IntroductionThe Framework for K-12 Science Education (the Framework) and the Next- Generation Science Standards (NGSS) define three dimensions of science: disciplinary core ideas, scientific and engineering practices, and crosscutting concepts and emphasize the integration of the three dimensions (3D) to reflect deep science understanding. The Framework also emphasizes the importance of using learning progressions (LPs) as roadmaps to guide assessment development. These assessments capable of measuring the integration of NGSS dimensions should probe the ability to explain phenomena and solve problems. This calls for the development of constructed response (CR) or open-ended assessments despite being expensive to score. Artificial intelligence (AI) technology such as machine learning (ML)-based approaches have been utilized to score and provide feedback on open-ended NGSS assessments aligned to LPs. ML approaches can use classifications resulting from holistic and analytic coding schemes for scoring short CR assessments. Analytic rubrics have been shown to be easier to evaluate for the validity of ML-based scores with respect to LP levels. However, a possible drawback of using analytic rubrics for NGSS-aligned CR assessments is the potential for oversimplification of integrated ideas. Here we describe how to deconstruct a 3D holistic rubric for CR assessments probing the levels of an NGSS-aligned LP for high school physical sciences. MethodsWe deconstruct this rubric into seven analytic categories to preserve the 3D nature of the rubric and its result scores and provide subsequent combinations of categories to LP levels. ResultsThe resulting analytic rubric had excellent human- human inter-rater reliability across seven categories (Cohen's kappa range 0.82-0.97). We found overall scores of responses using the combination of analytic rubric very closely agreed with scores assigned using a holistic rubric (99% agreement), suggesting the 3D natures of the rubric and scores were maintained. We found differing levels of agreement between ML models using analytic rubric scores and human-assigned scores. ML models for categories with a low number of positive cases displayed the lowest level of agreement. DiscussionWe discuss these differences in bin performance and discuss the implications and further applications for this rubric deconstruction approach.
Assessment developers are increasingly using the developing technology of machine learning in transforming how to assess students in their science learning. I argue that these algorithmic models further embed the structures of inequality that are pervasive in the development of science assessments in how they legitimize certain language practices that protect the hierarchical standing of status quo interests. My argument is situated within the broader emerging ethical challenges around this new technology. I apply a raciolinguistic equity analysis framework in critiquing the new black box that reinforces structural forms of discrimination against the linguistic repertoires of racially marginalized student populations. The article ends with me sharing a set of tactical shifts that can be deployed to form a more equitable and socially-just field of machine learning enhanced science assessments.
Machine learning (ML) has been increasingly employed in science assessment to facilitate automatic scoring efforts, although with varying degrees of success (i.e., magnitudes of machine-human score agreements [MHAs]). Little work has empirically examined the factors that impact MHA disparities in this growing field, thus constraining the improvement of machine scoring capacity and its wide applications in science education. We performed a meta-analysis of 110 studies of MHAs in order to identify the factors most strongly contributing to scoring success (i.e., high Cohen's kappa [kappa]). We empirically examined six factors proposed as contributors to MHA magnitudes: algorithm, subject domain, assessment format, construct, school level, and machine supervision type. Our analyses of 110 MHAs revealed substantial heterogeneity in kappa(mean=.64; range = .09-.97, taking weights into consideration). Using three-level random-effects modeling, MHA score heterogeneity was explained by the variability both within publications (i.e., the assessment task level: 82.6%) and between publications (i.e., the individual study level: 16.7%). Our results also suggest that all six factors have significant moderator effects on scoring success magnitudes. Among these, algorithm and subject domain had significantly larger effects than the other factors, suggesting that technical features and assessment external features might be primary targets for improving MHAs and ML-based science assessments.
This study provides a solid validity inferential network to guide the development, interpretation, and use of machine learning-based next-generation science assessments (NGSAs). Given that machine learning (ML) has been broadly implemented in the automatic scoring of constructed responses, essays, simulations, educational games, and interdisciplinary assessments to advance the evidence collection and inference of student science learning, we contend that additional validity issues arise for science assessments due to the involvement of ML. These emerging validity issues may not be addressed by prior validity frameworks developed for either non-science or non-ML assessments. We thus examine the changes brought in by ML to science assessments and identify seven critical validity issues of ML-based NGSAs: potential risk of misrepresenting the construct of interest, potential confounders due to that more variables may involve, nonalignment between interpretation and use of scores and designed learning goals, nonalignment between interpretation and use of scores and actual learning quality, nonalignment between machine scores and rubrics, limited generalizable ability of machine algorithmic models, and limited extrapolating ability of machine algorithmic models. Based on the seven validity issues identified, we propose a validity inferential network to address the cognitive, instructional, and inferential validity of ML-based NGSAs. To demonstrate the utility of this network, we present an exemplar of ML-based next-generation science assessments that was developed using a seven-step ML framework. We articulate how we used the validity inferential network to ensure accountable assessment design, as well as valid interpretation and use of machine scores.
As cutting-edge technologies, such as machine learning (ML), are increasingly involved in science assessments, it is essential to conceptualize how assessment practices are innovated by technologies. To partially meet this need, this article focuses on ML-based science assessments and elaborates on how ML innovates assessment practices in science education. The article starts with an articulation of the practice nature of assessment both of learning and for learning, identifying four essential assessment practices: identifying learning goals, eliciting performance, interpreting observations, and decision-making and action-taking. I then extend a three-dimensional framework for innovative assessments, including construct, functionality, and automaticity, and based on which to conceptualize innovative assessments in three levels: substitute, transform, and redefine. Using the framework, I elaborate on how the 10 articles included in this special issue, Applying Machine Learning in Science Assessment: Opportunity and Challenge, advanced our knowledge of the innovations that ML brought to science assessment practices. I contend that the 10 articles exemplify a great deal of effort to transform the four components of assessment practices: ML allows assessments to target complex, diverse, and structural constructs, and thus better approaching the three-dimensional science learning goals of the Next Generation Science Standards (NGSS Lead States, 2013); ML extends the approaches used to eliciting performance and collecting evidence; ML provides a means to better interpreting observations and using evidence; ML supports immediate and complex decision-making and action-taking. I conclude this article by pushing the field to consider the underlying educational theories that are needed for innovative assessment practices and the necessities of establishing a romance between assessment practices and the relevant educational theories, which I contend are the prominent challenges to forward innovative and ML-based assessment practices in science education.
This article is in response to the review entitled Identifying potential types of guidance for supporting student inquiry when using virtual and remote labs in science: a literature review (Zacharia et al. in Educ Technol Res Dev 63(2): 257-302). As COVID-19 alerts us to shift science education to digital when in-person schooling is not viable, one approach to facilitate this shift, as reviewed by Zacharia et al. (2015), is to involve students in computer supported inquiry learning (CoSIL) with appropriate guidance. As CoSIL guidance is critical to student success in CoSIL, Zacharia et al. (2015) contribute to our knowledge by systematically reviewing the forms and the efficacy of such guidance tools that are associated with each phase of scientific inquiry. With such knowledge we may develop decent guidance so that students can experience scientific inquiry virtually as they used to do in-person. Zacharia et al. (2015) indicated that the various guidance tools had increased the ease of use of CoSIL but failed in personalizing CoSIL to individual students. I agree with Zacharia et al. that the personalization of CoSIL guidance is vital. Further, I argue that the emergent machine learning may significantly increase the personalization of CoSIL without burdening teachers. I conclude the essay with suggestions to further investigate the cognitive needs of students in CoSIL and integrate the content, CoSIL, and guidance tools, as a way to move forward the personalization of CoSIL.


