Publications
Assessment is central to teaching and learning, and recently there has been a substantive shift from paper-and-pencil assessments towards technology delivered assessments such as computer-adaptive tests. Fairness is an important aspect of the assessment process, including design, administration, test-score interpretation, and data utility. The Universal Design for Learning (UDL) guidelines can inform assessment development to promote fairness; however, it is not explicitly clear how UDL and fairness may be linked through students’ conceptualizations of assessment fairness. This phenomenological study explores how middle grades students conceptualize and reason about the fairness of mathematics tests, including paper-and-pencil and technology-delivered assessments. Findings indicate that (a) students conceptualize fairness through unique notions related to educational opportunities and (b) students’ reason about fairness non-linearly. Implications of this study have potential to inform test developers and users about aspects of test fairness, as well as educators data usage from fixed-form, paper-and-pencil tests, and computer-adaptive, technology-delivered tests. © The Author(s) 2024.
Quantitative measures in mathematics education have informed policies and practices for over a century. Thus, it is critical that such measures in mathematics education have sufficient validity evidence to improve mathematics experiences for students. This article provides a systematic review of the validity evidence related to measures used in elementary mathematics education. The review includes measures that focus on elementary students as the unit of analyses and attends to validity as defined by current conceptions of measurement. Findings suggest that one in ten measures in mathematics education include rigorous evidence to support intended uses. Recommendations are made to support mathematics education researchers to continue to take steps to improve validity evidence in the design and use of quantitative measures.
Validity is a fundamental consideration of test development and test evaluation. The purpose of this study is to define and reify three key aspects of validity and validation, namely test-score interpretation, test-score use, and the claims supporting interpretation and use. This study employed a Delphi methodology to explore how experts in validity and validation conceptualize test-score interpretation, use, and claims. Definitions were developed through multiple iterations of data collection and analysis. By clarifying the language used when conducting validation, validation may be more accessible to a broader audience, including but not limited to test developers, test users, and test consumers. © 2023 The Authors. Educational Measurement: Issues and Practice published by Wiley Periodicals LLC on behalf of National Council on Measurement in Education.
[No abstract available]
Problem solving is a central focus of mathematics teaching and learning. If teachers are expected to support students' problem-solving development, then it reasons that teachers should also be able to solve problems aligned to grade level content standards. The purpose of this validation study is twofold: (1) to present evidence supporting the use of the Problem Solving Measures Grades 3-5 with preservice teachers (PSTs), and (2) to examine PSTs' abilities to solve problems aligned to grades 3-5 academic content standards. This study used Rasch measurement techniques to support psychometric analysis of the Problem Solving Measures when used with PSTs. Results indicate the Problem Solving Measures are appropriate for use with PSTs, and PSTs' performance on the Problem Solving Measures differed between first-year PSTs and end-of-program PSTs. Implications include program evaluation and the potential benefits of using K-12 student-level assessments as measures of PSTs' content knowledge.
The COVID-19 pandemic disrupted many school accountability systems that rely on student- level achievement data. Many states have encountered uncertainty about how to meet federal accountability requirements without typical school data. Prior research provides an abundance of evidence that student achievement is correlated to students' social background, which raises concerns about the predictive bias of accountability systems. The focus of this quantitative study is to explore the predictive ability of non-achievement based variables (i.e., students' social background) on measures of school accountability in one Midwest state. Results suggest that social background and community demographic variables have a significant impact on measures of school accountability, and might be interpreted cautiously. Implications for policy and future research are discussed.
This Research Commentary addresses the need for an instrument abstract???termed an Interpretation and Use Statement (IUS)???to be included when mathematics educators present instruments for use by others in journal articles and other communication venues (e.g., websites and administration manuals). We begin with presenting the need for IUSs, including the importance of a focus on interpretation and use. We then propose a set of elements???identified by a group of mathematics education researchers, instrument developers, and psychometricians???to be included in the IUS. We describe the development process, the recommended elements for inclusion, and two example IUSs. Last, we present why IUSs have the potential to benefit end users and the field of mathematics education.
Response process validity evidence provides a window into a respondent's cognitive processing. The purpose of this study is to describe a new data collection tool called a whole-class think aloud (WCTA). This work is performed as part of test development for a series of problem-solving measures to be used in elementary and middle grades. Data from third-grade students were collected in a 1-1 think-aloud setting and compared to data from similar students as part of WCTAs. Findings indicated that students performed similarly on the items when the two think-aloud settings were compared. Respondents also needed less encouragement to share ideas aloud during the WCTA compared to the 1-1 think aloud. They also communicated feeling more comfortable in the WCTA setting compared to the 1-1 think aloud. Drawing the findings together, WCTAs functioned as well if not better, than 1-1 think alouds for the purpose of contextualizing third-grade students' cognitive processes. Future studies using WCTAs are recommended to explore their limitations and other factors that might impact their success as data gathering tools.
Think alouds are valuable tools for academicians, test developers, and practitioners as they provide a unique window into a respondent's thinking during an assessment. The purpose of this special issue is to highlight novel ways to use think alouds as a means to gather evidence about respondents' thinking. An intended outcome from this special issue is that readers may better understand think alouds and feel better equipped to use them in practical and research settings.


