Publications
The goals of the current study were: 1) to modify and expand an existing spatial mathematical language coding system to include quantitative mathematical language terms and 2) to examine the extent to which preschool-aged children used spatial and quantitative mathematical language during a block play intervention. Participants included 24 preschool-aged children (Age M = 57.35 months) who were assigned to a block play intervention. Children participated in up to 14 sessions of 15-to-20-minute block play across seven weeks. Results demonstrated that spatial mathematical language terms were used with a higher raw frequency than quantitative mathematical language terms during the intervention sessions. However, once weighted frequencies were calculated to account for the number of codes in each category, spatial language was only used slightly more than quantitative language during block play. Similar patterns emerged between domains within the spatial and quantitative language categories. These findings suggest that both quantitative and spatial mathematical language usage should be evaluated when considering whether child activities can improve mathematical learning and spatial performance. Further, accounting for the number of codes within categories provided a more representative presentation of how mathematical language was used versus solely utilizing raw word counts. Implications for future research are discussed.
Teachers can make number talks more ambitious and increase learning opportunities for students by debugging errors, promoting multi-student thinking, and orienting to make connections. This article describes the basic number talks routine and examines some common tendencies, such as "Serial Sharing." It concludes with a vignette to illustrate these moves, along with ideas for how and why teachers might use them.
Generalizing is a critical aspect of mathematics learning, with researchers and policy documents highlighting generalizing as a core mathematical practice. It can also be challenging to foster in class settings, and teachers need access to better resources to teach generalizing, including an understanding of effective forms of instruction. This article proposes Classroom Supports for Generalizing (CSGs), investigating how multiple elements—such as tasks, teacher moves, student interactions, and representations—interact to meaningfully foster student generalizing. Drawing on class video data from a middle school teacher and two high school teachers, we present the CSG Framework, which identifies three categories of supports: Interactions for Generalizing, Structures for Generalizing, and Routines for Generalizing. © 2023 by The National Council of Teachers of Mathematics, Inc.
This Brief Report presents an example of assessment validation using an argument -based approach. The instrument we developed is a Brief Assessment of Students' Mature Number Sense, which measures a central goal in mathematics education. We chose to develop this assessment to provide an efficient way to measure the effect of instructional practices designed to improve students' number sense. Using an argument -based framework, we first identify our proposed interpretations and uses of student scores. We then outline our argument with three claims that provide evidence connecting students' responses on the assessment with its intended uses. Finally, we highlight why using argument -based validation benefits measure developers as well as the broader mathematics education community.
This article offers the construct unitizing predicates to name mental actions important for students’ reasoning about logic. To unitize a predicate is to conceptualize (possibly complex or multipart) conditions as a single property that every example has or does not have, thereby partitioning a universal set into examples and nonexamples. This explains the cognitive work that supports students to unify various statements with the same logical form, which is conventionally represented by replacing parts of statements with logical variables p or P(x). Using data from a constructivist teaching experiment with two undergraduate students, we document barriers to unitizing predicates and demonstrate how this activity influences students’ ability to render mathematical statements and proofs as having the same logical structure. © 2020 by The National Council of Teachers of Mathematics, Inc. www.nctm.org. All rights reserved.
The effectiveness of feedback in enhancing learning outcomes is well documented within Educational Data Mining (EDM). Various prior research have explored methodologies to enhance the effectiveness of feedback to students in various ways. Recent developments in Large Language Models (LLMs) have extended their utility in enhancing automated feedback systems. This study aims to explore the potential of LLMs in facilitating automated feedback in math education in the form of numeric assessment scores. We examine the effectiveness of LLMs in evaluating student responses and scoring the responses by comparing 3 different models: Llama, SBERT-Canberra, and GPT4 model. The evaluation requires the model to provide a quantitative score on the student's responses to open-ended math problems. We employ Mistral, a version of Llama catered to math, and fine-tune this model for evaluating student responses by leveraging a dataset of student responses and teacher-provided scores for middle-school math problems. A similar approach was taken for training the SBERT-Canberra model, while the GPT4 model used a zero-shot learning approach. We evaluate and compare the models' performance in scoring accuracy. This study aims to further the ongoing development of automated assessment and feedback systems and outline potential future directions for leveraging generative LLMs in building automated feedback systems. [This paper was published in: "Proceedings of the 17th International Conference on Educational Data Mining," edited by B. Paaßen and C. D. Epp, International Educational Data Mining Society, 2024, pp. 732-737. Funding for this paper was provided by the U.S. Department of Education's Graduate Assistance in Areas of National Need (GAANN).]
Gaming the system is a persistent problem in Computer-Based Learning Platforms. While substantial progress has been made in identifying and understanding such behaviors, effective interventions remain scarce. This study uses a method of causal moderation known as Fully Latent Principal Stratification to explore the impact of two types of interventions – gamification and manipulation of assistance access – on the learning outcomes of students who tend to game the system. The results indicate that gamification does not consistently mitigate these negative behaviors. One gamified condition had a consistently positive effect on learning regardless of students’ propensity to game the system, whereas the other had a negative effect on gamers. However, delaying access to hints and feedback may have a positive effect on the learning outcomes of those gaming the system. This paper also illustrates the potential for integrating detection and causal methodologies within educational data mining to evaluate effective responses to detected behaviors. © 2024 International Educational Data Mining Society. All rights reserved.
Common discourse conveys that to be an engineer, one must be smart. Our individual and collective beliefs about what constitutes smart behavior are shaped by our participation in the complex cultural practice of smartness. From the literature, we know that the criteria for being considered smart in our educational systems are biased. The emphasis on selecting and retaining only those who are deemed smart enough to be engineers perpetuates inequity in undergraduate engineering education. Less is known about what undergraduate students explicitly believe are the different ways of being smart in engineering or how those different ways of being a smart engineer are valued in introductory engineering classrooms. In this study, we explored the common beliefs of undergraduate engineering students regarding what it means to be smart in engineering. We also explored how the students personally valued those ways of being smart versus what they perceived as being valued in introductory engineering classrooms. Through our multi-phase, multi-method approach, we initially qualitatively characterized their beliefs into 11 different ways to be smart in engineering, based on a sample of 36 engineering students enrolled in first-year engineering courses. We then employed quantitative methods to uncover significant differences, with a 95% confidence interval, in six of the 11 ways of being smart between the values personally held by engineering students and what they perceived to be valued in their classrooms. Additionally, we qualitatively found that 1) students described grades as central to their classroom experience, 2) students described the classroom as a context where effortless achievement is associated with being smart, and 3) students described a lack of reward in the classroom for showing initiative and for considerations of social impact or helping others. As engineering educators strive to be more inclusive, it is essential to have a clear understanding and reflect on how students value different ways of being smart in engineering as well as consider how these values are embedded into teaching praxis.
The visual modality is central to both reception and expression of human creativity. Creativity assessment paradigms, such as structured drawing tasks Barbot (2018), seek to characterize this key modality of creative ideation. However, visual creativity assessment paradigms often rely on cohorts of expert or naive raters to gauge the level of creativity of the outputs. This comes at the cost of substantial human investment in both time and labor. To address these issues, recent work has leveraged the power of machine learning techniques to automatically extract creativity scores in the verbal domain (e.g., SemDis; Beaty & Johnson 53, 757-780, 2021). Yet, a comparably well-vetted solution for the assessment of visual creativity is missing. Here, we introduce AuDrA - an Automated Drawing Assessment platform to extract visual creativity scores from simple drawing productions. Using a collection of line drawings and human creativity ratings, we trained AuDrA and tested its generalizability to untrained drawing sets, raters, and tasks. Across four datasets, nearly 60 raters, and over 13,000 drawings, we found AuDrA scores to be highly correlated with human creativity ratings for new drawings on the same drawing task (r = .65 to .81; mean = .76). Importantly, correlations between AuDrA scores and human raters surpassed those between drawings' elaboration (i.e., ink on the page) and human creativity raters, suggesting that AuDrA is sensitive to features of drawings beyond simple degree of complexity. We discuss future directions, limitations, and link the trained AuDrA model and a tutorial (https://osf.io/kqn9v/) to enable researchers to efficiently assess new drawings.


