Publications
In this issue, Kindel et al. describe a new approach to managing survey data in service of the Fragile Families Challenge, which they call “treating metadata as data.” Although the approach they present is a good first step, a more ambitious proposal could improve survey data analysis even more substantially. The author recommends that data collection efforts distribute an open-source set of tools for working with a particular data set the author calls data-specific functions. The goal of these functions is to codify best practices for working with the data in a set of functions for commonly used statistical software. These functions would be jointly developed by the users and distributers of the data. Building such functions would both shorten the learning curve for new users and improve the quality of the data, by making tacit knowledge about problems with the data explicit and easy to act on. © SAGE Publications Inc.. All rights reserved.
Background Most science categories are hierarchically organized, with various high-level divisions comprising numerous subtypes. If we suppose that one's goal is to teach students to classify at the high level, past research has provided mixed evidence about whether an effective strategy is to require simultaneous classification learning of the subtypes. This past research was limited, however, either because authentic science categories were not tested, or because the procedures did not allow participants to form strong associations between subtype-level and high-level category names. Here we investigate a two-stage response-training procedure in which participants provide both a high-level and subtype-level response on most trials, with feedback provided at both levels. The procedure is tested in experiments in which participants learn to classify large sets of rocks that are representative of those taught in geoscience classes. Results The two-stage procedure yielded high-level classification performance that was as good as the performance of comparison groups who were trained solely at the high level. In addition, the two-stage group achieved far greater knowledge of the hierarchical structure of the categories than did the comparison controls. Conclusion In settings in which students are tasked with learning high-level names for rock types that are commonly taught in geoscience classes, it is best for students to learn simultaneously at the high and subtype levels (using training techniques similar to the presently investigated one). Beyond providing insights into the nature of category learning and representation, these findings have practical significance for improving science education.
Vision and Change challenged biology instructors to develop evidence-based instructional approaches that were grounded in the core concepts and competencies of biology. This call for reform provides an opportunity for new educational toots to be incorporated into biology education. In this essay, we advocate for learning progressions as one such educational tool. First, we address what learning progressions are and how they leverage research from the cognitive and learning sciences to inform instructional practices. Next, we use a published learning progression about carbon cycling to illustrate how learning progressions describe the maturation of student thinking about a key topic. Then, we discuss how learning progressions can inform undergraduate biology instruction, citing three particular learning progressions that could guide instruction about a number of key topics taught in introductory biology courses. Finally, we describe some challenges associated with learning progressions in undergraduate biology and some recommendations for how to address these challenges.
Faculty and peer interactions play a key role in shaping graduate student socialization. Yet, within the literature on graduate student socialization, researchers have primarily focused on understanding the nature and impact of faculty atone, and much less is known about how peer interactions also contribute to graduate student outcomes. Using a national sample of first-year biology doctoral students, this study reveals distinct categories that classify patterns of faculty and peer interaction. Further, we document inequities such that certain groups (e.g., underrepresented minority students) report constrained types of interactions with faculty and peers. Finally, we connect faculty and peer interaction patterns to student outcomes. Our findings reveal that, while the classification of faculty and peer interactions predicted affective and experiential outcomes (e.g., sense of belonging, satisfaction with academic development), it was not a consistent predictor of more central outcomes of the doctoral socialization process (e.g., research skills, commitment to degree). These and other findings are discussed, focusing on implications for future research, theory, and practice related to graduate training.
Economic downturns are known to spark periods of increased enrolment in traditional educational pursuits. The current study leverages 30-year longitudinal data from the Longitudinal Study of American Life (LSAL; N=1,556) to examine individual characteristics and experiences in adolescence, just prior to the Great Recession, and during it, to understand why some individuals chose to pursue new education or training in response to the recession whereas others did not. Indicators from adolescence include measures of self-esteem, locus of control, persistence, achievement in mathematics and achievement in science and were collected from 1987 to 1993. inclusive. Pre-recession indicators include level of education, occupational and marital status and were collected in 2007. Indicators of the impact of the recession were collected retrospectively in 2014 and include whether a job was lost, whether work hours were reduced, and whether there was difficulty making rent/mortgage payments. Binary logistic regression identified persistence in adolescence, pre-recession education level, reporting reduced hours and difficulty paying rent/mortgage during the recession as associated with the likelihood of pursuing new education during the recession. A follow-up analysis investigated whether the pursuit of additional education/training in response to the recession predicted the likelihood of being employed in 2017. Results indicate that obtaining new education during the recession was associated with later employment status, but the significance and direction of the effect depends on pre-recession education level. Implications of this longitudinal, life course analysis are discussed in addition to recommendation for future directions.
A growing number of methods aim to assess the challenging question of treatment effect variation in observational studies. This special section of Observational Studies reports the results of a workshop conducted at the 2018 Atlantic Causal Inference Conference designed to understand the similarities and differences across these methods. We invited eight groups of researchers to analyze a synthetic observational data set that was generated using a recent large-scale randomized trial in education. Overall, participants employed a diverse set of methods, ranging from matching and flexible outcome modeling to semiparametric estimation and ensemble approaches. While there was broad consensus on the topline estimate, there were also large differences in estimated treatment effect moderation. This highlights the fact that estimating varying treatment effects in observational studies is often more challenging than estimating the average treatment effect alone. We suggest several directions for future work arising from this workshop. © 2019 Carlos Carvalho, Avi Feller, Jared Murray, Spencer Woody, and David Yeager.
State accountability systems have been a primary school reform initiative in the US for the past 20 years, but often produce unintended negative consequences. In 2004, the Texas Education Agency (TEA) implemented the Performance Based Monitoring and Analysis System (PBMAS), which included an accountability indicator focused on the percentage of students found eligible for special education under the Individuals with Disabilities Education Act (IDEA), the nation’s special education law. From 2004 through 2016, the percentage of students found eligible for special education in Texas declined significantly, while the national rate held constant. Eventually, the U.S. Department of Education (ED) investigated TEA and the statewide implementation of IDEA. The purpose of this study is two-fold: (a) to evaluate the potential impact of the the PBMAS indicator on manipulation of special education identification practices; and (b) to describe how the indicator may have influenced school and district personnel. We highlight several concerning trends in state and district data and, through an analysis of publicly available reports from the ED, show how district and school personnel knowingly and unknowingly acted in ways that delayed and denied special education to potentially eligible students. We conclude with recommendations for TEA and implications for future research and policy. © 2019, Arizona State University. All rights reserved.
Randomized controlled trials (RCTs) admit unconfounded design-based inference--randomization largely justifies the assumptions underlying statistical effect estimates--but often have limited sample sizes. However, researchers may have access to big observational data on covariates and outcomes from RCT non-participants. For example, data from A/B tests conducted within an educational technology platform exist alongside historical observational data drawn from student logs. We outline a design-based approach to using such observational data for variance reduction in RCTs. First, we use the observational data to train a machine learning algorithm predicting potential outcomes using covariates, and use that algorithm to generate predictions for RCT participants. Then, we use those predictions, perhaps alongside other covariates, to adjust causal effect estimates with a flexible, design-based covariate-adjustment routine. In this way there is no danger of biases from the observational data leaking into the experimental estimates, which are guaranteed to be exactly unbiased regardless of whether the machine learning models are "correct" in any sense or whether the observational samples closely resemble RCT samples. We demonstrate the method in analyzing 33 randomized A/B tests, and show that it decreases standard errors relative to other estimators, sometimes substantially. [This is the online version of an article published in "Journal of Causal Inference." Additional funding was provided by the U.S. Department of Education's Graduate Assistance in Areas of National Need (GAANN) program.]
Efforts in the U.S. to design curriculum, instruction, and assessment based on Indigenous systems of knowledge and ways of teaching and assessing learning have been mounted wherever Indigenous peoples live. Yet, Western-style education in those places often continues to dominate, to the detriment of Indigenous students' engagement and school completion. Assessment, in particular, has long aroused great concern because many common assessments are not only ineffective but also destructive for Indigenous students-especially when they are used to make high-stakes decisions that affect students' life outcomes. Among such decisions are eligibility for passage from one grade to the next, high school graduation, and college admission. Much is known about how to make assessment culturally-responsive for Indigenous students, but it is often the case that successful programs and practices are jettisoned when new country-wide or state-wide policies are instituted. In the U.S., the most egregious recent case of public policy's interfering with highly successful education of American Indian and Alaska Native students was the No Child Left Behind Act of 2000. Driven by demands to attain high performance on standardized tests, teachers truncated or abandoned strong culture-based instruction in favor of instruction thought to prepare students to do well on the tests. This is just one example of how decision makers under external pressures tend to revert to best practices or the One Best Way, evoking historical movements to extinguish Indigenous languages and cultures. This article discusses obstacles to culturally-responsive assessment for Indigenous students, describes examples of efforts in the U.S. and elsewhere to improve assessment for Indigenous students, explores the concept of culturally-valid assessment, and interleaves recommendations for going forward constructively within various sections of the paper.
Diffusion MRI (dMRI) is a vital source of imaging data for identifying anatomical connections in the living human brain that form the substrate for information transfer between brain regions. dMRI can thus play a central role toward our understanding of brain function. The quantitative modeling and analysis of dMRI data deduces the features of neural fibers at the voxel level, such as direction and density. The modeling methods that have been developed range from deterministic to probabilistic approaches. Currently, the Ball-and-Stick model serves as a widely implemented probabilistic approach in the tractography toolbox of the popular FSL software package and FreeSurfer/TRACULA software package. However, estimation of the features of neural fibers is complex under the scenario of two crossing neural fibers, which occurs in a sizeable proportion of voxels within the brain. A Bayesian non-linear regression is adopted, comprised of a mixture of multiple non-linear components. Such models can pose a difficult statistical estimation problem computationally. To make the approach of Ball-and-Stick model more feasible and accurate, we propose a simplified version of Ball-and-Stick model that reduces parameter space dimensionality. This simplified model is vastly more efficient in the terms of computation time required in estimating parameters pertaining to two crossing neural fibers through Bayesian simulation approaches. Moreover, the performance of this new model is comparable or better in terms of bias and estimation variance as compared to existing models.


