Publications
Scholars have long recognized that teachers' social interactions play an important role in their learning and professional development. Still, while a growing body of research shows that teaching-focused social ties can give precollege educators access to valuable information, knowledge, and advice-or social capital-that improves professional practice and student learning, empirical, mixed methods studies on the phenomenon in the higher education sector are rare, and few investigate what conditions are necessary for these social ties to develop among college instructors. Focusing on college faculty in 17 associate- and baccalaureate-level institutions in one U.S. city, this study uses survey and interview data to explore the connections between structural and positional educator characteristics and the social networks, or compilations of social ties, in which faculty reported discussing teaching. Regression analyses of survey responses (n=244) indicate that fewer years of teaching experience, the time faculty take preparing to teach, discipline, and institution type are correlated with social network dimensions linked to improved professional practice. An inductive analysis of interview data from a subset of faculty (n=22) supplements survey findings with descriptions of how teaching experience, organizational support, and other factors constrain and reinforce the development of teaching-focused social ties. Results confirm and extend prior research indicating that the development of teaching-focused social networks and the accrual of ties linked to social capital demand faculty and organizational investment. Findings also suggest that leaders hoping to foster beneficial ties should tailor instructional initiatives to more closely align with faculty experience and time commitments.
The relationship between education policy and workforce policy has long been uneasy. It is widely believed in many quarters of American society that the U.S. education system is in decline and, what’s more, that it bears significant responsibility for a wide range of social ills, including stagnant wages, increasing inequality, high unemployment, and overall economic lethargy. However, as analyzed in this paper, the preponderance of evidence suggests that the U.S. education system has produced ample supplies of students to respond to STEM labor market demand. The “pipeline” of STEM-potential students is similarly strong and expanding.
Time management skills are an essential component of college student success, especially in online classes. Through a randomized control trial of students in a for-credit online course at a public 4-year university, we test the efficacy of a scheduling intervention aimed at improving students' time management. Results indicate the intervention had positive effects on initial achievement scores; students who were given the opportunity to schedule their lecture watching in advance scored about a third of a standard deviation better on the first quiz than students who were not given that opportunity. These effects are concentrated in students with the lowest self-reported time management skills. However, these effects diminish over time such that we see a marginally significant negative effect of treatment on the last week's quiz grade and no difference in overall course scores. We examine the effect of the intervention on plausible mechanisms to explain the observed achievement effects. We find no evidence that the intervention affected cramming, procrastination, or the time at which students did work.
Generating Automatically Labeled Data for Author Name Disambiguation: An Iterative Clustering Method
To train algorithms for supervised author name disambiguation, many studies have relied on hand-labeled truth data that are very laborious to generate. This paper shows that labeled data can be automatically generated using information features such as email address, coauthor names, and cited references that are available from publication records. For this purpose, high-precision rules for matching name instances on each feature are decided using an external-authority database. Then, selected name instances in target ambiguous data go through the process of pairwise matching based on the rules. Next, they are merged into clusters by a generic entity resolution algorithm. The clustering procedure is repeated over other features until further merging is impossible. Tested on 26 K instances out of the population of 228K author name instances, this iterative clustering produced accurately labeled data with pairwise F1=0.99. The labeled data represented the population data in terms of name ethnicity and co-disambiguating name group size distributions. In addition, trained on the labeled data, machine learning algorithms disambiguated 24K names in test data with performance of pairwise F1=0.90-0.92. Several challenges are discussed for applying this method to resolving author name ambiguity in large-scale scholarly data.
A Fast and Integrative Algorithm for Clustering Performance Evaluation in Author Name Disambiguation
Clustering results in author name disambiguation are often evaluated by measures such as Cluster-F, K-metric, Pairwise-F, Splitting and Lumping Error, and B-cubed. Although these measures have different evaluation approaches, this paper shows that they can be calculated in a single framework by a set of common steps that compare truth and predicted clusters through two hash tables recording information about name instances with their predicted cluster indices and frequencies of those indices per truth cluster. This integrative calculation reduces greatly calculation runtime, which is scalable to a clustering task involving millions of name instances within a few seconds. During the integration process, B-cubed and K-metric are shown to produce the same precision and recall scores. In addition, name instance pairs for Pairwise-F are counted using a heuristic, which enables the proposed method to surpass a state-of-the-art algorithm in speedy calculation. Details of the integrative calculation are described with examples and pseudo-code to assist scholars to implement each measure easily and validate the correctness of implementation. The integrative calculation will help scholars compare similarities and differences of multiple measures before they select ones that characterize best the clustering performances of their disambiguation methods.
This study evaluates the accuracy of previously published econometric forecasts for seven lodging sector variables that measure hotel activity in El Paso, Texas. The hotel forecasts have been generated annually using an econometric model of the El Paso metropolitan economy from 2006 forward. Predictive accuracy is evaluated relative to random walk benchmarks. Assessment is completed using both descriptive forecast error summary statistics as well as formal statistical tests. The econometric model outperforms the random walk benchmarks for a majority of the variables analyzed. However, statistical tests of forecast error differentials do not yield conclusive evidence in favor of the econometric historical track record. Tests of directional forecast accuracy also produce mixed results. Although the structural econometric model of hotel business conditions appears to provide useful predictive information, analysts and planners should also monitor recent history closely.
In item response theory (IRT), it is often necessary to perform restricted recalibration (RR) of the model: A set of (focal) parameters is estimated holding a set of (nuisance) parameters fixed. Typical applications of RR include expanding an existing item bank, linking multiple test forms, and associating constructs measured by separately calibrated tests. In the current work, we provide full statistical theory for RR of IRT models under the framework of pseudo-maximum likelihood estimation. We describe the standard error calculation for the focal parameters, the assessment of overall goodness-of-fit (GOF) of the model, and the identification of misfitting items. We report a simulation study to evaluate the performance of these methods in the scenario of adding a new item to an existing test. Parameter recovery for the focal parameters as well as Type I error and power of the proposed tests are examined. An empirical example is also included, in which we validate the pediatric fatigue short-form scale in the Patient-Reported Outcome Measurement Information System (PROMIS), compute global and local GOF statistics, and update parameters for the misfitting items. © 2019 The Psychometric Society.
The goal of this study was to use eye-tracking and log-file data to investigate the impact of prior knowledge on college students' (N = 194, with a subset of n = 30 for eye tracking and sequence mining analyses) fixations on (i.e., looking at) self-regulated learning-related areas of interest (i.e., specific locations on the interface) and on the sequences of engaging in cognitive and metacognitive self-regulated learning processes during learning with MetaTutor, an Intelligent Tutoring System that teaches students about the human circulatory system. Results revealed that there were no significant differences in fixations on single areas of interest by the prior knowledge group students were assigned to; however there were significant differences in fixations on pairs of areas of interest, as evidenced by eye-tracking data. Furthermore, there were significant differences in sequential patterns of engaging in cognitive and metacognitive self-regulated learning processes by students' prior knowledge group, as evidenced from log-file data. Specifically, students with high prior knowledge engaged in processes containing cognitive strategies and metacognitive strategies whereas students with low prior knowledge did not. These results have implications for designing adaptive intelligent tutoring systems that provide individualized scaffolding and feedback based on individual differences, such as levels of prior knowledge.
In decision making under risk, adults tend to overestimate small and underestimate large probabilities (Tversky & Kahneman, 1992). This inverse S-shaped distortion pattern is similar to that observed in a wide variety of proportion judgment tasks (see Hollands & Dyre, 2000, for review). In proportion judgment tasks, distortion patterns tend not to be fixed but rather to depend on the reference points to which the targets are compared. Here, we tested the novel hypothesis that probability distortion in decision making under risk might also be influenced by reference points in this case, references implied by the probability range. Adult participants were assigned to either a full-range (probabilities from 0-100%), upper-range (50-100%), or lower-range (0-50%) condition, where they indicated certainty equivalents for 176 hypothetical monetary gambles (e.g., a 50% chance of $100, otherwise $0). Using a modified cumulative prospect theory model, we found only minimal differences in probability distortion as a function of condition, suggesting no differences in use of reference points by condition, and broadly demonstrating the robustness of distortion pattern across contexts. However, we also observed deviations from the curve across all conditions that warrant further research.
Science consists of a body of knowledge and a set of processes by which the knowledge is produced. Although these have traditionally been treated separately in science instruction, there has been a shift to an integration of knowledge and processes, or set of practices, in how science should be taught and assessed. We explore whether a general overall mastery of the processes drives learning in new science content areas and if this overall mastery can be improved through engaged science learning. Through a review of literature, the paper conceptualizes this general process mastery as scientific sensemaking, defines the sub-dimensions, and presents a new measure of the construct centered in scenarios of general interest to young adolescents. Using a dataset involving over 2500 6th and 8th grade students, the paper shows that scientific sensemaking scores can predict content learning gains and that this relationship is consistent across student characteristics, content of instruction, and classroom environment. Further, students who are behaviorally and cognitively engaged during science classroom activities show greater growth in scientific sensemaking, showing a reciprocal relationship between sensemaking ability and effective science instruction. Findings from this work support early instruction on sensemaking activities to better position students to learn new scientific content.


