Publications
Studying 5.6 million biomedical science articles published over three decades, we reconcile conflicts in a long-standing interdisciplinary literature on scientists' life-cycle productivity by controlling for selective attrition and distinguishing between research quantity and quality. While research quality declines monotonically over the career, this decline is easily overlooked because higher ability authors have longer publishing careers. Our results have implications for broader questions of human capital accumulation over the career and federal research policies that shift funding to early-career researchers-while funding researchers at their most creative, these policies must be undertaken carefully because young researchers are less able on average.
Using unique, new, matched UMETRICS data on people employed on research projects and Author-ity data on biomedical publications, this paper shows that National Institutes of Health funding stimulates research by supporting the teams that conduct it. While faculty-both principal investigators (PIs) and other faculty-and their productivity are heavily affected by funding, so are trainees and staff. The largest effects of funding on research output are ripple effects on publications that do not include PIs. While funders focus on research output from projects, they would be well advised to consider how funding ripples through the wide range of people, including trainees and staff, employed on projects.
We study the effects of peer gender composition in STEM doctoral programs on persistence and degree completion. Leveraging unique new data and quasi-random variation in gender composition across cohorts within programs, we show that women entering cohorts with no female peers are 11.7 percentage points less likely to graduate within 6 years than their male counterparts. A 1 standard deviation increase in the percentage of female students differentially increases women’s probability of on-time graduation by 4.4 percentage points. These gender peer effects function primarily through changes in the probability of dropping out in a PhD program’s first year. © 2022 The University of Chicago. All rights reserved.
Attempts to find central influencers, opinion leaders, hubs, optimal seeds, or other important people who can hasten or slow diffusion or social contagion has long been a major research question in network science. We demonstrate that opinion leadership occurs only under conventional but implausible scope conditions. We demonstrate that a highly central node is a more effective seed for diffusion than a random node if nodes can only learn via the network. However, actors are also subject to external influences such as mass media and advertising. We find that diffusion is noticeably faster when it begins with a high centrality node, but that this advantage only occurs in the region of parameter space where external influence is constrained to zero and collapses catastrophically even at minimal levels of external influence. Importantly, nearly all prior agent-based research on choosing a seed or seeds implicitly occurs in the network influence only region of parameter space. We demonstrate this effect using preferential attachment, small world, and several empirical networks. These networks vary in how large the baseline opinion leadership effect is, but in all of them it collapses with the introduction of external influence. This implies that, in marketing and public health, advertising broadly may be underrated as a strategy for promoting network-based diffusion.
In author name disambiguation, author forenames are used to decide which name instances are disambiguated together and how much they are likely to refer to the same author. Despite such a crucial role of forenames, their effect on the performance of heuristic (string matching) and algorithmic disambiguation is not well understood. This study assesses the contributions of forenames in author name disambiguation using multiple labeled data sets under varying ratios and lengths of full forenames, reflecting real-world scenarios in which an author is represented by forename variants (synonym) and some authors share the same forenames (homonym). The results show that increasing the ratios of full forenames substantially improves both heuristic and machine-learning-based disambiguation. Performance gains by algorithmic disambiguation are pronounced when many forenames are initialized or homonyms are prevalent. As the ratios of full forenames increase, however, they become marginal compared to those by string matching. Using a small portion of forename strings does not reduce much the performances of both heuristic and algorithmic disambiguation methods compared to using full-length strings. These findings provide practical suggestions, such as restoring initialized forenames into a full-string format via record linkage for improved disambiguation performances. © 2019 ASIS&T
y Conference publications in computer science (CS) have attracted scholarly attention due to their unique status as a main research outlet, unlike other science fields where journals are dominantly used for communicating research findings. One frequent research question has been how different conference and journal publications are, considering an article as a unit of analysis. This study takes an author-based approach to analyze the publishing patterns of 517,763 scholars who have ever published both in CS conferences and journals for the last 57 years, as recorded in DBLP. The analysis shows that the majority of CS scholars tend to make their scholarly debut, publish more articles, and collaborate with more coauthors in conferences than in journals. Importantly, conference articles seem to serve as a distinct channel of scholarly communication, not a mere preceding step to journal publications: coauthors and title words of authors across conferences and journals tend not to overlap much. This study corroborates findings of previous studies on this topic from a distinctive perspective and suggests that conference authorship in CS calls for more special attention from scholars and administrators outside CS who have focused on journal publications to mine authorship data and evaluate scholarly performance.
Several studies have found that collaboration networks are scale-free, proposing that such networks can be modeled by specific network evolution mechanisms like preferential attachment. This study argues that collaboration networks can look more or less scale-free depending on the methods for resolving author name ambiguity in bibliographic data. Analyzing networks constructed from multiple datasets containing 3.4 M similar to 9.6 M publication records, this study shows that collaboration networks in which author names are disambiguated by the commonly used heuristic, i.e., forename-initial-based name matching, tend to produce degree distributions better fitted to power-law slopes with the typical scaling parameter (2 < alpha < 3) than networks disambiguated by more accurate algorithm-based methods. Such tendency is observed across collaboration networks generated under various conditions such as cumulative years, 5- and 1-year sliding windows, and random sampling, and through simulation, found to arise due mainly to artefactual entities created by inaccurate disambiguation. This cautionary study calls for special attention from scholars analyzing network data in which entities such as people, organization, and gene can be merged or split by improper disambiguation.
We identify two polar life cycles of scholarly creativity among Nobel laureate economists with Tinbergen falling broadly in the middle. Experimental innovators work inductively, accumulating knowledge from experience. Conceptual innovators work deductively, applying abstract principles. Innovators whose work is more conceptual do their most important work earlier in their careers than those whose work is more experimental. Our estimates imply that the probability that the most conceptual laureate publishes his single best work peaks at age 25 compared to the mid-50 s for the most experimental laureate. Thus, while experience benefits experimental innovators, newness to a field benefits conceptual innovators.
This paper uses transaction-based data to provide new insights into the link between the geographic proximity of businesses and associated economic activity. It develops two new measures of, and a set of stylized facts about, the distances between observed transactions between customers and vendors for a research intensive sector. Spending on research inputs is more likely with businesses physically closer to universities than those further away. Firms supplying a university project in one year are more likely to subsequently open an establishment near that university. Vendors who have supplied a project, are subsequently more likely to be a vendor on the same or related project.
A Fast and Integrative Algorithm for Clustering Performance Evaluation in Author Name Disambiguation
Clustering results in author name disambiguation are often evaluated by measures such as Cluster-F, K-metric, Pairwise-F, Splitting and Lumping Error, and B-cubed. Although these measures have different evaluation approaches, this paper shows that they can be calculated in a single framework by a set of common steps that compare truth and predicted clusters through two hash tables recording information about name instances with their predicted cluster indices and frequencies of those indices per truth cluster. This integrative calculation reduces greatly calculation runtime, which is scalable to a clustering task involving millions of name instances within a few seconds. During the integration process, B-cubed and K-metric are shown to produce the same precision and recall scores. In addition, name instance pairs for Pairwise-F are counted using a heuristic, which enables the proposed method to surpass a state-of-the-art algorithm in speedy calculation. Details of the integrative calculation are described with examples and pseudo-code to assist scholars to implement each measure easily and validate the correctness of implementation. The integrative calculation will help scholars compare similarities and differences of multiple measures before they select ones that characterize best the clustering performances of their disambiguation methods.


