Publications
A Fast and Integrative Algorithm for Clustering Performance Evaluation in Author Name Disambiguation
Clustering results in author name disambiguation are often evaluated by measures such as Cluster-F, K-metric, Pairwise-F, Splitting and Lumping Error, and B-cubed. Although these measures have different evaluation approaches, this paper shows that they can be calculated in a single framework by a set of common steps that compare truth and predicted clusters through two hash tables recording information about name instances with their predicted cluster indices and frequencies of those indices per truth cluster. This integrative calculation reduces greatly calculation runtime, which is scalable to a clustering task involving millions of name instances within a few seconds. During the integration process, B-cubed and K-metric are shown to produce the same precision and recall scores. In addition, name instance pairs for Pairwise-F are counted using a heuristic, which enables the proposed method to surpass a state-of-the-art algorithm in speedy calculation. Details of the integrative calculation are described with examples and pseudo-code to assist scholars to implement each measure easily and validate the correctness of implementation. The integrative calculation will help scholars compare similarities and differences of multiple measures before they select ones that characterize best the clustering performances of their disambiguation methods.
This technical note describes the results of a pilot approach to link administrative and survey data to better describe the richness and complexity of the research enterprise. In particular, we demonstrate how multiple funding channels can be studied by bringing together two disparate datasets: UMETRICS, which is based on university payroll and financial records, and the Survey of Earned Doctorates (SED), which is one of the most important US survey datasets about the doctoral workforce. We show how it is possible to link data on research funding and the doctorally qualified workforce to describe how many individuals are supported in different disciplines and by different agencies. We outline the potential for more work as the UMETRICS data expands to incorporate more linkages and more access is provided.
Social networks represent two different facets of social life: (1) stable paths for diffusion, or the spread of something through a connected population, and (2) random draws from an underlying social space, which indicate the relative positions of the people in the network to one another. The dual nature of networks creates a challenge: if the observed network ties are a single random draw, is it realistic to expect that diffusion only follows the observed network ties? This study takes a first step toward integrating these two perspectives by introducing a social space diffusion model. In the model, network ties indicate positions in social space, and diffusion occurs proportionally to distance in social space. Practically, the simulation occurs in two parts. First, positions are estimated using a statistical model (in this example, a latent space model). Then, second, the predicted probabilities of a tie from that model-representing the distances in social space-or a series of networks drawn from those probabilities-representing routine churn in the network-are used as weights in a weighted averaging framework. Using longitudinal data from high school friendship networks, the author explores the properties of the model. The author shows that the model produces smoothed diffusion results, which predict attitudes in future waves 10 percent better than a diffusion model using the observed network and up to 5 percent better than diffusion models using alternative, non-model-based smoothing approaches.
This paper uses transaction-based data to provide new insights into the link between the geographic proximity of businesses and associated economic activity. It develops two new measures of, and a set of stylized facts about, the distances between observed transactions between customers and vendors for a research intensive sector. Spending on research inputs is more likely with businesses physically closer to universities than those further away. Firms supplying a university project in one year are more likely to subsequently open an establishment near that university. Vendors who have supplied a project, are subsequently more likely to be a vendor on the same or related project.
Countries, research institutions, and scholars are interested in identifying and promoting high-impact and transformative scientific research. This paper presents a novel set of text-and citation-based metrics that can be used to identify high-impact and transformative works. The 11 metrics can be grouped into seven types: Radical-Generative, Radical-Destructive, Risky, Multidisciplinary, Wide Impact, Growing Impact, and Impact (overall). The metrics are exemplified, validated, and compared using a set of 10,778,696 MEDLINE articles matched to the Science Citation Index Expanded (TM). Articles are grouped into six 5-year periods (spanning 1983-2012) using publication year and into 6,159 fields constructed using comparable MeSH terms, with which each article is tagged. The analysis is conducted at the level of a field-period pair, of which 15,051 have articles and are used in this study. A factor analysis shows that transformativeness and impact are positively related (rho = .402), but represent distinct phenomena. Looking at the subcomponents of transformativeness, there is no evidence that transformative work is adopted slowly or that the generation of important new concepts coincides with the obsolescence of existing concepts. We also find that the generation of important new concepts and highly cited work is more risky. Finally, supporting the validity of our metrics, we show that work that draws on a wider range of research fields is used more widely.
In supervised machine learning for author name disambiguation, negative training data are often dominantly larger than positive training data. This paper examines how the ratios of negative to positive training data can affect the performance of machine learning algorithms to disambiguate author names in bibliographic records. On multiple labeled datasets, three classifiers-Logistic Regression, Naive Bayes, and Random Forest-are trained through representative features such as coauthor names, and title words extracted from the same training data but with various positive-to-negative training data ratios. Results show that increasing negative training data can improve disambiguation performance but with a few percent of performance gains and sometimes degrade it. Logistic and Naive Bayes learn optimal disambiguation models even with a base ratio (1:1) of positive and negative training data. Also, the performance improvement by Random Forest tends to quickly saturate roughly after 1:10 similar to 1:15. These findings imply that contrary to the common practice using all training data, name disambiguation algorithms can be trained using part of negative training data without degrading much disambiguation performance while increasing computational efficiency. This study calls for more attention from author name disambiguation scholars to methods for machine learning from imbalanced data.
Author name ambiguity in a digital library may affect the findings of research that mines authorship data of the library. This study evaluates author name disambiguation in DBLP, a widely used but insufficiently evaluated digital library for its disambiguation performance. In doing so, this study takes a triangulation approach that author name disambiguation for a digital library can be better evaluated when its performance is assessed on multiple labeled datasets with comparison to baselines. Tested on three types of labeled data containing 5000 to 6 M disambiguated names, DBLP is shown to assign author names quite accurately to distinct authors, resulting in pairwise precision, recall, and F1 measures around 0.90 or above overall. DBLP's author name disambiguation performs well even on large ambiguous name blocks but deficiently on distinguishing authors with the same names. Compared to other disambiguation algorithms, DBLP's disambiguation performance is quite competitive, possibly due to its hybrid disambiguation approach combining algorithmic disambiguation and manual error correction. A discussion follows on strengths and weaknesses of labeled datasets used in this study for future efforts to evaluate author name disambiguation on a digital library scale.
Social isolation is broadly associated with poor mental health and risky behaviors in adolescence, a time when peers are critical for healthy development. However, expectations for isolates' substance use remain unclear. Isolation in adolescence may signal deviant attitudes or spur self-medication, resulting in higher substance use. Conversely, isolates may lack access to substances, leading to lower use. Although treated as a homogeneous social condition for teens in much research, isolation represents a multifaceted experience with structurally distinct network components that present different risks for substance use. This study decomposes isolation into conceptually distinct dimensions that are then interacted to create a systematic typology of isolation subtypes representing different positions in the social space of the school. Each isolated position's association with cigarette, alcohol, and marijuana use is tested among 9(th) grade students (n = 10,310, 59% female, 83% white) using cross-sectional data from the PROSPER study. Different dimensions of isolation relate to substance use in distinct ways: unliked isolation is associated with lower alcohol use, whereas disengagement and outside orientation are linked to higher use of all three substances. Specifically, disengagement presents risks for cigarette and marijuana use among boys, and outside orientation is associated with cigarette use for girls. Overall, the adolescents disengaged from their school network who also identify close friends outside their grade are at greatest risk for substance use. This study indicates the importance of considering the distinct social positions of isolation to understand risks for both substance use and social isolation in adolescence.
The science and engineering workforce has aged rapidly in recent years, both in absolute terms and relative to the workforce as a whole. This is a potential concern if the large number of older scientists crowds out younger scientists, making it difficult for them to establish independent careers. In addition, scientists are believed to be most creative earlier in their careers, so the aging of the workforce may slow the pace of scientific progress. We develop and simulate a demographic model, which shows that a substantial majority of recent aging is a result of the aging of the large baby boom cohort of scientists. However, changes in behavior have also played a significant role, in particular, a decline in the retirement rate of older scientists, induced in part by the elimination of mandatory retirement in universities in 1994. Furthermore, the age distribution of the scientific workforce is still adjusting. Current retirement rates and other determinants of employment in science imply a steady-state mean age 2.3 y higher than the 2008 level of 48.6.


