Publications
A Bayesian framework for group testing under dilution effects has been developed, using lattice-based models. This work has particular relevance given the pressing public health need to enhance testing capacity for coronavirus disease 2019 and future pandemics, and the need for wide-scale and repeated testing for surveillance under constantly varying conditions. The proposed Bayesian approach allows for dilution effects in group testing and for general test response distributions beyond just binary outcomes. It is shown that even under strong dilution effects, an intuitive group testing selection rule that relies on the model order structure, referred to as the Bayesian halving algorithm, has attractive optimal convergence properties. Analogous look-ahead rules that can reduce the number of stages in classification by selecting several pooled tests at a time are proposed and evaluated as well. Group testing is demonstrated to provide great savings over individual testing in the number of tests needed, even for moderately high prevalence levels. However, there is a trade-off with higher number of testing stages, and increased variability. A web-based calculator is introduced to assist in weighing these factors and to guide decisions on when and how to pool under various conditions. High-performance distributed computing methods have also been implemented for considering larger pool sizes, when savings from group testing can be even more dramatic.
The COVID-19 pandemic underscored the necessity for disease surveillance using group testing. Novel Bayesian methods using lattice models were proposed, which offer substantial improvements in group testing efficiency by precisely quantifying uncertainty in diagnoses, acknowledging varying individual risk and dilution effects, and guiding optimally convergent sequential pooled test selections using a Bayesian Halving Algorithm. Computationally, however, Bayesian group testing poses considerable challenges as computational complexity grows exponentially with sample size. This can lead to shortcomings in reaching a desirable scale without practical limitations. We propose a new framework for scaling Bayesian group testing based on Spark: SBGT. We show that SBGT is lightning fast and highly scalable. In particular, SBGT is up to 376x, 1733x, and 1523x faster than the state-of-the-art framework in manipulating lattice models, performing test selections, and conducting statistical analyses, respectively, while achieving up to 97.9% scaling efficiency up to 4096 CPU cores. More importantly, SBGT fulfills our mission towards reaching applicable scale for guiding pooling decisions in wide-scale disease surveillance, and other large scale group testing applications.
The COVID-19 pandemic has necessitated disease surveillance using group testing. Novel Bayesian methods using lattice models were proposed, which offer substantial improvements in group testing efficiency by precisely quantifying uncertainty in diagnoses, acknowledging varying individual risk and dilution effects, and guiding optimally convergent sequential pooled test selections. Computationally, however, Bayesian group testing poses considerable challenges as computational complexity grows exponentially with sample size. HPC and big data stacks are needed for assessing computational and statistical performance across fluctuating prevalence levels at large scales. Here, we study how to design and optimize critical computational components of Bayesian group testing, including lattice model representation, test selection algorithms, and statistical analysis schemes, under the context of parallel computing. To realize this, we propose a high-performance Bayesian group testing framework named HiBGT, based on Apache Spark, which systematically explores the design space of Bayesian group testing and provides comprehensive heuristics on how to achieve highperformance, highly scalable Bayesian group testing. We show that HiBGT can perform large-scale test selections (> 250 state iterations) and accelerate statistical analyzes up to 15.9x (up to 363x with little trade-offs) through a varied selection of sophisticated parallel computing techniques while achieving near linear scalability using up to 924 CPU cores.
Introduction: Functional magnetic resonance imaging (fMRI) often involves long scanning durations to ensure the associated brain activity can be detected. However, excessive experimentation can lead to many undesirable effects, such as from learning and/or fatigue effects, discomfort for the subject, excessive motion artifacts and loss of sustained attention on task. Overly long experimentation can thus have a detrimental effect on signal quality and accurate voxel activation detection. Here, we propose dynamic experimentation with real-time fMRI using a novel statistically driven approach that invokes early stopping when sufficient statistical evidence for assessing the task-related activation is observed.Methods: Voxel-level sequential probability ratio test (SPRT) statistics based on general linear models (GLMs) were implemented on fMRI scans of a mathematical 1-back task from 12 healthy teenage subjects and 11 teenage subjects born extremely preterm (EPT). This approach is based on likelihood ratios and allows for systematic early stopping based on target statistical error thresholds. We adopt a two-stage estimation approach that allows for accurate estimates of GLM parameters before stopping is considered. Early stopping performance is reported for different first stage lengths, and activation results are compared with full durations. Finally, group comparisons are conducted with both early stopped and full duration scan data. Numerical parallelization was employed to facilitate completion of computations involving a new scan within every repetition time (TR).Results: Use of SPRT demonstrates the feasibility and efficiency gains of automated early stopping, with comparable activation detection as with full protocols. Dynamic stopping of stimulus administration was achieved in around half of subjects, with typical time savings of up to 33% (4 min on a 12 min scan). A group analysis produced similar patterns of activity for control subjects between early stopping and full duration scans. The EPT group, individually, demonstrated more variability in location and extent of the activations compared to the normal term control group. This was apparent in the EPT group results, reflected by fewer and smaller clusters.Conclusion: A systematic statistical approach for early stopping with real-time fMRI experimentation has been implemented. This dynamic approach has promise for reducing subject burden and fatigue effects.
A key area of research in epilepsy neurological disorder is the characterization of epileptic networks as they form and evolve during seizure events. In this paper, we describe the development and application of an integrative workflow to analyze functional and structural connectivity measures during seizure events using stereotactic electroencephalogram (SEEG) and diffusion weighted imaging data (DWI). We computed structural connectivity measures using electrode locations involved in recording SEEG signal data as reference points to filter fiber tracts. We used a new workflow-based tool to compute functional connectivity measures based on non-linear correlation coefficient, which allows the derivation of directed graph structures to represent coupling between signal data. We applied a hierarchical clustering based network analysis method over the functional connectivity data to characterize the organization of brain network into modules using data from 27 events across 8 seizures in a patient with refractory left insula epilepsy. The visualization of hierarchical clustering values as dendrograms shows the formation of connected clusters first within each insulae followed by merging of clusters across the two insula; however, there are clear differences between the network structures and clusters formed across the 8 seizures of the patient. The analysis of structural connectivity measures showed strong connections between contacts of certain electrodes within the same brain hemisphere with higher prevalence in the perisylvian/opercular areas. The combination of imaging and signal modalities for connectivity analysis provides information about a patient-specific dynamical functional network and examines the underlying structural connections that potentially influences the properties of the epileptic network. We also performed statistical analysis of the absolute changes in correlation values across all 8 seizures during a baseline normative time period and different seizure events, which showed decreased correlation values during seizure onset; however, the changes during ictal phases were varied.
Objective: The Cognition Battery of the National Institutes of Heath Toolbox is a commonly utilized set of assessments of neuropsychological abilities, evaluating executive function, attention, working memory, processing speed, and episodic memory. We highlight the utility of an advanced statistical model in providing nuanced characterization of neurocognition in an adolescent population. We propose that partially ordered set (POSET) models are well suited to analyze polyfactorial tasks and identify distinct profiles of cognitive functioning. Method: Two models were considered using POSET classification. The first modeled 5 distinct cognitive functions and allowed for multiple functions to contribute to task performance. The second simpler model involved only 2 broader-based functions without polyfactorial task specifications. Existing performance data from 745 adolescents aged 14-17 years were analyzed. Posterior probabilities of classification performance and the discriminatory properties of the estimated response distributions indicated how well the modeling approaches fit the data. Results: The larger first model resulted in 8 profiles or states characterized by combinations of high or low functioning in 5 distinct functions. The simpler second model involved 2 broader-based functions that resulted in 4 states. Comparing model fit criteria, we believe that the finer-grained first model may better reflect the cognitive constructs associated with the tasks. Notably, POSET modeling did not always provide adequate classification of working memory because of the limited design of the Cognition Battery. Conclusions: We demonstrate that the use of POSET models is a feasible approach for detailed analysis of neurocognitive data that can extract information on cognitive functions, even when provided with limited task batteries.
Diffusion MRI (dMRI) is a vital source of imaging data for identifying anatomical connections in the living human brain that form the substrate for information transfer between brain regions. dMRI can thus play a central role toward our understanding of brain function. The quantitative modeling and analysis of dMRI data deduces the features of neural fibers at the voxel level, such as direction and density. The modeling methods that have been developed range from deterministic to probabilistic approaches. Currently, the Ball-and-Stick model serves as a widely implemented probabilistic approach in the tractography toolbox of the popular FSL software package and FreeSurfer/TRACULA software package. However, estimation of the features of neural fibers is complex under the scenario of two crossing neural fibers, which occurs in a sizeable proportion of voxels within the brain. A Bayesian non-linear regression is adopted, comprised of a mixture of multiple non-linear components. Such models can pose a difficult statistical estimation problem computationally. To make the approach of Ball-and-Stick model more feasible and accurate, we propose a simplified version of Ball-and-Stick model that reduces parameter space dimensionality. This simplified model is vastly more efficient in the terms of computation time required in estimating parameters pertaining to two crossing neural fibers through Bayesian simulation approaches. Moreover, the performance of this new model is comparable or better in terms of bias and estimation variance as compared to existing models.


