SearcharxivSearch

arXiv subjects

David Swanson

Publications and source records attributed to David Swanson.

7 recordsLinked to original sources

Identifying expanding TCR clonotypes with a longitudinal Bayesian mixture model and their associations with cancer patient prognosis, metastasis-directed therapy, and VJ gene enrichment

Examination of T-cell receptor (TCR) clonality has become a way of understanding immunologic response to cancer and its interventions in recent years. An aspect of these analyses is determining which receptors expand or contract statistically significantly as a function of an exogenous perturbation such as therapeutic intervention. We characterize the commonly used Fisher's exact test approach for such analyses and propose an alternative formulation that does not necessitate pairwise, within-patient comparisons. We develop this flexible Bayesian longitudinal mixture model that accommodates variable length patient followup and handles missingness where present, not omitting data in estimation because of structural practicalities. Once clones are partitioned by the model into dynamic (expanding or contracting) and static categories, one can associate their counts or other characteristics with disease state, interventions, baseline biomarkers, and patient prognosis. We apply these developments to a cohort of prostate cancer patients who underwent randomized metastasis-directed therapy or not. Our analyses reveal a significant increase in clonal expansions among MDT patients and their association with later progressions both independent and within strata of MDT. Analysis of receptor motifs and VJ gene enrichment combinations using a high-dimensional penalized log-linear model we develop also suggests distinct biological characteristics of expanding clones, with and without inducement by MDT.

q-bio.QM

Variance component mixture modelling for longitudinal T-cell receptor clonal dynamics

Studies of T cells and their clonally unique receptors have shown promise in elucidating the association between immune response and human disease. Methods to identify T-cell receptor clones which expand or contract in response to certain therapeutic strategies have so far been limited to longitudinal pairwise comparisons of clone frequency with multiplicity adjustment. Here we develop a more general mixture model approach for arbitrary follow-up and missingness which partitions dynamic longitudinal clone frequency behavior from static. While it is common to mix on the location or scale parameter of a family of distributions, the model instead mixes on the parameterization itself, the dynamic component allowing for a variable, Gamma-distributed Poisson mean parameter over longitudinal follow-up, while the static component mean is time invariant. Leveraging conjugacy, one can integrate out the mean parameter for the dynamic and static components to yield distinct posterior predictive distributions whose expressions are a product of negative binomials and a single negative multinomial, respectively, each modified according to an offset for receptor read count normalization. An EM-algorithm is developed to estimate hyperparameters and component membership, and validity of the approach is demonstrated in simulation. The model identifies a statistically significant and clinically relevant increase in TCR clonal dynamism among metastasis-directed radiation therapy in a cohort of prostate cancer patients.

stat.ME

Localizing differences in smooths with simultaneous confidence bounds on the true discovery proportion

We demonstrate a method for localizing where two spline terms, or smooths, differ using a true discovery proportion (TDP) based interpretation. The procedure yields a statement on the proportion of some region where true differences exist between two smooths, which results from use of hypothesis tests on collections of basis coefficients parameterizing the smooths. The methodology avoids otherwise ad hoc means of making such statements like subsetting the data and then performing hypothesis tests on the truncated spline terms. TDP estimates are 1-alpha confidence bounded simultaneously. This means that the TDP estimate for a region is a lower bound on the proportion of actual difference, or true discoveries, in that region with high confidence regardless of the number of regions at which TDP is estimated. Our procedure is based on closed-testing using Simes local test. This local test requires that the `multivariate chi-sq test statistics of generalized Wishart type' underlying the method are positive regression dependent on subsets (PRDS), which we show. The method is well-powered because of a result on the off-diagonal decay structure of the covariance matrix of penalized B-splines of degree two or fewer. We demonstrate achievement of estimated TDP in simulation and analyze a study of walking gait of cerebral palsy patients.

stat.ME

Trua: Efficient Task Replication for Flexible User-defined Availability in Scientific Grids

Failure is inevitable in scientific computing. As scientific applications and facilities increase their scales over the last decades, finding the root cause of a failure can be very complex or at times nearly impossible. Different scientific computing customers have varying availability demands as well as a diverse willingness to pay for availability. In contrast to existing solutions that try to provide higher and higher availability in scientific grids, we propose a model called Task Replication for User-defined Availability (Trua). Trua provides flexible, user-defined, availability in scientific grids, allowing customers to express their desire for availability to computational providers. Trua differs from existing task replication approaches in two folds. First, it relies on the historic failure information collected from the virtual layer of the scientific grids. The reliability model for the failures can be represented with a bimodal Johnson distribution which is different from any existing distributions. Second, it adopts an anomaly detector to filter out anomalous failures; it additionally adopts novel selection algorithms to mitigate the effects of temporary and spatial correlations of the failures without knowing the root cause of the failures. We apply the Trua on real-world traces collected from the Open Science Grid (OSG). Our results show that the Trua can successfully meet user-defined availability demands.

cs.DC

Exploring Erasure Coding Techniques for High Availability of Intermediate Data

Scientific computing workflows generate enormous distributed data that is short-lived, yet critical for job completion time. This class of data is called intermediate data. A common way to achieve high data availability is to replicate data. However, an increasing scale of intermediate data generated in modern scientific applications demands new storage techniques to improve storage efficiency. Erasure Codes, as an alternative, can use less storage space while maintaining similar data availability. In this paper, we adopt erasure codes for storing intermediate data and compare its performance with replication. We also use the metric of Mean-Time-To-Data-Loss (MTTDL) to estimate the lifetime of intermediate data. We propose an algorithm to proactively relocate data redundancy from vulnerable machines to reliable ones to improve data availability with some extra network overhead. Furthermore, we propose an algorithm to assign redundancy units of data physically close to each other on the network to reduce the network bandwidth for reconstructing data when it is being accessed.

cs.DC

Discovering Job Preemptions in the Open Science Grid

The Open Science Grid(OSG) is a world-wide computing system which facilitates distributed computing for scientific research. It can distribute a computationally intensive job to geo-distributed clusters and process job's tasks in parallel. For compute clusters on the OSG, physical resources may be shared between OSG and cluster's local user-submitted jobs, with local jobs preempting OSG-based ones. As a result, job preemptions occur frequently in OSG, sometimes significantly delaying job completion time. We have collected job data from OSG over a period of more than 80 days. We present an analysis of the data, characterizing the preemption patterns and different types of jobs. Based on observations, we have grouped OSG jobs into 5 categories and analyze the runtime statistics for each category. we further choose different statistical distributions to estimate probability density function of job runtime for different classes.

cs.DC

The coarea formula for Sobolev mappings

We extend Federer's coarea formula to mappings $f$ belonging to the Sobolev class $W^{1,p}(R^n;R^m)$, $1 \le m < n$, $p>m$, and more generally, to mappings with gradient in the Lorentz space $L^{m,1}(R^n)$. This is accomplished by showing that the graph of $f$ in $R^{n+m}$ is a Hausdorff $n$-rectifiable set.

math.CA