SearcharxivSearch

arXiv subjects

Javier Cabrera

Publications and source records attributed to Javier Cabrera.

13 recordsLinked to original sources

From Cumulative Weights to Marginal Density Ratios: Per-Protocol Estimation in Sequential Target Trial Emulation

Sequential target trial emulation evaluates eligibility at multiple baseline times to emulate a sequence of randomized trials using observational data. Estimating per-protocol effects in this setting is challenging because treatment deviations and loss to follow-up induce selection among individuals who remain observed and adherent over time. Conventional inverse-probability methods address this selection using cumulative weights constructed from estimated adherence and censoring probabilities, but these weights can be highly variable, leading to unstable and imprecise effect estimates. We propose a different approach based on marginal density ratios (MDRs). The MDR directly compares the state distribution among individuals who would remain event-free under a target treatment strategy with the corresponding distribution among observed-adherent individuals. We use longitudinal g-computation to generate the target risk sets and a probabilistic classifier to estimate density ratios for reweighting the observed outcomes. Building on this approach, we also develop a doubly robust extension. Favorable performance across the simulation study suggests that MDR weighting is a promising alternative to cumulative longitudinal weights when its identification assumptions are plausible.

stat.ME

Advancing Evidence Generation in Biomedical Research Using Natural Hermite and Propensity Score Indices: Applications to External Control Arms

When it is not feasible to conduct randomized controlled trials (RCTs), the use of external control arms based on real-world data (RWD) may be a viable option. However, challenges arising from data heterogeneity must be addressed to ensure the reliability of trial results. We consider the use of Natural Hermite and propensity score indices to facilitate robust comparisons between RCTs and RWD studies. Illustrations are provided on the implementation and performance of the underlying algorithms using simulated data, as well as synthetic data from a clinical trial and RWD.

stat.AP

A Novel Two-stage Deming Regression Framework with Applications to Association Analysis between Clinical Risks

In healthcare, clinical risks are crucial for treatment decisions, yet the analysis of their associations is often overlooked. This gap is particularly significant when balancing risks that are weighed against each other, as in the case of atrial fibrillation (AF) patients facing stroke and bleeding risks with anticoagulant medication. While traditional regression models are ill-suited for this task due to standard errors in risk estimation, a novel two-stage Deming regression framework is proposed to address this issue, offering a more accurate tool for analyzing associations between variables observed with errors of known or estimated variances. The first stage is to obtain the variable values with variances of errors either by estimation or observation, followed by the second stage that fits a Deming regression model potentially subject to a transformation. The second stage accounts for the uncertainties associated with both independent and response variables, including known or estimated variances and additional unknown variances from the model. The complexity arising from different scenarios of uncertainty is handled by existing and advanced variations of Deming regression models. An important practical application is to support personalized treatment recommendations based on clinical risk associations that were identified by the proposed framework. The model's effectiveness is demonstrated by applying it to a real-world dataset of AF-diagnosed patients to explore the relationship between stroke and bleeding risks, providing crucial guidance for making informed decisions regarding anticoagulant medication. Furthermore, the model's versatility in addressing data containing multiple sources of uncertainty such as privacy-protected data suggests promising avenues for future research in regression analysis.

stat.AP

Data Nuggets: A Method for Reducing Big Data While Preserving Data Structure

Big data, with NxP dimension where N is extremely large, has created new challenges for data analysis, particularly in the realm of creating meaningful clusters of data. Clustering techniques, such as K-means or hierarchical clustering are popular methods for performing exploratory analysis on large datasets. Unfortunately, these methods are not always possible to apply to big data due to memory or time constraints generated by calculations of order PxN(N-1). To circumvent this problem, typically, the clustering technique is applied to a random sample drawn from the dataset: however, a weakness is that the structure of the dataset, particularly at the edges, is not necessarily maintained. We propose a new solution through the concept of "data nuggets", which reduce a large dataset into a small collection of nuggets of data, each containing a center, weight, and scale parameter. The data nuggets are then input into algorithms that compute methods such as principal components analysis and clustering in a more computationally efficient manner. We show the consistency of the data nuggets-based covariance estimator and apply the methodology of data nuggets to perform exploratory analysis of a flow cytometry dataset containing over one million observations using PCA and K-means clustering for weighted observations. Supplementary materials for this article are available online.

stat.ME

A New Projection Pursuit Index for Big Data

Visualization of extremely large datasets in static or dynamic form is a huge challenge because most traditional methods cannot deal with big data problems. A new visualization method for big data is proposed based on Projection Pursuit, Guided Tour and Data Nuggets methods, that will help display interesting hidden structures such as clusters, outliers, and other nonlinear structures in big data. The Guided Tour is a dynamic graphical tool for high-dimensional data combining Projection Pursuit and Grand Tour methods. It displays a dynamic sequence of low-dimensional projections obtained by using Projection Pursuit (PP) index functions to navigate the data space. Different PP indices have been developed to detect interesting structures of multivariate data but there are computational problems for big data using the original guided tour with these indices. A new PP index is developed to be computable for big data, with the help of a data compression method called Data Nuggets that reduces large datasets while maintaining the original data structure. Simulation studies are conducted and a real large dataset is used to illustrate the proposed methodology. Static and dynamic graphical tools for big data can be developed based on the proposed PP index to detect nonlinear structures.

stat.ME

Principal Phrase Mining

Extracting frequent words from a collection of texts is commonly performed in many subjects. However, as useful as it is to obtain a collection of commonly occurring words from texts, there is a need for more specific information to be obtained from texts in the form of most commonly occurring phrases. Despite this need, extracting frequent phrases is not commonly done due to inherent complications, the most significant being double-counting. Double-counting occurs when words or phrases are counted when they appear inside longer phrases that themselves are also counted, resulting in a selection of mostly meaningless phrases that are frequent only because they occur inside frequent super phrases. Several papers have been written on phrase mining that describe solutions to this issue; however, they either require a list of so-called quality phrases to be available to the extracting process, or they require human interaction to identify those quality phrases during the process. We present here a method that eliminates double-counting via a unique rectification process that does not require lists of quality phrases. In the context of a set of texts, we define a principal phrase as a phrase that does not cross punctuation marks, does not start with a stop word, with the exception of the stop words "not" and "no", does not end with a stop word, is frequent within those texts without being double counted, and is meaningful to the user. Our method identifies such principal phrases independently without human input, and enables their extraction from any texts within a reasonable amount of time.

cs.CL

Statistical Depth based Normalization and Outlier Detection of Gene Expression Data

Normalization and outlier detection belong to the preprocessing of gene expression data. We propose a natural normalization procedure based on statistical data depth which normalizes to the distribution of gene expressions of the most representative gene expression of the group. This differ from the standard method of quantile normalization, based on the coordinate-wise median array that lacks of the well-known properties of the one-dimensional median. The statistical data depth maintains those good properties. Gene expression data are known for containing outliers. Although detecting outlier genes in a given gene expression dataset has been broadly studied, these methodologies do not apply for detecting outlier samples, given the difficulties posed by the high dimensionality but low sample size structure of the data. The standard procedures used for detecting outlier samples are visual and based on dimension reduction techniques; instances are multidimensional scaling and spectral map plots. For detecting outlier genes in a given gene expression dataset, we propose an analytical procedure and based on the Tukey's concept of outlier and the notion of statistical depth, as previous methodologies lead to unassertive and wrongful outliers. We reveal the outliers of four datasets; as a necessary step for further research.

stat.ME

Bootstrap Confidence Intervals Using the Likelihood Ratio Test in Changepoint Detection

This study aims to evaluate the performance of power in the likelihood ratio test for changepoint detection by bootstrap sampling, and proposes a hypothesis test based on bootstrapped confidence interval lengths. Assuming i.i.d normally distributed errors, and using the bootstrap method, the changepoint sampling distribution is estimated. Furthermore, this study describes a method to estimate a data set with no changepoint to form the null sampling distribution. With the null sampling distribution, and the distribution of the estimated changepoint, critical values and power calculations can be made, over the lengths of confidence intervals.

stat.ME

Abstract Mining

We have developed an application that will take a "MEDLINE" output from the PubMed database and allows the user to cluster all non-trivial words of the abstracts of the PubMed output. The number of clusters to use can be selected by the user. A specific cluster may be selected, and the PMIDs and dates for all publications in the selected cluster are displayed underneath. See figure 2, where cluster 12 is selected. The application also has an "Abstracts" tab, where the abstracts for the selected cluster can be perused. Here, it is also possible to download a HTML file containing the PMID, date, title, and abstract for each publication in the selected cluster. A third tab is called "Titles", where all the titles for the selected cluster are displayed. Via a "Use Cluster" button, the selected Cluster can itself be clustered. A "Back" button allows the user to return to any previous state. Finally, it is also possible to exclude documents whose abstracts contain certain words (see figure 3). The application will allow researchers to enter general search terms in the PubMed search engine, then use the application to search for publications of special interest within those search terms.

cs.DL

Composite Inference for Gaussian Processes

Large-scale Gaussian process models are becoming increasingly important and widely used in many areas, such as, computer experiments, stochastic optimization via simulation, and machine learning using Gaussian processes. The standard methods, such as maximum likelihood estimation (MLE) for parameter estimation and the best linear unbiased predictor (BLUP) for prediction, are generally the primary choices in many applications. In spite of their merits, those methods are not feasible due to intractable computation when the sample size is huge. A novel method for the purposes of parameter estimation and prediction is proposed to solve the computational problems of large-scale Gaussian process based models, by separating the original dataset into tractable subsets. This method consistently combines parameter estimation and prediction by making full use of the dependence among conditional densities: a statistically efficient composite likelihood based on joint distributions of some well selected conditional densities is developed to estimate parameters and then "composite inference" is coined to make prediction for an unknown input point, based on its distributions conditional on each block subset. The proposed method transforms the intractable BLUP into a tractable convex optimization problem. It is also shown that the prediction given by the proposed method, called the best linear unbiased block predictor, has a minimum variance for a given separation of the dataset. Keywords: Large scale, Parallel computing, Composite likelihood, Spatial process

stat.ME

Clinical and Non-clinical Effects on Surgery Duration: Statistical Modeling and Analysis

Surgery duration is usually used as an input to the operation room (OR) allocation and surgery scheduling problems. A good estimation of surgery duration benefits the operation planning in ORs. In contrast, we would like to investigate whether the allocation decisions in turn influence surgery duration. Using almost two years of data from a large hospital in China, we find evidence in support of our conjecture. Surgery duration decreases with the number of surgeries a surgeon performs in a day. Numerically, surgery duration will decrease by 10 minutes on average if a surgeon performs one more surgery. Furthermore, we find a non-linear relationship between surgery duration and the number of surgeries allocated to an OR. Also, a surgery's duration is affected by its position in a sequence of surgeries performed by one surgeon. In addition, surgeons exhibit different patterns on the effects of surgery type and position. Since the findings are obtained from a particular data set, We do not claim the generalizability. Instead, the analysis in this paper provides insights into surgery duration study in ORs.

stat.AP

Predicting Surgery Duration from a New Perspective: Evaluation from a Database on Thoracic Surgery

BACKGROUND: Clinical factors influence surgery duration. This study also investigated non-clinical effects. METHODS: 22 months of data about thoracic operations in a large hospital in China were reviewed. Linear and nonlinear regression models were used to predict the duration of the operations. Interactions among predictors were also considered. RESULTS: Surgery duration decreased with the number of operations a surgeon performed in a day (P<0.001). Also, it was found that surgery duration decreased with the number of operations allocated to an OR as long as there were no more than four surgeries per day in the OR (P<0.001), but increased with the number of operations if it was more than four (P<0.01). The duration of surgery was affected by its position in a sequence of surgeries performed by a surgeon. In addition, surgeons exhibited different patterns of the effects of surgery type for surgeries in different positions in the day. CONCLUSIONS: Surgery duration was affected not only by clinical effects but also some non-clinical effects. Scheduling and allocation decisions significantly influenced surgery duration.

stat.AP

Estimating the proportion of differentially expressed genes in comparative DNA microarray experiments

DNA microarray experiments, a well-established experimental technique, aim at understanding the function of genes in some biological processes. One of the most common experiments in functional genomics research is to compare two groups of microarray data to determine which genes are differentially expressed. In this paper, we propose a methodology to estimate the proportion of differentially expressed genes in such experiments. We study the performance of our method in a simulation study where we compare it to other standard methods. Finally we compare the methods in real data from two toxicology experiments with mice.

stat.ME