SearcharxivSearch

arXiv subjects

Qianying Lin

Publications and source records attributed to Qianying Lin.

12 recordsLinked to original sources

Simultaneous confidence bands for cumulative hazard via exchangeable bootstrap and box calibration

Resampling-based simultaneous confidence bands for cumulative hazard functions often undercover in finite samples with right censoring. We study two aspects of the construction that can contribute to this gap, the resampling scheme and the calibration statistic, and propose a procedure that intervenes on both. The exchangeable bootstrap reweights the numerator and the denominator of the Nelson-Aalen ratio, preserving its ratio structure. The box-calibrated discrepancy constructs lower and upper step envelopes from adjacent values of the original and resampled Nelson-Aalen estimators and measures the resulting vertical discrepancy. We establish conditional weak convergence of the exchangeable bootstrap, prove that box calibration is first-order asymptotically equivalent to grid calibration, and show that the resulting band attains nominal coverage asymptotically. The box correction uses the same bootstrap paths and event-time grid as grid calibration; after each bootstrap path is formed, it requires only an additional linear pass over the event-time grid and therefore has negligible computational overhead. In simulations across a range of hazard shapes and censoring levels, the exchangeable bootstrap with box calibration is, in most configurations, closest to nominal coverage among the methods considered. A notable consequence is a ranking reversal: the ratio-preserving exchangeable bootstrap has the lowest coverage under grid calibration, yet is usually closest to the nominal level after box calibration. A melanoma data example illustrates the practical effect on the cumulative hazard bands. The proposed procedure operates on the original cumulative-hazard scale, requires no variance-stabilizing transformation, and permits inference from time zero.

stat.ME

Exact phylodynamic likelihood via structured Markov genealogy processes

We show that each member of a broad class of Markovian population models induces a unique stochastic process on the space of genealogies. We construct this genealogy process and derive exact expressions for the likelihood of an observed genealogy in terms of a filter equation, the structure of which is completely determined by the population model. We show that existing phylodynamic methods based on the coalescent and linear birth-death processes are special cases. We derive some properties of filter equations and describe a class of algorithms that can be used to numerically solve them. Importantly, because these algorithms rely only on simulation of the population model, they retain the plug-and-play property upon which simulation-based inference depends. Our results open the door to statistically efficient likelihood-based phylodynamic inference for a much wider class of models than has been possible.

q-bio.QM

A story of viral co-infection, co-transmission and co-feeding in ticks: how to compute an invasion reproduction number

With a single circulating vector-borne virus, the basic reproduction number incorporates contributions from tick-to-tick (co-feeding), tick-to-host and host-to-tick transmission routes. With two different circulating vector-borne viral strains, resident and invasive, and under the assumption that co-feeding is the only transmission route in a tick population, the invasion reproduction number depends on whether the model system of ordinary differential equations possesses the property of neutrality. We show that a simple model, with two populations of ticks infected with one strain, resident or invasive, and one population of co-infected ticks, does not have Alizon's neutrality property. We present model alternatives that are capable of representing the invasion potential of a novel strain by including populations of ticks dually infected with the same strain. The invasion reproduction number is analysed with the next-generation method and via numerical simulations.

q-bio.PE

Tunable robustness in power-law inference

Power-law probability distributions arise often in the social and natural sciences. Statistics have been developed for estimating the exponent parameter as well as gauging goodness-of-fit to a power law. Yet paradoxically, many famous power laws such as the distribution of wealth and earthquake magnitudes have not found good statistical support in data by modern methods. We show that measurement errors such as quantization and noise bias both maximum-likelihood estimators and goodness-of-fit measures. We address this issue using logarithmic binning and the corresponding discrete reference distribution for maximum likelihood estimators and Kolmogorov-Smirnov statistics. Using simulated errors, we validate that binning attenuates bias in parameter estimates and recalibrates goodness of fit to a power law by removing small errors from consideration. These benefits come at modest cost in statistical power, which can be compensated with larger sample sizes. We reanalyse three empirical cases of wealth, earthquake magnitudes and wildfire area and show that binning reverses statistical conclusions and aligns the statistical results with historical and scientific expectations. We explain through these cases how routine errors lead to incorrect conclusions and the necessity for more robust methods.

stat.ME

Sparse Attentive Memory Network for Click-through Rate Prediction with Long Sequences

Sequential recommendation predicts users' next behaviors with their historical interactions. Recommending with longer sequences improves recommendation accuracy and increases the degree of personalization. As sequences get longer, existing works have not yet addressed the following two main challenges. Firstly, modeling long-range intra-sequence dependency is difficult with increasing sequence lengths. Secondly, it requires efficient memory and computational speeds. In this paper, we propose a Sparse Attentive Memory (SAM) network for long sequential user behavior modeling. SAM supports efficient training and real-time inference for user behavior sequences with lengths on the scale of thousands. In SAM, we model the target item as the query and the long sequence as the knowledge database, where the former continuously elicits relevant information from the latter. SAM simultaneously models target-sequence dependencies and long-range intra-sequence dependencies with O(L) complexity and O(1) number of sequential updates, which can only be achieved by the self-attention mechanism with O(L^2) complexity. Extensive empirical results demonstrate that our proposed solution is effective not only in long user behavior modeling but also on short sequences modeling. Implemented on sequences of length 1000, SAM is successfully deployed on one of the largest international E-commerce platforms. This inference time is within 30ms, with a substantial 7.30% click-through rate improvement for the online A/B test. To the best of our knowledge, it is the first end-to-end long user sequence modeling framework that models intra-sequence and target-sequence dependencies with the aforementioned degree of efficiency and successfully deployed on a large-scale real-time industrial recommender system.

cs.IR

Learning-To-Ensemble by Contextual Rank Aggregation in E-Commerce

Ensemble models in E-commerce combine predictions from multiple sub-models for ranking and revenue improvement. Industrial ensemble models are typically deep neural networks, following the supervised learning paradigm to infer conversion rate given inputs from sub-models. However, this process has the following two problems. Firstly, the point-wise scoring approach disregards the relationships between items and leads to homogeneous displayed results, while diversified display benefits user experience and revenue. Secondly, the learning paradigm focuses on the ranking metrics and does not directly optimize the revenue. In our work, we propose a new Learning-To-Ensemble (LTE) framework RAEGO, which replaces the ensemble model with a contextual Rank Aggregator (RA) and explores the best weights of sub-models by the Evaluator-Generator Optimization (EGO). To achieve the best online performance, we propose a new rank aggregation algorithm TournamentGreedy as a refinement of classic rank aggregators, which also produces the best average weighted Kendall Tau Distance (KTD) amongst all the considered algorithms with quadratic time complexity. Under the assumption that the best output list should be Pareto Optimal on the KTD metric for sub-models, we show that our RA algorithm has higher efficiency and coverage in exploring the optimal weights. Combined with the idea of Bayesian Optimization and gradient descent, we solve the online contextual Black-Box Optimization task that finds the optimal weights for sub-models given a chosen RA model. RA-EGO has been deployed in our online system and has improved the revenue significantly.

cs.LG

Markov Genealogy Processes

We construct a family of genealogy-valued Markov processes that are induced by a continuous-time Markov population process. We derive exact expressions for the likelihood of a given genealogy conditional on the history of the underlying population process. These lead to a nonlinear filtering equation which can be used to design efficient Monte Carlo inference algorithms. We demonstrate these calculations with several examples. Existing full-information approaches for phylodynamic inference are special cases of the theory.

math.PR

Diversity Regularized Interests Modeling for Recommender Systems

With the rapid development of E-commerce and the increase in the quantity of items, users are presented with more items hence their interests broaden. It is increasingly difficult to model user intentions with traditional methods, which model the user's preference for an item by combining a single user vector and an item vector. Recently, some methods are proposed to generate multiple user interest vectors and achieve better performance compared to traditional methods. However, empirical studies demonstrate that vectors generated from these multi-interests methods are sometimes homogeneous, which may lead to sub-optimal performance. In this paper, we propose a novel method of Diversity Regularized Interests Modeling (DRIM) for Recommender Systems. We apply a capsule network in a multi-interest extractor to generate multiple user interest vectors. Each interest of the user should have a certain degree of distinction, thus we introduce three strategies as the diversity regularized separator to separate multiple user interest vectors. Experimental results on public and industrial data sets demonstrate the ability of the model to capture different interests of a user and the superior performance of the proposed approach.

cs.IR

The Sampled Moran Genealogy Process

We define the Sampled Moran Genealogy Process, a continuous-time Markov process on the space of genealogies with the demography of the classical Moran process, sampled through time. To do so, we begin by defining the Moran Genealogy Process using a novel representation. We then extend this process to include sampling through time. We derive exact conditional and marginal probability distributions for the sampled process under a stationarity assumption, and an exact expression for the likelihood of any sequence of genealogies it generates. This leads to some interesting observations pertinent to existing phylodynamic methods in the literature. Throughout, our proofs are original and make use of strictly forward-in-time calculations and are exact for all population sizes and sampling processes.

q-bio.PE

Religious Festivals and Influenza

Objectives Influenza outbreaks have been widely studied. However, the patterns between influenza and religious festivals remained unexplored. This study examined the patterns of influenza and Hanukkah in Israel, and that of influenza and Hajj in Bahrain, Egypt, Iraq, Jordan, Oman and Qatar. Method Influenza surveillance data of these seven countries from 2009 to 2017 were downloaded from the FluNet of the World Health Organization. Secondary data were collected for the countries' population, and the dates of Hajj and Hanukkah. We aggregated the weekly influenza A and B laboratory confirmations for each country over the study period. Weekly influenza A patterns and religious festival dates were further explored across the study period. Results We found that influenza A peaks closely followed Hanukkah in Israel in six out of seven years from 2010 to 2017. Aggregated influenza A peaks of the other six Middle East countries also occurred right after Hajj every year during the study period. Conclusions We predict that unless there is an emergence of new influenza strain, such influenza patterns are likely to persist in future years. Our results suggested that the optimal timing of mass influenza vaccination should take into considerations of the dates of these religious festivals.

q-bio.PE

Effects of Reactive Social Distancing on the 1918 Influenza Pandemic

The 1918 influenza pandemic was characterized by multiple epidemic waves. We investigated into reactive social distancing, a form of behavioral responses, and its effect on the multiple influenza waves in the United Kingdom. Two forms of reactive social distancing have been used in previous studies: Power function, which is a function of the proportion of recent influenza mortality in a population, and Hill function, which is a function of the actual number of recent influenza mortality. Using a simple epidemic model with a Power function and one common set of parameters, we provided a good model fit for the observed multiple epidemic waves in London boroughs, Birmingham and Liverpool. Our approach is different from previous studies where separate models are fitted to each city. We then applied these model parameters obtained from fitting three cities to all 334 administrative units in England and Wales and including the population sizes of individual administrative units. We computed the Pearson's correlation between the observed and simulated data for each administrative unit. We achieved a median correlation of 0.636, indicating our model predictions perform reasonably well. Our modelling approach which requires reduced number of parameters resulted in computational efficiency gain without over-fitting the model. Our works have both scientific and public health significance.

q-bio.PE

Spatio-temporal patterns of influenza B proportions

We study the spatio-temporal patterns of the proportion of influenza B out of laboratory confirmations of both influenza A and B, with data from 139 countries and regions downloaded from the FluNet compiled by the World Health Organization, from January 2006 to October 2015, excluding 2009. We restricted our analysis to 34 countries that reported more than 2000 confirmations for each of types A and B over the study period. We find that Pearson's correlation is 0.669 between effective distance from Mexico and influenza B proportion among the countries from January 2006 to October 2015. In the United States, influenza B proportion in the pre-pandemic period (2003-2008) negatively correlated with that in the post-pandemic era (2010-2015) at the regional level. Our study limitations are the country-level variations in both surveillance methods and testing policies. Influenza B proportion displayed wide variations over the study period. Our findings suggest that even after excluding 2009's data, the influenza pandemic still has an evident impact on the relative burden of the two influenza types. Future studies could examine whether there are other additional factors. This study has potential implications in prioritizing public health control measures.

q-bio.PE