Searcharxiv⌕ Search

arXiv subjects

Zhihao Qiao

Publications and source records attributed to Zhihao Qiao.

3 recordsLinked to original sources

Multi-Absorbing Phase-Type Distributions for Right-Censored Competing Risks Data

Phase-type (PH) distributions are versatile semi-parametric models for lifetime duration and can be used in survival and reliability analysis. In this paper we put forward methods and software for using PH distributions in a competing-risks model. The resulting multi-absorbing phase-type (MAPH) distribution records both the time until absorption and its cause. After formulating basic properties of this family, we develop an EM algorithm for parameter inference from exact event-time/cause observations and independently right-censored observations. For a censored subject, the E-step conditions on survival up to the censoring time and imputes both the latent transient path and its eventual cause of absorption; closed-form cause-resolved sufficient-statistic expectations are obtained from matrix exponentials. We illustrate applications and numerical properties and describe two accompanying Julia packages, in which both the exact-event and the censored E-step are implemented. Applied to intensive-care length-of-stay data in full, censored records included, the method attains a higher likelihood than the phase-type competing-risks fit previously published for those data.

stat.ME↗

Reinforcement Online Learning to Rank with Unbiased Reward Shaping

Online learning to rank (OLTR) aims to learn a ranker directly from implicit feedback derived from users' interactions, such as clicks. Clicks however are a biased signal: specifically, top-ranked documents are likely to attract more clicks than documents down the ranking (position bias). In this paper, we propose a novel learning algorithm for OLTR that uses reinforcement learning to optimize rankers: Reinforcement Online Learning to Rank (ROLTR). In ROLTR, the gradients of the ranker are estimated based on the rewards assigned to clicked and unclicked documents. In order to de-bias the users' position bias contained in the reward signals, we introduce unbiased reward shaping functions that exploit inverse propensity scoring for clicked and unclicked documents. The fact that our method can also model unclicked documents provides a further advantage in that less users interactions are required to effectively train a ranker, thus providing gains in efficiency. Empirical evaluation on standard OLTR datasets shows that ROLTR achieves state-of-the-art performance, and provides significantly better user experience than other OLTR approaches. To facilitate the reproducibility of our experiments, we make all experiment code available at https://github.com/ielab/OLTR.

cs.IR↗

Hidden Equations of Threshold Risk

We consider the problem of sensitivity of threshold risk, defined as the probability of a function of a random variable falling below a specified threshold level $δ>0.$ We demonstrate that for polynomial and rational functions of that random variable there exist at most finitely many risk critical points. The latter are those special values of the threshold parameter for which rate of change of risk is unbounded as $δ$ approaches these threshold values. We characterize candidates for risk critical points as zeroes of either the resolvent of a relevant $δ-$perturbed polynomial, or of its leading coefficient, or both. Thus the equations that need to be solved are themselves polynomial equations in $δ$ that exploit the algebraic properties of the underlying polynomial or rational functions. We name these important equations as "hidden equations of threshold risk".

math.PR↗