SearcharxivSearch

arXiv subjects

Hui-Mean Foo

Publications and source records attributed to Hui-Mean Foo.

4 recordsLinked to original sources

Sequential Certification of Threshold Decisions in Rare-Event Risk Prediction

Rare clinical outcomes pose a difficulty deeper than ordinary class imbalance: a penalized logistic model can return finite, stable-looking coefficients before the data support a reliable threshold decision. We formulate the accrual question as decision-targeted sequential certification. On a prespecified finite monitoring schedule and target set, asymptotic prediction bands are adjusted for simultaneous coverage, and certification is assessed only among profiles that might be referred, so that a large low-risk majority cannot trigger an uninformative stop. Under the working rare-event logistic model and stated regularity conditions, each certified decision is asymptotically model-conditionally correct with probability at least \(1-\alpha\) over the schedule, and the effective information scales with the number of genuine events. Simulations and a US linked birth/infant-death application illustrate the gap between a model that is merely estimable and one whose decisions are certifiable: the whole-population rule stopped while more than half of referral-relevant profiles remained ambiguous, whereas the decision-targeted rule did not certify by the 300{,}000-birth horizon despite stable temporal validation. Regularization makes a rare-event model estimable but does not substitute for genuine rare-event information.

stat.AP

Estimating the Conditional Forecast-Revision Scale in Sequential Models: Local-Smoothing Limits, Matched Models, and Cost--Accuracy Trade-offs

The \emph{conditional forecast-revision scale} $\It=\{\Var(\E[X_{t+1}\mid\F_t]\mid\F_{t-1})\}^{1/2}$ measures the history-specific size of the forecast update induced by observing $X_t$. Because it is a conditional second moment built from two unknown conditional means, it is not directly observed. We study which estimator of $\It$ should be used under different structural assumptions and computational budgets. The comparison includes a block bootstrap, a conditional-variance model, a fitted state-space model, two $O(1)$ streaming smoothers, and the forget gate of an already-trained recurrent network. An error decomposition separates one-step-prediction error from conditional-second-moment tracking error. We show that externally tuned lag-only smoothers can be inconsistent when $\It$ changes at the sampling scale, although they attain the usual $T^{-2/3}$ mean-squared-error rate ($T^{-1/3}$ for $\It$) under slow variation; a correctly specified state-space estimator escapes this limit by using the current state. In volatility-driven designs, a cheap conditional-variance model is more accurate and over one hundred times cheaper \emph{as a point estimator} than the implemented block bootstrap, whose value lies in the sampling distribution it provides rather than in point tracking. In state-driven designs, only the structurally matched filter recovers the fast variation. Read directly, a trained network's forget gate does not track $\It$ --- though a supervised linear probe on the full gate vector does, so $\It$ is linearly decodable but not available for free. These results yield a practical rule: identify the conditional-second-moment structure, match the estimator to it, and then choose the least costly adequate method.

stat.ME

Coupling Precipitation Forecasting and Early Warning with Reverse-Martingale Recurrent Neural Networks

Precipitation forecasts are judged by accuracy, but the decisions they support -- when to restrict water, when to warn of drought -- turn on noticing when a local regime is becoming abnormal, which forecast scores alone do not reveal. We ask whether one recurrent model can do both with little or no loss in forecast skill. We add a backward-coherence (reverse-martingale) penalty that keeps the network's hidden state smooth when read backward in time; the size of the resulting reconstruction defect becomes an online warning signal, monitored by a sequential change-point detector. The design is deliberately conservative. On real daily station data from four contrasting climates -- monsoonal Taiwan, semi-arid Texas, temperate Germany, and Mediterranean Anatolia (Turkey) -- the model matches a standard network's forecast skill everywhere, and makes the hidden state markedly steadier in every region. The novelty is the added information: on these real droughts the signal can alarm well ahead of the operational SPI-3 index, giving lead that neither the forecast nor the index provides. This benefit is not uniform across the four regions -- large in one, partial in two others, and near-absent in the fourth. We offer the hydroclimatic character of drought onset, whether it precedes or merely coincides with the rainfall deficit, as a plausible explanation to be tested in future work, supported by a controlled synthetic study with known onset times. The contribution is thus a new and conservative way to read precipitation records: no loss in forecast skill, a steadier model, and an early-warning signal beyond the standard index.

stat.AP

Optimal Stopping in Sequential Clinical Prediction

Most clinical prediction studies are developed from retrospective cohorts and reported as if all patient information were observed at once. In practice, clinicians face a more consequential question: \emph{when is there already enough information to stop testing and act?} A later stage can produce a better-looking model and still fail to justify the added delay, burden, or invasiveness of further workup. We formulate sequential clinical prediction as an \emph{optimal-stopping} problem under staged information, and illustrate the framework across four retrospective clinical datasets. The preferred stopping stage differed substantially by setting: sometimes fuller information justified waiting, whereas in other cases early or intermediate action was preferable. The key object is the patient-specific conditional risk trajectory: forward martingale structure represents coherent risk updating across stages, while reverse-martingale ideas describe information loss when a richer predictor is replaced by a simpler score. The results demonstrate that the best-performing model is not always the best stage for clinical decision-making.

stat.ME