Searcharxiv⌕ Search

arXiv · 2610.10321

Estimating Uncoded Crash Factors with Tabular Foundation and System One Models: Kumo Tabular and Jev

Abstract

Road safety programs count the coded fields of police crash records, while the officer's narrative, which often records factors the fields omit, is rarely read. A safety office thus cannot tell how much its counts miss or where to review. This study develops and evaluates a system that joins both views of the 5,601,890 Texas crashes from 2017 to 2025 into population estimates with stated validity. An in-context tabular foundation model, Kumo Tabular, reads the coded record of every crash, a calibrated System One model, Jev, reads the narratives of two probability samples, and human judgments recalibrate its probabilities. A multiwave predict-then-debias estimator joins the three tiers, and a second human tier drawn with recorded probabilities checks the estimates by design. For hydroplaning, medical episodes, fatigue, animals, and phone use, the narrative documents more injury crashes than the coded field, 15,074 against 7,340 for phone use, and the human check agrees with all fifteen estimates within its margin. A re-read list ranked by Kumo Tabular finds confirmed discordance 7 to 58 times as often as random reading. At the planning cost of human coding, one further round of human judgments would cut the root mean square relative half-width from 22.0 to 16.2 percent, against 21.2 for reading every narrative. Two calibrated readers of different views, joined by a sampling design, give a safety office counts, a discordance map, a validated re-read list, and a reading budget, with Kumo Tabular reading the table at 15 times the speed of TabPFN 3.5.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Amir Rafe, Subasish Das. 2026-10-07. Estimating Uncoded Crash Factors with Tabular Foundation and System One Models: Kumo Tabular and Jev. https://arxiv.org/abs/2610.10321

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Evaluating LiDAR Data Sources, Predictor Resolution, and Spatial Random Effects in Bayesian Change-of-Support Models for Forest Inventory

Forest managers need timely stand-level information to support operational planning, particularly in mixed-species, structurally heterogeneous forests facing climate-related disturbance. Model-based estimation combines sparse field data with remotely sensed predictors to estimate growing stock volume (GSV) for small areas. While uncrewed aerial vehicle laser scanning (ULS) offers flexible, high-resolution LiDAR acquisition, its advantages over conventional airborne laser scanning (ALS) remain unclear. We compared publicly available ALS and newly acquired ULS data using Bayesian change-of-support models to estimate GSV in a mixed-species forest in north-eastern Germany. We evaluated distributional LiDAR metrics and spatial random effects. ULS consistently outperformed ALS, achieving cross-validated RMSPEs of 68.5 m^3/ha and 79.5 m3/ha, respectively-a 13.8 % reduction in prediction error. Distributional metrics improved ULS models more strongly, reducing RMSPE by up to 10.1 %; spatial effects provided only minor gains at substantially higher computational cost. ULS also produced lower uncertainty in latent stand-mean GSV estimates. The ULS advantage may reflect both finer-scale canopy information and closer temporal alignment with field measurements. Timely, information rich LiDAR may therefore be more valuable for stand-level GSV estimation than increasingly complex spatial models. Temporally matched ALS-ULS comparisons are needed to isolate platform effects.

stat.AP↗

DeepAJM: Deep Association Joint Model for Irregularly Sampled data

Joint Models simultaneously model longitudinal and survival outcomes, leveraging patterns in patients' longitudinal trajectory to improve the prediction of survival outcomes. The classical parametric joint models, however, rely on fixed parametric assumptions, making them susceptible to bias under model misspecification and smaller sample sizes. We propose a deep joint model, DeepAJM, that does not require any parametric assumptions, while retaining a partially interpretable, per-longitudinal-outcome association structure. The joint model uses an encoder-decoder (sequence-to-sequence) architecture to learn the latent structure in patients' time-varying covariate trajectories. The model links the longitudinal processes to the survival processes through a learned interpretable association structure, in which each longitudinal output from the decoder gets remodulated by baseline covariates before it contributes to the risk scores from the survival head of the architecture. The model was evaluated on three datasets ( a cardiovascular-disease EHR cohort, a primary biliary cirrhosis (PBC2) dataset, and a simulated dataset) against a classical parametric joint model, TransformerJM, DA-LSTM and a Cox-based survival-only model. All models were assessed using C-index, integrated brier score (IBS), time-dependent AUROC, and time-dependent AUPRC. Our model achieved the best discrimination in terms of the C-index, time-dependent AUROC, and AUPRC across all datasets.

stat.AP↗

Bayesian Optimization for Dose Finding with Two Agents: Participant Allocation and Final Selection

In two-agent dose-finding trials, the next cohort should help identify a combination for final selection. We studied a constrained knowledge-gradient (cKG) rule with one-cohort lookahead that updates independent Gaussian-process models of efficacy and continuous toxicity, reapplies a probability criterion for mean toxicity, and evaluates the resulting selection. We derived a deterministic calculation over a fixed set of dose combinations, holding fitted model parameters fixed during each hypothetical update. We compared cKG with constrained expected improvement (cEI) and two toxicity-only rules, targeted mean squared error (tMSE) and entropy, in four synthetic scenarios. In the primary obstructive sleep apnea (OSA)-derived scenario, averaged equally over strata and five probability cutoffs, cKG assigned fewer participants to combinations above the true mean-toxicity limit than tMSE (17.92% versus 27.08%), but selected such combinations more often at trial completion (18.80% versus 11.85%). Compared with cEI, cKG had higher mean simulated reduction in the 4%-desaturation apnea-hypopnea index (AHI4) at final selection (7.46 versus 6.72 events/hour), more above-limit final selections (18.80% versus 10.50%), and more above-limit assignments (17.92% versus 15.10%). Across scenarios, its efficacy advantage over cEI was smaller under stricter toxicity criteria. Continuous outcomes, uncalibrated toxicity limits, and a rule that still selects a combination when none meets the criterion limit clinical interpretation. Allocation and final-selection toxicity should be reported separately, alongside efficacy.

stat.AP↗