SearcharxivSearch

arXiv · 2604.19819

Decision Traces: What Multi-System Data Fusion Reveals About Institutional Knowledge in Enterprise Hiring

Abstract

Enterprise hiring systems generate data across multiple disconnected platforms: applicant tracking systems (ATS) record candidate profiles, human resource information systems (HRIS) record performance outcomes, and behavioral assessments capture personality and behavioral dimensions. Each system operates independently, and the reasoning behind hiring decisions is lost when managers retire, transfer, or leave. Decision traces are structured evidence chains connecting screening inputs, assessment signals, and production outcomes. They have been theorized but never operationalized at production scale. We present, to our knowledge, the first such study: a deployment at a Fortune 500 insurance carrier (N=10,765 agents hired, 2022-2025), where connecting three siloed data systems produced three findings. First, of 8,181 unique skills parsed from ATS profiles (3,597 testable), not a single keyword predicts production after Bonferroni correction; 30 are significantly anti-predictive, and the median keyword is associated with 25% lower odds of production. Requiring insurance experience alone would reject 2,863 agents who produced $17.7M in annual premium credit. Second, personality-based behavioral assessment (Predictive Index) achieves AUC=0.647 standalone and AUC=0.735 when fused with ATS and behavioral scoring data. Third, speed-to-production follows a measurable economic constant of $54/day per agent unadjusted, or $35/day controlling for source channel and tenure, moderated by behavioral score: high-scored agents capture $114/day from speed acceleration versus $41/day for low-scored agents. These findings were invisible within any single system. We discuss implications for hiring system design, the limitations of keyword-based screening, and the conditions under which institutional knowledge can be captured and operationalized.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Saad Bin Shafiq. 2026-04-18. Decision Traces: What Multi-System Data Fusion Reveals About Institutional Knowledge in Enterprise Hiring. https://arxiv.org/abs/2604.19819

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Identification in Linear Quantile Panel Models

This paper studies identification in linear quantile panel models with unrestricted individual heterogeneity when the number of time periods is fixed and small. We impose strict exogeneity, whereby the conditional quantile restriction holds given the individual's complete regressor history and latent individual effect, but otherwise allow the disturbances to be arbitrarily dependent over time.

econ.EM

Experimental Design for Policy Choice

We show how to optimally design experiments when the resulting data will be used to choose a welfare-maximizing policy subject to constraints. A decision maker seeks to maximize Bayes expected welfare by choosing a policy whose effects depend on an unknown finite-dimensional parameter. The decision maker has access to a first wave of experimental data with a fixed design but may choose the design of a second wave that will be collected before choosing the policy. The resulting experimental design--policy choice problem is a very high-dimensional dynamic program that is generally intractable in finite samples. We propose a tractable approximation based on the limit experiment and show it is asymptotically optimal using a new asymptotic representation theorem for adaptive experiments with continuous treatments. We apply the method to a conditional cash transfer experiment and demonstrate the potential for large gains from tailoring the experiment to the policy choice.

econ.EM

Designing Spatial Treatments

Spatial treatments are interventions assigned to locations potentially distinct from those of the responding units. We study their optimal design under a general model in which a unit's response diminishes with distance to a treated site. Our estimand of interest is an ``uncontaminated'' effect equal to the average impact of a single intervention site over all hypothetical sites. We propose a novel design based on a Mat\'{e}rn point process which separates treatments by a distance of at least $r$. A larger choice of $r$ reduces bias by separating interventions but increases variance by reducing their numerosity. We choose $r$ to maximize the rate of convergence of a Horvitz-Thompson estimator and prove that this is minimax rate-optimal. We provide weak conditions under which the estimator is asymptotically normal and propose a variance estimator.

econ.EM