SearcharxivSearch

arXiv subjects

Xiaoya Xu

Publications and source records attributed to Xiaoya Xu.

2 recordsLinked to original sources

Uniform-Design Subsampling for Compute-Budgeted Double Machine Learning

Double machine learning (DML) combines orthogonal scores with flexible nuisance estimation, but repeated cross-fitting and repeated analysis can make full-data workflows expensive. When a fixed computational budget requires a working sample of size $r\ll n$, simple random subsampling may cover the covariate space poorly and produce unstable treated--control composition. We propose Uniform Design Double Machine Learning (UD-DML), which maps a common low-discrepancy skeleton through empirical marginal quantiles and matches each anchor to one treated and one control observation, without replacement within each treatment arm. We establish finite-sample integration and balance bounds, an exact selected-target decomposition, and a selection-aware leading-score central limit theorem on the root-$r$ scale. Under explicit moment, weighted-calibration, variance-stabilisation, and target-transport conditions, a residual-based variance estimator consistently studentises the original UD-DML estimator. Full-scale simulations show that the common-skeleton construction provides its clearest improvements over uniform subsampling under non-uniform covariate geometry and limited treated--control overlap, while maintaining empirical coverage near the nominal level. These results provide a principled working-sample design for repeated causal learning under an explicit computational budget.

stat.ME

Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness

Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlooks locale and cultural considerations. We therefore advocate for source-contrastive evaluation and instantiate it with Cultivar, a localised subset of FLORES, which enables locale-specific translation evaluation. When paired with unlocalised counterparts, performance discrepancy allows the probing of data contamination and localisation robustness. We benchmark 32 open-weight models and find that MT-specialised models are less robust, a few models potentially overfit FLORES, and models tend to translate US content better than that of other locales, regardless of language.

cs.CL