SearcharxivSearch

arXiv subjects

Lingfeng Lyu

Publications and source records attributed to Lingfeng Lyu.

3 recordsLinked to original sources

Conditional Evaluation of Language Models with Cheap Auxiliary Signals

Aggregate accuracy hides where models succeed and fail. Estimating conditional performance profiles from gold labels alone is expensive, while cheap auxiliary signals such as LLM-judge scores, pairwise comparisons, confidence scores, and judge-disagreement features can be collected for every benchmark item but are often biased or miscalibrated. We propose LACE (Local Augmented Control-Variate Evaluation), a semi-supervised estimator for conditional LLM evaluation. The key step is local centering: after subtracting the conditional mean of a cheap signal within the target profile region, any linear augmentation has zero conditional mean and therefore cannot change the estimand. The augmentation coefficient is used only for efficiency, and a local ridge control variate combines a gold-label residual mean from the labeled subset with a cheap-signal mean from the full item pool. We prove calibration-free identification, unbiasedness for grouped profiles, local oracle optimality within centered linear augmentations, and first-order adaptivity to the estimated coefficient. The resulting gain formula is governed by a population local $R^2$, which characterizes how the efficiency attainable from the cheap signals varies across profile values. We also derive corresponding estimators for direct paired model gaps and deployment-weighted scores. We empirically evaluate the primary performance-profile estimator on MATH-500, ScienceQA, MMLU, WinoGrande, HellaSwag, TruthfulQA, GSM8K, and ARC.

cs.LG

Transfer Learning for Moderate-Dimensional Ridge-Regularized Robust Linear Regression

This paper studies transfer learning for ridge-regularized robust linear regression in the moderate-dimensional regime, where the number of predictors is of the same order as the sample size and the regression coefficients are not assumed to be sparse. We propose Trans-RR, which combines a robust ridge estimator from a source study with a robust ridge correction based on the target study. Under mild assumptions, we characterize the asymptotic estimation error of the proposed estimator and show that leveraging source data can substantially improve estimation accuracy relative to the traditional single-study ridge-regularized robust estimator. Simulation results and a real-data analysis support the theory and illustrate both positive and negative transfer as the discrepancy between the source and target studies varies.

stat.ME

Preference-based Centrality and Ranking in General Metric Spaces

Ranking or assessing centrality in multivariate and non-Euclidean data is difficult because there is no canonical order and many depth notions become computationally fragile in high-dimensional or structured settings. We introduce a preference-based notion of centrality defined through population proximity comparisons with respect to a random reference draw, yielding a metric-intrinsic statistical functional that is well-defined on general metric spaces. Because the induced pairwise preferences may be non-transitive, we map them to a coherent one-dimensional score via a Bradley--Terry--Luce cross-entropy projection, viewed as a calibrated aggregation device rather than a correctly specified model. We develop two finite-sample estimators a convex M-estimator and a fast spectral estimator based on a comparison operator, and establish identifiability and consistency under mild conditions. Simulations and real-data examples, including high-dimensional and functional observations, illustrate that the proposed scores provide stable, interpretable rankings aligned with the underlying preference centrality.

stat.ME