SearcharxivSearch

arXiv subjects

Pengyun Wang

Publications and source records attributed to Pengyun Wang.

18 recordsLinked to original sources

CurveShift: Is Agent Progress Scalar? Separating Level from Shape

Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score. These summaries capture the overall performance, but they do not test whether progress is distributed differently across task difficulty. We find that most of the apparent shift in gains toward harder tasks does not reflect a change in the shape of the difficulty-response curve. On METR time-horizon data, a single Rasch model with rising ability reproduces this pattern, so it is largely explained by ceiling effects rather than a qualitative change in capability. This echoes how the choice of metric can make claimed emergent abilities look like a property of the models themselves. We then identify a smaller hard-task effect that survives this control. Isolating it is difficult on agentic benchmarks, because newer models are usually run with newer agentic harnesses, so a gain on hard tasks cannot be assigned to the model or its scaffold. We break the confound with LiveCodeBench, a public competitive programming benchmark that runs no agentic scaffold while pairing dated models with an exogenous difficulty ordering. After accounting for the rise in overall ability, models released after September 2024 still gain on the hardest problems beyond what their easy and medium performance predicts, by about +0.40 logits under our most conservative assumption, raising the hard-problem solve rate from roughly 18% to 25%. The effect is led by the strongest reasoning models and holds for hard tasks that need only short reasoning, not autonomy over long horizons. We present this as a result specific to competitive programming, since our clean identification rests on a single coding benchmark. We release the LiveCodeBench Difficulty Panel (66 dated models x 1,055 problems) and our analysis code.

cs.CL

Regularity, Phase Transitions, and Uniform Inference for Proximal Counterfactual Quantile Processes

This paper develops semiparametric theory for counterfactual distribution, quantile, and lower-tail risk processes under unmeasured confounding using proximal negative-control proxies. Rather than treating each threshold as a separate proximal mean problem with outcome $\mathbf 1\{Y\le y\}$, we study the continuum of inverse problems indexed by $y$. For each treatment arm $a$, the counterfactual CDF $F_a(y)=P\{Y(a)\le y\}$ is represented by the primal bridge equation $T_a h_{a,y}=g_{a,y}$ and the linear functional $\ell(h)=E\{h(W,X)\}$. The dual bridge $q_a$ solves $T_a^*q_a=1$, equivalently $E[\mathbf 1(A=a)q_a(Z,X)-1\mid W,X]=0$. We show that this dual equation, together with the minimal residual-moment condition required for the influence function to lie in $L_2(P_0)$, is the exact regularity boundary in a threshold-saturated observed-data proximal bridge model: $F_a(y)$ is pathwise differentiable if and only if a regular square-integrable dual bridge exists. The canonical gradient is \[ h_{a,y}(W,X)-F_a(y)+\mathbf 1(A=a)q_a(Z,X)\{\mathbf 1(Y\le y)-h_{a,y}(W,X)\}. \] A singular-system characterization gives a Picard-type phase transition: root-$n$ regular estimation is possible exactly when $\sum_j\ell_{a,j}^2/s_{a,j}^2<\infty$ and the residual moment is finite. Outside this region, finite-dimensional efficiency bounds diverge under residual-noise nondegeneracy, and Gaussian inverse benchmarks yield slower minimax rates. We further establish efficient CDF-process inference, cross-fitted uniform doubly robust expansions, finite-rank weak-proxy rate conditions, density-free simultaneous quantile bands by inversion of CDF bands, and lower-tail CVaR inference via a shortfall representation. The estimators rely on closed-form linear algebra, convex Tikhonov regularization, and isotonic projection for shape enforcement.

stat.ME

Nested Sensitivity Envelopes for Transported Quantile Treatment Effects

We study target-population quantile treatment effects when a source study may have unmeasured treatment confounding and may not transport to a target population after conditioning on observed covariates. The observed data consist of a source sample with treatment, outcome and covariates, and a target sample with covariates only. We impose two marginal sensitivity restrictions: an odds-ratio bound \(\Gam\) for source treatment assignment and a conditional likelihood-ratio bound \(\Lam\) for source-to-target potential-outcome distribution shift. For each treatment arm and threshold \(y\), we derive a closed-form sharp target counterfactual CDF envelope. The envelope nests a source marginal-sensitivity map inside a target outcome-shift map, preserving two normalizations and generally improving on a single product likelihood-ratio relaxation. We prove process-level sharpness, so the envelopes are attainable as entire CDFs and can be inverted to obtain sharp target quantile bounds and sharp interval-hull QTE bounds. We then develop semiparametric theory for these nonsmooth bound processes. On regular index sets, we give the canonical gradient, including the source propensity contribution required in observational studies, and construct cross-fitted Neyman-orthogonal one-step estimators with uniform Gaussian approximation. On full index sets with active-set ties or mass points, we use Hadamard directional differentiability and subsampling-valid inference, with a primitive finite-support route for the required weak convergence. Finally, we invert simultaneous monotone CDF bands to obtain honest confidence sets for quantile and QTE interval-hull processes, and formulate the two-dimensional \((\Gam,\Lam)\) breakdown frontier as level-set inference for interval-hull non-refutation.

stat.ME

Efficient Transported Distributional and Quantile Treatment Effects with Surrogate-Assisted Missing Primary Outcomes

We study target-population distributional and quantile treatment effects when a source study observes treatment and post-treatment surrogates for all source units but observes a long-run primary outcome only for a validation subset, while the target population contributes only baseline covariates. The target estimands are transported counterfactual distribution functions $ψ_a(y)=P(Y^a\le y\mid R=0)$, their quantiles $q_a(τ)$, and the quantile treatment effect $Δ(τ)=q_1(τ)-q_0(τ)$. The surrogate is not treated as a replacement endpoint and no Prentice-type surrogacy condition is imposed. Instead, the surrogate is used only to improve efficiency under missing-at-random primary-outcome sampling. We derive the nonparametric efficient influence function, which has three orthogonal components corresponding to target covariate sampling, the source surrogate process, and missing primary outcomes. This yields a closed-form cross-fitted one-step estimator after nuisance estimation. We establish identification, the canonical gradient, exact drift identities, ratio-level robustness, pointwise and uniform asymptotic linearity for transported CDFs, Bahadur representations for quantiles under explicit local inverse-map conditions, high-level multiplier-bootstrap simultaneous bands under explicit estimated-process and density conditions, and quantile-specific efficiency gains from observing surrogates. We also give lower-level nuisance-rate verification for a deliberately restricted class of analyzable bounded finite-dimensional or finite-rank implementations based on sieve ridge regression, ridge logistic regression, calibrated density-ratio estimation, finite-rank kernel ridge regression, and isotonic projection under explicit grid, eigenvalue, source, and entropy conditions.

stat.ME

A Comprehensive Graph Pooling Benchmark: Effectiveness, Robustness and Generalizability

Graph pooling has gained attention for its ability to obtain effective node and graph representations for various downstream tasks. Despite the recent surge in graph pooling approaches, there is a lack of standardized experimental settings and fair benchmarks to evaluate their performance. To address this issue, we have constructed a comprehensive benchmark that includes 17 graph pooling methods and 28 different graph datasets. This benchmark systematically assesses the performance of graph pooling methods in three dimensions, i.e., effectiveness, robustness, and generalizability. We first evaluate the performance of these graph pooling approaches across different tasks including graph classification, graph regression and node classification. Then, we investigate their performance under potential noise attacks and out-of-distribution shifts in real-world scenarios. We also involve detailed efficiency analysis, backbone analysis, parameter analysis and visualization to provide more evidence. Extensive experiments validate the strong capability and applicability of graph pooling approaches in various scenarios, which can provide valuable insights and guidance for deep geometric learning research. The source code of our benchmark is available at https://github.com/goose315/Graph_Pooling_Benchmark.

cs.LG

Reconciling Overt Bias and Hidden Bias in Sensitivity Analysis for Matched Observational Studies

Matching is one of the most widely used causal inference designs in observational studies, but post-matching confounding bias remains a critical concern. This bias includes overt bias from inexact matching on measured confounders and hidden bias from unmeasured confounders. Researchers routinely apply the famous Rosenbaum-type sensitivity analysis after matching to assess the impact of these biases on causal conclusions. In this work, we show that this approach is often conservative and may overstate sensitivity to confounding bias because the classical solution to the Rosenbaum sensitivity model may allocate hypothetical hidden bias in ways that contradict the overt bias observed in the matched dataset. To address this problem, we propose a new approach to Rosenbaum-type sensitivity analysis by ensuring compatibility between hidden and overt biases. Our approach does not need to add any additional assumptions (beyond mild regularity conditions) to Rosenbaum-type sensitivity analysis, and can produce uniformly more informative sensitivity analysis results than the conventional Rosenbaum-type sensitivity analysis. Computationally, our approach can be solved efficiently via iterative convex programming. Extensive simulations and a real data application demonstrate substantial gains in statistical power of sensitivity analysis. Importantly, our approach can also be applied to many other sensitivity analysis frameworks.

stat.ME

EgoLife: Towards Egocentric Life Assistant

We introduce EgoLife, a project to develop an egocentric life assistant that accompanies and enhances personal efficiency through AI-powered wearable glasses. To lay the foundation for this assistant, we conducted a comprehensive data collection study where six participants lived together for one week, continuously recording their daily activities - including discussions, shopping, cooking, socializing, and entertainment - using AI glasses for multimodal egocentric video capture, along with synchronized third-person-view video references. This effort resulted in the EgoLife Dataset, a comprehensive 300-hour egocentric, interpersonal, multiview, and multimodal daily life dataset with intensive annotation. Leveraging this dataset, we introduce EgoLifeQA, a suite of long-context, life-oriented question-answering tasks designed to provide meaningful assistance in daily life by addressing practical questions such as recalling past relevant events, monitoring health habits, and offering personalized recommendations. To address the key technical challenges of (1) developing robust visual-audio models for egocentric data, (2) enabling identity recognition, and (3) facilitating long-context question answering over extensive temporal information, we introduce EgoButler, an integrated system comprising EgoGPT and EgoRAG. EgoGPT is an omni-modal model trained on egocentric datasets, achieving state-of-the-art performance on egocentric video understanding. EgoRAG is a retrieval-based component that supports answering ultra-long-context questions. Our experimental studies verify their working mechanisms and reveal critical factors and bottlenecks, guiding future improvements. By releasing our datasets, models, and benchmarks, we aim to stimulate further research in egocentric AI assistants.

cs.CV

DELTA: Dual Consistency Delving with Topological Uncertainty for Active Graph Domain Adaptation

Graph domain adaptation has recently enabled knowledge transfer across different graphs. However, without the semantic information on target graphs, the performance on target graphs is still far from satisfactory. To address the issue, we study the problem of active graph domain adaptation, which selects a small quantitative of informative nodes on the target graph for extra annotation. This problem is highly challenging due to the complicated topological relationships and the distribution discrepancy across graphs. In this paper, we propose a novel approach named Dual Consistency Delving with Topological Uncertainty (DELTA) for active graph domain adaptation. Our DELTA consists of an edge-oriented graph subnetwork and a path-oriented graph subnetwork, which can explore topological semantics from complementary perspectives. In particular, our edge-oriented graph subnetwork utilizes the message passing mechanism to learn neighborhood information, while our path-oriented graph subnetwork explores high-order relationships from sub-structures. To jointly learn from two subnetworks, we roughly select informative candidate nodes with the consideration of consistency across two subnetworks. Then, we aggregate local semantics from its K-hop subgraph based on node degrees for topological uncertainty estimation. To overcome potential distribution shifts, we compare target nodes and their corresponding source nodes for discrepancy scores as an additional component for fine selection. Extensive experiments on benchmark datasets demonstrate that DELTA outperforms various state-of-the-art approaches. The code implementation of DELTA is available at https://github.com/goose315/DELTA.

cs.LG

OpenOOD v1.5: Enhanced Benchmark for Out-of-Distribution Detection

Out-of-Distribution (OOD) detection is critical for the reliable operation of open-world intelligent systems. Despite the emergence of an increasing number of OOD detection methods, the evaluation inconsistencies present challenges for tracking the progress in this field. OpenOOD v1 initiated the unification of the OOD detection evaluation but faced limitations in scalability and scope. In response, this paper presents OpenOOD v1.5, a significant improvement from its predecessor that ensures accurate and standardized evaluation of OOD detection methodologies at large scale. Notably, OpenOOD v1.5 extends its evaluation capabilities to large-scale data sets (ImageNet) and foundation models (e.g., CLIP and DINOv2), and expands its scope to investigate full-spectrum OOD detection which considers semantic and covariate distribution shifts at the same time. This work also contributes in-depth analysis and insights derived from comprehensive experimental results, thereby enriching the knowledge pool of OOD detection methodologies. With these enhancements, OpenOOD v1.5 aims to drive advancements and offer a more robust and comprehensive evaluation benchmark for OOD detection research.

cs.LG

Optically-Trapped Nanodiamond-Relaxometry Detection of Nanomolar Paramagnetic Spins in Aqueous Environments

Probing electrical and magnetic properties in aqueous environments remains a frontier challenge in nanoscale sensing. Our inability to do so with quantitative accuracy imposes severe limitations, for example, on our understanding of the ionic environments in a diverse array of systems, ranging from novel materials to the living cell. The Nitrogen-Vacancy (NV) center in fluorescent nanodiamonds (FNDs) has emerged as a good candidate to sense temperature, pH, and the concentration of paramagnetic species at the nanoscale, but comes with several hurdles such as particle-to-particle variation which render calibrated measurements difficult, and the challenge to tightly confine and precisely position sensors in aqueous environment. To address this, we demonstrate relaxometry with NV centers within optically-trapped FNDs. In a proof of principle experiment, we show that optically-trapped FNDs enable highly reproducible nanomolar sensitivity to the paramagnetic ion, (\mathrm{Gd}^{3+}). We capture the three distinct phases of our experimental data by devising a model analogous to nanoscale Langmuir adsorption combined with spin coherence dynamics. Our work provides a basis for routes to sense free paramagnetic ions and molecules in biologically relevant conditions.

quant-ph

Re-evaluating the impact of reduced malaria prevalence on birthweight in sub-Saharan Africa: A pair-of-pairs study via two-stage bipartite and non-bipartite matching

According to the WHO, in 2021, about 32% of pregnant women in sub-Saharan Africa were infected with malaria during pregnancy. Malaria infection during pregnancy can cause various adverse birth outcomes such as low birthweight. Over the past two decades, while some sub-Saharan African areas have experienced a large reduction in malaria prevalence due to improved malaria control and treatments, others have observed little change. Individual-level interventional studies have shown that preventing malaria infection during pregnancy can improve birth outcomes such as birthweight; however, it is still unclear whether natural reductions in malaria prevalence may help improve community-level birth outcomes. We conduct an observational study using 203,141 children's records in 18 sub-Saharan African countries from 2000 to 2018. Using heterogeneity of changes in malaria prevalence, we propose and apply a novel pair-of-pairs design via two-stage bipartite and non-bipartite matching to conduct a difference-in-differences study with a continuous measure of malaria prevalence, namely the Plasmodium falciparum parasite rate among children aged 2 to 10 ($\text{PfPR}_{2-10}$). The proposed novel statistical methodology allows us to apply difference-in-differences without dichotomizing $\text{PfPR}_{2-10}$, which can substantially increase the effective sample size, improve covariate balance, and facilitate the dose-response relationship during analysis. Our outcome analysis finds that among the pairs of clusters we study, the largest reduction in $\text{PfPR}_{2-10}$ over early and late years is estimated to increase the average birthweight by 98.899 grams (95% CI: $[39.002, 158.796]$), which is associated with reduced risks of several adverse birth or life-course outcomes. The proposed novel statistical methodology can be replicated in many other disease areas.

stat.AP

Generative Oversampling for Imbalanced Data via Majority-Guided VAE

Learning with imbalanced data is a challenging problem in deep learning. Over-sampling is a widely used technique to re-balance the sampling distribution of training data. However, most existing over-sampling methods only use intra-class information of minority classes to augment the data but ignore the inter-class relationships with the majority ones, which is prone to overfitting, especially when the imbalance ratio is large. To address this issue, we propose a novel over-sampling model, called Majority-Guided VAE~(MGVAE), which generates new minority samples under the guidance of a majority-based prior. In this way, the newly generated minority samples can inherit the diversity and richness of the majority ones, thus mitigating overfitting in downstream tasks. Furthermore, to prevent model collapse under limited data, we first pre-train MGVAE on sufficient majority samples and then fine-tune based on minority samples with Elastic Weight Consolidation(EWC) regularization. Experimental results on benchmark image datasets and real-world tabular data show that MGVAE achieves competitive improvements over other over-sampling methods in downstream classification tasks, demonstrating the effectiveness of our method.

cs.LG

Ti-MAE: Self-Supervised Masked Time Series Autoencoders

Multivariate Time Series forecasting has been an increasingly popular topic in various applications and scenarios. Recently, contrastive learning and Transformer-based models have achieved good performance in many long-term series forecasting tasks. However, there are still several issues in existing methods. First, the training paradigm of contrastive learning and downstream prediction tasks are inconsistent, leading to inaccurate prediction results. Second, existing Transformer-based models which resort to similar patterns in historical time series data for predicting future values generally induce severe distribution shift problems, and do not fully leverage the sequence information compared to self-supervised methods. To address these issues, we propose a novel framework named Ti-MAE, in which the input time series are assumed to follow an integrate distribution. In detail, Ti-MAE randomly masks out embedded time series data and learns an autoencoder to reconstruct them at the point-level. Ti-MAE adopts mask modeling (rather than contrastive learning) as the auxiliary task and bridges the connection between existing representation learning and generative Transformer-based methods, reducing the difference between upstream and downstream forecasting tasks while maintaining the utilization of original time series data. Experiments on several public real-world datasets demonstrate that our framework of masked autoencoding could learn strong representations directly from the raw data, yielding better performance in time series forecasting and classification tasks.

cs.LG

OpenOOD: Benchmarking Generalized Out-of-Distribution Detection

Out-of-distribution (OOD) detection is vital to safety-critical machine learning applications and has thus been extensively studied, with a plethora of methods developed in the literature. However, the field currently lacks a unified, strictly formulated, and comprehensive benchmark, which often results in unfair comparisons and inconclusive results. From the problem setting perspective, OOD detection is closely related to neighboring fields including anomaly detection (AD), open set recognition (OSR), and model uncertainty, since methods developed for one domain are often applicable to each other. To help the community to improve the evaluation and advance, we build a unified, well-structured codebase called OpenOOD, which implements over 30 methods developed in relevant fields and provides a comprehensive benchmark under the recently proposed generalized OOD detection framework. With a comprehensive comparison of these methods, we are gratified that the field has progressed significantly over the past few years, where both preprocessing methods and the orthogonal post-hoc methods show strong potential.

cs.CV

Multi-relation Message Passing for Multi-label Text Classification

A well-known challenge associated with the multi-label classification problem is modelling dependencies between labels. Most attempts at modelling label dependencies focus on co-occurrences, ignoring the valuable information that can be extracted by detecting label subsets that rarely occur together. For example, consider customer product reviews; a product probably would not simultaneously be tagged by both "recommended" (i.e., reviewer is happy and recommends the product) and "urgent" (i.e., the review suggests immediate action to remedy an unsatisfactory experience). Aside from the consideration of positive and negative dependencies, the direction of a relationship should also be considered. For a multi-label image classification problem, the "ship" and "sea" labels have an obvious dependency, but the presence of the former implies the latter much more strongly than the other way around. These examples motivate the modelling of multiple types of bi-directional relationships between labels. In this paper, we propose a novel method, entitled Multi-relation Message Passing (MrMP), for the multi-label classification problem. Experiments on benchmark multi-label text classification datasets show that the MrMP module yields similar or superior performance compared to state-of-the-art methods. The approach imposes only minor additional computational and memory overheads.

cs.LG

Label-Aware Distribution Calibration for Long-tailed Classification

Real-world data usually present long-tailed distributions. Training on imbalanced data tends to render neural networks perform well on head classes while much worse on tail classes. The severe sparseness of training instances for the tail classes is the main challenge, which results in biased distribution estimation during training. Plenty of efforts have been devoted to ameliorating the challenge, including data re-sampling and synthesizing new training instances for tail classes. However, no prior research has exploited the transferable knowledge from head classes to tail classes for calibrating the distribution of tail classes. In this paper, we suppose that tail classes can be enriched by similar head classes and propose a novel distribution calibration approach named as label-Aware Distribution Calibration LADC. LADC transfers the statistics from relevant head classes to infer the distribution of tail classes. Sampling from calibrated distribution further facilitates re-balancing the classifier. Experiments on both image and text long-tailed datasets demonstrate that LADC significantly outperforms existing methods.The visualization also shows that LADC provides a more accurate distribution estimation.

cs.LG

Mask-GVAE: Blind Denoising Graphs via Partition

We present Mask-GVAE, a variational generative model for blind denoising large discrete graphs, in which "blind denoising" means we don't require any supervision from clean graphs. We focus on recovering graph structures via deleting irrelevant edges and adding missing edges, which has many applications in real-world scenarios, for example, enhancing the quality of connections in a co-authorship network. Mask-GVAE makes use of the robustness in low eigenvectors of graph Laplacian against random noise and decomposes the input graph into several stable clusters. It then harnesses the huge computations by decoding probabilistic smoothed subgraphs in a variational manner. On a wide variety of benchmarks, Mask-GVAE outperforms competing approaches by a significant margin on PSNR and WL similarity.

cs.LG

Predicting Path Failure In Time-Evolving Graphs

In this paper we use a time-evolving graph which consists of a sequence of graph snapshots over time to model many real-world networks. We study the path classification problem in a time-evolving graph, which has many applications in real-world scenarios, for example, predicting path failure in a telecommunication network and predicting path congestion in a traffic network in the near future. In order to capture the temporal dependency and graph structure dynamics, we design a novel deep neural network named Long Short-Term Memory R-GCN (LRGCN). LRGCN considers temporal dependency between time-adjacent graph snapshots as a special relation with memory, and uses relational GCN to jointly process both intra-time and inter-time relations. We also propose a new path representation method named self-attentive path embedding (SAPE), to embed paths of arbitrary length into fixed-length vectors. Through experiments on a real-world telecommunication network and a traffic network in California, we demonstrate the superiority of LRGCN to other competing methods in path failure prediction, and prove the effectiveness of SAPE on path representation.

cs.LG