SearcharxivSearch

arXiv subjects

Wang Miao

Publications and source records attributed to Wang Miao.

At least 19 recordsLinked to original sources

Identifying the desert decision rule to assess and achieve fairness

We study fairness in decision-making when the data may encode systematic bias. Existing approaches typically impose fairness constraints while predicting the observed decision, which may itself be unfair. We propose a novel framework for characterising and addressing fairness issues by introducing the notion of desert decision, a latent variable representing the decision an individual rightfully deserves based on their actions, efforts, or abilities. This formulation shifts the prediction target from the potentially biased observed decision to the desert decision. We advocate achieving fair decision-making by predicting the desert decision and assessing unfairness by the discrepancy between desert and observed decisions. We establish nonparametric identification results under causally interpretable assumptions on the fairness of the desert decision and the unfairness mechanism of the observed decision. For estimation, we develop a sieve maximum likelihood estimator for the desert decision rule and an influence-function-based estimator for the degree of unfairness. Sensitivity analysis procedures are further proposed to assess robustness to violations of identifying assumptions. Our framework connects fairness with measurement error models, aligning predictive accuracy with fairness relative to an appropriate target, and providing a structural approach to modelling the unfairness mechanism.

stat.ME

Adaptive RAN Slicing Control via Reward-Free Self-Finetuning Agents

The integration of Generative AI models into AI-native network systems offers a transformative path toward achieving autonomous and adaptive control. However, the application of such models to continuous control tasks is impeded by intrinsic architectural limitations, including finite context windows, the lack of explicit reward signals, and the degradation of the long context. This paper posits that the key to unlocking robust continuous control is enabling agents to internalize experience by distilling it into their parameters, rather than relying on prompt-based memory. To this end, we propose a novel self-finetuning framework that enables agentic systems to learn continuously through direct interaction with the environment, bypassing the need for handcrafted rewards. Our framework implements a bi-perspective reflection mechanism that generates autonomous linguistic feedback to construct preference datasets from interaction history. A subsequent preference-based fine-tuning process distills long-horizon experiences into the model's parameters. We evaluate our approach on a dynamic Radio Access Network (RAN) slicing task, a challenging multi-objective control problem that requires the resolution of acute trade-offs between spectrum efficiency, service quality, and reconfiguration stability under volatile network conditions. Experimental results show that our framework outperforms standard Reinforcement Learning (RL) baselines and existing Large Language Model (LLM)-based agents in sample efficiency, stability, and multi-metric optimization. These findings demonstrate the potential of self-improving generative agents for continuous control tasks, paving the way for future AI-native network infrastructure.

cs.AI

Routing-Led Evolutionary Algorithm for Large-Scale Multi-Objective VNF Placement Problems

Modern data centers contain thousands of servers making them major consumers of electricity. To minimize their environmental impact, it is critical that we use their resources efficiently. In this paper we study how to discover the optimal placement of virtual network functions in large scale data centers. We propose a novel parallel metaheuristic, fast heuristic objective functions of the QoS and new memory efficient data structures for large networks. We further identify a simple, fast heuristic that can produce competitive solutions to very large problem instances. Using these new concepts, we are able to find high quality solutions for data centres with up to 64,000 servers.

cs.NE

Leveraging specificity for causal inference in observational studies

Hill's specificity criterion has been highly influential in biomedical and epidemiological research. However, it remains controversial and its application often relies on subjective and qualitative analysis without a comprehensive and rigorous causal theory. Focusing on unmeasured confounding adjustment with multiple treatments and multiple outcomes, this paper develops a formal and quantitative framework for leveraging specificity for causal inference in observational studies. The proposed framework introduces a causal specificity assumption, a quantitative measure of specificity, a hypothesis testing procedure, and identification and estimation strategies. Identification under a nonparametric outcome model is established. The causal specificity assumption concerns only the breadth of causal associations, in contrast to Hill's specificity that concerns observed associations and to existing confounding adjustment methods that rely on auxiliary variables (e.g., instrumental variables, negative controls) or independence structures (e.g., factor models). A sensitivity analysis procedure is proposed to assess robustness of the test against violations of this assumption. This framework is particularly suited to exposure- and outcome-wide studies, where joint causal discovery across multiple treatments and multiple outcomes is of interest. It also offers a potential tool for addressing other sources of bias, such as invalid instruments in Mendelian randomization and selection bias in missing data analysis.

stat.ME

Causal inference with dyadic data in randomized experiments

Estimating treatment effects in networked settings is a central challenge in online controlled experiments, particularly on social media platforms. We investigate a scenario where the unit-level outcome of interest comprises a series of dyadic outcomes that record pairwise interactions between units, spanning from point-to-point messaging at the microscale to bilateral trade flows at the macroscale. Because the response is defined at the dyadic level, the treatment assigned to one unit can affect the outcomes of all dyads that involve it, inducing a form of network interference. We propose a design-based causal inference framework for randomized experiments with dyadic outcomes. Within this framework, we propose estimators of the global average treatment effect under Bernoulli, complete, and cluster randomization, derive the convergence rates, and establish a central limit theorem for Bernoulli randomization. We further construct a class of variance estimators that are asymptotically conservative under transparent degree conditions. Numerical studies show that the proposed estimators can reduce bias and mean squared error relative to estimators based on unit-level outcomes in a range of finite-sample settings. We illustrate the methods using two large-scale experiments on WeChat, evaluating the impact of a recommendation algorithm and a calling feature.

stat.ME

Correcting nonignorable nonresponse bias in turnout estimation using callback data

Overestimation of turnout has long been an issue in election surveys, with nonresponse bias or voter overrepresentation identified as major sources of bias. However, adjusting for nonignorable nonresponse bias is substantially challenging. Based on the ANES Non-Response Follow-Up study concerning the 2020 U.S. presidential election, we investigate the role of callback data, that is, records of contact attempts in the survey course, in adjusting for nonresponse bias in the estimation of turnout. We propose a stableness of resistance assumption to account for nonignorable missingness in the outcome, which states that the impact of the missing outcome on the response propensity is stable in the first two call attempts. Under this assumption and by integrating with covariate information from the census data, we establish identifiability and develop estimation methods for turnout. Our methods produce estimates very close to the official turnout and successfully capture the trend of declining willingness to vote as response reluctance increases. This work highlights the importance of adjusting for nonignorable nonresponse bias and demonstrates the potential of widely available callback data for political surveys.

stat.ME

A generalized tetrad constraint for testing conditional independence given a latent variable

The tetrad constraint is widely used to test whether four observed variables are conditionally independent given a latent variable, based on the fact that if four observed variables following a linear model are mutually independent after conditioning on an unobserved variable, then products of covariances of any two different pairs of these four variables are equal. It is an important tool for discovering a latent common cause or distinguishing between alternative linear causal structures. However, the classical tetrad constraint fails in nonlinear models because the covariance of observed variables cannot capture nonlinear association. In this paper, we propose a generalized tetrad constraint, which establishes a testable implication for conditional independence given a latent variable in nonlinear and nonparametric models. In linear models, this constraint implies the classical tetrad constraint; in nonlinear models, it remains a necessary condition for conditional independence but the classical tetrad constraint no longer is. Based on this constraint, we further propose a formal test, which can control type I error and has power approaching unity under certain conditions. We illustrate the proposed approach via simulations and two real data applications on mental ability tests and on moral attitudes towards dishonesty.

stat.ME

Statistical inference for the probability of necessity for causal attribution

To answer questions of "causes of effects", the probability of necessity was previously introduced for assessing whether an observed outcome was caused by an earlier treatment. However, statistical inference for the probability of necessity is understudied due to several difficulties, which hinder its application in practice. The evaluation of the probability of necessity involves the joint distribution of potential outcomes, and thus it is generally not point identified and one can at best obtain lower and upper bounds even in randomized experiments, unless fairly stringent monotonicity assumptions on potential outcomes are made. Moreover, these bounds are non-smooth functionals of the observed data distribution and standard estimation and inference methods cannot be directly applied. In this paper, we investigate the statistical inference for the probability of necessity in general situations where it may not be point identified. We introduce a mild margin condition to tackle the non-smoothness, under which the bounds become pathwise differentiable. We establish the semiparametric efficiency theory and propose novel asymptotically efficient estimators of the lower and upper bounds, and further construct confidence intervals for the probability of necessity based on the proposed bound estimators. The resultant confidence intervals can effectively utilize the observed covariates to reduce lengths. The proposed approach has potential application in biomedical, epidemiological, and legal studies where understanding causal attribution beyond traditional causal effects is essential.

stat.ME

A robust regression approach to synthetic control with interference

Synthetic control methods are widely used for policy evaluation, but most existing approaches rule out interference among units, compromising validity when such effects are present. We develop a framework that accommodates contaminated donor pools and unknown interference patterns through two stages: factor-model adjustment for unobserved confounding, followed by robust regression in which direct and interference effects appear as a sparse outlier component. We study two asymptotic regimes. When the number of units is fixed and at least half are unaffected by interference, high-breakdown robust regression yields consistent identification of valid controls and asymptotically normal inference. When the number of units diverges, we allow for sparse large and dense weak interference, with robust M-estimation remaining valid even when the post-intervention period is short. Unlike existing approaches requiring prespecification of valid controls or parametric modeling of interference, our framework relies only on coarse sparsity information and enables formal inference on both direct and interference effects. We assess the proposed methods through simulations and two empirical applications. An analysis of the US embassy relocation to Jerusalem reveals significant interference effects on conflict outcomes in Jordan, and an analysis of Beijing's air pollution policy uncovers spatial interference patterns consistent with prevailing wind directions.

stat.ME

A confounding bridge approach for double negative control inference on causal effects

Unmeasured confounding is a key challenge for causal inference. In this paper, we establish a framework for unmeasured confounding adjustment with negative control variables. A negative control outcome is associated with the confounder but not causally affected by the exposure in view, and a negative control exposure is correlated with the primary exposure or the confounder but does not causally affect the outcome of interest. We introduce an outcome confounding bridge function that depicts the relationship between the confounding effects on the primary outcome and the negative control outcome, and we incorporate a negative control exposure to identify the bridge function and the average causal effect. We also consider the extension to the positive control setting by allowing for nonzero causal effect of the primary exposure on the control outcome. We illustrate our approach with simulations and apply it to a study about the short-term effect of air pollution on mortality. Although a standard analysis shows a significant acute effect of PM2.5 on mortality, our analysis indicates that this effect may be confounded, and after double negative control adjustment, the effect is attenuated toward zero.

stat.ME

On Doubly Robust Estimation with Nonignorable Missing Data Using Instrumental Variables

Suppose we are interested in the mean of an outcome that is subject to nonignorable nonresponse. This paper develops new semiparametric estimation methods with instrumental variables which affect nonresponse, but not the outcome. The proposed estimators remain consistent and asymptotically normal even under partial model misspecifications for two variation independent nuisance components. We evaluate the performance of the proposed estimators via a simulation study, and apply them in adjusting for missing data induced by HIV testing refusal in the evaluation of HIV seroprevalence in Mochudi, Botswana, using interviewer experience as an instrumental variable.

stat.ME

Causal Effect Identification and Inference with Endogenous Exposures and a Light-tailed Error

Endogeneity poses significant challenges in causal inference across various research domains. This paper proposes a novel approach to identify and estimate causal effects in the presence of endogeneity. We consider a structural equation with endogenous exposures and an additive error term. Assuming the light-tailedness of the error term, we show that the causal effect can be identified by contrasting extreme conditional quantiles of the outcome given the exposures. Unlike many existing results, our identification approach does not rely on additional parametric assumptions or auxiliary variables. Building on the identification result, we develop a new method that estimates the causal effect using extreme quantile regression. We establish the consistency of the proposed extreme-based estimator under a general additive structural equation and demonstrate its asymptotic normality in the linear model setting. Simulations and data analysis of an automobile sale dataset show the effectiveness of our method in handling endogeneity.

stat.ME

SpreadFGL: Edge-Client Collaborative Federated Graph Learning with Adaptive Neighbor Generation

Federated Graph Learning (FGL) has garnered widespread attention by enabling collaborative training on multiple clients for semi-supervised classification tasks. However, most existing FGL studies do not well consider the missing inter-client topology information in real-world scenarios, causing insufficient feature aggregation of multi-hop neighbor clients during model training. Moreover, the classic FGL commonly adopts the FedAvg but neglects the high training costs when the number of clients expands, resulting in the overload of a single edge server. To address these important challenges, we propose a novel FGL framework, named SpreadFGL, to promote the information flow in edge-client collaboration and extract more generalized potential relationships between clients. In SpreadFGL, an adaptive graph imputation generator incorporated with a versatile assessor is first designed to exploit the potential links between subgraphs, without sharing raw data. Next, a new negative sampling mechanism is developed to make SpreadFGL concentrate on more refined information in downstream tasks. To facilitate load balancing at the edge layer, SpreadFGL follows a distributed training manner that enables fast model convergence. Using real-world testbed and benchmark graph datasets, extensive experiments demonstrate the effectiveness of the proposed SpreadFGL. The results show that SpreadFGL achieves higher accuracy and faster convergence against state-of-the-art algorithms.

cs.LG

Automating the Selection of Proxy Variables of Unmeasured Confounders

Recently, interest has grown in the use of proxy variables of unobserved confounding for inferring the causal effect in the presence of unmeasured confounders from observational data. One difficulty inhibiting the practical use is finding valid proxy variables of unobserved confounding to a target causal effect of interest. These proxy variables are typically justified by background knowledge. In this paper, we investigate the estimation of causal effects among multiple treatments and a single outcome, all of which are affected by unmeasured confounders, within a linear causal model, without prior knowledge of the validity of proxy variables. To be more specific, we first extend the existing proxy variable estimator, originally addressing a single unmeasured confounder, to accommodate scenarios where multiple unmeasured confounders exist between the treatments and the outcome. Subsequently, we present two different sets of precise identifiability conditions for selecting valid proxy variables of unmeasured confounders, based on the second-order statistics and higher-order statistics of the data, respectively. Moreover, we propose two data-driven methods for the selection of proxy variables and for the unbiased estimation of causal effects. Theoretical analysis demonstrates the correctness of our proposed algorithms. Experimental results on both synthetic and real-world data show the effectiveness of the proposed approach.

cs.LG

Doubly Robust Proximal Synthetic Controls

To infer the treatment effect for a single treated unit using panel data, synthetic control methods construct a linear combination of control units' outcomes that mimics the treated unit's pre-treatment outcome trajectory. This linear combination is subsequently used to impute the counterfactual outcomes of the treated unit had it not been treated in the post-treatment period, and used to estimate the treatment effect. Existing synthetic control methods rely on correctly modeling certain aspects of the counterfactual outcome generating mechanism and may require near-perfect matching of the pre-treatment trajectory. Inspired by proximal causal inference, we obtain two novel nonparametric identifying formulas for the average treatment effect for the treated unit: one is based on weighting, and the other combines models for the counterfactual outcome and the weighting function. We introduce the concept of covariate shift to synthetic controls to obtain these identification results conditional on the treatment assignment. We also develop two treatment effect estimators based on these two formulas and the generalized method of moments. One new estimator is doubly robust: it is consistent and asymptotically normal if at least one of the outcome and weighting models is correctly specified. We demonstrate the performance of the methods via simulations and apply them to evaluate the effectiveness of a Pneumococcal conjugate vaccine on the risk of all-cause pneumonia in Brazil.

stat.ME

A maximin optimal approach for sampling designs in two-phase studies

Data collection costs can vary widely across variables in data science tasks. Two-phase designs can be employed to save data collection costs. This paper considers the two-phase studies where inexpensive variables are collected for all subjects in the first phase, and expensive variables are measured for a subsample of subjects in the second phase based on a predetermined sampling rule. The estimation efficiency under two-phase designs relies heavily on the sampling rule. Existing literature primarily focuses on designing sampling rules for estimating a scalar parameter in some parametric models or specific estimating problems. However, real-world scenarios are usually model-unknown and involve two-phase designs for model-free estimation of a scalar or multi-dimensional parameter. This paper proposes a maximin criterion to design an optimal sampling rule based on semiparametric efficiency bounds. The proposed method is model-free and applicable to general estimating problems. The resulting sampling rule can minimize the semiparametric efficiency bound when the parameter is scalar and improve the bound for every component when the parameter is multi-dimensional. Simulation studies demonstrate that the proposed designs reduce the variance of the resulting estimator in various settings. The implementation of the proposed design is illustrated in a real data analysis.

stat.ME

NAS-ASDet: An Adaptive Design Method for Surface Defect Detection Network using Neural Architecture Search

Deep convolutional neural networks (CNNs) have been widely used in surface defect detection. However, no CNN architecture is suitable for all detection tasks and designing effective task-specific requires considerable effort. The neural architecture search (NAS) technology makes it possible to automatically generate adaptive data-driven networks. Here, we propose a new method called NAS-ASDet to adaptively design network for surface defect detection. First, a refined and industry-appropriate search space that can adaptively adjust the feature distribution is designed, which consists of repeatedly stacked basic novel cells with searchable attention operations. Then, a progressive search strategy with a deep supervision mechanism is used to explore the search space faster and better. This method can design high-performance and lightweight defect detection networks with data scarcity in industrial scenarios. The experimental results on four datasets demonstrate that the proposed method achieves superior performance and a relatively lighter model size compared to other competitive methods, including both manual and NAS-based approaches.

cs.CV

Proximal Causal Inference without Uniqueness Assumptions

We consider identification and inference about a counterfactual outcome mean when there is unmeasured confounding using tools from proximal causal inference (Miao et al. [2018], Tchetgen Tchetgen et al. [2020]). Proximal causal inference requires existence of solutions to at least one of two integral equations. We motivate the existence of solutions to the integral equations from proximal causal inference by demonstrating that, assuming the existence of a solution to one of the integral equations, $\sqrt{n}$-estimability of a linear functional (such as its mean) of that solution requires the existence of a solution to the other integral equation. Solutions to the integral equations may not be unique, which complicates estimation and inference. We construct a consistent estimator for the solution set for one of the integral equations and then adapt the theory of extremum estimators to find from the estimated set a consistent estimator for a uniquely defined solution. A debiased estimator for the counterfactual mean is shown to be root-$n$ consistent, regular, and asymptotically semiparametrically locally efficient under additional regularity conditions.

stat.ME