SearcharxivSearch

arXiv subjects

Fan Li

Publications and source records attributed to Fan Li.

At least 19 recordsLinked to original sources

Nonparametric heterogeneous causal mediation with orthogonal machine learning

Causal mediation analysis decomposes the total effect of an intervention on an outcome into a direct pathway and an indirect pathway transmitted through a mediator, but standard methods typically summarize these pathways using population average effects. In many applications, however, the indirect effect may vary substantially across individual profiles. We propose an orthogonal statistical learning framework for estimating heterogeneous causal mediation effects conditional on individual characteristics. The method constructs a class of weighted Neyman orthogonal losses motivated by influence function representations of weighted population average effects. These losses directly target conditional mediation estimands whose minimizers are locally insensitive to nuisance estimation errors. We implement the resulting learners under a two-stage meta-learning framework with regularized linear sieves as second-stage smoothers, and introduce a combination of targeted learning and orthogonal learning designed to improve stability when mediator density ratios are unstable. We establish $L^2$ and uniform limit theory and develop pointwise and uniform confidence bands. Simulation studies show that the proposed orthogonal learners reduce the mean integrated squared error by more than $50\%$ compared with existing model-based methods and provide computationally efficient inference in nonlinear settings. The CARDIA, PSACR, and STAR analyses reveal heterogeneous mediated effects across cardiometabolic, psychological, and educational settings.

stat.ME

Cluster randomized crossover trials with very few clusters but multiple periods: which analyses for continuous outcomes should be used?

Cluster randomized crossover (CRXO) trials are often used when individual randomization is impractical and the number of available clusters is limited. However, statistical analysis of CRXO trials is complex because of the need to account for complex correlation structures over time. It becomes especially challenging when very few clusters are used because standard modeling assumptions may lead to unstable variance estimates, poor confidence interval coverage, and inflated type I error. This study evaluates individual-level mixed-effects and fixed-effects models with and without a cluster-period random effect, cluster-period summary analysis using normal- or \(t\)-based inference, and two-period crossover-difference estimators. Using extensive simulation studies under both nested exchangeable and discrete time decay correlation structures, we compare model performance in terms of bias, root mean squared error, coverage probability, type I error, and convergence. Across scenarios, all models produced approximately unbiased treatment effect estimates, but their inferential performance differed substantially. Models that explicitly accounted for cluster-period heterogeneity generally provided the most reliable control of coverage and type I error, whereas simpler exchangeable models performed adequately only when the true correlation structure closely matched their assumptions. Cluster-period level analysis performance improved with increasing numbers of periods but was unreliable in the sparsest designs. Overall, the findings suggest that in CRXO trials with very few clusters, accurate modeling of cluster-period correlation is more important than the choice between fixed and random cluster intercepts, and that results from extremely sparse designs should be interpreted with caution.

stat.ME

Deeply Interleaved Text-Image Contexts for Multimodal LLMs Assessment

Current evaluations and training of multimodal models predominantly focus on multi-image tasks, largely overlooking interleaved text-image scenarios. In such multi-image tasks, text typically serves merely as task instructions, lacking deep semantic interaction with the visual content. In contrast, realworld applications like text-image co-creation, character tracking, and spatial reconstruction require constant interaction between text and images. Consequently, models must possess a deep understanding of these interleaved contexts. To bridge this gap, we introduce a novel benchmark, TIC-Bench (deeply interleaved Text-Image Contexts), designed to evaluate the capability of models to integrate text-image clues and recover the ground truth facts within deeply interleaved contexts. This benchmark encompasses three core domains: Logical, Temporal, and Spatial Association, which are further categorized into eight specific types, comprising a total of 2,280 questions. We evaluated 10 state-of-the-art MLLMs and observed a substantial performance gap compared to human experts, together with persistent difficulties in integrating evidence distributed across interleaved visual and textual inputs. Ultimately, this benchmark provides a valuable analytical tool for assessing and advancing the ability of multimodal models to effectively integrate text and image information in deeply interleaved contexts. TIC-Bench is publicly available at https://huggingface.co/datasets/pino10010/TIC-Bench

cs.CV

Saturation in G: simple & robust causal inference in cluster randomized trials with informative cluster sizes

Cluster randomized trials (CRTs) can exhibit informative cluster sizes (ICS) where cluster size is associated with outcomes and/or treatment effects. Under ICS, the individual and cluster-average treatment effects (iATE, cATE) can diverge, and the conventional linear mixed-effects model (LMM) and generalized estimating equation (GEE) with an exchangeable working correlation can produce data-dependent weighted contrasts that are not consistent for either estimand. In these settings with ICS, we propose easy to implement "cluster-size saturated models with g-computation" (CS-g), which employ a simple two-step adjustment to standard practice: (1.) augment the appropriately weighted working LMM or GEE with a saturated continuous cluster-size main effect and treatment x cluster-size interaction, and (2.) apply g-computation to target an interpretable marginal estimand. We prove that the appropriately weighted cluster-size saturated LMM with g-computation and more general cluster-size saturated GEE with g-computation can consistently target the iATE and cATE, among a broad class of interpretable estimands, while allowing for ICS. Crucially, this consistency holds under arbitrary misspecification of other model components, including the functional form of the saturated cluster-size terms. Furthermore, we demonstrate exact finite-sample equivalence between these consistent CS-g estimators and their model-robust standardization counterparts. Across simulations with continuous and binary outcomes, the proposed CS-g estimators were unbiased, more efficient than other consistent estimators, and returned greater power to detect ICS. A re-analysis of the PPACT P-CRT further illustrates the approach. Altogether, CS-g offers a simple, robust, and efficient route to target interpretable marginal effects in P-CRTs with ICS.

stat.ME

Orthogonal double residual learning for optimal individualized treatment rules

Individualized treatment rules (ITRs) map baseline characteristics to treatment recommendations, with the optimal ITR maximizing expected reward or policy welfare. Indirect methods may require restrictive modeling assumptions, whereas direct methods can be sensitive to nuisance estimation error and limited overlap. We propose orthogonal double residual learning (ODRL), a two-stage, cross-fitted framework that directly targets the optimal ITR through cost-sensitive classification using the product of treatment and outcome residuals. To our knowledge, ODRL is the first direct method with a universally Neyman orthogonal objective requiring neither restrictive modeling assumptions nor inverse propensity score weighting. Thus, nuisance estimation errors affect regret through a second-order product, and ODRL remains robust under limited overlap. The Fisher consistent objective accommodates general decision rule sieves. We establish nonasymptotic high probability value function regret bounds relative to the Bayes classifier for VC classes, including linear rules and decision trees, and calibrated regret bounds for surrogate relaxations using support vector machines and deep ReLU neural networks. We further show that generic surrogate relaxations need not preserve orthogonality, whereas bounded score hinge learning does. Simulations demonstrate strong performance across complex and linear decision boundaries, limited overlap, and working model misspecification. Applications to the Right Heart Catheterization study and the Oxford Net Zero experiment illustrate interpretable treatment or policy recommendations. The \texttt{odrlITR} R package implements ODRL.

stat.ME

Can Large Language Models Explain Flight Safety Events? A Prior-Guided Semantic LLM-based Approach

Improving flight safety with flight data requires not only accurate detection of risk events, but more importantly, clear interpretation of their underlying causes at the level of pilot control behavior. Existing explainable AI techniques, such as feature importance maps, often require considerable domain knowledge to translate them into operationally meaningful explanations. Large Language Models (LLMs), which excel at language reasoning, bring a promising solution to this issue. However, applying LLMs in this domain presents key challenges such as modal inconsistency, limited classification ability, scarcity of task-specific data for fine-tuning, and lack of domain knowledge. To overcome these challenges, we propose FlightLLM, a prior-guided semantic LLM-based approach for interpretable flight safety analysis. Specifically, we first perform feature engineering to address modal inconsistency, combining statistical descriptors with physically meaningful flight indicators. This representation is further processed by a Semantic Discretization module, which converts abstract numerical patterns into qualitative descriptions that are more compatible with language reasoning. In addition, since LLMs are not inherently strong classifiers, CatBoost is incorporated as a statistical expert, and its prediction results are injected into the prompt as prior guidance. A contrastive few-shot learning strategy is further adopted to compensate for limited data. Finally, we design structured prompts to embed aviation-specific knowledge into the inference process. Using hard landing, a representative risk event with complex causal mechanisms, as an anchor point, we evaluate FlightLLM on a dataset of 704 real-world A320 flight samples. Experimental results show that the proposed approach achieves competitive classification performance while generating direct and reasonable explanations for event causes.

cs.AI

Doubly robust estimation of while-alive estimands in individually-randomized and cluster-randomized trials

Randomized trials in chronic disease settings often measure treatment benefit through recurrent non-fatal events that are truncated by death, where conventional summaries either discard recurrences, conflate the treatment effect with survival, or treat death as censoring and forfeit a causal interpretation. While-alive estimands measure event burden per unit time alive, but doubly robust estimation for the exposure-weighted while-alive rate remains undeveloped, particularly in cluster-randomized trials (CRTs). We develop a doubly robust estimator based on a local Nelson-Aalen representation, with augmented estimating equations targeting the marginal hazard of the terminal event and the weighted recurrent event rate among those alive; the estimator accommodates multiple event types through prespecified clinical weights and remains consistent if either the censoring model or the outcome working models are correctly specified. For CRTs, we define a new pair of individual-average and cluster-average estimands under informative cluster size, with inference based on cluster-level influence functions. We establish component-wise double robustness and asymptotic normality, corroborate the theory in simulations, and illustrate the methods with reanalyses of data from two completed randomized trials.

stat.ME

DualMiT-Net: Local-Global Transformer-Convolutional Fusion for Breast Mass Segmentation in Mammographic Regions of Interest

Breast mass segmentation is an important step in computer-aided mammography, but it remains difficult because masses can have low contrast, irregular shapes, and boundaries that blend with surrounding breast tissue. To address this problem, we present DualMiT-Net, a dual-branch network that uses both a focused view of the mass and a wider view of the surrounding tissue. The local branch uses a Mix Transformer (MiT-B5) encoder to learn mass shape, texture, and boundary information, while the global branch uses an EfficientNet-B5 encoder to learn surrounding breast context. Features from the two branches are shared at the deeper encoder levels and are then progressively fused in a single decoder. A spatial gate controls how much global information is added during decoding. We also evaluated four input representations and selected a percentile-windowed mammogram combined with a Gabor texture response. The model was trained and evaluated on the mass subset of the Curated Breast Imaging Subset of the Digital Database for Screening Mammography (CBIS-DDSM) using a patient-level split. Across three training runs, DualMiT-Net with exponential moving average weights achieved a mean Dice coefficient of 0.9375 and a mean Intersection over Union of 0.8834. It also achieved better Dice and IoU scores than six standard encoder-decoder baselines trained using the same data and training settings. These results show that combining local mass information with wider breast context can provide accurate and consistent breast mass segmentation.

cs.CV

Spin nematic liquid crystal and scalar spin chirality in tetragonal lattice YbMnBi$_2$

A spin nematic order, analogous to the nematic liquid crystal, characterizes the spontaneous breaking of spin-space rotational symmetry while preserving time-reversal ($T$) symmetry. In contrast, scalar spin chirality (SSC), a composite three-spin order, breaks $T$ symmetry and is known to induce an anomalous Hall effect (AHE). Although a spin nematic phase has been suggested in frustrated magnets and the square-lattice iridate, how it might affect magnetotransport properties is unknown. Here we use polarized neutron scattering to show that tetragonal $A$MnBi$_2$ ($A$ = Ca, Yb) is a strictly $c$-axis-aligned collinear antiferromagnet (C-type), with $T_N \approx 270$ K and 290 K, respectively. On cooling from 450 K to $T_N$, low-energy spin excitations in YbMnBi$_2$ spontaneously change from isotropic to anisotropic in spin space within the tetragonal plane, forming a dynamic spin nematic phase around 400 K due to heavy Yb-induced spin-orbit coupling, before gapping out below $T_N$. Similar measurements on CaMnBi$_2$ reveal isotropic paramagnetic scattering without a spin nematic phase above $T_N$. Under an in-plane magnetic field, the Yb$^{3+}$ moments may interact with the dynamic spin nematic phase to induce nonzero SSC, giving rise to AHE and an anomalous Nernst effect (ANE) in YbMnBi$_2$ that are absent in CaMnBi$_2$ above $T_N$. A symmetry-based Ginzburg-Landau analysis shows that coupling terms between the nematic order and SSC are allowed under an external magnetic field, which could explain the rapid increase of AHE with field in YbMnBi$_2$. Our results provide compelling evidence for dynamic SSC-induced AHE and ANE in the paramagnetic phase of a compensated collinear antiferromagnet, opening a new avenue for the physics of composite spin orders and room-temperature spintronics without magnetic order.

cond-mat.str-el

Estimating the Average Treatment Effect under Limited Overlap via Polynomial Approximation and Extrapolation

Estimating the average treatment effect (ATE) remains a fundamental challenge in observational studies in the presence of poor or limited covariate overlap. Although the inverse probability weighting (IPW) estimator is a widely used approach for estimating the ATE, its performance can deteriorate substantially when overlap is limited, often resulting in increased finite sample bias and unreliable confidence intervals. One common strategy is to shift attention from the original target estimand, the ATE, to alternative estimands that are less sensitive to extreme propensity scores; however, doing so changes the scientific question of interest. In this manuscript, we propose a novel ATE estimator that preserves the original target estimand, the ATE, while improving robustness to limited overlap. A key idea is that a class of estimands can be expressed by a polynomial function of a hyperparameter characterizing the estimands. Exploiting this structure, the proposed method computes IPW estimators for a sequence of such estimands, models these estimates using a polynomial function, and extrapolates to recover the ATE. We show that the estimator has consistency and asymptotic normality under weaker overlap conditions than required for the standard IPW estimator. Simulation studies demonstrate that the proposed method improves estimation accuracy and interval performance in settings with limited overlap. In addition to its theoretical and empirical advantages, the proposed approach has a clear interpretation and is easy to implement using standard statistical software.

stat.ME

DenialRAG: Single-Document RAG Poisoning via Embedded Parametric Denial

Retrieval-augmented generation (RAG) systems are vulnerable to corpus poisoning: an attacker who inserts a crafted document into the retrieval corpus can steer the underlying large language model (LLM) toward an attacker-chosen wrong answer. Prior single-document attacks typically avoid explicitly naming and refuting the correct answer inside the poisoned passage. In this paper, we examine a complementary design and propose \emph{DenialRAG}, a single-document poisoning attack that explicitly names the correct answer, denies it, and presents an attacker-controlled explanation for favoring the wrong answer. By placing both the correct answer and the corresponding poisoned answer inside the same retrieved passage, DenialRAG embeds the conflict directly into the context seen by the generator. We evaluate DenialRAG against four published single-document poisoning attacks across three open-domain question-answering datasets, eight target LLMs from four vendors, and five inference-time defenses. The results show that attack effectiveness is strongly model-dependent: DenialRAG achieves the highest attack success rate (ASR) on all three Mistral-7B datasets and remains effective on several other target LLMs, while other attacks dominate in some model regimes. Defense results show meaningful ASR reductions but non-uniform protection, with each defense leaving residual ASR in some settings. Component-level and cross-model analyses further identify the embedded denial as the most influential tested component and show that different poisoning mechanisms lose effectiveness at different rates across model groups. Together, these results show that RAG poisoning risk cannot be fully characterized by a single attack family or a single target model.

cs.CR

A Causality-aware Infer-diagnose-refine Framework for Test-time Modality Adaptation in VLA Models

Vision-language-action (VLA) models predict sequential actions to execute tasks specified by language instructions, conditioned on visual observations and proprioceptive states. However, how to fuse modalities in VLA models remains an open problem, since robot manipulation involves dynamic phases, such as long-distance movements and close-range interactions, in which the importance of visual observations may vary over time. In this paper, we propose an infer-diagnose-refine (IDR) framework, a model-agnostic framework that can be integrated with diverse VLA architectures for refining action predictions at test time. IDR first infers actions under factual and counterfactual scenarios of visual observations, and then diagnoses the causal effects of visual observations as the estimated dynamic importance, which is finally used to refine the action predictions in a training-free manner. We further design a causality-aware action refiner to realize the IDR framework, including zero-padding interventions for inferring counterfactual actions, norm-based quantification for diagnosing causal effects, and gated residual fusion for refining actions. Extensive experiments on both simulation benchmarks and real-world tasks show improvements in overall performance across multiple VLA backbones, demonstrating the efficacy of dynamically adjusting visual importance at test time.

cs.RO

DecoyFace: Beyond Obfuscation via Controllable and Imperceptible Identity Misdirection for Privacy-Preserving Face Recognition

Split face recognition reduces client-side computation but exposes intermediate features to feature inversion attacks and unauthorized analysis by honest-but-curious (HBC) servers. Existing privacy-preserving face recognition methods mainly aim to resist unauthorized reconstruction, typically producing features whose inversion yields visibly degraded results, which may reveal the existence of protection and motivate adaptive attacks. To address this issue, we propose DecoyFace, an imperceptible decoy-oriented framework that steers unauthorized reconstruction toward a plausible but incorrect identity while preserving recognition utility. The key idea is to decompose the intermediate representation into a reconstruction-sensitive subspace and its complementary subspace. The client injects decoy identity cues into the reconstruction-sensitive subspace, while limited recognition-relevant evidence from the true sample is retained in the complementary subspace. On the server side, an authorized canonicalization module suppresses decoy-dominant components and recovers a recognition-friendly representation. This design addresses both attacker-side inversion from intercepted features and HBC server-side reconstruction from canonicalized representations. Experiments show that DecoyFace preserves competitive recognition accuracy while substantially reducing identity leakage to 2.93% under U-Net attacks and 0.74% under Flow-Matching attacks while yielding visually plausible and imperceptible reconstructions, with over 99.78% face validity on LFW dataset.

cs.CV

The Resolution of Causal Heterogeneity

Causal subgroup analyses often report a small number of groups summarizing treatment effect heterogeneity, as if that number were a well-defined estimand. Outside genuinely latent class populations, however, a ``true'' subgroup count is model dependent rather than a population functional. We replace it with a new population estimand, the resolution profile, a functional of the causal feature law giving the fewest groups explaining a prescribed fraction of causal heterogeneity, defined for every population without latent structure. Inference is organized around one cross-fitted Bayesian-bootstrap posterior for a single structured moment process, its scores corrected with influence functions, so that paths, profiles, fixed-resolution summaries, and subgroup effects follow by composition. A uniform conditional Bernstein--von Mises theorem over a loss class containing the nonsmooth quantization losses shows this posterior merges with the efficient Gaussian limit under stated nuisance-rate and margin conditions. Subgroup-number uncertainty is not model selection but threshold nonregularity, the profile being an integer-valued threshold of a continuous path, discontinuous in the law at each knot. At these knots no single-valued selector is locally uniformly consistent over root-$n$ neighborhoods, and the set-valued report obtained by inverting a simultaneous band retains locally uniform validity over exactly the same perturbations. Simulations support the approximations, and an analysis of the MineThatData e-mail experiment illustrates the resolution-indexed report, in which two to three groups summarize the visit response while finer structure falls below a noise-floor diagnostic.

stat.ME

Learning-Driven Adaptive Audit Scheduling: A Sequential Decision Approach to Off-Chain Data Integrity

We model cryptographic auditing of off-chain data as a Constrained MDP (CMDP) under partial observability: the storage node's hidden type and corruption state make the problem a POMDP, while a miss-rate ceiling rho imposes an explicit security constraint. We propose DRQN-CMDP, a Deep Recurrent Q-Network whose GRU layer maintains a belief over the latent node type, paired with Lagrangian dual ascent that adapts the miss-rate penalty lambda automatically. A pairing-free homomorphic-MAC primitive supplies O(1) on-chain verification cost. Across 13 methods--four DQN variants, PPO, A2C, PPO-Lagrangian, a stateful Bayesian heuristic, three fixed-rule baselines, and an oracle-informed heuristic--DRQN-CMDP achieves a favourable balance: 83% lower gas than fixed high-frequency auditing, single-digit miss rate (7.5%), and moderate detection latency--a combination no other method matches across all three objectives simultaneously.

cs.AI

Reliable Associative Lookup in Content-Addressable Memory

Content Addressable Memory (CAM) is an important memory paradigm, which performs fast search by comparing an input query against all stored entries in parallel, achieving $O(1)$ lookup complexity. CAM is typically built upon conventional memory technologies, such as SRAM and Non-Volatile Memory (NVM). Accordingly, CAM can also be subject to the reliability challenges of these underlying technologies. In traditional memory systems, protection codes play a critical role in ensuring reliability and have been extensively studied. However, protection codes for CAM have remained largely unexplored. This paper takes an initial step toward addressing this longstanding gap by introducing a non-traditional code design.

cs.AR

InstanceControl: Controllable Complex Image Generation without Instance Labeling

Controllable image generation methods, such as ControlNet, have demonstrated a remarkable capacity to introduce visual conditions(e.g., depth maps) to guide image generation. However, these methods often struggle with complex multi-instance scenes, frequently leading to attribute confusion among instances. While recent approaches attempt to mitigate this via manual instance labeling, such requirements are labor-intensive. In this paper, we propose InstanceControl, a novel multi-instance controllable generation method that eliminates the need for instance labeling. We identify the primary bottleneck in existing methods as the inability to accurately associate instance descriptions with their corresponding regions within visual conditions. To address this, we leverage the Vision-Language Model (VLM) to establish instance-level correspondences between text prompts and visual conditions. Specifically, the VLM automatically parses instance descriptions from the text prompts and simultaneously predicts instance masks based on the visual conditions. Furthermore, since the predicted masks may contain noise, we introduce an adaptive mask refinement strategy that dynamically refines these instance masks during the generation process. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods, achieving superior fidelity and precise instance-level control.

cs.CV

Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation

The advancement of generative AI models capable of producing text and image marks a critical step forward in the realm of multimodal intelligence, particularly for tasks involving the interleaving of both modalities. To advance this intelligence to the next stage, it is crucial for models to autonomously generate free-form interleaved text-image sequences. In this paper, we introduce ILLUME-X, an advanced unified multimodal paradigm that enables high-quality, free-form interleaved text-image generation by improving multimodal data efficiency and stabilizing the multimodal training process. ILLUME-X comprises three key components: (i) an expanded training data pipeline optimized for interleaved text-image generation, (ii) a progressive training strategy with self-adaptive objectives for free-length multimodal token sequences, and (iii) an objective and comprehensive evaluation method ILScore for interleaved text-image sequences. Notably, our ILLUME-X outperforms previous unified models across multiple interleaved text-image generation tasks like style transfer, image decomposition and storytelling.

cs.CV