SearcharxivSearch

arXiv subjects

Wenhui Li

Publications and source records attributed to Wenhui Li.

At least 19 recordsLinked to original sources

Multi-Source Prediction-Powered Inference

Prediction-powered inference integrates a small gold-standard dataset with large pseudo-labeled data, whose labels are generated by machine learning methods, to enhance statistical inference. In modern applications, multiple data sources and diverse machine learning methods often give rise to multiple pseudo-labeled datasets, each encoding potentially different aspects of the underlying information. However, how to optimally combine multiple data sources and machine learning methods for statistical inference remains unclear. To address this problem, we propose a multi-source prediction-powered inference method by aggregating multiple pseudo-labeled datasets together, where the aggregation weights are estimated by minimizing the asymptotic volume of the resulting confidence region. We study both homogeneous settings, where the source and target distributions coincide, and heterogeneous settings, where distributional discrepancies arise between source and target distributions, including covariate shift and domain shift. Theoretically, we establish the asymptotic normality of the proposed estimator and show that the resulting confidence-region volume is asymptotically equivalent to the oracle optimal volume within the proposed weighting class. We further characterize when our method yields smaller confidence regions compared with both classical target-only inference and single-source prediction-powered inference. Simulation studies and a real-data application on dual-energy X-ray absorptiometry measured high body fat prevalence show that MPPI can reduce confidence-region volume while maintaining inferential validity in the settings considered.

stat.ME

Counterexamples regarding elementary symmetric partitions

Ballantine, Beck, and Merca defined the elementary symmetric partition map pre$_j$ that sends a partition $\lambda$ to a larger partition whose parts are the summands appearing in the evaluation of the $j$-th elementary symmetric polynomial on $\lambda$. They conjectured that pre$_j$ is injective on the set of partitions of $n$ with length $\ell \geq j$. The $\ell = j$ case was disproved by Devnani and Eyyunni; they instead conjectured the statement to be true for $\ell > j$. In this article, we answer this refined conjecture in the negative by proving that pre$_j$ is not injective on partitions of $n$ with length $2j$ for $j \geq 3$. We also prove that the analogous map prh$_j$ defined via the complete homogenous symmetric polynomial is injective on the set of all partitions.

math.CO

Estimation of Directed Acyclic Graphs by Frequentist Model Averaging

Directed acyclic graphs provide a fundamental tool for representing directed dependence structures in multivariate network data, and are widely used to model financial and economic networks. However, accurate and interpretable estimation remains challenging under graph structural uncertainty. We propose an optimal model averaging method for directed acyclic Gaussian graphs. With a set of candidate models varying by graph structures, we average estimates from candidate models using weights that minimize a penalized negative log-likelihood criterion. In contrast to existing approaches, we not only establish the asymptotic optimality, weight consistency, and parameter consistency of the proposed method, but also explicitly characterize how different candidate models affect the convergence rate. Moreover, we prove parameter consistency even when all candidate graph models are misspecified. Results from simulation studies and a real-data analysis on the banks' international liability data show the promise of the proposed method.

stat.ME

Enabling Agile Ambient IoT Networking via a Parameterized Hybrid Radio

The emergence of Ambient IoT signals a paradigm shift toward massive batteryless networking. However, the absence of an agile physical layer substrate remains a fundamental barrier to research and standardization. Current testbeds are hindered by decoupled radio paths, high static power, and cumbersome control methods, which stifle rapid protocol prototyping. In this paper, we present Janus, the first hybrid active-passive configurable radio architected for agile Ambient IoT networking. Janus introduces a parameterized architecture that unifies passive and active transmission into a single RF front end, abstracting complex physical layer behaviors into concise parameters. This design enables a system-level control plane for dynamic mode transitions and an energy management plane for fine-grained harvesting across multiple sources. We implement a compact PCB prototype and evaluate its performance across diverse protocol landscapes, including 3GPP A-IoT, IEEE 802.11 AMP, and Bluetooth SIG. Our experimental results demonstrate that Janus achieves communication performance on par with dedicated radios while significantly reducing configuration overhead. Ultimately, Janus serves as a versatile enabler for validating emerging protocols and accelerating the standardization of next-generation low-power networks.

cs.NI

LiftAvatar: Kinematic-Space Completion for Expression-Controlled 3D Gaussian Avatar Animation

We present LiftAvatar, a new paradigm that completes sparse monocular observations in kinematic space (e.g., facial expressions and head pose) and uses the completed signals to drive high-fidelity avatar animation. LiftAvatar is a fine-grained, expression-controllable large-scale video diffusion Transformer that synthesizes high-quality, temporally coherent expression sequences conditioned on single or multiple reference images. The key idea is to lift incomplete input data into a richer kinematic representation, thereby strengthening both reconstruction and animation in downstream 3D avatar pipelines. To this end, we introduce (i) a multi-granularity expression control scheme that combines shading maps with expression coefficients for precise and stable driving, and (ii) a multi-reference conditioning mechanism that aggregates complementary cues from multiple frames, enabling strong 3D consistency and controllability. As a plug-and-play enhancer, LiftAvatar directly addresses the limited expressiveness and reconstruction artifacts of 3D Gaussian Splatting-based avatars caused by sparse kinematic cues in everyday monocular videos. By expanding incomplete observations into diverse pose-expression variations, LiftAvatar also enables effective prior distillation from large-scale video generative models into 3D pipelines, leading to substantial gains. Extensive experiments show that LiftAvatar consistently boosts animation quality and quantitative metrics of state-of-the-art 3D avatar methods, especially under extreme, unseen expressions.

cs.CV

StealthMark: Harmless and Stealthy Ownership Verification for Medical Segmentation via Uncertainty-Guided Backdoors

Annotating medical data for training AI models is often costly and limited due to the shortage of specialists with relevant clinical expertise. This challenge is further compounded by privacy and ethical concerns associated with sensitive patient information. As a result, well-trained medical segmentation models on private datasets constitute valuable intellectual property requiring robust protection mechanisms. Existing model protection techniques primarily focus on classification and generative tasks, while segmentation models-crucial to medical image analysis-remain largely underexplored. In this paper, we propose a novel, stealthy, and harmless method, StealthMark, for verifying the ownership of medical segmentation models under black-box conditions. Our approach subtly modulates model uncertainty without altering the final segmentation outputs, thereby preserving the model's performance. To enable ownership verification, we incorporate model-agnostic explanation methods, e.g. LIME, to extract feature attributions from the model outputs. Under specific triggering conditions, these explanations reveal a distinct and verifiable watermark. We further design the watermark as a QR code to facilitate robust and recognizable ownership claims. We conducted extensive experiments across four medical imaging datasets and five mainstream segmentation models. The results demonstrate the effectiveness, stealthiness, and harmlessness of our method on the original model's segmentation performance. For example, when applied to the SAM model, StealthMark consistently achieved ASR above 95% across various datasets while maintaining less than a 1% drop in Dice and AUC scores, significantly outperforming backdoor-based watermarking methods and highlighting its strong potential for practical deployment. Our implementation code is made available at: https://github.com/Qinkaiyu/StealthMark.

cs.CV

Coexistence of stripe order and superconductivity in NaAlSi

Here, we report a scanning tunneling microscopy study on an s-wave superconductor NaAlSi, revealing the coexistence of stripe order and superconductivity. This stripe order manifests as a unidirectional spatial charge modulation with a commensurate period of four times the lattice constant. This modulation undergoes a phase shift in the differential conductance maps under opposite bias voltages, while its period remains approximately constant over an energy range of $\pm$50 meV. These features suggest that this stripe is likely a static charge order. Furthermore, we find that the stripe order imposes a periodic modulation on the intensity of the superconducting coherence peaks. This work provides new perspectives on the intricate interplay between stripe order and s-wave superconductivity.

cond-mat.supr-con

Metaphor-based Jailbreak Attacks on Text-to-Image Models

Text-to-image (T2I) models commonly incorporate defense mechanisms to prevent the generation of sensitive images. Unfortunately, recent jailbreak attacks have shown that adversarial prompts can effectively bypass these mechanisms and induce T2I models to produce sensitive content, revealing critical safety vulnerabilities. However, existing attack methods implicitly assume that the attacker knows the type of deployed defenses, which limits their effectiveness against unknown or diverse defense mechanisms. In this work, we reveal an underexplored vulnerability of T2I models to metaphor-based jailbreak attacks (MJA), which aims to attack diverse defense mechanisms without prior knowledge of their type by generating metaphor-based adversarial prompts. Specifically, MJA consists of two modules: an LLM-based multi-agent generation module (LMAG) and an adversarial prompt optimization module (APO). LMAG decomposes the generation of metaphor-based adversarial prompts into three subtasks: metaphor retrieval, context matching, and adversarial prompt generation. Subsequently, LMAG coordinates three LLM-based agents to generate diverse adversarial prompts by exploring various metaphors and contexts. To enhance attack efficiency, APO first trains a surrogate model to predict the attack results of adversarial prompts and then designs an acquisition strategy to adaptively identify optimal adversarial prompts. Extensive experiments on T2I models with various external and internal defense mechanisms demonstrate that MJA achieves stronger attack performance while using fewer queries, compared with six baseline methods. Additionally, we provide an in-depth vulnerability analysis suggesting that metaphor-based adversarial prompts evade safety mechanisms by inducing semantic ambiguity, while sensitive images arise from the model's probabilistic interpretation of concealed semantics.

cs.CR

T2I-RiskyPrompt: A Benchmark for Safety Evaluation, Attack, and Defense on Text-to-Image Model

Using risky text prompts, such as pornography and violent prompts, to test the safety of text-to-image (T2I) models is a critical task. However, existing risky prompt datasets are limited in three key areas: 1) limited risky categories, 2) coarse-grained annotation, and 3) low effectiveness. To address these limitations, we introduce T2I-RiskyPrompt, a comprehensive benchmark designed for evaluating safety-related tasks in T2I models. Specifically, we first develop a hierarchical risk taxonomy, which consists of 6 primary categories and 14 fine-grained subcategories. Building upon this taxonomy, we construct a pipeline to collect and annotate risky prompts. Finally, we obtain 6,432 effective risky prompts, where each prompt is annotated with both hierarchical category labels and detailed risk reasons. Moreover, to facilitate the evaluation, we propose a reason-driven risky image detection method that explicitly aligns the MLLM with safety annotations. Based on T2I-RiskyPrompt, we conduct a comprehensive evaluation of eight T2I models, nine defense methods, five safety filters, and five attack strategies, offering nine key insights into the strengths and limitations of T2I model safety. Finally, we discuss potential applications of T2I-RiskyPrompt across various research fields. The dataset and code are provided in https://github.com/datar001/T2I-RiskyPrompt.

cs.CR

PEARL: Performance-Enhanced Aggregated Representation Learning

Representation learning is a key technique in modern machine learning that enables models to identify meaningful patterns in complex data. However, different methods tend to extract distinct aspects of the data, and relying on a single approach may overlook important insights relevant to downstream tasks. This paper proposes a performance-enhanced aggregated representation learning method, which combines multiple representation learning approaches to improve the performance of downstream tasks. The framework is designed to be general and flexible, accommodating a wide range of loss functions commonly used in machine learning models. To ensure computational efficiency, we use surrogate loss functions to facilitate practical weight estimation. Theoretically, we prove that our method asymptotically achieves optimal performance in downstream tasks, meaning that the risk of our predictor is asymptotically equivalent to the theoretical minimum. Additionally, we derive that our method asymptotically assigns nonzero weights to correctly specified models. We evaluate our method on diverse tasks by comparing it with advanced machine learning models. The experimental results demonstrate that our method consistently outperforms baseline methods, showing its effectiveness and broad applicability in real-world machine learning scenarios.

stat.ML

Generalized optimal parameter-transfer learning through Mallows-type model averaging

In many economic applications, multiple source datasets are available, but their effective combination is challenging due to heterogeneity across datasets. To address this problem, we study a parameter-transfer framework that shares only source-side estimates and propose a Mallows-type model averaging method for combining target and source models in the parametric setting. The weights are obtained from a Mallows-type criterion that is unbiased for the target prediction risk up to a weight-independent term, extending the classical Mallows criterion to the parameter-transfer framework. We establish that the proposed weights are asymptotically optimal when the target model is misspecified, and asymptotically allocate weights only to informative sources when the target model is correctly specified. These guarantees do not require any source model to be correctly specified. We also consider extensions of the framework to semiparametric and panel data settings. Simulation studies and house price application further demonstrate the effectiveness of our approach.

stat.ME

Direct evidence of intrinsic Mott state and its layer-parity oscillation in a breathing kagome crystal down to monolayer

We report direct spectroscopic evidence of correlation-driven Mott states in layered Nb$_3$Cl$_8$ through combining scanning tunneling microscopy (STM) and dynamical mean-field theory. The Hubbard bands persist down to monolayer, providing the definitive evidence for the Mottness in Nb$_3$Cl$_8$. While the size of the Mott gap remains almost constant across all layers, a striking layer-parity-dependent oscillation emerges in the local density of states (LDOS) between even (n = 2,4,6) and odd layers (n = 1,3,5), which arises from the dimerization and correlation modulation of the obstructed atomic states, respectively. Our conclusions are supported by a critical technical advance in atomic-scale LDOS mapping for highly insulating systems. This work provides the definitive experimental verification of correlation-driven Mott ground states in Nb3Cl8 while establishing a general protocol for investigating the interplay of electronic correlation and interlayer coupling in layered insulators by using low-temperature STM technique.

cond-mat.str-el

Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning

Text-to-Image(T2I) models typically deploy safety filters to prevent the generation of sensitive images. Unfortunately, recent jailbreaking attack methods manually design instructions for the LLM to generate adversarial prompts, which effectively bypass safety filters while producing sensitive images, exposing safety vulnerabilities of T2I models. However, due to the LLM's limited understanding of the T2I model and its safety filters, existing methods require numerous queries to achieve a successful attack, limiting their practical applicability. To address this issue, we propose Reason2Attack(R2A), which aims to enhance the LLM's reasoning capabilities in generating adversarial prompts by incorporating the jailbreaking attack into the post-training process of the LLM. Specifically, we first propose a CoT example synthesis pipeline based on Frame Semantics, which generates adversarial prompts by identifying related terms and corresponding context illustrations. Using CoT examples generated by the pipeline, we fine-tune the LLM to understand the reasoning path and format the output structure. Subsequently, we incorporate the jailbreaking attack task into the reinforcement learning process of the LLM and design an attack process reward that considers prompt length, prompt stealthiness, and prompt effectiveness, aiming to further enhance reasoning accuracy. Extensive experiments on various T2I models show that R2A achieves a better attack success ratio while requiring fewer queries than baselines. Moreover, our adversarial prompts demonstrate strong attack transferability across both open-source and commercial T2I models.

cs.CR

Spectroscopic evidence of symmetry breaking in the superconducting vortices of UTe2

The recently discovered heavy-fermion superconductor, UTe2, is an excellent candidate for spin-triplet superconductors where electrons form spin-triplet Cooper pairs with spin S = 1 and odd parity. Unconventional superconductivity often hosts unconventional vortices. Yet, the vortex core and lattice in UTe2 have not been directly visualized and characterized. Here, by using ultralow-temperature scanning tunneling microscopy and spectroscopy, we study the superconducting vortices on the (0-11) surface termination of UTe2 with an out-of-plane external magnetic field. At the center of the vortex core, we observe a robust zero-energy vortex-core state which exhibits a cigar-shaped spatial distribution and extends to ~30 nm along the [100] direction (crystallographic a axis) of UTe2. Along the direction perpendicular to [100], the superconducting gap is deeper and the coherence peak on one side of the vortex core is stronger than on the opposite side, and they are even enhanced in comparison with those under zero field. Due to the anisotropy of magnetic susceptibility in UTe2, the asymmetric dI/dV spectra on the two sides of the vortex core result from the interplay between the magnetization-induced bound current and supercurrent around the vortex core. Our work reveals the important role of magnetization in the vortex behaviors of UTe2 and provides essential microscopic information for understanding its superconducting properties in magnetic field.

cond-mat.supr-con

Magnetic Bloch States at Integer Flux Quanta Induced by Super-moir\'e Potential in Graphene Aligned with Twisted Boron Nitride

Two-dimensional electron systems in both magnetic fields and periodic potentials are described by Hofstadter butterfly, a fundamental problem of solid-state physics. While moir\'e systems provide a powerful method to realize this spectrum, previous experiments, however, have been limited to fractional flux quanta regime due to the difficulty of building ~ 50 nm periodic modulations. Here, we demonstrate a super-moir\'e strategy to overcome this challenge. By aligning monolayer graphene (G) with 1.0{\deg} twisted hexagonal boron nitride (t-hBN), a 63.2 nm bichromatic G/t-hBN super-moir\'e is constructed, made possible by exploiting the electrostatic nature of t-hBN potential. Under magnetic field B, magnetic Bloch states at integer flux quanta (1-9) are achieved and observed as integer Brown-Zak oscillations, expanding the flux quanta from factions to integers. Theoretical analysis reproduces these experimental findings. This work opens new avenues to study unexplored Hofstadter butterfly, explore emergent topological order at integer flux quanta and engineer long-wavelength periodic modulations.

cond-mat.mes-hall

Domain Adaptation from Generated Multi-Weather Images for Unsupervised Maritime Object Classification

The classification and recognition of maritime objects are crucial for enhancing maritime safety, monitoring, and intelligent sea environment prediction. However, existing unsupervised methods for maritime object classification often struggle with the long-tail data distributions in both object categories and weather conditions. In this paper, we construct a dataset named AIMO produced by large-scale generative models with diverse weather conditions and balanced object categories, and collect a dataset named RMO with real-world images where long-tail issue exists. We propose a novel domain adaptation approach that leverages AIMO (source domain) to address the problem of limited labeled data, unbalanced distribution and domain shift in RMO (target domain), enhance the generalization of source features with the Vision-Language Models such as CLIP, and propose a difficulty score for curriculum learning to optimize training process. Experimental results shows that the proposed method significantly improves the classification accuracy, particularly for samples within rare object categories and weather conditions. Datasets and codes will be publicly available at https://github.com/honoria0204/AIMO.

cs.CV

Highly anisotropic Drude-weight-reduction and enhanced linear-dichroism in van der Waals Weyl semimetal Td-MoTe2 with coherent interlayer electronic transport

Weyl semimetal (WSM) states can be achieved by breaking spatial-inversion symmetry or time reversal symmetry. However, the anisotropy of the energy reduction contributing to the emergence of WSM states has seldom been investigated by experiments. A van der Waals metal MoTe2 exhibits a type-II WSM phase below the monoclinic-to-orthorhombic-phase-transition temperature Tc ~ 250 K. Here, we report a combined linearly-polarized optical-spectroscopy and electrical-transport study of MoTe2 at different temperatures. The Drude components in the a-axis, b-axis and c-axis optical conductivity spectra, together with the metallic out-of-plane and in-plane electrical resistivities, indicate the coherent inter-layer and in-plane charge transports. Moreover, the Drude weight in {\sigma}1a({\omega}), rather than the Drude weights in {\sigma}1b({\omega}) and {\sigma}1c({\omega}), decreases dramatically below Tc, which exhibits a highly anisotropic decrease in its Drude weight and thus suggests a strongly anisotropic reduction of the electronic kinetic energy in the WSM phase. Furthermore, below Tc, due to the in-plane anisotropic spectral-weight transfer from Drude component to high-energy region, the in-plane inter-band-absorption anisotropy increases remarkably around 770 meV, and has the largest value (~ 0.68) of normalized linear dichroism among the reported type-II WSMs. Our work sheds light on seeking new WSMs and developing novel photonic devices based on WSMs.

cond-mat.mtrl-sci

Gate-controlled superconducting switch in GaSe/NbSe$_2$ van der Waals heterostructure

The demand for low-power devices is on the rise as semiconductor engineering approaches the quantum limit and quantum computing continues to advance. Two-dimensional (2D) superconductors, thanks to their rich physical properties, hold significant promise for both fundamental physics and potential applications in superconducting integrated circuits and quantum computation. Here, we report a gate-controlled superconducting switch in GaSe/NbSe$_2$ van der Waals (vdW) heterostructure. By injecting high-energy electrons into NbSe$_2$ under an electric field, a non-equilibrium state is induced, resulting in significant modulation of the superconducting properties. Owing to the intrinsic polarization of ferroelectric GaSe, a much steeper subthreshold slope and asymmetric modulation are achieved, which is beneficial to the device performance. Based on these results, a superconducting switch is realized that can reversibly and controllably switch between the superconducting and normal state under an electric field. Our findings highlight a significant high-energy injection effect from band engineering in 2D vdW heterostructures combining superconductors and ferroelectric semiconductors, and demonstrate the potential applications for superconducting integrated circuits.

cond-mat.supr-con