Searcharxiv⌕ Search

arXiv subjects

Yunxiao Zhang

Publications and source records attributed to Yunxiao Zhang.

At least 19 recordsLinked to original sources

A Location-Invariant Estimator of Extremal Quantile Treatment Effects for Heavy-Tailed Distributions

Quantile treatment effects (QTEs) measure the effect of a treatment on the distribution of an outcome, and their estimation at extreme quantile levels is of central interest in applications where the target quantiles lie far beyond the range of the data. For heavy-tailed potential outcomes, existing extremal QTE estimators rely on extrapolation combined with a causal extreme value index (EVI) estimator, but the resulting estimator is not invariant under a common location shift of the potential outcome distributions, even though the population QTE is. We address this issue in two steps. First, we adapt the location-invariant Fraga estimator of the EVI to the causal setting using inverse propensity score weighting. Second, we replace the original extrapolation formula with a difference-based scheme, under which the location parameter cancels when quantile differences are taken. The resulting QTE estimator is therefore location invariant. We establish the consistency and asymptotic normality of the proposed extremal QTE estimators, and provide a consistent variance estimator, leading to asymptotically valid inference. A simulation study confirms the location invariance, the stability with respect to the threshold, and the coverage of the proposed methods.

cs.LG↗

Mitigating Rubric Interference in LLM Judges via On-Policy Self-Distillation

LLM judges increasingly evaluate responses against fine-grained rubric checklists. When a sample requires multiple rubrics, current methods typically assess each in a separate inference call. Evaluating all rubrics in a single pass is a natural alternative with greater efficiency, but we find that it introduces rubric interference: the verdict on one rubric shifts depending on which other rubrics are co-present. In a preliminary study, only one-third of samples receive fully consistent verdicts when evaluated under rubric sets of varying composition. We develop a measurement framework that probes interference through four controlled operations: rubric set expansion, subsetting, reordering, and noise injection. To mitigate interference without external supervision, we propose Self-Anchored Rubric Alignment (SARA). SARA uses a model's own single-rubric judgments as stable anchors and aligns multi-rubric reasoning with these anchors through on-policy self-distillation. We validate SARA on three datasets (HealthBench, FLASK, ResearchQA) and two model families (Qwen3, Llama-3.1). SARA consistently improves evaluation consistency while maintaining agreement with both base models and GPT-4.1 as a reference judge. Furthermore, the learned consistency transfers across datasets, confirming that SARA teaches a general capability rather than fitting dataset-specific patterns.

cs.LG↗

Structured but Fragile: On the Limits of LLMs in Cybersecurity Decision-Making

Large language models (LLMs) are increasingly used in cybersecurity workflows, yet it remains unclear whether they can perform structured security reasoning or merely rely on superficial cues and prior knowledge. We study this question in the context of defence selection over attack graphs derived from real-world threat scenarios, including ransomware, supply-chain compromise, cloud abuse, Kubernetes attacks, POS malware, and ICS/OT intrusion. Given a budget constraint, LLMs must select security controls to minimise attacker success. We compare their strategies against each other and against a game-theoretic optimization baseline used as a normative reference for structured reasoning. Our results show that LLMs exhibit conditional competence. When explicit attack-graph structure is provided, they often produce coherent strategies close to the optimization baseline. However, their capabilities are fragile. LLM behaviour becomes increasingly fragile with graph complexity and is highly sensitive to framing. Small prompt changes can substantially alter rankings, and merely relabeling a poor strategy as ``optimal'' dramatically improves its evaluation. We further observe a non-monotonic relationship between formal risk and LLM judgement: strategies closest to the optimum are not necessarily ranked highest by LLM evaluators. To further probe reasoning ability, we ask LLMs to generate solvers for the same optimization problem. While the generated implementations recover the correct high-level formulation, they scale poorly compared to a purpose-built solver. Overall, our findings show that LLMs can approximate structured cybersecurity reasoning under controlled representations, but do not apply it robustly. This has important implications for the design and evaluation of AI-assisted security decision-support systems.

cs.CR↗

Global multimode squeezing in a train of ultrashort pulses from unbalanced SU(1,1) interferometers

Time-domain multiplexed continuous-variable quantum states provide a promising route toward large-scale quantum networks. Existing platforms are based on continuous-wave pumped optical parametric systems, where the durations of temporal modes are on the order of nanoseconds. Here we demonstrate the time-domain multiplexed squeezing localized in a train of ultrashort pulses by exploiting unbalanced SU(1,1) interferometer (USUI) with a mode-locked laser serving as pump. Using the pulse-resolved measurement, we reveal the correlation structure of the state is unique and fundamentally different from previous approaches. To reach the ideal intensity squeezing, in principle, both the gain of USUI and mode number $M$ involved in joint measurement should tend to infinity, illustrating the feature of global multimode squeezing. We conduct proof-of-principle experiments, in which the temporal mode duration is down to 10 ps. We verify the intensity squeezing degree $R_d$ depends on both the gain of USUI and $M$. The results show $R_d$ improves with the increase of $M$ for $M<10$ and $R_d$ is lower than shot noise level by $\sim0.9$ dB for $M>10$ when the gain of USUI is fixed. Our investigations demonstrate the emission from high gain USUI is novel, which not only possesses the unique coherent feature but also enables the realization of ultra-large-scale quantum states.

quant-ph↗

Gap reopening as a possible signature of coupling between Majorana zero modes in Sn-(Bi,Sb)2(Te,S)3-based Josephson trijunctions

In the past two decades, enormous efforts have been made to search for possible platforms and schemes to implement topological quantum computation (TQC). In exploring the Fu-Kane scheme of TQC based on Josephson trijunctions constructed on topological insulators, the predicted Majorana phase diagram of an individual trijunction has already been verified experimentally. If Majorana zero modes indeed exist in this kind of trijunction, coupling between them in multiple trijunction devices should be further expected. In this study, we fabricated Josephson devices containing two adjacent Josephson trijunctions on the surface of Sn-(Bi, Sb)2(Te, S)3 and observed a possible signature of the coupling effect manifesting as the reopening of a minigap in both trijunctions where a closure would otherwise be expected if the trijunctions existed individually. While alternative interpretations cannot be fully ruled out, our findings provide experimental support for the validity of the Fu-Kane theory and provide further motivation for advancing the TQC scheme proposed by Fu and Kane.

cond-mat.supr-con↗

Benchmarking Audio Deepfake Detection Robustness in Real-world Communication Scenarios

Existing Audio Deepfake Detection (ADD) systems often struggle to generalise effectively due to the significantly degraded audio quality caused by audio codec compression and channel transmission effects in real-world communication scenarios. To address this challenge, we developed a rigorous benchmark to evaluate the performance of the ADD system under such scenarios. We introduced ADD-C, a new test dataset to evaluate the robustness of ADD systems under diverse communication conditions, including different combinations of audio codecs for compression and packet loss rates. Benchmarking three baseline ADD models on the ADD-C dataset demonstrated a significant decline in robustness under such conditions. A novel Data Augmentation (DA) strategy was proposed to improve the robustness of ADD systems. Experimental results demonstrated that the proposed approach significantly enhances the performance of ADD systems on the proposed ADD-C dataset. Our benchmark can assist future efforts towards building practical and robustly generalisable ADD systems.

eess.AS↗

3D Gaussian Splatting for Efficient Retrospective Dynamic Scene Novel View Synthesis with a Standardized Benchmark

Retrospective novel view synthesis (NVS) of dynamic scenes is fundamental to applications such as sports. Recent dynamic 3D Gaussian Splatting (3DGS) approaches introduce temporally coupled formulations to enforce motion coherence across time. In this paper, we argue that, in a synchronized multi-view (MV) setting typical of sports, the dynamic scene at each time step is already strongly geometrically constrained. We posit that the availability of calibrated, synchronized viewpoints provides sufficient spatial consistency, and therefore, explicit temporal coupling, or complex multi-body constraints seems unnecessary for retrospective NVS. To this end, we propose an approach tailored for synchronized MV dynamic scene. By initializing the SfM-derived point cloud at the start time and propagating optimized Gaussians over time, we show that efficient retrospective NVS can be achieved without imposing a temporal deformation constraint. Complementing our methodological contribution, we introduce a Dynamic MV dataset framework built on Blender for reproducible NeRF and 3DGS research. The framework generates high-quality, synchronized camera rigs and exports training-ready datasets in standard formats, eliminating inconsistencies in coordinate conventions and data pipelines. Using the framework, we construct a dynamic benchmark suite and evaluate representative NeRF and 3DGS approaches under controlled conditions. Together, we show that, under a synchronized MV setup, efficient retrospective dynamic scene NVS can be achieved using 3DGS. At the same time, the dataset-generation framework enables reproducible and principled benchmarking of dynamic NVS methods.

cs.CV↗

Confusion-Aware Spectral Regularizer for Long-Tailed Recognition

Long-tailed image classification remains a long-standing challenge, as real-world data typically follow highly imbalanced distributions where a few head classes dominate and many tail classes contain only limited samples. This imbalance biases feature learning toward head categories and leads to significant degradation on rare classes. Although recent studies have proposed re-sampling, re-weighting, and decoupled learning strategies, the improvement on the most underrepresented classes still remains marginal compared with overall accuracy. In this work, we present a confusion-centric perspective for long-tailed recognition that explicitly focuses on worst-class generalization. We first establish a new theoretical framework of class-specific error analysis, which shows that the worst-class error can be tightly upper-bounded by the spectral norm of the frequency-weighted confusion matrix and a model-dependent complexity term. Guided by this insight, we propose the Confusion-Aware Spectral Regularizer (CAR) that minimizes the spectral norm of the confusion matrix during training to reduce inter-class confusion and enhance tail-class generalization. To enable stable and efficient optimization, CAR integrates a Differentiable Confusion Matrix Surrogate and an EMA-based Confusion Estimator to maintain smooth and low-variance estimates across mini-batches. Extensive experiments across multiple long-tailed benchmarks demonstrates that CAR substantially improves both worst-class accuracy and overall performance. When combined with ConCutMix augmentation, CAR consistently surpasses exisiting state-of-the-art long-tailed learning methods under both the training-from-scratch setting (by 2.37% ~ 4.83%) and the fine-tuning-from-pretrained setting (by 2.42% ~ 4.17%) across ImageNet-LT, CIFAR100-LT, and iNaturalist datasets.

cs.CE↗

Time-Archival Camera Virtualization for Sports and Visual Performances

Camera virtualization -- an emerging solution to novel view synthesis -- holds transformative potential for visual entertainment, live performances, and sports broadcasting by enabling the generation of photorealistic images from novel viewpoints using images from a limited set of calibrated multiple static physical cameras. Despite recent advances, achieving spatially and temporally coherent and photorealistic rendering of dynamic scenes with efficient time-archival capabilities, particularly in fast-paced sports and stage performances, remains challenging for existing approaches. Recent methods based on 3D Gaussian Splatting (3DGS) for dynamic scenes could offer real-time view-synthesis results. Yet, they are hindered by their dependence on accurate 3D point clouds from the structure-from-motion method and their inability to handle large, non-rigid, rapid motions of different subjects (e.g., flips, jumps, articulations, sudden player-to-player transitions). Moreover, independent motions of multiple subjects can break the Gaussian-tracking assumptions commonly used in 4DGS, ST-GS, and other dynamic splatting variants. This paper advocates reconsidering a neural volume rendering formulation for camera virtualization and efficient time-archival capabilities, making it useful for sports broadcasting and related applications. By modeling a dynamic scene as rigid transformations across multiple synchronized camera views at a given time, our method performs neural representation learning, providing enhanced visual rendering quality at test time. A key contribution of our approach is its support for time-archival, i.e., users can revisit any past temporal instance of a dynamic scene and can perform novel view synthesis, enabling retrospective rendering for replay, analysis, and archival of live events, a functionality absent in existing neural rendering approaches and novel view synthesis...

cs.CV↗

Audio Deepfake Detection at the First Greeting: "Hi!"

This paper focuses on audio deepfake detection under real-world communication degradations, with an emphasis on ultra-short inputs (0.5-2.0s), targeting the capability to detect synthetic speech at a conversation opening, e.g., when a scammer says "Hi." We propose Short-MGAA (S-MGAA), a novel lightweight extension of Multi-Granularity Adaptive Time-Frequency Attention, designed to enhance discriminative representation learning for short, degraded inputs subjected to communication processing and perturbations. The S-MGAA integrates two tailored modules: a Pixel-Channel Enhanced Module (PCEM) that amplifies fine-grained time-frequency saliency, and a Frequency Compensation Enhanced Module (FCEM) to supplement limited temporal evidence via multi-scale frequency modeling and adaptive frequency-temporal interaction. Extensive experiments demonstrate that S-MGAA consistently surpasses nine state-of-the-art baselines while achieving strong robustness to degradations and favorable efficiency-accuracy trade-offs, including low RTF, competitive GFLOPs, compact parameters, and reduced training cost, highlighting its strong potential for real-time deployment in communication systems and edge devices.

eess.AS↗

Exchange operation of Majorana zero modes in topological insulator-based Josephson trijunctions

Majorana zero modes are anyons obeying non-Abelian exchange statistics distinct from fermions or bosons. While significant progresses have been achieved in the past two decades in searching for these exotic excitations in solid-state systems, their non-Abelian nature remains unverified, as definitive proof requires braiding operations. Here, we report preliminarily experimental advances in creating, manipulating, and exchanging the presumed Majorana zero modes in an envelope-shaped Josephson device composed of multiple trijunctions on a topological insulator surface. We observed the signatures of in-gap states migration consistent with the expectations of the Fu-Kane model, supporting the realization of an exchange operation. This work would establish a critical pathway toward ultimately braiding Majorana zero modes in the Fu-Kane scheme of topological quantum computation.

cond-mat.mes-hall↗

Multi-Granularity Adaptive Time-Frequency Attention Framework for Audio Deepfake Detection under Real-World Communication Degradations

The rise of highly convincing synthetic speech poses a growing threat to audio communications. Although existing Audio Deepfake Detection (ADD) methods have demonstrated good performance under clean conditions, their effectiveness drops significantly under degradations such as packet losses and speech codec compression in real-world communication environments. In this work, we propose the first unified framework for robust ADD under such degradations, which is designed to effectively accommodate multiple types of Time-Frequency (TF) representations. The core of our framework is a novel Multi-Granularity Adaptive Attention (MGAA) architecture, which employs a set of customizable multi-scale attention heads to capture both global and local receptive fields across varying TF granularities. A novel adaptive fusion mechanism subsequently adjusts and fuses these attention branches based on the saliency of TF regions, allowing the model to dynamically reallocate its focus according to the characteristics of the degradation. This enables the effective localization and amplification of subtle forgery traces. Extensive experiments demonstrate that the proposed framework consistently outperforms state-of-the-art baselines across various real-world communication degradation scenarios, including six speech codecs and five levels of packet losses. In addition, comparative analysis reveals that the MGAA-enhanced features significantly improve separability between real and fake audio classes and sharpen decision boundaries. These results highlight the robustness and practical deployment potential of our framework in real-world communication environments.

eess.AS↗

Procedure of tuning up a three-site artificial Kitaev chain based on transmon measurements

Artificial Kitaev chains (AKCs), formed of quantum dot-superconductor linear arrays, provide a promising platform for hosting Majorana bound states (MBSs) and implementing topological quantum computing. The main challenges along this research direction would include the tuning up of AKCs for hosting MBSs and the readout of the parity of the chains. In this work, we present a step-by-step procedure for tuning up a three-site AKC to its sweet spots based on the spectra of a transmon circuit which is integrated with the chain for the purpose of reading out the parity of the chain. The signatures of the transmon's plasma modes in each step, particular those related to the appearance of MBSs in the chain, will be given. We find that the sweet spots in a three-site AKC can be classified into three types based on the relative strengths of elastic cotunneling (ECT) and crossed Andreev reflection (CAR): ECT-dominated sweet spots, genuine sweet spots and CAR-dominated sweet spots. We show that the ECT-dominated and CAR-dominated sweet spots can be more conveniently accessed and utilized in transmon-based measurements.

cond-mat.mes-hall↗

Measurement of parity-dependent energy-phase relation of the low-energy states in a potential artificial Kitaev chain utilizing a transmon qubit

Artificial Kitaev chains have emerged as a promising platform for realizing topological quantum computing. Once the chains are formed and the Majorana zero modes are braided/fused, reading out the parity of the chains is essential for further verifying the non-Abelian property of the Majorana zero modes. Here we demonstrate the feasibility of using a superconducting transmon qubit, which incorporates an end of a four-site quantum dot-superconductor chain based on a Ge/Si nanowire, to directly detect the singlet/doublet state, and thus the parity of the entire chain. We also demonstrate that for multiple-dot chains there are two types of 0-π transitions between different charging states: the parity-flip 0-π transition and the parity-preserved 0-π transition. Furthermore, we show that the inter-dot coupling, hence the strengths of cross Andreev reflection and elastic cotunneling of electrons, can be adjusted by local electrostatic gating in chains fabricated on Ge/Si core-shell nanowires. Our exploration would be helpful for the ultimate realization of topological quantum computing based on artificial Kitaev chains.

cond-mat.mes-hall↗

Controllable creation of topological boundary states in topological-insulator-based Josephson corner junctions

Majorana zero modes (MZMs) in condensed matter systems have attracted great attention in the past two decades, due to their interesting physics and potential application in topological quantum computing (TQC). However, the topologically protected nature of MZMs still need more experimental verifications. In this study, we have realized controllable creation of a topological boundary state at the corner of topological insulator (TI)-based Josephson corner junctions. This state demonstrates protected existence across a broad region in parametric space, and exhibits a non-2π-period but 4π-period-compatible energy-phase relation. Our study suggests that TI-based Josephson junctions, as proposed in the Fu-Kane scheme of TQC, may provide a promising platform for hosting and braiding MZMs.

cond-mat.mes-hall↗

EDGE: Efficient Data Selection for LLM Agents via Guideline Effectiveness

Large Language Models (LLMs) have shown remarkable capabilities as AI agents. However, existing methods for enhancing LLM-agent abilities often lack a focus on data quality, leading to inefficiencies and suboptimal results in both fine-tuning and prompt engineering. To address this issue, we introduce EDGE, a novel approach for identifying informative samples without needing golden answers. We propose the Guideline Effectiveness (GE) metric, which selects challenging samples by measuring the impact of human-provided guidelines in multi-turn interaction tasks. A low GE score indicates that the human expertise required for a sample is missing from the guideline, making the sample more informative. By selecting samples with low GE scores, we can improve the efficiency and outcomes of both prompt engineering and fine-tuning processes for LLMs. Extensive experiments validate the performance of our method. Our method achieves competitive results on the HotpotQA and WebShop and datasets, requiring 75\% and 50\% less data, respectively, while outperforming existing methods. We also provide a fresh perspective on the data quality of LLM-agent fine-tuning.

cs.LG↗

Optical interference by amplitude measurement

Interference effects are usually observed by intensity measurement. Path indistinguishability by quantum complementarity principle requires projection of the interfering fields into a common indistinguishable mode before detection. On the other hand, the essence of wave interference is the addition of amplitudes of the interfering fields. Therefore, if amplitudes can be directly measured and added, interference can occur even though the interfering fields are in well-distinguishable modes. Here, we make a comprehensive study in both theory and experiment of a technique by homodyne measurement of field amplitudes to reveal interference. This works for both classical and quantum fields even though there exists distinguishability in the interfering paths of light. This directly challenges complementarity principle. We present a resolution of this issue from the viewpoint of measurement that emphasizes either particle or wave. This technique is particularly useful for recovering interference in unbalanced interferometers with path-imbalance beyond coherence length of the input field and can be applied to remote sensing to extend applicable range. Since the amplitude-based interference phenomena studied here are fundamentally different from the traditional intenisty-based interference phenomena, our approach leads to a new paradigm to study coherence between optical fields.

quant-ph↗

ITS: Implicit Thin Shell for Polygonal Meshes

In computer graphics, simplifying a polygonal mesh surface~$\mathcal{M}$ into a geometric proxy that maintains close conformity to~$\mathcal{M}$ is crucial, as it can significantly reduce computational demands in various applications. In this paper, we introduce the Implicit Thin Shell~(ITS), a concept designed to implicitly represent the sandwich-walled space surrounding~$\mathcal{M}$, defined as~$\{\textbf{x}\in\mathbb{R}^3|ε_1\leq f(\textbf{x}) \leq ε_2, ε_1< 0, ε_2>0\}$. Here, $f$ is an approximation of the signed distance function~(SDF) of~$\mathcal{M}$, and we aim to minimize the thickness~$ε_2-ε_1$. To achieve a balance between mathematical simplicity and expressive capability in~$f$, we employ a tri-variate tensor-product B-spline to represent~$f$. This representation is coupled with adaptive knot grids that adapt to the inherent shape variations of~$\mathcal{M}$, while restricting~$f$'s basis functions to the first degree. In this manner, the analytical form of~$f$ can be rapidly determined by solving a sparse linear system. Moreover, the process of identifying the extreme values of~$f$ among the infinitely many points on~$\mathcal{M}$ can be simplified to seeking extremes among a finite set of candidate points. By exhausting the candidate points, we find the extreme values~$ε_1<0$ and $ε_2>0$ that minimize the thickness. The constructed ITS is guaranteed to wrap~$\mathcal{M}$ rigorously, without any intersections between the bounding surfaces and~$\mathcal{M}$. ITS offers numerous potential applications thanks to its rigorousness, tightness, expressiveness, and computational efficiency. We demonstrate the efficacy of ITS in rapid inside-outside tests and in mesh simplification through the control of global error.

cs.GR↗