SearcharxivSearch

arXiv subjects

Xiaoying Huang

Publications and source records attributed to Xiaoying Huang.

8 recordsLinked to original sources

Channeling defect emission through topological Jackiw-Rebbi states in hBN

When two photonic structures of different topology are joined, an optical mode appears at their junction. Its existence depends only on this topological difference and not on the precise dimensions of either structure. Here we realize the photonic analogue of a Jackiw-Rebbi (JR) state in hexagonal boron nitride (hBN) hosting negatively charged boron vacancies, optically addressable spin defects. Two partially etched gratings, in which the order of the radiating and dark band-edge modes is reversed, are joined side by side. Angle-resolved reflectivity and photoluminescence reveal this reversal and a non-dispersive JR state at the junction. The state proves remarkably tolerant: despite fabrication errors of up to 20 nm, its wavelength lies within a 1 nm of the design, as predicted by a systematic numerical tolerance analysis. JR gratings therefore offer a robust, monolithic platform for controlling the emission of luminescent defects.

physics.optics

AT-ADD: A Benchmark and Challenge for Robust and All-Type Audio Deepfake Detection

Recent audio generation models can synthesize high-fidelity speech, environmental sound, singing voice, and music, creating new risks for multimedia trust. Existing audio deepfake detection (ADD) benchmarks remain predominantly speech-centric and often underrepresent realistic channel variation and diverse audio types. This paper presents AT-ADD, a large-scale benchmark and challenge designed to evaluate both robust speech deepfake detection and all-type audio deepfake detection. Track 1 evaluates binary speech detection under unseen generators, diverse recording conditions, signal perturbations, and replay effects. Track 2 evaluates type-agnostic real/fake detection over speech, sound, singing, and music when the audio type is unknown at test time. We detail the dataset construction, evaluation protocol, and reproducible baselines, and analyze the final systems submitted to the ACM Multimedia 2026 Grand Challenge. The strongest official baseline obtains 76.73% and 79.47% Macro-F1 on the Track 1 and Track 2 evaluation sets, respectively, whereas the winning challenge systems reach 90.71% and 96.10%. Beyond aggregate rankings, sample-level analysis of the top five submissions examines generator- and type-level difficulty, cross-system error complementarity, and ranking stability. The results show that large-scale self-supervised representations, condition-aware augmentation, multi-crop inference, and structured fusion or routing are central to generalization, while generator-specific robustness and consistent performance across diverse audio types remain unresolved.

cs.SD

Engineering of Dual Wavelength, Polarization Selective Metalenses in Silicon Carbide

Spin defects in silicon carbide (SiC) are promising candidates for integrated quantum photonics, offering long-lived spin states and near-infrared emission suitable for low-loss photonic integration and fibre-based quantum communication. However, light extraction from these defects remains challenging due the relatively high refractive index of SiC. Metalenses offer a compact approach to enhance light collection by engineering the wavefront directly at the material interface. Here, we design and fabricate monolithic metalenses from SiC bulk material that simultaneously operate at 860 and 1240 nm, matching with emission from the nitrogen vacancy and silicon vacancy colour centers. By independently engineering the phase response at both wavelengths, the metalens enables collection and polarization manipulation of the emitted light. We further employ the metalenses to demonstrate optically detected magnetic resonance of both defects simultaneously. These multifunctional metalenses provide a compact optical interface for scalable integrated SiC photonic devices.

quant-ph

AT-ADD: All-Type Audio Deepfake Detection Challenge Summary

This paper summarizes the ACM Multimedia 2026 AT-ADD Grand Challenge on all-type audio deepfake detection. AT-ADD contains two tracks: robust speech deepfake detection under realistic acoustic and channel variations, and type-agnostic detection over speech, environmental sound, singing voice, and music. We describe the challenge tasks, dataset and evaluation-set design, official leaderboard results, and common design patterns observed in participating systems. The best Track 1 system achieved 90.71% Macro-F1 on the final evaluation set, while the best Track 2 system achieved 96.10% Macro-F1. The final submissions show that strong systems commonly combine large-scale self-supervised audio representations, data augmentation, multi-crop inference, and structured fusion or routing. The results also reveal remaining challenges in generalization to unseen generators, robustness to realistic speech-domain distortions, and balanced performance across heterogeneous audio types.

cs.SD

StressEval: Failure-Driven Dynamic Benchmarking for Knowledge-Intensive Reasoning in Large Language Models

Static benchmarks for LLMs are increasingly compromised by contamination and overfitting especially on knowledge intensive reasoning tasks While recent dynamic benchmarks can alleviate staleness they often increase difficulty at the expense of answerability and controllability In this paper we propose StressEval a failure driven data synthesis framework that turns observed model failures into dynamic challenging and controllable test instances StressEval consists of three stages first it constructs a semi structured difficulty card that identifies the failed reasoning step and its root cause second it applies a dual perspective instance synthesis method that targets both knowledge gaps and reasoning breakdowns while preserving the underlying difficulty factors and third it applies a gating mechanism to retain only grounded unambiguous instances Seeding from multiple knowledge intensive reasoning datasets we employ StressEval to build Dynamic OneEval a focused suite of challenging dynamic benchmark Across several state of the art LLMs Dynamic OneEval yields substantially larger performance drops than the original benchmarks while retaining explicit difficulty factors enabling more actionable iteration

cs.CL

AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan

The rapid advancement of Audio Large Language Models (ALLMs) has enabled cost-effective, high-fidelity generation and manipulation of both speech and non-speech audio, including sound effects, singing voices, and music. While these capabilities foster creativity and content production, they also introduce significant security and trust challenges, as realistic audio deepfakes can now be generated and disseminated at scale. Existing audio deepfake detection (ADD) countermeasures (CMs) and benchmarks, however, remain largely speech-centric, often relying on speech-specific artifacts and exhibiting limited robustness to real-world distortions, as well as restricted generalization to heterogeneous audio types and emerging spoofing techniques. To address these gaps, we propose the All-Type Audio Deepfake Detection (AT-ADD) Grand Challenge for ACM Multimedia 2026, designed to bridge controlled academic evaluation with practical multimedia forensics. AT-ADD comprises two tracks: (1) Robust Speech Deepfake Detection, which evaluates detectors under real-world scenarios and against unseen, state-of-the-art speech generation methods; and (2) All-Type Audio Deepfake Detection, which extends detection beyond speech to diverse, unknown audio types and promotes type-agnostic generalization across speech, sound, singing, and music. By providing standardized datasets, rigorous evaluation protocols, and reproducible baselines, AT-ADD aims to accelerate the development of robust and generalizable audio forensic technologies, supporting secure communication, reliable media verification, and responsible governance in an era of pervasive synthetic audio.

cs.SD

High performance distributed feedback quantum dot lasers with laterally coupled dielectric grating

The combination of grating-based frequency-selective optical feedback mechanisms, such as distributed feedback (DFB) or distributed Bragg reflector (DBR) structures, with quantum dot (QD) gain materials is a main approach towards ultra-high-performance semiconductor lasers for many key novel applications, either as stand-alone sources or as on-chip sources in photonic integrated circuits. However, the fabrication of conventional buried Bragg grating structures on GaAs, GaAs/Si, GaSb and other material platforms have been met with major material regrowth difficulties. We report a novel and universal approach of introducing laterally coupled dielectric Bragg gratings to semiconductor lasers that allows highly controllable, reliable and strong coupling between the grating and the optical mode. We implement such a grating structure in a low-loss amorphous silicon material alongside GaAs lasers with InAs/GaAs QD gain layers. The resulting DFB laser arrays emit at pre-designed 0.8 THz LWDM frequency intervals in the 1300 nm band with record performance parameters, including side mode suppression ratios as high as 52.7 dB, continuous-wave output power of 27.7 mW (room-temperature) and 10 mW (at 70°C), and ultra-low relative intensity noise (RIN) of < -165 dB/Hz (2.5-25 GHz). The devices are also capable of operating isolator-free under very high external reflection levels of up to -12.3 dB whilst maintaining the high spectral and ultra-low RIN qualities. These results validate the novel laterally coupled dielectric grating as a technologically superior and potentially cost-effective approach for fabricating DFB and DBR lasers free of their semiconductor material constraints, thus universally applicable across different material platforms and wavelength bands.

physics.optics

A geometric method for spatiotemporal coherent structure analysis

We describe a geometric method to quantify wave patterns observed in the nervous system, which are non-stationary and with a mixture of spiral, target, plane and irregular waves. The method analyzes fluctuations of the energy angular distribution in two-dimensional Fourier spectrum of wave patterns, which reflects changes of the orientation distribution of wavefronts. We show that the number of the genuine peaks in generalized phase spectrum is close to the number of the coherent space-time clusters arising in wave patterns, and propose to use the number as a complexity measure.

nlin.PS