SearcharxivSearch

arXiv subjects

Yuhang Xiao

Publications and source records attributed to Yuhang Xiao.

9 recordsLinked to original sources

Electronic Toroidal Metals: Landau Theory and Magnetoelectric Fingerprints

Toroidal order is a higher-rank multipolar order whose intrinsic realization in itinerant-electron systems remains unexplored. Here, we develop a generic theory of the electronic toroidal metal (ETM), in which toroidal order emerges spontaneously from electronic degrees of freedom near the Fermi surfaces. Because candidate toroidal bilinears can overlap by symmetry with the uniform charge current, we formulate a projected instability criterion that removes the noncondensable current component. Within this current-orthogonal sector, we identify a soft collective mode in the $\mathcal{P}$-odd and $\mathcal{T}$-odd particle-hole channel and demonstrate that ETM arises as a Fermi-liquid instability. We then define the toroidal moment in an itinerant system through the antisymmetric magnetoelectric response tensor. We further establish an intimate connection between this response and the topology of pseudospin texture: the rearrangement of pseudospin vortices changes the winding structure of the Fermi surfaces and markedly enhances the magnetoelectric response. ETM also exhibits characteristic nonlinear charge transport, including a pronounced enhancement of its interband quantum-geometric contribution. Our results establish ETM as a distinct metallic Landau phase and provide a general framework for understanding toroidal order generated by itinerant electrons.

cond-mat.str-el

Bootstrapping Vision-Language Model for Hysteroscopic Surgical Scene Segmentation

Hysteroscopic surgical scene segmentation plays a pivotal role in understanding the hysteroscopic intraoperative environment as well as computer-assisted intervention. However, this task presents unique challenges due to the high morphological similarity among different lesions and the presence of artifacts such as specular reflections, motion blur, and fluid occlusions in surgical videos. In this work, we propose the first vision-language model (VLM)-based hysteroscopic surgical scene segmentation method, which performs pixel-wise localization for fifteen representative categories in hysteroscopic surgical scenes. Our VLM-hyster has a segmentation backbone that utilizes the pretrained image encoder for robust visual feature extraction, coupled with a transformer-based decoder for dense prediction. Moreover, we design category-specific text prompts and incorporate a masked distillation branch to filter out visual features with low correlation to the text prompts, enabling the model to focus more effectively on category-specific image regions and thereby enhancing segmentation performance. We collect a large multicentric hysteroscopic surgical scene dataset, containing 4,020 high-resolution images with detailed mask annotations, for model training and evaluation. Experimental results demonstrate that VLM-hyster substantially outperforms state-of-the-art AI models. Furthermore, extensive assessments by gynecologists, as well as multicentre and prospective validations, demonstrate VLM-hyster's robustness and generalizability. The results suggest that VLM-hyster earns considerable potential in enabling AI-assisted localization of surgical instruments and lesions in hysteroscopic surgeries. Code is available at https://github.com/viscom-tongji/VLM-hyster.

cs.CV

STSR: High-Fidelity Speech Super-Resolution via Spectral-Transient Context Modeling

Speech super-resolution (SR) reconstructs high-fidelity wideband speech from low-resolution inputs-a task that necessitates reconciling global harmonic coherence with local transient sharpness. While diffusion-based generative models yield impressive fidelity, their practical deployment is often stymied by prohibitive computational demands. Conversely, efficient time-domain architectures lack the explicit frequency representations essential for capturing long-range spectral dependencies and ensuring precise harmonic alignment. We introduce STSR, a unified end-to-end framework formulated in the MDCT domain to circumvent these limitations. STSR employs a Spectral-Contextual Attention mechanism that harnesses hierarchical windowing to adaptively aggregate non-local spectral context, enabling consistent harmonic reconstruction up to 48 kHz. Concurrently, a sparse-aware regularization strategy is employed to mitigate the suppression of transient components inherent in compressed spectral representations. STSR consistently outperforms state-of-the-art baselines in both perceptual fidelity and zero-shot generalization, providing a robust, real-time paradigm for high-quality speech restoration.

cs.SD

CogSR: Semantic-Aware Speech Super-Resolution via Chain-of-Thought Guided Flow Matching

Applying speech super-resolution (SR) to recordings with severely low sampling rates is a critical challenge in digital archiving and investigative audio recovery. In these scenarios, the input lacks essential acoustic cues. Consequently, existing generative models often fail; without sufficient context, they hallucinate phonetic content, guessing words based on probability rather than meaning. To address this, we propose CogSR, a framework designed specifically for high-precision, offline restoration. Our approach shifts the focus from simple signal mapping to cognitive reconstruction. By integrating a Large Audio-Language Model, we employ Chain-of-Thought reasoning to act as a semantic anchor, while explicit acoustic priors ensure the speaker's identity remains consistent. This guides a Rectified Flow backbone to synthesize high-frequency details that are not only realistic but linguistically accurate. Evaluations show that CogSR effectively eliminates ambiguity in severe degradation regimes, making it a robust solution for restoring high-value legacy and surveillance audio.

cs.SD

DreamDistribution: Learning Prompt Distribution for Diverse In-distribution Generation

The popularization of Text-to-Image (T2I) diffusion models enables the generation of high-quality images from text descriptions. However, generating diverse customized images with reference visual attributes remains challenging. This work focuses on personalizing T2I diffusion models at a more abstract concept or category level, adapting commonalities from a set of reference images while creating new instances with sufficient variations. We introduce a solution that allows a pretrained T2I diffusion model to learn a set of soft prompts, enabling the generation of novel images by sampling prompts from the learned distribution. These prompts offer text-guided editing capabilities and additional flexibility in controlling variation and mixing between multiple distributions. We also show the adaptability of the learned prompt distribution to other tasks, such as text-to-3D. Finally we demonstrate effectiveness of our approach through quantitative analysis including automatic evaluation and human assessment. Project website: https://briannlongzhao.github.io/DreamDistribution

cs.CV

Behavioral Bias of Vision-Language Models: A Behavioral Finance View

Large Vision-Language Models (LVLMs) evolve rapidly as Large Language Models (LLMs) was equipped with vision modules to create more human-like models. However, we should carefully evaluate their applications in different domains, as they may possess undesired biases. Our work studies the potential behavioral biases of LVLMs from a behavioral finance perspective, an interdisciplinary subject that jointly considers finance and psychology. We propose an end-to-end framework, from data collection to new evaluation metrics, to assess LVLMs' reasoning capabilities and the dynamic behaviors manifested in two established human financial behavioral biases: recency bias and authority bias. Our evaluations find that recent open-source LVLMs such as LLaVA-NeXT, MobileVLM-V2, Mini-Gemini, MiniCPM-Llama3-V 2.5 and Phi-3-vision-128k suffer significantly from these two biases, while the proprietary model GPT-4o is negligibly impacted. Our observations highlight directions in which open-source models can improve. The code is available at https://github.com/mydcxiao/vlm_behavioral_fin.

cs.CL

Quantum Geometric Effects on the Higgs Mode in Flat-band Superconductors

Flat-band systems are of great interest due to their strong electron correlations and unique band geometry. Recent studies have linked the properties of Cooper pairs in flat-band superconductors to the quantum metric. Unlike prior studies primarily focused on quasiparticle fluctuations, in this work, we investigate quantum geometric effects on a collective mode of the order parameter-the Higgs mode. We derive the quantum goemetric Higgs-mode correlation length and investigate whether the nonlinear electromagnetic response of the Higgs mode persists in flat band. It turns out that the quantum metric will replaces energy derivatives, playing a crucial role in third-harmonic generation(THG). We further numerically calculated the correlation length and THG in the Lieb lattice. In contrast to traditional single-band results, Higgs mode fluctuations contribute almost entirely to THG, with the quasiparticle contribution being negligible. This finding provides valuable guidance for detecting Higgs modes in flat-band systems through optical methods and reveals the profound influence of quantum geometry in flat-band systems.

cond-mat.supr-con

A generic model with unconventional Rashba bands and giant spin galvanic effect

In two-dimensional system, Rashba spin-orbit coupling can lift spin degeneracy and gives the opposite spin chirality of two split Fermi circles from two Rashba bands. Here, we propose a generic model which can produce unconventional Rashba bands. In such a case, the two Fermi circles from two bands have the same spin chirality. When various interactions are taken into account, many unique physics can emerge in case of unconventional Rashba bands in comparison with in case of conventional Rashba bands. For instance, we study the spin galvanic effect by considering two cases with potential impurity scattering and magnetic impurity scattering, respectively. In both cases, we find the efficiency of spin galvanic effect is strongly enhanced in unconventional Rashba bands in comparison with conventional Rashba bands. More intriguingly, we find the effeiciency of conventional Rashba bands is insensitive to potential or magnetic impurity scattering. However, such efficiency of uncoventional Rashba bands can be further enhanced by the magnetic impurity scattering in comparison with the potential impurity scattering. Thus, the unconventional Rashba bands can give giant spin galvanic effect. These results show that this model is useful to explore abnormal physics in the systems with unconventional Rashba bands.

cond-mat.mes-hall

One-bit Spectrum Sensing with the Eigenvalue Moment Ratio Approach

One-bit analog-to-digital converter (ADC), performing signal sampling as an extreme simple comparator, is an overwhelming technology for spectrum sensing due to its low-cost, low-power consumptions and high sampling rate. In this letter, we propose a novel one-bit sensing approach based on the eigenvalue moment ratio (EMR), which has been proved to be highly efficient for conventional multi-antenna spectrum sensing in $\infty$-bit situation. Particularly, we determine the asymptotic distribution of one-bit EMR under null hypothesis via the central limited theorem (CLT), allowing us to perform spectrum sensing with one-bit samples directly. Theoretical and simulation analysis show the new approach can provide reasonably good sensing performance at a low hardware cost.

cs.IT