SearcharxivSearch

arXiv subjects

Xi Shao

Publications and source records attributed to Xi Shao.

18 recordsLinked to original sources

CHARM: Charge Calibration and Acoustic Rescue for LLM-based Multimodal Sarcasm Detection

Sarcasm detection, the identification of discrepancies between literal and intended meaning, is a fundamental task in affective computing. However, zero-shot instruction-tuned Large Language Models (LLMs) systematically over-predict the positive (sarcastic) class across the entire capability spectrum, while the prosodic cues humans rely on remain underexploited and transfer unevenly across languages. We introduce CHARM (Charge Calibration and Acoustic Rescue for Multimodal Sarcasm Detection), a training-free framework that couples two modules. Bidirectional Charge Calibration (BiCAL) steers the LLM toward opposing sarcastic and literal verdicts along a symmetric axis of charged prompts; the induced directional biases cancel by construction, and a simple aggregation recovers an unbiased pragmatic signal. Acoustic Late-Fusion Rescue (ALFR) then fuses the calibrated votes with prosodic descriptors and LLM-generated auditory-perception probes through a shallow classifier, actively down-weighting saturated text votes in favour of acoustic evidence. Without fine-tuning any backbone, BiCAL attains the highest reported zero-shot text-only Macro-F1 of 0.787 on MUStARD, while ALFR lifts weak backbones by up to +0.382 Macro-F1 on CMMA. A Stouffer meta-analysis confirms statistical significance on MUStARD and CMMA (Z = 13.89 and Z = 34.64, respectively; p < 10^-43). Our analysis further uncovers a cross-cultural prosodic decoupling: low-level acoustics fail to transfer across languages, whereas high-level perceptual abstractions remain robust. Together, these components yield an explainable, cross-lingual multimodal detector.

cs.SD

Towards Event-Robust Acoustic Scene Classification

This paper introduces the Event-Shifted Acoustic Scene (ESAS) dataset, a novel benchmark for evaluating the robustness of Acoustic Scene Classification (ASC) systems against unknown sound events. Existing ASC datasets typically contain recordings of clean and consistent audio, while real-world environments often include diverse and unexpected sound events. To bridge this gap, ESAS simulates real-world acoustic variability by injecting foreground sound events into background scenes with the assistance of large language models. In this work, we present the construction methodology, dataset statistics, and evaluation protocols. Furthermore, a comprehensive evaluation of state-of-the-art ASC systems is conducted using the ESAS benchmark. Experimental results reveal that existing ASC models suffer significant performance degradation when facing the event-shift challenge. The introduction of the ESAS dataset aims to drive future research toward event-robust ASC.

cs.SD

Word-Anchored Temporal Forgery Localization

Current temporal forgery localization (TFL) approaches typically rely on temporal boundary regression or continuous frame-level anomaly detection paradigms to derive candidate forgery proposals. However, they suffer not only from feature granularity misalignment but also from costly computation. To address these issues, we propose word-anchored temporal forgery localization (WAFL), a novel paradigm that shifts the TFL task from temporal regression and continuous localization to discrete word-level binary classification. Specifically, we first analyze the essence of temporal forgeries and identify the minimum meaningful forgery units, word tokens, and then align data preprocessing with the natural linguistic boundaries of speech. To adapt powerful pre-trained foundation backbones for feature extraction, we introduce the forensic feature realignment (FFR) module, mapping representations from the pre-trained semantic space to a discriminative forensic manifold. This allows subsequent lightweight linear classifiers to efficiently perform binary classification and accomplish the TFL task. Furthermore, to overcome the extreme class imbalance inherent to forgery detection, we design the artifact-centric asymmetric (ACA) loss, which breaks the standard precision-recall trade-off by dynamically suppressing overwhelming authentic gradients while asymmetrically prioritizing subtle forensic artifacts. Extensive experiments demonstrate that WAFL significantly outperforms state-of-the-art approaches in localization performance under both in- and cross-dataset settings, while requiring substantially fewer learnable parameters and operating at high computational efficiency.

cs.CV

Learning Time-Graph Frequency Representation for Monaural Speech Enhancement

The Graph Fourier Transform (GFT) has recently demonstrated promising results in speech enhancement. However, existing GFT-based speech enhancement approaches often employ fixed graph topologies to build the graph Fourier basis, whose the representation lacks the adaptively and flexibility. In addition, they suffer from the numerical errors and instability introduced by matrix inversion in GFT based on both Singular Value Decomposition (GFT-SVD) and Eigen Vector Decomposition (GFT-EVD). Motivated by these limitations, this paper propose a simple yet effective learnable GFT-SVD framework for speech enhancement. Specifically, we leverage graph shift operators to construct a learnable graph topology and define a learnable graph Fourier basis by the singular value matrices using 1-D convolution (Conv-1D) neural layer. This eliminates the need for matrix inversion, thereby avoiding the associated numerical errors and stability problem.

eess.AS

Superconductivity at 22.3 K in Compressed Sodium-intercalated Graphite

Graphite intercalation compounds (GICs) have long been recognized as promising candidates for high-temperature superconductivity by intercalation or charge doping, yet experimental progress has stalled with transition temperatures (Tc) limited to 11.5 K at ambient pressure and 15.1 K at 7.5 GPa in calcium-intercalated graphite over decades. Here, we report robust superconductivity in sodium-intercalated graphite with Tc of 22.3 K, as demonstrated by clear zero-resistance behavior. Our approach involves simply room-temperature grinding of graphite with sodium, followed by slight compression up to 7.1 GPa, circumventing complex synthesis procedures. Through synchrotron X-ray diffraction combined with first-principles calculations, we identify the major superconducting phase as an orthorhombic stage-2 GIC structure with slightly over-stoichiometric composition (Na1+xC8). Electron-phonon coupling calculations reveal that superconductivity primarily emerges from the interactions between out-of-plane carbon electrons and low-frequency Na/C vibrations.The enhancement in Tc establishes sodium as superior for achieving higher-Tc in GICs and illustrates promising pathway for further optimization through compositional and structural tuning.

cond-mat.supr-con

Improving Acoustic Scene Classification with City Features

Acoustic scene recordings are often collected from a diverse range of cities. Most existing acoustic scene classification (ASC) approaches focus on identifying common acoustic scene patterns across cities to enhance generalization. However, the potential acoustic differences introduced by city-specific environmental and cultural factors are overlooked. In this paper, we hypothesize that the city-specific acoustic features are beneficial for the ASC task rather than being treated as noise or bias. To this end, we propose City2Scene, a novel framework that leverages city features to improve ASC. Unlike conventional approaches that may discard or suppress city information, City2Scene transfers the city-specific knowledge from pre-trained city classification models to scene classification model using knowledge distillation. We evaluate City2Scene on three datasets of DCASE Challenge Task 1, which include both scene and city labels. Experimental results demonstrate that city features provide valuable information for classifying scenes. By distilling city-specific knowledge, City2Scene effectively improves accuracy across a variety of lightweight CNN backbones, achieving competitive performance to the top-ranked solutions of DCASE Challenge in recent years.

cs.SD

The Spectral Behaviour and Variability of Narrow-line Seyfert 1 Galaxies with Australia Telescope Compact Array Observations

We present multi-frequency radio data for a sample of narrow-line Seyfert 1 galaxies. We first focus on the sub-class of gamma-ray emitting narrow-line Seyfert 1 galaxies, studying the long-term radio variability of five sources and comparing it to their gamma-ray state. We then extend the observations of the southern narrow-line Seyfert 1 galaxy sample of Chen et al. by observing several candidate narrow-line Seyfert 1 sources for the first time, and re-observing several other gamma-ray quiet sources to obtain a first indication of their radio variability. We find that the gamma-ray emitting narrow-line Seyfert 1 galaxies are highly variable radio emitters and that there are instances of contemporaneous flaring activity between the radio and gamma-ray band (PKS 0440$-$00, PMN J0948+0022 and PKS 1244$-$255). However, there are also cases of significant radio outbursts without gamma-ray counterparts (PMN J0948+0022 and PKS 2004$-$447). The five gamma-ray NLS1s favour flat or inverted radio spectra, although the spectral indices vary significantly over time. For the gamma-ray quiet sample, the difference between the previous observations at 5.5 GHz and new ATCA observations indicates that over half of the 14 sources exhibit apparent variability. In contrast to gamma-ray loud sources, gamma-ray quiet objects tend to have steep spectra especially in the lower radio band (887.5$-$1367.5 MHz), with a number of the variable sources having flatter spectra at higher radio frequencies.

astro-ph.GA

Leveraging Self-supervised Audio Representations for Data-Efficient Acoustic Scene Classification

Acoustic scene classification (ASC) predominantly relies on supervised approaches. However, acquiring labeled data for training ASC models is often costly and time-consuming. Recently, self-supervised learning (SSL) has emerged as a powerful method for extracting features from unlabeled audio data, benefiting many downstream audio tasks. This paper proposes a data-efficient and low-complexity ASC system by leveraging self-supervised audio representations extracted from general-purpose audio datasets. We introduce BEATs, an audio SSL pre-trained model, to extract the general representations from AudioSet. Through extensive experiments, it has been demonstrated that the self-supervised audio representations can help to achieve high ASC accuracy with limited labeled fine-tuning data. Furthermore, we find that ensembling the SSL models fine-tuned with different strategies contributes to a further performance improvement. To meet low-complexity requirements, we use knowledge distillation to transfer the self-supervised knowledge from large teacher models to an efficient student model. The experimental results suggest that the self-supervised teachers effectively improve the classification accuracy of the student model. Our best-performing system obtains an average accuracy of 56.7%.

cs.SD

STAA-Net: A Sparse and Transferable Adversarial Attack for Speech Emotion Recognition

Speech contains rich information on the emotions of humans, and Speech Emotion Recognition (SER) has been an important topic in the area of human-computer interaction. The robustness of SER models is crucial, particularly in privacy-sensitive and reliability-demanding domains like private healthcare. Recently, the vulnerability of deep neural networks in the audio domain to adversarial attacks has become a popular area of research. However, prior works on adversarial attacks in the audio domain primarily rely on iterative gradient-based techniques, which are time-consuming and prone to overfitting the specific threat model. Furthermore, the exploration of sparse perturbations, which have the potential for better stealthiness, remains limited in the audio domain. To address these challenges, we propose a generator-based attack method to generate sparse and transferable adversarial examples to deceive SER models in an end-to-end and efficient manner. We evaluate our method on two widely-used SER datasets, Database of Elicited Mood in Speech (DEMoS) and Interactive Emotional dyadic MOtion CAPture (IEMOCAP), and demonstrate its ability to generate successful sparse adversarial examples in an efficient manner. Moreover, our generated adversarial examples exhibit model-agnostic transferability, enabling effective adversarial attacks on advanced victim models.

cs.SD

Does a radio jet drive the massive multi-phase outflow in the ultra-luminous infrared galaxy IRAS 10565+2448?

We present new upgraded Giant Metrewave Radio Telescope (uGMRT) HI 21-cm observations of the ultra-luminous infrared galaxy IRAS 10565+2448, previously reported to show blueshifted, broad, and shallow HI absorption indicating an outflow. Our higher spatial resolution observations have localised this blueshifted outflow, which is $\sim$ 1.36 kpc southwest of the radio centre and has a blueshifted velocity of $\sim 148\,\rm km\,s^{-1}$ and a full width at half maximum (FWHM) of $\sim 581\,\rm km\,s^{-1}$. The spatial extent and kinematic properties of the HI outflow are consistent with the previously detected cold molecular outflows in IRAS 10565+2448, suggesting that they likely have the same driving mechanism and are tracing the same outflow. By combining the multi-phase gas observations, we estimate a total outflowing mass rate of at least $140\, \rm M_\odot \,yr^{-1}$ and a total energy loss rate of at least $8.9\times10^{42}\,\rm erg\,s^{-1}$, where the contribution from the ionised outflow is negligible, emphasising the importance of including both cold neutral and molecular gas when quantifying the impact of outflows. We present evidence of the presence of a radio jet and argue that this may play a role in driving the observed outflows. The modest radio luminosity $L_{\rm1.4GHz}$ $\sim1.3\times10^{23}\,{\rm W\,Hz^{-1}}$ of the jet in IRAS 10565+2448 implies that the jet contribution to driving outflows should not be ignored in low radio luminosity AGN.

astro-ph.GA

The radio structure of the $\gamma$-ray narrow-line Seyfert 1 galaxy SDSS J211852.96$-$073227.5

The $\gamma$-ray narrow-line Seyfert 1 (NLS1) galaxies can be considered to be the third class of $\gamma$-ray active galactic nuclei possessing relativistic jets. In this paper, we present multi-band high resolution Very Long Baseline Array (VLBA) images of the $\gamma$-ray NLS1, SDSS J211852.96$-$073227.5 (J2118$-$0732, $z=0.26$). We find a core-jet radio morphology and significant flux density variations in the radio core. The high brightness temperature estimated from VLBA images and core variability demonstrate that it exhibits substantial relativistic beaming effects. From considering radio emission in several bands, we find that the source has an inverted spectrum above 1 GHz but a steep spectrum at low frequencies from 74 MHz to 1 GHz; these may arise from the present activity and the old diffuse/extended emission, respectively. The core-jet morphology, significant flux density variations, and beaming effect make J2118$-$0732 resemble a blazar. Considering the low mass of its central black hole and ongoing merger environment, J2118$-$0732 may represent a low-mass, low-power counterpart of blazars, and may finally evolve to a blazar.

astro-ph.HE

Unusual phase transition of layer-stacked borophene under pressure

The 8-Pmmn borophene, a boron analogue of graphene, hosts tilted and anisotropic massless Dirac fermion quasiparticles owing to the presence of the distorted graphene-like sublattice. First-principles calculations show that the stacked 8-Pmmn borophene is transformed into the fused three-dimensional borophene under pressure, being accompanied by the partially bond-breaking and bond-reforming. Strikingly, the fused 8-Pmmn borophene inherits the Dirac band dispersion resulting in an unusual semimetal-semimetal transition. A simple tight-binding model derived from graphene qualitatively reveals the underlying physics due to the maximum preservation of graphene-like substructure after the phase transition, which contrasts greatly to the transformation of graphite into diamond associated with the semimetal-insulator transition.

cond-mat.mtrl-sci

An Encoder-Decoder Based Audio Captioning System With Transfer and Reinforcement Learning

Automated audio captioning aims to use natural language to describe the content of audio data. This paper presents an audio captioning system with an encoder-decoder architecture, where the decoder predicts words based on audio features extracted by the encoder. To improve the proposed system, transfer learning from either an upstream audio-related task or a large in-domain dataset is introduced to mitigate the problem induced by data scarcity. Besides, evaluation metrics are incorporated into the optimization of the model with reinforcement learning, which helps address the problem of ``exposure bias'' induced by ``teacher forcing'' training strategy and the mismatch between the evaluation metrics and the loss function. The resulting system was ranked 3rd in DCASE 2021 Task 6. Ablation studies are carried out to investigate how much each element in the proposed system can contribute to final performance. The results show that the proposed techniques significantly improve the scores of the evaluation metrics, however, reinforcement learning may impact adversely on the quality of the generated captions.

eess.AS

Helium Induced Nitrogen Salt at High Pressure

The energy landscape of helium-nitrogen mixtures is explored by ab initio evolutionary searches, which predicted several stable helium-nitrogen compounds in the pressure range from 25 to 100 GPa. In particular, the monoclinic structure of HeN$_{22}$ consists of neutral He atoms, partially ionic dimers N$_{2}$$^{δ-}$, and lantern-like cages N$_{20}$$^{δ+}$. The presence of helium not only greatly enhances structural diversity of nitrogen solids, but also tremendously lowers the formation pressure of nitrogen salt. The unique nitrogen framework of (HeN$_{20}$)$^{δ+}$N$_{2}$$^{δ-}$ may be quenchable to ambient pressure even after removing helium. The estimated energy density of N$_{20}$$^{δ+}$N$_{2}$$^{δ-}$ (10.44 kJ/g) is $\sim$2.4 times larger than that of trinitrotoluene (TNT), indicating a very promising high-energy-density material.

cond-mat.mtrl-sci

Curvature induced polarization and spectral index behavior for PKS 1502+106

A comprehensive study of multifrequency correlations can shed light on the nature of variation for blazars. In this work, we collect the long-term radio, optical and $γ$-ray light curves of PKS 1502+106. After performing the localized cross-correlation function analysis, we find that correlations between radio and $γ$-ray or $V$ band are beyond the $3σ$ significance level. The lag of the $γ$-ray relative to 15 GHz is $-60^{+5}_{-10}$ days, translating to a distance $3.18^{+0.50}_{-0.27}$ parsec (pc) between them. Within uncertainties, the locations of the $γ$-ray and optical emitting regions are roughly the same, and are away from the jet base within $1.2$ pc. The derived magnetic field in optical and $γ$-ray emitting regions is about $0.36$ G. The logarithm of $γ$-ray flux is significantly linearly correlated with that of $V$ band fluxes, which can be explained by the synchrotron self-Compton (SSC) process, the external Compton (EC) processes, or the combination of them. We find a significant linear correlation in the plot of $\log\prod$ (polarization degree) versus $\log νF_ν$ at $V$ band, and use the empirical relation $Π\sim \sin^n θ'$ ($θ'$ is the observing angle in the comoving frame blob) to explain it. The behaviors of color index (generally redder when brighter at the active state) and $γ$-ray spectral index (softer when brighter) could be well explained by the twisted jet model. These findings suggest that the curvature effect (mainly due to the change of the viewing angle) is dominant in the variation phenomena of fluxes, spectral indices, and polarization degrees for PKS 1502+106.

astro-ph.HE

Formation of copper boride on Cu(111)

Boron forms compounds with nearly all metals, with notable exception of copper and other group IB and IIB elements. Here, we report an unexpected discovery of ordered copper boride grown epitaxially on Cu(111) under ultrahigh vacuum. Scanning tunneling microscopy experiments combined with ab initio evolutionary structure prediction reveal a remarkably complex structure of 2D-Cu8B14. Strong intra-layer p-d hybridization and a large amount of charge transfer between Cu and B atoms are the key factors for the emergence of copper boride. This makes the discovered material unique and opens up the possibility of synthesizing ordered low-dimensional structures in similar immiscible systems.

cond-mat.mtrl-sci

Locations of optical and $γ$-ray emitting regions in the jet of PMN J2345-1555

We collect long term $γ$-ray, optical and radio $15$ GHz light curves of quasar object PMN J2345-1555. The correlation analyses between them are performed via the local cross-correlation function (LCCF). We found that all the optical $V$, $R$ band and the infrared $J$ band are correlated with the radio 15 GHz at beyond $3σ$ significance level, and the lag times are $-221.81^{+6.26}_{-6.72}$, $-201.38^{+6.42}_{-6.02}$ and $-192.27^{+8.26}_{-7.37}$ days, respectively. The $γ$-ray is strongly correlated with optical, but weakly correlated with the radio. We present that time lags between different frequencies can be used as an alternative parameter to derive the core-shift measurement. For this target, the magnetic field and particle density at 1 parsec in jet are derived to be $0.61$ Gauss and $1533/γ_{\rm min}$ cm$^{-3}$, respectively. The black hole mass and the 15 GHz core position in jet are estimated to be $10^{8.44} {\rm M}_{\odot}$ and $30$ parsec, respectively. The lag times enable us to derive that the optical and the $γ$-ray emitting regions coincide, which are located at $4.26^{+0.83}_{-0.79}$ pc away from 15 GHz core position in jet and beyond the broad line region (BLR). We found that a $3σ$ correlation between the color index and the radio light curve, which indicates that opacity may play an important role in the variation. The $δV-δR$ behaviors are complex, while the $R-J$ shows a bluer when brighter trend. As hinted from radio images, we proposed a positional dependent spectral index model to explain the color index behaviors, which is complementary for the shock in jet model. The curvature effects and contribution from accretion disk may also affect variables of blazars in many aspects.

astro-ph.HE

Unexpected Reconstruction of the alpha-Boron (111) Surface

We report on a novel reconstruction of the alpha-boron (111) surface, discovered using an ab initio evolution structure search, and reveal that it has an unexpected neat structure and much lower surface energy than the recently proposed (111)-I_R,(a) surface. For this reconstruction, every single interstitial boron atom forms bridges with the unique polar-covalent bonds between neighboring B_12 icosahedra, which perfectly meet the electron counting rule and are responsible for the reconstruction-induced metal-semiconductor transition. The peculiar charge transfer between the interstitial atoms and the icosahedra plays an important role in stabilizing the surface.

cond-mat.mtrl-sci