SearcharxivSearch

arXiv subjects

Xinyi Jiang

Publications and source records attributed to Xinyi Jiang.

At least 19 recordsLinked to original sources

M$^3$R-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding

Metaphor enables the understanding of abstract concepts through cross-domain mappings while conveying affective attitudes. In multimodal scenarios, visual and textual information jointly construct Target--Source mappings, requiring both conceptual understanding and cross-modal reasoning. However, existing benchmarks mainly evaluate metaphor understanding through isolated subtasks and lack evidence-grounded explanations, making it difficult to assess whether models establish mappings grounded in visual and textual cues.To address these limitations, we introduce M$^3$R-Bench, a unified and evidence-grounded benchmark containing 1,000 image--text instances with human-verified annotations. Guided by Conceptual Metaphor Theory and theories of nonliteral language understanding, M$^3$R-Bench provides joint annotations for metaphor occurrence, Target--Source mapping, sentiment, and stage-wise explanations following ``evidence identification--mapping establishment--sentiment inference.''Evaluations on M$^3$R-Bench reveal that existing models often overlook visual evidence, rely on superficial textual cues, and produce inaccurate Target--Source mappings, exposing a cross-modal evidence--mapping mismatch. To address this mismatch, we propose M$^3$R-Reasoner, which combines curriculum-based reasoning supervision with task-aware reinforcement learning to align model reasoning with metaphor interpretation. Experiments show that, with only an 8B-parameter backbone, M$^3$R-Reasoner outperforms larger proprietary MLLMs across four unified-task metrics and improves Visual Evidence and Sentiment Justification scores over GPT-5.5 by 28.45 and 30.11 points, respectively, while surpassing Claude-Sonnet-4.6 by 8.00 points in mean rubric score. The dataset and code are available at https://github.com/hongshi4/M3R-Bench.

cs.CL

Chemical and optical control of chiral-domain dynamics in 1T-TaS$_2$

Optical control of symmetry-breaking quantum phases is often constrained when one domain is strongly favored in equilibrium. This limitation is exemplified by the chiral charge-density-wave (CDW) order in 1T-TaS$_2$, where pristine samples predominantly select a single chirality. Here we show that free-energy landscape engineering through Ti substitution enables a distinct nonthermal pathway for ultrafast chiral-domain redistribution. Ti doping stabilizes coexisting chiral domains and tunes their relative stability, allowing femtosecond excitation to drive an asymmetric and anisotropic redistribution from the dominant toward the minority chirality. The subpicosecond response follows a $\sim2$ THz amplitude mode and is consistent with a phonon-assisted pathway involving transient domain-wall configurations. Our results establish free-energy landscape engineering as a strategy for selecting nonequilibrium transition pathways.

cond-mat.str-el

One Kiss: Emojis as Agents of Genre Flux in Generative Comics

Generative AI has made visual storytelling widely accessible, yet current prompt-based interactions often force users into a trade-off between precise control and creative flow. We present One Kiss, a co-creative comic generation system that introduces "Affective Steering". Instead of writing text prompts, users guide the tone of their story through emoji inputs, whose semantic ambiguity becomes a resource rather than a limitation. Unlike traditional text-to-image tools that rely on explicit descriptions, One Kiss uses a dual-stream input in which users define structural pacing by sketching panel frames and set atmospheric tone by pairing keywords with emojis. This mechanism enables "Genre Flux," where emotional inputs accumulate across panels and gradually shift the genre of a story. A preliminary study (N = 6) suggests that this soft steering approach may reframe the user's role from prompt engineer to narrative director, with ambiguity serving as a source of creative surprise rather than a loss of control.

cs.HC

Do Models See in Line with Human Vision? Probing the Correspondence Between LVLM Representations and EEG Signals

Large Vision Language Models (LVLMs) exhibit strong visual understanding and reasoning abilities. However, whether their internal representations reflect human visual cognition is still under-explored. In this paper, we address this by quantifying LVLM-brain alignment using image-evoked Electroencephalogram (EEG) signals, analyzing the effects of model architecture, scale, and image type. Specifically, by using ridge regression and representational similarity analysis, we compare visual representations from 32 open-source LVLMs with corresponding EEG responses. We observe a structured LVLM-brain correspondence: First, intermediate layers (8-16) show peak alignment with EEG activity in the 100-300 ms window, consistent with hierarchical human visual processing. Secondly, multimodal architectural design contributes 3.4 more to brain alignment than parameter scaling, and models with stronger downstream visual performance exhibit higher EEG similarity. Thirdly, spatiotemporal patterns further align with known cortical visual pathways. These results demonstrate that LVLMs learn human-aligned visual representations and establish neural alignment as a biologically grounded benchmark for evaluating and improving LVLMs. In addition, those results could provide insights that may inform the development of neuro-inspired applications.

cs.HC

Effect of a horizontal magnetic field on the melting of low Prandtl number liquid metal in a phase-change Rayleigh-B\'enard system

We investigated the low Prandtl number Rayleigh-B\'enard (RB) system with a melting top boundary under horizontal magnetic fields. This study is crucial to gain physical insights into the melting dynamics of thermal storage systems, which will help in controlling them. A three-dimensional solid-liquid phase change Rayleigh-B\'enard system of gallium was numerically simulated using the enthalpy-porosity method in a cubic domain. In the absence of a magnetic field, the melting process was clearly divided into four regimes: conduction, stable growth, coarsening, and chaotic regimes. We analyzed the flow and heat transfer characteristics in each regime and established scaling relations for the Nusselt number and liquid fraction. Under horizontal magnetic fields, a quasi-two-dimensional (Q2D) flow pattern was observed. We further examined the effects of magnetic field strength on different melting regimes. With increasing Hartmann number, a new flow mode emerged during the stable growth regime. The results also show that the magnetic field alters the relative duration of each melting regime in the overall melting process.

physics.flu-dyn

Mixed spin states for robust ferromagnetism in strained SrCoO$_3$ thin films

Epitaxial strain in transition-metal oxides can induce dramatic changes in electronic and magnetic properties. A recent study on the epitaxially strained SrCoO$_3$ thin films revealed persistent ferromagnetism even across a metal-insulator transition. This challenges the current theoretical predictions, and the nature of the local spin state underlying this robustness remains unresolved. Here, we employ high-resolution resonant inelastic x-ray scattering (RIXS) at the Co-$L_3$ edge to probe the spin states of strained SrCoO$_3$ thin films. Compared with CoO$_6$ cluster multiplet calculations, we identify a ground state composed of a mixed high- and low-spin configuration, distinct from the previously proposed intermediate-spin state. Our results demonstrate that the robustness of ferromagnetism arises from the interplay between this mixed spin state and the presence of ligand holes associated with negative charge transfer. These findings provide direct experimental evidence for a nontrivial magnetic ground state in SrCoO$_3$ and offer new pathways for designing robust ferromagnetic systems in correlated oxides.

cond-mat.str-el

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs

The rapid advancement of Large Language Models (LLMs) and Large Vision-Language Models (LVLMs) has enhanced our ability to process and generate human language and visual information. However, these models often struggle with complex, multi-step multi-modal instructions that require logical reasoning, dynamic feedback integration, and iterative self-correction. To address this, we propose CIMR: Contextualized Iterative Multimodal Reasoning, a novel framework that introduces a context-aware iterative reasoning and self-correction module. CIMR operates in two stages: initial reasoning and response generation, followed by iterative refinement using parsed multi-modal feedback. A dynamic fusion module deeply integrates textual, visual, and contextual features at each step. We fine-tune LLaVA-1.5-7B on the Visual Instruction Tuning (VIT) dataset and evaluate CIMR on the newly introduced Multi-modal Action Planning (MAP) dataset. CIMR achieves 91.5% accuracy, outperforming state-of-the-art models such as GPT-4V (89.2%), LLaVA-1.5 (78.5%), MiniGPT-4 (75.3%), and InstructBLIP (72.8%), demonstrating the efficacy of its iterative reasoning and self-correction capabilities in complex tasks.

cs.LG

Ultrafast Orbital-Selective Photodoping Melts Charge Order in Overdoped Bi-based Cuprates

High-temperature superconductivity in cuprates remains one of the enduring puzzles of condensed matter physics, with charge order (CO) playing a central yet elusive role, particularly in the overdoped regime. Here, we employ time-resolved X-ray absorption spectroscopy and resonant X-ray scattering at a free-electron laser to probe the transient electronic density of states and ultrafast CO dynamics in overdoped (Bi,Pb)$_{2.12}$Sr$_{1.88}$CuO$_{6+\delta}$. We reveal a striking pump laser wavelength dependence - the 800 nm light fails to suppress CO, whereas the 400 nm light effectively melts it. This behavior originates from the fact that 400 nm photons can promote electrons from the Zhang-Rice singlet band to the upper Hubbard band or apical oxygen states, while 800 nm photons lack the energy to excite electrons across the charge-transfer gap. The CO recovery time ($\sim$3 ps) matches that of the underdoped cuprates, indicating universal electronic instability in the phase diagram. Additionally, melting overdoped CO requires an order-of-magnitude higher fluence highlighting the role of lattice interactions. Our findings demonstrate orbital-selective photodoping and provide a route to ultrafast control of emergent quantum phases in correlated materials.

cond-mat.supr-con

Photo-induced Dynamics and Momentum Distribution of Chiral Charge Density Waves in 1T-TiSe$_{2}$

Exploring the photoinduced dynamics of chiral states offers promising avenues for advanced control of condensed matter systems. Photoinduced or photoenhanced chirality in 1T-TiSe$_{2}$ has been suggested as a fascinating platform for optical manipulation of chiral states. However, the mechanisms underlying chirality training and its interplay with the charge density wave (CDW) phase remain elusive. Here, we use time-resolved X-ray diffraction (tr-XRD) with circularly polarized pump lasers to probe the photoinduced dynamics of chirality in 1T-TiSe$_{2}$. We observe a notable ($\sim$20%) difference in CDW intensity suppression between left- and right-circularly polarized pumps. Additionally, we reveal momentum-resolved circular dichroism arising from domains of different chirality, providing a direct link between CDW and chirality. An immediate increase in CDW correlation length upon laser pumping is detected, suggesting the photoinduced expansion of chiral domains. These results both advance the potential of light-driven chirality by elucidating the mechanism driving chirality manipulation in TiSe$_2$, and they demonstrate that tr-XRD with circularly polarized pumps is an effective tool for chirality detection in condensed matter systems.

cond-mat.str-el

Using magnetic dynamics to measure the spin gap in a candidate Kitaev material

Materials potentially hosting Kitaev spin-liquid states are considered crucial for realizing topological quantum computing. However, the intricate nature of spin interactions within these materials complicates the precise measurement of low-energy spin excitations indicative of fractionalized excitations. Using Na$_{2}$Co$_2$TeO$_{6}$ as an example, we study these low-energy spin excitations using the time-resolved resonant elastic x-ray scattering (tr-REXS). Our observations unveil remarkably slow spin dynamics at the magnetic peak, whose recovery timescale is several nanoseconds. This timescale aligns with the extrapolated spin gap of $\sim$ 1 $\mu$eV, obtained by density matrix renormalization group (DMRG) simulations in the thermodynamic limit. The consistency demonstrates the efficacy of tr-REXS in discerning low-energy spin gaps inaccessible to conventional spectroscopic techniques.

cond-mat.str-el

Absence of localized $5d^1$ electrons in KTaO$_3$ interface superconductors

Recently, an exciting discovery of orientation-dependent superconductivity was made in two-dimensional electron gas (2DEG) at the interfaces of LaAlO$_3$/KTaO$_3$ (LAO/KTO) or EuO/KTaO$_3$ (EuO/KTO). The superconducting transition temperature can reach a $T_c$ of up to $\sim$ 2.2 K, which is significantly higher than its 3$d$ counterpart LaAlO$_3$/SrTiO$_3$ (LAO/STO) with a $T_c$ of $\sim$ 0.2 K. However, the underlying origin remains to be understood. To uncover the nature of electrons in KTO-based interfaces, we employ x-ray absorption spectroscopy (XAS) and resonant inelastic x-ray spectroscopy (RIXS) to study LAO/KTO and EuO/KTO with different orientations. We reveal the absence of $dd$ orbital excitations in all the measured samples. Our RIXS results are well reproduced by calculations that considered itinerant $5d$ electrons hybridized with O $2p$ electrons. This suggests that there is a lack of localized Ta $5d^1$ electrons in KTO interface superconductors, which is consistent with the absence of magnetic hysteresis observed in magneto-resistance (MR) measurements. These findings offer new insights into our understanding of superconductivity in Ta $5d$ interface superconductors and their potential applications.

cond-mat.supr-con

Physical Layer Security Techniques for Future Wireless Networks

The broadcast nature of wireless communication systems makes wireless transmission extremely susceptible to eavesdropping and even malicious interference. Physical layer security technology can effectively protect the private information sent by the transmitter from being listened to by illegal eavesdroppers, thus ensuring the privacy and security of communication between the transmitter and legitimate users. The development of mobile communication presents new challenges to physical layer security research. This paper provides a comprehensive survey of the physical layer security research on various promising mobile technologies, including directional modulation (DM), spatial modulation (SM), covert communication, intelligent reflecting surface (IRS)-aided communication, and so on. Finally, future trends and the unresolved technical challenges are summarized in physical layer security for mobile communications.

cs.CR

Beamforming and Transmit Power Design for Intelligent Reconfigurable Surface-aided Secure Spatial Modulation

Intelligent reflecting surface (IRS) is a promising solution to build a programmable wireless environment for future communication systems, in which the reflector elements steer the incident signal in fully customizable ways by passive beamforming. In this paper, an IRS-aided secure spatial modulation (SM) is proposed, where the IRS perform passive beamforming and information transfer simultaneously by adjusting the on-off states of the reflecting elements. We formulate an optimization problem to maximize the average secrecy rate (SR) by jointly optimizing the passive beamforming at IRS and the transmit power at transmitter under the consideration that the direct pathes channels from transmitter to receivers are obstructed by obstacles. As the expression of SR is complex, we derive a newly fitting expression (NASR) for the expression of traditional approximate SR (TASR), which has simpler closed-form and more convenient for subsequent optimization. Based on the above two fitting expressions, three beamforming methods, called maximizing NASR via successive convex approximation (Max-NASR-SCA), maximizing NASR via dual ascent (Max-NASR-DA) and maximizing TASR via semi-definite relaxation (Max-TASR-SDR) are proposed to improve the SR performance. Additionally, two transmit power design (TPD) methods are proposed based on the above two approximate SR expressions, called Max-NASR-TPD and Max-TASR-TPD. Simulation results show that the proposed Max-NASR-DA and Max-NASR-SCA IRS beamformers harvest substantial SR performance gains over Max-TASR-SDR. For TPD, the proposed Max-NASR-TPD performs better than Max-TASR-TPD. Particularly, the Max-NASR-TPD has a closed-form solution.

cs.IT

Estimation of Covariance Matrix of Interference for Secure Spatial Modulation against a Malicious Full-duplex Attacker

In a secure spatial modulation with a malicious full-duplex attacker, how to obtain the interference space or channel state information (CSI) is very important for Bob to cancel or reduce the interference from Mallory. In this paper, different from existing work with a perfect CSI, the covariance matrix of malicious interference (CMMI) from Mallory is estimated and is used to construct the null-space of interference (NSI). Finally, the receive beamformer at Bob is designed to remove the malicious interference using the NSI. To improve the estimation accuracy, a rank detector relying on Akaike information criterion (AIC) is derived. To achieve a high-precision CMMI estimation, two methods are proposed as follows: principal component analysis-eigenvalue decomposition (PCA-EVD), and joint diagonalization (JD). The proposed PCA-EVD is a rank deduction method whereas the JD method is a joint optimization method with improved performance in low signal to interference plus noise ratio (SINR) region at the expense of increased complexities. Simulation results show that the proposed PCA-EVD performs much better than the existing method like sample estimated covariance matrix (SCM) and EVD in terms of normalized mean square error (NMSE) and secrecy rate (SR). Additionally, the proposed JD method has an excellent NMSE performance better than PCA-EVD in the low SINR region (SINR < 0dB) while in the high SINR region PCA-EVD performs better than JD.

cs.IT

Fast Ambiguous DOA Elimination Method of DOA Measurement for Hybrid Massive MIMO Receiver

DOA estimation for massive multiple-input multiple-output (MIMO) system can provide ultra-high-resolution angle estimation. However, due to the high computational complexity and cost of all digital MIMO systems, a hybrid analog digital (HAD) structure MIMO was proposed. In this paper, a fast ambiguous phase elimination method is proposed to solve the problem of direction-finding ambiguity caused by the HAD MIMO. Only two-data-blocks are used to realize DOA estimation. Simulation results show that the proposed method can greatly reduce the estimation delay with a slight performance loss.

cs.IT

Virufy: A Multi-Branch Deep Learning Network for Automated Detection of COVID-19

Fast and affordable solutions for COVID-19 testing are necessary to contain the spread of the global pandemic and help relieve the burden on medical facilities. Currently, limited testing locations and expensive equipment pose difficulties for individuals trying to be tested, especially in low-resource settings. Researchers have successfully presented models for detecting COVID-19 infection status using audio samples recorded in clinical settings [5, 15], suggesting that audio-based Artificial Intelligence models can be used to identify COVID-19. Such models have the potential to be deployed on smartphones for fast, widespread, and low-resource testing. However, while previous studies have trained models on cleaned audio samples collected mainly from clinical settings, audio samples collected from average smartphones may yield suboptimal quality data that is different from the clean data that models were trained on. This discrepancy may add a bias that affects COVID-19 status predictions. To tackle this issue, we propose a multi-branch deep learning network that is trained and tested on crowdsourced data where most of the data has not been manually processed and cleaned. Furthermore, the model achieves state-of-art results for the COUGHVID dataset [16]. After breaking down results for each category, we have shown an AUC of 0.99 for audio samples with COVID-19 positive labels.

cs.SD

Spatial Modulation: an Attractive Secure Solution to Future Wireless Network

As a green and secure wireless transmission method, secure spatial modulation (SM) is becoming a hot research area. Its basic idea is to exploit both the index of activated transmit antenna and amplitude phase modulation signal to carry messages, improve security, and save energy. In this paper, we review its crucial challenges: transmit antenna selection (TAS), artificial noise (AN) projection, power allocation (PA) and joint detection at the desired receiver. As the size of signal constellation tends to medium-scale or large-scale, the complexity of traditional maximum likelihood detector becomes prohibitive. To reduce this complexity, a low-complexity maximum likelihood (ML) detector is proposed. To further enhance the secrecy rate (SR) performance, a deep-neural-network (DNN) PA strategy is proposed. Simulation results show that the proposed low-complexity ML detector, with a lower-complexity, has the same bit error rate performance as the joint ML method while the proposed DNN method strikes a good balance between complexity and SR performance.

cs.IT

Virufy: Global Applicability of Crowdsourced and Clinical Datasets for AI Detection of COVID-19 from Cough

Rapid and affordable methods of testing for COVID-19 infections are essential to reduce infection rates and prevent medical facilities from becoming overwhelmed. Current approaches of detecting COVID-19 require in-person testing with expensive kits that are not always easily accessible. This study demonstrates that crowdsourced cough audio samples recorded and acquired on smartphones from around the world can be used to develop an AI-based method that accurately predicts COVID-19 infection with an ROC-AUC of 77.1% (75.2%-78.3%). Furthermore, we show that our method is able to generalize to crowdsourced audio samples from Latin America and clinical samples from South Asia, without further training using the specific samples from those regions. As more crowdsourced data is collected, further development can be implemented using various respiratory audio samples to create a cough analysis-based machine learning (ML) solution for COVID-19 detection that can likely generalize globally to all demographic groups in both clinical and non-clinical settings.

cs.SD