SearcharxivSearch

arXiv subjects

Xue Zhao

Publications and source records attributed to Xue Zhao.

11 recordsLinked to original sources

Element-Aware Group Learning for E-Commerce Image Generation

Recent advances in image generation and editing have made prompt quality a key bottleneck for e-commerce creatives. Vision-language models (VLMs) can generate image-editing prompts from product images and metadata, but further improving their prompt-writing capabilities requires post-training with feedback from the generated images. Group Relative Policy Optimization (GRPO) is a natural framework for such outcome-level reward optimization. However, it assigns credit only at the full-prompt level, even though image quality often depends on specific design elements such as composition, background, and the presentation of selling points. Existing fine-grained credit assignment methods typically require step-level supervision or learned critics. To address this, we propose EAGLE-GRPO (Element-Aware Group Learning for E-Commerce Image Generation), which decomposes the group-centered reward over predefined elements. We cast element-level credit assignment as a kernel ridge regression problem and derive a closed-form solution, without additional rollouts or separate credit-assignment models. This yields interpretable per-element advantages and more precise policy updates. Experiments show that EAGLE-GRPO sustains performance gains over more training steps before plateauing and generates prompts that produce higher-quality e-commerce images than competitive VLM prompt-writing baselines.

cs.CV

Artificial Intelligence for Instability in Inorganic Perovskites: From Mechanism Discovery to Engineering Strategies

Three-dimensional all-inorganic halide perovskites, represented by CsPbX$_3$ (X = Cl, Br, I), have attracted broad interest in photovoltaics, photodetectors, and light-emitting devices because of their outstanding optoelectronic properties. Their practical deployment, however, remains limited by instability under thermal, chemical, optical, and electrical stress. Conventional studies have established important experimental and theoretical foundations, but they still struggle with multimodal data, coupled degradation pathways, protocol dependence, sparse statistics, and uncertainty quantification. Artificial intelligence (AI) offers a practical route to address these limitations. This review summarizes recent progress in AI-assisted studies of instability in 3D CsPbX$_3$ and organizes the discussion around four linked tasks, including stability discrimination and diagnosis, microscopic mechanism analysis, consequence and reliability modeling, and engineering stability enhancement. We further discuss the main limitations of current methods, especially in data quality, protocol consistency, benchmark design, interpretability, and transferability across domains. Finally, we outline future directions for the field, including standardized data infrastructures, interpretable cross-scale models, and tighter integration of AI with automated experiments and physics-based modeling. The aim of this review is to provide a coherent and practically useful framework for researchers seeking to use AI to understand, predict, and mitigate instability in inorganic perovskites.

cond-mat.mtrl-sci

A compact setup for 87Rb optical tweezer arrays

We describe a simple and compact experimental setup for optical tweezer arrays of 87Rb atoms. This setup includes a compact vacuum system, a single cooling laser, a simple tweezer laser, and a flexible control system. The small vacuum system with only 40 cm length takes advantage of the high atomic flux two-dimensional magneto-optical trap (2D MOT) while maintaining a low background pressure in the 3D MOT chamber ensuring sufficient lifetime of the trapped atoms. Atom number of the laser cooled sample of 2e7 and temperature of 92 uK is achieved. The flexible control system with real-time waveform generator modules (RWG) provides precise control of all the RF devices, and enables real-time feedback control of both the global and individual beams in optical tweezer arrays. An optical tweezer array with 25x25 homogeneous traps is demonstrated. This simple and compact demo setup makes it more accessible to experimental quantum physics.

cond-mat.quant-gas

Learning to Hear by Seeing: It's Time for Vision Language Models to Understand Artistic Emotion from Sight and Sound

Emotion understanding is critical for making Large Language Models (LLMs) more general, reliable, and aligned with humans. Art conveys emotion through the joint design of visual and auditory elements, yet most prior work is human-centered or single-modality, overlooking the emotion intentionally expressed by the artwork. Meanwhile, current Audio-Visual Language Models (AVLMs) typically require large-scale audio pretraining to endow Visual Language Models (VLMs) with hearing, which limits scalability. We present Vision Anchored Audio-Visual Emotion LLM (VAEmotionLLM), a two-stage framework that teaches a VLM to hear by seeing with limited audio pretraining and to understand emotion across modalities. In Stage 1, Vision-Guided Audio Alignment (VG-Align) distills the frozen visual pathway into a new audio pathway by aligning next-token distributions of the shared LLM on synchronized audio-video clips, enabling hearing without a large audio dataset. In Stage 2, a lightweight Cross-Modal Emotion Adapter (EmoAdapter), composed of the Emotion Enhancer and the Emotion Supervisor, injects emotion-sensitive residuals and applies emotion supervision to enhance cross-modal emotion understanding. We also construct ArtEmoBenchmark, an art-centric emotion benchmark that evaluates content and emotion understanding under audio-only, visual-only, and audio-visual inputs. VAEmotionLLM achieves state-of-the-art results on ArtEmoBenchmark, outperforming audio-only, visual-only, and audio-visual baselines. Ablations show that the proposed components are complementary.

cs.CV

ICH-Qwen: A Large Language Model Towards Chinese Intangible Cultural Heritage

The intangible cultural heritage (ICH) of China, a cultural asset transmitted across generations by various ethnic groups, serves as a significant testament to the evolution of human civilization and holds irreplaceable value for the preservation of historical lineage and the enhancement of cultural self-confidence. However, the rapid pace of modernization poses formidable challenges to ICH, including threats damage, disappearance and discontinuity of inheritance. China has the highest number of items on the UNESCO Intangible Cultural Heritage List, which is indicative of the nation's abundant cultural resources and emphasises the pressing need for ICH preservation. In recent years, the rapid advancements in large language modelling have provided a novel technological approach for the preservation and dissemination of ICH. This study utilises a substantial corpus of open-source Chinese ICH data to develop a large language model, ICH-Qwen, for the ICH domain. The model employs natural language understanding and knowledge reasoning capabilities of large language models, augmented with synthetic data and fine-tuning techniques. The experimental results demonstrate the efficacy of ICH-Qwen in executing tasks specific to the ICH domain. It is anticipated that the model will provide intelligent solutions for the protection, inheritance and dissemination of intangible cultural heritage, as well as new theoretical and practical references for the sustainable development of intangible cultural heritage. Furthermore, it is expected that the study will open up new paths for digital humanities research.

cs.CL

A stable phase-locking-free single beam optical lattice with multiple configurations

Optical lattices formed by interfering laser beams are widely used to trap and manipulate atoms for quantum simulation, metrology, and computation. To stabilize optical lattices in experiments, it is usually challenging to implement delicate phase-locking systems with complicated optics and electronics to reduce the relative phase fluctuation of multiple laser beams. Here we report a phase-locking-free scheme to implement optical lattices by passing a single laser beam through a prism with n-fold symmetric facets and large apex angles. The scheme ensures a stable optical lattice since the interference occurs among different deflected parts of a single laser beam without any moving component. Various lattice configurations, including a triangular lattice and a quasi-crystal lattice with ten-fold symmetry are demonstrated. In both cases, stability measurements show a change of lattice constant in less than 1.14%, and a drift of lattice position in less than 1.61%.

quant-ph

An efficient method to generate near-ideal hollow beams of different shapes for box potential of quantum gases

Ultracold quantum gases are usually prepared in conservative traps for quantum simulation experiments. The atomic density inhomogeneity, together with the consequent position-dependent energy and time scales of cold atoms in traditional harmonic traps, makes it difficult to manipulate and detect the sample at a better level. These problems are partially solved by optical box traps of blue-detuned hollow beams. However, generating a high-quality hollow beam with high light efficiency for the box trap is challenging. Here, we present a scheme that combines the fixed optics, including axicons and prisms, to pre-shape a Gaussian beam into a hollow beam, with a digital micromirror device (DMD) to improve the quality of the hollow beam further, providing a nearly ideal optical potential of various shapes for preparing highly homogeneous cold atoms. The highest power-law exponent of potential walls can reach a value over 100, and the light efficiency from a Gaussian to a hollow beam is also improved compared to direct optical shaping by a mask or a DMD. Combined with a one-dimensional optical lattice, a nearly ideal two-dimensional uniform quantum gas with different geometrical boundaries can be prepared for exploring quantum many-body physics to an unprecedented level.

cond-mat.quant-gas

GujiBERT and GujiGPT: Construction of Intelligent Information Processing Foundation Language Models for Ancient Texts

In the context of the rapid development of large language models, we have meticulously trained and introduced the GujiBERT and GujiGPT language models, which are foundational models specifically designed for intelligent information processing of ancient texts. These models have been trained on an extensive dataset that encompasses both simplified and traditional Chinese characters, allowing them to effectively handle various natural language processing tasks related to ancient books, including but not limited to automatic sentence segmentation, punctuation, word segmentation, part-of-speech tagging, entity recognition, and automatic translation. Notably, these models have exhibited exceptional performance across a range of validation tasks using publicly available datasets. Our research findings highlight the efficacy of employing self-supervised methods to further train the models using classical text corpora, thus enhancing their capability to tackle downstream tasks. Moreover, it is worth emphasizing that the choice of font, the scale of the corpus, and the initial model selection all exert significant influence over the ultimate experimental outcomes. To cater to the diverse text processing preferences of researchers in digital humanities and linguistics, we have developed three distinct categories comprising a total of nine model variations. We believe that by sharing these foundational language models specialized in the domain of ancient texts, we can facilitate the intelligent processing and scholarly exploration of ancient literary works and, consequently, contribute to the global dissemination of China's rich and esteemed traditional culture in this new era.

cs.CL

Conformational change-modulated spin transport at the single-molecule level in carbon systems --Invited for the Third Carbon Special Topic

Controlling the spin transport at the single-molecule level, especially without the use of ferromagnetic contacts, becomes a focus of research in spintronics. Inspired by the progress on atomic-level molecular synthesis, through first-principles calculations, we investigate the spin-dependent electronic transport of graphene nanoflakes with side-bonded functional groups, contacted by atomic carbon chain electrodes. It is found that, by rotating the functional group, the spin polarization of the transmission at the Fermi level could be switched between completely polarized and unpolarized states. Moreover, the transition between spin-up and spin-down polarized states can also be achieved, operating as a dual-spin filter. Further analysis shows that, it is the spin-dependent shift of density of states, caused by the rotation, that triggers the shift of transmission peaks, and then results in the variation of spin polarization. Such a feature is found to be robust to the of the nanoflake and the electrode material, showing great application potential. Those findings may throw light on the development of spintronic devices.

physics.atm-clus

Microscopic origin of magnetization reversal in exchange-coupled ferro-/ferrimagnetic bilayers

In this study, the magnetic reversal process of exchange-coupled bilayer systems, consisting of a ferrimagnetic TbFeCo alloy layer and a ferromagnetic [Co/Ni/Pt]N multilayer, was investigated. In particular, minor loop studies, probing solely the reversal characteristics of the softer ferromagnetic layer, reveal two distinct reversal mechanisms, which depend strongly on the thickness of the ferromagnetic layer. For thick layers, irreversible switching of the macroscopic minor loop is observed. The underlying microscopic origin of this reversal process was studied in detail by high-resolution magnetic force microscopy, showing that the reversal is triggered by in-plane domain walls propagating through the ferromagnetic layer. In contrast, thin ferromagnetic layers show a hysteresis-free reversal, which is nucleation-dominated due to grain-to-grain variations in magnetic anisotropy of the Co/Ni/Pt multilayer and an inhomogeneous exchange coupling with the magnetically hard TbFeCo layer, as confirmed by micromagnetic simulations.

cond-mat.mtrl-sci

Bimodal Magnetic Force Microscopy with Capacitive Tip-Sample Distance Control

A single-passage, bimodal magnetic force microscopy technique optimized for scanning samples with arbitrary topography is discussed. A double phase-locked loop (PLL) system is used to mechanically excite a high quality factor cantilever under vacuum conditions on its first mode and via an oscillatory tip-sample potential on its second mode. The obtained second mode oscillation amplitude is then used as a proxy for the tip-sample distance, and for the control thereof. With appropriate $z$-feedback parameters two data sets reflecting the magnetic tip-sample interaction and the sample topography are simultaneously obtained.

cond-mat.mes-hall