SearcharxivSearch

arXiv subjects

Yuchi Zhang

Publications and source records attributed to Yuchi Zhang.

18 recordsLinked to original sources

Process-Knowledge-Embedded Safe DRL for Real-Time Dispatch of Process Loads in Industrial Microgrids

Steelmaking process loads (SPLs) are flexible resources that enhance local renewable-energy utilization and reduce electricity procurement costs in industrial microgrids. However, strong multistage coupling makes current decisions affect subsequent feasibility, challenging conventional deep reinforcement learning to reduce costs while maintaining process feasibility throughout production. This paper proposes a process-knowledge-embedded safe deep reinforcement learning framework for the real-time dispatch of SPLs in industrial microgrids. Specifically, a lossless active-frontier action space is constructed, and a process-distance-guided action-processing mechanism reallocates excluded-action probabilities according to process distance and the actor's safe-action preference. Recursive process feasibility is established to guarantee admissible execution and feasible continuation. Furthermore, the expected process-correction distance is incorporated into PPO through a correction budget and a primal-dual update to internalize process knowledge into the raw policy, while a derived bound quantifies the raw policy's dependence on safety processing. Case studies using real-world data demonstrate zero process losses, electricity-cost reductions of 49.2% and 25.9% relative to rule-based scheduling and rolling MILP, respectively, within an acceptable computation time.

eess.SY

GS-Playground: A High-Throughput Photorealistic Simulator for Vision-Informed Robot Learning

Embodied AI research is undergoing a shift toward vision-centric perceptual paradigms. While massively parallel simulators have catalyzed breakthroughs in proprioception-based locomotion, their potential remains largely untapped for vision-informed tasks due to the prohibitive computational overhead of large-scale photorealistic rendering. Furthermore, the creation of simulation-ready 3D assets heavily relies on labor-intensive manual modeling, while the significant sim-to-real physical gap hinders the transfer of contact-rich manipulation policies. To address these bottlenecks, we propose GS-Playground, a multi-modal simulation framework designed to accelerate end-to-end perceptual learning. We develop a novel high-performance parallel physics engine, specifically designed to integrate with a batch 3D Gaussian Splatting (3DGS) rendering pipeline to ensure high-fidelity synchronization. Our system achieves a breakthrough throughput of 10^4 FPS at 640x480 resolution, significantly lowering the barrier for large-scale visual RL. Additionally, we introduce an automated Real2Sim workflow that reconstructs photorealistic, physically consistent, and memory-efficient environments, streamlining the generation of complex simulation-ready scenes. Extensive experiments on locomotion, navigation, and manipulation demonstrate that GS-Playground effectively bridges the perceptual and physical gaps across diverse embodied tasks. Project homepage: https://gsplayground.github.io.

cs.RO

ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents

Multi-turn LLM agents are increasingly important for solving complex, interactive tasks, and reinforcement learning (RL) is a key ingredient for improving their long-horizon behavior. However, RL training requires generating large numbers of sandboxed rollout trajectories, and existing infrastructures often couple rollout orchestration with the training loop, making systems hard to migrate and maintain. Under the rollout-as-a-service philosophy, we present ProRL Agent , a scalable infrastructure that serves the full agentic rollout lifecycle through an API service. ProRL Agent also provides standardized and extensible sandbox environments that support diverse agentic tasks in rootless HPC settings. We validate ProRL Agent through RL training on software engineering, math, STEM, and coding tasks. ProRL Agent is open-sourced and integrated as part of NVIDIA NeMo Gym.

cs.AI

Bridging Scale Discrepancies in Robotic Control via Language-Based Action Representations

Recent end-to-end robotic manipulation research increasingly adopts architectures inspired by large language models to enable robust manipulation. However, a critical challenge arises from severe distribution shifts between robotic action data, primarily due to substantial numerical variations in action commands across diverse robotic platforms and tasks, hindering the effective transfer of pretrained knowledge. To address this limitation, we propose a semantically grounded linguistic representation to normalize actions for efficient pretraining. Unlike conventional discretized action representations that are sensitive to numerical scales, the motion representation specifically disregards numeric scale effects, emphasizing directionality instead. This abstraction mitigates distribution shifts, yielding a more generalizable pretraining representation. Moreover, using the motion representation narrows the feature distance between action tokens and standard vocabulary tokens, mitigating modality gaps. Multi-task experiments on two benchmarks demonstrate that the proposed method significantly improves generalization performance and transferability in robotic manipulation tasks.

cs.RO

MoE-GraphSAGE-Based Integrated Evaluation of Transient Rotor Angle and Voltage Stability in Power Systems

The large-scale integration of renewable energy and power electronic devices has increased the complexity of power system stability, making transient stability assessment more challenging. Conventional methods are limited in both accuracy and computational efficiency. To address these challenges, this paper proposes MoE-GraphSAGE, a graph neural network framework based on the MoE for unified TAS and TVS assessment. The framework leverages GraphSAGE to capture the power grid's spatiotemporal topological features and employs multi-expert networks with a gating mechanism to model distinct instability modes jointly. Experimental results on the IEEE 39-bus system demonstrate that MoE-GraphSAGE achieves superior accuracy and efficiency, offering an effective solution for online multi-task transient stability assessment in complex power systems.

eess.SY

Cavity-Enhanced Rydberg Atomic Superheterodyne Receiver

High-sensitivity measurements of the microwave electric field are important in applications of communication and metrology. \replaced{The sensitivity of traditional Rydberg superheterodyne receivers in free space is effectively determined by the signal-to-noise ratio (SNR), which is often considered equivalent to sensitivity in practical sensing applications.}{The sensitivity of the traditional Rydberg superheterodyne receivers in free space is limited by signal-to-noise contrast.} In this work, we demonstrate a cavity-enhanced receiver, where an optical cavity significantly amplifies the interaction between the probe light and cesium atoms, which substantially improves the signal-to-noise ratio via enhancing the expansion coefficient \( \kappa \). \added{Here, $\kappa$ is the edge slope of the single peak obtained by fitting the double-peak EIT-AT spectrum, characterizing the response of the probe light to the frequency detuning of the coupling laser.}The sensitivity is thus boosted by a factor of approximately 19 dB. This study highlights the pivotal role of optical cavities in advancing Rydberg-based detection systems, offering a promising approach for high-sensitivity microwave electric field measurements.

quant-ph

Conversational Recommender System and Large Language Model Are Made for Each Other in E-commerce Pre-sales Dialogue

E-commerce pre-sales dialogue aims to understand and elicit user needs and preferences for the items they are seeking so as to provide appropriate recommendations. Conversational recommender systems (CRSs) learn user representation and provide accurate recommendations based on dialogue context, but rely on external knowledge. Large language models (LLMs) generate responses that mimic pre-sales dialogues after fine-tuning, but lack domain-specific knowledge for accurate recommendations. Intuitively, the strengths of LLM and CRS in E-commerce pre-sales dialogues are complementary, yet no previous work has explored this. This paper investigates the effectiveness of combining LLM and CRS in E-commerce pre-sales dialogues, proposing two collaboration methods: CRS assisting LLM and LLM assisting CRS. We conduct extensive experiments on a real-world dataset of Ecommerce pre-sales dialogues. We analyze the impact of two collaborative approaches with two CRSs and two LLMs on four tasks of Ecommerce pre-sales dialogue. We find that collaborations between CRS and LLM can be very effective in some cases.

cs.CL

Towards Personalized Review Summarization by Modeling Historical Reviews from Customer and Product Separately

Review summarization is a non-trivial task that aims to summarize the main idea of the product review in the E-commerce website. Different from the document summary which only needs to focus on the main facts described in the document, review summarization should not only summarize the main aspects mentioned in the review but also reflect the personal style of the review author. Although existing review summarization methods have incorporated the historical reviews of both customer and product, they usually simply concatenate and indiscriminately model this two heterogeneous information into a long sequence. Moreover, the rating information can also provide a high-level abstraction of customer preference, it has not been used by the majority of methods. In this paper, we propose the Heterogeneous Historical Review aware Review Summarization Model (HHRRS) which separately models the two types of historical reviews with the rating information by a graph reasoning module with a contrastive loss. We employ a multi-task framework that conducts the review sentiment classification and summarization jointly. Extensive experiments on four benchmark datasets demonstrate the superiority of HHRRS on both tasks.

cs.CL

AudioViewer: Learning to Visualize Sounds

A long-standing goal in the field of sensory substitution is to enable sound perception for deaf and hard of hearing (DHH) people by visualizing audio content. Different from existing models that translate to hand sign language, between speech and text, or text and images, we target immediate and low-level audio to video translation that applies to generic environment sounds as well as human speech. Since such a substitution is artificial, without labels for supervised learning, our core contribution is to build a mapping from audio to video that learns from unpaired examples via high-level constraints. For speech, we additionally disentangle content from style, such as gender and dialect. Qualitative and quantitative results, including a human study, demonstrate that our unpaired translation approach maintains important audio features in the generated video and that videos of faces and numbers are well suited for visualizing high-dimensional audio features that can be parsed by humans to match and distinguish between sounds and words. Code and models are available at https://chunjinsong.github.io/audioviewer

cs.HC

10-Hertz quantum light source generation on the cesium D2 line using single photon modulation

Generation of quantum light source is a promising technique to overcome the standard quantum limit in precision measurement. Here, we demonstrate an experimental generation of quadrature squeezing resonating on the cesium D2 line down to 10 Hz for the first time. The maximum squeezing in audio frequency band is 5.57 dB. Moreover, we have presented a single-photon modulation locking to control the squeezing angle, while effectively suppressing the influence of laser noise on low-frequency squeezing. The whole system operates steadily for hours. The generated low-frequency quantum light source can be applied in quantum metrology,light-matter interaction investigation and quantum memory in the audio frequency band and even below.

quant-ph

Towards Generalized Models for Task-oriented Dialogue Modeling on Spoken Conversations

Building robust and general dialogue models for spoken conversations is challenging due to the gap in distributions of spoken and written data. This paper presents our approach to build generalized models for the Knowledge-grounded Task-oriented Dialogue Modeling on Spoken Conversations Challenge of DSTC-10. In order to mitigate the discrepancies between spoken and written text, we mainly employ extensive data augmentation strategies on written data, including artificial error injection and round-trip text-speech transformation. To train robust models for spoken conversations, we improve pre-trained language models, and apply ensemble algorithms for each sub-task. Typically, for the detection task, we fine-tune \roberta and ELECTRA, and run an error-fixing ensemble algorithm. For the selection task, we adopt a two-stage framework that consists of entity tracking and knowledge ranking, and propose a multi-task learning method to learn multi-level semantic information by domain classification and entity selection. For the generation task, we adopt a cross-validation data process to improve pre-trained generative language models, followed by a consensus decoding algorithm, which can add arbitrary features like relative \rouge metric, and tune associated feature weights toward \bleu directly. Our approach ranks third on the objective evaluation and second on the final official human evaluation.

cs.CL

Precognition in Task-oriented Dialogue Understanding: Posterior Regularization by Future Context

Task-oriented dialogue systems have become overwhelmingly popular in recent researches. Dialogue understanding is widely used to comprehend users' intent, emotion and dialogue state in task-oriented dialogue systems. Most previous works on such discriminative tasks only models current query or historical conversations. Even if in some work the entire dialogue flow was modeled, it is not suitable for the real-world task-oriented conversations as the future contexts are not visible in such cases. In this paper, we propose to jointly model historical and future information through the posterior regularization method. More specifically, by modeling the current utterance and past contexts as prior, and the entire dialogue flow as posterior, we optimize the KL distance between these distributions to regularize our model during training. And only historical information is used for inference. Extensive experiments on two dialogue datasets validate the effectiveness of our proposed method, achieving superior results compared with all baseline models.

cs.CL

High-Order Continuous-Variable Coherence of Phase-Dependent Squeezed State

We study continuous variable coherence of phase-dependent squeezed state based on an extended Hanbury Brown-Twiss scheme. High-order coherence is continuously varied by adjusting squeezing parameter $r$, displacement $α$, and squeezing phase $θ$. We also analyze effects of background noise $γ$ and detection efficiency $η$ on the measurements. As the squeezing phase shifts from 0 to $π$, the photon statistics of the squeezed state continuously change from the anti-bunching ($g^{(n)}<1$) to super-bunching ($g^{(n)}>n!$) which shows a transition from particle nature to wave nature. The experiment feasibility is also examined. It provides a practical method to generate phase-dependent squeezed states with high-order continuous-variable coherence by tuning squeezing phase $θ$. The controllable coherence source can be applied to sensitivity improvement in gravitational wave detection and quantum imaging.

quant-ph

HeteroQA: Learning towards Question-and-Answering through Multiple Information Sources via Heterogeneous Graph Modeling

Community Question Answering (CQA) is a well-defined task that can be used in many scenarios, such as E-Commerce and online user community for special interests. In these communities, users can post articles, give comment, raise a question and answer it. These data form the heterogeneous information sources where each information source have their own special structure and context (comments attached to an article or related question with answers). Most of the CQA methods only incorporate articles or Wikipedia to extract knowledge and answer the user's question. However, various types of information sources in the community are not fully explored by these CQA methods and these multiple information sources (MIS) can provide more related knowledge to user's questions. Thus, we propose a question-aware heterogeneous graph transformer to incorporate the MIS in the user community to automatically generate the answer. To evaluate our proposed method, we conduct the experiments on two datasets: $\text{MSM}^{\text{plus}}$ the modified version of benchmark dataset MS-MARCO and the AntQA dataset which is the first large-scale CQA dataset with four types of MIS. Extensive experiments on two datasets show that our model outperforms all the baselines in terms of all the metrics.

cs.CL

Determination of weak squeezed vacuum state through photon statistics measurement

Weak squeezed vacuum light, especially resonant to the atomic transition, plays an important role in quantum storage and generation of various quantum sources. However, the general homodyne detection (HD) cannot determine weak squeezing due to the low signal to noise ratio and the limited resolution of the HD system. Here we provide an alternative method based on photon statistics measurement to determine the weak squeezing of the squeezed vacuum light generated from an optical parametric oscillator working far below the threshold. The approach is established the relationship between the squeezing parameter and the second-order degree of coherence. The theoretical analysis agrees well with the experiment results. The advantage of this method is that it provides a feasible and reliable experimental measure to determine the weak squeezing with high precision and the measurement is independent on the detection efficiency. This method can be used to measure other quantum features for various quantum states with extremely weak non-classicality.

quant-ph

Improve Diverse Text Generation by Self Labeling Conditional Variational Auto Encoder

Diversity plays a vital role in many text generating applications. In recent years, Conditional Variational Auto Encoders (CVAE) have shown promising performances for this task. However, they often encounter the so called KL-Vanishing problem. Previous works mitigated such problem by heuristic methods such as strengthening the encoder or weakening the decoder while optimizing the CVAE objective function. Nevertheless, the optimizing direction of these methods are implicit and it is hard to find an appropriate degree to which these methods should be applied. In this paper, we propose an explicit optimizing objective to complement the CVAE to directly pull away from KL-vanishing. In fact, this objective term guides the encoder towards the "best encoder" of the decoder to enhance the expressiveness. A labeling network is introduced to estimate the "best encoder". It provides a continuous label in the latent space of CVAE to help build a close connection between latent variables and targets. The whole proposed method is named Self Labeling CVAE~(SLCVAE). To accelerate the research of diverse text generation, we also propose a large native one-to-many dataset. Extensive experiments are conducted on two tasks, which show that our method largely improves the generating diversity while achieving comparable accuracy compared with state-of-art algorithms.

stat.ML

Magnetically guided Cesium interferometer for inertial sensing

In this paper we demonstrate a magnetically guided Cesium (Cs) atom interferometer in the Talbot-Lau regime for inertial sensing with two interferometer schemes, Mach-Zenhder and Ramsey-Borde. The recoil frequency of the Cs atoms and the acceleration along the waveguide symmetry axis is measured. An acceleration measurement uncertainty of $7\times10^{-5}$ m/s$^{2}$ is achieved. We also realize an enclosed area of $0.018$ mm$^{2}$ for rotation measurement. As the first reported magnetically guided Cs atom interferometer, the system limitation and its advantages are discussed.

physics.atom-ph

Losses-based test of wave-particle duality with Mach-Zehnder interferometers

Wave-particle duality of photons with losses in the Mach-Zehnder interferometer (MZI) is investigated experimentally and theoretically. The experiment is done with the standard MZI with the beam splitter or the beam merger being continuously varied. The losses are deliberately introduced either inside the MZI (the two arms between the beam splitter and beam mergers) or outside the MZI (after the beam merger). It is proved that the unbalanced losses have great influence on the predictability $P$ (particle nature) and visibility $V$ (wave nature). For the former case the duality inequality holds while for the later the duality inequality is ``violated''. We get $P^2+V^2>1$. This ``violation'' could be eliminated in principle by switching the two paths and detectors and then averaging the results. The observed results can be exactly explained theoretically. The experiment is done with coherent beam, instead of single photons, and we have proved that they are exactly equivalent in duality experiment with MZI.

quant-ph