SearcharxivSearch

arXiv subjects

Wenjing Zhu

Publications and source records attributed to Wenjing Zhu.

11 recordsLinked to original sources

MTDrive: Multi-turn Interactive Reinforcement Learning for Autonomous Driving

Trajectory planning is a core task in autonomous driving, requiring the prediction of safe and comfortable paths across diverse scenarios. Integrating Multi-modal Large Language Models (MLLMs) with Reinforcement Learning (RL) has shown promise in addressing "long-tail" scenarios. However, existing methods are constrained to single-turn reasoning, limiting their ability to handle complex tasks requiring iterative refinement. To overcome this limitation, we present MTDrive, a multi-turn framework that enables MLLMs to iteratively refine trajectories based on environmental feedback. MTDrive introduces Multi-Turn Group Relative Policy Optimization (mtGRPO), which mitigates reward sparsity by computing relative advantages across turns. We further construct an interactive trajectory understanding dataset from closed-loop simulation to support multi-turn training. Experiments on the NAVSIM benchmark demonstrate superior performance compared to existing methods, validating the effectiveness of our multi-turn reasoning paradigm. Additionally, we implement system-level optimizations to reduce data transfer overhead caused by high-resolution images and multi-turn sequences, achieving 2.5x training throughput. Our data, models, and code will be made available soon.

cs.RO

Skipformer: A Skip-and-Recover Strategy for Efficient Speech Recognition

Conformer-based attention models have become the de facto backbone model for Automatic Speech Recognition tasks. A blank symbol is usually introduced to align the input and output sequences for CTC or RNN-T models. Unfortunately, the long input length overloads computational budget and memory consumption quadratically by attention mechanism. In this work, we propose a "Skip-and-Recover" Conformer architecture, named Skipformer, to squeeze sequence input length dynamically and inhomogeneously. Skipformer uses an intermediate CTC output as criteria to split frames into three groups: crucial, skipping and ignoring. The crucial group feeds into next conformer blocks and its output joint with skipping group by original temporal order as the final encoder output. Experiments show that our model reduces the input sequence length by 31 times on Aishell-1 and 22 times on Librispeech corpus. Meanwhile, the model can achieve better recognition accuracy and faster inference speed than recent baseline models. Our code is open-sourced and available online.

cs.CL

Learning a Structural Causal Model for Intuition Reasoning in Conversation

Reasoning, a crucial aspect of NLP research, has not been adequately addressed by prevailing models including Large Language Model. Conversation reasoning, as a critical component of it, remains largely unexplored due to the absence of a well-designed cognitive model. In this paper, inspired by intuition theory on conversation cognition, we develop a conversation cognitive model (CCM) that explains how each utterance receives and activates channels of information recursively. Besides, we algebraically transformed CCM into a structural causal model (SCM) under some mild assumptions, rendering it compatible with various causal discovery methods. We further propose a probabilistic implementation of the SCM for utterance-level relation reasoning. By leveraging variational inference, it explores substitutes for implicit causes, addresses the issue of their unobservability, and reconstructs the causal representations of utterances through the evidence lower bounds. Moreover, we constructed synthetic and simulated datasets incorporating implicit causes and complete cause labels, alleviating the current situation where all available datasets are implicit-causes-agnostic. Extensive experiments demonstrate that our proposed method significantly outperforms existing methods on synthetic, simulated, and real-world datasets. Finally, we analyze the performance of CCM under latent confounders and propose theoretical ideas for addressing this currently unresolved issue.

cs.CL

How to Enhance Causal Discrimination of Utterances: A Case on Affective Reasoning

Our investigation into the Affective Reasoning in Conversation (ARC) task highlights the challenge of causal discrimination. Almost all existing models, including large language models (LLMs), excel at capturing semantic correlations within utterance embeddings but fall short in determining the specific causal relationships. To overcome this limitation, we propose the incorporation of \textit{i.i.d.} noise terms into the conversation process, thereby constructing a structural causal model (SCM). It explores how distinct causal relationships of fitted embeddings can be discerned through independent conditions. To facilitate the implementation of deep learning, we introduce the cogn frameworks to handle unstructured conversation data, and employ an autoencoder architecture to regard the unobservable noise as learnable "implicit causes." Moreover, we curate a synthetic dataset that includes i.i.d. noise. Through comprehensive experiments, we validate the effectiveness and interpretability of our approach. Our code is available in https://github.com/Zodiark-ch/mater-of-our-EMNLP2023-paper.

cs.CL

Determination of the resonant parameters of excited vector strangenia with the $e^{+}e^{-}\toηϕ$ data

We determine the resonant parameters of the vector states $ϕ(1680)$ and $ϕ(2170)$, by doing a combined fit to the $e^{+}e^{-}\to ηϕ$ cross sections from threshold to $2.85~\rm GeV$ measured by BaBar, Belle, BESIII and CMD-3 experiments. The mass $(1678^{+5}_{-3} \pm 7)~\rm MeV/c^2$ and the width $(156\pm 5 \pm 9)~\rm MeV$ are obtained for the $ϕ(1680)$, and the mass $(2169\pm 5 \pm 6)~\rm MeV/c^2$ and the width $(96^{+17}_{-14} \pm 9)~\rm MeV$ for the $ϕ(2170)$. The statistical significance of $ϕ(2170)$ is $7.2σ$. Depending on the interference between the $ϕ(1680)$, $ϕ(2170)$ and a non-resonant $ηϕ$ amplitude in the nominal fit, we obtain four solutions and $Γ^{e^{+}e^{-}}_{ϕ(1680)}\cdot \mathcal{B}[ϕ(1680)\toηϕ] = (79 \pm 4 \pm 16)$, $(127\pm 5 \pm 12)$, $(65^{+5}_{-4} \pm 13)$ or $(215 ^{+8}_{-5} \pm 11)~\rm eV$, and $Γ^{e^{+}e^{-}}_{ϕ(2170)}\cdot \mathcal{B}[ϕ(2170)\toηϕ] = (0.56^{+0.03}_{-0.02} \pm 0.07)$, $(0.36^{+0.05}_{-0.03} \pm 0.07)$, $(38 \pm 1 \pm 5)$ or $(41\pm 2 \pm 6)~\rm eV$, respectively. We also search for the production of $X(1750)\to ηϕ$ and the significance is only $2.0σ$, then we determine the upper limit of $Γ^{e^{+}e^{-}}_{X(1750)}\cdot \mathcal{B}[X(1750)\toηϕ]$ at $90\%$ confidence level.

hep-ex

Special Transition and Extraordinary Phase on the Surface of a Two-Dimensional Quantum Heisenberg Antiferromagnet

Continuous phase transitions exhibit richer critical phenomena on the surface than in the bulk, because distinct surface universality classes can be realized at the same bulk critical point by tuning the surface interactions. The exploration of surface critical behavior provides a window looking into higher-dimensional boundary conformal field theories. In this work, we study the surface critical behavior of a two-dimensional (2D) quantum critical Heisenberg model by tuning the surface coupling strength, and discover a direct special transition on the surface from the ordinary phase into an extraordinary phase. The extraordinary phase has a long-range antiferromagnetic order on the surface, in sharp contrast to the logarithmic decaying spin correlations in the 3D classical O(3) model. The special transition point has a new set of critical exponents, $y_{s}=0.86(4)$ and $η_{\parallel}=-0.33(1)$, which are distinct from the special transition of the classical O(3) model and indicate a new surface universality class of the 3D O(3) Wilson-Fisher theory.

cond-mat.str-el

A Multi-Stage Triple-Path Method for Speech Separation in Noisy and Reverberant Environments

In noisy and reverberant environments, the performance of deep learning-based speech separation methods drops dramatically because previous methods are not designed and optimized for such situations. To address this issue, we propose a multi-stage end-to-end learning method that decouples the difficult speech separation problem in noisy and reverberant environments into three sub-problems: speech denoising, separation, and de-reverberation. The probability and speed of searching for the optimal solution of the speech separation model are improved by reducing the solution space. Moreover, since the channel information of the audio sequence in the time domain is crucial for speech separation, we propose a triple-path structure capable of modeling the channel dimension of audio sequences. Experimental results show that the proposed multi-stage triple-path method can improve the performance of speech separation models at the cost of little model parameter increment.

cs.SD

Multi-Dimensional and Multi-Scale Modeling for Speech Separation Optimized by Discriminative Learning

Transformer has shown advanced performance in speech separation, benefiting from its ability to capture global features. However, capturing local features and channel information of audio sequences in speech separation is equally important. In this paper, we present a novel approach named Intra-SE-Conformer and Inter-Transformer (ISCIT) for speech separation. Specifically, we design a new network SE-Conformer that can model audio sequences in multiple dimensions and scales, and apply it to the dual-path speech separation framework. Furthermore, we propose Multi-Block Feature Aggregation to improve the separation effect by selectively utilizing information from the intermediate blocks of the separation network. Meanwhile, we propose a speaker similarity discriminative loss to optimize the speech separation model to address the problem of poor performance when speakers have similar voices. Experimental results on the benchmark datasets WSJ0-2mix and WHAM! show that ISCIT can achieve state-of-the-art results.

cs.SD

Speech Emotion Recognition with Global-Aware Fusion on Multi-scale Feature Representation

Speech Emotion Recognition (SER) is a fundamental task to predict the emotion label from speech data. Recent works mostly focus on using convolutional neural networks~(CNNs) to learn local attention map on fixed-scale feature representation by viewing time-varied spectral features as images. However, rich emotional feature at different scales and important global information are not able to be well captured due to the limits of existing CNNs for SER. In this paper, we propose a novel GLobal-Aware Multi-scale (GLAM) neural network (The code is available at https://github.com/lixiangucas01/GLAM) to learn multi-scale feature representation with global-aware fusion module to attend emotional information. Specifically, GLAM iteratively utilizes multiple convolutional kernels with different scales to learn multiple feature representation. Then, instead of using attention-based methods, a simple but effective global-aware fusion module is applied to grab most important emotional information globally. Experiments on the benchmark corpus IEMOCAP over four emotions demonstrates the superiority of our proposed model with 2.5% to 4.5% improvements on four common metrics compared to previous state-of-the-art approaches.

cs.SD

Surface critical behaviors of coupled Haldane chains

The special surface transition at (2+1)-dimensional quantum critical point is precluded in corresponding classical critical point. The mechanism of such behavior which is only found in dimerized Heisenberg models so far is still under debate. To illuminate the role of symmetry protected topological (SPT) phase in inducing such nonordinary behaviors, we study a system on a two-dimensional square lattice consisted by interacting spin-1 Haldane chains, which has a genuine SPT phase--the Haldane phase--at weak interchain interactions and a quantum critical point belonging to the classical 3D O(3) universality class to the Néel phase. Different from models studied previously, there is no dimerization in the current model. Cutting the system along the chain direction or perpendicular to the chain direction exposes two different surfaces. Using unbiased quantum Monte Carlo simulations, we find that the two different types of surface show completely different surface critical behaviors at the bulk critical point, resulted from different surface states in the SPT phase. For the system with surfaces along the chain direction, the surface critical behavior is of ordinary type of the bulk 3D O(3) critical point, while for the surfaces perpendicular to the chain direction, the surface critical behavior is nonordinary, consistent with special transitions found in dimerized Heisenberg models. Our numerical results demonstrate that the gapless surface state in the gapped SPT phase together with the gapless mode of critical point is a pure quantum scenario that leads to the nonordinary transition.

cond-mat.str-el

SQLFlow: A Bridge between SQL and Machine Learning

Industrial AI systems are mostly end-to-end machine learning (ML) workflows. A typical recommendation or business intelligence system includes many online micro-services and offline jobs. We describe SQLFlow for developing such workflows efficiently in SQL. SQL enables developers to write short programs focusing on the purpose (what) and ignoring the procedure (how). Previous database systems extended their SQL dialect to support ML. SQLFlow (https://sqlflow.org/sqlflow ) takes another strategy to work as a bridge over various database systems, including MySQL, Apache Hive, and Alibaba MaxCompute, and ML engines like TensorFlow, XGBoost, and scikit-learn. We extended SQL syntax carefully to make the extension working with various SQL dialects. We implement the extension by inventing a collaborative parsing algorithm. SQLFlow is efficient and expressive to a wide variety of ML techniques -- supervised and unsupervised learning; deep networks and tree models; visual model explanation in addition to training and prediction; data processing and feature extraction in addition to ML. SQLFlow compiles a SQL program into a Kubernetes-native workflow for fault-tolerable execution and on-cloud deployment. Current industrial users include Ant Financial, DiDi, and Alibaba Group.

cs.DB