SearcharxivSearch

arXiv subjects

Ruofei Zhang

Publications and source records attributed to Ruofei Zhang.

At least 19 recordsLinked to original sources

Dense Cores in the Vicinity of an HII Region

Massive stars strongly influence their surroundings through radiative and mechanical feedback, but its effects on dense gas structures at sub-pc scales remain poorly constrained. We investigate how feedback from a newly formed massive star affects dense cores in the filamentary molecular cloud IRAS 18530+0215. We analyze ALMA Band 6 observations of 1.3 mm dust continuum and DCN, N$_2$D$^+$, and $^{13}$CS line emission, together with VLA K-band continuum and NH$_3$ observations. Dense cores are identified with astrodendro, and their temperatures, masses, velocity dispersions, and virial parameters are derived. The dynamical state of the ultra-compact H II region is examined through energy and pressure estimates. The H II region has a radius of $\sim$0.1 pc and an expansion velocity of $\sim$2.5 km s$^{-1}$, corresponding to a shell dynamical age of $\sim$0.06 Myr. DCN and $^{13}$CS cores are concentrated near the H II region, whereas N$_2$D$^+$ cores preferentially lie farther away. Core temperatures and velocity dispersions decrease with projected distance from the H II region. Virial parameters increase within the inner $\sim$0.3 pc but decline sharply beyond this scale, while core masses show no significant trend with distance. Strong star formation signatures are found at $\sim$0.2 pc, whereas more distant regions still host quiescent, cold dense cores. The compact H II region appears trapped or choked within $\sim$0.1 pc, while its feedback extends to at least $\sim$0.3 pc. Within this region, feedback enhances core velocity dispersions, gas temperatures, and virial parameters, with no evidence that it promotes the formation of more massive dense cores.

astro-ph.GA

Trace the Self-Gravitating Gas Using CO Isotopologues

Recent studies have shown that the star formation rate (SFR) correlates tightly and linearly with the mass of gravitationally bound gas, which can be delineated from the power-law tail of the column-density probability distribution function ($N$-PDF) derived from dust emission observations. This relationship holds across four orders of magnitude within the Milky Way--spanning low-mass to high-mass star-forming regions and encompassing the extreme environment of the Central Molecular Zone. Building on this framework, we present a new approach for estimating the mass of gravitationally bound gas in molecular clouds using multi-line CO isotopologue observations. Our sample includes 16 molecular clouds with robust detections in $^{12}$CO, $^{13}$CO, and C$^{18}$O $J$ = 1-0, spanning both massive inner Galaxy clouds and nearby star-forming regions. We find that the $N$-PDFs derived from combined CO isotopologue data recover the characteristic log-normal plus power-law profiles seen in dust-based studies. The mass and spatial distribution of the self-gravitating structures estimated from both dust-based and CO-based methods agree well throughout the sample. This indicates that the CO isotopologue combination can robustly trace the self-gravitating component via the $N$-PDF method and provides a reliable, scalable, and velocity-resolved alternative to dust emission for identifying the star-forming gas in molecular clouds.

astro-ph.GA

Tails of Gravity: Persistence of Star Formation in the CMZ Environment

We characterize star-forming gas in six molecular clouds (Sgr B1-off, Sgr B2, Sgr C, the 20 km s$^{-1}$ and 50 km s$^{-1}$ molecular clouds, and the Brick) in the Galactic central molecular zone (CMZ), and compare their star-forming activities with those in molecular clouds outside the CMZ. Using multi-band continuum observations taken from ${\it Planck}$, ${\it Herschel}$, JCMT/SCUBA-2, and CSO/SHARC2, we derived 8.5" resolution column density maps for the CMZ clouds and evaluated the column density probability distribution functions (N-PDFs). With the archival Atacama Large Millimeter/submillimeter Array (ALMA) 1.3 mm dust continuum data, we further evaluated the mass of the most massive cores ($M_{\rm core}^{\rm ma x}$). We find that the N-PDFs of four of the selected CMZ clouds are well described by a piecewise log-normal + power-law function, while the N-PDFs of the remaining two can be approximated by log-normal functions. In the first four targets, the masses in the power-law component ($M_{\rm gas}^{\rm bound}$), $M_{\rm core}^{\rm max}$, and star formation rate (SFR) are correlated. These correlations are very similar to those derived from low-mass clouds in the Solar neighborhood and massive star-forming regions on the Galactic disk. These findings lead to our key hypotheses: (1) In the extreme environment of the CMZ, the power-law component in the N-PDF also represents self-gravitationally bound gas structures, and (2) evolution and star-forming activities of self-gravitationally bound gas structures may be self-regulated, insensitive to the exterior environment on $\gtrsim$5-10 pc scales.

astro-ph.GA

Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice

Simultaneous Interpretation (SI) represents one of the most daunting frontiers in the translation industry, with product-level automatic systems long plagued by intractable challenges: subpar transcription and translation quality, lack of real-time speech generation, multi-speaker confusion, and translated speech inflation, especially in long-form discourses. In this study, we introduce Seed-LiveInterpret 2.0, an end-to-end SI model that delivers high-fidelity, ultra-low-latency speech-to-speech generation with voice cloning capabilities. As a fully operational product-level solution, Seed-LiveInterpret 2.0 tackles these challenges head-on through our novel duplex speech-to-speech understanding-generating framework. Experimental results demonstrate that through large-scale pretraining and reinforcement learning, the model achieves a significantly better balance between translation accuracy and latency, validated by human interpreters to exceed 70% correctness in complex scenarios. Notably, Seed-LiveInterpret 2.0 outperforms commercial SI solutions by significant margins in translation quality, while slashing the average latency of cloned speech from nearly 10 seconds to a near-real-time 3 seconds, which is around a near 70% reduction that drastically enhances practical usability.

cs.CL

The Double-Episode Jet Genesis of the eROSITA and Fermi Bubbles

The Fermi and eROSITA bubbles are giant gamma-ray and X-ray lobes in the Milky Way, extending up to $\sim$50{\deg} and ~$\sim$80{\deg} in galactic latitude, respectively, yet their origins remain debated. Using three-dimensional magnetohydrodynamic simulations, we investigate a scenario in which two temporally separated episodes of active galactic nucleus (AGN) jets launched from the Galactic center produce the bubbles, with each structure bounded by a forward shock. Our simulations reveal that the first jet pair, launched 15 Myr ago, forms the outer eROSITA bubbles (extending to $\sim$18 kpc), while the second, launched 5 Myr ago, creates the nested Fermi bubbles ($\sim$10 kpc height). This model broadly reproduces the observed elongated morphology, multi-band X-ray surface brightness distribution, O VIII/O VII line ratios, radio ridge structures, and gamma-ray emissions of the bubbles. Cosmic-ray electrons are accelerated \textit{in situ} at the shock fronts, explaining the sharp edges and nearly uniform gamma-ray surface brightness distribution of Fermi bubbles. The results suggest that the eROSITA and Fermi bubbles encode a time-resolved record of episodic AGN activity in the Galactic center, providing a physically motivated framework for interpreting their multi-wavelength properties.

astro-ph.HE

Healing Unsafe Dialogue Responses with Weak Supervision Signals

Recent years have seen increasing concerns about the unsafe response generation of large-scale dialogue systems, where agents will learn offensive or biased behaviors from the real-world corpus. Some methods are proposed to address the above issue by detecting and replacing unsafe training examples in a pipeline style. Though effective, they suffer from a high annotation cost and adapt poorly to unseen scenarios as well as adversarial attacks. Besides, the neglect of providing safe responses (e.g. simply replacing with templates) will cause the information-missing problem of dialogues. To address these issues, we propose an unsupervised pseudo-label sampling method, TEMP, that can automatically assign potential safe responses. Specifically, our TEMP method groups responses into several clusters and samples multiple labels with an adaptively sharpened sampling strategy, inspired by the observation that unsafe samples in the clusters are usually few and distribute in the tail. Extensive experiments in chitchat and task-oriented dialogues show that our TEMP outperforms state-of-the-art models with weak supervision signals and obtains comparable results under unsupervised learning settings.

cs.CL

MERGE: Fast Private Text Generation

The drastic increase in language models' parameters has led to a new trend of deploying models in cloud servers, raising growing concerns about private inference for Transformer-based models. Existing two-party privacy-preserving techniques, however, only take into account natural language understanding (NLU) scenarios. Private inference in natural language generation (NLG), crucial for applications like translation and code completion, remains underexplored.In addition, previous privacy-preserving techniques suffer from convergence issues during model training and exhibit poor inference speed when used with NLG models due to the neglect of time-consuming operations in auto-regressive generations. To address these issues, we propose a fast private text generation framework for Transformer-based language models, namely MERGE.MERGE reuses the output hidden state as the word embedding to bypass the embedding computation and reorganize the linear operations in the Transformer module to accelerate the forward procedure. Extensive experiments show that MERGE achieves a 26.5x speedup to the vanilla encrypted model under the sequence length 512, and reduces 80\% communication cost, with an up to 10x speedup to state-of-the-art approximated models.

cs.CL

SwiftPruner: Reinforced Evolutionary Pruning for Efficient Ad Relevance

Ad relevance modeling plays a critical role in online advertising systems including Microsoft Bing. To leverage powerful transformers like BERT in this low-latency setting, many existing approaches perform ad-side computations offline. While efficient, these approaches are unable to serve cold start ads, resulting in poor relevance predictions for such ads. This work aims to design a new, low-latency BERT via structured pruning to empower real-time online inference for cold start ads relevance on a CPU platform. Our challenge is that previous methods typically prune all layers of the transformer to a high, uniform sparsity, thereby producing models which cannot achieve satisfactory inference speed with an acceptable accuracy. In this paper, we propose SwiftPruner - an efficient framework that leverages evolution-based search to automatically find the best-performing layer-wise sparse BERT model under the desired latency constraint. Different from existing evolution algorithms that conduct random mutations, we propose a reinforced mutator with a latency-aware multi-objective reward to conduct better mutations for efficiently searching the large space of layer-wise sparse models. Extensive experiments demonstrate that our method consistently achieves higher ROC AUC and lower latency than the uniform sparse baseline and state-of-the-art search methods. Remarkably, under our latency requirement of 1900us on CPU, SwiftPruner achieves a 0.86% higher AUC than the state-of-the-art uniform sparse baseline for BERT-Mini on a large scale real-world dataset. Online A/B testing shows that our model also achieves a significant 11.7% cut in the ratio of defective cold start ads with satisfactory real-time serving latency.

cs.IR

A Self-Paced Mixed Distillation Method for Non-Autoregressive Generation

Non-Autoregressive generation is a sequence generation paradigm, which removes the dependency between target tokens. It could efficiently reduce the text generation latency with parallel decoding in place of token-by-token sequential decoding. However, due to the known multi-modality problem, Non-Autoregressive (NAR) models significantly under-perform Auto-regressive (AR) models on various language generation tasks. Among the NAR models, BANG is the first large-scale pre-training model on English un-labeled raw text corpus. It considers different generation paradigms as its pre-training tasks including Auto-regressive (AR), Non-Autoregressive (NAR), and semi-Non-Autoregressive (semi-NAR) information flow with multi-stream strategy. It achieves state-of-the-art performance without any distillation techniques. However, AR distillation has been shown to be a very effective solution for improving NAR performance. In this paper, we propose a novel self-paced mixed distillation method to further improve the generation quality of BANG. Firstly, we propose the mixed distillation strategy based on the AR stream knowledge. Secondly, we encourage the model to focus on the samples with the same modality by self-paced learning. The proposed self-paced mixed distillation algorithm improves the generation quality and has no influence on the inference latency. We carry out extensive experiments on summarization and question generation tasks to validate the effectiveness. To further illustrate the commercial value of our approach, we conduct experiments on three generation tasks in real-world advertisements applications. Experimental results on commercial data show the effectiveness of the proposed model. Compared with BANG, it achieves significant BLEU score improvement. On the other hand, compared with auto-regressive generation method, it achieves more than 7x speedup.

cs.CL

Taming Sparsely Activated Transformer with Stochastic Experts

Sparsely activated models (SAMs), such as Mixture-of-Experts (MoE), can easily scale to have outrageously large amounts of parameters without significant increase in computational cost. However, SAMs are reported to be parameter inefficient such that larger models do not always lead to better performance. While most on-going research focuses on improving SAMs models by exploring methods of routing inputs to experts, our analysis reveals that such research might not lead to the solution we expect, i.e., the commonly-used routing methods based on gating mechanisms do not work better than randomly routing inputs to experts. In this paper, we propose a new expert-based model, THOR (Transformer witH StOchastic ExpeRts). Unlike classic expert-based models, such as the Switch Transformer, experts in THOR are randomly activated for each input during training and inference. THOR models are trained using a consistency regularized loss, where experts learn not only from training data but also from other experts as teachers, such that all the experts make consistent predictions. We validate the effectiveness of THOR on machine translation tasks. Results show that THOR models are more parameter efficient in that they significantly outperform the Transformer and MoE models across various settings. For example, in multilingual translation, THOR outperforms the Switch Transformer by 2 BLEU scores, and obtains the same BLEU score as that of a state-of-the-art MoE model that is 18 times larger. Our code is publicly available at: https://github.com/microsoft/Stochastic-Mixture-of-Experts.

cs.CL

KFCNet: Knowledge Filtering and Contrastive Learning Network for Generative Commonsense Reasoning

Pre-trained language models have led to substantial gains over a broad range of natural language processing (NLP) tasks, but have been shown to have limitations for natural language generation tasks with high-quality requirements on the output, such as commonsense generation and ad keyword generation. In this work, we present a novel Knowledge Filtering and Contrastive learning Network (KFCNet) which references external knowledge and achieves better generation performance. Specifically, we propose a BERT-based filter model to remove low-quality candidates, and apply contrastive learning separately to each of the encoder and decoder, within a general encoder--decoder architecture. The encoder contrastive module helps to capture global target semantics during encoding, and the decoder contrastive module enhances the utility of retrieved prototypes while learning general features. Extensive experiments on the CommonGen benchmark show that our model outperforms the previous state of the art by a large margin: +6.6 points (42.5 vs. 35.9) for BLEU-4, +3.7 points (33.3 vs. 29.6) for SPICE, and +1.3 points (18.3 vs. 17.0) for CIDEr. We further verify the effectiveness of the proposed contrastive module on ad keyword generation, and show that our model has potential commercial value.

cs.CL

FastSeq: Make Sequence Generation Faster

Transformer-based models have made tremendous impacts in natural language generation. However the inference speed is a bottleneck due to large model size and intensive computing involved in auto-regressive decoding process. We develop FastSeq framework to accelerate sequence generation without accuracy loss. The proposed optimization techniques include an attention cache optimization, an efficient algorithm for detecting repeated n-grams, and an asynchronous generation pipeline with parallel I/O. These optimizations are general enough to be applicable to Transformer-based models (e.g., T5, GPT2, and UniLM). Our benchmark results on a set of widely used and diverse models demonstrate 4-9x inference speed gain. Additionally, FastSeq is easy to use with a simple one-line code change. The source code is available at https://github.com/microsoft/fastseq.

cs.CL

EL-Attention: Memory Efficient Lossless Attention for Generation

Transformer model with multi-head attention requires caching intermediate results for efficient inference in generation tasks. However, cache brings new memory-related costs and prevents leveraging larger batch size for faster speed. We propose memory-efficient lossless attention (called EL-attention) to address this issue. It avoids heavy operations for building multi-head keys and values, cache for them is not needed. EL-attention constructs an ensemble of attention results by expanding query while keeping key and value shared. It produces the same result as multi-head attention with less GPU memory and faster inference speed. We conduct extensive experiments on Transformer, BART, and GPT-2 for summarization and question generation tasks. The results show EL-attention speeds up existing models by 1.6x to 5.3x without accuracy loss.

cs.CL

ProphetNet-X: Large-Scale Pre-training Models for English, Chinese, Multi-lingual, Dialog, and Code Generation

Now, the pre-training technique is ubiquitous in natural language processing field. ProphetNet is a pre-training based natural language generation method which shows powerful performance on English text summarization and question generation tasks. In this paper, we extend ProphetNet into other domains and languages, and present the ProphetNet family pre-training models, named ProphetNet-X, where X can be English, Chinese, Multi-lingual, and so on. We pre-train a cross-lingual generation model ProphetNet-Multi, a Chinese generation model ProphetNet-Zh, two open-domain dialog generation models ProphetNet-Dialog-En and ProphetNet-Dialog-Zh. And also, we provide a PLG (Programming Language Generation) model ProphetNet-Code to show the generation performance besides NLG (Natural Language Generation) tasks. In our experiments, ProphetNet-X models achieve new state-of-the-art performance on 10 benchmarks. All the models of ProphetNet-X share the same model structure, which allows users to easily switch between different models. We make the code and models publicly available, and we will keep updating more pre-training models and finetuning scripts.

cs.CL

Mask Attention Networks: Rethinking and Strengthen Transformer

Transformer is an attention-based neural network, which consists of two sublayers, namely, Self-Attention Network (SAN) and Feed-Forward Network (FFN). Existing research explores to enhance the two sublayers separately to improve the capability of Transformer for text representation. In this paper, we present a novel understanding of SAN and FFN as Mask Attention Networks (MANs) and show that they are two special cases of MANs with static mask matrices. However, their static mask matrices limit the capability for localness modeling in text representation learning. We therefore introduce a new layer named dynamic mask attention network (DMAN) with a learnable mask matrix which is able to model localness adaptively. To incorporate advantages of DMAN, SAN, and FFN, we propose a sequential layered structure to combine the three types of layers. Extensive experiments on various tasks, including neural machine translation and text summarization demonstrate that our model outperforms the original Transformer.

cs.CL

TextGNN: Improving Text Encoder via Graph Neural Network in Sponsored Search

Text encoders based on C-DSSM or transformers have demonstrated strong performance in many Natural Language Processing (NLP) tasks. Low latency variants of these models have also been developed in recent years in order to apply them in the field of sponsored search which has strict computational constraints. However these models are not the panacea to solve all the Natural Language Understanding (NLU) challenges as the pure semantic information in the data is not sufficient to fully identify the user intents. We propose the TextGNN model that naturally extends the strong twin tower structured encoders with the complementary graph information from user historical behaviors, which serves as a natural guide to help us better understand the intents and hence generate better language representations. The model inherits all the benefits of twin tower models such as C-DSSM and TwinBERT so that it can still be used in the low latency environment while achieving a significant performance gain than the strong encoder-only counterpart baseline models in both offline evaluations and online production system. In offline experiments, the model achieves a 0.14% overall increase in ROC-AUC with a 1% increased accuracy for long-tail low-frequency Ads, and in the online A/B testing, the model shows a 2.03% increase in Revenue Per Mille with a 2.32% decrease in Ad defect rate.

cs.CL

BANG: Bridging Autoregressive and Non-autoregressive Generation with Large Scale Pretraining

In this paper, we propose BANG, a new pretraining model to Bridge the gap between Autoregressive (AR) and Non-autoregressive (NAR) Generation. AR and NAR generation can be uniformly regarded as to what extent previous tokens can be attended, and BANG bridges AR and NAR generation by designing a novel model structure for large-scale pretraining. The pretrained BANG model can simultaneously support AR, NAR and semi-NAR generation to meet different requirements. Experiments on question generation (SQuAD 1.1), summarization (XSum) and dialogue generation (PersonaChat) show that BANG improves NAR and semi-NAR performance significantly as well as attaining comparable performance with strong AR pretrained models. Compared with the semi-NAR strong baselines, BANG achieves absolute improvements of 14.01 and 5.24 in the overall scores of SQuAD 1.1 and XSum, respectively. In addition, BANG achieves absolute improvements of 10.73, 6.39 and 5.90 in the overall scores of SQuAD, XSUM and PersonaChat respectively compared with the strong NAR baselines.

cs.CL

An Enhanced Knowledge Injection Model for Commonsense Generation

Commonsense generation aims at generating plausible everyday scenario description based on a set of provided concepts. Digging the relationship of concepts from scratch is non-trivial, therefore, we retrieve prototypes from external knowledge to assist the understanding of the scenario for better description generation. We integrate two additional modules, namely position indicator and scaling module, into the pretrained encoder-decoder model for prototype modeling to enhance the knowledge injection procedure. We conduct experiment on CommonGen benchmark, and experimental results show that our method significantly improves the performance on all the metrics.

cs.CL