SearcharxivSearch

arXiv subjects

Yijia Liu

Publications and source records attributed to Yijia Liu.

At least 19 recordsLinked to original sources

MISO: Model-Internal-State-Guided Optimization for Ranking Models

Ranking models are repeatedly refined within established model families, yet the choice of which component to scale, replace, or retire is often guided by expensive trial-and-error. We present Model Internal State Optimization (MISO), a systems workflow that uses model internal states (MIS), including parameters, activations, gradients, and normalization statistics, to prioritize such local optimization decisions. MISO extracts MIS from a trained ranking model, aggregates them into ranking, alignment, and comparison signals, and converts those signals into a small set of interpretable candidate edits. Because MIS are re-extracted after each retraining cycle, MISO naturally supports an adaptive optimization workflow that tracks evolving model behavior as data distributions and system requirements shift over time. In an ads ranking case study, MISO improves normalized entropy while requiring substantially fewer validation runs than expert-driven and black-box scaling workflows, offering a practical middle ground between manual tuning and opaque automated search.

cs.IR

NO molecule in massive star forming regions

Context. Among diatomic molecules composed of the abundant elements C, N and O, NO has been detected far less than the well studied CN and CO, making it a crucial yet under-observed component in nitrogen-containing chemical networks. NO was thought to serve as a potential tracer of shocks, as evidenced with orders abundance enhancements reported in literature. Aims. Large-sample observations for NO molecule in widespread interstellar environments are needed to confirm if the enhancement of NO is due to shock chemistry or not. Methods. Single-point survey for NO lines around 150 GHz was carried out by Arizona Radio Observatory 12-meter telescope towards a sample of 36 massive star forming regions containing SiO emission, which include three evolutionary stages: 4 IRDCs, 6 protostars and 26 H II regions. Results. The NO emission was detected in 28 sources with a detection rate of 78%. Beam-averaged NO column densities and abundances relative to H2 were derived from integrated intensities of two main hyperfine lines. Correlations between NO and SiO in integrated intensity and relative abundance are similar to the corresponding correlations of c-C3H2, indicating that NO enrichment may not significantly involve pronounced shock activities, which coincides with the trend in line widths: NO is close to c-C3H2, both smaller than H2CO and far smaller than SiO. Conclusions. Observational evidence does not strongly support significant NO enhancement by shock chemistry in the observed sources, indicating that the formation of NO does not necessarily require shocks.

astro-ph.GA

The evolution of C4H and c-C3H2 in molecular cores

Linear C4H and cyclic c-C3H2, as small unsaturated hydrocarbons, are the key precursors to complex organic molecules and are critical components of the interstellar medium. We present on-the-fly mapping observations of C4H 9-8 lines, c-C3H2 2-1, H13CO+ 1-0, and H42 toward a sample of 22 massive star-forming regions using the IRAM 30m telescope. Our aim is to further explore the evolution of these carbon-chain molecules by combining observational results obtained in cold cores. We employed H13CO+ 1-0 and H42 as tracers to probe the positions of molecular cloud cores and ionised hydrogen regions (HII regions), respectively. One chemical model in particular, which includes gas, dust grain surface, and icy mantle phases for C4H and c-C3H2 molecules, was used to make comparisons with observed abundances. From mapping observations targeting 31 regions across 22 sources, C4H 9-8 (J = 19/2-17/2) and C4H 9-8 (J = 17/2-15/2) were detected in only 17 regions, while H13CO+ 1-0 and c-C3H2 2-1 were successfully detected in all 31 regions. We find that the emission of C4H 9-8 and c-C3H2 2-1 is concentrated at the edges of H42 emission regions. The C4H/H13CO+ and c-C3H2/H13CO+ relative abundance ratios range from 0.17 to 1.77 and 1.42 to 6.69, respectively, with a median C4H/c-C3H2 ratio of 0.13. By combining the observational results of cold cores, we find that C4H/H13CO+ and c-C3H2/H13CO+ ratios show a strong decreasing trend as molecular cores evolve. The decreasing trends in C4H/H13CO+ and c-C3H2/H13CO+ ratios imply that small unsaturated hydrocarbons can be consumed and converted into other organic molecules during the evolution of molecular cores. The spatial concentration of C4H and c-C3H2 emission at the edges of H42 regions further supports their role as precursors in the chemical pathways that lead to complex organic molecules in the interstellar medium.

astro-ph.GA

Target-Aware Early Stage Ranking

Early Stage Ranking (ESR) in large-scale recommendation systems is dominated by ''user--item decoupling'' Two Tower architectures, which scale efficiently but cannot capture fine-grained, target-aware user--item interactions directly. We propose Target-Aware Early Stage Ranking (TESR), which augments the Two Tower with a Mixture of Attention (MoA) module trained as a request-level sequence modeling over user history. MoA combines (i) Hard Matching Attention (HMA) to capture explicit categorical-ID level overlap signals between user history and candidate item, (ii) target-aware HSTU attention for implicit affinities conditioned on the candidate, and (iii) target dependent and independent cross-attention for symmetric user-item contextualization. On top of this, a Multi-Logit Parameterized Gating (MLPG) head amplifies these signals at scoring time. To keep latency within ESR budgets, we co-design the architecture with FP8 quantization, custom kernels, and a Torch Inductor compilation path. On a production deployment, TESR delivers consistent offline NE wins and online topline gains, and is, to our knowledge, the first deployment of full target-aware attention sequence modeling in an ESR stage at this scale.

cs.LG

FreeDriveRF: Monocular RGB Dynamic NeRF without Poses for Autonomous Driving via Point-Level Dynamic-Static Decoupling

Dynamic scene reconstruction for autonomous driving enables vehicles to perceive and interpret complex scene changes more precisely. Dynamic Neural Radiance Fields (NeRFs) have recently shown promising capability in scene modeling. However, many existing methods rely heavily on accurate poses inputs and multi-sensor data, leading to increased system complexity. To address this, we propose FreeDriveRF, which reconstructs dynamic driving scenes using only sequential RGB images without requiring poses inputs. We innovatively decouple dynamic and static parts at the early sampling level using semantic supervision, mitigating image blurring and artifacts. To overcome the challenges posed by object motion and occlusion in monocular camera, we introduce a warped ray-guided dynamic object rendering consistency loss, utilizing optical flow to better constrain the dynamic modeling process. Additionally, we incorporate estimated dynamic flow to constrain the pose optimization process, improving the stability and accuracy of unbounded scene reconstruction. Extensive experiments conducted on the KITTI and Waymo datasets demonstrate the superior performance of our method in dynamic scene modeling for autonomous driving.

cs.CV

Spatial distribution of C4H and c-C3H2 in cold molecular cores

C$_4$H and $c$-C$_3$H$_2$, as unsaturated hydrocarbon molecules, are important for forming large organic molecules in the interstellar medium. We present mapping observations of C$_4$H ($N$=9$-8$) lines, $c$-C$_3$H$_2$ ($J_{Ka,Kb}$=2$_{1,2}$-1$_{0,1}$) %at 85338.894 MHz and H$^{13}$CO$^+$ ($J$=1$-0$) %at 86754.2884 MHz toward 19 nearby cold molecular cores in the Milky Way with the IRAM 30m telescope. C$_4$H 9--8 was detected in 13 sources, while $c$-C$_3$H$_2$ was detected in 18 sources. The widely existing C$_4$H and $c$-C$_3$H$_2$ molecules in cold cores provide material to form large organic molecules. Different spatial distributions between C$_4$H 9--8 and $c$-C$_3$H$_2$ 2--1 were found. The relative abundances of these three molecules were obtained under the assumption of local thermodynamic equilibrium conditions with a fixed excitation temperature. The abundance ratio of C$_4$H to $c$-C$_3$H$_2$ ranged from 0.34 $\pm$ 0.09 in G032.93+02 to 4.65 $\pm$ 0.50 in G008.67+22. A weak correlation between C$_4$H/H$^{13}$CO$^+$ and $c$-C$_3$H$_2$/H$^{13}$CO$^+$ abundance ratios was found, with a correlation coefficient of 0.46, which indicates that there is no tight astrochemical connection between C$_4$H and $c$-C$_3$H$_2$ molecules.

astro-ph.GA

Dense Outflowing Molecular Gas in Massive Star-forming Regions

Dense outflowing gas, traced by transitions of molecules with large dipole moment, is important for understanding mass loss and feedback of massive star formation. HCN 3-2 and HCO$^+$ 3-2 are good tracers of dense outflowing molecular gas, which are closely related to active star formation. In this study, we present on-the-fly (OTF) mapping observations of HCN 3-2 and HCO$^+$ 3-2 toward a sample of 33 massive star-forming regions using the 10-m Submillimeter Telescope (SMT). With the spatial distribution of line wings of HCO$^+$ 3-2 and HCN 3-2, outflows are detected in 25 sources, resulting in a detection rate of 76$\%$. The optically thin H$^{13}$CN and H$^{13}$CO$^+$ 3-2 lines are used to identify line wings as outflows and estimate core mass. The mass $M_{out}$, momentum $P_{out}$, kinetic energy $E_{K}$, force $F_{out}$ and mass loss rate $\dot M_{out}$ of outflow and core mass, are obtained for each source. A sublinear tight correlation is found between the mass of dense molecular outflow and core mass, with an index of $\sim$ 0.8 and a correlation coefficient of 0.88.

astro-ph.GA

Implicit Generative Prior for Bayesian Neural Networks

Predictive uncertainty quantification is crucial for reliable decision-making in various applied domains. Bayesian neural networks offer a powerful framework for this task. However, defining meaningful priors and ensuring computational efficiency remain significant challenges, especially for complex real-world applications. This paper addresses these challenges by proposing a novel neural adaptive empirical Bayes (NA-EB) framework. NA-EB leverages a class of implicit generative priors derived from low-dimensional distributions. This allows for efficient handling of complex data structures and effective capture of underlying relationships in real-world datasets. The proposed NA-EB framework combines variational inference with a gradient ascent algorithm. This enables simultaneous hyperparameter selection and approximation of the posterior distribution, leading to improved computational efficiency. We establish the theoretical foundation of the framework through posterior and classification consistency. We demonstrate the practical applications of our framework through extensive evaluations on a variety of tasks, including the two-spiral problem, regression, 10 UCI datasets, and image classification tasks on both MNIST and CIFAR-10 datasets. The results of our experiments highlight the superiority of our proposed framework over existing methods, such as sparse variational Bayesian and generative models, in terms of prediction accuracy and uncertainty quantification.

cs.LG

On pseudo-Anosov autoequivalences

Motivated by results of Thurston, we prove that any autoequivalence of a triangulated category induces a filtration by triangulated subcategories, provided the existence of Bridgeland stability conditions. The filtration is given by the exponential growth rate of masses under iterates of the autoequivalence, and only depends on the choice of a connected component of the stability manifold. We then propose a new definition of pseudo-Anosov autoequivalences, and prove that our definition is more general than the one previously proposed by Dimitrov, Haiden, Katzarkov, and Kontsevich. We construct new examples of pseudo-Anosov autoequivalences on the derived categories of quintic Calabi-Yau threefolds and quiver Calabi-Yau categories. Finally, we prove that certain pseudo-Anosov autoequivalences on quiver 3-Calabi-Yau categories act hyperbolically on the space of Bridgeland stability conditions.

math.AG

VECO: Variable and Flexible Cross-lingual Pre-training for Language Understanding and Generation

Existing work in multilingual pretraining has demonstrated the potential of cross-lingual transferability by training a unified Transformer encoder for multiple languages. However, much of this work only relies on the shared vocabulary and bilingual contexts to encourage the correlation across languages, which is loose and implicit for aligning the contextual representations between languages. In this paper, we plug a cross-attention module into the Transformer encoder to explicitly build the interdependence between languages. It can effectively avoid the degeneration of predicting masked words only conditioned on the context in its own language. More importantly, when fine-tuning on downstream tasks, the cross-attention module can be plugged in or out on-demand, thus naturally benefiting a wider range of cross-lingual tasks, from language understanding to generation. As a result, the proposed cross-lingual model delivers new state-of-the-art results on various cross-lingual understanding tasks of the XTREME benchmark, covering text classification, sequence labeling, question answering, and sentence retrieval. For cross-lingual generation tasks, it also outperforms all existing cross-lingual models and state-of-the-art Transformer variants on WMT14 English-to-German and English-to-French translation datasets, with gains of up to 1~2 BLEU.

cs.CL

Lattice-BERT: Leveraging Multi-Granularity Representations in Chinese Pre-trained Language Models

Chinese pre-trained language models usually process text as a sequence of characters, while ignoring more coarse granularity, e.g., words. In this work, we propose a novel pre-training paradigm for Chinese -- Lattice-BERT, which explicitly incorporates word representations along with characters, thus can model a sentence in a multi-granularity manner. Specifically, we construct a lattice graph from the characters and words in a sentence and feed all these text units into transformers. We design a lattice position attention mechanism to exploit the lattice structures in self-attention layers. We further propose a masked segment prediction task to push the model to learn from rich but redundant information inherent in lattices, while avoiding learning unexpected tricks. Experiments on 11 Chinese natural language understanding tasks show that our model can bring an average increase of 1.5% under the 12-layer setting, which achieves new state-of-the-art among base-size models on the CLUE benchmarks. Further analysis shows that Lattice-BERT can harness the lattice structures, and the improvement comes from the exploration of redundant information and multi-granularity representations. Our code will be available at https://github.com/alibaba/pretrained-language-models/LatticeBERT.

cs.CL

Improving Biomedical Pretrained Language Models with Knowledge

Pretrained language models have shown success in many natural language processing tasks. Many works explore incorporating knowledge into language models. In the biomedical domain, experts have taken decades of effort on building large-scale knowledge bases. For example, the Unified Medical Language System (UMLS) contains millions of entities with their synonyms and defines hundreds of relations among entities. Leveraging this knowledge can benefit a variety of downstream tasks such as named entity recognition and relation extraction. To this end, we propose KeBioLM, a biomedical pretrained language model that explicitly leverages knowledge from the UMLS knowledge bases. Specifically, we extract entities from PubMed abstracts and link them to UMLS. We then train a knowledge-aware language model that firstly applies a text-only encoding layer to learn entity representation and applies a text-entity fusion encoding to aggregate entity representation. Besides, we add two training objectives as entity detection and entity linking. Experiments on the named entity recognition and relation extraction from the BLURB benchmark demonstrate the effectiveness of our approach. Further analysis on a collected probing dataset shows that our model has better ability to model medical knowledge.

cs.CL

Few-shot Slot Tagging with Collapsed Dependency Transfer and Label-enhanced Task-adaptive Projection Network

In this paper, we explore the slot tagging with only a few labeled support sentences (a.k.a. few-shot). Few-shot slot tagging faces a unique challenge compared to the other few-shot classification problems as it calls for modeling the dependencies between labels. But it is hard to apply previously learned label dependencies to an unseen domain, due to the discrepancy of label sets. To tackle this, we introduce a collapsed dependency transfer mechanism into the conditional random field (CRF) to transfer abstract label dependency patterns as transition scores. In the few-shot setting, the emission score of CRF can be calculated as a word's similarity to the representation of each label. To calculate such similarity, we propose a Label-enhanced Task-Adaptive Projection Network (L-TapNet) based on the state-of-the-art few-shot classification model -- TapNet, by leveraging label name semantics in representing labels. Experimental results show that our model significantly outperforms the strongest few-shot learning baseline by 14.64 F1 scores in the one-shot setting.

cs.CL

Entity-Consistent End-to-end Task-Oriented Dialogue System with KB Retriever

Querying the knowledge base (KB) has long been a challenge in the end-to-end task-oriented dialogue system. Previous sequence-to-sequence (Seq2Seq) dialogue generation work treats the KB query as an attention over the entire KB, without the guarantee that the generated entities are consistent with each other. In this paper, we propose a novel framework which queries the KB in two steps to improve the consistency of generated entities. In the first step, inspired by the observation that a response can usually be supported by a single KB row, we introduce a KB retrieval component which explicitly returns the most relevant KB row given a dialogue history. The retrieval result is further used to filter the irrelevant entities in a Seq2Seq response generation model to improve the consistency among the output entities. In the second step, we further perform the attention mechanism to address the most correlated KB column. Two methods are proposed to make the training feasible without labeled retrieval data, which include distant supervision and Gumbel-Softmax technique. Experiments on two publicly available task oriented dialog datasets show the effectiveness of our model by outperforming the baseline systems and producing entity-consistent responses.

cs.CL

Cross-Lingual BERT Transformation for Zero-Shot Dependency Parsing

This paper investigates the problem of learning cross-lingual representations in a contextual space. We propose Cross-Lingual BERT Transformation (CLBT), a simple and efficient approach to generate cross-lingual contextualized word embeddings based on publicly available pre-trained BERT models (Devlin et al., 2018). In this approach, a linear transformation is learned from contextual word alignments to align the contextualized embeddings independently trained in different languages. We demonstrate the effectiveness of this approach on zero-shot cross-lingual transfer parsing. Experiments show that our embeddings substantially outperform the previous state-of-the-art that uses static embeddings. We further compare our approach with XLM (Lample and Conneau, 2019), a recently proposed cross-lingual language model trained with massive parallel data, and achieve highly competitive results.

cs.CL

Few-Shot Sequence Labeling with Label Dependency Transfer and Pair-wise Embedding

While few-shot classification has been widely explored with similarity based methods, few-shot sequence labeling poses a unique challenge as it also calls for modeling the label dependencies. To consider both the item similarity and label dependency, we propose to leverage the conditional random fields (CRFs) in few-shot sequence labeling. It calculates emission score with similarity based methods and obtains transition score with a specially designed transfer mechanism. When applying CRF in the few-shot scenarios, the discrepancy of label sets among different domains makes it hard to use the label dependency learned in prior domains. To tackle this, we introduce the dependency transfer mechanism that transfers abstract label transition patterns. In addition, the similarity methods rely on the high quality sample representation, which is challenging for sequence labeling, because sense of a word is different when measuring its similarity to words in different sentences. To remedy this, we take advantage of recent contextual embedding technique, and further propose a pair-wise embedder. It provides additional certainty for word sense by embedding query and support sentence pairwisely. Experimental results on slot tagging and named entity recognition show that our model significantly outperforms the strongest few-shot learning baseline by 11.76 (21.2%) and 12.18 (97.7%) F1 scores respectively in the one-shot setting.

cs.CL

An AMR Aligner Tuned by Transition-based Parser

In this paper, we propose a new rich resource enhanced AMR aligner which produces multiple alignments and a new transition system for AMR parsing along with its oracle parser. Our aligner is further tuned by our oracle parser via picking the alignment that leads to the highest-scored achievable AMR graph. Experimental results show that our aligner outperforms the rule-based aligner in previous work by achieving higher alignment F1 score and consistently improving two open-sourced AMR parsers. Based on our aligner and transition system, we develop a transition-based AMR parser that parses a sentence into its AMR graph directly. An ensemble of our parsers with only words and POS tags as input leads to 68.4 Smatch F1 score.

cs.CL

Towards Better UD Parsing: Deep Contextualized Word Embeddings, Ensemble, and Treebank Concatenation

This paper describes our system (HIT-SCIR) submitted to the CoNLL 2018 shared task on Multilingual Parsing from Raw Text to Universal Dependencies. We base our submission on Stanford's winning system for the CoNLL 2017 shared task and make two effective extensions: 1) incorporating deep contextualized word embeddings into both the part of speech tagger and parser; 2) ensembling parsers trained with different initialization. We also explore different ways of concatenating treebanks for further improvements. Experimental results on the development data show the effectiveness of our methods. In the final evaluation, our system was ranked first according to LAS (75.84%) and outperformed the other systems by a large margin.

cs.CL