Searcharxiv⌕ Search

arXiv subjects

Yu Meng

Publications and source records attributed to Yu Meng.

At least 73 records · Page 4Linked to original sources

Unchosen Experts Can Contribute Too: Unleashing MoE Models' Power by Self-Contrast

Mixture-of-Experts (MoE) has emerged as a prominent architecture for scaling model size while maintaining computational efficiency. In MoE, each token in the input sequence activates a different subset of experts determined by a routing mechanism. However, the unchosen experts in MoE models do not contribute to the output, potentially leading to underutilization of the model's capacity. In this work, we first conduct exploratory studies to demonstrate that increasing the number of activated experts does not necessarily improve and can even degrade the output quality. Then, we show that output distributions from an MoE model using different routing strategies substantially differ, indicating that different experts do not always act synergistically. Motivated by these findings, we propose Self-Contrast Mixture-of-Experts (SCMoE), a training-free strategy that utilizes unchosen experts in a self-contrast manner during inference. In SCMoE, the next-token probabilities are determined by contrasting the outputs from strong and weak activation using the same MoE model. Our method is conceptually simple and computationally lightweight, as it incurs minimal latency compared to greedy decoding. Experiments on several benchmarks (GSM8K, StrategyQA, MBPP and HumanEval) demonstrate that SCMoE can consistently enhance Mixtral 8x7B's reasoning capability across various domains. For example, it improves the accuracy on GSM8K from 61.79 to 66.94. Moreover, combining SCMoE with self-consistency yields additional gains, increasing major@20 accuracy from 75.59 to 78.31.

cs.CL↗

SimPO: Simple Preference Optimization with a Reference-Free Reward

Direct Preference Optimization (DPO) is a widely used offline preference optimization algorithm that reparameterizes reward functions in reinforcement learning from human feedback (RLHF) to enhance simplicity and training stability. In this work, we propose SimPO, a simpler yet more effective approach. The effectiveness of SimPO is attributed to a key design: using the average log probability of a sequence as the implicit reward. This reward formulation better aligns with model generation and eliminates the need for a reference model, making it more compute and memory efficient. Additionally, we introduce a target reward margin to the Bradley-Terry objective to encourage a larger margin between the winning and losing responses, further improving the algorithm's performance. We compare SimPO to DPO and its latest variants across various state-of-the-art training setups, including both base and instruction-tuned models such as Mistral, Llama 3, and Gemma 2. We evaluate on extensive chat-based evaluation benchmarks, including AlpacaEval 2, MT-Bench, and Arena-Hard. Our results demonstrate that SimPO consistently and significantly outperforms existing approaches without substantially increasing response length. Specifically, SimPO outperforms DPO by up to 6.4 points on AlpacaEval 2 and by up to 7.5 points on Arena-Hard. Our top-performing model, built on Gemma-2-9B-it, achieves a 72.4% length-controlled win rate on AlpacaEval 2, a 59.1% win rate on Arena-Hard, and ranks 1st on Chatbot Arena among <10B models with real user votes.

cs.CL↗

MoNTA: Accelerating Mixture-of-Experts Training with Network-Traffc-Aware Parallel Optimization

The Mixture of Experts (MoE) is an advanced model architecture in the industry that combines multiple specialized expert models from various domains into a single supermodel. This approach enables the model to scale without significantly increasing the computational costs of training and inference, while maximizing model performance. However, current distributed training frameworks do not consider the ultimate optimization of communication, especially for large base models. This paper proposes a network-traffic-aware parallel optimization method that selects the optimal parallel strategy based on the communication volume, and the training cluster's inter-node and intra-node network topologies. Compared to the DeepSpeed, MoNTA achieves an 8x increase in AllToAll communication performance under 8-card tensor parallelism. Compared to the baseline, training a 2x70B model using 16 A800 cards, with an 8K sequence, results in a 13% overall latency performance improvement. Project Page: https://github.com/EnflameTechnology/DeepSpeed.

cs.LG↗

Grasping the Essentials: Tailoring Large Language Models for Zero-Shot Relation Extraction

Relation extraction (RE) aims to identify semantic relationships between entities within text. Despite considerable advancements, existing models predominantly require extensive annotated training data, which is both costly and labor-intensive to collect. Moreover, these models often struggle to adapt to new or unseen relations. Few-shot learning, aiming to lessen annotation demands, typically provides incomplete and biased supervision for target relations, leading to degraded and unstable performance. To accurately and explicitly describe relation semantics while minimizing annotation demands, we explore the definition only zero-shot RE setting where only relation definitions expressed in natural language are used to train a RE model. We introduce REPaL, comprising three stages: (1) We leverage large language models (LLMs) to generate initial seed instances from relation definitions and an unlabeled corpus. (2) We fine-tune a bidirectional Small Language Model (SLM) with initial seeds to learn relations for the target domain. (3) We expand pattern coverage and mitigate bias from initial seeds by integrating feedback from the SLM's predictions on the unlabeled corpus and the synthesis history. To accomplish this, we leverage the multi-turn conversation ability of LLMs to generate new instances in follow-up dialogues, informed by both the feedback and synthesis history. Studies reveal that definition-oriented seed synthesis enhances pattern coverage whereas indiscriminately increasing seed quantity leads to performance saturation. Experiments on two datasets show REPaL significantly improved cost-effective zero-shot performance by large margins.

cs.CL↗

First lattice QCD calculation of $J/ψ$ semileptonic decay containing $D$ and $D_s$ particles

We perform the first lattice calculation on the semileptonic decay of $J/ψ$ using the (2+1)-flavor Wilson-clover gauge ensembles generated by CLQCD collaboration. Three gauge ensembles with different lattice spacings, from 0.0519 fm to 0.1053 fm, and pion masses, $m_π\sim$ 300 MeV, are utilized. After a naive continuum extrapolation using three lattice spacings, we obtain $\operatorname{Br}(J/ψ\rightarrow D_s eν_e)=1.90(6)(5)_{V_{cs}}\times 10^{-10}$ and $\operatorname{Br}(J/ψ\rightarrow D eν_e)=1.21(6)(9)_{V_{cd}}\times 10^{-11}$, where the first errors are statistical, and the second come from the uncertainties of CKM matrix element $V_{cs(d)}$. The ratios of the branching fractions between lepton $μ$ and $e$ are also calculated as $R_{J/ψ}(D_s)=0.97002(8)$ and $R_{J/ψ}(D)=0.97423(15)$ after performing a continuum limit including only $a^2$ term. The ratios provide necessary theoretical support for the future experimental test of lepton flavor universality.

hep-lat↗

Two Tales of Persona in LLMs: A Survey of Role-Playing and Personalization

The concept of persona, originally adopted in dialogue literature, has re-surged as a promising framework for tailoring large language models (LLMs) to specific context (e.g., personalized search, LLM-as-a-judge). However, the growing research on leveraging persona in LLMs is relatively disorganized and lacks a systematic taxonomy. To close the gap, we present a comprehensive survey to categorize the current state of the field. We identify two lines of research, namely (1) LLM Role-Playing, where personas are assigned to LLMs, and (2) LLM Personalization, where LLMs take care of user personas. Additionally, we introduce existing methods for LLM personality evaluation. To the best of our knowledge, we present the first survey for role-playing and personalization in LLMs under the unified view of persona. We continuously maintain a paper collection to foster future endeavors: https://github.com/MiuLab/PersonaLLM-Survey

cs.CL↗

Graph Chain-of-Thought: Augmenting Large Language Models by Reasoning on Graphs

Large language models (LLMs), while exhibiting exceptional performance, suffer from hallucinations, especially on knowledge-intensive tasks. Existing works propose to augment LLMs with individual text units retrieved from external knowledge corpora to alleviate the issue. However, in many domains, texts are interconnected (e.g., academic papers in a bibliographic graph are linked by citations and co-authorships) which form a (text-attributed) graph. The knowledge in such graphs is encoded not only in single texts/nodes but also in their associated connections. To facilitate the research of augmenting LLMs with graphs, we manually construct a Graph Reasoning Benchmark dataset called GRBench, containing 1,740 questions that can be answered with the knowledge from 10 domain graphs. Then, we propose a simple and effective framework called Graph Chain-of-thought (Graph-CoT) to augment LLMs with graphs by encouraging LLMs to reason on the graph iteratively. Each Graph-CoT iteration consists of three sub-steps: LLM reasoning, LLM-graph interaction, and graph execution. We conduct systematic experiments with three LLM backbones on GRBench, where Graph-CoT outperforms the baselines consistently. The code is available at https://github.com/PeterGriffinJin/Graph-CoT.

cs.CL↗

Establishing Knowledge Preference in Language Models

Language models are known to encode a great amount of factual knowledge through pretraining. However, such knowledge might be insufficient to cater to user requests, requiring the model to integrate external knowledge sources and adhere to user-provided specifications. When answering questions about ongoing events, the model should use recent news articles to update its response; when asked to provide recommendations, the model should prioritize user specifications over retrieved product reviews; when some facts are edited in the model, the updated facts should override all prior knowledge learned by the model even if they are conflicting. In all of the cases above, the model faces a decision between its own parametric knowledge, (retrieved) contextual knowledge, and user instruction knowledge. In this paper, we (1) unify such settings into the problem of knowledge preference and define a three-level preference hierarchy over these knowledge sources; (2) compile a collection of existing datasets IfQA, MQuAKE, and MRQA covering a combination of settings (with/without user specifications, with/without context documents) to systematically evaluate how well models obey the intended knowledge preference; and (3) propose a dataset synthesis method that composes diverse question-answer pairs with user assumptions and related context to directly fine-tune LMs for instilling the hierarchy of knowledge. We demonstrate that a 7B model, fine-tuned on only a few thousand examples automatically generated by our proposed method, effectively achieves superior performance (more than 18% improvement across all evaluation benchmarks) in adhering to the desired knowledge preference hierarchy.

cs.CL↗

Learning Multiplex Representations on Text-Attributed Graphs with One Language Model Encoder

In real-world scenarios, texts in a graph are often linked by multiple semantic relations (e.g., papers in an academic graph are referenced by other publications, written by the same author, or published in the same venue), where text documents and their relations form a multiplex text-attributed graph. Mainstream text representation learning methods use pretrained language models (PLMs) to generate one embedding for each text unit, expecting that all types of relations between texts can be captured by these single-view embeddings. However, this presumption does not hold particularly in multiplex text-attributed graphs. Along another line of work, multiplex graph neural networks (GNNs) directly initialize node attributes as a feature vector for node representation learning, but they cannot fully capture the semantics of the nodes' associated texts. To bridge these gaps, we propose METAG, a new framework for learning Multiplex rEpresentations on Text-Attributed Graphs. In contrast to existing methods, METAG uses one text encoder to model the shared knowledge across relations and leverages a small number of parameters per relation to derive relation-specific representations. This allows the encoder to effectively capture the multiplex structures in the graph while also preserving parameter efficiency. We conduct experiments on nine downstream tasks in five graphs from both academic and e-commerce domains, where METAG outperforms baselines significantly and consistently. The code is available at https://github.com/PeterGriffinJin/METAG.

cs.CL↗

Experimental test of generalized multipartite entropic uncertainty relations

Entropic uncertainty relation (EUR) formulates the restriction of the inherent uncertainty of quantum mechanics from the information-theoretic perspective. A tighter lower bound for uncertainty relations can provide information-theoretic security to quantum communication protocols. Recently, a generalized EUR (GEUR) for the measurement of multiple observables in arbitrary many-body systems has been formulated. Here, we experimentally test this GEUR using a four-photon entangled state with a controllable decoherence channel and show that for the tripartite scenario, the GEUR improves the entropic bound from Renes--Boileau's famous results. As an application, we further demonstrate an improvement of the secure key rate in quantum key distribution from the GEUR. Our results extend the test of EURs into multipartite regimes and may find applications in practical quantum cryptography tasks.

quant-ph↗

Lattice QCD calculation of the $D_s^{*}$ radiative decay with (2+1)-flavor Wilson-clover ensembles

We perform a lattice calculation on the radiative decay of $D_s^*$ using the (2+1)-flavor Wilson-clover gauge ensembles generated by CLQCD collaboration. A method allowing us to calculate the form factor with zero transfer momentum is proposed and applied to the radiative transition $D_s^*\rightarrow D_sγ$ and the Dalitz decay $D_s^*\rightarrow D_s e^+e^-$. After a continuum extrapolation using three lattice spacings, we obtain $Γ(D_s^*\rightarrow D_s γ)=0.0549(54)$ keV, where the error is purely statistical. The result is consistent with previous lattice calculations but with a error reduced to only a fifth of the before. The Dalitz decay rate is also calculated for the first time and the ratio with the radiative transition is found to be $R_{ee}=0.624(3)\%$. A total decay width of $D_s^*$ can then be determined as 0.0587(54) keV taking into account the experimental branching fraction. Combining with the most recent experimental measurement on the branching fraction of the purely leptonic decay $D_s^{+,*}\rightarrow e^+ν_e$, we obtain the quantity $f_{D_s^*}|V_{cs}|=(190.5^{+55.1}_{-41.7_{\textrm{stat.}}}\pm 12.6_{\textrm{syst.}})$ MeV, where the stat. is only the statistical error from the experiment, and syst. results from the experimental systematic uncertainty and the lattice statistical error. Our result leads to an improved systematic uncertainty compared to $42.7_{\textrm{syst.}}$ obtained using previous lattice prediction of total decay width $0.070(28)$ keV as the input.

hep-lat↗

Evaluating Large Language Models at Evaluating Instruction Following

As research in large language models (LLMs) continues to accelerate, LLM-based evaluation has emerged as a scalable and cost-effective alternative to human evaluations for comparing the ever increasing list of models. This paper investigates the efficacy of these ``LLM evaluators'', particularly in using them to assess instruction following, a metric that gauges how closely generated text adheres to the given instruction. We introduce a challenging meta-evaluation benchmark, LLMBar, designed to test the ability of an LLM evaluator in discerning instruction-following outputs. The authors manually curated 419 pairs of outputs, one adhering to instructions while the other diverging, yet may possess deceptive qualities that mislead an LLM evaluator, e.g., a more engaging tone. Contrary to existing meta-evaluation, we discover that different evaluators (i.e., combinations of LLMs and prompts) exhibit distinct performance on LLMBar and even the highest-scoring ones have substantial room for improvement. We also present a novel suite of prompting strategies that further close the gap between LLM and human evaluators. With LLMBar, we hope to offer more insight into LLM evaluators and foster future research in developing better instruction-following models.

cs.CL↗

Rapid Mobile App Development for Generative AI Agents on MIT App Inventor

The evolution of Artificial Intelligence (AI) stands as a pivotal force shaping our society, finding applications across diverse domains such as education, sustainability, and safety. Leveraging AI within mobile applications makes it easily accessible to the public, catalyzing its transformative potential. In this paper, we present a methodology for the rapid development of AI agent applications using the development platform provided by MIT App Inventor. To demonstrate its efficacy, we share the development journey of three distinct mobile applications: SynchroNet for fostering sustainable communities; ProductiviTeams for addressing procrastination; and iHELP for enhancing community safety. All three applications seamlessly integrate a spectrum of generative AI features, leveraging OpenAI APIs. Furthermore, we offer insights gleaned from overcoming challenges in integrating diverse tools and AI functionalities, aiming to inspire young developers to join our efforts in building practical AI agent applications.

cs.SE↗

Representation Deficiency in Masked Language Modeling

Masked Language Modeling (MLM) has been one of the most prominent approaches for pretraining bidirectional text encoders due to its simplicity and effectiveness. One notable concern about MLM is that the special $\texttt{[MASK]}$ symbol causes a discrepancy between pretraining data and downstream data as it is present only in pretraining but not in fine-tuning. In this work, we offer a new perspective on the consequence of such a discrepancy: We demonstrate empirically and theoretically that MLM pretraining allocates some model dimensions exclusively for representing $\texttt{[MASK]}$ tokens, resulting in a representation deficiency for real tokens and limiting the pretrained model's expressiveness when it is adapted to downstream data without $\texttt{[MASK]}$ tokens. Motivated by the identified issue, we propose MAE-LM, which pretrains the Masked Autoencoder architecture with MLM where $\texttt{[MASK]}$ tokens are excluded from the encoder. Empirically, we show that MAE-LM improves the utilization of model dimensions for real token representations, and MAE-LM consistently outperforms MLM-pretrained models across different pretraining settings and model sizes when fine-tuned on the GLUE and SQuAD benchmarks.

cs.CL↗

A universal programmable Gaussian Boson Sampler for drug discovery

Gaussian Boson Sampling (GBS) exhibits a unique ability to solve graph problems, such as finding cliques in complex graphs. It is noteworthy that many drug discovery tasks can be viewed as the clique-finding process, making them potentially suitable for quantum computation. However, to perform these tasks in their quantum-enhanced form, a large-scale quantum hardware with universal programmability is essential, which is yet to be achieved even with the most advanced GBS devices. Here, we construct a time-bin encoded GBS photonic quantum processor that is universal, programmable, and software-scalable. Our processor features freely adjustable squeezing parameters and can implement arbitrary unitary operations with a programmable interferometer. Using our processor, we have demonstrated the clique-finding task in a 32-node graph, where we found the maximum weighted clique with approximately twice the probability of success compared to classical sampling. Furthermore, a multifunctional quantum pharmaceutical platform is developed. This GBS processor is successfully used to execute two different drug discovery methods, namely molecular docking and RNA folding prediction. Our work achieves the state-of-the-art in GBS circuitry with its distinctive universal and programmable architecture which advances GBS towards real-world applications.

quant-ph↗

Isospin-$\frac{1}{2}$ $Dπ$ scattering and the $D_0^*$ resonance

Preliminary lattice QCD results for $Dπ$ scattering in isospin $I=\frac{1}{2}$ channel are presented. Utilizing the $N_f=2+1$ Wilson-Clover configuration at two volumes ($L^3 \times T=32^3 \times 96$ and $48^3 \times 96$) with the same lattice spacing ($a=0.07746(18)$ fm) and pion mass ($m_π\approx 303$ MeV), various two-particle operators in both the COM and the moving frames are constructed and the corresponding finite-volume spectra are determined from their correlation functions. The $S$ and $P$-wave scattering phase shifts are then extracted using the Lüscher approach, assuming negligible contributions from higher partial waves. A virtual state associated with the $D_0^*$ is also identified.

hep-lat↗

SCStory: Self-supervised and Continual Online Story Discovery

We present a framework SCStory for online story discovery, that helps people digest rapidly published news article streams in real-time without human annotations. To organize news article streams into stories, existing approaches directly encode the articles and cluster them based on representation similarity. However, these methods yield noisy and inaccurate story discovery results because the generic article embeddings do not effectively reflect the story-indicative semantics in an article and cannot adapt to the rapidly evolving news article streams. SCStory employs self-supervised and continual learning with a novel idea of story-indicative adaptive modeling of news article streams. With a lightweight hierarchical embedding module that first learns sentence representations and then article representations, SCStory identifies story-relevant information of news articles and uses them to discover stories. The embedding module is continuously updated to adapt to evolving news streams with a contrastive learning objective, backed up by two unique techniques, confidence-aware memory replay and prioritized-augmentation, employed for label absence and data scarcity problems. Thorough experiments on real and the latest news data sets demonstrate that SCStory outperforms existing state-of-the-art algorithms for unsupervised online story discovery.

cs.CL↗

A Spatially resolved X-ray Polarization map of the Vela Pulsar Wind Nebula

In this paper, we present a full spatially resolved polarization map for the Vela Pulsar Wind Nebula (PWN) observed by IXPE. By employing effective background discrimination techniques, our results show a remarkably high degree of local polarization in the outskirt region, exceeding 60% (55%) with a probability of 95% (99%), which approaches the upper limit predicted by the synchrotron emission mechanism. The high degree of polarization suggests that the turbulent magnetic energy is at most 33% of the ordered one. In addition, the X-ray polarization map exhibits a toroidal magnetic field pattern that is consistent with the field revealed by radio observations across the entire nebula. This consistency reveals that the observed X-ray and radio emissions are radiated by electrons from the same magnetic field. Different from the Crab PWN, the consistency observed in the Vela PWN may be attributed to the interaction between the reverse shock of supernova blast wave and the PWN, which leads to a displacement between the synchrotron-cooled nebula and the fresh nebula close to the pulsar. These findings deepen our understanding of the structure and evolution of the Vela PWN, and the magnetohydrodynamic interaction in PWNe.

astro-ph.HE↗