SearcharxivSearch

arXiv subjects

Yanping Huang

Publications and source records attributed to Yanping Huang.

At least 19 recordsLinked to original sources

Measurement of $\Xi^-/\bar{\Xi}^{+}$ production in jets from $Z$ boson decays with the DELPHI open data

The production rates of $\Xi^{-}/\bar{\Xi}^{+}$ baryons in energy-ranked jets produced in $Z\to\text{hadrons}$ decays are measured using $3.2$ million hadronic $Z$ events recorded by the DELPHI experiment. Jets are reconstructed using the Durham algorithm with $y_{\text{cut}}=0.005$. Quark- and gluon-enriched jet samples are obtained by ranking the jet energies in three-jet events. The softest jet are found to produce fewer $\Xi^{-}/\bar{\Xi}^{+}$ and less energetic baryons than the other jets. The ratio of $\Xi^{-}/\bar{\Xi}^{+}$ production rates in gluon and quark jets, each normalized to the corresponding mean charged-particle multiplicity, is measured to be $1.21 \pm 0.18~\mathrm{(stat.)} \pm 0.26~\mathrm{(syst.)}$. The result is consistent with the JETSET expectation and the OPAL measurements of $K_S^0$ and $\Lambda$ productions in $Z$ decays. This study presents the first measurement of the gluon-to-quark production ratio for baryons containing two $s$ quarks, providing new insights into strange-quark production and hadronization. Future $e^{+}e^{-}$ colliders such as CEPC and FCCee will provide much larger $Z$-boson samples and will allow far more precise studies of the subject.

hep-ex

New Physics Search at the CEPC: a General Perspective

The Circular Electron-Positron Collider (CEPC), a proposed next-generation Higgs factory, provides new opportunities to explore physics beyond the Standard Model (SM). With its clean electron-positron collision environment and the ability to collect large samples of Higgs, W, and Z bosons, the CEPC enables precision measurements and searches for new physics. This white paper outlines the CEPC's discovery potential, including studies of exotic decays of the Higgs, Z, and top quarks, dark matter and dark sector phenomena, long-lived particles, supersymmetry, and neutrino-related signatures. Advanced detector technologies and reconstruction techniques, such as one-to-one correspondence reconstruction and jet origin identification, significantly improve sensitivity to rare and weakly interacting processes. The CEPC is particularly well suited to probe the electroweak phase transition and test models of electroweak baryogenesis and dark sector interactions. In addition, global fit analyses highlight the CEPC's complementary role in constraining a wide range of new physics scenarios. These features position the CEPC as a powerful tool for exploring the next frontier in fundamental particle physics in the post-Higgs discovery era.

hep-ex

Discovery of a Glueball-like particle X(2370) at BESIII

Radiative decays of the $J/\psi$ particle are of gluon-rich environment, providing an ideal place for hunting glueballs. The $X(2370)$ particle was first discovered in $J/\psi\to \gamma \pi^+\pi^-\eta^{\prime}$ process in 2011 with the BESIII experiment at BEPCII Collider, and later it was confirmed in $J/\psi\to\gamma K\bar{K}\eta^{\prime}$ decays. In 2024, with a sample of 10 billion $J/\psi$ events collected at the BESIII detector, the spin-parity of the $X(2370)$ was determined to be $0^{-+}$ for the first time in the partial wave analysis of $J/\psi\to \gamma K^{0}_{S}K^{0}_{S} \eta^{\prime}$ process. Recently, new decay modes of $X(2370)\to K^{0}_{S}K^{0}_{S}\pi^0$, $\pi^0\pi^0\eta$ and $a^0\pi^0$ were observed. The mass, spin-parity quantum numbers, production and decay properties of the $X(2370)$ particle are consistent with the features of the lightest pseudoscalar glueball.

hep-ex

Autellix: An Efficient Serving Engine for LLM Agents as General Programs

Large language model (LLM) applications are evolving beyond simple chatbots into dynamic, general-purpose agentic programs, which scale LLM calls and output tokens to help AI agents reason, explore, and solve complex tasks. However, existing LLM serving systems ignore dependencies between programs and calls, missing significant opportunities for optimization. Our analysis reveals that programs submitted to LLM serving engines experience long cumulative wait times, primarily due to head-of-line blocking at both the individual LLM request and the program. To address this, we introduce Autellix, an LLM serving system that treats programs as first-class citizens to minimize their end-to-end latencies. Autellix intercepts LLM calls submitted by programs, enriching schedulers with program-level context. We propose two scheduling algorithms-for single-threaded and distributed programs-that preempt and prioritize LLM calls based on their programs' previously completed calls. Our evaluation demonstrates that across diverse LLMs and agentic workloads, Autellix improves throughput of programs by 4-15x at the same latency compared to state-of-the-art systems, such as vLLM.

cs.LG

Flavor Physics at the CEPC: a General Perspective

We discuss the landscape of flavor physics at the Circular Electron-Positron Collider (CEPC), based on the nominal luminosity outlined in its Technical Design Report. The CEPC is designed to operate in multiple modes to address a variety of tasks. At the $Z$ pole, the expected production of 4 Tera $Z$ bosons will provide unique and highly precise measurements of $Z$ boson couplings, while the substantial number of boosted heavy-flavored quarks and leptons produced in clean $Z$ decays will facilitate investigations into their flavor physics with unprecedented precision. We investigate the prospects of measuring various physics benchmarks and discuss their implications for particle theories and phenomenological models. Our studies indicate that, with its highlighted advantages and anticipated excellent detector performance, the CEPC can explore beauty and $\tau$ physics in ways that are superior to or complementary with the Belle II and Large-Hadron-Collider-beauty experiments, potentially enabling the detection of new physics at energy scales of 10 TeV and above. This potential also extends to the observation of yet-to-be-discovered rare and exotic processes, as well as testing fundamental principles such as lepton flavor universality, lepton and baryon number conservation, etc., making the CEPC a vibrant platform for flavor physics research. The $WW$ threshold scan, Higgs-factory operation and top-pair productions of the CEPC further enhance its merits in this regard, especially for measuring the Cabibbo-Kobayashi-Maskawa matrix elements, and Flavor-Changing-Neutral-Current physics of Higgs boson and top quarks. We outline the requirements for detector performance and considerations for future development to achieve the anticipated scientific goals.

hep-ex

Measurements of decay branching fractions of the Higgs boson to hadronic final states at the CEPC

The Circular Electron Positron Collider (CEPC) is a large-scale particle accelerator designed to collide electrons and positrons at high energies. One of the primary goals of the CEPC is to achieve high-precision measurements of the properties of the Higgs boson, facilitated by the large number of Higgs bosons that can be produced with significantly low contamination. The measurements of Higgs boson branching fractions into $b\overline{b} /c\overline{c} /gg$ and $\tau\overline{\tau} /WW^{*} /ZZ^{*} $, where the $W$ or $Z$ bosons decay hadronically, are presented in the context of the CEPC experiment, assuming a scenario with 5600 fb$^{-1}$ of collision data at a center-of-mass energy of 240 GeV. In this study the Higgs bosons are produced in association with a $Z$ boson, with the $Z$ boson decaying into a pair of muons $(\mu^{+}\mu^{-})$, which have high efficiency and high resolution. In order to separate all decay channels simultaneously with high accuracy, the Particle Flow Network (PFN), a graph-based machine learning model, is considered. The precise classification provided by the PFN is employed in measuring the branching fractions using the migration matrix method, which accurately corrects for detector effects in each decay channel. The statistical uncertainty of the measured branching ratio is estimated to be 0.55% in $H\to b\overline{b}$ final state, and approximately 1.5%-16% in $H\to c\overline{c} /gg/\tau\overline{\tau}/WW^{*} /ZZ^{*} $ final states. In addition, the main sources of systematic uncertainties to the measurement of the branching fractions are discussed.

hep-ex

Determination of Strong Coupling Constant from Inclusive Semileptonic Decays of Charmed Mesons

Employing the heavy quark expansion model with the kinetic scheme, we evaluate $\alpha_S(m_c^2)$, the strong coupling constant at the charm quark mass $m_c$ with data on inclusive semileptonic decays of charmed mesons. Using the experimental values of semileptonic decay widths of the $D^0$ and the $D^+$, the value of $\alpha_{s}(m_c^{2})$ is determined to be $0.445\pm0.009\pm0.114$, where the first uncertainty is experimental and the second systematic. This reported $\alpha_{s}(m_c^{2})$ is in good agreement with the value of $\alpha_{s}(m_c^{2})$ calculated by running $\alpha_S(m_Z^2)$ at the $Z^0$ boson mass $m_Z$ with the renormalization group evolution equation. In addition, values of $\alpha_{s}(m_c^{2})$ obtained individually from each of the $D^0$, $D^+$, and $D_s^+$ mesons are found to be consistent being of the same origin.

hep-ph

Stylus: Automatic Adapter Selection for Diffusion Models

Beyond scaling base models with more data or parameters, fine-tuned adapters provide an alternative way to generate high fidelity, custom images at reduced costs. As such, adapters have been widely adopted by open-source communities, accumulating a database of over 100K adapters-most of which are highly customized with insufficient descriptions. This paper explores the problem of matching the prompt to a set of relevant adapters, built on recent work that highlight the performance gains of composing adapters. We introduce Stylus, which efficiently selects and automatically composes task-specific adapters based on a prompt's keywords. Stylus outlines a three-stage approach that first summarizes adapters with improved descriptions and embeddings, retrieves relevant adapters, and then further assembles adapters based on prompts' keywords by checking how well they fit the prompt. To evaluate Stylus, we developed StylusDocs, a curated dataset featuring 75K adapters with pre-computed adapter embeddings. In our evaluation on popular Stable Diffusion checkpoints, Stylus achieves greater CLIP-FID Pareto efficiency and is twice as preferred, with humans and multimodal models as evaluators, over the base model. See stylus-diffusion.github.io for more.

cs.CV

Double Dome and Reemergence of Superconductivity in Pristine 6R-TaS2 under Pressure

Investigating the implications of interlayer coupling on superconductivity is essential for comprehending the intrinsic mechanisms of high temperature superconductors. Van der Waals heterojunctions have attracted extensive research due to their exotic interlayer coupling. Here, we present a natural heterojunction superconductor of 6R-TaS2 that demonstrates a double-dome of superconductivity, in addition to, the reemergence of superconducting under high pressures. Our first principles calculation shows that the first dome of superconductivity in 6R-TaS2 can be attributed to changes in interlayer coupling and charge transfer. The second superconducting dome and the reemergence of superconductivity can be ascribed to changes in the density of states resulting from Fermi surface reconstruction, in which the DOS of T-layer and S p-orbitals play a crucial role. We have reported the first observation in TMDs that non-metallic atoms playing a dominant role in the reemergence of superconducting and the influence of two Lifshitz transitions on superconducting properties.

cond-mat.supr-con

Controlled Decoding from Language Models

KL-regularized reinforcement learning (RL) is a popular alignment framework to control the language model responses towards high reward outcomes. We pose a tokenwise RL objective and propose a modular solver for it, called controlled decoding (CD). CD exerts control through a separate prefix scorer module, which is trained to learn a value function for the reward. The prefix scorer is used at inference time to control the generation from a frozen base model, provably sampling from a solution to the RL objective. We empirically demonstrate that CD is effective as a control mechanism on popular benchmarks. We also show that prefix scorers for multiple rewards may be combined at inference time, effectively solving a multi-objective RL problem with no additional training. We show that the benefits of applying CD transfer to an unseen base model with no further tuning as well. Finally, we show that CD can be applied in a blockwise decoding fashion at inference-time, essentially bridging the gap between the popular best-of-K strategy and tokenwise control through reinforcement learning. This makes CD a promising approach for alignment of language models.

cs.LG

SPAE: Semantic Pyramid AutoEncoder for Multimodal Generation with Frozen LLMs

In this work, we introduce Semantic Pyramid AutoEncoder (SPAE) for enabling frozen LLMs to perform both understanding and generation tasks involving non-linguistic modalities such as images or videos. SPAE converts between raw pixels and interpretable lexical tokens (or words) extracted from the LLM's vocabulary. The resulting tokens capture both the semantic meaning and the fine-grained details needed for visual reconstruction, effectively translating the visual content into a language comprehensible to the LLM, and empowering it to perform a wide array of multimodal tasks. Our approach is validated through in-context learning experiments with frozen PaLM 2 and GPT 3.5 on a diverse set of image understanding and generation tasks. Our method marks the first successful attempt to enable a frozen LLM to generate image content while surpassing state-of-the-art performance in image understanding tasks, under the same setting, by over 25%.

cs.CV

Brainformers: Trading Simplicity for Efficiency

Transformers are central to recent successes in natural language processing and computer vision. Transformers have a mostly uniform backbone where layers alternate between feed-forward and self-attention in order to build a deep network. Here we investigate this design choice and find that more complex blocks that have different permutations of layer primitives can be more efficient. Using this insight, we develop a complex block, named Brainformer, that consists of a diverse sets of layers such as sparsely gated feed-forward layers, dense feed-forward layers, attention layers, and various forms of layer normalization and activation functions. Brainformer consistently outperforms the state-of-the-art dense and sparse Transformers, in terms of both quality and efficiency. A Brainformer model with 8 billion activated parameters per token demonstrates 2x faster training convergence and 5x faster step time compared to its GLaM counterpart. In downstream task evaluation, Brainformer also demonstrates a 3% higher SuperGLUE score with fine-tuning compared to GLaM with a similar number of activated parameters. Finally, Brainformer largely outperforms a Primer dense model derived with NAS with similar computation per token on fewshot evaluations.

cs.LG

Lifelong Language Pretraining with Distribution-Specialized Experts

Pretraining on a large-scale corpus has become a standard method to build general language models (LMs). Adapting a model to new data distributions targeting different downstream tasks poses significant challenges. Naive fine-tuning may incur catastrophic forgetting when the over-parameterized LMs overfit the new data but fail to preserve the pretrained features. Lifelong learning (LLL) aims to enable information systems to learn from a continuous data stream across time. However, most prior work modifies the training recipe assuming a static fixed network architecture. We find that additional model capacity and proper regularization are key elements to achieving strong LLL performance. Thus, we propose Lifelong-MoE, an extensible MoE (Mixture-of-Experts) architecture that dynamically adds model capacity via adding experts with regularized pretraining. Our results show that by only introducing a limited number of extra experts while keeping the computation cost constant, our model can steadily adapt to data distribution shifts while preserving the previous knowledge. Compared to existing lifelong learning approaches, Lifelong-MoE achieves better few-shot performance on 19 downstream NLP tasks.

cs.CL

PaLM 2 Technical Report

We introduce PaLM 2, a new state-of-the-art language model that has better multilingual and reasoning capabilities and is more compute-efficient than its predecessor PaLM. PaLM 2 is a Transformer-based model trained using a mixture of objectives. Through extensive evaluations on English and multilingual language, and reasoning tasks, we demonstrate that PaLM 2 has significantly improved quality on downstream tasks across different model sizes, while simultaneously exhibiting faster and more efficient inference compared to PaLM. This improved efficiency enables broader deployment while also allowing the model to respond faster, for a more natural pace of interaction. PaLM 2 demonstrates robust reasoning capabilities exemplified by large improvements over PaLM on BIG-Bench and other reasoning tasks. PaLM 2 exhibits stable performance on a suite of responsible AI evaluations, and enables inference-time control over toxicity without additional overhead or impact on other capabilities. Overall, PaLM 2 achieves state-of-the-art performance across a diverse set of tasks and capabilities. When discussing the PaLM 2 family, it is important to distinguish between pre-trained models (of various sizes), fine-tuned variants of these models, and the user-facing products that use these models. In particular, user-facing products typically include additional pre- and post-processing steps. Additionally, the underlying models may evolve over time. Therefore, one should not expect the performance of user-facing products to exactly match the results reported in this report.

cs.CL

AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving

Model parallelism is conventionally viewed as a method to scale a single large deep learning model beyond the memory limits of a single device. In this paper, we demonstrate that model parallelism can be additionally used for the statistical multiplexing of multiple devices when serving multiple models, even when a single model can fit into a single device. Our work reveals a fundamental trade-off between the overhead introduced by model parallelism and the opportunity to exploit statistical multiplexing to reduce serving latency in the presence of bursty workloads. We explore the new trade-off space and present a novel serving system, AlpaServe, that determines an efficient strategy for placing and parallelizing collections of large deep learning models across a distributed cluster. Evaluation results on production workloads show that AlpaServe can process requests at up to 10x higher rates or 6x more burstiness while staying within latency constraints for more than 99% of requests.

cs.LG

Massively Multilingual Shallow Fusion with Large Language Models

While large language models (LLM) have made impressive progress in natural language processing, it remains unclear how to utilize them in improving automatic speech recognition (ASR). In this work, we propose to train a single multilingual language model (LM) for shallow fusion in multiple languages. We push the limits of the multilingual LM to cover up to 84 languages by scaling up using a mixture-of-experts LLM, i.e., generalist language model (GLaM). When the number of experts increases, GLaM dynamically selects only two at each decoding step to keep the inference computation roughly constant. We then apply GLaM to a multilingual shallow fusion task based on a state-of-the-art end-to-end model. Compared to a dense LM of similar computation during inference, GLaM reduces the WER of an English long-tail test set by 4.4% relative. In a multilingual shallow fusion task, GLaM improves 41 out of 50 languages with an average relative WER reduction of 3.85%, and a maximum reduction of 10%. Compared to the baseline model, GLaM achieves an average WER reduction of 5.53% over 43 languages.

cs.CL

Expected $H\to\mu^{+}\mu^{-}$ measurement precision with $e^{+}e^{-}\to Z(q\bar{q})H$ production at the CEPC

A search for the dimuon decay of the Standard Model Higgs boson is performed using the Monte Carlo simulated events to mimic data corresponding to an integrated luminosity of 5.6 ab$^{-1}$ collected with the Circular Electron-Positron Collider detector in $e^{+}e^{-}$ collisions at $\sqrt{s}=240$ GeV. The paper studies $e^{+}e^{-}\to ZH,\,Z\to q\bar{q},\,H\to\mu^{+}\mu^{-}$ process, and the expected significance considering only the data statistical uncertainty over the background-only hypothesis for a Higgs boson with a mass of 125 GeV is found to be 6.1$\sigma$, corresponding to the precision of 19%. The systematic impacts from the background Monte Carlo statistical fluctuations are estimated to be negligible. The dependence of the measurement accuracy on the muon momentum resolution of the CEPC detector has been investigated. It is found that the muon momentum resolution has to be better than 204 MeV to discover the $H\to\mu\mu$ process at the nominal integrated luminosity. And if the resolution is 100% worse than the designed parameter, the integrated luminosity is needed to be greater than 7.2 ab$^{-1}$ to reach 5$\sigma$ significance.

hep-ex

Scaling Instruction-Finetuned Language Models

Finetuning language models on a collection of datasets phrased as instructions has been shown to improve model performance and generalization to unseen tasks. In this paper we explore instruction finetuning with a particular focus on (1) scaling the number of tasks, (2) scaling the model size, and (3) finetuning on chain-of-thought data. We find that instruction finetuning with the above aspects dramatically improves performance on a variety of model classes (PaLM, T5, U-PaLM), prompting setups (zero-shot, few-shot, CoT), and evaluation benchmarks (MMLU, BBH, TyDiQA, MGSM, open-ended generation). For instance, Flan-PaLM 540B instruction-finetuned on 1.8K tasks outperforms PALM 540B by a large margin (+9.4% on average). Flan-PaLM 540B achieves state-of-the-art performance on several benchmarks, such as 75.2% on five-shot MMLU. We also publicly release Flan-T5 checkpoints, which achieve strong few-shot performance even compared to much larger models, such as PaLM 62B. Overall, instruction finetuning is a general method for improving the performance and usability of pretrained language models.

cs.LG