SearcharxivSearch

arXiv subjects

Youngjun Kim

Publications and source records attributed to Youngjun Kim.

3 recordsLinked to original sources

LaDiMo: Layer-wise Distillation Inspired MoEfier

The advent of large language models has revolutionized natural language processing, but their increasing complexity has led to substantial training costs, resource demands, and environmental impacts. In response, sparse Mixture-of-Experts (MoE) models have emerged as a promising alternative to dense models. Since training MoE models from scratch can be prohibitively expensive, recent studies have explored leveraging knowledge from pre-trained non-MoE models. However, existing approaches have limitations, such as requiring significant hardware resources and data. We propose a novel algorithm, LaDiMo, which efficiently converts a Transformer-based non-MoE model into a MoE model with minimal additional training cost. LaDiMo consists of two stages: layer-wise expert construction and routing policy decision. By harnessing the concept of Knowledge Distillation, we compress the model and rapidly recover its performance. Furthermore, we develop an adaptive router that optimizes inference efficiency by profiling the distribution of routing weights and determining a layer-wise policy that balances accuracy and latency. We demonstrate the effectiveness of our method by converting the LLaMA2-7B model to a MoE model using only 100K tokens, reducing activated parameters by over 20% while keeping accuracy. Our approach offers a flexible and efficient solution for building and deploying MoE models.

cs.CL

HyperCLOVA X Technical Report

We introduce HyperCLOVA X, a family of large language models (LLMs) tailored to the Korean language and culture, along with competitive capabilities in English, math, and coding. HyperCLOVA X was trained on a balanced mix of Korean, English, and code data, followed by instruction-tuning with high-quality human-annotated datasets while abiding by strict safety guidelines reflecting our commitment to responsible AI. The model is evaluated across various benchmarks, including comprehensive reasoning, knowledge, commonsense, factuality, coding, math, chatting, instruction-following, and harmlessness, in both Korean and English. HyperCLOVA X exhibits strong reasoning capabilities in Korean backed by a deep understanding of the language and cultural nuances. Further analysis of the inherent bilingual nature and its extension to multilingualism highlights the model's cross-lingual proficiency and strong generalization ability to untargeted languages, including machine translation between several language pairs and cross-lingual inference tasks. We believe that HyperCLOVA X can provide helpful guidance for regions or countries in developing their sovereign LLMs.

cs.CL

Multiflagellarity leads to the size-independent swimming speed of peritrichous bacteria

To swim through a viscous fluid, a flagellated bacterium must overcome the fluid drag on its body by rotating a flagellum or a bundle of multiple flagella. Because the drag increases with the size of bacteria, it is expected theoretically that the swimming speed of a bacterium inversely correlates with its body length. Nevertheless, despite extensive research, the fundamental size-speed relation of flagellated bacteria remains unclear with different experiments reporting conflicting results. Here, by critically reviewing the existing evidence and synergizing our own experiments of large sample sizes, hydrodynamic modeling and simulations, we demonstrate that the average swimming speed of \textit{Escherichia coli}, a premier model of peritrichous bacteria, is independent of their body length. Our quantitative analysis shows that such a counterintuitive relation is the consequence of the collective flagellar dynamics dictated by the linear correlation between the body length and the number of flagella of bacteria. Notably, our study reveals how bacteria utilize the increasing number of flagella to regulate the flagellar motor load. The collective load sharing among multiple flagella results in a lower load on each flagellar motor and therefore faster flagellar rotation, which compensates for the higher fluid drag on the longer bodies of bacteria. Without this balancing mechanism, the swimming speed of monotrichous bacteria generically decreases with increasing body length, a feature limiting the size variation of the bacteria. Altogether, our study resolves a long-standing controversy over the size-speed relation of flagellated bacteria and provides new insights into the functional benefit of multiflagellarity in bacteria.

physics.flu-dyn