SearcharxivSearch

arXiv subjects

Taeho Kim

Publications and source records attributed to Taeho Kim.

At least 19 recordsLinked to original sources

MonoMoE: An Efficient Fused Mega-kernel for Quantized MoE Decoding

Mixture-of-Experts (MoE) layers increase model capacity without proportionally increasing arithmetic, but their sparse expert computation is difficult to execute efficiently during autoregressive decode. Existing grouped and batched GEMMs are token-major: they construct expert-local token tiles and obtain parallelism from the token dimension. When few tokens reach each expert, this organization incurs tile padding and preprocessing, launches short-lived grids that underutilize memory bandwidth, and exposes quantization, activation, and reduction as separate stages. We present \textbf{MonoMoE}, a weight-major persistent megakernel for block-wise quantized MoE decode. MonoMoE places the complete decode-step token tile on the fine-grained tensor-core $N$ dimension and partitions CTAs over expert-weight tiles, eliminating expert-local token materialization and reducing padded arithmetic. A persistent grid fuses routing, top-$k$ selection, quantization, both expert projections, activation, and reduction in one launch; warp specialization and readiness flags overlap auxiliary work with the dominant expert-weight stream. MonoMoE is integrated with vLLM and supports multiple model shapes through generated kernel specializations and offline schedule tuning. On NVIDIA H200 GPUs, MonoMoE accelerates the complete routed-MoE operator by up to $\mathbf{1.54\times}$ over vLLM Triton Grouped GEMM, is $\mathbf{2.20}$--$\mathbf{3.84\times}$ faster than FlashMoE-FP8 adaptation across various models and batch sizes, and reduces end-to-end time per output token by up to $\mathbf{18.7\%}$ across the evaluated FP8 models, while preserving task accuracy. The MonoMoE implementation and supporting artifacts are open source and available in the \href{https://github.com/flashinfer-ai/flashinfer/tree/main/csrc/fused_moe/monomoe}{FlashInfer repository}.

cs.AR

KernelSight-LM: A Kernel-Level LLM Inference Simulator

As large language models (LLMs) move into production serving, practitioners must rapidly evaluate inference performance across diverse hardware, models, and serving parameters to meet cost and latency targets. However, the end-to-end behavior of LLMs couples serving-layer policies with low-level GPU kernel execution and rapidly evolving architectures, forcing slow, deployment-specific benchmarking that is hard to generalize. We present KernelSight-LM, a fine-grained inference simulator that models token-level execution and produces kernel-level latency breakdowns. It decomposes each serving step into a roofline kernel model with a learned efficiency term, a communication model, and a host-overhead model, composed through a discrete-event scheduler that also captures mechanisms like prefix caching and continuous batching. KernelSight-LM offers two prediction tiers that trade target-GPU data for accuracy. The cross-generation tier uses no target-GPU measurements, only hardware specifications and kernel microbenchmarks from previously profiled GPUs, and predicts per-kernel latency on an unseen GPU generation to 12.1% error, a 1.8x improvement over the roofline baseline (22.0%). A second target-measured tier adds one model-agnostic kernel-microbenchmark sweep on the target GPU, sharpening per-kernel error to 3.8%, a 7.3x improvement over a comparable baseline (27.7%). Both tiers require far less target-GPU data than the prior systems they extend. In our simulator, these predictions yield end-to-end median (p50) errors across six model families of 15.4%, 12.8%, and 3.0% (TTFT, TPOT, throughput) in the cross-generation tier and 14.3%, 6.2%, and 2.7% in the target-measured tier, matching dedicated profiling tools while collecting far less on-device data. Beyond prediction, its kernel-level bottleneck breakdowns support hardware/software co-design and capacity planning.

cs.PF

On the Missing Red Giants near the Galactic Center

There is a long-acknowledged deficiency of bright red giants relative to fainter old stars within a few arc seconds of Sgr A*. We explore whether this could be due to tidal stripping by the central black hole. This requires putting the stars onto highly eccentric orbits, for which we evaluate diffusion by both scalar resonant and non-resonant relaxation of the orbital angular momentum. We conclude that tidal stripping does not discriminate sufficiently between main-sequence and red giant stars. While the tidal loss cone increases with stellar radius, the rate of diffusion into the loss cone increases only logarithmically, whereas the lifetime on the red giant branch decreases more rapidly than $R_*^{-1}$. In agreement with previous studies, we find that stellar collisions are a more likely explanation for the deficiency of bright red giants relative to fainter ones.

astro-ph.GA

AssurAI: Experience with Constructing Korean Socio-cultural Datasets to Discover Potential Risks of Generative AI

The rapid evolution of generative AI necessitates robust safety evaluations. However, current safety datasets are predominantly English-centric, failing to capture specific risks in non-English, socio-cultural contexts such as Korean, and are often limited to the text modality. To address this gap, we introduce AssurAI, a new quality-controlled Korean multimodal dataset for evaluating the safety of generative AI. First, we define a taxonomy of 35 distinct AI risk factors, adapted from established frameworks by a multidisciplinary expert group to cover both universal harms and relevance to the Korean socio-cultural context. Second, leveraging this taxonomy, we construct and release AssurAI, a large-scale Korean multimodal dataset comprising 11,480 instances across text, image, video, and audio. Third, we apply the rigorous quality control process used to ensure data integrity, featuring a two-phase construction (i.e., expert-led seeding and crowdsourced scaling), triple independent annotation, and an iterative expert red-teaming loop. Our pilot study validates AssurAI's effectiveness in assessing the safety of recent LLMs. We release AssurAI to the public to facilitate the development of safer and more reliable generative AI systems for the Korean community.

cs.AI

Versatile and Fast Location-Based Private Information Retrieval with Fully Homomorphic Encryption over the Torus

Location-based services often require users to share sensitive locational data, raising privacy concerns due to potential misuse or exploitation by untrusted servers. In response, we present VeLoPIR, a versatile location-based private information retrieval (PIR) system designed to preserve user privacy while enabling efficient and scalable query processing. VeLoPIR introduces three operational modes-interval validation, coordinate validation, and identifier matching-that support a broad range of real-world applications, including information and emergency alerts. To enhance performance, VeLoPIR incorporates multi-level algorithmic optimizations with parallel structures, achieving significant scalability across both CPU and GPU platforms. We also provide formal security and privacy proofs, confirming the system's robustness under standard cryptographic assumptions. Extensive experiments on real-world datasets demonstrate that VeLoPIR achieves up to 11.55 times speed-up over a prior baseline. The implementation of VeLoPIR is publicly available at https://github.com/PrivStatBool/VeLoPIR.

cs.CR

Data-Driven Sequential Sampling for Tail Risk Mitigation

Given a finite collection of stochastic alternatives, we study the problem of sequentially allocating a fixed sampling budget to identify the optimal alternative with a high probability, where the optimal alternative is defined as the one with the smallest value of extreme tail risk. We particularly consider a situation where these alternatives generate heavy-tailed losses whose probability distributions are unknown and may not admit any specific parametric representation. In this setup, we propose data-driven sequential sampling policies that maximize the rate at which the likelihood of falsely selecting suboptimal alternatives decays to zero. We rigorously demonstrate the superiority of the proposed methods over existing approaches, which is further validated via numerical studies.

stat.ME

Quantum interference and occupation control in high harmonic generation from monolayer $WS_2$

Two-dimensional hexagonal materials such as transition metal dichalcogenides exhibit valley degrees of freedom, offering fascinating potential for valley-based quantum computing and optoelectronics. In nonlinear optics, the K and K' valleys provide excitation resonances that can be used for ultrafast control of excitons, Bloch oscillations, and Floquet physics. Under intense laser fields, however, the role of coherent carrier dynamics away from the K/K' valleys is largely unexplored. In this study, we observe quantum interferences in high harmonic generation from monolayer $WS_2$ as laser fields drive electrons from the valleys across the full Brillouin zone. In the perturbative regime, interband resonances at the valleys enhance high harmonic generation through multi-photon excitations. In the strong-field regime, the high harmonic spectrum is sensitively controlled by light-driven quantum interferences between the interband valley resonances and intraband currents originating from electrons occupying various points in the Brillouin zone, also away from K/K' valleys such as $\Gamma$ and M. Our experimental observations are in strong agreement with quantum simulations, validating their interpretation. This work proposes new routes for harnessing laser-driven quantum interference in two-dimensional hexagonal systems and all-optical techniques to occupy and read-out electronic structures in the full Brillouin zone via strong-field nonlinear optics, advancing quantum technologies.

physics.optics

Optimizing Input Data Collection for Ranking and Selection

We study a ranking and selection (R&S) problem when all solutions share common parametric Bayesian input models updated with the data collected from multiple independent data-generating sources. Our objective is to identify the best system by designing a sequential sampling algorithm that collects input and simulation data given a budget. We adopt the most probable best (MPB) as the estimator of the optimum and show that its posterior probability of optimality converges to one at an exponential rate as the sampling budget increases. Assuming that the input parameters belong to a finite set, we characterize the $\epsilon$-optimal static sampling ratios for input and simulation data that maximize the convergence rate. Using these ratios as guidance, we propose the optimal sampling algorithm for R&S (OSAR) that achieves the $\epsilon$-optimal ratios almost surely in the limit. We further extend OSAR by adopting the kernel ridge regression to improve the simulation output mean prediction. This not only improves OSAR's finite-sample performance, but also lets us tackle the case where the input parameters lie in a continuous space with a strong consistency guarantee for finding the optimum. We numerically demonstrate that OSAR outperforms a state-of-the-art competitor.

stat.ME

LLMem: Estimating GPU Memory Usage for Fine-Tuning Pre-Trained LLMs

Fine-tuning pre-trained large language models (LLMs) with limited hardware presents challenges due to GPU memory constraints. Various distributed fine-tuning methods have been proposed to alleviate memory constraints on GPU. However, determining the most effective method for achieving rapid fine-tuning while preventing GPU out-of-memory issues in a given environment remains unclear. To address this challenge, we introduce LLMem, a solution that estimates the GPU memory consumption when applying distributed fine-tuning methods across multiple GPUs and identifies the optimal method. We conduct GPU memory usage estimation prior to fine-tuning, leveraging the fundamental structure of transformer-based decoder models and the memory usage distribution of each method. Experimental results show that LLMem accurately estimates peak GPU memory usage on a single GPU, with error rates of up to 1.6%. Additionally, it shows an average error rate of 3.0% when applying distributed fine-tuning methods to LLMs with more than a billion parameters on multi-GPU setups.

cs.AI

Interplay of valley, layer and band topology towards interacting quantum phases in moir\'e bilayer graphene

In Bernal-stacked bilayer graphene (BBG), the Landau levels give rise to an intimate connection between valley and layer degrees of freedom. Adding a moir\'e superlattice potential enriches the BBG physics with the formation of topological minibands - potentially leading to tunable exotic quantum transport. Here, we present magnetotransport measurements of a high-quality bilayer graphene-hexagonal boron nitride (hBN) heterostructure. The zero-degree alignment generates a strong moir\'e superlattice potential for the electrons in BBG and the resulting Landau fan diagram of longitudinal and Hall resistance displays a Hofstadter butterfly pattern with a high level of detail. We demonstrate that the intricate relationship between valley and layer degrees of freedom controls the topology of moir\'e-induced bands, significantly influencing the energetics of interacting quantum phases in the BBG superlattice. We further observe signatures of field-induced correlated insulators, helical edge states and clear quantizations of interaction-driven topological quantum phases, such as symmetry broken Chern insulators.

cond-mat.mes-hall

Self-Supervised Learning from Non-Object Centric Images with a Geometric Transformation Sensitive Architecture

Most invariance-based self-supervised methods rely on single object-centric images (e.g., ImageNet images) for pretraining, learning features that invariant to geometric transformation. However, when images are not object-centric, the semantics of the image can be significantly altered due to cropping. Furthermore, as the model becomes insensitive to geometric transformations, it may struggle to capture location information. For this reason, we propose a Geometric Transformation Sensitive Architecture designed to be sensitive to geometric transformations, specifically focusing on four-fold rotation, random crop, and multi-crop. Our method encourages the student to be sensitive by predicting rotation and using targets that vary with those transformations through pooling and rotating the teacher feature map. Additionally, we use patch correspondence loss to encourage correspondence between patches with similar features. This approach allows us to capture long-term dependencies in a more appropriate way than capturing long-term dependencies by encouraging local-to-global correspondence, which occurs when learning to be insensitive to multi-crop. Our approach demonstrates improved performance when using non-object-centric images as pretraining data compared to other methods that train the model to be insensitive to geometric transformation. We surpass DINO[Caron et al.[2021b]] baseline in tasks including image classification, semantic segmentation, detection, and instance segmentation with improvements of 4.9 $Top-1 Acc$, 3.3 $mIoU$, 3.4 $AP^b$, and 2.7 $AP^m$. Code and pretrained models are publicly available at: https://github.com/bok3948/GTSA

cs.CV

Maximum Agreement Linear Predictors

This paper studies predictor functions motivated by maximizing a measure of agreement with the predictand. Specifically, it examines distributional properties and predictive performance of the estimated maximum agreement linear predictor (MALP), the linear predictor maximizing Lin's concordance correlation coefficient (CCC) between the predictor and the predictand. It is compared and contrasted, theoretically and through computer experiments, with the estimated least-squares linear predictor (LSLP), with respect to some performance measures. Finite-sample and asymptotic properties are obtained, and confidence intervals and prediction intervals are also presented. Predictors are illustrated using two real data sets: an eye data set and a body fat data set. Results indicate that the estimated MALP is a viable alternative to the estimated LSLP if one desires a predictor whose predicted values possesses higher agreement with the predictand values, as measured by the CCC.

stat.ME

Tensor Slicing and Optimization for Multicore NPUs

Although code generation for Convolution Neural Network (CNN) models has been extensively studied, performing efficient data slicing and parallelization for highly-constrai\-ned Multicore Neural Processor Units (NPUs) is still a challenging problem. Given the size of convolutions' input/output tensors and the small footprint of NPU on-chip memories, minimizing memory transactions while maximizing parallelism and MAC utilization are central to any effective solution. This paper proposes a TensorFlow XLA/LLVM compiler optimization pass for Multicore NPUs, called Tensor Slicing Optimization (TSO), which: (a) maximizes convolution parallelism and memory usage across NPU cores; and (b) reduces data transfers between host and NPU on-chip memories by using DRAM memory burst time estimates to guide tensor slicing. To evaluate the proposed approach, a set of experiments was performed using the NeuroMorphic Processor (NMP), a multicore NPU containing 32 RISC-V cores extended with novel CNN instructions. Experimental results show that TSO is capable of identifying the best tensor slicing that minimizes execution time for a set of CNN models. Speed-ups of up to 21.7\% result when comparing the TSO burst-based technique to a no-burst data slicing approach. To validate the generality of the TSO approach, the algorithm was also ported to the Glow Machine Learning framework. The performance of the models were measured on both Glow and TensorFlow XLA/LLVM compilers, revealing similar results.

cs.PF

Effect of Annealing Temperature on Minimum Domain Size of Ferroelectric Hafnia

Here, we optimized the annealing temperature of HZO/TiN thin film heterostructure via multiscale analysis of remnant polarization, crystallographic phase, minimum ferroelectric domain size, and average grain size. We found that the remnant polarization was closely related to the relative amount of the orthorhombic phase whereas the minimum domain size was to the relative amount of the monoclinic phase. The minimum domain size was obtained at the annealing temperature of 500$^\cird$C while the optimum remnant polarization and capacitance at the annealing temperature of 600$^\circ$C. We conclude that the minimum domain size is more important than the sheer magnitude of remnant polarization considering the retention and fatigue of switchable polarization in nanoscale ferroelectric devices. Our results are expected to contribute to the development of ultra-low-power logic transistors and next-generation non-volatile memory devices.

physics.app-ph

Gate-tunable quantum pathways of high harmonic generation in graphene

Under strong laser fields, electrons in solids radiate high-harmonic fields by travelling through quantum pathways in Bloch bands in the sub-laser-cycle timescales. Understanding these pathways in the momentum space through the high-harmonic radiation can enable an all-optical ultrafast probe to observe coherent lightwave-driven processes and measure electronic structures as recently demonstrated for semiconductors. However, such demonstration has been largely limited for semimetals because the absence of the bandgap hinders an experimental characterization of the exact pathways. In this study, by combining electrostatic control of chemical potentials with HHG measurement, we resolve quantum pathways of massless Dirac fermions in graphene under strong laser fields. Electrical modulation of HHG reveals quantum interference between the multi-photon interband excitation channels. As the light-matter interaction deviates beyond the perturbative regime, elliptically polarized laser fields efficiently drive massless Dirac fermions via an intricate coupling between the interband and intraband transitions, which is corroborated by our theoretical calculations. Our findings pave the way for strong-laser-field tomography of Dirac electrons in various quantum semimetals and their ultrafast electronics with a gate control.

cond-mat.mes-hall

Selection of the Most Probable Best

We consider an expected-value ranking and selection (R&S) problem where all k solutions' simulation outputs depend on a common parameter whose uncertainty can be modeled by a distribution. We define the most probable best (MPB) to be the solution that has the largest probability of being optimal with respect to the distribution and design an efficient sequential sampling algorithm to learn the MPB when the parameter has a finite support. We derive the large deviations rate of the probability of falsely selecting the MPB and formulate an optimal computing budget allocation problem to find the rate-maximizing static sampling ratios. The problem is then relaxed to obtain a set of optimality conditions that are interpretable and computationally efficient to verify. We devise a series of algorithms that replace the unknown means in the optimality conditions with their estimates and prove the algorithms' sampling ratios achieve the conditions as the simulation budget increases. Furthermore, we show that the empirical performances of the algorithms can be significantly improved by adopting the kernel ridge regression for mean estimation while achieving the same asymptotic convergence results. The algorithms are benchmarked against a state-of-the-art contextual R&S algorithm and demonstrated to have superior empirical performances.

stat.ME

CPrune: Compiler-Informed Model Pruning for Efficient Target-Aware DNN Execution

Mobile devices run deep learning models for various purposes, such as image classification and speech recognition. Due to the resource constraints of mobile devices, researchers have focused on either making a lightweight deep neural network (DNN) model using model pruning or generating an efficient code using compiler optimization. Surprisingly, we found that the straightforward integration between model compression and compiler auto-tuning often does not produce the most efficient model for a target device. We propose CPrune, a compiler-informed model pruning for efficient target-aware DNN execution to support an application with a required target accuracy. CPrune makes a lightweight DNN model through informed pruning based on the structural information of subgraphs built during the compiler tuning process. Our experimental results show that CPrune increases the DNN execution speed up to 2.73x compared to the state-of-the-art TVM auto-tune while satisfying the accuracy requirement.

cs.LG

Quantune: Post-training Quantization of Convolutional Neural Networks using Extreme Gradient Boosting for Fast Deployment

To adopt convolutional neural networks (CNN) for a range of resource-constrained targets, it is necessary to compress the CNN models by performing quantization, whereby precision representation is converted to a lower bit representation. To overcome problems such as sensitivity of the training dataset, high computational requirements, and large time consumption, post-training quantization methods that do not require retraining have been proposed. In addition, to compensate for the accuracy drop without retraining, previous studies on post-training quantization have proposed several complementary methods: calibration, schemes, clipping, granularity, and mixed-precision. To generate a quantized model with minimal error, it is necessary to study all possible combinations of the methods because each of them is complementary and the CNN models have different characteristics. However, an exhaustive or a heuristic search is either too time-consuming or suboptimal. To overcome this challenge, we propose an auto-tuner known as Quantune, which builds a gradient tree boosting model to accelerate the search for the configurations of quantization and reduce the quantization error. We evaluate and compare Quantune with the random, grid, and genetic algorithms. The experimental results show that Quantune reduces the search time for quantization by approximately 36.5x with an accuracy loss of 0.07 ~ 0.65% across six CNN models, including the fragile ones (MobileNet, SqueezeNet, and ShuffleNet). To support multiple targets and adopt continuously evolving quantization works, Quantune is implemented on a full-fledged compiler for deep learning as an open-sourced project.

cs.LG