SearcharxivSearch

arXiv subjects

Youngjae Kim

Publications and source records attributed to Youngjae Kim.

At least 19 recordsLinked to original sources

DiaLSM: Towards Write-Stall-Free Performance via Shard-based LSM-tree

Log-Structured Merge-tree (LSM) aims to achieve high write throughput, but is known to experience the write stall problems when subjected to sustained write pressure. We quantify the occurrence probability and average duration of write stalls in LSM using a queuing model in the write--flush--compaction pipeline, moving beyond existing empirical analysis. The proposed model demonstrates that a monolithic LSM with a single pipeline cannot eliminate write stalls, revealing that internal sharding within the LSM offers an opportunity for fundamental write stall mitigation. To break this structural bottleneck, we propose DiaLSM, an internally shard-based LSM architecture. Instead of forcing all writes through one pipeline, DiaLSM splits the write--flush--compaction path into multiple independent shards and employs dynamic fallback redirection, allowing writes to proceed even when some shards stall. Implemented on RocksDB, DiaLSM achieves up to 2.4x higher throughput, 94% lower stalls, and significantly lower latency than state-of-the-art methods ADOC and Sub-Compaction, as demonstrated by db_bench, YCSB, and Sysbench OLTP evaluations.

cs.DB

BAFF: Bid-Aware Filter Family for Mitigating Training Data Interference in RTB A/B Tests

In online A/B tests for real-time bidding (RTB), control and treatment models are typically trained on a shared serving log that includes data generated by the counterpart model. This shared-log training biases each model's training data through two channels: the counterpart model may have selected a different ad from the ad-candidate pool (ad-ranking disagreement) and may have bid a different price (bid-pricing disagreement), potentially distorting the A/B test outcome. Log-splitting eliminates the bias but sacrifices training data; log-sharing retains all data but leaves the bias unaddressed. We formalize the Bid-Aware Filter Family (BAFF), a class of (k,l)-parameterized hard filters that controls tolerance to each channel independently, providing a structured search space between these two extremes. We further propose a three-stage online measurement protocol that enables evaluating data-sharing strategies by their deviation from an interference-free reference model in production. In offline simulation, a (k,l) sweep surfaces operating points with smaller deviation from the interference-free reference model than both log-sharing and log-splitting. In a live RTB deployment on a demand-side platform (DSP), filter-based variants preserve the reference model's business metrics (e.g., CPC, CTR) more closely than both baselines. The best operating point is setting-dependent, underscoring the practical value of the search space itself.

cs.LG

Floquet spintronics: tuning the current-induced spin polarization of topological surface states with light

Topological surface states are a promising platform for spintronics due to spin-momentum locking. Spin-momentum locking can induce a net spin polarization in topological surface states via an electric current, a phenomenon known as the Edelstein effect. In this work, using Floquet theory, we show that the current-induced spin polarization of topological surface states can be tuned by illuminating them with light, thereby modifying their spin texture in momentum space. Specifically, the electric spin susceptibility of topological surface states can be controlled and even reversed by varying the electric-field strength of high-frequency, circularly polarized light.

cond-mat.mes-hall

Light-Wave Engineering for Selective Polarization of a Single $\mathbf{Q}$ Valley in Transition Metal Dichalcogenides

The selective control of specific momentum valleys lies at the core of valleytronics, a field that has thus far focused primarily on the $\mathbf{K}$ and $\mathbf{K'}$ valleys in transition metal dichalcogenides (TMDs). However, direct optical access to other low-lying yet conventionally inaccessible valleys such as the sixfold degenerate $\mathbf{Q}$ valleys has remained an outstanding challenge, fundamentally limiting the exploitation of the full valley degree of freedom for information processing. Here, we theoretically introduce an emergent light-wave valley selection rule that enables deterministic and high fidelity excitation of any single $\mathbf{Q}$ valley in monolayer TMDs. By coherently combining a circularly polarized pump pulse with a linearly polarized driver pulse, we engineer distinct quantum pathways that unambiguously excited electrons into a targeted $\mathbf{Q}$ valley, completely decoupled from the conventional $\mathbf{K}/\mathbf{K'}$ valleys. This all-optical scheme achieves near-unity ($\sim$100\%) valley polarization across an exceptionally broad ultrafast window, from the terahertz ($10^{12}$~Hz) to petahertz ($10^{15}$~Hz) regimes, enabling single $\mathbf{Q}$ valley polarization on femtosecond timescales. Our findings establish a new paradigm of light-wave quantum metrology in valleytronics, unlocking the $\mathbf{Q}$-valley subspace for scalable multi-state valley information processing.

cond-mat.str-el

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference

The increasing deployment of Large Language Model (LLM) inference on edge AI systems demands efficient execution under tight memory budgets. A key challenge arises from Key-Value (KV) caches, which often exceed available device memory. Although NVMe-based offloading offers scalable capacity, existing file-based designs rely heavily on the kernel page cache, leading to cache thrashing, unpredictable latency, and high software overhead under memory pressure. We present DUAL-BLADE, a dual-path KV residency framework that dynamically assigns KV tensors to either a page-cache path or an NVMe-direct path based on runtime memory availability. The NVMe-direct path bypasses the filesystem by mapping KV tensors to contiguous logical block address (LBA) regions, enabling low-overhead direct storage access. DUAL-BLADE further incorporates adaptive pipeline parallelism to overlap storage I/O with GPU DMA, improving inference throughput. Our evaluation shows that DUAL-BLADE substantially mitigates I/O bottlenecks, reducing prefill and decode latency by up to 33.1% and 42.4%, respectively, while improving SSD utilization by 2.2x across diverse memory budgets.

cs.DC

Data-driven oscillator model for multi-frequency turbulent flows

The complex dynamics of high-dimensional oscillatory flows can be simplified using phase-reduction analysis, providing a deeper understanding of the flow response to external perturbations. Although phase-based modeling and analysis have been utilized in recent studies on oscillatory fluid flows, their usages are still limited to single-frequency flows due to difficulties in addressing chaotic characteristics induced by multiple frequencies of turbulent flows. In order to overcome this limitation, we propose a data-driven framework that models the dynamics of multi-frequency turbulent flows based on a set of oscillators. The representative oscillators are extracted from the flow field data by training specially designed autoencoders. The oscillator dynamics are modeled through a machine-learning technique using neural networks to accurately predict the multi-frequency oscillatory behavior of turbulent flows. We verify the oscillator-based model of the multi-frequency turbulent flow by applying the proposed data-driven method to the three-dimensional supersonic turbulent flow over a cavity. We show that the extracted oscillators represent the dominant large-scale flow features and reflect the physical characteristics of the turbulent cavity flow. The data-driven oscillator dynamics model accurately forecasts the oscillatory behavior of the turbulent cavity flow for a long period. The proposed data-driven method for reduced-order modeling of turbulent flows with oscillators will enable deeper investigations of perturbation dynamics and control of turbulent flows.

physics.flu-dyn

Ultrafast Current Switching from Quantum Geometry in Semimetals

Technological progress towards next-generation electronics critically relies on achieving faster switching with reduced energy consumption. Because device operation speeds are fundamentally constrained by the intrinsic properties of constituent materials, identifying systems with inherently superior switching capabilities is essential. Here, we propose that semimetallic systems characterized by non-trivial quantum geometry, including quadratic band-touching semimetals and singular flat bands, can serve as a promising platform for ultrafast switching at voltages compatible with modern electronics. We show that, in such quantum geometric semimetals, an electric current is generated instantaneously upon application of a moderate external electric field, reaching its steady-state value. As a consequence, the current exhibits rapid and stable on-off switching behaviour under periodic optical pulse trains, demonstrating robustness under experimentally feasible conditions. In terms of switching speed, this quantum geometric semimetal outperforms conventional metals, semiconductors, and graphene. We identify the microscopic origin of this behaviour as interband coupling governed by the Hilbert-Schmidt quantum distance, together with a finite density of states at the band-touching point. This mechanism further leads to a universal classification of conductivity for both gapless and gapped quantum geometric semimetals. Finally, first-principles calculations suggest realistic material platforms, including bilayer graphene, cyclic graphene, monolayer bismuth and V3F8-in which the predicted instantaneous current switching can be directly realized, further supported by time-dependent density functional theory simulations performed for representative systems.

cond-mat.str-el

High Photovoltaic Efficiency in Bulk-Stacked One-Dimensional GeSe$_{2}$ van der Waals Crystal

Germanium diselenide (GeSe$_{2}$) has recently attracted substantial interest as a rare example of one-dimensional (1D) van der Waals material. Here, we investigate the photovoltaic potential of bulk-stacked GeSe$_{2}$ chains using first-principles calculations within the $GW0$ approximation and the Bethe-Salpeter equation (BSE) to capture quasiparticle and excitonic effects. The bulk GeSe$_{2}$ exhibits indirect GW band gaps of 1.92 eV (type-I) and 1.08 eV (type-II). Optical calculations show markedly stronger visible-light absorption in type-II, yielding a spectroscopically limited maximum efficiency (SLME) of ~25.6% at a 0.5 $μ$m thickness. Phonon and room-temperature ab initio molecular dynamics analyses indicate that type-II is dynamically stable, whereas type-I shows imaginary phonon modes, suggesting a propensity for structural distortion. These results identify type-II GeSe2 as a promising stable absorber for thin-film photovoltaics with enhanced flexibility compared to typical 2D vdW systems.

cond-mat.mtrl-sci

AFLL: Real-time Load Stabilization for MMO Game Servers Based on Circular Causality Learning

Massively Multiplayer Online (MMO) game servers must handle thousands of simultaneous players while maintaining sub-100ms response times. When server load exceeds capacity, traditional approaches either uniformly throttle all message types regardless of importance (damaging gameplay) or apply fixed heuristic rules that fail to adapt to dynamic workloads. This paper presents AFLL (Adaptive Feedback Loop Learning), a real-time load stabilization system that learns the causal relationship between outgoing server messages and subsequent incoming client requests. AFLL employs backpropagation to continuously adjust message type weights, enabling predictive throttling that blocks low-priority messages before overload occurs while guaranteeing critical message delivery. Through controlled experiments with 1,000 concurrent players, AFLL reduced average CPU time by 48.3% (13.2ms to 6.8ms), peak CPU time by 51.7% (54.0ms to 26.1ms), and thread contention by 64.4% (19.6% to 7.0%), while maintaining zero learning overhead through background computation and caching optimizations. The system achieved remarkable reproducibility (CV < 2% across all metrics) and identified a three-stage causal chain linking message blocking to load reduction. AFLL demonstrates that circular causality learning enables practical real-time adaptation for latency-critical systems.

cs.DC

Floquet Chern Insulators and Radiation-Induced Zero Resistance in Irradiated Graphene

Recent advances in optics and time-resolved techniques have facilitated the exploration of new states of matter under nonequilibrium conditions. Here, we predict that irradiated graphene can host two novel nonequilibrium steady states of matter with zero resistance when exposed to circularly polarized light: (i) Floquet Chern insulators and (ii) a radiation-induced zero-resistance state with spontaneous formation of an inhomogeneous current distribution. Specifically, we calculate nonequilibrium anomalous Hall and longitudinal conductivities to map the nonequilibrium phase diagram of irradiated graphene as a function of the driving frequency and the electric-field strength of circularly polarized light. As a result, Floquet Chern insulators are found to occur at high driving frequencies above the graphene band width. By contrast, at low driving frequencies below the graphene band width, the nonequilibrium anomalous Hall conductivity deviates from the expected quantized values, and the nonequilibrium longitudinal conductivity exhibits highly irregular behavior, including negative resistance. It is predicted that the thermodynamically unstable negative resistance will trigger a catastrophic breakdown, inducing a zero-resistance state with spontaneous formation of an inhomogeneous current distribution, similar to the radiation-induced zero-resistance state observed in quantum Hall systems.

cond-mat.mes-hall

Progressive Facial Granularity Aggregation with Bilateral Attribute-based Enhancement for Face-to-Speech Synthesis

For individuals who have experienced traumatic events such as strokes, speech may no longer be a viable means of communication. While text-to-speech (TTS) can be used as a communication aid since it generates synthetic speech, it fails to preserve the user's own voice. As such, face-to-voice (FTV) synthesis, which derives corresponding voices from facial images, provides a promising alternative. However, existing methods rely on pre-trained visual encoders, and finetune them to align with speech embeddings, which strips fine-grained information from facial inputs such as gender or ethnicity, despite their known correlation with vocal traits. Moreover, these pipelines are multi-stage, which requires separate training of multiple components, thus leading to training inefficiency. To address these limitations, we utilize fine-grained facial attribute modeling by decomposing facial images into non-overlapping segments and progressively integrating them into a multi-granular representation. This representation is further refined through multi-task learning of speaker attributes such as gender and ethnicity at both the visual and acoustic domains. Moreover, to improve alignment robustness, we adopt a multi-view training strategy by pairing various visual perspectives of a speaker in terms of different angles and lighting conditions, with identical speech recordings. Extensive subjective and objective evaluations confirm that our approach substantially enhances face-voice congruence and synthesis stability.

cs.SD

A GND-based back stress model for reverse loading in metal sheets with consideration of GNB

Accurate prediction of springback and formability in sheet metal forming requires understanding reverse loading behavior under complex loading path changes, such as tension followed by compression. However, for ultra-thin sheets experimental characterization of such behavior is difficult due to compressive instability like plastic buckling. This study presents a crystal plasticity finite element method (CPFEM) incorporating a physically motivated back stress model based on geometrically necessary dislocations (GNDs) and boundaries (GNBs). The model captures grain size effects, including the Hall-Petch and Bauschinger effects, through a single grain size-dependent back stress parameter, enabling reverse loading prediction using only tensile data from specimens with different grain sizes. The back stress parameter was calibrated by fitting tensile stress-strain curves from two microstructures - one as-received and one annealed. Without using Tension-Compression (T-C) data for calibration, the model accurately predicted reverse loading behavior in low-carbon steel (0.64 mm thick) and Tension-Bending (T-B) responses in ultra-thin SUS316 (0.083 mm thick) when the developed theory was incorporated to an upscaled anisotropic hardening model. Identifiability analysis confirmed that the model parameters are uniquely determined by the available data. This physically interpretable framework provides an efficient and robust means to predict reverse loading in thin metal sheets, overcoming experimental limitations.

cond-mat.mtrl-sci

Facilitating Personalized TTS for Dysarthric Speakers Using Knowledge Anchoring and Curriculum Learning

Dysarthric speakers experience substantial communication challenges due to impaired motor control of the speech apparatus, which leads to reduced speech intelligibility. This creates significant obstacles in dataset curation since actual recording of long, articulate sentences for the objective of training personalized TTS models becomes infeasible. Thus, the limited availability of audio data, in addition to the articulation errors that are present within the audio, complicates personalized speech synthesis for target dysarthric speaker adaptation. To address this, we frame the issue as a domain transfer task and introduce a knowledge anchoring framework that leverages a teacher-student model, enhanced by curriculum learning through audio augmentation. Experimental results show that the proposed zero-shot multi-speaker TTS model effectively generates synthetic speech with markedly reduced articulation errors and high speaker fidelity, while maintaining prosodic naturalness.

cs.SD

CALL: Context-Aware Low-Latency Retrieval in Disk-Based Vector Databases

Embedding models capture both semantic and syntactic structures of queries, often mapping different queries to similar regions in vector space. This results in non-uniform cluster access patterns in modern disk-based vector databases. While existing approaches optimize individual queries, they overlook the impact of cluster access patterns, failing to account for the locality effects of queries that access similar clusters. This oversight increases cache miss penalty. To minimize the cache miss penalty, we propose CALL, a context-aware query grouping mechanism that organizes queries based on shared cluster access patterns. Additionally, CALL incorporates a group-aware prefetching method to minimize cache misses during transitions between query groups and latency-aware cluster loading. Experimental results show that CALL reduces the 99th percentile tail latency by up to 33% while consistently maintaining a higher cache hit ratio, substantially reducing search latency.

cs.DB

OASIS: Object-based Analytics Storage for Intelligent SQL Query Offloading in Scientific Tabular Workloads

Computation-Enabled Object Storage (COS) systems, such as MinIO and Ceph, have recently emerged as promising storage solutions for post hoc, SQL-based analysis on large-scale datasets in High-Performance Computing (HPC) environments. By supporting object-granular layouts, COS facilitates column-oriented access and supports in-storage execution of data reduction operators, such as filters, close to where the data resides. Despite growing interest and adoption, existing COS systems exhibit several fundamental limitations that hinder their effectiveness. First, they impose rigid constraints on output data formats, limiting flexibility and interoperability. Second, they support offloading for only a narrow set of operators and expressions, restricting their applicability to more complex analytical tasks. Third--and perhaps most critically--they fail to incorporate design strategies that enable compute offloading optimized for the characteristics of deep storage hierarchies. To address these challenges, this paper proposes OASIS, a novel COS system that features: (i) flexible and interoperable output delivery through diverse formats, including columnar layouts such as Arrow; (ii) broad support for complex operators (e.g., aggregate, sort) and array-aware expressions, including element-wise predicates over array structures; and (iii) dynamic selection of optimal execution paths across internal storage layers, guided by operator characteristics and data movement costs. We implemented a prototype of OASIS and integrated it into the Spark analytics framework. Through extensive evaluation using real-world scientific queries from HPC workflows, OASIS achieves up to a 32.7% performance improvement over Spark configured with existing COS-based storage systems.

cs.DB

OpenCXD: An Open Real-Device-Guided Hybrid Evaluation Framework for CXL-SSDs

The advent of Compute Express Link (CXL) enables SSDs to participate in the memory hierarchy as large-capacity, byte-addressable memory devices. These CXL-enabled SSDs (CXL-SSDs) offer a promising new tier between DRAM and traditional storage, combining NAND flash density with memory-like access semantics. However, evaluating the performance of CXL-SSDs remains difficult due to the lack of hardware that natively supports the CXL.mem protocol on SSDs. As a result, most prior work relies on hybrid simulators combining CPU models augmented with CXL.mem semantics and SSD simulators that approximate internal flash behaviors. While effective for early-stage exploration, this approach cannot faithfully model firmware-level interactions and low-level storage dynamics critical to CXL-SSD performance. In this paper, we present OpenCXD, a real-device-guided hybrid evaluation framework that bridges the gap between simulation and hardware. OpenCXD integrates a cycle-accurate CXL.mem simulator on the host side with a physical OpenSSD platform running real firmware. This enables in-situ firmware execution triggered by simulated memory requests. Through these contributions, OpenCXD reflects device-level phenomena unobservable in simulation-only setups, providing critical insights for future firmware design tailored to CXL-SSDs.

cs.AR

Shared Disk KV Cache Management for Efficient Multi-Instance Inference in RAG-Powered LLMs

Recent large language models (LLMs) face increasing inference latency as input context length and model size continue to grow. In particular, the retrieval-augmented generation (RAG) technique, which enhances LLM responses by incorporating external knowledge, exacerbates this issue by significantly increasing the number of input tokens. This expansion in token length leads to a substantial rise in computational overhead, particularly during the prefill stage, resulting in prolonged time-to-first-token (TTFT). To address this issue, this paper proposes a method to reduce TTFT by leveraging a disk-based key-value (KV) cache to lessen the computational burden during the prefill stage. We also introduce a disk-based shared KV cache management system, called Shared RAG-DCache, for multi-instance LLM RAG service environments. This system, together with an optimal system configuration, improves both throughput and latency under given resource constraints. Shared RAG-DCache exploits the locality of documents related to user queries in RAG, as well as the queueing delay in LLM inference services. It proactively generates and stores disk KV caches for query-related documents and shares them across multiple LLM instances to enhance inference performance. In experiments on a single host equipped with 2 GPUs and 1 CPU, Shared RAG-DCache achieved a 15~71% increase in throughput and up to a 12~65% reduction in latency, depending on the resource configuration.

cs.AI

Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading

LLM inference is essential for applications like text summarization, translation, and data analysis, but the high cost of GPU instances from Cloud Service Providers (CSPs) like AWS is a major burden. This paper proposes InferSave, a cost-efficient VM selection framework for cloud based LLM inference. InferSave optimizes KV cache offloading based on Service Level Objectives (SLOs) and workload charac teristics, estimating GPU memory needs, and recommending cost-effective VM instances. Additionally, the Compute Time Calibration Function (CTCF) improves instance selection accuracy by adjusting for discrepancies between theoretical and actual GPU performance. Experiments on AWS GPU instances show that selecting lower-cost instances without KV cache offloading improves cost efficiency by up to 73.7% for online workloads, while KV cache offloading saves up to 20.19% for offline workloads.

cs.LG