SearcharxivSearch

arXiv subjects

Dohyeon Kim

Publications and source records attributed to Dohyeon Kim.

14 recordsLinked to original sources

Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts

Mixture-of-Experts (MoE) architectures have emerged as a powerful paradigm for scaling model capacity while preserving efficient inference in large foundation models. However, most MoE models use a fixed top-$k$ expert selection policy, assigning the same expert budget to every token even when fewer experts may be sufficient. Inference-time dynamic top-$k$ routing can reduce computation without retraining, but existing methods often overlook the distributional shift caused by deviating from the training-time routing configuration. We show that reducing the number of activated experts consistently increases the RMS scale and variance of SMoE outputs, inducing a representation mismatch that contributes to downstream performance degradation in addition to the loss of expert capacity. To address this correctable component, we propose Layer-wise Distribution Alignment (LDA), a lightweight inference-time correction that uses layer-wise calibration statistics to align reduced-routing representations with the default configuration. Across multiple SMoE LLMs, benchmarks, and routing strategies, LDA recovers much of the performance lost induced by the distributional shift under reduced routing while preserving sparse-inference efficiency with negligible overhead.

cs.LG

Quantitative Target Convergence and Uniform-in-Time Propagation of Chaos for Langevin-Regularized SVGD

We establish quantitative convergence to the target and uniform-in-time propagation of chaos for Langevin-regularized Stein variational gradient descent. The Stein interaction need not be small relative to the confining Langevin drift and does not generally yield a contractive particle coupling. At the mean-field level, the Stein and Langevin components dissipate the same relative entropy in the kernel-induced Stein and $2$-Wasserstein geometries, producing the squared kernel Stein discrepancy and relative Fisher information. Under a log-Sobolev inequality for the target, this yields exponential last-iterate convergence. We also derive a finite-particle entropy identity relative to the product target, giving exponential-in-time convergence of the empirical measure up to polynomial sampling errors. For propagation of chaos, we develop two complementary finite-time approaches. A synchronous coupling, combined with exponential moment estimates for the nonlinear mean-field diffusion, yields explicit single-exponential bounds in Wasserstein distance and kernel Stein discrepancy (KSD). Moving-product entropy gives joint-law relative entropy control relative to the evolving mean-field product law and, through entropy superadditivity and concentration, fixed-marginal relative entropy and total variation bounds and empirical KSD estimates. Under an additional $T_2$ inequality for the initial law, it also yields Wasserstein bounds. Combining these finite-time estimates with target convergence at a logarithmic cutoff time gives polynomial uniform-in-time propagation of chaos rates in expectation for empirical KSD and $W_2^2$, and for fixed-marginal total variation and $W_2^2$. All bounds control the last iterate in physical time. We also compare the two finite-time mechanisms and identify regimes in which each gives the sharper polynomial exponent.

stat.ML

Long-time Stability and Convergence of Particle Swarm Optimization

Particle Swarm Optimization (PSO) is a global optimization algorithm defined by an interacting set of particles evolving over the search space. Heuristically motivated, its theoretical analysis remains limited due to the second-order, stochastic, and highly nonlinear nature of the dynamics. In this paper, we connect classical PSO stability analysis under the stagnation assumption with more recent mean-field methods, providing new quantitative estimates for the time-discrete algorithm. We study in particular a regularized PSO model without memory, with non-degenerate noise by adding a noise floor to the original model. Studying such a surrogate model allows us to identify quantitative conditions under which the dynamics is stable and converges toward a small neighborhood of a global minimizer. We do so by first studying the Schur stability of the linearized dynamics, then analyzing the convergence properties of a nonlinear mean-field system via a Laplace principle, and finally establishing a quantitative error bound for the mean-field approximation of order $N^{-1/2}$.

math.OC

Uniform-in-time propagation of chaos for Second-Order Consensus-Based Optimization

We study second-order Consensus-Based Optimization (CBO), a derivative-free global optimization algorithm in which the consensus force and the multiplicative exploratory noise act on particle velocities. We prove quantitative uniform-in-time propagation of chaos for the unmodified second-order CBO dynamics, together with an almost uniform-in-time stability estimate for the microscopic particle system. The proof is not a direct adaptation of the first-order CBO argument. Although both first- and second-order CBO have multiplicative noise that degenerates near consensus and a shift-invariant weighted interaction, the kinetic model has an additional structural obstruction: the consensus mechanism and the stochastic forcing act only on the velocity variable, while the position variable evolves by transport. Thus spatial concentration has to be recovered indirectly through velocity dissipation. Moreover, the shift-invariant interaction leaves a translation mode that is not directly damped by the consensus force, so a standard synchronous coupling in the Euclidean phase-space distance does not close uniformly in time. The main idea of the paper is to introduce shifted internal variables that separate the contracting fluctuation modes from the undamped translation mode. In these variables we build a Lyapunov functional with a position-velocity cross term and prove exponential decay of centered moments. This decay is the mechanism that makes the time-dependent coupling coefficient integrable. Combining it with uniform-in-time raw moment bounds, concentration inequalities, stability estimates for the weighted mean, and a Monte Carlo estimate, we obtain the classical Monte Carlo rate for propagation of chaos uniformly in time. The system-to-system stability estimate avoids the sampling error and yields the faster rate \(O(J^{-q})\).

math.PR

RLDX-1 Technical Report

While Vision-Language-Action models (VLAs) have shown remarkable progress toward human-like generalist robotic policies through the versatile intelligence (i.e. broad scene understanding and language-conditioned generalization) inherited from pre-trained Vision-Language Models, they still struggle with complex real-world tasks requiring broader functional capabilities (e.g. motion awareness, long-term memory, and physical sensing). To address this, we introduce RLDX-1, a general-purpose robotic policy for dexterous manipulation built on the Multi-Stream Action Transformer (MSAT), an architecture that unifies these capabilities by integrating heterogeneous modalities through modality-specific streams with cross-modal joint self-attention. RLDX-1 further combines this architecture with system-level design choices, including data synthesis for rare manipulation scenarios, learning procedures specialized for human-like manipulation, and inference optimizations for real-time deployment. Through empirical evaluation, we show that RLDX-1 consistently outperforms recent frontier VLAs (e.g. $\pi_{0.5}$ and GR00T N1.6) across both simulation benchmarks and real-world tasks that require broad functional capabilities beyond general versatility. In particular, RLDX-1 shows superiority in ALLEX humanoid tasks by achieving success rates of 86.8% while $\pi_{0.5}$ and GR00T N1.6 achieve around 40%, highlighting the ability of RLDX-1 to control a high-DoF humanoid robot under diverse functional demands. Together, these results position RLDX-1 as a promising step toward reliable VLAs for complex, contact-rich, and dynamic real-world dexterous manipulation.

cs.RO

EgoX: Egocentric Video Generation from a Single Exocentric Video

Egocentric perception enables humans to experience and understand the world directly from their own point of view. Translating exocentric (third-person) videos into egocentric (first-person) videos opens up new possibilities for immersive understanding but remains highly challenging due to extreme camera pose variations and minimal view overlap. This task requires faithfully preserving visible content while synthesizing unseen regions in a geometrically consistent manner. To achieve this, we present EgoX, a novel framework for generating egocentric videos from a single exocentric input. EgoX leverages the pretrained spatio temporal knowledge of large-scale video diffusion models through lightweight LoRA adaptation and introduces a unified conditioning strategy that combines exocentric and egocentric priors via width and channel wise concatenation. Additionally, a geometry-guided self-attention mechanism selectively attends to spatially relevant regions, ensuring geometric coherence and high visual fidelity. Our approach achieves coherent and realistic egocentric video generation while demonstrating strong scalability and robustness across unseen and in-the-wild videos.

cs.CV

How Far Can LLMs Emulate Human Behavior?: A Strategic Analysis via the Buy-and-Sell Negotiation Game

With the rapid advancement of Large Language Models (LLMs), recent studies have drawn attention to their potential for handling not only simple question-answer tasks but also more complex conversational abilities and performing human-like behavioral imitations. In particular, there is considerable interest in how accurately LLMs can reproduce real human emotions and behaviors, as well as whether such reproductions can function effectively in real-world scenarios. However, existing benchmarks focus primarily on knowledge-based assessment and thus fall short of sufficiently reflecting social interactions and strategic dialogue capabilities. To address these limitations, this work proposes a methodology to quantitatively evaluate the human emotional and behavioral imitation and strategic decision-making capabilities of LLMs by employing a Buy and Sell negotiation simulation. Specifically, we assign different personas to multiple LLMs and conduct negotiations between a Buyer and a Seller, comprehensively analyzing outcomes such as win rates, transaction prices, and SHAP values. Our experimental results show that models with higher existing benchmark scores tend to achieve better negotiation performance overall, although some models exhibit diminished performance in scenarios emphasizing emotional or social contexts. Moreover, competitive and cunning traits prove more advantageous for negotiation outcomes than altruistic and cooperative traits, suggesting that the assigned persona can lead to significant variations in negotiation strategies and results. Consequently, this study introduces a new evaluation approach for LLMs' social behavior imitation and dialogue strategies, and demonstrates how negotiation simulations can serve as a meaningful complementary metric to measure real-world interaction capabilities-an aspect often overlooked in existing benchmarks.

cs.AI

Uniform-in-time propagation of chaos for Consensus-Based Optimization

We study the derivative-free global optimization algorithm Consensus-Based Optimization (CBO), establishing uniform-in-time propagation of chaos as well as an almost uniform-in-time stability result for the microscopic particle system. Moreover, we prove almost sure exponential convergence of the microscopic CBO system around a point close to the global minimizer. The proof of these results is based on a novel stability estimate for the weighted mean and on a quantitative concentration inequality for the microscopic particle system around the empirical mean. Our propagation of chaos result recovers the classical Monte Carlo rate, with a prefactor that depends explicitly on the parameters of the problem. Notably, in the case of CBO with anisotropic noise, this prefactor is independent of the problem dimension.

math.PR

MirrorCBO: A consensus-based optimization method in the spirit of mirror descent

In this work we propose MirrorCBO, a consensus-based optimization (CBO) method which generalizes standard CBO in the same way that mirror descent generalizes gradient descent. For this we apply the CBO methodology to a swarm of dual particles and retain the primal particle positions by applying the inverse of the mirror map, which we parametrize as the subdifferential of a strongly convex function $\phi$. In this way, we combine the advantages of a derivative-free non-convex optimization algorithm with those of mirror descent. As a special case, the method extends CBO to optimization problems with convex constraints. Assuming bounds on the Bregman distance associated to $\phi$, we provide asymptotic convergence results for MirrorCBO with explicit exponential rate. Another key contribution is an exploratory numerical study of this new algorithm across different application settings, focusing on (i) sparsity-inducing optimization, and (ii) constrained optimization, demonstrating the competitive performance of MirrorCBO. We observe empirically that the method can also be used for optimization on (non-convex) submanifolds of Euclidean space, can be adapted to mirrored versions of other recent CBO variants, and that it inherits from mirror descent the capability to select desirable minimizers, like sparse ones. We also include an overview of recent CBO approaches for constrained optimization and compare their performance to MirrorCBO.

math.OC

Mono-cluster flocking and uniform-in-time stability of the discrete Motsch-Tadmor model

The Motsch-Tadmor (MT) model is a variant of the Cucker-Smale model with a normalized communication weight function. The normalization poses technical challenges in analyzing the collective behavior due to the absence of conservation of momentum. We study three quantitative estimates for the discrete-time MT model considering the first-order Euler discretization. First, we provide a sufficient framework leading to the asymptotic mono-cluster flocking. The proposed framework is given in terms of coupling strength, communication weight function, and initial data. Second, we show that the continuous transition from the discrete MT model to the continuous MT model can be made uniformly in time using the finite-time convergence result and asymptotic flocking estimate. Third, we present uniform-in-time stability estimates for the discrete MT model. We also provide several numerical examples and compare them with analytical results.

math.NA

Statistical Analysis by Semiparametric Additive Regression and LSTM-FCN Based Hierarchical Classification for Computer Vision Quantification of Parkinsonian Bradykinesia

Bradykinesia, characterized by involuntary slowing or decrement of movement, is a fundamental symptom of Parkinson's Disease (PD) and is vital for its clinical diagnosis. Despite various methodologies explored to quantify bradykinesia, computer vision-based approaches have shown promising results. However, these methods often fall short in adequately addressing key bradykinesia characteristics in repetitive limb movements: "occasional arrest" and "decrement in amplitude." This research advances vision-based quantification of bradykinesia by introducing nuanced numerical analysis to capture decrement in amplitudes and employing a simple deep learning technique, LSTM-FCN, for precise classification of occasional arrests. Our approach structures the classification process hierarchically, tailoring it to the unique dynamics of bradykinesia in PD. Statistical analysis of the extracted features, including those representing arrest and fatigue, has demonstrated their statistical significance in most cases. This finding underscores the importance of considering "occasional arrest" and "decrement in amplitude" in bradykinesia quantification of limb movement. Our enhanced diagnostic tool has been rigorously tested on an extensive dataset comprising 1396 motion videos from 310 PD patients, achieving an accuracy of 80.3%. The results confirm the robustness and reliability of our method.

cs.CV

A Reliable, Self-Adaptive Face Identification Framework via Lyapunov Optimization

Realtime face identification (FID) from a video feed is highly computation-intensive, and may exhaust computation resources if performed on a device with a limited amount of resources (e.g., a mobile device). In general, FID performs better when images are sampled at a higher rate, minimizing false negatives. However, performing it at an overwhelmingly high rate exposes the system to the risk of a queue overflow that hampers the system's reliability. This paper proposes a novel, queue-aware FID framework that adapts the sampling rate to maximize the FID performance while avoiding a queue overflow by implementing the Lyapunov optimization. A preliminary evaluation via a trace-based simulation confirms the effectiveness of the framework.

cs.DC

Modelling the Intrusive feelings of advanced driver assistance systems based on vehicle activity log data: a case study for the lane keeping assistance system

Although the automotive industry has been among the sectors that best-understands the importance of drivers' affect, the focus of design and research in the automotive field has long emphasized the visceral aspects of exterior and interior design. With the adoption of Advanced Driver Assistance Systems (ADAS), endowing 'semi-autonomy' to the vehicles, however, the scope of affective design should be expanded to include the behavioural aspects of the vehicle. In such a 'shared-control' system wherein the vehicle can intervene in the human driver's operations, a certain degree of 'intrusive feelings' are unavoidable. For example, when the Lane Keeping Assistance System (LKAS), one of the most popular examples of ADAS, operates the steering wheel in a dangerous situation, the driver may feel interrupted or surprised because of the abrupt torque generated by LKAS. This kind of unpleasant experience can lead to prolonged negative feelings such as irritation, anxiety, and distrust of the system. Therefore, there are increasing needs of investigating the driver's affective responses towards the vehicle's dynamic behaviour. In this study, four types of intrusive feelings caused by LKAS were identified to be proposed as a quantitative performance indicator in designing the affectively satisfactory behaviour of LKAS. A metric as well as a statistical data analysis method to quantitatively measure the intrusive feelings through the vehicle sensor log data.

cs.HC

Usability of the Size, Spacing, and Depth of Virtual Buttons on Head-Mounted Displays

Virtual reality (VR) allows users to see and manipulate virtual scenes and items through input devices, like head-mounted displays. In this study, the effects of button size, spacing, and depth on the usability of virtual buttons in VR environments were investigated. Task completion time, number of errors, and subjective preferences were collected to test different levels of the button size, spacing, and depth. The experiment was conducted in a desktop setting with Oculus Rift and Leap motion. A total of 18 subjects performed a button selection task. The optimal levels of button size and spacing within the experimental conditions are 25 mm and between 5 mm and 9 mm, respectively. Button sizes of 15 mm with 1-mm spacing were too small to be used in VR environments. A trend of decreasing task completion time and the number of errors was observed as button size and spacing increased. However, large size and spacing may cause fatigue, due to continuous extension of the arms. For depth effects, the touch method took a shorter task completion time. However, the push method recorded a smaller number of errors, owing to the visual push-feedback. In this paper, we discuss advantages and disadvantages in detail. The results can be applied to many different application areas with VR HMD.

cs.HC