SearcharxivSearch

arXiv subjects

Kihyun Kim

Publications and source records attributed to Kihyun Kim.

At least 19 recordsLinked to original sources

Safety in Batches? Understanding and Mitigating Safety Failures in Batch Prompting

Batch prompting is a practical inference strategy for large language models, but its safety implications remain underexplored. We show that the success of batch prompting for utility does not extend to safety: a harmful question that is reliably refused in isolation can elicit a harmful response when embedded in a batch of benign questions. We identify this as a distinct safety failure mode -- not reducible to known vulnerabilities such as in-context learning or long-context effects -- and analyze its causes from two complementary perspectives: alignment signal weakening and refusal signal dilution. Across widely used open-source and frontier commercial models, batch prompting consistently achieves high attack success rates as a simple black-box attack. We further show that batch-aware preference optimization effectively mitigates the vulnerability. These findings highlight a blind spot in current safety alignment and point to batch-aware alignment as a necessary step toward robust deployment. Code is available at https://github.com/96kihyun/batch_jailbreak

cs.CR

Fine-Grained Multi Image Object Hallucination Benchmark

Multimodal Large Language Models (MLLMs) are increasingly deployed in multi-image scenarios requiring complex reasoning across visual contexts. However, current MLLMs remain fundamentally limited by object hallucination-generating plausible yet factually inconsistent descriptions about objects. Existing benchmarks, designed primarily for single-image settings or providing only high-level multi-image assessments, cannot systematically diagnose how visual complexity and reasoning demands trigger hallucination. To address this gap, we introduce MIOH, a fine-grained multi-image object hallucination benchmark that systematically evaluates object hallucination across four foundational tasks (existence, counting, attribute, position) through three multi-image reasoning patterns (comprehensive, comparative, selective) under three controlled adversarial pressures (visual context scale, perceptual difficulty, contextual bias). Through evaluation of 29 models, we reveal that even state-of-the-art systems like GPT-5 and Gemini-2.5-Pro exhibit distinct failure patterns across different reasoning patterns and tasks. Our evaluation reveals that hallucination stems not merely from perceptual failures but from integration-stage limitations when maintaining object representations across multiple images. MIOH provides a controlled framework for analyzing multi-image object hallucination and serves as a critical evaluation tool for developing more reliable multimodal AI systems.

cs.CV

Enhanced Detection of Tiny Objects in Aerial Images

While one-stage detectors like YOLOv8 offer fast training speed, they often under-perform on detecting small objects as a trade-off. This becomes even more critical when detecting tiny objects in aerial imagery due to low-resolution targets and cluttered backgrounds. To address this, we introduce four enhancement strategies-input image resolution adjustment, data augmentation, attention mechanisms, and an alternative gating function for attention modules-that can be easily implemented on YOLOv8. We demonstrate that image size enlargement and the proper use of augmentation can lead to enhancement. Additionally, we designed a Mixture of Orthogonal Neural-modules Network (MoonNet) pipeline which consists of multiple attention-module-augmented CNNs. Two well-known attention modules, Squeeze-and-Excitation (SE) Block and Convolutional Block Attention Module (CBAM), were integrated into the backbone of YOLOv8 to form the MoonNet design, and the MoonNet backbone obtained improved detection accuracy compared to the original YOLOv8 backbone and single-type attention-module-augmented backbones. MoonNet further proved its adaptability and potential by achieving state-of-the-art performance on a tiny-object benchmark when integrated with the YOLC model. Our code is available at: https://github.com/Kihyun11/MoonNet

cs.CV

PHASOR: Phase-Anchored Universal Action Representations for Humanoid Embodiments

Learning a good action embedding space is fundamental to scalable robot policy learning, yet existing methods treat action latents as task-specific intermediates rather than first-class representations. The resulting latents are unstructured, embodiment-specific, and weakly tied to motion semantics, limiting interpretability, controllability, and transferability across robots. We position the action embedding space itself as a first-class design target, with downstream policy quality emerging from representation quality. Exploiting motion's intrinsic periodicity, we factorize it into a phase manifold that captures cyclic structure via FFT-parametric coefficients, together with a pose branch that conditions the manifold on non-periodic configuration detail. Combined with motion-semantic distillation, this factorized structure yields a cross-embodiment motion manifold that is interpretable and embodiment-agnostic by design. Anchoring multiple humanoid robots to a shared human-pretrained manifold then produces a unified action embedding space across diverse platforms, achieving strong cross-embodiment retrieval and consistent gains on downstream robot tasks.

cs.RO

Inverse Reinforcement Learning without an Optimal Demonstrator: A Feasible Reward Set Approach

Inverse reinforcement learning (IRL) typically assumes demonstrations from a single optimal demonstrator, but in many applications data come from multiple imperfect demonstrators with heterogeneous suboptimality levels. We study reward learning in this setting through a feasible-reward-set framework: for each demonstrator, we encode its declared suboptimality level as a linear constraint and intersect the resulting feasible sets across demonstrators. Our theoretical analysis shows that the joint feasible set shrinks monotonically as data are added, and we give an exact characterization of when a new demonstrator strictly tightens it. We further establish two recovery guarantees for the feasible reward set of the ground-truth optimal demonstrator: one bound depends on closeness to the optimal occupancy, while the other requires only sufficient coverage and no near-optimal demonstrator. On the practical side, we introduce strategies to address the inherent reward ambiguity in the obtained reward set and provide an offline algorithm with function approximation for high-dimensional environments. Experiments in tabular grid-world and large language model (LLM) fine-tuning settings are consistent with the theoretical predictions and demonstrate the effectiveness of the proposed framework over baselines.

cs.LG

STAR-Teaming: A Strategy-Response Multiplex Network Approach to Automated LLM Red Teaming

While Large Language Models (LLMs) are widely used, they remain susceptible to jailbreak prompts that can elicit harmful or inappropriate responses. This paper introduces STAR-Teaming, a novel black-box framework for automated red teaming that effectively generates such prompts. STAR-Teaming integrates a Multi-Agent System (MAS) with a Strategy-Response Multiplex Network and employs network-driven optimization to sample effective attack strategies. This network-based approach recasts the intractable high-dimensional embedding space into a tractable structure, yielding two key advantages: it enhances the interpretability of the LLM's strategic vulnerabilities, and it streamlines the search for effective strategies by organizing the search space into semantic communities, thereby preventing redundant exploration. Empirical results demonstrate that STAR-Teaming significantly surpasses existing methods, achieving a higher attack success rate (ASR) at a lower computational cost. Extensive experiments validate the effectiveness and explainability of the Multiplex Network. The code is available at https://github.com/selectstar-ai/STAR-Teaming-paper.

cs.CL

Construction of infinite time bubble tower solutions to critical wave maps equation

We construct infinite time bubble tower solutions to the critical wave maps equation taking values in the two-sphere. More precisely, for any integers $k\geq3$ and $J\geq1$, we construct a solution that is global in one time direction, has $k$-corotational symmetry, and asymptotically decomposes into $J$-many concentric bubbles of alternating signs with asymptotically vanishing radiation. The scales of each bubble are of order $t^{-α_{j}}$ with $α_{j}=(\frac{k}{k-2})^{j-1}-1$. This shows the existence of multi-bubble solutions with an arbitrary number of bubbles in soliton resolution, provided that $k\geq3$, global existence in one time direction, and alternating signs are considered. Our proof is based on modulation analysis with the method of backward construction. The key new ingredient is a Morawetz-type functional that provides suitable monotonicity estimates for solutions around multi-bubble configurations.

math.AP

Beyond RLHF and NLHF: Population-Proportional Alignment under an Axiomatic Framework

Conventional preference learning methods often prioritize opinions held more widely when aggregating preferences from multiple evaluators. This may result in policies that are biased in favor of some types of opinions or groups and susceptible to strategic manipulation. To address this issue, we develop a novel preference learning framework capable of aligning aggregate opinions and policies proportionally with the true population distribution of evaluator preferences. Grounded in social choice theory, our approach infers the feasible set of evaluator population distributions directly from pairwise comparison data. Using these estimates, the algorithm constructs a policy that satisfies foundational axioms from social choice theory, namely monotonicity and Pareto efficiency, as well as our newly-introduced axioms of population-proportional alignment and population-bounded manipulability. Moreover, we propose a soft-max relaxation method that smoothly trades off population-proportional alignment with the selection of the Condorcet winner (which beats all other options in pairwise comparisons). Finally, we validate the effectiveness and scalability of our approach through experiments on both tabular recommendation tasks and large language model alignment.

cs.AI

CAGE: A Framework for Culturally Adaptive Red-Teaming Benchmark Generation

Existing red-teaming benchmarks, when adapted to new languages via direct translation, fail to capture socio-technical vulnerabilities rooted in local culture and law, creating a critical blind spot in LLM safety evaluation. To address this gap, we introduce CAGE (Culturally Adaptive Generation), a framework that systematically adapts the adversarial intent of proven red-teaming prompts to new cultural contexts. At the core of CAGE is the Semantic Mold, a novel approach that disentangles a prompt's adversarial structure from its cultural content. This approach enables the modeling of realistic, localized threats rather than testing for simple jailbreaks. As a representative example, we demonstrate our framework by creating KoRSET, a Korean benchmark, which proves more effective at revealing vulnerabilities than direct translation baselines. CAGE offers a scalable solution for developing meaningful, context-aware safety benchmarks across diverse cultures. Our dataset and evaluation rubrics are publicly available at https://github.com/selectstar-ai/CAGE-paper. (WARNING: This paper contains model outputs that can be offensive in nature.)

cs.CY

Blow-up dynamics for radial self-dual Chern-Simons-Schrödinger equation with prescribed asymptotic profile

We construct finite energy blow-up solutions for the radial self-dual Chern-Simons-Schrödinger equation with a continuum of blow-up rates. Our result stands in stark contrast to the rigidity of blow-up of $H^{3}$ solutions proved by the first author for equivariant index $m \geq 1$, where the soliton-radiation interaction is too weak to admit the present blow-up scenarios. It is optimal (up to an endpoint) in terms of the range of blow-up rates and the regularity of the asymptotic profiles in view of the authors' previous proof of $H^{1}$ soliton resolution for the self-dual Chern-Simons-Schrödinger equation in any equivariance class. Our approach is a backward construction combined with modulation analysis, starting from prescribed asymptotic profiles and deriving the corresponding blow-up rates from their strong interaction with the soliton. In particular, our work may be seen as an adaptation of the method of Jendrej-Lawrie-Rodriguez (developed for energy critical equivariant wave maps) to the Schrödinger case. However, the Schrödinger nature of the equation (in particular, the lack of finite speed of propagation) and the optimal range (up to the $H^{1}$-endpoint) of our blow-up construction give rise to new challenges. Notably, the construction of (approximate) radiation from the prescribed asymptotic profile is one of our key novelties and might be of independent interest.

math.AP

Rigidity results in multi-bubble dynamics for non-radial energy-critical heat equation

This paper concerns the classification of asymptotic behaviors in multi-bubble dynamics for the energy-critical nonlinear heat equations in large dimensions $N\geq7$ without symmetry. This multi-bubble dynamics appears naturally at least for a sequence of times in view of soliton resolution. We assume each bubble is given by the scalings and translations of $\pm W$ with (localized) non-colliding conditions for a sequence of times, where $W$ is the ground state. The case of one soliton was previously established and in particular there is no blow-up. We consider the case of $J\geq2$ solitons, where we expect only infinite-time blow-up. We are able to identify three different scenarios, where we have a continuous-in-time resolution with an unexpected universal blow-up speed. The first one is when one scaling is much larger than the others. In this case, one bubble does not concentrate (hence stabilize) and the other bubbles concentrate with the universal blow-up speed $t^{-2/(N-6)}$ together with strong sign constraints. Next, assuming we are not in the first scenario, we establish a non-degenerate condition on the positions of bubbles to obtain that all bubbles concentrate with the universal blow-up speed $t^{-1/(N-4)}$. The last case we consider is a degenerate, but not too much degenerate, scenario. Here again, we obtain that all bubbles concentrate with the universal blow-up speed $t^{-1/(N-3)}$. This last rate has not been discovered before. Our theorem covers the case of four or less bubbles and we provide the construction of examples. To our knowledge, this is the first classification result in the non-radial multi-bubble dynamics, where both the scales, positions, and signs enter the dynamics nontrivially.

math.AP

Classification of single-bubble blow-up solutions for Calogero--Moser derivative nonlinear Schrödinger equation

We study the Calogero--Moser derivative nonlinear Schrödinger equation (CM-DNLS), a mass-critical and completely integrable dispersive model. Recent works established finite-time blow-up constructions and soliton resolution, describing the asymptotic behaviors of blow-up solutions. In this paper, we go beyond soliton resolution and provide a sharp classification of finite-time blow-up dynamics in the \textit{single-bubble} regime. Assuming that a solution blows up at time $0<T<\infty$ with a single-soliton profile, we determine all possible blow-up rates. For initial data in $H^{2L+1}(\mathbb{R})$ with $L\ge1$, we prove a dichotomy: either the solution lies in a \emph{quantized regime}, where the scaling parameter satisfies \[ λ(t)\sim (T-t)^{2k},\qquad 1\le k\le L, \] with convergent phase and translation parameters, or it lies in an \emph{exotic regime}, where the blow-up rate satisfies $λ(t)\lesssim (T-t)^{2L+\frac 32}$. To our knowledge, this is the first classification result for quantized blow-up dynamics in the class of dispersive models. We provide a framework for identifying the quantized blow-up rates in classification problems. The proof relies on a modulation analysis combined with the hierarchy of conservation laws provided by the complete integrability of (CM-DNLS). However, it does not use \emph{more refined integrability-based techniques}, such as the inverse scattering method, the method of commuting flows, or the explicit formula. As a result, our analysis applies beyond the chiral solutions.

math.AP

Construction of smooth chiral finite-time blow-up solutions to Calogero--Moser derivative nonlinear Schrödinger equation

We consider the Calogero--Moser derivative nonlinear Schrödinger equation (CM-DNLS), which is an $L^{2}$-critical nonlinear Schrödinger equation with explicit solitons, self-duality, and pseudo-conformal symmetry. More importantly, this equation is known to be completely integrable in the Hardy space $L_{+}^{2}$ and the solutions in this class are referred to as \emph{chiral} solutions. A rigorous PDE analysis of this equation with complete integrability was recently initiated by Gérard and Lenzmann. Our main result constructs smooth, chiral, and finite energy finite-time blow-up solutions with mass arbitrarily close to that of a soliton, answering the global regularity question for chiral solutions raised by Gérard and Lenzmann. The blow-up rate obtained for these solutions is different from the pseudo-conformal rate. Our proof also gives a construction of a codimension one set of smooth finite energy initial data (but without addressing chirality) leading to the same blow-up dynamics. Our blow-up construction in the Hardy space might also be contrasted with the global well-posedness of the derivative nonlinear Schrödinger equation (DNLS), which is another integrable $L^{2}$-critical Schrödinger equation. The overall scheme of our proof is the forward construction of blow-up dynamics with modulation analysis. We begin with developing a linear theory for the near-soliton dynamics. We discover a nontrivial conjugation identity, which unveils a surprising connection from the linearized (CM-DNLS) to the 1D free Schrödinger equation, which is a crucial ingredient for overcoming the difficulties from the nonlocal nonlinearity. Another principal challenge in this work, the slow decay of the soliton, is overcome by introducing a trick of decomposing solutions depending on topologies, which we believe is of independent interest.

math.AP

DiaTool-DPO: Multi-Turn Direct Preference Optimization for Tool-Augmented Large Language Models

Tool-Augmented Larage Language Models (TA-LLMs) have shown promise in real-world applications, but face challenges in handling incomplete queries and out-of-scope requests. While existing approaches rely mainly on Supervised Fine-Tuning with expert trajectories, we propose DiaTool-DPO, a novel method that enhances TA-LLM's dialogue capabilities through Direct Preference Optimization. We model TA-LLM interactions as a Markov Decision Process with 5 distinct dialogue states and categorize user queries into 3 types based on their state transition trajectories. We automatically construct paired trajectory datasets of correct and incorrect dialogue flows and introduce a specialized objective loss for dialogue control. Our comprehensive evaluation demonstrates that DiaTool-DPO approaches GPT-4o's performance (94.8% in information gathering, 91% in tool call rejection) with substantial improvements over baseline (44% and 9.6% respectively) while maintaining core functionality. Our approach opens new possibilities for developing TA-LLMs that can handle diverse real-world scenarios without requiring additional expert demonstrations or human labeling.

cs.CL

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts

Optimization-based jailbreaks typically adopt the Toxic-Continuation setting in large vision-language models (LVLMs), following the standard next-token prediction objective. In this setting, an adversarial image is optimized to make the model predict the next token of a toxic prompt. However, we find that the Toxic-Continuation paradigm is effective at continuing already-toxic inputs, but struggles to induce safety misalignment when explicit toxic signals are absent. We propose a new paradigm: Benign-to-Toxic (B2T) jailbreak. Unlike prior work, we optimize adversarial images to induce toxic outputs from benign conditioning. Since benign conditioning contains no safety violations, the image alone must break the model's safety mechanisms. Our method outperforms prior approaches, transfers in black-box settings, and complements text-based jailbreaks. These results reveal an underexplored vulnerability in multimodal alignment and introduce a fundamentally new direction for jailbreak approaches.

cs.CV

Towards Scalable Human-aligned Benchmark for Text-guided Image Editing

A variety of text-guided image editing models have been proposed recently. However, there is no widely-accepted standard evaluation method mainly due to the subjective nature of the task, letting researchers rely on manual user study. To address this, we introduce a novel Human-Aligned benchmark for Text-guided Image Editing (HATIE). Providing a large-scale benchmark set covering a wide range of editing tasks, it allows reliable evaluation, not limited to specific easy-to-evaluate cases. Also, HATIE provides a fully-automated and omnidirectional evaluation pipeline. Particularly, we combine multiple scores measuring various aspects of editing so as to align with human perception. We empirically verify that the evaluation of HATIE is indeed human-aligned in various aspects, and provide benchmark results on several state-of-the-art models to provide deeper insights on their performance.

cs.CV

Shared Disk KV Cache Management for Efficient Multi-Instance Inference in RAG-Powered LLMs

Recent large language models (LLMs) face increasing inference latency as input context length and model size continue to grow. In particular, the retrieval-augmented generation (RAG) technique, which enhances LLM responses by incorporating external knowledge, exacerbates this issue by significantly increasing the number of input tokens. This expansion in token length leads to a substantial rise in computational overhead, particularly during the prefill stage, resulting in prolonged time-to-first-token (TTFT). To address this issue, this paper proposes a method to reduce TTFT by leveraging a disk-based key-value (KV) cache to lessen the computational burden during the prefill stage. We also introduce a disk-based shared KV cache management system, called Shared RAG-DCache, for multi-instance LLM RAG service environments. This system, together with an optimal system configuration, improves both throughput and latency under given resource constraints. Shared RAG-DCache exploits the locality of documents related to user queries in RAG, as well as the queueing delay in LLM inference services. It proactively generates and stores disk KV caches for query-related documents and shares them across multiple LLM instances to enhance inference performance. In experiments on a single host equipped with 2 GPUs and 1 CPU, Shared RAG-DCache achieved a 15~71% increase in throughput and up to a 12~65% reduction in latency, depending on the resource configuration.

cs.AI

Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading

LLM inference is essential for applications like text summarization, translation, and data analysis, but the high cost of GPU instances from Cloud Service Providers (CSPs) like AWS is a major burden. This paper proposes InferSave, a cost-efficient VM selection framework for cloud based LLM inference. InferSave optimizes KV cache offloading based on Service Level Objectives (SLOs) and workload charac teristics, estimating GPU memory needs, and recommending cost-effective VM instances. Additionally, the Compute Time Calibration Function (CTCF) improves instance selection accuracy by adjusting for discrepancies between theoretical and actual GPU performance. Experiments on AWS GPU instances show that selecting lower-cost instances without KV cache offloading improves cost efficiency by up to 73.7% for online workloads, while KV cache offloading saves up to 20.19% for offline workloads.

cs.LG