SearcharxivSearch

arXiv subjects

Jeongho Kim

Publications and source records attributed to Jeongho Kim.

At least 19 recordsLinked to original sources

MoNe: Modular Neural Memory for Efficient Long Context Inference

We present MoNe, a lightweight modular neural memory that attaches to any frozen pretrained Transformer to enable long-context inference without retraining. MoNe reads context in fixed-size segments via test-time learning of fast-weight neural memory networks with layer-localized gradient updates; at inference, the memory generates keys and values from the query tokens alone, with no context tokens re-read. This two-phase design decouples inference cost from context length, achieving $O(N)$ preprocessing and $O(1)$ query cost with peak GPU memory that does not grow with $N$. At 128K tokens, MoNe reduces both compute and peak GPU memory by approximately 80% compared to ICL with only 6.4% parameter overhead. MoNe generalizes to context lengths far beyond the backbone's native window, achieving strong performance on needle-in-a-haystack and word extraction benchmarks from RULER, where ICL degrades sharply.

cs.AI

Time-Asymptotic Stability of the Stationary Solution to the Impermeable Wall Problem for the radially Symmetric Navier-Stokes-Korteweg Equations

We study the asymptotic behavior of the initial-boundary value problem for the radially symmetric Navier--Stokes--Korteweg (NSK) equations defined on the exterior domain $\Omega = \{x\in\R^n~|~|x|> 1\}$. In particular, we consider the impermeable wall problem, where the velocity at the boundary $\{x\in\R^n~|~|x|=1\}$ is set to be zero. We show that, if the initial data is a small perturbation of the stationary solution, and the boundary data are sufficiently small, then there exists a global-in-time strong solution to the radially symmetric NSK equations, and it converges to the stationary solution time-asymptotically. Our method is based on elementary energy estimates with a combination of carefully designed energy functionals.

math.AP

Time-asymptotic stability of viscous shocks for the outflow problem of one-dimensional compressible fluids of Korteweg type

We study the time-asymptotic stability of viscous-dispersive shock waves for the outflow problem of the barotropic Navier--Stokes--Korteweg equations, which describe viscous fluids with internal capillarity. Assuming that the far-field state is subsonic or transonic and that the velocity at the boundary is larger than the far-field velocity, we prove that the solution converges to the corresponding viscous-dispersive shock wave as $t \to +\infty$, provided that the shock amplitude and the initial perturbation are sufficiently small. The proof is based on the method of $a$-contraction with shifts (for viscous equations) introduced in \cite{KV17,KV21,KVW23}. A main difficulty comes from controlling the boundary effect of the viscous-dispersive shock wave, as well as the influence of capillarity near the boundary.

math.AP

Stationary solutions to the spherically symmetric compressible fluid with capillarity effect

We consider the spherically symmetric Navier--Stokes--Korteweg (NSK) system on the exterior domain $\Omega=\{x\in\mathbb{R}^n~|~|x|>1\}$ with $n\ge2$ when the boundary and far-field data are given. We show that, if the boundary data are sufficiently small, then there exists a unique smooth stationary solution to the spherically symmetric NSK system with impermeable wall, inflow, and outflow boundary conditions. We also establish the decay rate of the stationary solutions. Precisely, the stationary solution for the impermeable wall problem exponentially decays to the far-field states, while that of the inflow/outflow problem algebraically decays. Finally, we investigate the asymptotic convergences of the stationary solution for the impermeable wall problem as the capillarity coefficient vanishes. Numerical results validate that our theoretical convergence rate of the stationary solution is optimal.

math.AP

Quantitative Hydrodynamic Limit of the Chern--Simons--Higgs System

We study the hydrodynamic limit of the Chern--Simons--Higgs system, a relativistic gauge field model involving the Chern--Simons interaction. We introduce a single scaling parameter capturing both the non-relativistic (infinite speed of light) and semi-classical (vanishing Planck constant) regimes. This unified scaling allows us to justify the simultaneous non-relativistic and semi-classical limit, while retaining the nontrivial influence of the Chern--Simons gauge structure. Using a modulated energy method, we establish quantitative convergence rates toward the corresponding compressible Euler--Chern--Simons system as the scaling parameter tends to zero.

math.AP

Memory-Efficient Fine-Tuning Diffusion Transformers via Dynamic Patch Sampling and Block Skipping

Diffusion Transformers (DiTs) have significantly enhanced text-to-image (T2I) generation quality, enabling high-quality personalized content creation. However, fine-tuning these models requires substantial computational complexity and memory, limiting practical deployment under resource constraints. To tackle these challenges, we propose a memory-efficient fine-tuning framework called DiT-BlockSkip, integrating timestep-aware dynamic patch sampling and block skipping by precomputing residual features. Our dynamic patch sampling strategy adjusts patch sizes based on the diffusion timestep, then resizes the cropped patches to a fixed lower resolution. This approach reduces forward & backward memory usage while allowing the model to capture global structures at higher timesteps and fine-grained details at lower timesteps. The block skipping mechanism selectively fine-tunes essential transformer blocks and precomputes residual features for the skipped blocks, significantly reducing training memory. To identify vital blocks for personalization, we introduce a block selection strategy based on cross-attention masking. Evaluations demonstrate that our approach achieves competitive personalization performance qualitatively and quantitatively, while reducing memory usage substantially, moving toward on-device feasibility (e.g., smartphones, IoT devices) for large-scale diffusion transformers.

cs.CV

Contraction of viscous-dispersive shocks: Zero viscosity-capillarity limits

We prove the contraction property of any large solution perturbed from a viscous-dispersive shock wave of the Navier--Stokes--Korteweg (NSK) system. The contraction holds up to a dynamical shift, since the contraction is measured by the relative entropy that is locally $L^2$. We use the contraction property to show the global existence of large solution perturbed from a viscous-dispersive shock wave. To prove the contraction property, we first employ the effective velocity to transform the NSK system into the system of two degenerate parabolic equations, then apply the method of $a$-contraction with shifts. The contraction property does not depend on the strengths of viscosity and capillarity. Based on this uniformity, we show the existence of zero viscosity-capillarity limits of solutions to the NSK system, on which Riemann shocks are unique and stable up to shifts.

math.AP

InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion

Recent advances in diffusion models have enabled impressive video editing capabilities, yet production-grade Video Object Insertion (VOI) remains challenging due to inadequate 4D scene understanding and a lack of proper optical interactions, such as shadows and reflections. To address these limitations, we present InsertAnywhere, a comprehensive VOI framework that achieves geometrically grounded object placement and optics-aware video synthesis. Our approach first leverages a 4D-aware mask generation module that allows users to anchor an object's 3D pose in a single frame. The framework automatically propagates this placement across the video, accurately handling local scene dynamics and occlusions. To synthesize realistic physical lighting interactions, we introduce Optics-Aware Representation Alignment, a novel strategy that utilizes an extended mask to guide feature extraction, enabling optical effects to seamlessly extend beyond the inserted object's boundary. Finally, to overcome the lack of training data for such phenomena, we construct and open-source ROSE++, a specialized quadruplet dataset tailored for the supervised learning of optical effects. Extensive experiments demonstrate that InsertAnywhere produces geometrically plausible and photometrically realistic insertions in complex real-world scenarios, significantly outperforming existing research and commercial generative tools.

cs.CV

Infinite-Homography as Robust Conditioning for Camera-Controlled Video Generation

Recent progress in video diffusion models has spurred growing interest in camera-controlled novel-view video generation for dynamic scenes, aiming to provide creators with cinematic camera control capabilities in post-production. A key challenge in camera-controlled video generation is ensuring fidelity to the specified camera pose, while maintaining view consistency and reasoning about occluded geometry from limited observations. To address this, existing methods either train trajectory-conditioned video generation model on trajectory-video pair dataset, or estimate depth from the input video to reproject it along a target trajectory and generate the unprojected regions. Nevertheless, existing methods struggle to generate camera-pose-faithful, high-quality videos for two main reasons: (1) reprojection-based approaches are highly susceptible to errors caused by inaccurate depth estimation; and (2) the limited diversity of camera trajectories in existing datasets restricts learned models. To address these limitations, we present InfCam, a depth-free, camera-controlled video-to-video generation framework with high pose fidelity. The framework integrates two key components: (1) infinite homography warping, which encodes 3D camera rotations directly within the 2D latent space of a video diffusion model. Conditioning on this noise-free rotational information, the residual parallax term is predicted through end-to-end training to achieve high camera-pose fidelity; and (2) a data augmentation pipeline that transforms existing synthetic multiview datasets into sequences with diverse trajectories and focal lengths. Experimental results demonstrate that InfCam outperforms baseline methods in camera-pose accuracy and visual fidelity, generalizing well from synthetic to real-world data. Link to our project page:https://emjay73.github.io/InfCam/

cs.CV

Local well-posedness and asymptotic analysis of a nonlocal incompressible Navier--Stokes--Korteweg system

We consider a relaxed formulation of the inhomogeneous incompressible Navier--Stokes--Korteweg system, where the classical third-order capillarity term is replaced by a nonlocal approximation. We first establish the local-in-time well-posedness of the relaxed system, under standard regularity and positivity assumptions on the initial data. The existence time is uniform with respect to both the capillarity coefficient and the relaxation parameter. We then study two asymptotic limits of the system: the nonlocal-to-local limit as the relaxation parameter tends to infinity, and the vanishing capillarity limit. In each case, we prove convergence of the solution to that of the corresponding target system. Our analysis provides a rigorous justification for the use of nonlocal relaxation models in approximating capillarity-driven incompressible fluid flows.

math.AP

Multi-objective CFD optimization of an intermediate diffuser stage for PediaFlow pediatric ventricular assist device

Background: Computational fluid dynamics (CFD) has become an essential design tool for ventricular assist devices (VADs), where the goal of maximizing performance often conflicts with biocompatibility. This tradeoff becomes even more pronounced in pediatric applications due to the stringent size constraints imposed by the smaller patient population. This study presents an automated CFD-driven shape optimization of a new intermediate diffuser stage for the PediaFlow pediatric VAD, positioned immediately downstream of the impeller to improve pressure recovery. Methods: We adopted a multi-objective optimization approach to maximize pressure recovery while minimizing hemolysis. The proposed diffuser stage was isolated from the rest of the flow domain, enabling efficient evaluation of over 450 design variants using Sobol sequence, which yielded a Pareto front of non-dominated solutions. The selected best candidate was further refined using local T-search algorithm. We then incorporated the optimized front diffuser into the full pump for CFD verification and in vitro validation. Results: We identified critical dependencies where longer blades increased pressure recovery but also hemolysis, while the wrap angle showed a strong parabolic relationship with pressure recovery but a monotonic relationship with hemolysis. Counterintuitively, configurations with fewer blades (2-3) consistently outperformed those with more blades (4-5) in both metrics. The optimized two-blade design enabled operation at lower pump speeds (14,000 vs 16,000 RPM), improving hydraulic efficiency from 26.3% to 32.5% and reducing hemolysis by 31%. Conclusion: This approach demonstrates that multi-objective CFD optimization can systematically explore complex design spaces while balancing competing priorities of performance and hemocompatibility for pediatric VADs.

physics.med-ph

Memory-Efficient Personalization of Text-to-Image Diffusion Models via Selective Optimization Strategies

Memory-efficient personalization is critical for adapting text-to-image diffusion models while preserving user privacy and operating within the limited computational resources of edge devices. To this end, we propose a selective optimization framework that adaptively chooses between backpropagation on low-resolution images (BP-low) and zeroth-order optimization on high-resolution images (ZO-high), guided by the characteristics of the diffusion process. As observed in our experiments, BP-low efficiently adapts the model to target-specific features, but suffers from structural distortions due to resolution mismatch. Conversely, ZO-high refines high-resolution details with minimal memory overhead but faces slow convergence when applied without prior adaptation. By complementing both methods, our framework leverages BP-low for effective personalization while using ZO-high to maintain structural consistency, achieving memory-efficient and high-quality fine-tuning. To maximize the efficacy of both BP-low and ZO-high, we introduce a timestep-aware probabilistic function that dynamically selects the appropriate optimization strategy based on diffusion timesteps. This function mitigates the overfitting from BP-low at high timesteps, where structural information is critical, while ensuring ZO-high is applied more effectively as training progresses. Experimental results demonstrate that our method achieves competitive performance while significantly reducing memory consumption, enabling scalable, high-quality on-device personalization without increasing inference latency.

cs.CV

From Wardrobe to Canvas: Wardrobe Polyptych LoRA for Part-level Controllable Human Image Generation

Recent diffusion models achieve personalization by learning specific subjects, allowing learned attributes to be integrated into generated images. However, personalized human image generation remains challenging due to the need for precise and consistent attribute preservation (e.g., identity, clothing details). Existing subject-driven image generation methods often require either (1) inference-time fine-tuning with few images for each new subject or (2) large-scale dataset training for generalization. Both approaches are computationally expensive and impractical for real-time applications. To address these limitations, we present Wardrobe Polyptych LoRA, a novel part-level controllable model for personalized human image generation. By training only LoRA layers, our method removes the computational burden at inference while ensuring high-fidelity synthesis of unseen subjects. Our key idea is to condition the generation on the subject's wardrobe and leverage spatial references to reduce information loss, thereby improving fidelity and consistency. Additionally, we introduce a selective subject region loss, which encourages the model to disregard some of reference images during training. Our loss ensures that generated images better align with text prompts while maintaining subject integrity. Notably, our Wardrobe Polyptych LoRA requires no additional parameters at the inference stage and performs generation using a single model trained on a few training samples. We construct a new dataset and benchmark tailored for personalized human image generation. Extensive experiments show that our approach significantly outperforms existing techniques in fidelity and consistency, enabling realistic and identity-preserving full-body synthesis.

cs.CV

MultiHuman-Testbench: Benchmarking Image Generation for Multiple Humans

Generation of images containing multiple humans, performing complex actions, while preserving their facial identities, is a significant challenge. A major factor contributing to this is the lack of a dedicated benchmark. To address this, we introduce MultiHuman-Testbench, a novel benchmark for rigorously evaluating generative models for multi-human generation. The benchmark comprises 1,800 samples, including carefully curated text prompts, describing a range of simple to complex human actions. These prompts are matched with a total of 5,550 unique human face images, sampled uniformly to ensure diversity across age, ethnic background, and gender. Alongside captions, we provide human-selected pose conditioning images which accurately match the prompt. We propose a multi-faceted evaluation suite employing four key metrics to quantify face count, ID similarity, prompt alignment, and action detection. We conduct a thorough evaluation of a diverse set of models, including zero-shot approaches and training-based methods, with and without regional priors. We also propose novel techniques to incorporate image and region isolation using human segmentation and Hungarian matching, significantly improving ID similarity. Our proposed benchmark and key findings provide valuable insights and a standardized tool for advancing research in multi-human image generation. The dataset and evaluation codes will be available at https://github.com/Qualcomm-AI-research/MultiHuman-Testbench.

cs.CV

ReCDAP: Relation-Based Conditional Diffusion with Attention Pooling for Few-Shot Knowledge Graph Completion

Knowledge Graphs (KGs), composed of triples in the form of (head, relation, tail) and consisting of entities and relations, play a key role in information retrieval systems such as question answering, entity search, and recommendation. In real-world KGs, although many entities exist, the relations exhibit a long-tail distribution, which can hinder information retrieval performance. Previous few-shot knowledge graph completion studies focused exclusively on the positive triple information that exists in the graph or, when negative triples were incorporated, used them merely as a signal to indicate incorrect triples. To overcome this limitation, we propose Relation-Based Conditional Diffusion with Attention Pooling (ReCDAP). First, negative triples are generated by randomly replacing the tail entity in the support set. By conditionally incorporating positive information in the KG and non-existent negative information into the diffusion process, the model separately estimates the latent distributions for positive and negative relations. Moreover, including an attention pooler enables the model to leverage the differences between positive and negative cases explicitly. Experiments on two widely used datasets demonstrate that our method outperforms existing approaches, achieving state-of-the-art performance. The code is available at https://github.com/hou27/ReCDAP-FKGC.

cs.AI

Derivation of nonlinear aggregation-diffusion equation from a kinetic BGK-type equation

This paper investigates the diffusion limit of a kinetic BGK-type equation, focusing on its relaxation to a nonlinear aggregation-diffusion equation, where the diffusion exhibits a porous-medium-type nonlinearity. Unlike previous studies by Dolbeault et al. [Arch. Ration. Mech. Anal., 186, (2007), 133-158] and Addala and Tayeb [J. Hyperbolic Differ. Equ., 16, (2019), 131-156], which required bounded initial data, our work considers initial data that need not be bounded. We develop new techniques for handling weak entropy solutions that satisfy the natural bounds associated with the kinetic entropy inequality. Our proof employs the relative entropy method and various compactness arguments to establish the convergence and properties of these solutions.

math.AP

Convergence to Superposition of Boundary Layer, Rarefaction and Shock for the Inflow Problem of the 1D Navier--Stokes Equations

We establish the asymptotic stability of solutions to the inflow problem for the one-dimensional barotropic Navier--Stokes equations in half space. When the boundary value is located at the subsonic regime, all the possible thirteen asymptotic patterns are classified in \cite{M01}. We consider the most complicated pattern, the superposition of the boundary layer solution, the 1-rarefaction wave, and the viscous 2-shock waves. In this superposition, the boundary layer is degenerate and large. We prove that, if the strengths of the rarefaction wave and shock wave are small, and if the initial data is a small perturbation of the superposition, then the solution asymptotically converges to the superposition up to a dynamical shift for the shock. As a corollary, our result implies the asymptotic stability for the simpler case where the superposition consists of the degenerate boundary layer solution and the viscous 2-shock. Therefore, we complete the study of the asymptotic stability of the inflow problem for the 1D barotropic Navier--Stokes equations for subsonic boundary values.

math.AP

Time-asymptotic stability of composite wave for the one-dimensional compressible fluid of Kortwewg type

We study the asymptotic stability of a composition of rarefaction and shock waves for the one-dimensional barotropic compressible fluid of Korteweg type, called the Navier-Stokes-Korteweg(NSK) system. Precisely, we show that the solution to the NSK system asymptotically converges to the composition of the rarefaction wave and shifted viscous-dispersive shock wave, under certain smallness assumption on the initial perturbation and strength of the waves. Our method is based on the method of $a$-contraction with shift developed by Kang and Vasseur \cite{KV16}, successfully applied to obtain contraction or stability of nonlinear waves for hyperbolic systems.

math.AP