SearcharxivSearch

arXiv subjects

Jiwon Kang

Publications and source records attributed to Jiwon Kang.

14 recordsLinked to original sources

Quantum mutual information statistics for detecting dependence-structure change points in time series

Detecting when the dependence between two components of a multivariate time series changes, while the marginals drift freely, requires a dependence-specific statistic. We take the inferential object to be a density operator -- the trace-normalised second moment of unit-norm random Fourier features of ranks -- rather than a probability distribution. Partial traces recover the marginal operators exactly, so von Neumann entropies yield a quantum mutual information (QMI) statistic computed from prefix sums of small matrices, without density estimation, matrix inversion, or a tuned parameter. We develop the inference it needs: a segment-separable cost that drives penalised optimal partitioning, its split gain a Holevo information; finite-sample exact calibration by joint pair permutation, a block-permutation form for serially dependent series, and an exact, provably consistent exchangeability diagnostic that selects between them. We also prove a weighted chi-square boundary law, at the segment length and not its square root, for the rank-based statistic exactly as computed. In 500 replicates QMI detects nonlinear, correlation-free dependence changes with more power than the Hilbert-Schmidt independence criterion, distance correlation, Spearman, and empirical-copula statistics on the same ranks, by at least 15 percentage points wherever any statistic detects the change. Its false-alarm rate stays near nominal under marginal drift, where the empirical-copula statistic reaches 0.87. On eight years of hourly Korean weather observations, a two-stage segment-and-certify procedure finds dependence-change candidates above chance (five of 27 at p $\le$ 0.05 against 1.4 expected); stage two certifies one as a pure coupling change and reclassifies eight as marginal-driven.

stat.ME

Transferability Between Understanding and Generation in Unified Multimodal Models

Unified Multimodal Models (UMMs) integrate image understanding and generation within a single architecture, yet how the two tasks interact remains understudied. We investigate $\boldsymbol{\mathsf{transferability}}$ in UMMs: whether training a capability on one task improves the same capability on the other without explicit supervision. Through controlled experiments, we empirically find that transferability depends on architecture-models with fully shared transformer backbone and a unified visual encoder exhibit consistent cross-task transfer, while loosely coupled designs show little or none. Leveraging this transferability, we propose a practical training strategy. The most straightforward way to improve a target generative capability (e.g., counting) is to fine-tune generation directly, but this can degrade visual quality due to distribution shift. Instead, we train the corresponding understanding task and let it transfer into generation, which improves capability-specific generative performance while minimizing distribution shift. We validate this across three capabilities-counting, spatial relation, and text recognition/generation-showing that cross-task transferability can be systematically exploited in UMMs.

cs.CV

Irrationality of rapidly converging series: a problem of Erd\H{o}s and Graham

Answering a question of Erd\H{o}s and Graham, we show that the double exponential growth condition $\limsup_{n\to\infty}a_n^{1/\phi^n}=\infty$ for a strictly increasing sequence of positive integers $\{a_n\}_{n=1}^\infty$ is sufficient for the series $\sum_{n=1}^\infty 1/(a_n a_{n+1})$ to have an irrational sum; here $\phi$ denotes the golden ratio. We also provide a positive generalization to $\sum_{n=1}^\infty 1/(a_n^{w_0}\cdots a_{n+d-1}^{w_{d-1}})$, and a negative result showing that some of its instances are essentially optimal. The original problem was autonomously solved by the AI agent \emph{Aletheia}, powered by Gemini Deep Think, while the remaining material is largely a product of human-AI interactions.

math.NT

Semi-Autonomous Mathematics Discovery with Gemini: A Case Study on the Erd\H{o}s Problems

We present a case study in semi-autonomous mathematics discovery, using Gemini to systematically evaluate 700 conjectures labeled 'Open' in Bloom's Erd\H{o}s Problems database. We employ a hybrid methodology: AI-driven natural language verification to narrow the search space, followed by human expert evaluation to gauge correctness and novelty. We address 13 problems that were marked 'Open' in the database: 5 through seemingly novel autonomous solutions, and 8 through identification of previous solutions in the existing literature. Our findings suggest that the 'Open' status of the problems was through obscurity rather than difficulty. We also identify and discuss issues arising in applying AI to math conjectures at scale, highlighting the difficulty of literature identification and the risk of ''subconscious plagiarism'' by AI. We reflect on the takeaways from AI-assisted efforts on the Erd\H{o}s Problems.

cs.AI

APPLE: Attribute-Preserving Pseudo-Labeling for Diffusion-Based Face Swapping

Face swapping aims to transfer the identity of a source face onto a target face while preserving target-specific attributes such as pose, expression, lighting, skin tone, and makeup. However, since real ground truth for face swapping is unavailable, achieving both accurate identity transfer and high-quality attribute preservation remains challenging. Recent diffusion-based approaches attempt to improve visual fidelity through conditional inpainting on masked target images, but the masked condition removes crucial appearance cues, resulting in plausible yet misaligned attributes. To address this limitation, we propose APPLE (Attribute-Preserving Pseudo-Labeling), a fully diffusion-based teacher-student framework for attribute-preserving face swapping. Our approach introduces a teacher design to produce pseudo-labels aligned with the target attributes through (1) a conditional deblurring formulation that improves the preservation of global attributes such as skin tone and illumination, and (2) an attribute-aware inversion scheme that further enhances fine-grained attribute preservation such as makeup. APPLE conditions the student on clean pseudo-labels rather than degraded masked inputs, enabling more faithful attribute preservation. As a result, APPLE achieves state-of-the-art performance in attribute preservation while maintaining competitive identity transferability.

cs.CV

Where and How to Perturb: On the Design of Perturbation Guidance in Diffusion and Flow Models

Recent guidance methods in diffusion models steer reverse sampling by perturbing the model to construct an implicit weak model and guide generation away from it. Among these approaches, attention perturbation has demonstrated strong empirical performance in unconditional scenarios where classifier-free guidance is not applicable. However, existing attention perturbation methods lack principled approaches for determining where perturbations should be applied, particularly in Diffusion Transformer (DiT) architectures where quality-relevant computations are distributed across layers. In this paper, we investigate the granularity of attention perturbations, ranging from the layer level down to individual attention heads, and discover that specific heads govern distinct visual concepts such as structure, style, and texture quality. Building on this insight, we propose "HeadHunter", a systematic framework for iteratively selecting attention heads that align with user-centric objectives, enabling fine-grained control over generation quality and visual attributes. In addition, we introduce SoftPAG, which linearly interpolates each selected head's attention map toward an identity matrix, providing a continuous knob to tune perturbation strength and suppress artifacts. Our approach not only mitigates the oversmoothing issues of existing layer-level perturbation but also enables targeted manipulation of specific visual styles through compositional head selection. We validate our method on modern large-scale DiT-based text-to-image models including Stable Diffusion 3 and FLUX.1, demonstrating superior performance in both general quality enhancement and style-specific guidance. Our work provides the first head-level analysis of attention perturbation in diffusion models, uncovering interpretable specialization within attention layers and enabling practical design of effective perturbation strategies.

cs.CV

Analog Quantum Simulation of Dirac Hamiltonians in Circuit QED Using Rabi Driven Qubits

Quantum simulators hold promise for solving many intractable problems. However, a major challenge in quantum simulation, and quantum computation in general, is to solve problems with limited physical hardware. Currently, this challenge is tackled by designing dedicated devices for specific models, thereby allowing to reduce control requirements and simplify the construction. Here, we suggest a new method for quantum simulation in circuit QED, that provides versatility in model design and complete control over its parameters with minimal hardware requirements. We show how these features manifest through examples of quantum simulation of Dirac dynamics, which is relevant to the study of both high-energy physics and 2D materials. We conclude by discussing the advantages and limitations of the proposed method.

quant-ph

Identity-preserving Distillation Sampling by Fixed-Point Iterator

Score distillation sampling (SDS) demonstrates a powerful capability for text-conditioned 2D image and 3D object generation by distilling the knowledge from learned score functions. However, SDS often suffers from blurriness caused by noisy gradients. When SDS meets the image editing, such degradations can be reduced by adjusting bias shifts using reference pairs, but the de-biasing techniques are still corrupted by erroneous gradients. To this end, we introduce Identity-preserving Distillation Sampling (IDS), which compensates for the gradient leading to undesired changes in the results. Based on the analysis that these errors come from the text-conditioned scores, a new regularization technique, called fixed-point iterative regularization (FPR), is proposed to modify the score itself, driving the preservation of the identity even including poses and structures. Thanks to a self-correction by FPR, the proposed method provides clear and unambiguous representations corresponding to the given prompts in image-to-image editing and editable neural radiance field (NeRF). The structural consistency between the source and the edited data is obviously maintained compared to other state-of-the-art methods.

cs.CV

A Noise is Worth Diffusion Guidance

Diffusion models excel in generating high-quality images. However, current diffusion models struggle to produce reliable images without guidance methods, such as classifier-free guidance (CFG). Are guidance methods truly necessary? Observing that noise obtained via diffusion inversion can reconstruct high-quality images without guidance, we focus on the initial noise of the denoising pipeline. By mapping Gaussian noise to `guidance-free noise', we uncover that small low-magnitude low-frequency components significantly enhance the denoising process, removing the need for guidance and thus improving both inference throughput and memory. Expanding on this, we propose \ours, a novel method that replaces guidance methods with a single refinement of the initial noise. This refined noise enables high-quality image generation without guidance, within the same diffusion pipeline. Our noise-refining model leverages efficient noise-space learning, achieving rapid convergence and strong performance with just 50K text-image pairs. We validate its effectiveness across diverse metrics and analyze how refined noise can eliminate the need for guidance. See our project page: https://cvlab-kaist.github.io/NoiseRefine/.

cs.CV

Relaxing Accurate Initialization Constraint for 3D Gaussian Splatting

3D Gaussian splatting (3DGS) has recently demonstrated impressive capabilities in real-time novel view synthesis and 3D reconstruction. However, 3DGS heavily depends on the accurate initialization derived from Structure-from-Motion (SfM) methods. When the quality of the initial point cloud deteriorates, such as in the presence of noise or when using randomly initialized point cloud, 3DGS often undergoes large performance drops. To address this limitation, we propose a novel optimization strategy dubbed RAIN-GS (Relaing Accurate Initialization Constraint for 3D Gaussian Splatting). Our approach is based on an in-depth analysis of the original 3DGS optimization scheme and the analysis of the SfM initialization in the frequency domain. Leveraging simple modifications based on our analyses, RAIN-GS successfully trains 3D Gaussians from sub-optimal point cloud (e.g., randomly initialized point cloud), effectively relaxing the need for accurate initialization. We demonstrate the efficacy of our strategy through quantitative and qualitative comparisons on multiple datasets, where RAIN-GS trained with random point cloud achieves performance on-par with or even better than 3DGS trained with accurate SfM point cloud. Our project page and code can be found at https://ku-cvlab.github.io/RAIN-GS.

cs.CV

Self-Evolving Neural Radiance Fields

Recently, neural radiance field (NeRF) has shown remarkable performance in novel view synthesis and 3D reconstruction. However, it still requires abundant high-quality images, limiting its applicability in real-world scenarios. To overcome this limitation, recent works have focused on training NeRF only with sparse viewpoints by giving additional regularizations, often called few-shot NeRF. We observe that due to the under-constrained nature of the task, solely using additional regularization is not enough to prevent the model from overfitting to sparse viewpoints. In this paper, we propose a novel framework, dubbed Self-Evolving Neural Radiance Fields (SE-NeRF), that applies a self-training framework to NeRF to address these problems. We formulate few-shot NeRF into a teacher-student framework to guide the network to learn a more robust representation of the scene by training the student with additional pseudo labels generated from the teacher. By distilling ray-level pseudo labels using distinct distillation schemes for reliable and unreliable rays obtained with our novel reliability estimation method, we enable NeRF to learn a more accurate and robust geometry of the 3D scene. We show and evaluate that applying our self-training framework to existing models improves the quality of the rendered images and achieves state-of-the-art performance in multiple settings.

cs.CV

Test for parameter change in the presence of outliers: the density power divergence based approach

This study considers the problem of testing for a parameter change in the presence of outliers. For this, we propose a robust test using the objective function of minimum density power divergence estimator (MDPDE) by Basu et al. (Biometrika, 1998), and then derive its limiting null distribution. Our test procedure can be naturally extended to any parametric model to which MDPDE can be applied. To illustrate this, we apply our test procedure to GARCH models. We demonstrate the validity and robustness of the proposed test through a simulation study. In a real data application to the Hang Seng index, our test locates some change-points that are not detected by the previous tests such as the score test and the residual-based CUSUM test.

math.ST

A robust approach for testing parameter change in Poisson autoregressive models

Parameter change test has been an important issue in time series analysis. The problem has also been actively explored in the field of integer-valued time series, but the testing in the presence of outliers has not yet been extensively investigated. This study considers the problem of testing for parameter change in Poisson autoregressive models particularly when observations are contaminated by outliers. To lessen the impact of outliers on testing procedure, we propose a test based on the density power divergence, which is introduced by Basu et al. (Biometrika, 1998), and derive its limiting null distribution. Monte Carlo simulation results demonstrate validity and strong robustness of the proposed test.

stat.ME

Robust Estimation in Stochastic Frontier Models

This study proposes a robust estimator for stochastic frontier models by integrating the idea of Basu et al. [1998, Biometrika 85, 549-559] into such models. We verify that the suggested estimator is strongly consistent and asymptotic normal under regularity conditions and investigate robust properties. We use a simulation study to demonstrate that the estimator has strong robust properties with little loss in asymptotic efficiency relative to the maximum likelihood estimator. A real data analysis is performed for illustrating the use of the estimator.

stat.ME