SearcharxivSearch

arXiv subjects

Truong Vu

Publications and source records attributed to Truong Vu.

14 recordsLinked to original sources

Modified Scattering and Asymptotics for Perturbed One-Dimensional Cubic NLS

We study the long-time dynamics of small solutions to the one-dimensional nonlinear Schr\"odinger equation \[ i\partial_t v+\partial_x^2v-\beta\abs{v}^2v +\cW(x)\abs{v}^4v+\gamma i\partial_xv=0, \] where $\cW$ is spatially localized. The cubic nonlinearity is long range and produces the logarithmic phase correction, whereas the localized quintic term is short range at leading order. We construct a global forward modified wave operator for small complex asymptotic profiles and prove quantitative final-state estimates. For small data in the weighted energy space, we also establish global existence, sharp $t^{-1/2}$ decay, and forward modified scattering with a unique asymptotic profile. The principal new phenomenon occurs beyond this leading law. The exact Duhamel tail generated by the localized quintic term admits a quantitative inner scaling limit at the distinguished frequency $\zeta=-\gamma/2$ on the scale $\abs{\zeta+\gamma/2}\sim t^{-1/2}$. Its universal shape is explicit and depends on the value of the scattering profile on the distinguished ray and on the zeroth moment of $\cW$. When both quantities are nonzero, the limit is nontrivial, belongs optimally to $C^{2,1}_{\mathrm{loc}}$, and is not $C^3$ at the center. Away from the corresponding self-similar ray $\xi=-\gamma$, we construct rigorously defined higher-order outer expansions through every integer order allowed by the decay of $\cW$, and to each fixed finite order when $\cW$ is rapidly decreasing.

math.AP

Stability of Finite-Batch Particle Mean-Field Variational Inference Beyond Strong Convexity

We study the implementable finite-batch particle algorithm for mean-field variational inference as a fully discrete stochastic approximation of the projected Wasserstein dynamics. The target potential is globally smooth but need not be strongly convex. The departure from contractivity is quantified by the curvature defect \[ \mathfrak d_\alpha(x,y) = \bigl[\alpha\|x-y\|^2- \langle\nabla V(x)-\nabla V(y),x-y\rangle\bigr]_+, \] which is the additive loss in the one-step Euler contraction estimate. We prove a non-asymptotic Wasserstein stability bound that separates initialization, product-empirical approximation, finite-batch drift error, time discretization, and the defects accumulated along the coupled trajectories. Under the uniform bound $\mathfrak d_\alpha\leq\beta$, the particle iterates remain within $O(\sqrt{\beta/\alpha})$ of any MFVI minimizer, up to explicit errors in the particle number, batch size, and step size. The proof uses a stationary comparison array whose population law is an MFVI minimizer but whose particle-level law is a random product empirical measure, and it controls the resulting projected-drift discrepancy explicitly. We also give coordinatewise defect estimates and structural conditions for dimension-independent projected-drift sensitivity, construct an arbitrary-dimensional smooth nonconvex benchmark with a closed-form MFVI minimizer, and explain why polynomially growing drifts require a modification of the untamed explicit scheme.

math.NA

Root Dynamics of Differentiated Polynomials with Rotationally Invariant Structure

The dynamics of polynomial roots under repeated differentiation has recently been conjectured to converge to a limiting measure governed by specific nonlinear PDEs, the conjectures being shown in some particular settings. For rotationally invariant initial distributions, a deterministic structured sampling model placing roots on concentric circles was recently introduced by Galligo, Najnudel, and Vu. In this paper, the authors proved convergence under the technical growth condition $m_n / (n \log n) \to \infty$, where $n$ is the number of circles and $m_n$ is the number of points per circle. In this paper, we significantly improve this result by relaxing the growth condition to $m_n / \log n \to \infty$, thus allowing for regimes where the number of points per circle grows proportionally to the number of circles. The key innovation is a refined upper bound on the root magnitudes after differentiation. This sharper estimate prevents the rapid accumulation of errors over multiple differentiations, fully validating a recent conjecture regarding the robustness of the sampling scheme.

math.PR

Attention-Based Prototype Calibration for Multi-Rater Few-Shot Medical Image Segmentation

Few-shot medical image segmentation methods typically assume a single ground-truth annotation, overlooking systematic variability across expert raters commonly observed in clinical datasets. We propose an attention-based prototype calibration framework for few-shot multi-rater segmentation that models rater-specific deviations from a consensus representation in prototype space. A lightweight yet principled attention operator directly refines rater prototypes without modifying the backbone feature extractor, making the approach fully compatible with existing prototype-based few-shot segmentation methods. This design preserves semantic consistency while enabling personalized segmentation outputs with minimal computational overhead. Experiments on multi-rater medical imaging datasets demonstrate consistent improvements over baseline prototype approaches, highlighting the effectiveness of structured prototype calibration for modeling annotation variability.

cs.CV

Gradient Flows of Interfacial Energies: Curvature Agents and Incompressibility

We present a framework for the gradient flow of sharp-interface surface energies that couple to embedded curvature active agents. We use a penalty method to develop families of locally incompressible gradient flows that couple interface stretching or compression to local flux of interfacial mass. We establish the convergence of the penalty method to an incompressible flow both formally for a broad family of surface energies and rigorously for a more narrow class of surface energies. We present an analysis, including a $Γ$-limit, of an Allen-Cahn type model for a coupled surface agent curvature energy.

math.AP

Dynamics of rotationally invariant polynomial root sets under iterated differentiations

We associate to an $N$-sample of a given rotationally invariant probability measure $μ_0$ with compact support in the complex plane, a polynomial $P_N$ with roots given by the sample. Then, for $t \in (0,1)$, we consider the empirical measure $μ_t^{N}$ associated to the root set of the $\lfloor t N\rfloor$-th derivative of $P_N$. A question posed by O'Rourke and Steinerberger [21], reformulated as a conjecture by Hoskins and Kabluchko [10], and recently reaffirmed by Campbell, O'Rourke and Renfrew [5], states that under suitable conditions of regularity on $μ_0$, for an i.i.d. sample, $μ_t^{N}$ converges to a rotationally invariant probability measure $μ_t$ when $N$ tends to infinity, and that $(1-t)μ_t$ has a radial density $x \mapsto ψ(x,t)$ satisfying the following partial differential equation: \begin{equation} \label{PDErotational} \frac{ \partial ψ(x,t) }{\partial t} = \frac{ \partial}{\partial x} \left( \frac{ ψ(x,t) }{ \frac{1}{x} \int_0^x ψ(y,t) dy } \right). \end{equation} In [10], this equation is reformulated as an equation on the distribution function $Ψ_t$ of the radial part of $(1-t) μ_t$: \begin{equation} \label{equationPsixtabstract} \frac{\partial Ψ_t (x)}{\partial t} = x \frac{\frac{\partial Ψ_t (x)}{\partial x} } {Ψ_t(x)} - 1. \end{equation} Restricting our study to a specific family of $N$-samplings, we are able to prove a variant of the conjecture above. We also emphasize the important differences between the two-dimensional setting and the one-dimensional setting, illustrated in our Theorem 2.1.

math.PR

The Fourier coefficients of the holomorphic multiplicative chaos in the limit of large frequency

The holomorphic multiplicative chaos (HMC) is a holomorphic analogue of the Gaussian multiplicative chaos. It arises naturally as the limit in large matrix size of the characteristic polynomial of Haar unitary, and more generally circular-$β$-ensemble, random matrices. We consider the Fourier coefficients of the holomorphic multiplicative chaos in the $L^1$-phase, and we show that appropriately normalized, this converges in distribution to a complex normal random variable, scaled by the total mass of the Gaussian multiplicative chaos measure on the unit circle. We further generalize this to a process convergence, showing the joint convergence of consecutive Fourier coefficients. As a corollary, we derive convergence in law of the secular coefficients of sublinear index of the circular-$β$-ensemble for all $β> 2$.

math.PR

EditScout: Locating Forged Regions from Diffusion-based Edited Images with Multimodal LLM

Image editing technologies are tools used to transform, adjust, remove, or otherwise alter images. Recent research has significantly improved the capabilities of image editing tools, enabling the creation of photorealistic and semantically informed forged regions that are nearly indistinguishable from authentic imagery, presenting new challenges in digital forensics and media credibility. While current image forensic techniques are adept at localizing forged regions produced by traditional image manipulation methods, current capabilities struggle to localize regions created by diffusion-based techniques. To bridge this gap, we present a novel framework that integrates a multimodal Large Language Model (LLM) for enhanced reasoning capabilities to localize tampered regions in images produced by diffusion model-based editing methods. By leveraging the contextual and semantic strengths of LLMs, our framework achieves promising results on MagicBrush, AutoSplice, and PerfBrush (novel diffusion-based dataset) datasets, outperforming previous approaches in mIoU and F1-score metrics. Notably, our method excels on the PerfBrush dataset, a self-constructed test set featuring previously unseen types of edits. Here, where traditional methods typically falter, achieving markedly low scores, our approach demonstrates promising performance.

cs.CV

Stable Messenger: Steganography for Message-Concealed Image Generation

In the ever-expanding digital landscape, safeguarding sensitive information remains paramount. This paper delves deep into digital protection, specifically focusing on steganography. While prior research predominantly fixated on individual bit decoding, we address this limitation by introducing ``message accuracy'', a novel metric evaluating the entirety of decoded messages for a more holistic evaluation. In addition, we propose an adaptive universal loss tailored to enhance message accuracy, named Log-Sum-Exponential (LSE) loss, thereby significantly improving the message accuracy of recent approaches. Furthermore, we also introduce a new latent-aware encoding technique in our framework named \Approach, harnessing pretrained Stable Diffusion for advanced steganographic image generation, giving rise to a better trade-off between image quality and message recovery. Throughout experimental results, we have demonstrated the superior performance of the new LSE loss and latent-aware encoding technique. This comprehensive approach marks a significant step in evolving evaluation metrics, refining loss functions, and innovating image concealment techniques, aiming for more robust and dependable information protection.

cs.CV

Anti-concentration applied to roots of randomized derivatives of polynomials

Let $(Z^{(n)}_k)_{1 \leq k \leq n}$ be a random set of points and let $μ_n$ be its \emph{empirical measure}: $$μ_n = \frac{1}{n} \sum_{k=1}^n δ_{Z^{(n)}_k}. $$ Let $$P_n(z) := (z - Z^{(n)}_1)\cdots (z - Z^{(n)}_n)\quad \text{and}\quad Q_n (z) := \sum_{k=1}^n γ^{(n)}_k \prod_{1 \leq j \leq n, j \neq k} (z- Z^{(n)}_j), $$ where $(γ^{(n)}_k)_{1 \leq k \leq n}$ are independent, i.i.d. random variables with Gamma distribution of parameter $β/2$, for some fixed $β> 0$. We prove that in the case where $μ_n$ almost surely tends to $μ$ when $n \rightarrow \infty$, the empirical measure of the complex zeros of the \emph{randomized derivative} $Q_n$ also converges almost surely to $μ$ when $n$ tends to infinity. Furthermore, for $k = o(n / \log n)$, we obtain that the zeros of the $k-$th \emph{randomized derivative} of $P_n$ converge to the limiting measure $μ$ in the same sense. We also derive the same conclusion for a variant of the randomized derivative related to the unit circle.

math.PR

LP-OVOD: Open-Vocabulary Object Detection by Linear Probing

This paper addresses the challenging problem of open-vocabulary object detection (OVOD) where an object detector must identify both seen and unseen classes in test images without labeled examples of the unseen classes in training. A typical approach for OVOD is to use joint text-image embeddings of CLIP to assign box proposals to their closest text label. However, this method has a critical issue: many low-quality boxes, such as over- and under-covered-object boxes, have the same similarity score as high-quality boxes since CLIP is not trained on exact object location information. To address this issue, we propose a novel method, LP-OVOD, that discards low-quality boxes by training a sigmoid linear classifier on pseudo labels retrieved from the top relevant region proposals to the novel text. Experimental results on COCO affirm the superior performance of our approach over the state of the art, achieving $\textbf{40.5}$ in $\text{AP}_{novel}$ using ResNet50 as the backbone and without external datasets or knowing novel classes during training. Our code will be available at https://github.com/VinAIResearch/LP-OVOD.

cs.CV

Dataset Diffusion: Diffusion-based Synthetic Dataset Generation for Pixel-Level Semantic Segmentation

Preparing training data for deep vision models is a labor-intensive task. To address this, generative models have emerged as an effective solution for generating synthetic data. While current generative models produce image-level category labels, we propose a novel method for generating pixel-level semantic segmentation labels using the text-to-image generative model Stable Diffusion (SD). By utilizing the text prompts, cross-attention, and self-attention of SD, we introduce three new techniques: class-prompt appending, class-prompt cross-attention, and self-attention exponentiation. These techniques enable us to generate segmentation maps corresponding to synthetic images. These maps serve as pseudo-labels for training semantic segmenters, eliminating the need for labor-intensive pixel-wise annotation. To account for the imperfections in our pseudo-labels, we incorporate uncertainty regions into the segmentation, allowing us to disregard loss from those regions. We conduct evaluations on two datasets, PASCAL VOC and MSCOCO, and our approach significantly outperforms concurrent work. Our benchmarks and code will be released at https://github.com/VinAIResearch/Dataset-Diffusion

cs.CV

Face Swapping as A Simple Arithmetic Operation

We propose a novel high-fidelity face swapping method called "Arithmetic Face Swapping" (AFS) that explicitly disentangles the intermediate latent space W+ of a pretrained StyleGAN into the "identity" and "style" subspaces so that a latent code in W+ is the sum of an "identity" code and a "style" code in the corresponding subspaces. Via our disentanglement, face swapping (FS) can be regarded as a simple arithmetic operation in W+, i.e., the summation of a source "identity" code and a target "style" code. This makes AFS more intuitive and elegant than other FS methods. In addition, our method can generalize over the standard face swapping to support other interesting operations, e.g., combining the identity of one source with styles of multiple targets and vice versa. We implement our identity-style disentanglement by learning a neural network that maps a latent code to a "style" code. We provide a condition for this network which theoretically guarantees identity preservation of the source face even after a sequence of face swapping operations. Extensive experiments demonstrate the advantage of our method over state-of-the-art FS methods in producing high-quality swapped faces. Our source code was made public at https://github.com/truongvu2000nd/AFS

cs.CV