SearcharxivSearch

arXiv subjects

Yijun Ding

Publications and source records attributed to Yijun Ding.

6 recordsLinked to original sources

FlexID: Training-Free Flexible Identity Injection via Intent-Aware Modulation for Text-to-Image Generation

Personalized text-to-image generation aims to seamlessly integrate specific identities into textual descriptions. However, existing training-free methods often rely on rigid visual feature injection, creating a conflict between identity fidelity and textual adaptability. To address this, we propose FlexID, a novel training-free framework utilizing intent-aware modulation. FlexID orthogonally decouples identity into two dimensions: a Semantic Identity Projector (SIP) that injects high-level priors into the language space, and a Visual Feature Anchor (VFA) that ensures structural fidelity within the latent space. Crucially, we introduce a Context-Aware Adaptive Gating (CAG) mechanism that dynamically modulates the weights of these streams based on editing intent and diffusion timesteps. By automatically relaxing rigid visual constraints when strong editing intent is detected, CAG achieves synergy between identity preservation and semantic variation. Extensive experiments on IBench demonstrate that FlexID achieves a state-of-the-art balance between identity consistency and text adherence, offering an efficient solution for complex narrative generation.

cs.CV

DVI: Disentangling Semantic and Visual Identity for Training-Free Personalized Generation

Recent tuning-free identity customization methods achieve high facial fidelity but often overlook visual context, such as lighting, skin texture, and environmental tone. This limitation leads to ``Semantic-Visual Dissonance,'' where accurate facial geometry clashes with the input's unique atmosphere, causing an unnatural ``sticker-like'' effect. We propose **DVI (Disentangled Visual-Identity)**, a zero-shot framework that orthogonally disentangles identity into fine-grained semantic and coarse-grained visual streams. Unlike methods relying solely on semantic vectors, DVI exploits the inherent statistical properties of the VAE latent space, utilizing mean and variance as lightweight descriptors for global visual atmosphere. We introduce a **Parameter-Free Feature Modulation** mechanism that adaptively modulates semantic embeddings with these visual statistics, effectively injecting the reference's ``visual soul'' without training. Furthermore, a **Dynamic Temporal Granularity Scheduler** aligns with the diffusion process, prioritizing visual atmosphere in early denoising stages while refining semantic details later. Extensive experiments demonstrate that DVI significantly enhances visual consistency and atmospheric fidelity without parameter fine-tuning, maintaining robust identity preservation and outperforming state-of-the-art methods in IBench evaluations.

cs.CV

Invertibility of Multi-Energy X-ray Transform

Purpose: The goal is to provide a sufficient condition on the invertibility of a multi-energy (ME) X-ray transform. The energy-dependent X-ray attenuation profiles can be represented by a set of coefficients using the Alvarez-Macovski (AM) method. An ME X-ray transform is a mapping from $N$ AM coefficients to $N$ noise-free energy-weighted measurements, where $N\geq2$. Methods: We apply a general invertibility theorem which tests whether the Jacobian of the mapping $J(\mathbf A)$ has zero values over the support of the mapping. The Jacobian of an arbitrary ME X-ray transform is an integration over all spectral measurements. A sufficient condition of $J(\mathbf A)\neq0$ for all $\mathbf A$ is that the integrand of $J(\mathbf A)$ is $\geq0$ (or $\leq0$) everywhere. Note that the trivial case of the integrand equals to zero everywhere is ignored. With symmetry, we simplified the integrand of the Jacobian into three factors that are determined by the total attenuation, the basis functions, and the energy-weighting functions, respectively. The factor related to total attenuation is always positive, hence the invertibility of the X-ray transform can be determined by testing the signs of the other two factors. Furthermore, we use the Cramer-Rao lower bound (CRLB) to characterize the noise-induced estimation uncertainty and provide a maximum-likelihood (ML) estimator. Conclusions: We have provided a framework to study the invertibility of an arbitrary ME X-ray transform and proved the global invertibility for four types of systems.

physics.med-ph

Bounds on mutual information of mixture data for classification tasks

The data for many classification problems, such as pattern and speech recognition, follow mixture distributions. To quantify the optimum performance for classification tasks, the Shannon mutual information is a natural information-theoretic metric, as it is directly related to the probability of error. The mutual information between mixture data and the class label does not have an analytical expression, nor any efficient computational algorithms. We introduce a variational upper bound, a lower bound, and three estimators, all employing pair-wise divergences between mixture components. We compare the new bounds and estimators with Monte Carlo stochastic sampling and bounds derived from entropy bounds. To conclude, we evaluate the performance of the bounds and estimators through numerical simulations.

eess.SP

X-ray measurement model incorporating energy-correlated material variability and its application in information-theoretic system analysis

Extending our prior work, we propose a multi-energy X-ray measurement model incorporating material variability with energy correlations to enable the analysis and exploration of the performance of X-ray imaging and sensing systems. Based on this measurement model, we provide analytical expressions for bounds on the probability of error, $P_e$, to quantify the performance limits of an X-ray measurement system for binary classification task. We analyze the performance of a prototypical X-ray measurement system to demonstrate the utility of our proposed material variability measurement model.

eess.SP

Quantifying allowable motion to achieve safe dose escalation in pancreatic SBRT

Tumor motion plays a key role in the safe delivery of Stereotactic Body Radiotherapy (SBRT) for pancreatic cancer. The purpose of this study was to use tumor motion data measured in patients to establish limits on motion magnitude for safe delivery of pancreatic SBRT. Using 91 sets of pancreatic tumor motion data measured in patients, we calculated motion-convolved dose for 25 pancreatic cancer patients, and established the maximum amount of motion allowable while satisfying error thresholds on key dose metrics. In our patient cohort, the mean [min-max] allowable motion for 33/40/50 Gy to the PTV was 11.9 [6.3-22.4], 10.4 [5.2-19.1] and 9.0 [4.2-16.0] mm, respectively. Maximum allowable motion decreased as dose was escalated, and was smaller in patients with larger tumors. The effects of motion on pancreatic SBRT are highly variable between patients and there is potential to allow more motion in certain patients, even in dose-escalated scenarios. In our dataset, a conservative limit of 6.3 mm would ensure safe treatment of all patients treated to 33 Gy in 5 fractions.

physics.med-ph