SearcharxivSearch

arXiv subjects

Jaehoon Kang

Publications and source records attributed to Jaehoon Kang.

16 recordsLinked to original sources

Edge-Inference Governors Need Memory-Clock State

On integrated edge SoCs whose memory fabric is governed independently of the compute clocks, frequency-aware latency estimators let deadline-aware DVFS governors schedule ML inference by modeling latency over CPU and GPU clocks -- but they do not condition on the memory clock (EMC), a deployment state that decides whether a governor meets its deadlines and at what energy. We show this with a deployed, measured governor on Jetson Orin: an EMC-blind GPU-only fit misses 25-28% of cycles at tight deadlines, whereas an EMC-aware two-cell refit holds misses to <=0.9% under a 2% QoS budget -- selecting a budget-feasible operating point proactively, where latency-only reactive calibration cannot repair memory-clock-dependent slope error. Across six models on two Orin SKUs, the core MobileNetV2 and ViT-Small results replicate on both boards; detection and LLM deployments reproduce the failure on the NX. Sustained deployment requires two further state layers: decode horizon, where KV-cache growth erodes tight-deadline feasibility over long responses, and GPU co-tenancy, whose occupancy opens queueing tails. A contract-admission policy composes the three layers from bounded probes and deploys live, including a joint cell decoding 2,000 tokens against an active GPU co-tenant where tenancy-blind admission misses 32% of tokens. Under the reference-probe maximum-guard accounting, every accepted contract is decisively measured-feasible. Fresh-probe stress tests then expose a probe-variance failure mode at knife-edge admissions; a dispersion-banded guard, validated on held-out trials, eliminates the observed probe-induced low-clock selections (held-out aggregate 1.19%); and one bin of headroom yielded observed per-launch compliance across eight fresh launches. A governor's guarantee is thus a measured conservatism ladder -- state repair, probe-variance banding, and headroom for launch realization.

cs.PF

GLASS: GRPO-Trained LoRA for Acoustic Style Steering in Zero-Shot Text-to-Speech

We propose GLASS, a framework for composable acoustic style control in zero-shot autoregressive text-to-speech (TTS) that learns controls from post-generation rewards rather than style labels. In zero-shot TTS, a speaker prompt often entangles speaker identity with prosodic attributes such as speaking rate and pitch, making it difficult to change style without changing the prompt itself. GLASS instead treats each acoustic attribute as a reward-defined control direction. For each control axis, GLASS freezes the TTS backbone and trains one lightweight LoRA adapter with Group Relative Policy Optimization (GRPO), using speech-token length and mean F0 as style rewards and WER as an intelligibility anchor. Because each control is represented as a LoRA weight update, independently trained adapters can be swapped, interpolated, and composed through linear LoRA arithmetic without retraining the backbone. Experiments on speaking rate and pitch control show targeted style shifts while preserving naturalness, speaker similarity, and intelligibility, and demonstrate smooth interpolation and multi-axis composition across independently trained adapters.

cs.SD

DialBGM: A Benchmark for Background Music Recommendation from Everyday Multi-Turn Dialogues

Selecting an appropriate background music (BGM) that supports natural human conversation is a common production step in media and interactive systems. In this paper, we introduce dialogue-conditioned BGM recommendation, where a model should select non-intrusive, fitting music for a multi-turn conversation that often contains no music descriptors. To study this novel problem, we present DialBGM, a benchmark of 1,200 open-domain daily dialogues, each paired with four candidate music clips and annotated with human preference rankings. Rankings are determined by background suitability criteria, including contextual relevance, non-intrusiveness, and consistency. We evaluate a wide range of open-source and proprietary models, including audio-language models and multimodal LLMs, and show that current models fall far short of human judgments; no model exceeds 35% Hit@1 when selecting the top-ranked clip. DialBGM provides a standardized benchmark for developing discourse-aware methods for BGM selection and for evaluating both retrieval-based and generative models.

cs.AI

Unlocking Fine-Grained and Within-Utterance Speaking Style Control in Prompt-Based Text-to-Speech Models

While prompt-based text-to-speech (TTS) models enable natural language-driven speaking style control, they often provide limited fine-grained control and apply a single global style across an utterance. This restricts practical use cases that require continuous style attribute interpolation across utterances and time-varying style transitions within a single utterance. In this paper, we propose novel techniques to achieve both capabilities in existing prompt-based TTS models. For inter-utterance style interpolation, we compute direction vectors between contrastive style prompts in the embedding space and perform simple interpolation, enabling smooth transitions between style characteristics. For intra-utterance style transition, we first identify a strong attention bias toward early tokens in autoregressive TTS decoders, causing the initial audio realization to dominate subsequent generation. To mitigate this effect, we introduce KV-cache swapping and sliding-window attention masking. Experiments demonstrate that our proposed inter-utterance interpolation achieves a 99-100% success rate in gender conversion, up to 36 Hz pitch variation, and up to 1.6 syllables-per-second speed change. Our intra-utterance transition maintains a speaker similarity of 0.81-0.91 and achieves perceptual smoothness scores of 3.48-4.48.

cs.CL

WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models

Recent decoder-only autoregressive text-to-speech (AR-TTS) models produce high-fidelity speech, but their memory and compute costs scale quadratically with sequence length due to full self-attention. In this paper, we propose WAND, Windowed Attention and Knowledge Distillation, a framework that adapts pretrained AR-TTS models to operate with constant computational and memory complexity. WAND separates the attention mechanism into two: persistent global attention over conditioning tokens and local sliding-window attention over generated tokens. To stabilize fine-tuning, we employ a curriculum learning strategy that progressively tightens the attention window. We further utilize knowledge distillation from a full-attention teacher to recover high-fidelity synthesis quality with high data efficiency. Evaluated on three modern AR-TTS models, WAND preserves the original quality while achieving up to 66.2% KV cache memory reduction and length-invariant, near-constant per-step latency.

cs.CL

P2VA: Converting Persona Descriptions into Voice Attributes for Fair and Controllable Text-to-Speech

While persona-driven large language models (LLMs) and prompt-based text-to-speech (TTS) systems have advanced significantly, a usability gap arises when users attempt to generate voices matching their desired personas from implicit descriptions. Most users lack specialized knowledge to specify detailed voice attributes, which often leads TTS systems to misinterpret their expectations. To address these gaps, we introduce Persona-to-Voice-Attribute (P2VA), the first framework enabling voice generation automatically from persona descriptions. Our approach employs two strategies: P2VA-C for structured voice attributes, and P2VA-O for richer style descriptions. Evaluation shows our P2VA-C reduces WER by 5% and improves MOS by 0.33 points. To the best of our knowledge, P2VA is the first framework to establish a connection between persona and voice synthesis. In addition, we discover that current LLMs embed societal biases in voice attributes during the conversion process. Our experiments and findings further provide insights into the challenges of building persona-voice systems.

eess.AS

A regularity theory for evolution equations with space-time anisotropic non-local operators in mixed-norm Sobolev spaces

In this article, we study the regularity of solutions to inhomogeneous time-fractional evolution equations involving anisotropic non-local operators in mixed-norm Sobolev spaces of variable order, with non-trivial initial conditions. The primary focus is on space-time non-local equations where the spatial operator is the infinitesimal generator of a vector of independent subordinate Brownian motions, making it the sum of subdimensional non-local operators. A representative example of such an operator is $(\Delta_{x})^{\beta_{1}/2}+(\Delta_{y})^{\beta_{2}/2}$. We establish existence, uniqueness, and precise estimates for solutions in corresponding Sobolev spaces. Due to singularities arising in the Fourier transforms of our operators, traditional methods involving Fourier analysis are not directly applicable. Instead, we employ a probabilistic approach to derive solution estimates. Additionally, we identify the optimal initial data space using generalized real interpolation theory.

math.AP

Heat kernel estimates for Markov processes of direction-dependent type

We prove sharp pointwise heat kernel estimates for symmetric Markov processes associated with symmetric Dirichlet forms that are local with respect to some coordinates and nonlocal with respect to the remaining coordinates. The main theorem is a robustness result like the famous estimate for the fundamental solution of second order differential operators, obtained by Donald G. Aronson. Analogous to his result, we show that the corresponding translation-invariant process and the one given by the general Dirichlet form share the same pointwise points.

math.PR

A regularity theory for parabolic equations with anisotropic non-local operators in $L_{q}(L_{p})$ spaces

In this paper, we present an $L_q(L_p)$-regularity theory for parabolic equations of the form: $$ \partial_t u(t,x)=\mathcal{L}^{\vec{a},\vec{b}}(t)u(t,x)+f(t,x),\quad u(0,x)=0. $$ Here, $\mathcal{L}^{\vec{a},\vec{b}}(t)$ represents anisotropic non-local operators encompassing the singular anisotropic fractional Laplacian with measurable coefficients: $$ \mathcal{L}^{\vec{a},\vec{0}}(t)u(x)=\sum_{i=1}^{d} \int_{\mathbb{R}}\left( u(x^{1},\dots,x^{i-1},x^{i}+y^{i},x^{i+1},\dots,x^{d}) - u(x) \right) \frac{a_{i}(t,y^{i})}{|y^{i}|^{1+α_{i}}} \mathrm{d}y^{i} . $$ To address the anisotropy of the operator, we employ a probabilistic representation of the solution and Calderón-Zygmund theory. As applications of our results, we demonstrate the solvability of elliptic equations with anisotropic non-local operators and parabolic equations with isotropic non-local operators.

math.AP

An $L_{q}(L_{p})$-regularity theory for parabolic equations with integro-differential operators having low intensity kernels

In this article, we present the existence, uniqueness, and regularity of solutions to parabolic equations with non-local operators $$ \partial_{t}u(t,x) = \mathcal{L}^{a}u(t,x) + f(t,x), \quad t>0 $$ in $L_{q}(L_{p})$ spaces. Our spatial operator $\mathcal{L}^{a}$ is an integro-differential operator of the form $$ \int_{\mathbb{R}^{d}} \left( u(x+y)-u(x) -\nabla u(x) \cdot y \mathrm{1}_{|y|\leq 1} \right) a(t,y) j_{d}(|y|)dy. $$ Here, $a(t,y)$ is a merely bounded measurable coefficient, and we employed the theory of additive process to handle it. We investigate conditions on $j_{d}(r)$ which yield $L_{q}(L_{p})$-regularity of solutions. Our assumptions on $j_d$ are general so that $j_d(r)$ may be comparable to $r^{-d}\ell(r^{-1})$ for a function $\ell$ which is slowly varying at infinity. For example, we can take $\ell(r)=\log{(1+r^{\alpha})}$ or $\ell(r) = \min{\{r^{\alpha},1\}}$ ($\alpha\in(0,2)$). Indeed, our result covers the operators whose Fourier multiplier $\psi(\xi)$ does not have any scaling condition for $|\xi|\geq 1$. Furthermore, we give some examples of operators, which cannot be covered by previous results where smoothness or scaling conditions on $\psi$ are considered.

math.AP

Heat kernel estimates and their stabilities for symmetric jump processes with general mixed polynomial growths on metric measure spaces

In this paper, we consider a symmetric pure jump Markov process $X$ on a metric measure space with volume doubling conditions. Our focus is on estimating the transition density $p(t,x,y)$ of $X$ and studying its stability when the jumping kernel exhibits general mixed polynomial growth. Unlike previous work, in our setting, the rate function governing the jump growth may not be comparable to the scale function that determines whether $p(t,x,y)$ has near-diagonal or off-diagonal estimates. Under the assumption that lower scaling index of scale function is greater than $1$, we establish stabilities of heat kernel estimates. Additionally, if the metric measure space admits a conservative diffusion process with a transition density satisfying sub-Gaussian bounds, we generalize heat kernel estimates from [3, Theorems 1.2 and 1.4] using the rate function and the function $F$ related to walk dimension of underlying space. As an application, we prove the equivalence between a finite moment condition based on $F$ and a generalized Khintchine-type law of iterated logarithm at infinity for symmetric Markov processes.

math.PR

An $L_q(L_p)$-theory for time-fractional diffusion equations with nonlocal operators generated by Lévy processes with low intensity of small jumps

We investigate an $L_{q}(L_{p})$-regularity ($1<p,q<\infty$) theory for space-time nonlocal equations of the type $\partial^α_{t}u = \mathcal{L}u +f$. Here, $\partial^α_{t}$ is the Caputo fractional derivative of order $α\in(0,1)$ and $\mathcal{L}$ is an integro-differential operator $$ \mathcal{L}u(x) = \int_{\mathbb{R}^{d}} \left( u(x)-u(x+y) -\nabla u (x) \cdot y \mathbf{1}_{|y|\leq 1} \right) j_{d}(|y|)dy $$ which is the infinitesimal generator of an isotropic unimodal Lévy process. We assume that the jump kernel $j_{d}(r)$ is comparable to $r^{-d} \ell(r^{-1})$, where $\ell$ is a continuous function satisfying $$ C_{1}\left(\frac{R}{r}\right)^{δ_{1}} \leq \frac{\ell(R)}{\ell(r)} \leq C_{2} \left( \frac{R}{r} \right)^{δ_{2}} \quad \text{for}\;\; \,1\leq r\leq R<\infty, $$ where $0\leq δ_{1}\leq δ_{2}<2$. Hence, $\ell$ can be slowly varying at infinity. Our result covers $\mathcal{L}$ whose Fourier multiplier $Ψ(ξ)$ satisfies $Ψ(ξ)\asymp -\log{(1+|ξ|^β)}$ for $β\in (0,2]$ and $Ψ(ξ) \asymp-(\log(1+|ξ|^{β/4}))^{2}$ for $β\in(0,2)$ by taking $\ell(r) \asymp 1$ and $\ell(r) \asymp \log{(1+r^β)}$ for $r\geq1$ respectively. In this article, we use the Calderón-Zygmund approach and function space theory for operators having slowly varying symbols.

math.AP

Heat kernel estimates for anisotropic symmetric jump processes

We show two-sided bounds of heat kernel for anisotropic non-singular symmetric pure jump Markov process whose jump kernel $J(x,y)$ is comparable to $\frac{{\bf 1}_{\mathcal{V}}(x-y)}{|x-y|^{d+α}}$, where $\mathcal{V}$ is a union of symmetric cones, $0<α<2$ and $x,y\in\mathbb{R}^d$.

math.PR

Estimates of Dirichlet heat kernels for unimodal Lévy processes with low intensity of small jumps

In this paper, we study transition density functions for pure jump unimodal Lévy processes killed upon leaving an open set $D$. Under some mild assumptions on the Lévy density, we establish two-sided Dirichlet heat kernel estimates when the open set $D$ is $C^{1, 1}$. Our result covers the case that the Lévy densities of unimodal Lévy processes are regularly varying functions whose indices are equal to the Euclidean dimension. This is the first results on two-sided Dirichlet heat kernel estimates for Lévy processes such that the weak lower scaling index of the Lévy densities is not necessarily strictly bigger than the Euclidean dimension.

math.PR

Heat kernel estimates for symmetric jump processes with mixed polynomial growths

In this paper, we study the transition densities of pure-jump symmetric Markov processes in $ {\mathbb R}^d$, whose jumping kernels are comparable to radially symmetric functions with mixed polynomial growths. Under some mild assumptions on their scale functions, we establish sharp two-sided estimates of transition densities (heat kernel estimates) for such processes. This is the first study on global heat kernel estimates of jump processes (including non-Lévy processes) whose weak scaling index is not necessarily strictly less than 2. As an application, we proved that the finite second moment condition on such symmetric Markov process is equivalent to the Khintchine-type law of iterated logarithm at the infinity.

math.PR

Tangential limits for harmonic functions with respect to $ϕ(Δ)$ : stable and beyond

In this paper, we discuss tangential limits for regular harmonic functions with respect to $ϕ(Δ):=-ϕ(-Δ)$ in the $C^{1,1}$ open set $D$ in $\mathbb{R}^d$, where $ϕ$ is the complete Bernstein function and $d \ge 2$. When the exterior function $f$ is local $L^p$-Hölder continuous of order $β$ on $D^c$ with $ p\in(1,\infty]$ and $β>1/p$, for a large class of Bernstein function $ϕ$, we show that the regular harmonic function $u_f$ with respect to $ϕ(Δ)$, whose value is $f$ on $D^c$, converges a.e. through a certain parabola that depends on $ϕ$ and $ϕ'$. Our result includes the case $ϕ(λ)=\log(1+λ^{α/2})$. Our proofs use both the probabilistic and analytic methods.

math.PR