SearcharxivSearch

arXiv subjects

Jaehun Lee

Publications and source records attributed to Jaehun Lee.

At least 19 recordsLinked to original sources

Mesoscopic eigenvalue statistics for correlated random matrices

We prove a mesoscopic central limit theorem for linear eigenvalue statistics of correlated Hermitian random matrices. The class considered here includes Wigner and Wigner-type matrices, as well as models whose entry correlations decay polynomially in the distance between index pairs. The proof combines a multivariate cumulant expansion with multi-resolvent local laws and a detailed analysis of the resulting variance kernel on the operator-level.

math.PR

3DLS: A 3D Logic-Stacked Architecture for Disaggregated LLM Serving

Large language model (LLM) serving increasingly combines prefill-decode (PD) disaggregation with tensor parallelism (TP) to support large models and long contexts. In conventional 2D/2.5D chiplet architectures, layer-wise prefill-to-decode KV-cache transfer decode-side TP collectives share the same lateral die-to-die (D2D) interconnect, creating mixed-traffic contention on the decode critical path. This contention increases communication latency, prolongs token generation intervals, and degrades end-to-end serving performance. We propose 3DLS, a logic-on-logic 3D-stacked chiplet architecture that separates traffic classes by routing KV-cache transfers through vertical interconnects while preserving decode-side TP collectives on the lateral D2D fabric. 3DLS achieves up to 1.49$\times$ throughput and 60.2\% lower end-to-end (E2E) latency over the shared-fabric planar baseline, and still achieves up to 1.17$\times$ throughput and 31.4\% lower E2E latency over a workload-aware priority-managed planar baseline. These results highlight that physical isolation is an important design principle for future chiplet-based PD-disaggregated LLM serving systems.

cs.AR

EHRNote-ChatQA: A Benchmark for Evidence-Grounded Multi-Turn Clinical Question Answering over Longitudinal Discharge Summaries

Discharge summaries are crucial clinical documents containing the context of a patient's overall hospital stay, and are routinely reviewed by medical experts for patient readmission, ongoing care, and diagnostic decision-making. When reviewing them, medical experts often must iteratively synthesize information across multiple summaries while verifying the evidence supporting each answer. Although large language models (LLMs) are increasingly explored for clinical question answering, existing benchmarks do not sufficiently reflect this setting: they often evaluate exam-style medical knowledge or focus on single-turn question answering with limited evidence-grounding evaluation. We introduce EHRNote-ChatQA, the first benchmark for evidence-grounded multi-turn clinical question answering over patients' multiple discharge summaries. Built from de-identified MIMIC-IV discharge summaries, EHRNote-ChatQA contains 967 patient-level multi-turn samples spanning one to five notes and 16,072 medical-expert-verified QA pairs (8,036 content questions, each paired with an evidence-grounding question) across eight clinical categories. The benchmark is constructed through an expert-informed pipeline combining discharge-summary structuring schema, expert-curated multi-turn QA templates, and LLM-based generation, followed by review and revision of every single QA sample by 11 medical experts. Benchmarking 22 open- and closed-source LLMs reveals several challenges, including that LLMs struggle more with evidence grounding than content answering, multi-turn errors compound across turns, and single-turn clinical QA performance does not reliably transfer to this setting. These findings establish EHRNote-ChatQA as a rigorous and practical benchmark for evaluating clinical QA systems. The dataset will be made publicly available through PhysioNet credentialed access.

cs.CL

MASQ: Accelerating Masked Diffusion via Stage-Wise Multi-Precision Quantization

Masked diffusion enables region-specific image synthesis but suffers from computational redundancy, since the entire image is processed each timestep even though only the masked region requires generation. To address this, we introduce MASQ, a hardware-software co-designed accelerator for masked diffusion. Our approach performs stage-wise MXINT8/4/2 precision assignment that dynamically reflects spatial and semantic importance, complemented by timestep-aware scheduling and optimized non-matrix operations. MASQ features a block-wise multi-precision compute engine and mask management unit, efficiently handling our approach. It achieves up to 16.06x and 5.39x speedup and 4.18x and 4.93x energy-efficiency gain over A100 and Orin NX, respectively, while preserving quality.

cs.AR

Signal detection from spiked noise via asymmetrization

The signal plus noise model $H=S+Y$ is a fundamental model in signal detection when a low rank signal $S$ is polluted by noise $Y$. In the high-dimensional setting, one often uses the leading singular values and corresponding singular vectors of $H$ to conduct the statistical inference of the signal $S$. Especially, when $Y$ consists of iid random entries, the singular values of $S$ can be estimated from those of $H$ as long as the signal $S$ is strong enough. However, when the $Y$ entries are heteroscedastic or heavy-tailed, this standard approach may fail. Especially in this work, we consider a situation that can easily arise with heteroscedastic or heavy-tailed noise but is particularly difficult to address using the singular value approach, namely, when the noise $Y$ itself may create spiked singular values. It has been a recurring question how to distinguish the signal $S$ from the spikes in $Y$, as this seems impossible by examining the leading singular values of $H$. Inspired by the work \cite{CCF21}, we turn to study the eigenvalues of an asymmetrized model when two samples $H_1=S+Y_1$ and $H_2=S+Y_2$ are available. We show that by looking into the leading eigenvalues (in magnitude) of the asymmetrized model $H_1H_2^*$, one can easily detect $S$. We will primarily discuss the heteroscedastic case and then discuss the extension to the heavy-tailed case. As a byproduct, we also derive the fundamental result regarding the outlier of non-Hermitian random matrix in \cite{Tao} under the minimal 2nd moment condition.

math.ST

Dichotomy in the small-time asymptotics of spectral heat content for L\'evy processes

We establish a dichotomy in the small-time asymptotic behavior of the spectral heat content (SHC) for symmetric, but not necessarily isotropic, L\'evy processes whose L\'evy density satisfies a weak lower scaling condition near zero. This dichotomy is governed by whether the process has unbounded or bounded variation. In the unbounded variation case, the leading asymptotic behavior of the SHC is determined by the expected supremum of the process projected in the normal direction near the boundary. In contrast, for processes with bounded variation, the SHC decays linearly in time. Our main result, Theorem \ref{thm:main}, extends and unifies key results from \cite{GPS19}, \cite{KP24}, and \cite{PS22}, covering a broader class of non-isotropic L\'evy processes and offering a streamlined proof.

math.PR

Phase transition for the bottom singular vector of rectangular random matrices

In this paper, we consider the rectangular random matrix $X=(x_{ij})\in \mathbb{R}^{N\times n}$ whose entries are iid with tail $\mathbb{P}(|x_{ij}|>t)\sim t^{-α}$ for some $α>0$. We consider the regime $N(n)/n\to \mathsf{a}>1$ as $n$ tends to infinity. Our main interest lies in the right singular vector corresponding to the smallest singular value, which we will refer to as the "bottom singular vector", denoted by $\mathfrak{u}$. In this paper, we prove the following phase transition regarding the localization length of $\mathfrak{u}$: when $α<2$ the localization length is $O(n/\log n)$; when $α>2$ the localization length is of order $n$. Similar results hold for all right singular vectors around the smallest singular value. The variational definition of the bottom singular vector suggests that the mechanism for this localization-delocalization transition when $α$ goes across $2$ is intrinsically different from the one for the top singular vector when $α$ goes across $4$.

math.PR

Phase transition for the smallest eigenvalue of covariance matrices

In this paper, we study the smallest non-zero eigenvalue of the sample covariance matrices $\mathcal{S}(Y)=YY^*$, where $Y=(y_{ij})$ is an $M\times N$ matrix with iid mean $0$ variance $N^{-1}$ entries. We prove a phase transition for its distribution, induced by the fatness of the tail of $y_{ij}$'s. More specifically, we assume that $y_{ij}$ is symmetrically distributed with tail probability $\mathbb{P}(|\sqrt{N}y_{ij}|\geq x)\sim x^{-α}$ when $x\to \infty$, for some $α\in (2,4)$. We show the following conclusions: (i). When $α>\frac83$, the smallest eigenvalue follows the Tracy-Widom law on scale $N^{-\frac23}$; (ii). When $2<α<\frac83$, the smallest eigenvalue follows the Gaussian law on scale $N^{-\fracα{4}}$; (iii). When $α=\frac83$, the distribution is given by an interpolation between Tracy-Widom and Gaussian; (iv). In case $α\leq \frac{10}{3}$, in addition to the left edge of the MP law, a deterministic shift of order $N^{1-\fracα{2}}$ shall be subtracted from the smallest eigenvalue, in both the Tracy-Widom law and the Gaussian law. Overall speaking, our proof strategy is inspired by \cite{ALY} which is originally done for the bulk regime of the Lévy Wigner matrices. In addition to various technical complications arising from the bulk-to-edge extension, two ingredients are needed for our derivation: an intermediate left edge local law based on a simple but effective matrix minor argument, and a mesoscopic CLT for the linear spectral statistic with asymptotic expansion for its expectation.

math.PR

Probabilistic Imputation for Time-series Classification with Missing Data

Multivariate time series data for real-world applications typically contain a significant amount of missing values. The dominant approach for classification with such missing values is to impute them heuristically with specific values (zero, mean, values of adjacent time-steps) or learnable parameters. However, these simple strategies do not take the data generative process into account, and more importantly, do not effectively capture the uncertainty in prediction due to the multiple possibilities for the missing values. In this paper, we propose a novel probabilistic framework for classification with multivariate time series data with missing values. Our model consists of two parts; a deep generative model for missing value imputation and a classifier. Extending the existing deep generative models to better capture structures of time-series data, our deep generative model part is trained to impute the missing values in multiple plausible ways, effectively modeling the uncertainty of the imputation. The classifier part takes the time series data along with the imputed missing values and classifies signals, and is trained to capture the predictive uncertainty due to the multiple possibilities of imputations. Importantly, we show that naïvely combining the generative model and the classifier could result in trivial solutions where the generative model does not produce meaningful imputations. To resolve this, we present a novel regularization technique that can promote the model to produce useful imputation values that help classification. Through extensive experiments on real-world time series data with missing values, we demonstrate the effectiveness of our method.

cs.LG

General Law of iterated logarithm for Markov processes: Limsup law

In this paper, we discuss general criteria of limsup law of iterated logarithm (LIL) for continuous-time Markov processes. We consider minimal assumptions for LILs to hold at zero(at infinity, respectively) in general metric measure spaces. We establish LILs under local assumptions near zero (near infinity, respectively) on uniform bounds of the expectations of first exit times from balls in terms of a function $ϕ$ and uniform bounds on the tails of the jumping kernel in terms of a function $ψ$. The main result is that a simple ratio test in terms of the functions $ϕ$ and $ψ$ completely determines whether there exists a positive non-decreasing function $Ψ$ such that $\limsup |X_t|/Ψ(t)$ is positive and finite a.s., or not. Our results cover a large class of subordinate diffusions, jump processes with mixed polynomial local growths, jump processes with singular jumping kernels and random conductance models with long range jumps.

math.PR

Higher order fluctuations of extremal eigenvalues of sparse random matrices

We consider extremal eigenvalues of sparse random matrices, a class of random matrices including the adjacency matrices of Erdős-Rényi graphs $\mathcal{G}(N,p)$. Recently, it was shown that the leading order fluctuations of extremal eigenvalues are given by a single random variable associated with the total degree of the graph (Ann. Probab., 48(2):916-962, 2020; Probab. Theory Related Fields, 180:985-1056, 2021). We construct a sequence of random correction terms to capture higher (sub-leading) order fluctuations of extremal eigenvalues in the regime $N^ε < pN < N^{1/3-ε}$. Using these random correction terms, we prove a local law up to a shifted edge and recover the rigidity of extremal eigenvalues under some corrections for $pN>N^ε$.

math.PR

Extremal spectral behavior of weighted random $d$-regular graphs

Analyzing the spectral behavior of random matrices with dependency among entries is a challenging problem. The adjacency matrix of the random $d$-regular graph is a prominent example that has attracted immense interest. A crucial spectral observable is the extremal eigenvalue, which reveals useful geometric properties of the graph. According to the Alon's conjecture, which was verified by Friedman, the (nontrivial) extremal eigenvalue of the random $d$-regular graph is approximately $2\sqrt{d-1}$. In the present paper, we analyze the extremal spectrum of the random $d$-regular graph (with $d\ge 3$ fixed) equipped with random edge-weights, and precisely describe its phase transition behavior with respect to the tail of edge-weights. In addition, we establish that the extremal eigenvector is always localized, showing a sharp contrast to the unweighted case where all eigenvectors are delocalized. Our method is robust and inspired by a sparsification technique developed in the context of Erdős-Rényi graphs (Ganguly and Nam, '22), which can also be applied to analyze the spectrum of general random matrices whose entries are dependent.

math.PR

Heat kernel estimates and their stabilities for symmetric jump processes with general mixed polynomial growths on metric measure spaces

In this paper, we consider a symmetric pure jump Markov process $X$ on a metric measure space with volume doubling conditions. Our focus is on estimating the transition density $p(t,x,y)$ of $X$ and studying its stability when the jumping kernel exhibits general mixed polynomial growth. Unlike previous work, in our setting, the rate function governing the jump growth may not be comparable to the scale function that determines whether $p(t,x,y)$ has near-diagonal or off-diagonal estimates. Under the assumption that lower scaling index of scale function is greater than $1$, we establish stabilities of heat kernel estimates. Additionally, if the metric measure space admits a conservative diffusion process with a transition density satisfying sub-Gaussian bounds, we generalize heat kernel estimates from [3, Theorems 1.2 and 1.4] using the rate function and the function $F$ related to walk dimension of underlying space. As an application, we prove the equivalence between a finite moment condition based on $F$ and a generalized Khintchine-type law of iterated logarithm at infinity for symmetric Markov processes.

math.PR

Probabilistic Reasoning at Scale: Trigger Graphs to the Rescue

The role of uncertainty in data management has become more prominent than ever before, especially because of the growing importance of machine learning-driven applications that produce large uncertain databases. A well-known approach to querying such databases is to blend rule-based reasoning with uncertainty. However, techniques proposed so far struggle with large databases. In this paper, we address this problem by presenting a new technique for probabilistic reasoning that exploits Trigger Graphs (TGs) -- a notion recently introduced for the non-probabilistic setting. The intuition is that TGs can effectively store a probabilistic model by avoiding an explicit materialization of the lineage and by grouping together similar derivations of the same fact. Firstly, we show how TGs can be adapted to support the possible world semantics. Then, we describe techniques for efficiently computing a probabilistic model, and formally establish the correctness of our approach. We also present an extensive empirical evaluation using a prototype called LTGs. Our comparison against other leading engines shows that LTGs is not only faster, even against approximate reasoning techniques, but can also reason over probabilistic databases that existing engines cannot scale to.

cs.DB

Laws of the iterated logarithm for occupation times of Markov processes

In this paper, we discuss the laws of the iterated logarithm (LIL) for occupation times of Markov processes $Y$ in general metric measure space both near zero and near infinity under some minimal assumptions. We first establish LILs of (truncated) occupation times on balls $B(x,r)$ of radii $r$ up to an function $Φ(r)$, which is an iterated logarithm of mean exit time of $Y$, by showing that the function $Φ$ is optimal. Our first result on LILs of occupation times covers both near zero and near infinity regardless of transience and recurrence of the process. Our assumptions are truly local in particular at zero and the function $Φ$ in our truncated occupation times $r \mapsto\int_0^{ Φ(x,r)} {\bf 1}_{B(x,r)}(Y_s)ds$ depends on space variable $x$ too. We also prove that a similar LIL for total occupation times $r \mapsto\int_0^\infty {\bf 1}_{B(x,r)}(Y_s)ds$ holds when the process is transient. Then we establish LIL concerning large time behaviors of occupation times $t \mapsto \int_0^t {\bf 1}_{A}(Y_s)ds$ under an additional condition that guarantees the recurrence of the process. Our results cover a large class of Feller (Levy-like) processes, random conductance models with long range jumps, jump processes with mixed polynomial local growths and jump processes with singular jumping kernels.

math.PR

General Law of iterated logarithm for Markov processes: Liminf laws

Continuing from arXiv:2102.01917v2, in this paper, we discuss general criteria and forms of liminf laws of iterated logarithm (LIL) for continuous-time Markov processes. Under some minimal assumptions, which are weaker than those in arXiv:2102.01917v2, we establish liminf LIL at zero (at infinity, respectively) in general metric measure spaces. In particular, our assumptions for liminf law of LIL at zero and the form of liminf LIL are truly local so that we can cover highly space-inhomogenous cases. Our results cover all examples in arXiv:2102.01917v2 including random conductance models with long range jumps. Moreover, we show that the general form of liminf law of LIL at zero holds for a large class of jump processes whose jumping measures have logarithmic tails and Feller processes with symbols of varying order which are not covered before.

math.PR

Noise sensitivity for the top eigenvector of a sparse random matrix

We investigate the noise sensitivity of the top eigenvector of a sparse random symmetric matrix. Let $v$ be the top eigenvector of an $N\times N$ sparse random symmetric matrix with an average of $d$ non-zero centered entries per row. We resample $k$ randomly chosen entries of the matrix and obtain another realization of the random matrix with top eigenvector $v^{[k]}$. Building on recent results on sparse random matrices and a noise sensitivity analysis previously developed for Wigner matrices, we prove that, if $d\geq N^{2/9}$, with high probability, when $k \ll N^{5/3}$, the vectors $v$ and $v^{[k]}$ are almost collinear and, on the contrary, when $k\gg N^{5/3}$, the vectors $v$ and $v^{[k]}$ are almost orthogonal. A similar result holds for the eigenvector associated to the second largest eigenvalue of the adjacency matrix of an Erdős-Rényi random graph with average degree $d \geq N^{2/9}$.

math.PR

Noise sensitivity of second-top eigenvectors of Erdős-Rényi graphs and sparse matrices

We consider eigenvectors of adjacency matrices of Erdős-Rényi graphs and study the variation of their directions by resampling the entries randomly. Let $\mathbf{v}$ be the eigenvector associated with the second-largest eigenvalue of the Erdős-Rényi graphs. After choosing $k$ entries of the given matrix randomly and resampling them, we obtain another eigenvector $\mathbf{w}$ corresponding to the second-largest eigenvalue of the matrix obtained from the resampling procedure. We prove that, in a certain sparsity regime, $\mathbf{w}$ is "almost" orthogonal to $\mathbf{v}$ with high probability if $k\gg N^{5/3}$. On the other hand, if $k\ll q^2 N^{2/3}$, where $q$ is the sparsity parameter, we observe that $\mathbf{v}$ and $\mathbf{w}$ are "almost" collinear. This extends the recent work of Bordenave, Lugosi and Zhivotovskiy to the Erdős-Rényi model.

math.PR