SearcharxivSearch

arXiv subjects

Jungin Lee

Publications and source records attributed to Jungin Lee.

At least 19 recordsLinked to original sources

The average number of rational preperiodic points of polynomials over $\mathbb{Q}$

Let $M_d(X)$ denote the average number of rational preperiodic points among degree-$d$ polynomials over $\mathbb{Q}$ with vanishing $z^{d-1}$ coefficient, constant term $1$, and height at most $X$. We prove that for every integer $d\ge3$ and every real number $C>4\sqrt{2}$, $M_d(X) \ll_d X^{-1} \exp(C\sqrt{\log X / \log \log X})$. For $d=2$, we prove the optimal bound $M_2(X)\ll X^{-1}$.

math.NT

An upper bound for counting algebraic tori over $\mathbb{Q}$ by Artin conductor

Let $N_n^{\mathrm{tor}}(X)$ be the number of isomorphism classes of $n$-dimensional algebraic tori over $\mathbb{Q}$ whose Artin conductor is bounded by $X$. We prove that there is an absolute constant $C>0$ such that, for every positive integer $n \ge 2$, $N_n^{\mathrm{tor}}(X)\ll_n X^{\exp(C(\log n)^2)}$. The proofs of the main results were developed through an iterative dialogue with ChatGPT 5.6 Pro.

math.NT

A mod $p$ determinant criterion for Cohen--Lenstra convergence of random $p$-adic matrices with prescribed zero patterns

We study the distribution of cokernels of Haar-random matrices over the $p$-adic integers with prescribed zero patterns, motivated by the Cohen--Lenstra heuristics. A central feature of our approach is that the asymptotic cokernel distribution is governed by the reductions modulo $p$ of these matrices, viewed as random matrices over the finite field $\mathbb{F}_p$. For several families of support patterns arising from stair-shaped zero regions, including general stair-shaped patterns, band matrices, and matrices with two symmetric stair-shaped zero regions, we show that convergence of the cokernel distribution to the Cohen--Lenstra distribution is equivalent to an asymptotic nonsingularity condition over $\mathbb{F}_p$. We further propose a conjecture for general support patterns and give examples showing that analogous rank-$r$ criteria fail for $r\ge 1$.

math.NT

Universality of the cokernels of random $p$-adic matrices with inhomogeneously balanced columns

In this paper, we prove universality of the distribution of the cokernels of random $p$-adic matrices with inhomogeneously balanced columns. More precisely, let $u \ge 0$ be an integer, and for each positive integer $n$, let $A(n)$ be a random $n \times (n+u)$ matrix over $\mathbb{Z}_p$ whose $i$-th column is $\alpha_n(i)$-balanced. We prove that if $\sum_{i=1}^{n+u} \exp(-\epsilon \alpha_n(i)n) \to 0$ as $n \to \infty$ for every $\epsilon>0$, then the cokernels of $A(n)$ converge in distribution, as $n \to \infty$, to the same limiting law as the cokernels of Haar-random $n \times (n+u)$ matrices over $\mathbb{Z}_p$. This extends a universality theorem of Nguyen and Wood to random $p$-adic matrices with inhomogeneously balanced columns.

math.NT

Sharp threshold for universality of cokernels of classical random matrix models over the $p$-adic integers

We prove that $\frac{\log n}{n}$ is the sharp threshold for universality of the distribution of cokernels of random matrices over $\mathbb{Z}_p$. More precisely, let $\alpha_n = \frac{c\log n}{n}$ for a constant $c>0$ and let $A(n)$ be an $\alpha_n$-balanced random matrix over $\mathbb{Z}_p$. For non-symmetric, symmetric, and alternating matrix models, we prove that if $c>1$, then the limiting distribution of the cokernel of $A(n)$ coincides with the universal distribution of the corresponding symmetry type, whereas universality fails at the critical scale $c=1$. This improves earlier universality results, which required $\alpha_n \gg \frac{\log n}{n}$, to the optimal threshold. As an application, we generalize the universality result for Sylow $p$-subgroups of sandpile groups of Erd\H{o}s-R\'enyi random graphs to a broader class of Erd\H{o}s-R\'enyi graph sequences. Our approach is based on a unified framework that simultaneously treats all symmetry types of random matrices as well as the random graph model, rather than handling each case separately.

math.CO

Sharp threshold for universality of cokernels of random matrices over finite fields

In this paper, we determine the sharp threshold for universality of cokernels of random matrices over finite fields. More precisely, we prove the following: given any constant $c>1$, let $(A(n))_{n \ge 1}$ be a sequence of random $n \times n$ matrices over $\mathbb{F}_p$ such that, for all sufficiently large $n$, the entries of $A(n)$ are independent and take any given value of $\mathbb{F}_p$ with probability at most $1 - \frac{c \log n}{n}$. Then the cokernels of $A(n)$ converge in distribution, as $n \to \infty$, to the same limiting law as the cokernels of uniform random $n \times n$ matrices over $\mathbb{F}_p$. This answers an open problem posed by Wood.

math.PR

On simultaneously preperiodic points for one-parameter families of polynomials in characteristic $p$

For a field $L$ of characteristic $p$, a polynomial $f \in \overline{\mathbb{F}}_p[x]$ and $\alpha, \beta \in L$, let $\mathrm{Prep}(f;\alpha,\beta)$ be the set of all $\lambda \in \overline{L}$ such that both $\alpha$ and $\beta$ are preperiodic under the action of $f_{\lambda}(x) := f(x) + \lambda$. Ghioca and Hsia proved that for certain families of polynomials, this set is infinite if and only if $f(\alpha)=f(\beta)$ or $\alpha, \beta \in \overline{\mathbb{F}}_p$. Building on their work, we determine when $\mathrm{Prep}(f;\alpha,\beta)$ is infinite for most of the remaining binomial cases that were left open. Specifically, let $f(x)=c_1 x^{d_1} + c_2 x^{d_2} \in \overline{\mathbb{F}}_p[x]$, where $c_i \in \overline{\mathbb{F}}_p^*$, $1 \le d_1 < d_2$ and $d_i=p^{\ell_i}s_i$ with $\ell_i \ge 0$ and $p \nmid s_i$. We prove that if $p^{\ell_2}(s_2-1) < p^{\ell_1}(s_1-1)$, then $\mathrm{Prep}(f;\alpha,\beta)$ is infinite if and only if $f(\alpha)=f(\beta)$ or $\alpha, \beta \in \overline{\mathbb{F}}_p$. The key idea of the proof is to use the parameters $\lambda_{\overline{\alpha}} := \overline{\alpha} - f(\overline{\alpha})$ associated to suitable elements $\overline{\alpha} \in \overline{L}$ satisfying $f(\overline{\alpha})=f(\alpha)$. As an application, we extend the work of Asgarli and Ghioca on the colliding orbits problem to binomials satisfying $s_2>1$ and $p^{\ell_2}(s_2-1) < p^{\ell_1}(s_1-1)$.

math.NT

Distribution of the cokernels of determinantal row-sparse matrices

We study the distribution of the cokernels of random row-sparse integral matrices $A_n$ according to the determinantal measure from a structured matrix $B_n$ with a parameter $k_n \ge 3$. Under a mild assumption on the growth rate of $k_n$, we prove that the distribution of the $p$-Sylow subgroup of the cokernel of $A_n$ converges to that of Cohen--Lenstra for every prime $p$. Our result extends the work of A. M\'esz\'aros which established convergence to the Cohen--Lenstra distribution when $p \ge 5$ and $k_n=3$ for all positive integers $n$.

math.PR

ValSub: Subsampling Validation Data to Mitigate Forgetting during ASR Personalization

Automatic Speech Recognition (ASR) is widely used within consumer devices such as mobile phones. Recently, personalization or on-device model fine-tuning has shown that adaptation of ASR models towards target user speech improves their performance over rare words or accented speech. Despite these gains, fine-tuning on user data (target domain) risks the personalized model to forget knowledge about its original training distribution (source domain) i.e. catastrophic forgetting, leading to subpar general ASR performance. A simple and efficient approach to combat catastrophic forgetting is to measure forgetting via a validation set that represents the source domain distribution. However, such validation sets are large and impractical for mobile devices. Towards this, we propose a novel method to subsample a substantially large validation set into a smaller one while maintaining the ability to estimate forgetting. We demonstrate the efficacy of such a dataset in mitigating forgetting by utilizing it to dynamically determine the number of ideal fine-tuning epochs. When measuring the deviations in per user fine-tuning epochs against a 50x larger validation set (oracle), our method achieves a lower mean-absolute-error (3.39) compared to randomly selected subsets of the same size (3.78-8.65). Unlike random baselines, our method consistently tracks the oracle's behaviour across three different forgetting thresholds.

eess.AS

persoDA: Personalized Data Augmentation for Personalized ASR

Data augmentation (DA) is ubiquitously used in training of Automatic Speech Recognition (ASR) models. DA offers increased data variability, robustness and generalization against different acoustic distortions. Recently, personalization of ASR models on mobile devices has been shown to improve Word Error Rate (WER). This paper evaluates data augmentation in this context and proposes persoDA; a DA method driven by user's data utilized to personalize ASR. persoDA aims to augment training with data specifically tuned towards acoustic characteristics of the end-user, as opposed to standard augmentation based on Multi-Condition Training (MCT) that applies random reverberation and noises. Our evaluation with an ASR conformer-based baseline trained on Librispeech and personalized for VOICES shows that persoDA achieves a 13.9% relative WER reduction over using standard data augmentation (using random noise & reverberation). Furthermore, persoDA shows 16% to 20% faster convergence over MCT.

eess.AS

Random $p$-adic matrices with fixed zero entries and the Cohen--Lenstra distribution

In this paper, we study the distribution of the cokernels of random $p$-adic matrices with fixed zero entries. Let $X_n$ be a random $n \times n$ matrix over $\mathbb{Z}_p$ in which some entries are fixed to be zero and the other entries are i.i.d. copies of a random variable $\xi \in \mathbb{Z}_p$. We consider the minimal number of random entries of $X_n$ required for the cokernel of $X_n$ to converge to the Cohen--Lenstra distribution. When $\xi$ is given by the Haar measure, we prove a lower bound of the number of random entries and prove its converse-type result using random regular bipartite multigraphs. When $\xi$ is a general random variable, we determine the minimal number of random entries. Let $M_n$ be a random $n \times n$ matrix over $\mathbb{Z}_p$ with $k$-step stairs of zeros and the other entries given by independent random $\epsilon$-balanced variables valued in $\mathbb{Z}_p$. We prove that the cokernel of $M_n$ converges to the Cohen--Lenstra distribution under a mild assumption. This extends Wood's universality theorem on random $p$-adic matrices.

math.NT

Intersection of orbits for polynomials in characteristic $p$

In [GTZ08, GTZ12], the following result was established: given polynomials $f,g\in\mathbb{C}[x]$ of degrees larger than $1$, if there exist $\alpha,\beta\in\mathbb{C}$ such that their corresponding orbits $\mathcal{O}_f(\alpha)$ and $\mathcal{O}_g(\beta)$ (under the action of $f$, respectively of $g$) intersect in infinitely many points, then $f$ and $g$ must share a common iterate, i.e., $f^m=g^n$ for some $m,n\in\mathbb{N}$. If one replaces $\mathbb{C}$ with a field $K$ of characteristic $p$, then the conclusion fails; we provide numerous examples showing the complexity of the problem over a field of positive characteristic. We advance a modified conjecture regarding polynomials $f$ and $g$ which admit two orbits with infinite intersection over a field of characteristic $p$. Then we present various partial results, along with connections with another deep conjecture in the area, the dynamical Mordell-Lang conjecture.

math.NT

A converse of dynamical Mordell--Lang conjecture in positive characteristic

In this paper, we prove the converse of the dynamical Mordell--Lang conjecture in positive characteristic: For every subset $S \subseteq \mathbb{N}_0$ which is a union of finitely many arithmetic progressions along with finitely many $p$-sets of the form $\left \{ \sum_{j=1}^{m} c_j p^{k_jn_j} : n_j \in \mathbb{N}_0 \right \}$ ($c_j \in \mathbb{Q}$, $k_j \in \mathbb{N}_0$), there exist a split torus $X = \mathbb{G}_m^k$ defined over $K=\overline{\mathbb{F}_p}(t)$, an endomorphism $\Phi$ of $X$, $\alpha \in X(K)$ and a closed subvariety $V \subseteq X$ such that $\left \{ n \in \mathbb{N}_0 : \Phi^n(\alpha) \in V(K) \right \} = S$.

math.NT

Locality enhanced dynamic biasing and sampling strategies for contextual ASR

Automatic Speech Recognition (ASR) still face challenges when recognizing time-variant rare-phrases. Contextual biasing (CB) modules bias ASR model towards such contextually-relevant phrases. During training, a list of biasing phrases are selected from a large pool of phrases following a sampling strategy. In this work we firstly analyse different sampling strategies to provide insights into the training of CB for ASR with correlation plots between the bias embeddings among various training stages. Secondly, we introduce a neighbourhood attention (NA) that localizes self attention (SA) to the nearest neighbouring frames to further refine the CB output. The results show that this proposed approach provides on average a 25.84% relative WER improvement on LibriSpeech sets and rare-word evaluation compared to the baseline.

eess.AS

Consistency Based Unsupervised Self-training For ASR Personalisation

On-device Automatic Speech Recognition (ASR) models trained on speech data of a large population might underperform for individuals unseen during training. This is due to a domain shift between user data and the original training data, differed by user's speaking characteristics and environmental acoustic conditions. ASR personalisation is a solution that aims to exploit user data to improve model robustness. The majority of ASR personalisation methods assume labelled user data for supervision. Personalisation without any labelled data is challenging due to limited data size and poor quality of recorded audio samples. This work addresses unsupervised personalisation by developing a novel consistency based training method via pseudo-labelling. Our method achieves a relative Word Error Rate Reduction (WERR) of 17.3% on unlabelled training data and 8.1% on held-out data compared to a pre-trained model, and outperforms the current state-of-the art methods.

eess.AS

Joint distribution of the cokernels of random $p$-adic matrices II

In this paper, we study the combinatorial relations between the cokernels $\text{cok}(A_n+px_iI_n)$ ($1 \le i \le m$) where $A_n$ is an $n \times n$ matrix over the ring of $p$-adic integers $\mathbb{Z}_p$, $I_n$ is the $n \times n$ identity matrix and $x_1, \cdots, x_m$ are elements of $ \mathbb{Z}_p$ whose reductions modulo $p$ are distinct. For a positive integer $m \le 4$ and given $x_1, \cdots, x_m \in \mathbb{Z}_p$, we determine the set of $m$-tuples of finitely generated $\mathbb{Z}_p$-modules $(H_1, \cdots, H_m)$ for which $(\text{cok}(A_n+px_1I_n), \cdots, \text{cok}(A_n+px_mI_n)) = (H_1, \cdots, H_m)$ for some matrix $A_n$. We also prove that if $A_n$ is an $n \times n$ Haar random matrix over $\mathbb{Z}_p$ for each positive integer $n$, then the joint distribution of $\text{cok}(A_n+px_iI_n)$ ($1 \le i \le m$) converges as $n \rightarrow \infty$.

math.CO

Mixed moments and the joint distribution of random groups

We study the joint distribution of random abelian and non-abelian groups. In the abelian case, we prove several universality results for the joint distribution of the multiple cokernels for random $p$-adic matrices. In the non-abelian case, we compute the joint distribution of random groups given by the quotients of the free profinite group by random relations. In both cases, we generalize the known results on the distribution of the cokernels of random $p$-adic matrices and random groups. Our proofs are based on the observation that mixed moments determine the joint distribution of random groups, which extends the works of Wood for abelian groups and Sawin for non-abelian groups.

math.NT

Universality of the cokernels of random $p$-adic Hermitian matrices

In this paper, we study the distribution of the cokernel of a general random Hermitian matrix over the ring of integers $\mathcal{O}$ of a quadratic extension $K$ of $\mathbb{Q}_p$. For each positive integer $n$, let $X_n$ be a random $n \times n$ Hermitian matrix over $\mathcal{O}$ whose upper triangular entries are independent and their reductions are not too concentrated on certain values. We show that the distribution of the cokernel of $X_n$ always converges to the same distribution which does not depend on the choices of $X_n$ as $n \rightarrow \infty$ and provide an explicit formula for the limiting distribution. This answers Open Problem 3.16 from the ICM 2022 lecture note of Wood in the case of the ring of integers of a quadratic extension of $\mathbb{Q}_p$.

math.NT