SearcharxivSearch

arXiv subjects

Shifan Zhao

Publications and source records attributed to Shifan Zhao.

16 recordsLinked to original sources

On Ramanujan Primes for Hecke-Maass Cusp Forms

For a primitive Hecke-Maass cusp form $\phi$ of level $N$ with the $n$-th Hecke eigenvalue $\lambda_{\phi}(n)$ and a prime number $p\nmid N$, the celebrated Ramanujan conjecture at $p$ asserts the following sharp upper bound: \[ |\lambda_{\phi}(p)| \leq 2. \] In this work, we determine an upper bound for the least prime $p$ at which the Ramanujan conjecture holds for two or three distinct primitive Hecke-Maass cusp forms simultaneously. Moreover, given a set of distinct primitive Hecke-Maass cusp forms $\{\phi_i\}$, we also provide a lower bound for the lower natural density of the set of primes at which the Ramanujan conjecture holds for at least one of the $\phi_i$'s.

math.NT

Landau-Siegel zeros of Rankin-Selberg $L$-functions

We establish standard zero-free regions with no exceptional Landau-Siegel zeros for Rankin-Selberg $L$-functions and triple product $L$-functions in several new families for which modularity is not yet known.

math.NT

Low-lying zeros of Hilbert modular $L$-functions weighted by powers of central $L$-values

Let $\mathcal{F}(\textbf{k},\mathfrak{q})$ be the set of primitive Hilbert modular forms of weight $\textbf{k}$ and prime level $\mathfrak{q}$, with trivial central character. We study the one-level density of low-lying zeros of $L(s,\pi)$ weighted by powers of central $L$-values $L(1/2,\pi)^r$, where $\pi$ runs through $\mathcal{F}(\textbf{k},\mathfrak{q})$. For $r=1,2,3$, we show that the resulting distributions $W_r$ match with predictions from Random Matrix Theory. For general $r \geq 1$, we also formulate a conjectural formula for $W_r$ based on the ``recipe'' method.

math.NT

Landau-Siegel Zeros of Triple Product L-functions

Let $F$ be a number field. Let $\pi_1,\pi_2$ be cuspidal automorphic representations of $GL_2(\mathbb{A}_F)$, and let $\pi$ be a cuspidal automorphic representation of either $GL_2(\mathbb{A}_F)$ or $GL_3(\mathbb{A}_F)$. When $(\pi_1,\pi_2,\pi)$ is of general type, we show that the triple product $L$-function $L(s,\pi_1 \times \pi_2 \times \pi)$ on either $GL(2) \times GL(2) \times GL(2)$ or $GL(2) \times GL(2) \times GL(3)$ has a standard zero-free region with no exceptional Landau-Siegel zero. Moreover, when $(\pi_1,\pi_2,\pi)$ is not of general type, we give precise conditions when $L(s,\pi_1 \times \pi_2 \times \pi)$ could possibly have exceptional Landau-Siegel zeros.

math.NT

Early Risk Prediction of Pediatric Cardiac Arrest from Electronic Health Records via Multimodal Fused Transformer

Early prediction of pediatric cardiac arrest (CA) is critical for timely intervention in high-risk intensive care settings. We introduce PedCA-FT, a novel transformer-based framework that fuses tabular view of EHR with the derived textual view of EHR to fully unleash the interactions of high-dimensional risk factors and their dynamics. By employing dedicated transformer modules for each modality view, PedCA-FT captures complex temporal and contextual patterns to produce robust CA risk estimates. Evaluated on a curated pediatric cohort from the CHOA-CICU database, our approach outperforms ten other artificial intelligence models across five key performance metrics and identifies clinically meaningful risk factors. These findings underscore the potential of multimodal fusion techniques to enhance early CA detection and improve patient care.

cs.LG

Relative Trace Formula and Uniform non-vanishing of Central $L$-values of Hilbert Modular Forms

Let $\mathcal{F}(\mathbf{k},\mathfrak{q})$ be the set of normalized Hilbert newforms of weight $\mathbf{k}$ and prime level $\mathfrak{q}$. In this paper, utilizing regularized relative trace formulas, we establish a positive proportion of $\#\{\pi\in\mathcal{F}(\mathbf{k},\mathfrak{q}):L(1/2,\pi)\neq 0\}$ as $\#\mathcal{F}(\mathbf{k},\mathfrak{q})\to+\infty$. Moreover, our result matches the strength of the best known results in both the level and weight aspects.

math.NT

Efficient Two-Stage Gaussian Process Regression Via Automatic Kernel Search and Subsampling

Gaussian Process Regression (GPR) is widely used in statistics and machine learning for prediction tasks requiring uncertainty measures. Its efficacy depends on the appropriate specification of the mean function, covariance kernel function, and associated hyperparameters. Severe misspecifications can lead to inaccurate results and problematic consequences, especially in safety-critical applications. However, a systematic approach to handle these misspecifications is lacking in the literature. In this work, we propose a general framework to address these issues. Firstly, we introduce a flexible two-stage GPR framework that separates mean prediction and uncertainty quantification (UQ) to prevent mean misspecification, which can introduce bias into the model. Secondly, kernel function misspecification is addressed through a novel automatic kernel search algorithm, supported by theoretical analysis, that selects the optimal kernel from a candidate set. Additionally, we propose a subsampling-based warm-start strategy for hyperparameter initialization to improve efficiency and avoid hyperparameter misspecification. With much lower computational cost, our subsampling-based strategy can yield competitive or better performance than training exclusively on the full dataset. Combining all these components, we recommend two GPR methods-exact and scalable-designed to match available computational resources and specific UQ requirements. Extensive evaluation on real-world datasets, including UCI benchmarks and a safety-critical medical case study, demonstrates the robustness and precision of our methods.

cs.LG

An Adaptive Factorized Nyström Preconditioner for Regularized Kernel Matrices

The spectrum of a kernel matrix significantly depends on the parameter values of the kernel function used to define the kernel matrix. This makes it challenging to design a preconditioner for a regularized kernel matrix that is robust across different parameter values. This paper proposes the Adaptive Factorized Nyström (AFN) preconditioner. The preconditioner is designed for the case where the rank k of the Nyström approximation is large, i.e., for kernel function parameters that lead to kernel matrices with eigenvalues that decay slowly. AFN deliberately chooses a well-conditioned submatrix to solve with and corrects a Nyström approximation with a factorized sparse approximate matrix inverse. This makes AFN efficient for kernel matrices with large numerical ranks. AFN also adaptively chooses the size of this submatrix to balance accuracy and cost.

math.NA

NLTGCR: A class of Nonlinear Acceleration Procedures based on Conjugate Residuals

This paper develops a new class of nonlinear acceleration algorithms based on extending conjugate residual-type procedures from linear to nonlinear equations. The main algorithm has strong similarities with Anderson acceleration as well as with inexact Newton methods - depending on which variant is implemented. We prove theoretically and verify experimentally, on a variety of problems from simulation experiments to deep learning applications, that our method is a powerful accelerated iterative algorithm.

math.NA

Weighted low-lying zeros of L-functions attached to Siegel modular forms

In this paper, we study weighted low-lying zeros of spinor and standard $L$-functions attached to degree 2 Siegel modular forms. We show the symmetry type of weighted low-lying zeros of spinor $L$-functions is symplectic, for test functions whose Fourier transform have support in $(-1,1)$, extending the previous range $(-\frac{4}{15},\frac{4}{15})$ by E. Kowalski, A. Saha and J. Tsimerman . We then show the symmetry type of weighted low-lying zeros of standard $L$-functions is also symplectic. We further extend the range of support by performing an average over weight. As an application, we discuss non-vanishing of central values of those $L$-functions.

math.NT

MuG: A Multimodal Classification Benchmark on Game Data with Tabular, Textual, and Visual Fields

Previous research has demonstrated the advantages of integrating data from multiple sources over traditional unimodal data, leading to the emergence of numerous novel multimodal applications. We propose a multimodal classification benchmark MuG with eight datasets that allows researchers to evaluate and improve their models. These datasets are collected from four various genres of games that cover tabular, textual, and visual modalities. We conduct multi-aspect data analysis to provide insights into the benchmark, including label balance ratios, percentages of missing features, distributions of data within each modality, and the correlations between labels and input modalities. We further present experimental results obtained by several state-of-the-art unimodal classifiers and multimodal classifiers, which demonstrate the challenging and multimodal-dependent properties of the benchmark. MuG is released at https://github.com/lujiaying/MUG-Bench with the data, tutorials, and implemented baselines.

cs.LG

On Möbius functions from automorphic forms and a generalized Sarnak's conjecture

In this paper, we consider Möbius functions associated with two types of $L$-functions: Rankin-Selberg $L$-functions of symmetric powers of distinct holomorphic cusp forms and $L$-functions of Maass cusp forms. We show that these Möbius functions are weakly orthogonal to bounded sequences. As a direct corollary, a generalized Sarnak's conjecture holds for these two types of Möbius functions.

math.NT

MedDiff: Generating Electronic Health Records using Accelerated Denoising Diffusion Model

Due to patient privacy protection concerns, machine learning research in healthcare has been undeniably slower and limited than in other application domains. High-quality, realistic, synthetic electronic health records (EHRs) can be leveraged to accelerate methodological developments for research purposes while mitigating privacy concerns associated with data sharing. The current state-of-the-art model for synthetic EHR generation is generative adversarial networks, which are notoriously difficult to train and can suffer from mode collapse. Denoising Diffusion Probabilistic Models, a class of generative models inspired by statistical thermodynamics, have recently been shown to generate high-quality synthetic samples in certain domains. It is unknown whether these can generalize to generation of large-scale, high-dimensional EHRs. In this paper, we present a novel generative model based on diffusion models that is the first successful application on electronic health records. Our model proposes a mechanism to perform class-conditional sampling to preserve label information. We also introduce a new sampling strategy to accelerate the inference speed. We empirically show that our model outperforms existing state-of-the-art synthetic EHR generation methods.

cs.LG

On Siegel Zeros of Symmetric Power L-functions

Let $f$ be a holomorphic cusp form of even weight $k$ for the modular group $SL(2,\mathbb{Z})$, which is assumed to be a common eigenfunction for all Hecke operators. For positive integer $n$, let $\text{Sym}^n(f)$ be the symmetric nth power lifting of $f$ , which was shown by Newton and Thorne to be automorphic and cuspidal. In this paper, we construct certain auxiliary $L$-functions to show that Siegel zeros of $\text{Sym}^n(f)$ do not exist, for each given $n$, utilizing the above functoriality result. As an application, we give a lower bound of those symmetric power $L$-functions at $s=1$ of logarithm power type.

math.NT

An Efficient Nonlinear Acceleration method that Exploits Symmetry of the Hessian

Nonlinear acceleration methods are powerful techniques to speed up fixed-point iterations. However, many acceleration methods require storing a large number of previous iterates and this can become impractical if computational resources are limited. In this paper, we propose a nonlinear Truncated Generalized Conjugate Residual method (nlTGCR) whose goal is to exploit the symmetry of the Hessian to reduce memory usage. The proposed method can be interpreted as either an inexact Newton or a quasi-Newton method. We show that, with the help of global strategies like residual check techniques, nlTGCR can converge globally for general nonlinear problems and that under mild conditions, nlTGCR is able to achieve superlinear convergence. We further analyze the convergence of nlTGCR in a stochastic setting. Numerical results demonstrate the superiority of nlTGCR when compared with several other competitive baseline approaches on a few problems. Our code will be available in the future.

cs.LG

GDA-AM: On the effectiveness of solving minimax optimization via Anderson Acceleration

Many modern machine learning algorithms such as generative adversarial networks (GANs) and adversarial training can be formulated as minimax optimization. Gradient descent ascent (GDA) is the most commonly used algorithm due to its simplicity. However, GDA can converge to non-optimal minimax points. We propose a new minimax optimization framework, GDA-AM, that views the GDAdynamics as a fixed-point iteration and solves it using Anderson Mixing to con-verge to the local minimax. It addresses the diverging issue of simultaneous GDAand accelerates the convergence of alternating GDA. We show theoretically that the algorithm can achieve global convergence for bilinear problems under mild conditions. We also empirically show that GDA-AMsolves a variety of minimax problems and improves GAN training on several datasets

cs.LG