SearcharxivSearch

arXiv subjects

Tianming Zhu

Publications and source records attributed to Tianming Zhu.

13 recordsLinked to original sources

On clique-to-clique densities

Let $k_r(G)$ denote the number of $r$-cliques in a graph $G$ and let $F_r(\cdot)$ be the Lovász--Simonovits $r$-clique density function. For any integers $2\le s<t$, we determine the asymptotically sharp lower bound on $k_t(G)$ in an $n$-vertex graph $G$ with a prescribed number $k_s(G)$, by showing that \[ \frac{k_t(G)}{n^t}\ge F_t\!\left(F_s^{-1}\!\left(\frac{k_s(G)}{n^s}\right)\right), \] where $F_s^{-1}$ denotes the generalized inverse. This strengthens Bollobás's piecewise-linear interpolation bound and, in the case $s=2$, recovers Reiher's clique density theorem via a new inductive proof.

math.CO

On a hypergraph Turán problem of Balogh-Bohman-Bollobás-Zhao

Let $S$ and $T$ be disjoint sets with $|S|=i$ and $|T|=r-1$ for $2\le i\le r-1$, and let $B_i^{(r)}$ be the $r$-graph on $S\cup T$ whose edges are the $r$-subsets containing $S$ or $T$. We study the deficit $q_{r,i}:=1-π(B_i^{(r)})$ in its Turán density. Balogh, Bohman, Bollobás, and Zhao previously obtained bounds for these deficits with logarithmic gaps near both ends of the sequence $B_i^{(r)}$, namely, when $i=O(1)$ or $i=r-O(1)$. We close these gaps by showing that, as $r\to\infty$, for every fixed integer $a\ge1$, $q_{r,a+1}=Θ_a(r^{-a})$, and for every fixed integer $b\ge2$, $q_{r,r-b}=Θ_b(r^{-b}\log r)$.

math.CO

Upper Bounds on Turán Densities via Extremal Set Theory

We exhibit, in a systematic way, connections between hypergraph Turán problems and extremal set theory. More specifically, we construct natural families of uniform hypergraphs for which the upper bounds on their Turán densities reduce to classical problems in extremal set theory, including the Erdős--Ko--Rado theorem, $L$-intersecting families, and the Erdős matching problem.

math.CO

The inducibility of Turán graphs

Let $I(F,n)$ denote the maximum number of induced copies of a graph $F$ in an $n$-vertex graph. The inducibility of $F$, defined as $i(F)=\lim_{n\to \infty} I(F,n)/\binom{n}{v(F)}$, is a central problem in extremal graph theory. In this work, we investigate the inducibility of Turán graphs $F$. This topic has been extensively studied in the literature, including works of Pippenger--Golumbic, Brown--Sidorenko, Bollobás--Egawa--Harris--Jin, Mubayi, Reiher, and the first author, and Yuster. Broadly speaking, these results resolve or asymptotically resolve the problem when the part sizes of $F$ are either sufficiently large or sufficiently small (at most four). We complete this picture by proving that for every Turán graph $F$ and sufficiently large $n$, the value $I(F,n)$ is attained uniquely by the $m$-partite Turán graph on $n$ vertices, where $m$ is given explicitly in terms of the number of parts and vertices of $F$. This confirms a conjecture of Bollobás--Egawa--Harris--Jin from 1995, and we also establish the corresponding stability theorem. Moreover, we prove an asymptotic analogue for $I_{k+1}(F,n)$, the maximum number of induced copies of $F$ in an $n$-vertex $K_{k+1}$-free graph, thereby completely resolving a recent problem of Yuster. Finally, our results extend to a broader class of complete multipartite graphs in which the largest and smallest part sizes differ by at most on the order of the square root of the smallest part size.

math.CO

Odd hypergraph Mantel theorems

A classical result of Sidorenko (1989) shows that the Turán density of every $r$-uniform hypergraph with three edges is bounded from above by $1/2$. For even $r$, this bound is tight, as demonstrated by Mantel's theorem on triangles and Frankl's theorem on expanded triangles. In this note, we prove that for odd $r$, the bound $1/2$ is never attained, thereby answering a question of Keevash and revealing a fundamental difference between hypergraphs of odd and even uniformity. Moreover, our result implies that the expanded triangles form the unique class of three-edge hypergraphs whose Turán density attains $1/2$.

math.CO

A note on hypergraph extensions of Mantel's theorem

Chao and Yu introduced an entropy method for hypergraph Turán problems, and used it to show that the family of $\lfloor k/2\rfloor$ $k$-uniform tents have Turán density $k!/k^k$. Il'kovič and Yan improved this by reducing to a subfamily of $\lceil k/e\rceil$ tents. In this note, enhancing Il'kovič-Yan's result, we give a significantly shorter entropy proof, with optimal bounds within this framework.

math.CO

The general linear hypothesis testing problem for multivariate functional data with applications

As technology continues to advance at a rapid pace, the prevalence of multivariate functional data (MFD) has expanded across diverse disciplines, spanning biology, climatology, finance, and numerous other fields of study. Although MFD are encountered in various fields, the development of methods for hypotheses on mean functions, especially the general linear hypothesis testing (GLHT) problem for such data has been limited. In this study, we propose and study a new global test for the GLHT problem for MFD, which includes the one-way FMANOVA, post hoc, and contrast analysis as special cases. The asymptotic null distribution of the test statistic is shown to be a chi-squared-type mixture dependent of eigenvalues of the heteroscedastic covariance functions. The distribution of the chi-squared-type mixture can be well approximated by a three-cumulant matched chi-squared-approximation with its approximation parameters estimated from the data. By incorporating an adjustment coefficient, the proposed test performs effectively irrespective of the correlation structure in the functional data, even when dealing with a relatively small sample size. Additionally, the proposed test is shown to be root-n consistent, that is, it has a nontrivial power against a local alternative. Simulation studies and a real data example demonstrate finite-sample performance and broad applicability of the proposed test.

stat.ME

Modified Tests of Linear Hypotheses Under Heteroscedasticity for Multivariate Functional Data with Finite Sample Sizes

As big data continues to grow, statistical inference for multivariate functional data (MFD) has become crucial. Although recent advancements have been made in testing the equality of mean functions, research on testing linear hypotheses for mean functions remains limited. Current methods primarily consist of permutation-based tests or asymptotic tests. However, permutation-based tests are known to be time-consuming, while asymptotic tests typically require larger sample sizes to maintain an accurate Type I error rate. This paper introduces three finite-sample tests that modify traditional MANOVA methods to tackle the general linear hypothesis testing problem for MFD. The test statistics rely on two symmetric, nonnegative-definite matrices, approximated by Wishart distributions, with degrees of freedom estimated via a U-statistics-based method. The proposed tests are affine-invariant, computationally more efficient than permutation-based tests, and better at controlling significance levels in small samples compared to asymptotic tests. A real-data example further showcases their practical utility.

stat.ME

HDNRA: An R package for HDLSS location testing with normal-reference approaches

The challenge of location testing for high-dimensional data in statistical inference is notable. Existing literature suggests various methods, many of which impose strong regularity conditions on underlying covariance matrices to ensure asymptotic normal distribution of test statistics, leading to difficulties in size control. To address this, a recent set of tests employing the normal-reference approach has been proposed. Moreover, the availability of tests for high-dimensional location testing in R packages implemented in C++ is limited. This paper introduces the latest methods utilizing normal-reference approaches to test the equality of mean vectors in high-dimensional samples with potentially different covariance matrices. We present an R package named HDNRA to illustrate the implementation of these tests, extending beyond the two-sample problem to encompass general linear hypothesis testing (GLHT). The package offers easy and user-friendly access to these tests, with its core implemented in C++ using Rcpp, OpenMP and RcppArmadillo for efficient execution. Theoretical properties of these normal-reference tests are revisited, and examples based on real datasets using different tests are provided.

stat.AP

Uniquely colorable hypergraphs

An $r$-uniform hypergraph is uniquely $k$-colorable if there exists exactly one partition of its vertex set into $k$ parts such that every edge contains at most one vertex from each part. For integers $k \ge r \ge 2$, let $Φ_{k,r}$ denote the minimum real number such that every $n$-vertex $k$-partite $r$-uniform hypergraph with positive codegree greater than $Φ_{k,r} \cdot n$ and no isolated vertices is uniquely $k$-colorable. A classic result by of Bollobás\cite{Bol78} established that $Φ_{k,2} = \frac{3k-5}{3k-2}$ for every $k \ge 2$. We consider the uniquely colorable problem for hypergraphs. Our main result determines the precise value of $Φ_{k,r}$ for all $k \ge r \ge 3$. In particular, we show that $Φ_{k,r}$ exhibits a phase transition at approximately $k = \frac{4r-2}{3}$, a phenomenon not seen in the graph case. As an application of the main result, combined with a classic theorem by Frankl--Füredi--Kalai, we derive general bounds for the analogous problem on minimum positive $i$-degrees for all $1\leq i<r$, which are tight for infinitely many cases.

math.CO

A fast and accurate kernel-based independence test with applications to high-dimensional and functional data

Testing the dependency between two random variables is an important inference problem in statistics since many statistical procedures rely on the assumption that the two samples are independent. To test whether two samples are independent, a so-called HSIC (Hilbert--Schmidt Independence Criterion)-based test has been proposed. Its null distribution is approximated either by permutation or a Gamma approximation. In this paper, a new HSIC-based test is proposed. Its asymptotic null and alternative distributions are established. It is shown that the proposed test is root-n consistent. A three-cumulant matched chi-squared approximation is adopted to approximate the null distribution of the test statistic. By choosing a proper reproducing kernel, the proposed test can be applied to many different types of data including multivariate, high-dimensional, and functional data. Three simulation studies and two real data applications show that in terms of level accuracy, power, and computational cost, the proposed test outperforms several existing tests for multivariate, high-dimensional, and functional data.

stat.ME

Two-sample Behrens--Fisher problems for high-dimensional data: a normal reference F-type test

The problem of testing the equality of mean vectors for high-dimensional data has been intensively investigated in the literature. However, most of the existing tests impose strong assumptions on the underlying group covariance matrices which may not be satisfied or hardly be checked in practice. In this article, an F-type test for two-sample Behrens--Fisher problems for high-dimensional data is proposed and studied. When the two samples are normally distributed and when the null hypothesis is valid, the proposed F-type test statistic is shown to be an F-type mixture, a ratio of two independent chi-square-type mixtures. Under some regularity conditions and the null hypothesis, it is shown that the proposed F-type test statistic and the above F-type mixture have the same normal and non-normal limits. It is then justified to approximate the null distribution of the proposed F-type test statistic by that of the F-type mixture, resulting in the so-called normal reference F-type test. Since the F-type mixture is a ratio of two independent chi-square-type mixtures, we employ the Welch--Satterthwaite chi-square-approximation to the distributions of the numerator and the denominator of the F-type mixture respectively, resulting in an approximation F-distribution whose degrees of freedom can be consistently estimated from the data. The asymptotic power of the proposed F-type test is established. Two simulation studies are conducted and they show that in terms of size control, the proposed F-type test outperforms two existing competitors. The proposed F-type test is also illustrated by a real data example.

math.ST

Two-Sample Test for High-Dimensional Covariance Matrices: a normal-reference approach

Testing the equality of the covariance matrices of two high-dimensional samples is a fundamental inference problem in statistics. Several tests have been proposed but they are either too liberal or too conservative when the required assumptions are not satisfied which attests that they are not always applicable in real data analysis. To overcome this difficulty, a normal-reference test is proposed and studied in this paper. It is shown that under some regularity conditions and the null hypothesis, the proposed test statistic and a chi-square-type mixture have the same limiting distribution. It is then justified to approximate the null distribution of the proposed test statistic using that of the chi-square-type mixture. The distribution of the chi-square-type mixture can be well approximated using a three-cumulant matched chi-square-approximation with its approximation parameters consistently estimated from the data. The asymptotic power of the proposed test under a local alternative is also established. Simulation studies and a real data example demonstrate that in terms of size control, the proposed test outperforms the existing competitors substantially.

math.ST