SearcharxivSearch

arXiv subjects

Zichao Dong

Publications and source records attributed to Zichao Dong.

At least 19 recordsLinked to original sources

Set families: restricted distances via restricted intersections

Denote by $f_D(n)$ the maximum size of a set family $\mathcal{F}$ on $[n] = \{1, \dots, n\}$ with distance set $D$. That is, $|A \bigtriangleup B| \in D$ holds for every pair of distinct sets $A, B \in \mathcal{F}$. Kleitman's celebrated discrete isodiametric inequality states that $f_D(n)$ is maximized at Hamming balls of radius $d/2$ when $D = \{1, \dots, d\}$. We study the generalization where $D$ is a set of arithmetic progression and determine $f_D(n)$ asymptotically for all homogeneous $D$. In the special case when $D$ is an interval, our result confirms a conjecture of Huang, Klurman, and Pohoata. Moreover, we demonstrate a dichotomy in the growth of $f_D(n)$, showing linear growth in $n$ when $D$ is a non-homogeneous arithmetic progression. Different from previous combinatorial and spectral approaches, we deduce our results by converting the restricted distance problems to restricted intersection problems. Our proof ideas can be adapted to prove upper bounds on $t$-distance sets in Hamming cubes (also known as binary $t$-codes), which has been extensively studied by algebraic combinatorialists community, improving previous bounds from polynomial methods and optimization approaches.

math.CO

Sharp bounds on $k$-wise generalizations of oddtowns and eventowns

For $\boldsymbolα = (α_1, \dots, α_k) \in \mathbb{F}_2^k$, an $\boldsymbolα$-town is a set family in which every $i$-wise intersection has parity $α_i$. Denote by $f_{\boldsymbolα}(n)$ the maximum size of an $\boldsymbolα$-town on $[n]$. The classical oddtown and eventown problems study the cases $\boldsymbolα = (1, 0)$ and $(0, 0)$, respectively. We determine the sharp asymptotics of $f_{\boldsymbolα}(n)$ for all $\boldsymbolα$, answering questions of Johnston--O'Neill and Wei--Zhang--Ge. We also study a symmetric variant $g_{\boldsymbolα}(n)$, in which $i$-wise intersection sizes $|F_1 \cap \dots \cap F_i|$ are replaced by $i$-wise intersection-union sizes $|F_1 \cap \dots \cap F_i| + |F_1 \cup \dots \cup F_i|$.

math.CO

A dichotomy for hypergraph Zarankiewicz problems on axis-parallel boxes

We study the Zarankiewicz problem for $r$-partite, $r$-uniform intersection hypergraphs arising from $r$ families of axis-parallel boxes in $\mathbb{R}^d$ with prescribed directions $F_1, \dots, F_r \subseteq \{1, \dots, d\}$. This extends the problems studied by Chan and Har-Peled on points and $d$-dimensional boxes in $\mathbb{R}^d$, corresponding to $(F_1,F_2)=(\varnothing,[d])$, as well as by Chan, Keller, and Smorodinsky on $r$ families of $d$-dimensional boxes, corresponding to $(F_1,\dots,F_r)=([d],\dots,[d])$. Our main result establishes a sharp dichotomy for the Zarankiewicz number in this setting: it is either $Θ_r(tn^{r-1})$ or at least $Ω\bigl( tn^{r-1} \cdot \frac{\log n}{\log\log n} \bigr)$, depending only on a simple set-theoretic condition on $(F_1,\dots,F_r)$, which we call $2$-coherence. Informally, $2$-coherence captures whether the configuration contains an underlying two-dimensional incidence structure, which is precisely what gives rise to the extra polylogarithmic factor. Our proof proceeds via a sequence of reductions and a geometric slicing argument that reduces the problem to planar incidence bounds.

math.CO

Bipartite Turán problems via graph gluing

For graphs $H_1$ and $H_2$, if we glue them by identifying a given pair of vertices $u \in V(H_1)$ and $v \in V(H_2)$, what is the extremal number of the resulting graph $H_1^u \odot H_2^v$? In this paper, we study this problem and show that interestingly it is equivalent to an old question of Erdős and Simonovits on the Zarankiewicz problem. When $H_1, H_2$ are copies of a same bipartite graph $H$ and $u, v$ come from a same part, we prove that $\operatorname{ex}(n, H_1^u \odot H_2^v) = Θ\bigl( \operatorname{ex}(n, H) \bigr)$. As a corollary, we provide a short self-contained disproof of a conjecture of Erdős, which was recently disproved by Janzer.

math.CO

Saturation results around the Erdős--Szekeres problem

In this paper, we consider saturation problems related to the celebrated Erdős--Szekeres convex polygon problem. For each $n \ge 7$, we construct a planar point set of size $(7/8) \cdot 2^{n-2}$ which is saturated for convex $n$-gons. That is, the set contains no $n$ points in convex position while the addition of any new point creates such a configuration. This demonstrates that the saturation number is smaller than the Ramsey number for the Erdős--Szekeres problem. The proof also shows that the original Erdős--Szekeres construction is indeed saturated. Our construction is based on a similar improvement for the saturation version of the cups-versus-caps theorem. Moreover, we consider the generalization of the cups-versus-caps theorem to monotone paths in ordered hypergraphs. In contrast to the geometric setting, we show that this abstract saturation number is always equal to the corresponding Ramsey number.

math.CO

Large grid subsets without many cospherical points

Motivated by intuitions from projective algebraic geometry, we provide a novel construction of subsets of the $d$-dimensional grid $[n]^d$ of size $n - o(n)$ with no $d + 2$ points on a sphere or a hyperplane. For $d = 2$, this improves the previously best known lower bound of $n/4$ toward the Erdős--Purdy problem due to Thiele in 1995. For $d \ge 3$, this improves the recent $Ω\bigl( n^{\frac{3}{d+1}-o(1)} \bigr)$ bound due to Suk and White, confirming their conjectured $Ω\bigl( n^{\frac{d}{d+1}} \bigr)$ bound in a strong sense, and asymptotically resolves the generalized Erdős--Purdy problem posed by Brass, Moser, and Pach.

math.CO

Induced rational exponents and bipartite subgraphs in $K_{s, s}$-free graphs

In this paper, we study a general phenomenon that many extremal results for bipartite graphs can be transferred to the induced setting when the host graph is $K_{s, s}$-free. As manifestations of this phenomenon, we prove that every rational $\frac{a}{b} \in (1, 2), \, a, b \in \mathbb{N}_+$, can be achieved as Turán exponent of a family of at most $2^a$ induced forbidden bipartite graphs, extending a result of Bukh and Conlon [JEMS 2018]. Our forbidden family is a subfamily of theirs which is substantially smaller. A key ingredient, which is yet another instance of this phenomenon, is supersaturation results for induced trees and cycles in $K_{s, s}$-free graphs. We also provide new evidence to a recent conjecture of Hunter, Milojević, Sudakov, and Tomon [JCTB 2025] by proving optimal bounds for the maximum size of $K_{s, s}$-free graphs without an induced copy of theta graphs or prism graphs, whose Turán exponents were determined by Conlon [BLMS 2019] and by Gao, Janzer, Liu, and Xu [IJM 2025+].

math.CO

Convex polytopes in restricted point sets in $\mathbb{R}^d$

For a finite point set $P \subset \mathbb{R}^d$, denote by $\text{diam}(P)$ the ratio of the largest to the smallest distances between pairs of points in $P$. Let $c_{d, α}(n)$ be the largest integer $c$ such that any $n$-point set $P \subset \mathbb{R}^d$ in general position, satisfying $\text{diam}(P) < α\sqrt[d]{n}$, contains an $c$-point convex independent subset. We determine the asymptotics of $c_{d, α}(n)$ as $n \to \infty$ by showing the existence of positive constants $β= β(d, α)$ and $γ= γ(d)$ such that $βn^{\frac{d-1}{d+1}} \le c_{d, α}(n) \le γn^{\frac{d-1}{d+1}}$ for $α\geq 2$.

math.CO

Many cliques with small degree powers

Suppose $0 < p \le \infty$. For a simple graph $G$ with a vertex-degree sequence $d_1, \dots, d_n$ satisfying $(d_1^p + \dots + d_n^p)^{1/p} \le C$, we prove asymptotically sharp upper bounds on the number of $t$-cliques in $G$. This result bridges the $p = 1$ case, which is the notable Kruskal--Katona theorem, and the $p = \infty$ case, known as the Gan--Loh--Sudakov conjecture, and resolved by Chase. In particular, we demonstrate that the extremal construction exhibits a dichotomy between a single clique and multiple cliques at $p_0 = t - 1$. Our proof employs the entropy method.

math.CO

Empty red-red-blue triangles

Let $P$ be a $2n$-point set in the plane that is in general position. We prove that every red-blue bipartition of $P$ into $R$ and $B$ with $|R| = |B| = n$ generates $Ω(n^{3/2})$ red-red-blue empty triangles.

math.CO

MV-DETR: Multi-modality indoor object detection by Multi-View DEtecton TRansformers

We introduce a novel MV-DETR pipeline which is effective while efficient transformer based detection method. Given input RGBD data, we notice that there are super strong pretraining weights for RGB data while less effective works for depth related data. First and foremost , we argue that geometry and texture cues are both of vital importance while could be encoded separately. Secondly, we find that visual texture feature is relatively hard to extract compared with geometry feature in 3d space. Unfortunately, single RGBD dataset with thousands of data is not enough for training an discriminating filter for visual texture feature extraction. Last but certainly not the least, we designed a lightweight VG module consists of a visual textual encoder, a geometry encoder and a VG connector. Compared with previous state of the art works like V-DETR, gains from pretrained visual encoder could be seen. Extensive experiments on ScanNetV2 dataset shows the effectiveness of our method. It is worth mentioned that our method achieve 78\% AP which create new state of the art on ScanNetv2 benchmark.

cs.CV

LVIC: Multi-modality segmentation by Lifting Visual Info as Cue

Multi-modality fusion is proven an effective method for 3d perception for autonomous driving. However, most current multi-modality fusion pipelines for LiDAR semantic segmentation have complicated fusion mechanisms. Point painting is a quite straight forward method which directly bind LiDAR points with visual information. Unfortunately, previous point painting like methods suffer from projection error between camera and LiDAR. In our experiments, we find that this projection error is the devil in point painting. As a result of that, we propose a depth aware point painting mechanism, which significantly boosts the multi-modality fusion. Apart from that, we take a deeper look at the desired visual feature for LiDAR to operate semantic segmentation. By Lifting Visual Information as Cue, LVIC ranks 1st on nuScenes LiDAR semantic segmentation benchmark. Our experiments show the robustness and effectiveness. Codes would be make publicly available soon.

cs.CV

PeP: a Point enhanced Painting method for unified point cloud tasks

Point encoder is of vital importance for point cloud recognition. As the very beginning step of whole model pipeline, adding features from diverse sources and providing stronger feature encoding mechanism would provide better input for downstream modules. In our work, we proposed a novel PeP module to tackle above issue. PeP contains two main parts, a refined point painting method and a LM-based point encoder. Experiments results on the nuScenes and KITTI datasets validate the superior performance of our PeP. The advantages leads to strong performance on both semantic segmentation and object detection, in both lidar and multi-modal settings. Notably, our PeP module is model agnostic and plug-and-play. Our code will be publicly available soon.

cs.CV

Rainbow even cycles

We prove that every family of (not necessarily distinct) even cycles $D_1, \dotsc, D_{\lfloor 1.2(n-1) \rfloor+1}$ on some fixed $n$-vertex set has a rainbow even cycle (that is, a set of edges from distinct $D_i$'s, forming an even cycle). This resolves an open problem of Aharoni, Briggs, Holzman and Jiang. Moreover, the result is best possible for every positive integer $n$.

math.CO

HuBo-VLM: Unified Vision-Language Model designed for HUman roBOt interaction tasks

Human robot interaction is an exciting task, which aimed to guide robots following instructions from human. Since huge gap lies between human natural language and machine codes, end to end human robot interaction models is fair challenging. Further, visual information receiving from sensors of robot is also a hard language for robot to perceive. In this work, HuBo-VLM is proposed to tackle perception tasks associated with human robot interaction including object detection and visual grounding by a unified transformer based vision language model. Extensive experiments on the Talk2Car benchmark demonstrate the effectiveness of our approach. Code would be publicly available in https://github.com/dzcgaara/HuBo-VLM.

cs.RO

OG: Equip vision occupancy with instance segmentation and visual grounding

Occupancy prediction tasks focus on the inference of both geometry and semantic labels for each voxel, which is an important perception mission. However, it is still a semantic segmentation task without distinguishing various instances. Further, although some existing works, such as Open-Vocabulary Occupancy (OVO), have already solved the problem of open vocabulary detection, visual grounding in occupancy has not been solved to the best of our knowledge. To tackle the above two limitations, this paper proposes Occupancy Grounding (OG), a novel method that equips vanilla occupancy instance segmentation ability and could operate visual grounding in a voxel manner with the help of grounded-SAM. Keys to our approach are (1) affinity field prediction for instance clustering and (2) association strategy for aligning 2D instance masks and 3D occupancy instances. Extensive experiments have been conducted whose visualization results and analysis are shown below. Our code will be publicly released soon.

cs.CV

OVO: Open-Vocabulary Occupancy

Semantic occupancy prediction aims to infer dense geometry and semantics of surroundings for an autonomous agent to operate safely in the 3D environment. Existing occupancy prediction methods are almost entirely trained on human-annotated volumetric data. Although of high quality, the generation of such 3D annotations is laborious and costly, restricting them to a few specific object categories in the training dataset. To address this limitation, this paper proposes Open Vocabulary Occupancy (OVO), a novel approach that allows semantic occupancy prediction of arbitrary classes but without the need for 3D annotations during training. Keys to our approach are (1) knowledge distillation from a pre-trained 2D open-vocabulary segmentation model to the 3D occupancy network, and (2) pixel-voxel filtering for high-quality training data generation. The resulting framework is simple, compact, and compatible with most state-of-the-art semantic occupancy prediction models. On NYUv2 and SemanticKITTI datasets, OVO achieves competitive performance compared to supervised semantic occupancy prediction approaches. Furthermore, we conduct extensive analyses and ablation studies to offer insights into the design of the proposed framework. Our code is publicly available at https://github.com/dzcgaara/OVO.

cs.CV

A simple proof of the Gan-Loh-Sudakov conjecture

We give a new unified proof that any simple graph on $n$ vertices with maximum degree at most $Δ$ has no more than $a\binom{Δ+1}{t}+\binom{b}{t}$ cliques of size $t \ (t \ge 3)$, where $n = a(Δ+1)+b \ (0 \le b \le Δ)$.

math.CO