SearcharxivSearch

arXiv subjects

Cheng Chi

Publications and source records attributed to Cheng Chi.

At least 19 recordsLinked to original sources

StageWAM: Joint-Embedding Stage Prediction for World-Action Models in Robot Manipulation

Generalist robot policies aim to map multimodal observations and linguistic task instructions to actions across diverse tasks. However, existing methods typically represent the future as a fixed, short video-action chunk. This short-term future captures local scene evolution for action execution, but it does not explicitly describe the stage-level future that specifies how a task should progress from its current stage to the next. We therefore distinguish two complementary futures for robot manipulation: a short-term physical future to capture local scene evolution and a stage-level semantic future to represent task progress. We introduce StageWAM, which augments a Motus-based World Action Model (WAM) with Stage-JEPA, a goal-conditioned Joint-Embedding Predictive Architecture (JEPA) predictor. Given the current observation and task instruction, Stage-JEPA uses a frozen V-JEPA2 encoder to extract the current-state representation and predicts the latent target of the next inferred stage. Across 50 RoboTwin 2.0 tasks in clean and randomized environments, StageWAM achieves 90.25% overall success and reduces the mean number of execution steps in successful rollouts by 5.97% relative to the strongest baseline.

cs.RO

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling

Robust robot control benefits from explicitly modeling state transitions, but video-generation world action models (WAMs) introduce substantial deployment cost. Existing latent WAMs avoid explicit future generation, but often compress predictive representations or separate predictive modeling from the representations used for action generation. We introduce JEPA-WAM, a latent WAM built in a pretrained V-JEPA space, which couples latent transition prediction with continuous action generation through a shared predictor. JEPA-WAM predicts a spatially structured joint current-future target that captures task-shared visual temporal structure between current and future observations, while preserving dense patch-level correspondence. Through the shared predictor, transition supervision directly shapes the backbone, from which dedicated representations are extracted for action prediction. The same design can also be instantiated in pretrained VLA policies while preserving their original perception and action pathways. On LIBERO-Plus, JEPA-WAM achieves 79.2%, the best result without large-scale robot-policy pretraining, while its pretrained $\pi_{0.5}$ instantiation reaches 86.3%, achieving the best overall performance. Experiments on RoboTwin 2.0 and real-world bimanual manipulation further demonstrate strong generalization under visual and spatial shifts.

cs.RO

On a conjecture of Kolokolnikov on algebraic connectivity

For a graph $G$, let $\alpha(G)$ be the second smallest eigenvalue of the Laplacian matrix of $G$, also known as the algebraic connectivity. Algebraic connectivity plays an important role in characterizing the connectivity of graphs and convergence properties of networks. Kolokolnikov conjectured that among all graphs on $n$ vertices with exactly $2n-4$ edges, $\alpha(G)\leq 2$ and one of the maximizers is the complete bipartite graph whose two parts have sizes two and $n-2$, respectively. In this paper, we completely resolve this conjecture.

math.CO

MobileWAM: Bridging World Action Models to Mobile Manipulation with Chain-of-Foresight

World action models (WAMs) built on video generation backbones are a rising recipe for robot learning, yet remain confined to tabletop manipulation. Mobile manipulation demands simultaneous locomotion and whole-body manipulation amid scene-scale dynamics, yet is still dominated by dynamics-blind visual encoders with hand-crafted coordination. We bridge this gap with MobileWAM, a mixture-of-transformers architecture that fuses a pretrained video diffusion transformer with a lightweight action expert through layerwise joint attention, translating internet-scale motion priors into whole-body control. To reconcile the heterogeneous dynamics of moving and manipulating, each feed-forward layer of the action expert becomes a three-expert mixture of shared, locomotion, and manipulation experts, softly routed by the motion intent in the action tokens. To densify supervision, we further propose Chain-of-Foresight (CoF): intermediate representations sequentially predict a chain of future latent chunks, each step conditioned on its predecessor. CoF pairs naturally with our decoupled video--action denoising scheme. At deployment, the WAM serves as a pure current-frame encoder; foresight acts only through gradients, so at inference the foresight chain and video generation are discarded, leaving only policy-level cost. MobileWAM surpasses state-of-the-art mobile manipulation policies on ManiSkill-HAB and fine-tunes to a real ARX Lift2 mobile manipulator across diverse tasks with strong generalization. Code will be released soon.

cs.CV

On three open problems in zero-sum Ramsey numbers

Let $K_N^{(r)}$ denote the $N$-vertex complete $r$-uniform hypergraph. For an $r$-uniform hypergraph $H$ and an integer $k\geq2$, the $k$-color Ramsey number $R(H,k)$ is the least integer $N$ such that every $k$-edge-coloring of $K_N^{(r)}$ contains a monochromatic copy of $H$. When $k\mid\esize(H)$, the zero-sum Ramsey number $R(H,\mathbb Z_k)$ is the least integer $N$ such that every edge-labeling of $K_N^{(r)}$ by elements of $\mathbb Z_k$ contains a copy of $H$ whose edge labels sum to $0$ in $\mathbb Z_k$. We settle two conjectures and a problem concerning these two Ramsey numbers. First, Caro and Provstgaard proposed exact values for the zero-sum Ramsey numbers over $\mathbb Z_2$ of delta-systems with an even number of edges. We determine these numbers and thereby prove their conjecture. Second, for a forest $F$ with $m$ edges, let $tF$ denote the disjoint union of $t$ copies of $F$. Caro conjectured that $R(tF,\mathbb Z_{mt})=R(tF,2)$ for all sufficiently large $t$. We show that this conjecture does not hold for double stars. Caro also asked whether there exists a tree $T$ with $m$ edges such that $R(T,\mathbb Z_m)>R(T,2)$. We answer this question affirmatively by constructing an infinite family of such trees.

math.CO

Extremal Families for the Erd\H{o}s--Kleitman Problem: The Missing Constructions

For integers $n\ge s\ge2$, let $e(n,s)$ be the maximum size of a family $\mathcal F\subseteq2^{[n]}$ with no $s$ pairwise disjoint members. The problem of determining $e(n,s)$, now called the Erd\H{o}s--Kleitman problem, is closely related to the well-known Erd\H{o}s matching problem. Frankl and Kupavskii posed a meta-conjecture predicting that the maximum is always attained by a weighted family. Fix $m\ge3$, write $n=ms+c$ with $0\le c 0$, $\beta=\beta(m,k)>0$ and an integer $s_0=s_0(m,k)$ such that, for all integers $s\ge s_0$ and all integers $c$ with $0\le c<s$, the only extremal families for $e(n,s)$ are the families $\mathcal H^k(m,s,\ell;A)$ with $A\in\binom{[n]}{a_k}$ whenever $\beta s^{(k-1)/k}\le c\le \alpha s^{k/(k+1)}$. In particular, this result determines an infinite number of new extremal families for the Erd\H{o}s--Kleitman problem and verifies the Frankl--Kupavskii meta-conjecture in these ranges. This also provides a quantitative extension of the result of Kupavskii and Sokolov on the extremality of $\mathcal H^1$.

math.CO

A note on zero-sum Ramsey numbers of complete graphs

For a graph $H$ with $3\mid e(H)$, the zero-sum Ramsey number $R(H,\Z_3)$ is the least integer $N$ such that every labeling of the edges of $K_N$ by elements of $\Z_3$ contains a copy of $H$ whose edge labels sum to zero. We determine the last previously unresolved infinite family in the complete-graph case modulo $3$. More precisely, we prove that \(R(K_n,\Z_3)=n+3\) for every $n\ge 10$ satisfying $n\equiv 1\pmod 3$. Consequently, for $k\ge 1$, \(R(K_{9k+7},\Z_3)=9k+10\), resolving a problem of Caro and Mifsud.

math.CO

On Zero-sum Ramsey numbers of complete bipartite graphs

For an integer $q\ge 2$ and a graph $F$ satisfying $q\mid e(F)$, the zero-sum Ramsey number $R(F,\mathbb Z_q)$ is the least integer $n$ such that every edge-labeling $w\colon E(K_n)\to \mathbb Z_q$ contains a copy of $F$ whose edge-label sum is zero in $\mathbb Z_q$. Write $K_{s,t}$ for the complete bipartite graph with $s$ vertices on one side and $t$ vertices on the other side. We prove that for every $q\ge2$, there is an explicit threshold $S(q)$ such that $R(K_{s,qk},\mathbb Z_q)=s+qk$ for all $s\ge S(q)$ and all $k\ge1$. We also determine the zero-sum Ramsey number of $K_{s,3k}$ over $\mathbb Z_3$ for all $s\ge2$ and $k\ge1$. We prove that $R(K_{s,3k},\mathbb Z_3)=s+3k$, except when $s=2$ and $k\ge1$, or when $s\in\{3,4,5,7\}$ and $k=1$. In these exceptional cases, $R(K_{s,3k},\mathbb Z_3)=s+3k+1$. In particular, this shows that the threshold $S(q)$ is best possible for \(q=3\).

math.CO

Xiaomi Auto World Model: A Joint World Model Integrating Reconstruction and Generation for Autonomous Driving

This report presents a unified technical system addressing the two core capabilities of world models for autonomous driving: world representation and world generation. For world representation, we propose WorldRec, a feed-forward reconstruction architecture driven by sparse scene queries. WorldRec initializes structured queries in 3D space, leveraging them to aggregate cross-view, cross-temporal features, thereby naturally enforcing spatial consistency across frames and yielding compact yet high-fidelity 3D Gaussian scene representations. For world generation, we propose WorldGen, a two-stage training framework of bidirectional pretraining followed by causal fine-tuning through three progressive stages (Teacher Forcing, ODE distillation, and DMD), enabling high-quality online causal video generation in as few as 4 denoising steps. Building on both modules, we further introduce the JWM, which deeply integrates WorldRec and WorldGen to achieve synergistic gains in generation stability, cross-frame consistency, and visual fidelity, providing a solid foundation for closed-loop simulation, data synthesis, and end-to-end training in autonomous driving.

cs.CV

On zero-sum Ramsey numbers of cycles and wheels

For an integer $q\ge 2$ and a graph $F$ with $q\mid e(F)$, let $R(F,\Z_q)$ be the least integer $n$ such that every edge-labeling $w\colon E(K_n)\to \Z_q$ contains a copy of $F$ whose edge-label sum is zero in $\Z_q$. Write $C_{qk}$ for the cycle on $qk$ vertices. We prove that $R(C_{qk},\Z_q)\le \max\{R(C_{2q},\Z_q),qk+q-1\}$ via an insertion argument rooted in the classic Erd\H{o}s-Ginzburg-Ziv theorem. Combined with Pikhurko's result, we obtain $R(C_{qk},\Z_q)\le \max\{35q^2,qk+q-1\}$ for every $q\ge 3$. We also show that $R(C_{qk},\Z_q)\ge qk+q-1$ for odd $q\ge 3$. Hence, for every fixed odd $q\ge 3$ and every $k\ge 35q$, we obtain the exact value $R(C_{qk},\Z_q)=qk+q-1$. For even $q\ge 4$, the same method gives $qk+\frac q2-1\le R(C_{qk},\Z_q)\le \max\{35q^2,qk+q-1\}$, leaving an additive gap of order $q/2$ when $k$ is large. Moreover, for the case $q=3$, we prove that \(R(C_{3k}, \mathbb{Z}_3) = 3k + 2\) for all \(k \ge 2\). Extending our techniques beyond cycles, we also resolve the zero-sum Ramsey number for wheel graphs \(W_m = C_m + K_1\), proving that \(R(W_{3k}, \mathbb{Z}_3) = 3k + 1\) for all \(k \ge 2\).

math.CO

PointForward: Feedforward Driving Reconstruction through Point-Aligned Representations

High-fidelity reconstruction of driving scenes is crucial for autonomous driving. While recent feedforward 3D Gaussian Splatting (3DGS) methods enable fast reconstruction, their per-pixel Gaussian prediction paradigm often suffers from multi-view inconsistency and layering artifacts. Moreover, existing methods often model dynamic instances via dense flow prediction, which lacks explicit cross-view correspondence and instance-level consistency. In this paper, we propose PointForward, a feedforward driving reconstruction framework through point-aligned representations. Unlike pixel-aligned methods, we initialize sparse 3D queries in world space and aggregate multi-view image information via spatial-temporal fusion onto these queries, enforcing explicit cross-view consistency in a single feedforward pass. To handle scene dynamics, we introduce scene graphs that explicitly organize moving instances during reconstruction. By leveraging 3D bounding boxes, our method enables instance-level motion propagation and temporally consistent dynamic representations. Extensive experiments demonstrate that PointForward achieves state-of-the-art performance on large-scale driving benchmarks. The code will be available upon the publication of the paper.

cs.CV

New Extremal Ranges and Constructions of the Erd\H{o}s--Kleitman Problem

For integers $n\ge s\ge2$, let $e(n,s)$ denote the maximum size of a family $\mathcal F\subseteq2^{[n]}$ with no $s$ pairwise disjoint members. The problem of determining $e(n,s)$, now called the Erd\H{o}s--Kleitman problem, is the non-uniform analogue of the Erd\H{o}s matching conjecture. We prove that for every fixed $m\ge3$, there exist constants $\beta_m$ and $\delta_m$ such that for sufficiently large $s$, the extremal families for $e(ms+c,s)$ are \[ \mathcal P'(m,s,\ell;L'):=\binom{L'}m\cup\binom{[ms+c]}{\ge m+1} \] for some $L'$ with $\ell=s-c$ and $|L'|=m\ell-1$, when $\beta_m s^{(m-1)/m}\le c\le \delta_m s$. This determines the extremal families in an unknown range when $\ell$ is large, complementing our earlier work on the range when $\ell$ is small. Moreover, for $m=3$, we sharpen this to the asymptotically optimal range. Let \[ t(s)=\frac{17-18s+\sqrt{49-852s+1284s^2}}{20}=0.8916\cdots s+O(1) \] We prove that \(\mathcal P'(3,s,\ell;L')\) is the unique extremal family when $t(s)<\ell<s-((4/3)^{1/3}+o(1))s^{2/3}$. Note that the lower bound \(t(s)\) of $\ell$ is exact, while the the constant \((4/3)^{1/3}\) in the upper bound of $\ell$ is best possible. Kupavskii and Sokolov introduced four candidate extremal families and conjectured that the value of $e(n,s)$ is the maximum of their sizes. We disprove this conjecture by constructing a new family $\mathcal R(m,s,\ell)$ that is larger than each of their four proposed candidates when $\alpha_{\mathrm R}s^{1/2}\le c\le \beta_{\mathrm R}s^{(m-1)/m}$ for some constants $\alpha_{\mathrm R}$ and $\beta_{\mathrm R}$. This also shows that the exponent $(m-1)/m$ in the first result is tight.

math.CO

A solution to Frankl and Kupavskii's conjecture concerning Erd\H{o}s-Kleitman matching problem

For integers $n\ge s\ge2$, let $e(n,s)$ be the maximum size of a family $\mathcal F\subseteq2^{[n]}$ with no $s$ pairwise disjoint members. The study of determining $e(n,s)$ is closely related to its uniform counterpart, the well-known Erd\H{o}s matching conjecture. Frankl and Kupavskii conjectured an exact formula for $e((m+1)s-\ell,s)$ when $1\le \ell\le \lceil s/2\rceil$. We prove that for every fixed $m\ge3$ and sufficiently large $s$, the extremal families for $e((m+1)s-\ell,s)$ are $P(m,s,\ell;L)\coloneqq\{A\subseteq [n]\colon |A|+|A\cap L|\ge m+1\}$ for some $L$ with $|L|=\ell-1$ when $1\le \ell\le (\frac{m+1}{2m+1}-o(1))s$. In particular, this confirms the Frankl--Kupavskii conjecture for every fixed $m\ge3$ and all sufficiently large $s$. For $m=3$, we determine the whole range of $\ell$ for which $P(3,s,\ell;L)$ is extremal, generalizing a theorem of Kupavskii and Sokolov.

math.CO

Anti-Ramsey numbers for cancellative configurations in p-graphs

We study edge-colorings of the complete $p$-graph on $n$ vertices that contain no three edges $A,B,C$ of distinct colors such that the symmetric difference of $A$ and $B$ is contained in $C$. For $p\ge3$ and $n\ge p+1$, we show that every such coloring contains at most $1+\floor{n/p}$ colors and characterize the extremal colorings, generalizing a theorem of Erd\H{o}s, Simonovits and S\'os. %\cite{erdos1975}. When $p=3$, the condition $A\triangle B\subseteq C$ implies $|A\triangle B|=2$, and the three edges necessarily form a copy of $F_4\coloneqq\{abc,abd,bcd\}$ or $F_5\coloneqq\{abc,abd,cde\}$. For $n\ge5$, we show that every rainbow $F_5$-free edge-coloring is rainbow cancellative. For rainbow $F_4$-free colorings, we construct colorings with $m(n)+1$ colors for all $n\ge4$, where $m(n)$ is the size of a maximum partial Steiner triple system of order $n$ and satisfies $m(n)=n^2/6+O(n)$, improving the linear lower bound by Budden and Stiles. %\cite{budden}. Moreover, for $n=2^s-1$, we obtain $\ar(n,F_4)\ge m(n)+n^2/42+o(n^2)=4n^2/21+o(n^2)$ via a construction based on independent sets in the Grassmann graph. We also prove that $\ar(n,F_4)\le (5n^2-8n)/21$ for $n\ge4$, improving the quadratic coefficient in the upper bound of Budden and Stiles from $1/4$ to $5/21$.

math.CO

The Tur\'{a}n number of the Cartesian product of a star and an edge

Let $C_k$ denote the cycle of length $k$, $S_t$ be a star with $t$ edges. And let $B_t$ be the graph consisting of $t$ copies of $C_4$ sharing one fixed edge. Equivalently, $B_t=K_2 \mathbin{\square} S_t$, which is the Cartesian product of a star with $t$ edges and an edge. Recently, Gao, Janzer, Liu and Xu [\textit{Israel J. Math. 269(2025)}] proved that the Tur\'an number of $K_2\mathbin{\square} C_{2l}$ is $\Theta(n^{\frac{3}{2}})$ for every $l\ge 4$. In this paper, we obtain upper and lower estimates for the Tur\'an number of $B_t$ in both the general and bipartite settings for every $t\geq 2$. For the lower bound, we use random construction based on the extremal structure of $C_4$. These results imply that $\frac{1}{2\sqrt{2}}\leq \lim_{t\to \infty} \frac{\mathrm{ex}(n,B_t)}{\sqrt{t}}\leq \frac{1}{2}$, and $\frac{1}{4}\leq \lim_{t\to \infty} \frac{\mathrm{ex}_{bip}(n,B_t)}{\sqrt{t}}\leq \frac{1}{2\sqrt{2}}.$ In the case of $B_2$, we obtain sharper estimates. We show that the Tur\'an number of $B_2$ is approximately between $(0.518+o(1))n^{\frac{3}{2}}$ and $(0.603+o(1))n^{\frac{3}{2}}$. And in the bipartite setting, it is approximately between $(0.385+o(1))n^{\frac{3}{2}}$ and $(0.468+o(1))n^{\frac{3}{2}}$. Moreover, in the bipartite setting, we give a more general result, which shows that for every tree $T$ with $t$ edges, the bipartite Tur\'an number of $K_2\mathbin{\square}T$ is at most $\frac{\sqrt{t}}{2\sqrt{2}}(1+o(1))n^{\frac{3}{2}}$.

math.CO

UniDA3D: A Unified Domain-Adaptive Framework for Multi-View 3D Object Detection

Camera-only 3D object detection is critical for autonomous driving, offering a cost-effective alternative to LiDAR based methods. In particular, multi-view 3D object detection has emerged as a promising direction due to its balanced trade-off between performance and cost. However, existing methods often suffer significant performance degradation under complex environmental conditions such as nighttime, fog, and rain, primarily due to their reliance on training data collected mostly in ideal conditions. To address this challenge, we propose UniDA3D, a unified domain-adaptive multi-view 3D object detector designed for robust perception under diverse adverse conditions. UniDA3D formulates nighttime, rainy, and foggy scenes as a unified multi target domain adaptation problem and leverages a novel query guided domain discrepancy mitigation (QDDM) module to align object features between source and target domains at both batch and global levels via query-centric adversarial and contrastive learning. Furthermore, we introduce a domain-adaptive teacher student training pipeline with an exponential-moving-average teacher and dynamically updated high-quality pseudo labels to enhance consistency learning and suppress background noise in unlabeled target domains. In contrast to prior approaches that require separate training for each condition, UniDA3D performs a single unified training process across multiple domains, enabling robust all-weather 3D perception. On a synthesized multi-view 3D benchmark constructed by generating nighttime, rainy, and foggy counterparts from nuScenes (nuScenes-Night, nuScenes-Rain, and nuScenes-Haze), UniDA3D consistently outperforms state of-the-art camera-only multi-view 3D detectors under extreme conditions, achieving substantial gains in mAP and NDS while maintaining real-time inference efficiency.

cs.CV

PRM-as-a-Judge: A Dense Evaluation Paradigm for Fine-Grained Robotic Auditing

Current robotic evaluation is still largely dominated by binary success rates, which collapse rich execution processes into a single outcome and obscure critical qualities such as progress, efficiency, and stability. To address this limitation, we propose PRM-as-a-Judge, a dense evaluation paradigm that leverages Process Reward Models (PRMs) to audit policy execution directly from trajectory videos by estimating task progress from observation sequences. Central to this paradigm is the OPD (Outcome-Process-Diagnosis) metric system, which explicitly formalizes execution quality via a task-aligned progress potential. We characterize dense robotic evaluation through two axiomatic properties: macro-consistency, which requires additive and path-consistent aggregation, and micro-resolution, which requires sensitivity to fine-grained physical evolution. Under this formulation, potential-based PRM judges provide a natural instantiation of dense evaluation, with macro-consistency following directly from the induced scalar potential. We empirically validate the micro-resolution property using RoboPulse, a diagnostic benchmark specifically designed for probing micro-scale progress discrimination, where several trajectory-trained PRM judges outperform discriminative similarity-based methods and general-purpose foundation-model judges. Finally, leveraging PRM-as-a-Judge and the OPD metric system, we conduct a structured audit of mainstream policy paradigms across long-horizon tasks, revealing behavioral signatures and failure modes that are invisible to outcome-only metrics.

cs.RO

LassoFlexNet: Flexible Neural Architecture for Tabular Data

Despite their dominance in vision and language, deep neural networks often underperform relative to tree-based models on tabular data. To bridge this gap, we incorporate five key inductive biases into deep learning: robustness to irrelevant features, axis alignment, localized irregularities, feature heterogeneity, and training stability. We propose \emph{LassoFlexNet}, an architecture that evaluates the linear and nonlinear marginal contribution of each input via Per-Feature Embeddings, and sparsely selects relevant variables using a Tied Group Lasso mechanism. Because these components introduce optimization challenges that destabilize standard proximal methods, we develop a \emph{Sequential Hierarchical Proximal Adaptive Gradient optimizer with exponential moving averages (EMA)} to ensure stable convergence. Across $52$ datasets from three benchmarks, LassoFlexNet matches or outperforms leading tree-based models, achieving up to a $10$\% relative gain, while maintaining Lasso-like interpretability. We substantiate these empirical results with ablation studies and theoretical proofs confirming the architecture's enhanced expressivity and structural breaking of undesired rotational invariance.

stat.ML