SearcharxivSearch

arXiv subjects

Xiaochen Zhao

Publications and source records attributed to Xiaochen Zhao.

At least 19 recordsLinked to original sources

Largest Sperner families with restricted differences

Let $L$ be a fixed set of positive integers. A family $\mathcal{F}\subseteq 2^{[n]}$ is called $L$-differencing if $\lvert A\setminus B\rvert\in L$ for every ordered pair of distinct members $A,B\in\mathcal{F}$. A longstanding conjecture of Frankl, proposed in 1985, asserts that every $L$-differencing family has size at most $\binom{n}{|L|}$. We resolve this conjecture asymptotically for every fixed $L$, and obtain the exact answer in the only case in which the conjectured bound could be tight. (1) If $L\ne [s]$ and $n$ is large, then every $L$-differencing family satisfies $\lvert \mathcal{F}\rvert \le \left(\frac{s}{s+1}+o_L(1)\right)\binom{n}{s}$. (2) If $L=[s]$ and $n\ge 2s-1$, then $\lvert \mathcal{F}\rvert\le\binom{n}{s}$, with equality only for $\binom{[n]}{s}$ and $\binom{[n]}{n-s}$. The first result follows by reducing directed differences to restricted Hamming distances. For the exact result, we develop a new homogeneous polynomial method, which might be of independent interest.

math.CO

Recursive Lifting Beyond the Ahlswede--Khachatrian Construction

For the Erd\H{o}s--Frankl--Pach problem on uniform set systems of bounded VC-dimension, the Ahlswede--Khachatrian/Mubayi--Zhao construction has long served as the standard lower-bound benchmark. We develop a recursive lifting method that goes beyond this benchmark in every dimension \(d\ge3\), proving that for every \(d\ge3\) and \(n\ge d+3\), \[ M_d(n)\ge \binom{n-1}{d}+\binom{n-4}{d-2}+M_{d-3}(n-5). \] The proof is elementary and proceeds through explicit trace obstructions. We also record a further recursive improvement in the concluding remarks.

math.CO

Recursive lower bounds for uniform set systems of bounded VC-dimension

For integers $n\ge d+1$, let $\mathsf{M}_d(n)$ denote the maximum size of a $(d+1)$-uniform family on an $n$-element ground set with VC-dimension at most $d$. For $n\ge2d+2$, the classical construction of Ahlswede and Khachatrian, later generalized by Mubayi and Zhao, gives \[ \mathsf{M}_d(n)\ge \binom{n-1}{d}+\binom{n-4}{d-2}. \] We introduce a two-cover lifting construction and prove the recursive lower bound \[ \mathsf{M}_d(n)\ge \binom{n-1}{d}+\binom{n-4}{d-2}+\mathsf{M}_{d-3}(n-5) \] for every $d\ge 3$ and $n\ge d+3$. Consequently, \[ \mathsf{M}_d(n)\ge \binom{n-1}{d}+\binom{n-4}{d-2}+\binom{n-6}{d-3}. \] Thus the Mubayi--Zhao conjecture on the exact value of $\mathsf{M}_d(n)$ for $n\ge2(d+2)$ is false for any $d\ge 3$. The proof is elementary and proceeds entirely through an explicit analysis of traces.

math.CO

Universal Fourier-inertia bounds for prescribed even distances

The study of set families with restricted Hamming distances is a classical topic of extremal combinatorics and coding theory. Let \(H=\{A\subseteq[n]: |A|\text{ is even}\}\) be the even subcube. Let \(\ell_1,\dots,\ell_t\) be distinct positive integers and set \(\mathcal L=\{2\ell_1,\dots,2\ell_t\}\). We prove that, for all sufficiently large \(n\), every family \(\mathcal F\subseteq H\) satisfying \( |A\triangle B|\in\mathcal L \) for all \(A\ne B\in\mathcal F\) has \[ |\mathcal F|\le \sum_{i=0}^{t}\binom{n-1}{i}. \] This is best possible as a universal bound, with equality attained at the distance set \(\mathcal L=\{2,4,\dots,2t\}\). Our proof uses a Fourier-inertia argument based on a universal low/high boundary-layer sign pattern for the Fourier transform of the distance-polynomial kernel on the even subcube: the prescribed distances enter only through lower-order Fourier terms, while the leading boundary-layer signs depend solely on \(t\). This even-subcube result immediately yields an odd-subcube analogue and, through parity reductions, provides bounds for arbitrary distance sets. In particular, this approach recovers the classical interval bounds of Kleitman and the corresponding interval bounds of Huang--Klurman--Pohoata, while offering a direct spectral proof of the maximality of \(\{2,4,\dots,2t\}\) among all fixed \(t\)-distance sets.

math.CO

Monochromatic unit equilateral triangle on low-dimensional spheres

A result of Matou\v{s}ek and R\"odl in 1995 states that for every $\varepsilon>0$ and every triangle $T$ with circumradius $\rho(T)$, there exists a dimension $n=n(\varepsilon,T)$ such that every $2$-coloring of the $n$-dimensional sphere of radius $\rho(T)+\varepsilon$, namely $\mathbb{S}^{n}(\rho(T)+\varepsilon)$, contains a monochromatic congruent copy of $T$. In this paper, we determine the exact threshold dimension for the unit equilateral triangle on the sphere $\mathbb{S}^{n}(1/\sqrt{2})$: there exists a $2$-coloring of $\mathbb{S}^{2}(1/\sqrt{2})$ with no monochromatic unit equilateral triangle, whereas every $2$-coloring of $\mathbb{S}^{3}(1/\sqrt{2})$ contains one. Along the way, we also establish several further Euclidean Ramsey-type results on low-dimensional spheres, including asymmetric and isosceles variants.

math.CO

GeoDiff4D: Geometry-Aware Diffusion for 4D Head Avatar Reconstruction

Reconstructing photorealistic and animatable 4D head avatars from a single portrait image remains a fundamental challenge in computer vision. While diffusion models have enabled remarkable progress in image and video generation for avatar reconstruction, existing methods primarily rely on 2D priors and struggle to achieve consistent 3D geometry. We propose a novel framework that leverages geometry-aware diffusion to learn strong geometry priors for high-fidelity head avatar reconstruction. Our approach jointly synthesizes portrait images and corresponding surface normals, while a pose-free expression encoder captures implicit expression representations. Both synthesized images and expression latents are incorporated into 3D Gaussian-based avatars, enabling photorealistic rendering with accurate geometry. Extensive experiments demonstrate that our method substantially outperforms state-of-the-art approaches in visual quality, expression fidelity, and cross-identity generalization, while supporting real-time rendering.

cs.CV

HyMem: Hybrid Memory Architecture with Dynamic Retrieval Scheduling

Large language model (LLM) agents demonstrate strong performance in short-text contexts but often underperform in extended dialogues due to inefficient memory management. Existing approaches face a fundamental trade-off between efficiency and effectiveness: memory compression risks losing critical details required for complex reasoning, while retaining raw text introduces unnecessary computational overhead for simple queries. The crux lies in the limitations of monolithic memory representations and static retrieval mechanisms, which fail to emulate the flexible and proactive memory scheduling capabilities observed in humans, thus struggling to adapt to diverse problem scenarios. Inspired by the principle of cognitive economy, we propose HyMem, a hybrid memory architecture that enables dynamic on-demand scheduling through multi-granular memory representations. HyMem adopts a dual-granular storage scheme paired with a dynamic two-tier retrieval system: a lightweight module constructs summary-level context for efficient response generation, while an LLM-based deep module is selectively activated only for complex queries, augmented by a reflection mechanism for iterative reasoning refinement. Experiments show that HyMem achieves strong performance on both the LOCOMO and LongMemEval benchmarks, outperforming full-context while reducing computational cost by 92.6\%, establishing a state-of-the-art balance between efficiency and performance in long-term memory management.

cs.AI

Plan-X: Instruct Video Generation via Semantic Planning

Diffusion Transformers have demonstrated remarkable capabilities in visual synthesis, yet they often struggle with high-level semantic reasoning and long-horizon planning. This limitation frequently leads to visual hallucinations and mis-alignments with user instructions, especially in scenarios involving complex scene understanding, human-object interactions, multi-stage actions, and in-context motion reasoning. To address these challenges, we propose Plan-X, a framework that explicitly enforces high-level semantic planning to instruct video generation process. At its core lies a Semantic Planner, a learnable multimodal language model that reasons over the user's intent from both text prompts and visual context, and autoregressively generates a sequence of text-grounded spatio-temporal semantic tokens. These semantic tokens, complementary to high-level text prompt guidance, serve as structured "semantic sketches" over time for the video diffusion model, which has its strength at synthesizing high-fidelity visual details. Plan-X effectively integrates the strength of language models in multimodal in-context reasoning and planning, together with the strength of diffusion models in photorealistic video synthesis. Extensive experiments demonstrate that our framework substantially reduces visual hallucinations and enables fine-grained, instruction-aligned video generation consistent with multimodal context.

cs.CV

Enhanced Self-Distillation Framework for Efficient Spiking Neural Network Training

Spiking Neural Networks (SNNs) exhibit exceptional energy efficiency on neuromorphic hardware due to their sparse activation patterns. However, conventional training methods based on surrogate gradients and Backpropagation Through Time (BPTT) not only lag behind Artificial Neural Networks (ANNs) in performance, but also incur significant computational and memory overheads that grow linearly with the temporal dimension. To enable high-performance SNN training under limited computational resources, we propose an enhanced self-distillation framework, jointly optimized with rate-based backpropagation. Specifically, the firing rates of intermediate SNN layers are projected onto lightweight ANN branches, and high-quality knowledge generated by the model itself is used to optimize substructures through the ANN pathways. Unlike traditional self-distillation paradigms, we observe that low-quality self-generated knowledge may hinder convergence. To address this, we decouple the teacher signal into reliable and unreliable components, ensuring that only reliable knowledge is used to guide the optimization of the model. Extensive experiments on CIFAR-10, CIFAR-100, CIFAR10-DVS, and ImageNet demonstrate that our method reduces training complexity while achieving high-performance SNN training. Our code is available at https://github.com/Intelli-Chip-Lab/enhanced-self-distillation-framework-for-snn.

cs.CV

X-Streamer: Unified Human World Modeling with Audiovisual Interaction

We introduce X-Streamer, an end-to-end multimodal human world modeling framework for building digital human agents capable of infinite interactions across text, speech, and video within a single unified architecture. Starting from a single portrait, X-Streamer enables real-time, open-ended video calls driven by streaming multimodal inputs. At its core is a Thinker-Actor dual-transformer architecture that unifies multimodal understanding and generation, turning a static portrait into persistent and intelligent audiovisual interactions. The Thinker module perceives and reasons over streaming user inputs, while its hidden states are translated by the Actor into synchronized multimodal streams in real time. Concretely, the Thinker leverages a pretrained large language-speech model, while the Actor employs a chunk-wise autoregressive diffusion model that cross-attends to the Thinker's hidden states to produce time-aligned multimodal responses with interleaved discrete text and audio tokens and continuous video latents. To ensure long-horizon stability, we design inter- and intra-chunk attentions with time-aligned multimodal positional embeddings for fine-grained cross-modality alignment and context retention, further reinforced by chunk-wise diffusion forcing and global identity referencing. X-Streamer runs in real time on two A100 GPUs, sustaining hours-long consistent video chat experiences from arbitrary portraits and paving the way toward unified world modeling of interactive digital humans.

cs.CV

X-UniMotion: Animating Human Images with Expressive, Unified and Identity-Agnostic Motion Latents

We present X-UniMotion, a unified and expressive implicit latent representation for whole-body human motion, encompassing facial expressions, body poses, and hand gestures. Unlike prior motion transfer methods that rely on explicit skeletal poses and heuristic cross-identity adjustments, our approach encodes multi-granular motion directly from a single image into a compact set of four disentangled latent tokens -- one for facial expression, one for body pose, and one for each hand. These motion latents are both highly expressive and identity-agnostic, enabling high-fidelity, detailed cross-identity motion transfer across subjects with diverse identities, poses, and spatial configurations. To achieve this, we introduce a self-supervised, end-to-end framework that jointly learns the motion encoder and latent representation alongside a DiT-based video generative model, trained on large-scale, diverse human motion datasets. Motion-identity disentanglement is enforced via 2D spatial and color augmentations, as well as synthetic 3D renderings of cross-identity subject pairs under shared poses. Furthermore, we guide motion token learning with auxiliary decoders that promote fine-grained, semantically aligned, and depth-aware motion embeddings. Extensive experiments show that X-UniMotion outperforms state-of-the-art methods, producing highly expressive animations with superior motion fidelity and identity preservation.

cs.CV

X-Actor: Emotional and Expressive Long-Range Portrait Acting from Audio

We present X-Actor, a novel audio-driven portrait animation framework that generates lifelike, emotionally expressive talking head videos from a single reference image and an input audio clip. Unlike prior methods that emphasize lip synchronization and short-range visual fidelity in constrained speaking scenarios, X-Actor enables actor-quality, long-form portrait performance capturing nuanced, dynamically evolving emotions that flow coherently with the rhythm and content of speech. Central to our approach is a two-stage decoupled generation pipeline: an audio-conditioned autoregressive diffusion model that predicts expressive yet identity-agnostic facial motion latent tokens within a long temporal context window, followed by a diffusion-based video synthesis module that translates these motions into high-fidelity video animations. By operating in a compact facial motion latent space decoupled from visual and identity cues, our autoregressive diffusion model effectively captures long-range correlations between audio and facial dynamics through a diffusion-forcing training paradigm, enabling infinite-length emotionally-rich motion prediction without error accumulation. Extensive experiments demonstrate that X-Actor produces compelling, cinematic-style performances that go beyond standard talking head animations and achieves state-of-the-art results in long-range, audio-driven emotional portrait acting.

cs.CV

X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention

We propose X-NeMo, a novel zero-shot diffusion-based portrait animation pipeline that animates a static portrait using facial movements from a driving video of a different individual. Our work first identifies the root causes of the key issues in prior approaches, such as identity leakage and difficulty in capturing subtle and extreme expressions. To address these challenges, we introduce a fully end-to-end training framework that distills a 1D identity-agnostic latent motion descriptor from driving image, effectively controlling motion through cross-attention during image generation. Our implicit motion descriptor captures expressive facial motion in fine detail, learned end-to-end from a diverse video dataset without reliance on pretrained motion detectors. We further enhance expressiveness and disentangle motion latents from identity cues by supervising their learning with a dual GAN decoder, alongside spatial and color augmentations. By embedding the driving motion into a 1D latent vector and controlling motion via cross-attention rather than additive spatial guidance, our design eliminates the transmission of spatial-aligned structural clues from the driving condition to the diffusion backbone, substantially mitigating identity leakage. Extensive experiments demonstrate that X-NeMo surpasses state-of-the-art baselines, producing highly expressive animations with superior identity resemblance. Our code and models are available for research.

cs.CV

Temporal Separation with Entropy Regularization for Knowledge Distillation in Spiking Neural Networks

Spiking Neural Networks (SNNs), inspired by the human brain, offer significant computational efficiency through discrete spike-based information transfer. Despite their potential to reduce inference energy consumption, a performance gap persists between SNNs and Artificial Neural Networks (ANNs), primarily due to current training methods and inherent model limitations. While recent research has aimed to enhance SNN learning by employing knowledge distillation (KD) from ANN teacher networks, traditional distillation techniques often overlook the distinctive spatiotemporal properties of SNNs, thus failing to fully leverage their advantages. To overcome these challenge, we propose a novel logit distillation method characterized by temporal separation and entropy regularization. This approach improves existing SNN distillation techniques by performing distillation learning on logits across different time steps, rather than merely on aggregated output features. Furthermore, the integration of entropy regularization stabilizes model optimization and further boosts the performance. Extensive experimental results indicate that our method surpasses prior SNN distillation strategies, whether based on logit distillation, feature distillation, or a combination of both. The code will be available on GitHub.

cs.CV

Efficient Logit-based Knowledge Distillation of Deep Spiking Neural Networks for Full-Range Timestep Deployment

Spiking Neural Networks (SNNs) are emerging as a brain-inspired alternative to traditional Artificial Neural Networks (ANNs), prized for their potential energy efficiency on neuromorphic hardware. Despite this, SNNs often suffer from accuracy degradation compared to ANNs and face deployment challenges due to fixed inference timesteps, which require retraining for adjustments, limiting operational flexibility. To address these issues, our work considers the spatio-temporal property inherent in SNNs, and proposes a novel distillation framework for deep SNNs that optimizes performance across full-range timesteps without specific retraining, enhancing both efficacy and deployment adaptability. We provide both theoretical analysis and empirical validations to illustrate that training guarantees the convergence of all implicit models across full-range timesteps. Experimental results on CIFAR-10, CIFAR-100, CIFAR10-DVS, and ImageNet demonstrate state-of-the-art performance among distillation-based SNNs training methods. Our code is available at https://github.com/Intelli-Chip-Lab/snn\_temporal\_decoupling\_distillation.

cs.LG

The Frankl-Pach upper bound is not tight for any uniformity

For any positive integers $n\ge d+1\ge 3$, what is the maximum size of a $(d+1)$-uniform set system in $[n]$ with VC-dimension at most $d$? In 1984, Frankl and Pach initiated the study of this fundamental problem and provided an upper bound $\binom{n}{d}$ via an elegant algebraic proof. Surprisingly, in 2007, Mubayi and Zhao showed that when $n$ is sufficiently large and $d$ is a prime power, the Frankl-Pach upper bound is not tight. They also remarked that their method requires $d$ to be a prime power, and asked for new ideas to improve the Frankl-Pach upper bound without extra assumptions on $n$ and $d$. In this paper, we provide an improvement for any $d\ge 2$ and $n\ge 2d+2$, which demonstrates that the long-standing Frankl-Pach upper bound $\binom{n}{d}$ is not tight for any uniformity. Our proof combines a simple yet powerful polynomial method and structural analysis.

math.CO

Stereo-Talker: Audio-driven 3D Human Synthesis with Prior-Guided Mixture-of-Experts

This paper introduces Stereo-Talker, a novel one-shot audio-driven human video synthesis system that generates 3D talking videos with precise lip synchronization, expressive body gestures, temporally consistent photo-realistic quality, and continuous viewpoint control. The process follows a two-stage approach. In the first stage, the system maps audio input to high-fidelity motion sequences, encompassing upper-body gestures and facial expressions. To enrich motion diversity and authenticity, large language model (LLM) priors are integrated with text-aligned semantic audio features, leveraging LLMs' cross-modal generalization power to enhance motion quality. In the second stage, we improve diffusion-based video generation models by incorporating a prior-guided Mixture-of-Experts (MoE) mechanism: a view-guided MoE focuses on view-specific attributes, while a mask-guided MoE enhances region-based rendering stability. Additionally, a mask prediction module is devised to derive human masks from motion data, enhancing the stability and accuracy of masks and enabling mask guiding during inference. We also introduce a comprehensive human video dataset with 2,203 identities, covering diverse body gestures and detailed annotations, facilitating broad generalization. The code, data, and pre-trained models will be released for research purposes.

cs.CV

Algebraic approach to stability results for Erd\H{o}s-Ko-Rado theorem

Celebrated results often unfold like episodes in a long-running series. In the field of extremal set thoery, Erd\H{o}s, Ko, and Rado in 1961 established that any $k$-uniform intersecting family on $[n]$ has a maximum size of $\binom{n-1}{k-1}$, with the unique extremal structure being a star. In 1967, Hilton and Milner followed up with a pivotal result, showing that if such a family is not a star, its size is at most $\binom{n-1}{k-1} - \binom{n-k-1}{k-1} + 1$, and they identified the corresponding extremal structures. In recent years, Han and Kohayakawa, Kostochka and Mubayi, and Huang and Peng have provided the second and third levels of stability results in this line of research. In this paper, we provide a unified approach to proving the stability result for the Erd\H{o}s-Ko-Rado theorem at any level. Our framework primarily relies on a robust linear algebra method, which leverages appropriate non-shadows to effectively handle the structural complexities of these intersecting families.

math.CO