Searcharxiv⌕ Search

arXiv subjects

Hong Liu

Publications and source records attributed to Hong Liu.

At least 37 records · Page 2Linked to original sources

Structural Reductions for Monochromatic Matchings and Ramsey Tilings

The Alon--Frankl--Lovász theorem determines the chromatic number of Kneser hypergraphs; equivalently, it gives the sharp minimum size of a monochromatic matching in every edge-colouring of a complete uniform hypergraph. Its known general proofs are topological. We introduce a topology-free structural framework. It reduces every colouring of a pseudorandom $t$-graph, with only $o(n)$ loss in the largest monochromatic matching, to a colouring of $K_n^{(t)}$ whose vertex set has at most $r$ parts and whose edge colours depend only on intersection profiles. Together with a stability analysis at the critical scale, we prove an exact robust form: there exists $c=c(r,t)>0$ such that, if a $t$-graph satisfies $δ(\mathcal G)\ge(1-c)\binom{n-1}{t-1}$, then every $r$-colouring of $\mathcal G$ contains a monochromatic matching of the exact optimal size for all sufficiently large $n$. This gives a topology-free proof of the Alon--Frankl--Lovász theorem for large $n$, a sparse random transference theorem, and the exact value predicted by Meunier's stable Kneser conjecture throughout the range covered by our robust AFL theorem. We further develop the framework for Ramsey graph tilings. For a graph $H$, let $Rt_r(H;K_n)$ be the minimum, over all $r$-edge-colourings of $K_n$, of the largest monochromatic $H$-tiling. We prove $$ Rt_r(H;K_n)=(β_{r,H}+o(1))n, $$ where $β_{r,H}$ is effectively computable from finitely many rational linear programs depending only on $H$ and $r$. An additional multipartite Ramsey argument is needed to reconstruct a consistent coloured template. This gives an effective asymptotic solution to the multicolour Ramsey-tiling problem, extending the classical two-colour theorem of Burr, Erdős and Spencer. We also determine explicit constants for several natural families.

math.CO↗

EATR-Stereo: Embodiment-Aware Token Routing of Paired Stereo Evidence for Humanoid Vision-Language-Action Control

Long-horizon humanoid vision--language--action (VLA) control with head-mounted stereo cameras requires visual interfaces that can exploit complementary views while maintaining compatibility with pretrained representations. Existing interfaces often discard complementary stereo evidence or fuse additional observations without preserving the native primary-view pathway and adapting auxiliary information to robot embodiment. We present EATR-Stereo, an embodiment-aware token-routing framework that retains primary-view tokens and constructs primary-aligned Cross-View Auxiliary Tokens (CVATs) by querying the synchronized auxiliary-view token sequence. A body-segmented proprioceptive encoder further conditions token-wise auxiliary usage on robot configuration history, enabling selective incorporation of stereo evidence during action generation. The routed auxiliary stream augments the language and primary-visual context of a pretrained VLA while keeping its vision--language model frozen. On a 33-DoF physical humanoid with a 37-D proprioceptive state, we evaluate nine configurations in over-100-s search--approach--grasp--place--return tasks. EATR-Stereo achieves 60.0% full-task success, 100.0% grasp success, and 80.0% stage success. Under severe asymmetric occlusion, it improves recovery to 80% compared with 30% for CVAT alone. Ablation studies further show the importance of preserving primary tokens and combining cross-view auxiliary features with structured proprioceptive routing. These results demonstrate that selectively routed paired stereo evidence improves spatial grounding for reliable long-horizon humanoid VLA control.

cs.RO↗

MoE-Enhanced Explainable Deep Manifold Transformation for Complex Data Embedding and Visualization

Dimensionality reduction (DR) plays a crucial role in various fields, including data engineering and visualization, by simplifying complex datasets while retaining essential information. However, achieving both high DR accuracy and strong explainability remains a fundamental challenge, especially for users dealing with high-dimensional data. Traditional DR methods often face a trade-off between precision and transparency, where optimizing for performance can lead to reduced explainability, and vice versa. This limitation is especially prominent in real-world applications such as image, tabular, and text data analysis, where both accuracy and explainability are critical. To address these challenges, this work introduces the MoE-based Explainable Deep Manifold Transformation (DMT-ME). The proposed approach combines a geometry-aware hyperbolic mapper with Mixture of Experts (MoE) models, where sparse expert specialization provides the main representational gain and the hyperbolic component offers an additional refinement for structurally complex data. DMT-ME enhances DR accuracy primarily through MoE-based sparse routing and structure-aware matching, while also improving explainability by explicitly linking input data, embedding outcomes, and key features through the MoE structure. Extensive experiments demonstrate that DMT-ME consistently achieves superior performance in both DR accuracy and model explainability, making it a robust solution for complex data analysis. The code is available at https://github.com/zangzelin/code_dmtme.

cs.LG↗

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention

The quadratic cost of self-attention makes long-context inference prohibitively expensive, and proxy-based block-sparse attention has become a practical remedy. Existing methods typically rely on a proxy to predict a binary sparse mask and a kernel to consume this mask and perform sparse attention computation. Such an approach is effective under moderate budgets. However, as the budget tightens, the estimated proxy inevitably drops some salient blocks, while the kernel can only apply the sparse mask mechanically, leading to an evident drop in model accuracy. We propose CoSA, a two-stage training-free Sparse Attention under proxy-kernel CO-design, which couples a Kernel-Aware Proxy (KAP) with an Ordered-Skipping Kernel (OSK). In the first stage, the KAP selects blocks under a moderate budget and produces an ordered mask that prescribes the order in which KV pages are visited in the kernel inner loop. In the second stage, the OSK applies this mask and skips more blocks under a tightened budget given online-softmax statistics. Across mainstream LLM backbones and long-context benchmarks, CoSA attains higher accuracy at lower budgets. Impressively, CoSA achieves a 4.93$\times$ attention speedup and reduces end-to-end Time-to-First-Token by 2.53$\times$ under a context length of 128K with negligible performance degradation. Code is available at https://github.com/Tencent/AngelSlim.

cs.CL↗

LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer Indexing

DeepSeek Sparse Attention (DSA) enables efficient long-context modeling through its Lightning Indexer. However, practical deployment remains constrained by the indexer's expensive $O(L^2)$ scoring overhead and the hardware-inefficient, discontinuous memory-access patterns induced by its outputs. To address these system-level bottlenecks, we introduce LongCat Sparse Attention (LSA), a hardware-algorithm co-designed framework comprising three complementary and orthogonal strategies: (1) Streaming-Aware Indexing, which selectively converts scattered KV entries into hardware-aligned contiguous layouts to enable coalesced HBM access; (2) Cross-Layer Indexing, which amortizes indexing overhead by reusing the results produced by a single layer across consecutive layers, supported by cross-layer distillation; and (3) Hierarchical Indexing, which adopts a coarse-to-fine scoring scheme to progressively narrow the candidate set for each query, thereby substantially reducing indexing computation. Extensive scaling experiments, ranging from 69B-A3B to 560B-A27B models, demonstrate that LSA consistently achieves performance on par with full attention across both general-purpose and long-context benchmarks. Moreover, LSA supports native training with context lengths of up to one million tokens and underpins the development of LongCat-2.0 (1.6T-A48B). To facilitate further research, we also introduce and open-source LongCat-Flash-Lite-Sparse (69B-A3B), which integrates LSA into LongCat-Flash-Lite and incorporates an updated long-context training corpus.

cs.AI↗

FOCUS: FP4 Optimization via Coupled-Relaxation and Dual-Granularity Scaling

Large language models (LLMs) achieve remarkable performance but are expensive to deploy due to their enormous size. FP4 quantization, with formats such as MXFP4 and NVFP4, offers an appealing solution with native hardware support on modern accelerators. However, maintaining accuracy under FP4 precision remains difficult. A key bottleneck lies in scale optimization: existing methods tightly couple the quantization and dequantization scales, forcing both to conform to the discrete low-precision format required by hardware, such as E8M0 in MXFP4. Yet the quantization scale is never stored and need not obey this constraint, suggesting a significant untapped optimization space. In this work, we propose FOCUS, a post-training quantization framework with end-to-end scale learning for FP4 Optimization via Coupled-Relaxation and Dual-Granularity Scaling. Coupled-Relaxation Scaling (CRS) relaxes the tight coupling between quantization and dequantization scales with a learnable full-precision coefficient, enabling more effective optimization without breaking hardware compliance. Dual-Granularity Scaling (DGS) further refines the quantization scale at a finer sub-block granularity, allowing more precise adaptation to local weight distributions. Experiments across multiple LLM families and benchmarks show that FOCUS achieves state-of-the-art FP4 accuracy under both MXFP4 and NVFP4 formats, while introducing no additional inference overhead. Code and quantized models will be released at https://github.com/tencent/AngelSlim.

cs.AI↗

Strong invariants and Tverberg numbers in convexity spaces

Helly, Carathéodory, and Radon numbers encode three kinds of finite certificates in a convexity space: for the emptiness of an intersection, for membership in a convex hull, and for the existence of intersecting hulls. We study exact versions of these certificates, in which a subfamily must preserve the whole intersection or a subset must preserve the whole hull. Our first main result shows that, for finite configurations in an arbitrary convexity space, five a priori different boundedness conditions are equivalent: VC-dimension, strong Helly number, strong Carathéodory number, comatching number, and strong Radon number (with the expected additive-one shift). We also obtain equivalent layered Tverberg-type decompositions and colorful consequences. The common mechanism is exposed by the bipartite incidence graph between points and a generating family. For finite spaces, the unique minimal generator yields a natural dual convexity space; we characterize double dualization and prove that the strong parameters are duality invariant. The same model gives a polynomial-size, $O(t^4)$, realization of Bukh's counterexample to the Calder-Eckhoff partition conjecture. Finally, we obtain the first Tverberg bound for separable convexity spaces that is simultaneously linear in the number of parts and polynomial in the Radon number. If an $S_3$-separable convexity space has Helly number $h$ and its halfspaces have VC-dimension $d$, then $r_t=O(dh\log h)\,t$; in particular, Radon number $r$ gives $r_t=O(r^2\log r)\,t$. The bound attains the weak-Eckhoff scale $O(rt)$ whenever the Helly number is bounded. For axis-parallel box convexity in $\mathbb{R}^k$, gives the optimal order $r_t=O(rt)$ uniformly in every dimension. This appears to be the first dimension-uniform estimate of weak-Eckhoff order for box convexity, whereas the previous direct theory was confined to dimension three.

math.CO↗

AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding

Speculative decoding accelerates large language model inference without changing the target distribution, but no single drafting structure performs best across real-world workloads. Autoregressive multi-token prediction (MTP) is a lightweight, stable proposal mechanism, whereas block-parallel diffusion amortizes drafting latency over much longer candidate sequences; the better choice depends strongly on the output distribution. We present AngelSpec, a unified training framework for MTP and block-parallel speculative decoding that addresses this heterogeneity at three levels. At the training level, rather than fitting one universal drafter to a uniform data mixture, we co-specialize structure and data: the MTP drafter is trained on diverse conversational data for high-entropy open-ended chat, and the block-diffusion drafter on code and mathematics data for longer predictable continuations. At the architecture level, we propose DFly, a block-diffusion framework combining a hybrid target-conditioning backbone with a predecessor-conditioned autoregressive head, improving target-feature utilization and intra-block dependency modeling while keeping generation parallel. At the inference level, both acceptance length and verification cost vary with domain, request, online load, and hardware, so DFly treats verification as a shared batch-level resource: it reallocates compute toward high-confidence prefixes across requests and combines expected utility with a profiled cost model to adapt verification depth online. Across the Hy3 series, DFly raises the average accepted length on Hy3-A21B by roughly 30% and attains the highest average throughput at every tested concurrency from 4 to 64, a 1.98-2.40x speedup over autoregressive decoding and 10.5-11.8% higher throughput than DFlash. We release AngelSpec to support training and extending these methods.

cs.CL↗

Probing Stringy Horizons with Pole-Skipping in Non-Maximal Chaotic Systems

In this paper, we study pole-skipping in non-maximally quantum chaotic systems. Using Rindler conformal field theories and the large-$q$ SYK chain as illustrative examples, we argue that the pole-skipping points of few-body operators organize into trajectories in the complex frequency-momentum plane, with the leading trajectory encoding the quantum Lyapunov exponent. We further propose that these trajectories admit a natural interpretation as Regge trajectories of stringy excitations in a dual stringy black hole geometry. From this perspective, pole-skipping for an individual operator can be viewed as tracking the stringy horizon through the response of a single excitation. Our results suggest that pole-skipping reflects intrinsic properties of quantum chaotic systems and may be deeply connected to the structure of horizons in the stringy regime.

hep-th↗

PIVOT: Efficient Query-Group Indexing for Token-Level Sparse Attention

Token-level sparse attention, as implemented by DeepSeek Sparse Attention (DSA) in production systems, makes the downstream attention efficient but shifts the bottleneck to the indexer that feeds it. To select the top-k tokens for each query, the indexer must still score every preceding token, incurring a cost of O(L^2) per layer for a sequence of length L. We observe that this per-query scan is largely redundant: nearby queries select highly overlapping top-k tokens, and the indexer scores are long-tailed along the key axis. We exploit these properties in PIVOT, Proxy Indexing Via One full-prefix Traversal, a training-free, drop-in replacement for the DSA indexer that shares one prefix scan across a group of nearby queries. PIVOT aggregates a group into a single proxy query, performs one shared full-prefix scan to obtain a candidate set, and then selects a top-k for each query from that set. Two variants trade speed for fidelity: PIVOT-Reuse shares the proxy top-k across the group for maximum speed, whereas PIVOT-Refine re-scores the candidate set with the indexer of each query and then selects an individual top-k, matching the dense indexer at a small additional cost. A single algorithm covers both inference phases, differing only in how groups are formed: fixed-size groups of consecutive queries in prefill, and the queries decoded together in one multi-token prediction (MTP) step in decode. On DeepSeek-V3.2 and GLM-5.1 across LongBench and RULER, PIVOT matches the accuracy of the dense DSA indexer while accelerating it by up to 4x and reducing end-to-end latency by up to 1.6x at long context.

cs.CL↗

Fabric Pneumatic Artificial Muscles Based on the Drawstring Principle

Pneumatic artificial muscles have wide applications in robotics and industrial fields. Conventional pneumatic artificial muscles generate extra radial deformation during axial contraction, which severely wastes available working space. Inspired by the widely adopted drawstring principle in textile products, this paper proposes a novel drawstring fabric pneumatic artificial muscle (DPAM). Unlike traditional counterparts, the proposed DPAM produces no extra radial deformation during contraction, greatly improving structural compactness. The DPAM exhibits outstanding mechanical performance: a load capacity over 800 times its self-weight, a maximum contraction ratio of 44%, and a power density up to 4.98 kW/kg, alongside excellent scalability. Two representative application scenarios, bionic robots and industrial production lines, are demonstrated to validate its practicability. The DPAM can be easily expanded within a two-dimensional plane, as verified by the fabricated DPAM matrix. This work not only presents a high-performance novel pneumatic artificial muscle but also inspires researchers to draw design inspiration from conventional textile structures to address existing challenges in soft robotics.

cs.RO↗

D-cut: Adaptive Verification Depth Pruning for Batched Speculative Decoding

Speculative decoding accelerates large language model (LLM) inference without compromising output quality. Recent parallel drafting methods further improve single-request performance by decoupling draft length from drafting latency, enabling longer drafts and higher mean accepted tokens (MAT). However, under high request concurrency, long drafts waste substantial computation on rejected tokens, increasing verification cost and potentially making speculative decoding slower than autoregressive decoding. We present D-Cut, an adaptive pruning method that selects draft tokens jointly across the batch and concentrates the verification budget on tokens most likely to be accepted. D-Cut is motivated by two observations. First, acceptance lengths vary considerably across concurrent requests; D-Cut therefore performs cross-request pruning, allocating the verification budget adaptively according to draft confidence. Second, verification cost depends strongly on the deployment environment, including GPU architecture and parallelism strategy; D-Cut incorporates a runtime cost model to adapt its pruning depth to the target environment. Experiments on dense and mixture-of-experts (MoE) models show that, under high concurrency, D-Cut improves the average speedup from \(1.26\times\) to \(1.65\times\), restores acceleration in dense-model configurations where long-draft baselines are slower than autoregressive decoding, and achieves up to \(3.0\times\) speedup over autoregressive decoding on MoE models.

cs.CL↗

Tree suspensions and transfer functions for single degree Turán spectra

For integers $1\le \ell<k$, let $Π^k_\ell$ denote the single-forbidden $\ell$-degree Turán spectrum of $k$-uniform hypergraphs. We introduce transfer functions for this spectrum: explicit functions $f$ such that, for every $F$, there is another single $k$-graph $F^*$ with $π_\ell(F^*)=f(π_\ell(F))$. This gives a mechanism for producing new single-forbidden densities while retaining full control of the resulting value. Our transfer functions are realized by a new family of suspension-type operations, called tree suspensions. From these operations we obtain three explicit maps: one acting on $Π^k_\ell$ for every $1\le\ell<k$, a second acting when $\ell\ge k/2$, and a third acting in the ordinary Turán case $\ell=1$. The common feature is a robust tree structure which gives the lower bound by a two-part construction and, in the regimes above, admits a matching embedding or Lagrangian upper bound. As a first application, the universal transfer function propagates accumulation points. Using the recent zero-accumulation results for $\ell\ge2$ together with the ordinary Turán accumulation result of Conlon and Schülke, we prove that $Π^k_\ell$ has infinitely many accumulation points for every $k\ge3$ and every $1\le\ell<k$. This recovers, in particular, the known infinitude of accumulation points in the ordinary and codegree spectra. As a second application, combining two independent transfer functions forces algebraic degrees to grow. For every $k\ge3$ and every $\ell\in\{1,\lceil k/2\rceil,\ldots,k-2\}$, the spectrum $Π^k_\ell$ contains algebraic numbers of arbitrarily large degree over $\mathbb Q$. Thus the arithmetic complexity previously known for finite forbidden families already occurs in the single-forbidden spectrum, both for ordinary Turán density and for a broad range of degree Turán densities.

math.CO↗

A Higher-Order Clique Density Theorem

Reiher's clique density theorem determines the sharp lower envelope for the density of $K_r$ at fixed edge density. We prove a higher-order version in which the prescribed quantity is itself a clique density. For every $3\le s<r$, we determine the minimum possible $K_r$-density among graphons with prescribed $K_s$-density. For $s\ge3$ the constraint is genuinely nonlinear and leaves the edge density undetermined; nevertheless, on the positive range the sharp lower boundary is the classical multipartite edge-to-clique profile, reparametrised by $K_s$-density. We also prove stability on the positive branches of this profile: at every interior point, near extremality forces cut-distance closeness to the corresponding extremal family at the induced edge density.

math.CO↗

Tight Staircase Bounds for Cyclic Subsets below Dirac's Threshold

Let $\operatorname{Cyc}(G)$ denote the number of cyclic subsets in a graph $G$, which are subsets that induce a Hamiltonian subgraph. Draganić, Keevash and Müyesser recently proved that every regular Dirac graph has $Ω(2^n)$ cyclic subsets, resolving a problem of Erdős and Faudree. We determine the sharp asymptotic lower bound throughout the linear range below Dirac's threshold. Let $G$ be an $n$-vertex $d$-regular graph with $d=Ω(n)$ and $d<n/2$, then $$ \operatorname{Cyc}(G)\ge (q-o(1))2^{n/q}, \quad \text{where } \quad q=\left\lfloor \frac{n}{d+1}\right\rfloor \ge 2. $$ This bound is asymptotically best possible, including the leading coefficient $q$, as witnessed at the staircase levels by the disjoint union of $q$ equal cliques. Consequently, the optimal exponential rate changes by discrete jumps as $d$ crosses the thresholds $n/k$, rather than varying smoothly with $d$. We also prove the optimal exponential rate at the Dirac boundary: every $n$-vertex $n/2$-regular graph satisfies $\operatorname{Cyc}(G)\ge 2^{(1-o(1))n},$ which is sharp up to a subexponential factor by $K_{n/2,n/2}$.

math.CO↗

On the spectrum and structure of blowup thresholds

The chromatic threshold of Erdős and Simonovits asks when a minimum-degree condition forces every \(H\)-free graph to have bounded chromatic number. Thomassen's homomorphism threshold strengthens this by requiring a bounded \(H\)-free homomorphic image. The recently introduced blowup threshold \(δ_{\mathrm B}(H)\) asks for a still more rigid conclusion: when must every sufficiently dense maximal \(H\)-free graph be an actual blowup of a bounded graph? Thus the blowup threshold measures when quotient-level structure can be upgraded to exact bounded-template structure. We show that, although chromatic and homomorphism thresholds are often hard to separate, the stronger blowup threshold diverges from the chromatic threshold in several fundamental ways. First, we prove that \(δ_{\mathrm B}(H)>0\) for every non-bipartite graph \(H\). Hence, unlike the chromatic threshold, the blowup threshold never vanishes outside the bipartite world. Second, we prove that $δ_{\mathrm B}$ is not monotone under taking induced subgraphs. This shows that the blowup threshold is sensitive to global features of the forbidden graph and cannot be classified by a direct analogue of the monotonicity-based strategy used for chromatic thresholds. Third, we prove that $δ_{\mathrm B}(H)=\frac{1}{4}$ for a natural family of \(3\)-chromatic constrained blowups of odd cycles. This gives a new exact blowup-threshold value beyond the chromatic-threshold spectrum.

math.CO↗

Vanishing orders, suspensions and zero degree Turán densities

For integers $1\le \ell<k$, the $\ell$-degree Turán density $π_\ell(F)$ measures the minimum $\ell$-degree threshold that forces a copy of a fixed $k$-uniform hypergraph $F$, generalizing both the classical Turán density $π_1$ and the codegree Turán density $π_{k-1}$. Motivated by Erdős' characterization of $k$-graphs with zero Turán density, we study the structural implications of vanishing $\ell$-degree Turán density. Our main result concerns the case $\ell=2$. We prove that, for every $k\ge3$, if a $k$-graph $F$ satisfies $π_2(F)=0$, then $F$ admits a $2$-vanishing order, that is, a global vertex ordering under which all edges align canonically with respect to their pairs. This extends to all uniformities a structural phenomenon previously known for $3$-graphs, and gives a higher-degree analogue of the classical fact that $π_1(F)=0$ forces $F$ to be $k$-partite. In particular, the absence of a $2$-vanishing order is a structural obstruction to vanishing $2$-degree Turán density. We also establish a suspension principle connecting consecutive degree parameters. Given a $(k-1)$-graph $F$, let $\mathcal{S}_F$ be the $k$-graph obtained by adding an apex vertex $v$ and replacing each edge $e\in E(F)$ with $v\cup e$. We show that, for $2\le \ell<k$, $π_{\ell}(\mathcal{S}_F)=0$ if and only if $π_{\ell-1}(F)=0$. This provides a bridge between different degree Turán densities and allows vanishing results to be lifted across uniformities and degree parameters. As an application, we prove that except the classical Turán density, all other degree Turán densities accumulate at zero. The proof of our main result combines random geometric building blocks, a design-theoretic gluing scheme, and random sparsification to reconcile positive $2$-degree with local vanishing structure.

math.CO↗

The sharp threshold for rainbow stackings of random edge-colourings

A rainbow stacking of $m$ independent, uniformly random $r$-edge-colourings of $K_n$ is a tuple of vertex permutations that superimposes the colourings such that no two edges of the same colour overlap. The study of the critical palette size $r$ required for the existence of such stackings was recently initiated by Alon, Defant, and Kravitz [Bull. Lond. Math. Soc., 57, 2025], who bounded the phase transition within a constant-order window around $\frac{m\binom{n}{2}}{2\log(n!)}$. We determine the constant term in this transition. For every fixed $m\ge2$ and every function $ω(n)\to\infty$, with high probability there is no rainbow stacking if $$r\le \frac{m\binom{n}{2}}{2\log(n!)}+\frac{2m-1}{6}-\frac{ω(n)}{(\log n)^2},$$ while with high probability there is one if $$r\ge \frac{m\binom{n}{2}}{2\log(n!)}+\frac{2m-1}{6}+\frac{ω(n)}{(\log n)^2}.$$ Our proof combines a chromatic-polynomial expansion for an auxiliary conflict graph with a refined estimate of the associated weighted permutation sum. Our result yields the exact threshold $\Big\lceil \frac{m\binom{n}{2}}{2\log(n!)}+\frac{2m-1}{6}\Big\rceil$ for a density-one set of integers $n$, resolving a problem of Alon, Defant and Kravitz.

math.CO↗