SearcharxivSearch

arXiv subjects

Heng Li

Publications and source records attributed to Heng Li.

At least 19 recordsLinked to original sources

ReconPlusGen: Injecting Reconstruction Prior into Multi-view 3D Generation through Noise Inversion and Modulation

Qualitative results and an illustration of our core idea. Top left: reconstruction results on benchmark images. Top right: reconstruction results on real-world images. Bottom: illustration of reconstruction-guided noise initialization and modulation. Given multiple input images, we predict a point cloud in canonical space, deterministically inject the predicted geometry into the diffusion process through noise inversion, and modulate the resulting noise to preserve the generative flexibility required to complete unobserved regions and refine visible geometry.

cs.CV

MotionCanvas: Learning Implicit Motion Planning from Composable Kinematic Cues

Professional character animation requires both natural motion and precise, versatile control. For example, it is common for the creators to define the timing of a specified action, to control the motion range of the character's arm swing, and the route the character walks through, like specifying various kinematic motion cues on a ``motion canvas''. This motivates us to propose MotionCanvas, a model that supports \emph{cue-conditioned implicit motion planning} to faithfully and coherently connect all cues, dense or sparse, full or partial, into one full-body motion sequence. Specifically, MotionCanvas represents heterogeneous kinematic cues on a shared motion canvas, where position and rotation values are specified across body joints and time. A shared flow-matching model generates motion conditioned on this canvas, with optional language and input motion; cue imputation keeps the specified canvas values fixed in both training and sampling. To learn coherent completion across different cue sets, we train with a compositional cue sampler that varies when cues are applied, which positions or rotations are specified, and how they are combined. Together, these designs enable a single generator to synthesize globally coherent actions that jointly satisfy compatible heterogeneous cues. We test this planning ability with temporal, root, and body-part cues---alone and in combination---and language-guided editing. We naturally extend this evaluation to sequential generation and motion repair, since both require the same ability to organize coherent motion from kinematic cues. Across these evaluations, MotionCanvas establishes state-of-the-art results in controlled-motion quality, mixed-cue adherence, sequential generation, instruction editing, and motion repair while preserving its text-to-motion capability.

cs.MM

ToolLoop: Closed-Loop Tool-Use Data Synthesis via Decomposed Generation and Dynamic Self-Feedback

High-quality tool-use data is critical for training language models to interact effectively with external tools. However, existing synthetic approaches typically follow a generate-then-filter paradigm with static post-hoc verification, often yielding inefficient data with imbalanced feature distributions. We propose ToolLoop, a closed-loop framework that decomposes synthesis into three progressive stages: (1) sampling function name combinations as ground truth; (2) backward derivation of user queries; and (3) forward derivation of tool calls. At each stage, dynamic self-feedback iteratively guides the model toward high-quality generation, realizing a transition from generate-then-filter to generate-verify-refine. On the Berkeley Function Calling Leaderboard (BFCL), a 4B parameter model trained with our 11K synthetic examples achieves 86.40% accuracy in non-reasoning mode, while an Isolate variant that removes BFCL-overlapping candidate functions still reaches 86.07\%. Cross-benchmark evaluation on ACEBench further demonstrates strong generalization, with 72.1% overall accuracy using only 18.3% of baseline training data.

cs.CL

Turing universality, computability, and incompleteness in hypergraph Tur\'an theory

Given a finite family $\mathcal F$ of forbidden $r$-graphs, the Tur\'an problem asks for the maximum asymptotic edge density of $\mathcal F$-free $r$-graphs and the structure of near-extremal examples. We show that both questions can encode arbitrary computation. Fix a universal Turing machine $\mathsf U$. For every sufficiently large fixed $r$, there is a rational $\tau_r\in(0,1)$ such that, from each binary word $\beta$, one can construct a finite family $\mathcal F_{r,\beta}$ with $\pi(\mathcal F_{r,\beta})=\tau_r$ if $\mathsf U$ does not halt on $\beta$, and $\pi(\mathcal F_{r,\beta})>\tau_r$ otherwise. The same dichotomy governs extremal structure. We construct finite families $\mathcal G_{r,\beta}$ such that nonhalting gives a unique extremal limit and Erd\H{o}s--Simonovits stability, whereas halting gives two nonempty compact extremal phases separated by the sign of a fixed continuous statistic. Hence uniqueness and connectedness of the extremal space, symmetry breaking, two-phase behavior, and stability are all undecidable. The reductions are effective and verifiable in ZFC by finite certificates. Consequently, for every consistent computably axiomatized extension of ZFC and every sufficiently large fixed $r$, there is a finite family $\mathcal F$ for which the true equality $\pi(\mathcal F)=\tau_r$ is neither provable nor refutable; analogous independence holds for the five structural properties above. We also obtain effective approximation, classify exact comparison complexity, and show that the smallest improvement witnesses have Busy-Beaver growth, with no uniform computable positive lower bound on the density gain.

math.CO

Low-Rank Velocity Fields as a Structural Prior for Unsupervised 4D Medical Image Interpolation

Endpoint-only unsupervised 4D medical image interpolation synthesizes intermediate volumes from sparsely sampled sequences with only the start and end volumes available for training; however, this weakly constrained setting often yields intermediates with unstable boundaries and non-physiological motion, limiting interpretability and downstream analysis. We propose low-rank velocity fields as a structural prior, constraining motion to a structured Tucker low-rank velocity field space that decomposes motion into globally shared spatial bases and a compact sample-specific core, thereby encouraging spatially correlated, anatomy-consistent deformation while suppressing voxel-wise high-frequency artifacts. To capture global coordination and local non-rigid details, we model motion in a coarse-to-fine multi-scale scheme and compose scale-wise deformations at inference to synthesize volumes at arbitrary times. We further provide a theoretical analysis showing that, under Tucker parameterization, low-rank parameters control the smoothness energy of the velocity field, explaining why low-rank modeling promotes smoother motion. Experiments on ACDC and 4D-Lung demonstrate state-of-the-art performance, remaining competitive with methods trained with intermediate-frame supervision, and producing intermediates with improved structural coherence and more stable anatomical contours.

cs.CV

Phase-Aligned Finite-Fourier Periodic Deformation for 4D Medical Image Interpolation

4D medical image interpolation aims to recover missing volumes from sparsely observed time points and is important for dynamic anatomical analysis in applications such as cardiac MRI and thoracic CT, where motion is often repetitive or near-periodic over clinically relevant intervals. A key challenge is that this structure is not always encoded directly in deformation representations for interpolation. In addition, physiological motion is often non-uniform, so equal temporal intervals do not necessarily correspond to equal amounts of anatomical change. To address these issues, we formulate interpolation as learning a continuous deformation process with a phase-structured prior. Given two endpoint volumes, we parameterize a phase-conditioned velocity field with a finite Fourier basis, which embeds near-periodic motion patterns directly into the deformation space and supports continuous querying at arbitrary target times. We further introduce a phase-aligned temporal reparameterization that maps normalized within-interval time to a latent motion phase according to deformation variation intensity, thereby better modeling non-uniform motion progression. Intermediate volumes are then synthesized by continuously warping both endpoints, followed by bidirectional fusion and lightweight residual refinement. Experiments on ACDC and 4D-Lung show that the proposed method achieves state-of-the-art performance over existing baselines while producing anatomically plausible and coherent intermediate volumes from sparse observations.

cs.CV

Superlinear Lower Bounds for Monochromatic Path Partitions

In 1989, Gy\'arf\'as conjectured that the vertex set of every $r$-edge-coloured complete graph can be partitioned into at most $r$ vertex-disjoint monochromatic paths. Erd\H{o}s, Gy\'arf\'as, and Pyber subsequently proposed the analogous conjecture for monochromatic cycles. Pokrovskiy proved Gy\'arf\'as's conjecture for $r=3$, while disproving the conjecture of Erd\H{o}s, Gy\'arf\'as, and Pyber for every $r\ge3$ by constructing colourings that require at least $r+1$ monochromatic cycles. In this paper, we disprove Gy\'arf\'as's conjecture in a quantitatively strong superlinear form: for every sufficiently large $r$, there exists an $r$-edge-coloured complete graph that requires at least $(1-o(1))r\log\log r$ vertex-disjoint monochromatic paths. Consequently, the monochromatic cycle-partition number is also superlinear in $r$. Our construction also disproves two conjectures of Pokrovskiy: one on monochromatic cycle coverings and the other on path coverings in the balanced bipartite setting.

math.CO

Forcing Quasirandomness via Rooted F-Densities

Let $F$ be a finite graph with at least one edge, and let $W$ be a graphon. We show that if the density of $F$ rooted at each edge is almost everywhere constant, then either $t(F,W)=0$ or $W$ is constant. For edge-transitive $F$, one rooted equation suffices. This recovers the edge-rooted triangle theorem of Reiher and Schacht. In their terminology, our result also shows that every clique is $2$-forcing, answering a question they posed. We give an explicit stability estimate when $W$ is bounded away from zero. Our proof has two steps: an entropy argument turns constant rooted densities into an additive identity for $\log W$, and a Hoeffding decomposition determines all solutions of that identity. The same method gives exact classifications and quantitative stability estimates for symmetric uniform hyperkernels, dissociated Aldous--Hoover hypergraphons, directed kernels, and tournamentons.

math.CO

Understanding the Energy Impact of Software Refactoring: A Workload-Aware Study of Controlled Examples and Real-World Commits

Refactoring improves software maintainability while preserving functional behavior, yet behavior preservation does not imply energy neutrality. Existing studies primarily examine isolated refactorings under fixed or simple workloads, leaving the effects of workload variation, real-world refactoring practices, explanatory factors, and energy regression identification insufficiently understood. We present the first large-scale empirical study of the energy impact of refactoring across two complementary Java benchmarks: a Micro-benchmark, comprising 68 refactoring types evaluated under diverse workloads, and a Practical-benchmark, containing 481 real-world refactoring commits from 430 GitHub projects. Using repeated paired energy measurements, we analyze workload sensitivity, refactoring patterns, explanatory factors, and the effectiveness of metric- and LLM-based regression identification. In the Micro-benchmark, 199 of 384 refactoring-workload pairs (51.8%) exhibit statistically significant energy differences, and 45.3% of refactoring instances change energy-impact classification across workloads. In the Practical-benchmark, only 36 commits (7.5%) show significant energy changes, although two-thirds differ by at least 10%. Refactoring type alone is insufficient to predict energy outcomes, while certain recurring refactoring combinations are associated with energy reductions. Changes in execution time consistently explain energy variation in the controlled benchmark but correlate weakly with energy changes in real-world commits. Our findings highlight the need for workload-diverse evaluation of the energy impact of refactoring; neither existing metric-based approaches nor LLM-based predictors can reliably identify refactoring-induced energy regressions, motivating the development of more accurate techniques for predicting the energy impact of refactoring.

cs.SE

Intervals of uniform Tur\'an densities

We prove that the set $\Pi_{\therefore,\infty}$ of uniform Tur\'an densities of possibly infinite families of $3$-graphs contains a terminal interval: there exists $\delta>0$ such that $[1-\delta,1]\subseteq\Pi_{\therefore,\infty}$. Consequently, $\Pi_{\therefore,\infty}$ has positive Lebesgue measure and Hausdorff dimension $1$.

math.CO

LLM Serving in the Wild: An Empirical Study of Frameworks, Methods, and System Designs

Large Language Models (LLMs) are integrated into software systems and AI services, making efficient LLM serving a concern for software engineering. Serving LLMs is challenging because inference requires computation, memory, GPU resources, and execution while maintaining latency and throughput. Although prior research has proposed LLM inference, optimization, and serving techniques and frameworks, little is known about how they are adopted in practice. In this study, we investigate the use of LLM serving frameworks and serving methods in open-source software systems. We identify and analyze five LLM-specific frameworks: vLLM, SGLang, TensorRT-LLM, LMDeploy, and FlashInfer. We examine how these frameworks and techniques are adopted individually and in combination, how adoption varies across categories of LLMs, and how repositories differ in intent, focus, use case, and architectural design. Our results show that vLLM is the most visible framework in popularity and adoption, while parallel computation, memory management, and network pruning are the most frequently used serving-method categories. Multi-framework usage is limited, suggesting that developers rely on a single serving framework; however, combined frameworks connect complementary capabilities across the serving stack. Framework adoption varies across model families, modalities, model sizes, domain specializations, and deployment settings. Repository-level analysis shows that LLM serving frameworks support applications and architectures, including Reinforcement Learning (RL)-based reasoning, multimodal generation and understanding, microservices, and cloud infrastructure. Overall, this study provides a large-scale empirical characterization of LLM serving framework adoption in practice and offers insights for researchers, framework maintainers, and practitioners working on LLM systems.

cs.SE

Nearly Sharp Bounds for Lattice Coverings by Convex Bodies

For an $n$-dimensional convex body $K$, let $\theta_L(K)$ denote its lattice covering density, and let $\Theta_L^{\mathrm{conv}}(n)$ and $\Theta_L^{\mathrm{sym}}(n)$ be the corresponding worst-case quantities over all convex bodies and over origin-symmetric convex bodies, respectively. Before this work, these quantities were known only to lie between a lower bound of order $n$ and an upper bound of order $n^2$, so even their polynomial order was undetermined. We prove that there are absolute constants $c,C>0$ such that \[ c n\log n \le \Theta_L^{\mathrm{sym}}(n) \le \Theta_L^{\mathrm{conv}}(n) \le Cn\log n\,(\log\log n)^{10/3+o(1)}. \] Thus both worst-case quantities are $n\log n\,(\log n)^{o(1)}$, and the upper and lower bounds differ by a factor at most $(\log\log n)^{10/3+o(1)}$. For the upper bound, a vertical--horizontal amplification based on weighted Boolean cubes combines covering estimates for low-codimensional sections into an exact lattice covering of an arbitrary convex body. For the lower bound, a random-slab construction and Poisson witnesses on flat tori show, with positive probability, that the resulting body admits no lattice covering of density below $c n\log n$.

math.MG

Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes

Speculative Decoding (SD) accelerates large language model inference by allowing a lightweight draft model to propose tokens that are subsequently verified in parallel by a larger target model. Recent approaches introduce lossy verification schemes to further improve efficiency by relaxing strict distributional matching. Yet such relaxation silently rewrites the decoding distribution, and the resulting acceleration can come at the cost of unstable, sometimes severely degraded generation quality. In this work, we present a principled analysis of the distributions induced by lossy verification methods. We show that many seemingly distinct approaches differ only superficially and can be unified into two categories: truncation-based verification and collaborative verification. We further construct a diagnostic evaluation framework across curated benchmarks. For truncation-based methods, we identify a fundamental pitfall-performance can degrade significantly compared to the true truncation sampling baseline due to distributional distortion. For collaborative verification, we reveal that well-designed relaxation principles, namely overshoot suppression and supervision quality, matter far more than the linear interpolation between draft and target. Our code is available at https://github.com/ZhouYuxuanYX/Fast-HSD.

cs.CL

Backend-Aware Graph Learning for Denoising Outcome Distributions in Quantum Program Testing

Testing quantum programs on NISQ (Noisy Intermediate-Scale Quantum) backends is challenging because the noise disturbs outcome distributions and can affect pass/fail decisions. We present Q-BRIDGE, a graph learning-based approach that converts noisy observations into denoised distributions suitable for oracle-based verification. Q-BRIDGE uses a graph transformer architecture to encode a transpiled quantum circuit, capturing the characteristics of its gates and their connectivity; the physical backend information is encoded together with the logical structure of the circuit. An additional conditioning layer, based on FiLM (Feature-Wise Linear Modulation), takes the encoding as input and integrates noisy observations to produce denoised outcomes. We evaluate Q-BRIDGE on 23 IBM noise backends and 6 circuit families representative of practical workloads. In the first setting, we train a separate Q-BRIDGE model for each backend; in the second setting, we train a single general model shared across all backends. Across both settings, Q-BRIDGE outperforms the state-of-the-art baseline in noise mitigation by a large margin. In testing scenarios with noisy executions, Q-BRIDGE achieves 93.97%-94.90% precision and 82.50%-83.51% recall in detecting bug-induced test failures, significantly outperforming the state-of-the-art baseline. These results indicate that considering the graph structure of the transpiled circuits and the physical characteristics of specific quantum backends is a practical route to more reliable noise-aware quantum program testing.

cs.SE

FSE: Continual Learning for Named Entity Recognition by Fast-Slow Experts

Continual Learning for Named Entity Recognition (CLNER) enable models to incrementally learn new entity types without forgetting previously acquired ones. However, existing methods suffer from catastrophic forgetting and insufficient exploitation of shared information across tasks. This paper proposes FSE, a Fast-Slow Experts enhanced span-based NER model for CLNER. The shared fast expert learns token-level links to efficiently filter out unlikely spans, while the task-specific slow expert performs span classification only on the remaining candidates. It stabilizes learning by promoting knowledge sharing across tasks and maintains plasticity by reducing learning burden at each task. A length-decay negative sampling strategy to mitigate span imbalance is also introduced. Extensive experiments on OntoNotes and FewNERD synthestic datasets demonstrate that FSE achieves state-of-the-art performance in CLNER scenarios, with effectiveness of each component, empirical evidence of faster convergence and expected functionality of both experts.

cs.CL

Grasp, Handover, Rotate: Bimanual Object Reorientation via Compositional Diffusion and Energy-Based Optimization

Bimanual object reorientation - picking an object, handing it over between two arms, and placing it in a desired target pose - is valuable when direct placement from the initial grasp is infeasible due to collisions, kinematic constraints, or poor final orientation. However, achieving this under multiple competing objectives remains challenging. We introduce BiCompoDiff, a compositional diffusion and energy-based framework that jointly optimizes grasp selection, handover, regrasp, and motion planning under multiple constraints. By combining a pretrained grasp diffusion model with bimanual planning energy-based models (EBMs), our method injects gradient guidance during reverse diffusion to enforce collision avoidance, trajectory smoothness (via differentiable inverse kinematics), handover feasibility, and regrasp safety. Annealed MCMC sampling further refines grasp poses over the composite energy landscape. Experiments across diverse simulated household reorientation tasks demonstrate that BiCompoDiff achieves over 20% higher success rates and up to 37% smoother trajectories (measured by joint displacement) compared to strong sampling-based baselines. Real-world validation confirms effective sim-to-real transfer and robust performance on challenging scenes.

cs.RO

Glob3R: Global Structure-from-Motion with 3D Foundation Models

Recent 3D geometric foundation models, such as VGGT, provide robust feed-forward 3D reconstruction by directly predicting camera poses and 3D scene points from input images. However, their results remain inaccurate, and scaling them to long sequences or large unordered image sets typically requires chunk-wise processing, which can introduce drift and inconsistency. We present Glob3R, a global SfM-style reconstruction built on 3D foundation models. Our key idea is to explicitly optimize feed-forward geometric predictions. To this end, we augment a frozen Pi3X backbone with a lightweight dense matching head that predicts image warps between selected reference frames and neighboring views. These dense warps are converted into sparse but reliable multi-view feature tracks, which provide correspondence constraints for global optimization. We further introduce a keyframe-based sliding-window association strategy that propagates tracks and relative poses across overlapping windows, enabling scalable reconstruction. Finally, we perform global motion averaging and bundle adjustment to refine camera poses, reduce scale inconsistencies, and recover dense scene geometry. Extensive experiments on indoor, outdoor, large-scale driving, and unordered SfM benchmarks demonstrate that Glob3R achieves robust and accurate reconstruction. It consistently improves over feed-forward foundation-model baselines and recent scalable reconstruction methods, while being more robust than classical SfM pipelines. The refined poses also lead to higher-quality neural rendering, validating the benefit of combining foundation-model priors with global geometric optimization. Project page: https://junyuandeng.github.io/Glob3r

cs.CV

A Higher-Order Clique Density Theorem

Reiher's clique density theorem determines the sharp lower envelope for the density of $K_r$ at fixed edge density. We prove a higher-order version in which the prescribed quantity is itself a clique density. For every $3\le s<r$, we determine the minimum possible $K_r$-density among graphons with prescribed $K_s$-density. For $s\ge3$ the constraint is genuinely nonlinear and leaves the edge density undetermined; nevertheless, on the positive range the sharp lower boundary is the classical multipartite edge-to-clique profile, reparametrised by $K_s$-density. We also prove stability on the positive branches of this profile: at every interior point, near extremality forces cut-distance closeness to the corresponding extremal family at the induced edge density.

math.CO