SearcharxivSearch

arXiv subjects

Qin Zhao

Publications and source records attributed to Qin Zhao.

At least 19 recordsLinked to original sources

Spectral extrema of 1-planar graphs with no short cycles or small cliques

The spectral Tur\'an type problem, initiated by Nikiforov in 2007, aims to determine the graphs among $n$-vertex $H$-free graphs having maximum spectral radius. In this paper, we study this problem for $1$-planar graphs, i.e., graphs that admit a drawing in the plane such that each edge is crossed at most once. Recently, Xu and Chang proved that the graphs among all $n$-vertex $K_5$-free $1$-planar graphs having maximum spectral radius lie within a small family of candidates. First, this paper explicitly identifies the unique spectral extremal graph among the $n$-vertex $K_5$-free $1$-planar graphs. Second, it establishes a structural reduction theorem: For any forbidden subgraph $F$ with $\delta(F)\ge2$ that is contained in $K_2\vee P_{n-2}^{2+}$ but not in $K_2\vee I_{n-2}$, every spectral extremal $F$-free $1$-planar graph contains a spanning complete bipartite graph $K_{2,n-2}$, where $P^{2+}_{n-2}$ is obtained from a path $u_1u_2\dots u_{n-2}$ by adding edge $u_1u_{n-2}$ and all edges $u_iu_{i+2}$ for $1\le i\le n-4$, and $I_{n-2}$ denotes the empty graph on $n-2$ vertices. As applications, the graph among all $n$-vertex $C_5$-free (resp. $2C_5$-free) $1$-planar graphs having maximum spectral radius is determined. These results extend spectral Tur\'{a}n type problems for $1$-planar graphs from cliques to cycles and their disjoint union.

math.CO

Uncovering and Mitigating Positional Blind Spots in Vision-Language-Action Models

Recent Vision-Language-Action (VLA) models achieve promising performance in robotic manipulation, typically measured by success rates aggregated over predefined object configurations, an evaluation that implicitly assumes spatially uniform competence across the workspace. However, this assumption does not hold: even with the instruction and every other scene factor held fixed, merely relocating a task-irrelevant distractor can sharply raise the failure probability within localized, spatially coherent regions, which we term Positional Blind Spots (PBS). In this paper, we propose a two-stage black-box framework to uncover and mitigate PBS. During the uncovering stage, we grid the workspace and apply a one-sided log-likelihood-ratio test to localize PBS cells with significantly elevated risk. During the mitigation stage, we fine-tune the policy via LoRA on demonstrations collected from these PBS regions, improving competence there while largely preserving performance across the rest of the workspace. We evaluate our framework on five state-of-the-art VLA policies across two benchmarks, and find that PBS are pervasive and spatially concentrated in all of them, with failure rates up to 0.58. Our search strategy achieves an average F1-score of 0.678, outperforming random search and adaptive sampling baselines by 0.268 and 0.178, respectively. Guided by the discovered regions, targeted fine-tuning reduces the overall failure rate by 40.00%--85.19%.

cs.RO

RoleMix: Unifying Sequential and Non-Sequential Features via Semantic Tokenization for Post-Click Conversion Rate Prediction

Post-click conversion rate (PCVR) prediction is central to industrial recommendation, but remains challenged by the structural mismatch between sparse, unordered multi-field features and long, domain-specific behavior histories. Existing models often process these signals through separate pathways and fuse them late, weakening semantic roles and limiting cross-signal refinement. We propose RoleMix, a unified interaction architecture that represents sequential and non-sequential evidence through a shared, role-preserving token interface. Non-sequential fields are converted into explicit semantic tokens that preserve user, item, pairwise, dense, contextual, and cross-feature roles, while long behavior domains are compressed into item- and context-aware sequence-query tokens through two-stage hierarchical window attention. The resulting global, semantic, and sequence-query tokens are jointly refined by stacked UniMixing-Lite blocks for PCVR prediction. On the large-scale KDD Cup 2026 Tencent UniRec Challenge, RoleMix achieves 83.648% online AUC, outperforming the official industrial baseline by 1.953%. Ablation studies show that semantic tokenization yields the largest isolated gain, highlighting a key principle for large-scale PCVR modeling: preserving field semantics at the token-interface level is as important as scaling the interaction backbone.

cs.AI

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

Despite the recent promise in robot control, video generative models suffer from a domain mismatch due to their primary focus on content creation. For example, their design inherently prioritizes visual fidelity and creativity over computational efficiency and physical realism. In this work, we present LingBot-Video, a DiT-based video pretraining paradigm specifically tailored for embodied intelligence. From the architecture perspective, we adopt the Mixture-of-Experts (MoE), instead of dense, framework to achieve a better trade-off between modeling capacity and inference efficiency, and manage to scale it up from scratch. From the data perspective, we construct a data profiling engine that augments standard internet videos with extensive robot-oriented footage, encompassing manipulation, navigation, and egocentric perspectives, to equip the base model with an intrinsic understanding of actions and world dynamics. From the training perspective, we develop a multi-dimensional reward system to enforce the alignment regarding physical rationality and task completion, going beyond standard criteria such as aesthetics, prompt-following, and motion consistency. Comprehensive evaluations validate its performance and efficiency as a video foundation model. We contribute LingBot-Video as the inaugural large-scale, open-source MoE video foundation model to the community, in a pioneering effort to bridge digital creativity and physical actuation.

cs.CV

Tele-Catch: Adaptive Teleoperation for Dexterous Dynamic 3D Object Catching

Teleoperation is a key paradigm for transferring human dexterity to robots, yet most prior work targets objects that are initially static, such as grasping or manipulation. Dynamic object catch, where objects move before contact, remains underexplored. Pure teleoperation in this task often fails due to timing, pose, and force errors, highlighting the need for shared autonomy that combines human input with autonomous policies. To this end, we present Tele-Catch, a systematic framework for dexterous hand teleoperation in dynamic object catching. At its core, we design DAIM, a dynamics-aware adaptive integration mechanism that realizes shared autonomy by fusing glove-based teleoperation signals into the diffusion policy denoising process. It adaptively modulates control based on the interaction object state. To improve policy robustness, we introduce DP-U3R, which integrates unsupervised geometric representations from point cloud observations into diffusion policy learning, enabling geometry-aware decision making. Extensive experiments demonstrate that Tele-Catch significantly improves accuracy and robustness in dynamic catching tasks, while also exhibiting consistent gains across distinct dexterous hand embodiments and previously unseen object categories.

cs.RO

$\mathrm{L}^{2}$--convergence of the time-splitting scheme for nonlinear Dirac equation in 1+1 dimensions

We study the time-splitting scheme for approximating solutions to the Cauchy problem of the nonlinear Dirac equation in 1+1 dimensions. Under the assumption that the initial data for the scheme are convergent in $\mathrm{L}^{2}(\mathbb{R})$, we prove that the approximate solutions constructed by the corresponding time-splitting scheme are strongly convergent in $\mathrm{L}^{2}(\mathbb{R}\times[0,T])$ to the global strong solution of the nonlinear Dirac equation for any $T>0$. To achieve this, we first establish the pointwise estimates for time-splitting solutions. Based on these estimates, a modified Glimm-type functional is carefully designed to show that it is uniformly bounded in time, which yields $\mathrm{L}^2$ stability estimates for the scheme. Furthermore, we prove that the set of time-splitting solutions is relatively compact in $\mathrm{C}([0,T];\mathrm{L}^{2}(\mathbb{R}))$ for any $T>0$. Finally, we show that the limit of any convergent subsequence of the time-splitting solutions is the strong solution to the Cauchy problem of the nonlinear Dirac equation.

math.AP

SynthVerse: A Large-Scale Diverse Synthetic Dataset for Point Tracking

Point tracking aims to follow visual points through complex motion, occlusion, and viewpoint changes, and has advanced rapidly with modern foundation models. Yet progress toward general point tracking remains constrained by limited high-quality data, as existing datasets often provide insufficient diversity and imperfect trajectory annotations. To this end, we introduce SynthVerse, a large-scale, diverse synthetic dataset specifically designed for point tracking. SynthVerse includes several new domains and object types missing from existing synthetic datasets, such as animated-film-style content, embodied manipulation, scene navigation, and articulated objects. SynthVerse substantially expands dataset diversity by covering a broader range of object categories and providing high-quality dynamic motions and interactions, enabling more robust training and evaluation for general point tracking. In addition, we establish a highly diverse point tracking benchmark to systematically evaluate state-of-the-art methods under broader domain shifts. Extensive experiments and analyses demonstrate that training with SynthVerse yields consistent improvements in generalization and reveal limitations of existing trackers under diverse settings.

cs.CV

From Frames to Sequences: Temporally Consistent Human-Centric Dense Prediction

In this work, we focus on the challenge of temporally consistent human-centric dense prediction across video sequences. Existing models achieve strong per-frame accuracy but often flicker under motion, occlusion, and lighting changes, and they rarely have paired human video supervision for multiple dense tasks. We address this gap with a scalable synthetic data pipeline that generates photorealistic human frames and motion-aligned sequences with pixel-accurate depth, normals, and masks. Unlike prior static data synthetic pipelines, our pipeline provides both frame-level labels for spatial learning and sequence-level supervision for temporal learning. Building on this, we train a unified ViT-based dense predictor that (i) injects an explicit human geometric prior via CSE embeddings and (ii) improves geometry-feature reliability with a lightweight channel reweighting module after feature fusion. Our two-stage training strategy, combining static pretraining with dynamic sequence supervision, enables the model first to acquire robust spatial representations and then refine temporal consistency across motion-aligned sequences. Extensive experiments show that we achieve state-of-the-art performance on THuman2.1 and Hi4D and generalize effectively to in-the-wild videos.

cs.CV

Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation

We propose Ming-Flash-Omni, an upgraded version of Ming-Omni, built upon a sparser Mixture-of-Experts (MoE) variant of Ling-Flash-2.0 with 100 billion total parameters, of which only 6.1 billion are active per token. This architecture enables highly efficient scaling (dramatically improving computational efficiency while significantly expanding model capacity) and empowers stronger unified multimodal intelligence across vision, speech, and language, representing a key step toward Artificial General Intelligence (AGI). Compared to its predecessor, the upgraded version exhibits substantial improvements across multimodal understanding and generation. Notably, it achieves strong performance on vision-language understanding benchmarks, with overall scores on par with Gemini 2.5 Pro, and enables seamless switching among multimodal tasks in multi-turn interactions. In speech, it achieves strong performance in contextual and dialect-aware ASR while enabling joint, continuous-generation of speech, sound, and music. In vision, it introduces generative semantic segmentation that achieves competitive standalone performance and enhances spatial control and editing consistency, alongside marked improvements in identity preservation, and high-fidelity in-image text rendering. Together, these capabilities demonstrate that a single unified model can serve as a practical foundation for general-purpose multimodal intelligence.

cs.CV

Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation

Existing speech models suffer from competing requirements on token representations by understanding and generation tasks. This discrepancy in representation prevents speech language models from performing instruction-based free-form editing. To solve this challenge, we introduce a novel framework that unifies speech understanding, generation, and editing. The core of our unified model is a unified continuous speech tokenizer MingTok-Audio, the first continuous tokenizer to effectively integrate semantic and acoustic features, which makes it suitable for both understanding and generation tasks. Based on this unified continuous audio tokenizer, we developed the speech language model Ming-UniAudio, which achieved a balance between generation and understanding capabilities. Ming-UniAudio sets new state-of-the-art (SOTA) records on 8 out of 12 metrics on the ContextASR benchmark. Notably, for Chinese voice cloning, it achieves a highly competitive Seed-TTS-WER of 0.95. Leveraging this foundational model, we further trained a dedicated speech editing model Ming-UniAudio-Edit, the first speech language model that enables universal, free-form speech editing guided solely by natural language instructions, handling both semantic and acoustic modifications without timestamp condition. To rigorously assess the editing capability and establish a foundation for future research, we introduce Ming-Freeform-Audio-Edit, the first comprehensive benchmark tailored for instruction-based free-form speech editing, featuring diverse scenarios and evaluation dimensions spanning semantic correctness, acoustic quality, and instruction alignment. We open-sourced the continuous audio tokenizer, the unified foundational model, and the free-form instruction-based editing model to facilitate the development of unified audio understanding, generation, and manipulation.

cs.CL

Controllability and mixing for acoustic wave motions

This paper concerns the dynamical behaviors of acoustic wave motion driven by a force acting through the boundary. If the boundary force is a suitable control, we show that the dynamical system associated to the acoustic wave motion is exactly controllable. Furthermore, when it is random perturbation of white noise type, we prove that the corresponding stochastic system is strong mixing. The bridge between these two problems is the observability inequality, which will be established in this work.

math.AP

Planar Tur\'an number of quasi-double stars

Given a graph H, we call a graph $\textit{H-free}$ if it does not contain H as a subgraph. The planar Tur\'an number of a graph H, denoted by $ex_{\mathcal{P}}(n, H)$, is the maximum number of edges in a planar H-free graph on n vertices. A (h,k)-quasi-double star $W_{h,k}$, obtained from a path $P_3=v_1v_2v_3$ by adding h leaves and k leaves to the vertices $v_1$ and $v_3$, respectively, is a subclass of caterpillars. In this paper, we study $ex_{\mathcal{P}}(n,W_{h,k})$ for all $1\le h\le 2\le k\le 5$, and obtain some tight bounds $ex_{\mathcal{P}}(n,W_{h,k})\leq\frac{3(h+k)}{h+k+2}n$ for $3\le h+k\le 5$ with equality holds if $(h+k+2)\mid n$, and $ex_{\mathcal{P}}(n,W_{1,5})\le \frac{5}{2}n$ with equality holds if $12\mid n$. Also we show that $\frac{9}{4}n\le ex_{\mathcal{P}}(n,W_{2,4})\le \frac{5}{2}n$ and $\frac{5}{2}n\le ex_{\mathcal{P}}(n,W_{2,5})\le \frac{17}{6}n$, respectively.

math.CO

TeleOpBench: A Simulator-Centric Benchmark for Dual-Arm Dexterous Teleoperation

Teleoperation is a cornerstone of embodied-robot learning, and bimanual dexterous teleoperation in particular provides rich demonstrations that are difficult to obtain with fully autonomous systems. While recent studies have proposed diverse hardware pipelines-ranging from inertial motion-capture gloves to exoskeletons and vision-based interfaces-there is still no unified benchmark that enables fair, reproducible comparison of these systems. In this paper, we introduce TeleOpBench, a simulator-centric benchmark tailored to bimanual dexterous teleoperation. TeleOpBench contains 30 high-fidelity task environments that span pick-and-place, tool use, and collaborative manipulation, covering a broad spectrum of kinematic and force-interaction difficulty. Within this benchmark we implement four representative teleoperation modalities-(i) MoCap, (ii) VR device, (iii) arm-hand exoskeletons, and (iv) monocular vision tracking-and evaluate them with a common protocol and metric suite. To validate that performance in simulation is predictive of real-world behavior, we conduct mirrored experiments on a physical dual-arm platform equipped with two 6-DoF dexterous hands. Across 10 held-out tasks we observe a strong correlation between simulator and hardware performance, confirming the external validity of TeleOpBench. TeleOpBench establishes a common yardstick for teleoperation research and provides an extensible platform for future algorithmic and hardware innovation. Codes is now available at https://github.com/cyjdlhy/TeleOpBench .

cs.RO

SIGMAN:Scaling 3D Human Gaussian Generation with Millions of Assets

3D human digitization has long been a highly pursued yet challenging task. Existing methods aim to generate high-quality 3D digital humans from single or multiple views, but remain primarily constrained by current paradigms and the scarcity of 3D human assets. Specifically, recent approaches fall into several paradigms: optimization-based and feed-forward (both single-view regression and multi-view generation with reconstruction). However, they are limited by slow speed, low quality, cascade reasoning, and ambiguity in mapping low-dimensional planes to high-dimensional space due to occlusion and invisibility, respectively. Furthermore, existing 3D human assets remain small-scale, insufficient for large-scale training. To address these challenges, we propose a latent space generation paradigm for 3D human digitization, which involves compressing multi-view images into Gaussians via a UV-structured VAE, along with DiT-based conditional generation, we transform the ill-posed low-to-high-dimensional mapping problem into a learnable distribution shift, which also supports end-to-end inference. In addition, we employ the multi-view optimization approach combined with synthetic data to construct the HGS-1M dataset, which contains $1$ million 3D Gaussian assets to support the large-scale training. Experimental results demonstrate that our paradigm, powered by large-scale training, produces high-quality 3D human Gaussians with intricate textures, facial details, and loose clothing deformation.

cs.CV

Transonic Shocks for 2-D Steady Euler Flows with Large Gravity in a Nozzle for Polytropic Gases

In this paper, we are concerned with the existence of transonic shock solutions for two-dimensional (2-d) steady Euler flows of polytropic gases with the vertical gravity in a horizontal nozzle under a pressure condition imposed at the exit of the nozzle. The acceleration of the gravity g is assumed to take a generic value. We first show that the existence of special transonic shock solutions with the flow states depending only on the variable in the gravity direction can be established if and only if the Mach number of the incoming flow satisfies certain conditions. However, the shock position of the special solutions is arbitrary in the nozzle. We determine the shock position and establish the existence of transonic shock solution when the boundary data are small perturbations of the special shock solutions under certain conditions. Mathematically, the perturbation problem can be formulated as a free boundary problem of a nonlinear system of hyperbolic-elliptic mixed type and composite. Key difficulties in the analysis mainly comes from the vertical gravity. Methods and techniques are developed in this paper to deal with these key difficulties. Finally, it turns out that the vertical gravity plays a dominant role in the mechanism determining the shock position.

math.AP

GAS: Generative Avatar Synthesis from a Single Image

We present a unified and generalizable framework for synthesizing view-consistent and temporally coherent avatars from a single image, addressing the challenging task of single-image avatar generation. Existing diffusion-based methods often condition on sparse human templates (e.g., depth or normal maps), which leads to multi-view and temporal inconsistencies due to the mismatch between these signals and the true appearance of the subject. Our approach bridges this gap by combining the reconstruction power of regression-based 3D human reconstruction with the generative capabilities of a diffusion model. In a first step, an initial 3D reconstructed human through a generalized NeRF provides comprehensive conditioning, ensuring high-quality synthesis faithful to the reference appearance and structure. Subsequently, the derived geometry and appearance from the generalized NeRF serve as input to a video-based diffusion model. This strategic integration is pivotal for enforcing both multi-view and temporal consistency throughout the avatar's generation. Empirical results underscore the superior generalization ability of our proposed method, demonstrating its effectiveness across diverse in-domain and out-of-domain in-the-wild datasets.

cs.CV

Transonic shock solutions for steady 3-D axisymmetric full Euler flows with large swirl velocity in a finite cylindrical nozzle

This paper concerns the existence and location of three-dimensional axisymmetric transonic shocks with large swirl velocity for shock solutions of the steady compressible full Euler system in a cylindrical nozzle with prescribed receiver pressure. As far as we know, it is the first mathematical result on the three-dimensional transonic shock with either large vorticity or large swirl velocity. One of the key difficulties is the fact that the Euler system is elliptic-hyperbolic composite for the flow behind the shock front, and its elliptic part and hyperbolic part are strongly coupled in the lower order terms because of the large swirl velocity, such that they cannot be simply decoupled in the principal parts as the case for the flow without swirls or with small swirl velocity. New decomposition techniques for the elliptic-hyperbolic composite system are developed to deal with this difficulty, and the solvability condition for the boundary value problem is deduced to determine the location of the shock front. It finally turns out that the non-zero swirl velocity, which brings new challenging difficulties in the analysis, plays an essential and fundamental role in determining the location of the shock front. Another key difficulty in the analysis is that there are no readily established shock solutions available. Non-trivial special shock solutions are first constructed as background solutions under the assumption that the flow parameters depend only on the radial distance to the symmetric axis. Necessary and sufficient conditions for the existence of such solutions are also found. Even though the shock location can be arbitrarily shifted for the special shock solutions, it will be shown that the shock solution with a determined shock position can be established as the boundary data are perturbations of one of the established special shock solutions under certain conditions.

math.AP

Improving Few-shot and Zero-shot Entity Linking with Coarse-to-Fine Lexicon-based Retriever

Few-shot and zero-shot entity linking focus on the tail and emerging entities, which are more challenging but closer to real-world scenarios. The mainstream method is the ''retrieve and rerank'' two-stage framework. In this paper, we propose a coarse-to-fine lexicon-based retriever to retrieve entity candidates in an effective manner, which operates in two layers. The first layer retrieves coarse-grained candidates by leveraging entity names, while the second layer narrows down the search to fine-grained candidates within the coarse-grained ones. In addition, this second layer utilizes entity descriptions to effectively disambiguate tail or new entities that share names with existing popular entities. Experimental results indicate that our approach can obtain superior performance without requiring extensive finetuning in the retrieval stage. Notably, our approach ranks the 1st in NLPCC 2023 Shared Task 6 on Chinese Few-shot and Zero-shot Entity Linking.

cs.CL