SearcharxivSearch

arXiv subjects

Zhengqi Bai

Publications and source records attributed to Zhengqi Bai.

3 recordsLinked to original sources

SenseNova-U1.5: Towards Native Unified Visual Intelligence

We launch SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture. We strengthen its visual interface through spatially coherent patch reconstruction and scale its training with carefully curated generation and editing data, improved task formulation, structural prompt enhancement, and native resolutions of up to 4K. For post-training, we optimize specialized experts for visual aesthetics, bilingual text rendering, infographic generation, and image editing, and consolidate their capabilities through multi-expert on-policy distillation. Across extensive evaluations, SenseNova-U1.5 largely advances image fidelity, text rendering, complex composition, multi-reference editing, and interleaved generation, while improving instruction following and preserving subject identity, geometry, and unmodified regions. Despite limited exposure to structured formats in its generation data, SenseNova-U1.5 generalizes effectively to long, complex, and structured visual instructions, further proving that multimodal understanding can transfer to visual planning and creation. Together, these findings position native unified modelling as a promising path towards systems that perceive, reason and create within a fully end-to-end framework. We will open-source training code, including supervised fine-tuning, reinforcement learning, and on-policy distillation.

cs.CV

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Recent large vision-language models (VLMs) remain fundamentally constrained by a persistent dichotomy: understanding and generation are treated as distinct problems, leading to fragmented architectures, cascaded pipelines, and misaligned representation spaces. We argue that this divide is not merely an engineering artifact, but a structural limitation that hinders the emergence of native multimodal intelligence. Hence, we introduce SenseNova-U1, a native unified multimodal paradigm built upon NEO-unify, in which understanding and generation evolve as synergistic views of a single underlying process. We launch two native unified variants, SenseNova-U1-8B-MoT and SenseNova-U1-A3B-MoT, built on dense (8B) and mixture-of-experts (30B-A3B) understanding baselines, respectively. Designed from first principles, they rival top-tier understanding-only VLMs across text understanding, vision-language perception, knowledge reasoning, agentic decision-making, and spatial intelligence. Meanwhile, they deliver strong semantic consistency and visual fidelity, excelling in conventional or knowledge-intensive any-to-image (X2I) synthesis, complex text-rich infographic generation, and interleaved vision-language generation, with or without think patterns. Beyond performance, we show detailed model design, data preprocessing, pre-/post-training, and inference strategies to support community research. Last but not least, preliminary evidence demonstrates that our models extend beyond perception and generation, performing strongly in vision-language-action (VLA) and world model (WM) scenarios. This points toward a broader roadmap where models do not translate between modalities, but think and act across them in a native manner. Multimodal AI is no longer about connecting separate systems, but about building a unified one and trusting the necessary capabilities to emerge from within.

cs.CV

Geometric Quantum Gates of Non-closed Paths Under Counterdiabatic Driving

Non-adiabatic and non-closed evolutionary paths play a significant role in the fidelity of quantum gates. We propose a high-fidelity quantum control framework based on the quasi-topological number ($ν_{\text{qua}}$), which extends the traditional Chern number to characterize geometric responses in non-closed paths. By introducing a counterdiabatic gauge potential (AGP) that dynamically suppresses non-adiabatic transitions and reconstructs path curvature, we demonstrate that $ν_{\text{qua}}$ -a relative homotopy invariant of compact manifolds in parameter space-quantifies the robustness of geometric phases during open-path quantum evolution. This integer invariant ensures gauge-invariant suppression of decoherence errors arising from dynamical phase coupling. By introducing nonlinear parametric ring paths, we address the defects caused by intermediate states in the Rydberg atomic system. Numerical simulations in the Kitaev superconducting chain and 2D transverse-field Ising model confirm that our protocol achieves quantum gate fidelity exceeding $\mathcal{F} > 0.9999$. We bridges geometric quantum control with topological protection, offering a universal approach to noise-resistant quantum computing.

quant-ph