SearcharxivSearch

arXiv subjects

Junxiang Xu

Publications and source records attributed to Junxiang Xu.

14 recordsLinked to original sources

SenseNova-U1.5: Towards Native Unified Visual Intelligence

We launch SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture. We strengthen its visual interface through spatially coherent patch reconstruction and scale its training with carefully curated generation and editing data, improved task formulation, structural prompt enhancement, and native resolutions of up to 4K. For post-training, we optimize specialized experts for visual aesthetics, bilingual text rendering, infographic generation, and image editing, and consolidate their capabilities through multi-expert on-policy distillation. Across extensive evaluations, SenseNova-U1.5 largely advances image fidelity, text rendering, complex composition, multi-reference editing, and interleaved generation, while improving instruction following and preserving subject identity, geometry, and unmodified regions. Despite limited exposure to structured formats in its generation data, SenseNova-U1.5 generalizes effectively to long, complex, and structured visual instructions, further proving that multimodal understanding can transfer to visual planning and creation. Together, these findings position native unified modelling as a promising path towards systems that perceive, reason and create within a fully end-to-end framework. We will open-source training code, including supervised fine-tuning, reinforcement learning, and on-policy distillation.

cs.CV

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be understood or outputs to be rendered, but first-class substrates for problem solving beyond language. Yet progress remains bottlenecked by the lack of scalable training tasks, reliable feedback, and controlled comparisons across generative substrates. In this work, we introduce VBVR-Pro, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable. 1) Task scaling. VBVR-Pro turns visual reasoning into a controlled task space of 300 procedurally generated tasks. Models trained on VBVR-Pro show strong transfer beyond the proposed suite across seven external visual reasoning benchmarks such as RISE-Video, MME-CoF-Pro, and BabyVision. 2) Verifiable rewards. VBVR-Pro provides verifiable reward scorers for task-grounded evaluation. Through a systematic study of leading MLLMs as judges, we identify recurring failure modes of the prevalent VLM-as-a-judge paradigm. In contrast, the proposed scorers are grounded in deterministic, task-specific rules, achieve fine-grained alignment with human judgments. Importantly, they serve as reliable reward signals for large-scale multi-task reinforcement learning and demonstrate stronger post-RL performance across visual reasoning tasks. 3) Mechanism study. VBVR-Pro enables controlled modality studies across more than 30 image, video, and interleaved generators. Our analysis shows that video generation remains strongest for tasks requiring persistent spatiotemporal state tracking, while interleaved generation provides a compute-efficient alternative. Critically, ablations and probing suggest the presence of vision-native trajectories that are crucial to visual reasoning. We release all data, models, scorers, and code.

cs.CV

ResiliFlow: An Open Transport World Model for Infrastructure Perception and Disaster Resilience

Transport resilience work is often split across separate data preparation scripts, network models, simulation tools, image inspection systems and reports. This fragmentation makes it difficult to move from an observation to a tested and reviewable decision. We introduce ResiliFlow, an open transport world model concept and an implemented platform for infrastructure resilience, response and recovery. The platform connects two workspaces. Disaster Transport Resilience Analysis provides six map-centred functions for critical-road and critical-area identification, recovery prioritisation, disruption routing, resilience testing and scenario simulation. AI-based Transport Infrastructure Perception and Decision Support organises street-level and satellite evidence, detects visible road, footpath and kerb conditions, and prepares these observations for human-reviewed intervention planning. Both workspaces share an eight-step cycle of perception, prediction, model development, verification, execution, decision, feedback and memory. Research Validation records assumptions and checks, while a local Assistant and an optional multi-provider large language model Copilot translate user questions into bounded calls to executable tools. We document the platform architecture, representative mathematical models, interface evidence and computer-vision learning results. Examples show accurate recognition across eight visible-condition classes, while compact error analysis demonstrates how difficult cases guide continued learning. ResiliFlow shows how transport models can become an inspectable, reusable and question-led system rather than a collection of disconnected analyses. The accompanying release is intended to support research collaboration, public scrutiny and extension under institutional review.

math.OC

Quantum percolation theory for dynamic propagation connectivity of transport networks

Connectivity degradation in transport networks under structural disturbance is a central problem in network resilience research. Existing methods rely mainly on percolation theory and topological connectivity measures. They focus on whether paths exist and whether connected components fragment. These approaches cannot capture functional degradation where network topology remains intact but propagation ability has already declined substantially. This paper introduces quantum percolation theory into transport network connectivity analysis and proposes Dynamic Propagation Connectivity (DPC) as a new measure that characterises network propagation ability under disturbance. By mapping a transport network under disturbance into a propagation operator system, this paper establishes a spectral analysis framework for DPC and defines the time-averaged participation index as its core quantification. This paper provides a series of rigorous theoretical results. DPC remains constant under homogeneous disturbance and degrades under heterogeneous disturbance. This paper establishes a quantitative relationship between the degradation rate, the minimum eigenvalue spacing of the propagation operator, and heterogeneous deviation strength. This paper proves a separation theorem between DPC and algebraic connectivity. It derives an analytical expression for DPC and a second-order perturbation approximation on the ring graph. Numerical experiments on three transport benchmark networks verify all theoretical conclusions and confirm degradation monotonicity, separation from algebraic connectivity, and degradation amplification by network size. This paper provides a theoretical framework for transport network resilience assessment that goes beyond topological connectivity.

quant-ph

Quantum percolation based dynamic propagation connectivity for critical-area identification in transport networks

Transport networks often lose functionality through gradual degradation in link operating conditions before topological disconnection occurs. Link-centred and binary percolation measures identify important facilities or connectivity failures, but they provide limited information on which spatial areas cause the largest loss of network-wide propagation capability. This paper develops a Dynamic Propagation Connectivity (DPC) metric based on quantum percolation for critical-area identification in transport networks. Time-varying link travel times are converted into continuous propagation strengths, which define a Hermitian propagation operator at each observation time. Candidate regions are then evaluated by a regional degradation experiment that measures the resulting loss of DPC. The method is applied to a benchmark Sioux Falls network and six Florida road networks during the post-Hurricane Irma disruption and recovery period, using 1,281 five-minute observation times. The benchmark confirms that the regional DPC score identifies a predefined structurally critical corridor. In the Florida networks, the identified critical areas differ from regions selected by link count, local degradation, edge betweenness, algebraic connectivity, and classical percolation. In Networks 1 to 4, DPC and classical percolation rankings have negative Spearman correlations, showing that continuous propagation degradation and binary fragmentation reveal different vulnerability patterns. Robustness tests under alternative travel time scaling, degradation strength, and grid size show stable results, with mean rank agreement between 0.84 and 0.96. The findings extend transport resilience analysis based on percolation from binary connectivity loss to continuous propagation degradation and provide a spatial diagnostic tool for regional monitoring, emergency planning, and recovery prioritisation.

quant-ph

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Recent large vision-language models (VLMs) remain fundamentally constrained by a persistent dichotomy: understanding and generation are treated as distinct problems, leading to fragmented architectures, cascaded pipelines, and misaligned representation spaces. We argue that this divide is not merely an engineering artifact, but a structural limitation that hinders the emergence of native multimodal intelligence. Hence, we introduce SenseNova-U1, a native unified multimodal paradigm built upon NEO-unify, in which understanding and generation evolve as synergistic views of a single underlying process. We launch two native unified variants, SenseNova-U1-8B-MoT and SenseNova-U1-A3B-MoT, built on dense (8B) and mixture-of-experts (30B-A3B) understanding baselines, respectively. Designed from first principles, they rival top-tier understanding-only VLMs across text understanding, vision-language perception, knowledge reasoning, agentic decision-making, and spatial intelligence. Meanwhile, they deliver strong semantic consistency and visual fidelity, excelling in conventional or knowledge-intensive any-to-image (X2I) synthesis, complex text-rich infographic generation, and interleaved vision-language generation, with or without think patterns. Beyond performance, we show detailed model design, data preprocessing, pre-/post-training, and inference strategies to support community research. Last but not least, preliminary evidence demonstrates that our models extend beyond perception and generation, performing strongly in vision-language-action (VLA) and world model (WM) scenarios. This points toward a broader roadmap where models do not translate between modalities, but think and act across them in a native manner. Multimodal AI is no longer about connecting separate systems, but about building a unified one and trusting the necessary capabilities to emerge from within.

cs.CV

Quantum Optimisation for Transport Vulnerability Identification

Transport network vulnerability analysis plays a crucial role in safeguarding urban resilience. Traditional vulnerability identification approaches have provided valuable insights, yet they face two major limitations. First, the number of disruption scenarios increases combinatorially with the number of disrupted links considered simultaneously, making classical approaches computationally prohibitive. Second, most studies approximate the impacts of multiple simultaneous link failures through linear aggregation, which fails to capture the nonlinear interaction effects observed in real networks. To address these gaps, we reformulate the bi-level Mixed-Integer Nonlinear Programming (MINLP) model into a quantum-compatible Quadratic Unconstrained Binary Optimisation (QUBO) structure, enabling parallel exploration of complex disruption scenarios while incorporating nonlinear interaction effects. We develop a hybrid optimisation framework that integrates the quantum optimisation algorithm with the Frank-Wolfe method to validate the model's effectiveness on the small-scale network. Then, we further verify the framework through the D-Wave hardware across benchmark networks of different scales, including Sioux Falls, Anaheim, Chicago Sketch, and Berlin Full, to examine scalability and feasibility. The results show that this framework achieves strong solvability and stability. In particular, optimisation for large and larger networks is completed within minutes (Approximately 2.8 minutes for the 914-link, 9.8 minutes for the 2950-link, and 31.2 minutes for the 6018-link on D-Wave), demonstrating a computational efficiency improvement by one to two orders of magnitude compared with classical metaheuristic algorithms. These findings highlight the feasibility and potential of applying quantum computing to network vulnerability identification and open a new avenue for resilience-oriented planning.

math.OC

Quantum optimisation in cities: Limitations and prospects of urban transport systems

Recently, quantum computing has gained attention in urban studies as a tool for complex transport planning problems, but its role remains unclear. This paper reviews quantum computing research in urban transport planning and highlights major limits in scalability, robustness, constraint handling, and engineering feasibility.Stable and reproducible advantages of quantum optimisation in real urban systems have yet to be shown. By comparing quantum methods with established classical optimisation methods, it is found that decomposition methods, metaheuristics, and reinforcement learning already provide transparent, scalable, and policy-interpretable solutions for medium and large-sized urban transport networks. In contrast, the contribution of quantum methods largely lies in the exploratory analysis of limited, discrete combinatorial subproblems rather than full system-level optimisation. It is argued in this paper for a shift from technology-driven application narrative towards problem-driven method selection. From an urban transport planning perspective, we have identified the specific problem types where the exploratory use of quantum computing may be relevant, including critical link and node vulnerability identification, combinatorial screening of congestion and failure scenarios, disaster-related condition analysis, constrained path option selection, and small-scale facility location and investment option assessment. It is concluded that hybrid frameworks represent a more realistic pathway for integrating quantum computing into urban transport research, in which classical methods ensure systemlevel consistency and policy interpretability while quantum methods support local combinatorial exploration. Until stable engineering advantages are demonstrated, public agencies and researchers should prioritise method validation, scenario suitability, and cross-disciplinary collaboration.

math.OC

Data-driven identification of critical links in transport networks using quantum annealing

In urban transport systems, time-varying demand and network conditions cause the importance of infrastructure elements to evolve, requiring the identification of period-specific critical links to support systemlevel risk and resilience analysis. However, static or time-averaged network analyses struggle to capture the temporal variation of infrastructure importance at the city scale. To address this gap, this study proposes a time-dependent critical link identification framework for large-scale urban transport networks. The problem is formulated as a Quadratic Unconstrained Binary Optimisation (QUBO) model and solved using quantum annealing on D-Wave hardware. Empirical analysis using real-world traffic data reveals a strong temporal concentration of critical links. Rather than persistently influencing system performance, critical links emerge mainly within a small number of key time windows, during which even limited disruptions can lead to substantial network delay amplification. These findings demonstrate the value of time-dependent analysis for risk screening, stress testing, and resilience-oriented transport management.

math.OC

MonoArt: Progressive Structural Reasoning for Monocular Articulated 3D Reconstruction

Reconstructing articulated 3D objects from a single image requires jointly inferring object geometry, part structure, and motion parameters from limited visual evidence. A key difficulty lies in the entanglement between motion cues and object structure, which makes direct articulation regression unstable. Existing methods address this challenge through multi-view supervision, retrieval-based assembly, or auxiliary video generation, often sacrificing scalability or efficiency. We present MonoArt, a unified framework grounded in progressive structural reasoning. Rather than predicting articulation directly from image features, MonoArt progressively transforms visual observations into canonical geometry, structured part representations, and motion-aware embeddings within a single architecture. This structured reasoning process enables stable and interpretable articulation inference without external motion templates or multi-stage pipelines. Extensive experiments on PartNet-Mobility demonstrate that OM achieves state-of-the-art performance in both reconstruction accuracy and inference speed. The framework further generalizes to robotic manipulation and articulated scene reconstruction.

cs.CV

Demystifying Video Reasoning

Recent advances in video generation have revealed an unexpected phenomenon: diffusion-based video models exhibit non-trivial reasoning capabilities. Prior work attributes this to a Chain-of-Frames (CoF) mechanism, where reasoning is assumed to unfold sequentially across video frames. In this work, we challenge this assumption and uncover a fundamentally different mechanism. We show that reasoning in video models instead primarily emerges along the diffusion denoising steps. Through qualitative analysis and targeted probing experiments, we find that models explore multiple candidate solutions in early denoising steps and progressively converge to a final answer, a process we term Chain-of-Steps (CoS). Beyond this core mechanism, we identify several emergent reasoning behaviors critical to model performance: (1) working memory that supports tasks requiring consistent reference, such as object permanence; (2) self-correction and enhancement, allowing recovery from incorrect intermediate solutions; and (3) perception before action, where early steps establish semantic grounding and later steps perform structured manipulation. Moreover, analysis of Diffusion Transformer layers shows that middle layers conduct key reasoning procedures. Motivated by these insights, we present a simple Training-Free Ensemble (TFE) as a proof-of-concept, demonstrating how reasoning can be improved by ensembling latent trajectories from identical models with different random seeds. Overall, our work provides the first systematic dissection of the mechanisms underlying video reasoning, offering a foundation to guide future research in better exploiting the inherent reasoning dynamics of video models as a new substrate for intelligence.

cs.CV

Scaling Spatial Intelligence with Multimodal Foundation Models

Despite remarkable progress, multimodal foundation models still exhibit surprising deficiencies in spatial intelligence. In this work, we explore scaling up multimodal foundation models to cultivate spatial intelligence within the SenseNova-SI family, built upon established multimodal foundations including visual understanding models (i.e., Qwen3-VL and InternVL3) and unified understanding and generation models (i.e., Bagel). We take a principled approach to constructing high-performing and robust spatial intelligence by systematically curating SenseNova-SI-8M: eight million diverse data samples under a rigorous taxonomy of spatial capabilities. SenseNova-SI demonstrates unprecedented performance across a broad range of spatial intelligence benchmarks: 68.8% on VSI-Bench, 43.3% on MMSI, 85.7% on MindCube, 54.7% on ViewSpatial, 47.7% on SITE, 63.9% on BLINK, 55.5% on 3DSR, and 72.0% on EmbSpatial, while maintaining strong general multimodal understanding (e.g., 84.9% on MMBench-En). More importantly, we analyze the impact of data scaling, discuss early signs of emergent generalization capabilities enabled by diverse data training, analyze the risk of overfitting and language shortcuts, present a preliminary study on spatial chain-of-thought reasoning, and validate the potential downstream application. All newly trained multimodal foundation models are publicly released.

cs.CV

General KAM theorems and their applications to invariant tori with prescribed frequencies

In this paper we develop some new KAM-technique to prove two general KAM theorems for nearly integrable hamiltonian systems without assuming any non-degeneracy condition. Many of KAM-type results (including the classical KAM theorem) are special cases of our theorems under some non-degeneracy condition and some smoothness condition. Moreover, we can obtain some interesting results about KAM tori with prescribed frequencies.

math.DS