SearcharxivSearch

arXiv subjects

Zheng Tang

Publications and source records attributed to Zheng Tang.

At least 19 recordsLinked to original sources

The 10th AI City Challenge

The 10th AI City Challenge, held with ECCV 2026, marks a decade of community benchmarking for intelligent transportation, smart cities, and physical AI. Since its 2017 start with vehicle detection, classification, and tracking, the challenge has grown into a broad benchmark suite for multi-camera perception, multimodal reasoning, synthetic-to-real learning, generative forecasting, and privacy-preserving evaluation. The 2026 edition continued this growth with 325 registered teams, up from 245 in 2025, and participation from 26 countries and regions, up from 15. Its six primary tracks cover multi-camera 3D perception, transportation safety captioning and VQA, traffic anomaly reasoning, text-based person anomaly search, generative traffic video forecasting, and cross-city object detection. Track 3 further includes two out-of-domain leaderboards, submitted as Tracks 7 and 8, for fisheye traffic-violation understanding and pedestrian situated-intent VQA. This paper summarizes the challenge setup, datasets, evaluation protocols, leaderboard results, and workshop papers. Across tracks, successful systems combine foundation models with geometric grounding, retrieval or reranking, synthetic-data design, domain adaptation, and controlled inference.

cs.CV

From Detection to Understanding: TAR and TAR-Bench for Multi-Task Traffic Anomaly Reasoning

We present TAR (Traffic Anomaly Reasoning) and TAR-Bench datasets, resources for training and evaluating video-language models beyond anomaly detection. TAR contains 44,040 chain-of-thought training annotations across 10 tasks for 3,670 CCTV videos ($\sim$26 hours) from eight public datasets. Its evaluation component, TAR-Bench, contains 960 human-curated test annotations for 80 held-out clips trimmed from 17 public YouTube videos. TAR's training annotations are produced with MAVEN, which consolidates multi-scale video evidence into structured event descriptions before generating question-answer pairs and reasoning traces. On TAR-Bench, eleven vision-language models reveal that strong question-answering accuracy does not reliably predict temporal or scene reasoning ability. Multi-task fine-tuning on TAR yields consistent gains, with the full 10-task model improving aggregate score by 21.4 points over its zero-shot baseline. TAR and TAR-Bench provide the official training and in-domain evaluation data for AI City Challenge 2026 Track 3. The dataset is available at https://huggingface.co/datasets/nvidia/PhysicalAI-Traffic-Anomaly-Reasoning

cs.CV

3D Topologically Polarized Elastic Metamaterials Enable Asymmetric Energy Isolation at Low Frequencies

Topologically polarized elasticity has been extensively studied in lower-dimensions, yet its three-dimensional (3D) counterpart remains largely unexplored. Here, we demonstrate omnidirectional topological elasticity in 3D structures that incorporate bending stiffness, which elevates zero-frequency topological mechanical states into finite-frequency phononic modes. These modes are localized at a single boundary, creating a pronounced stiffness contrast in both static and finite-frequency dynamic regimes. This three-dimensional structure exhibits highly polarized mechanical behavior across all spatial dimensions, establishing omnidirectional asymmetric topological elasticity. Experimental and numerical results confirm robust, asymmetric energy isolation, arising from the interplay between bulk topological polarization and boundary-localized surface modes. Our findings establish a paradigm for 3D metamaterials, with promising applications in vibration shielding and directional wave manipulation.

cond-mat.soft

Finding accurate eigenvalues and eigenvectors of positive semi-definite matrices given a subspace

We revisit a classical problem in numerical linear algebra: given an $k$-dimensional subspace $\mathcal{Q}$ that approximates the leading eigenspace of an $n\times n$ positive semi-definite matrix $A$, the goal is to extract high-accuracy eigenvalues. The Rayleigh-Ritz (RR) method is the standard algorithm for the task, which has been shown to be optimal in several ways (when $A$ is symmetric, not necessarily positive semi-definite $A\succeq 0$). In this paper, we show that when $A \succeq 0$, alternative methods can outperform RR, while having the same computational complexity, that is, the main cost is in computing $AQ$, plus an $O(nk^2)$ term. In particular, we advocate the use of Nystr{\"o}m's method, showing that the approximate eigenvalues always have higher accuracy than RR, and the improvement can be arbitrarily large. The difference is significant, especially when $A$ has a fast-decaying spectrum. A similar improvement is numerically observed for the purpose of approximating the leading eigenvectors. In contrast, when the target eigenvalues are the trailing ones, the situation is reversed, and the Nystr{\"o}m method performs poorly; we suggest a remedy for this situation.

math.NA

Model Optimization for Multi-Camera 3D Detection and Tracking

Outside-in multi-camera perception is increasingly important in indoor environments, where networks of static cameras must support multi-target tracking under occlusion and heterogeneous viewpoints. We evaluate Sparse4D, a query-based spatiotemporal 3D detection and tracking framework that fuses multi-view features in a shared world frame and propagates sparse object queries via instance memory. We study reduced input frame rates, post-training quantization (INT8 and FP8), transfer to the WILDTRACK benchmark, and Transformer Engine mixed-precision fine-tuning. To better capture identity stability, we report Average Track Duration (AvgTrackDur), which measures identity persistence in seconds. Sparse4D remains stable under moderate FPS reductions, but below 2 FPS, identity association collapses even when detections are stable. Selective quantization of the backbone and neck offers the best speed-accuracy trade-off, while attention-related modules are consistently sensitive to low precision. On WILDTRACK, low-FPS pretraining yields large zero-shot gains over the base checkpoint, while small-scale fine-tuning provides limited additional benefit. Transformer Engine mixed precision reduces latency and improves camera scalability, but can destabilize identity propagation, motivating stability-aware validation.

cs.CV

A Unified 3D Object Perception Framework for Real-Time Outside-In Multi-Camera Systems

Accurate 3D object perception and multi-target multi-camera (MTMC) tracking are fundamental for the digital transformation of industrial infrastructure. However, transitioning "inside-out" autonomous driving models to "outside-in" static camera networks presents significant challenges due to heterogeneous camera placements and extreme occlusion. In this paper, we present an adapted Sparse4D framework specifically optimized for large-scale infrastructure environments. Our system leverages absolute world-coordinate geometric priors and introduces an occlusion-aware ReID embedding module to maintain identity stability across distributed sensor networks. To bridge the Sim2Real domain gap without manual labeling, we employ a generative data augmentation strategy using the NVIDIA COSMOS framework, creating diverse environmental styles that enhance the model's appearance-invariance. Evaluated on the AI City Challenge 2025 benchmark, our camera-only framework achieves a state-of-the-art HOTA of $45.22$. Furthermore, we address real-time deployment constraints by developing an optimized TensorRT plugin for Multi-Scale Deformable Aggregation (MSDA). Our hardware-accelerated implementation achieves a $2.15\times$ speedup on modern GPU architectures, enabling a single Blackwell-class GPU to support over 64 concurrent camera streams.

cs.CV

Deep g-Pricing for CSI 300 Index Options with Volatility Trajectories and Market Sentiment

Option pricing in real markets faces fundamental challenges. The Black--Scholes--Merton (BSM) model assumes constant volatility and uses a linear generator $g(t,x,y,z)=-ry$, while lacking explicit behavioral factors, resulting in systematic departures from observed dynamics. This paper extends the BSM model by learning a nonlinear generator within a deep Forward--Backward Stochastic Differential Equation (FBSDE) framework. We propose a dual-network architecture where the value network $u_\theta$ learns option prices and the generator network $g_\phi$ characterizes the pricing mechanism, with the hedging strategy $Z_t=\sigma_t X_t \nabla_x u_\theta$ obtained via automatic differentiation. The framework adopts forward recursion from a learnable initial condition $Y_0=u_\theta(0,\cdot)$, naturally accommodating volatility trajectory and sentiment features. Empirical results on CSI 300 index options show that our method reduces Mean Absolute Error (MAE) by 32.2\% and Mean Absolute Percentage Error (MAPE) by 35.3\% compared with BSM. Interpretability analysis indicates that architectural improvements are effective across all option types, while the information advantage is asymmetric between calls and puts. Specifically, call option improvements are primarily driven by sentiment features, whereas put options show more balanced contributions from volatility trajectory and sentiment features. This finding aligns with economic intuition regarding option pricing mechanisms.

q-fin.CP

The 9th AI City Challenge

The ninth AI City Challenge continues to advance real-world applications of computer vision and AI in transportation, industrial automation, and public safety. The 2025 edition featured four tracks and saw a 17% increase in participation, with 245 teams from 15 countries registered on the evaluation server. Public release of challenge datasets led to over 30,000 downloads to date. Track 1 focused on multi-class 3D multi-camera tracking, involving people, humanoids, autonomous mobile robots, and forklifts, using detailed calibration and 3D bounding box annotations. Track 2 tackled video question answering in traffic safety, with multi-camera incident understanding enriched by 3D gaze labels. Track 3 addressed fine-grained spatial reasoning in dynamic warehouse environments, requiring AI systems to interpret RGB-D inputs and answer spatial questions that combine perception, geometry, and language. Both Track 1 and Track 3 datasets were generated in NVIDIA Omniverse. Track 4 emphasized efficient road object detection from fisheye cameras, supporting lightweight, real-time deployment on edge devices. The evaluation framework enforced submission limits and used a partially held-out test set to ensure fair benchmarking. Final rankings were revealed after the competition concluded, fostering reproducibility and mitigating overfitting. Several teams achieved top-tier results, setting new benchmarks in multiple tasks.

cs.CV

Dynamical generation of geometric squeezing in interacting Bose-Einstein condensates

When the rotating frequency of a non-interacting Bose-Einstein condensate (BEC) confined in a weak anisotropic harmonic potential is suddenly quenched to its trapping frequency, the condensate evolves from its ground state to a single-mode squeezed state with exponentially growing quantum fluctuation anisotropy. Such a squeezed state is called the geometrically squeezed state. However, for interacting BECs with two-body collisions, a similar quench only results in quantum fluctuations oscillating periodically without squeezing. In this work, we identify superfluid stability as the key factor behind this non-squeezing phenomenon, with the periodic oscillations arising from collective excitations of a stable collective excitation mode. By strategically breaking the stability criteria, we propose a dynamical approach for generating squeezing that can exponentially suppress quantum fluctuations in a relatively short time, surpassing the efficiency of existing experimental preparation schemes.

cond-mat.quant-gas

Four-mode quantum sensing and Fisher information in a spin-orbit-coupled Bose gas

Multi-mode squeezing and entanglement are important resources in quantum metrology and sensing. For spin-1/2 Bose-Einstein condensates subject to spin-orbit coupling (SOC), previous studies on spin squeezing have been limited to two-mode systems. In this work, we demonstrate that such a system can naturally construct a four-mode model spanning an $\mathfrak{su}(4)$ algebra with six SU(2) subspaces. Using spin squeezing parameters and quantum Fisher information matrices, we analyze the dynamical evolution of coherent spin states. The results show that, beyond two-mode models, the SOC-induced four-mode couplings give rise to richer entanglement-enhanced sensing approaching the Heisenberg limit across various SU(2) subspaces. Additionally, by tuning a single system parameter (the Raman Rabi frequency), one can selectively control the optimal measurement directions across different subspaces.

cond-mat.quant-gas

Berry-Esseen bound for the Moment Estimation of the fractional Ornstein-Uhlenbeck model under fixed step size discrete observations

Let the Ornstein-Uhlenbeck process $\{X_t,\,t\geq 0\}$ driven by a fractional Brownian motion $B^H$ described by $d X_t=-\theta X_t dt+ d B_t^H,\, X_0=0$ with known parameter $H\in (0,\frac34)$ be observed at discrete time instants $t_k=kh, k=1,2,\dots, n $. If $\theta>0$ and if the step size $h>0$ is arbitrarily fixed, we derive Berry-Ess\'{e}en bound for the ergodic type estimator (or say the moment estimator) $\hat{\theta}_n$, i.e., the Kolmogorov distance between the distribution of $\sqrt{n}(\hat{\theta}_n-\theta)$ and its limit distribution is bounded by a constant $C_{\theta, H,h}$ times $n^{-\frac12}$ and $ n^{4H-3}$ when $H\in (0,\,\frac58]$ and $H\in (\frac58,\,\frac34)$, respectively. This result greatly improve the previous result in literature where $h$ is forced to go zero. Moreover, we extend the Berry-Esseen bound to the Ornstein-Uhlenbeck model driven by a lot of Gaussian noises such as the sub-bifractional Brownian motion and others. A few ideas of the present paper come from Haress and Hu (2021), Sottinen and Viitasaari (2018), and Chen and Zhou (2021).

math.PR

Dynamic Noise Preference Optimization: Self-Improvement of Large Language Models with Self-Synthetic Data

Although LLMs have achieved significant success, their reliance on large volumes of human-annotated data has limited their potential for further scaling. In this situation, utilizing self-generated synthetic data has become crucial for fine-tuning LLMs without extensive human annotation. However, current methods often fail to ensure consistent improvements across iterations, with performance stagnating after only minimal updates. To overcome these challenges, we introduce Dynamic Noise Preference Optimization (DNPO), which combines dynamic sample labeling for constructing preference pairs with controlled, trainable noise injection during preference optimization. Our approach effectively prevents stagnation and enables continuous improvement. In experiments with Llama-3.2-3B and Zephyr-7B, DNPO consistently outperforms existing methods across multiple benchmarks. Additionally, with Zephyr-7B, DNPO shows a significant improvement in model-generated data quality, with a 29.4% win-loss rate gap compared to the baseline in GPT-4 evaluations.

cs.CL

ToMoE: Converting Dense Large Language Models to Mixture-of-Experts through Dynamic Structural Pruning

Large Language Models (LLMs) have demonstrated remarkable abilities in tackling a wide range of complex tasks. However, their huge computational and memory costs raise significant challenges in deploying these models on resource-constrained devices or efficiently serving them. Prior approaches have attempted to alleviate these problems by permanently removing less important model structures, yet these methods often result in substantial performance degradation due to the permanent deletion of model parameters. In this work, we tried to mitigate this issue by reducing the number of active parameters without permanently removing them. Specifically, we introduce a differentiable dynamic pruning method that pushes dense models to maintain a fixed number of active parameters by converting their MLP layers into a Mixture of Experts (MoE) architecture. Our method, even without fine-tuning, consistently outperforms previous structural pruning techniques across diverse model families, including Phi-2, LLaMA-2, LLaMA-3, and Qwen-2.5.

cs.LG

MCBLT: Multi-Camera Multi-Object 3D Tracking in Long Videos

Object perception from multi-view cameras is crucial for intelligent systems, particularly in indoor environments, e.g., warehouses, retail stores, and hospitals. Most traditional multi-target multi-camera (MTMC) detection and tracking methods rely on 2D object detection, single-view multi-object tracking (MOT), and cross-view re-identification (ReID) techniques, without properly handling important 3D information by multi-view image aggregation. In this paper, we propose a 3D object detection and tracking framework, named MCBLT, which first aggregates multi-view images with necessary camera calibration parameters to obtain 3D object detections in bird's-eye view (BEV). Then, we introduce hierarchical graph neural networks (GNNs) to track these 3D detections in BEV for MTMC tracking results. Unlike existing methods, MCBLT has impressive generalizability across different scenes and diverse camera settings, with exceptional capability for long-term association handling. As a result, our proposed MCBLT establishes a new state-of-the-art on the AICity'24 dataset with $81.22$ HOTA, and on the WildTrack dataset with $95.6$ IDF1.

cs.CV

Fully-Polarized Topological Isostatic Metamaterials in Three Dimensions

Topological surface states are unique to topological materials and are immune to disturbances. In isostatic lattices, mechanical topological floppy modes exhibit softness depending on the polarization relative to the terminating surface. However, in three dimensions, the polarization of topological floppy modes is disrupted by the ubiquitous mechanical Weyl lines. Here, we demonstrate, both theoretically and experimentally, the fully-polarized topological mechanical phases free of Weyl lines. Floppy modes emerge exclusively on a particular surface of the three-dimensional isostatic structure, leading to the strongly asymmetric stiffness between opposing boundaries. Additionally, uniform soft strains can reversibly shift the lattice configuration to Weyl phases, reducing the stiffness contrast to a trivially comparable level. Our work demonstrates the fully-polarized topological mechanical phases in three dimensions, and paves the way towards engineering soft and adaptive metamaterials.

cond-mat.soft

Radiance Field Learners As UAV First-Person Viewers

First-Person-View (FPV) holds immense potential for revolutionizing the trajectory of Unmanned Aerial Vehicles (UAVs), offering an exhilarating avenue for navigating complex building structures. Yet, traditional Neural Radiance Field (NeRF) methods face challenges such as sampling single points per iteration and requiring an extensive array of views for supervision. UAV videos exacerbate these issues with limited viewpoints and significant spatial scale variations, resulting in inadequate detail rendering across diverse scales. In response, we introduce FPV-NeRF, addressing these challenges through three key facets: (1) Temporal consistency. Leveraging spatio-temporal continuity ensures seamless coherence between frames; (2) Global structure. Incorporating various global features during point sampling preserves space integrity; (3) Local granularity. Employing a comprehensive framework and multi-resolution supervision for multi-scale scene feature representation tackles the intricacies of UAV video spatial scales. Additionally, due to the scarcity of publicly available FPV videos, we introduce an innovative view synthesis method using NeRF to generate FPV perspectives from UAV footage, enhancing spatial perception for drones. Our novel dataset spans diverse trajectories, from outdoor to indoor environments, in the UAV domain, differing significantly from traditional NeRF scenarios. Through extensive experiments encompassing both interior and exterior building structures, FPV-NeRF demonstrates a superior understanding of the UAV flying space, outperforming state-of-the-art methods in our curated UAV dataset. Explore our project page for further insights: https://fpv-nerf.github.io/.

cs.CV

Paraphrase and Aggregate with Large Language Models for Minimizing Intent Classification Errors

Large language models (LLM) have achieved remarkable success in natural language generation but lesser focus has been given to their applicability in decision making tasks such as classification. We show that LLMs like LLaMa can achieve high performance on large multi-class classification tasks but still make classification errors and worse, generate out-of-vocabulary class labels. To address these critical issues, we introduce Paraphrase and AGgregate (PAG)-LLM approach wherein an LLM generates multiple paraphrases of the input query (parallel queries), performs multi-class classification for the original query and each paraphrase, and at the end aggregate all the classification labels based on their confidence scores. We evaluate PAG-LLM on two large multi-class classication datasets: CLINC, and Banking and show 22.7% and 15.1% error reduction. We show that PAG-LLM is especially effective for hard examples where LLM is uncertain, and reduces the critical misclassification and hallucinated label generation errors

cs.CL

Partial confinement in a quantum-link simulator

Confinement/deconfinement, captivating attributes of high-energy elementary particles, have recently garnered wide attention in quantum simulations based on cold atoms. Yet, the partial confinement, an intermediate state between the confinement and deconfinement, remains underexplored. The partial confinement encapsulates the phenomenon that the confining behavior of charged particles is contingent upon their relative positions. In this paper, we demonstrate that the spin-1 quantum link model provides an excellent platform for exploring partial confinement. We conduct a comprehensive investigation of the physics emerging from partial confinement in both the context of equilibrium and non-equilibrium dynamics. Potential experimental setups using cold atoms are also discussed. Our work offers a simple and feasible routine for the study of confinement-related physics in the state-of-the-art artificial quantum systems subject to gauge symmetries.

cond-mat.quant-gas