SearcharxivSearch

arXiv subjects

Weijun Zhang

Publications and source records attributed to Weijun Zhang.

At least 19 recordsLinked to original sources

Geometry and Gradient-based Partitioning for Panoramic Outdoor Reconstruction

Scaling 3D Gaussian Splatting (3DGS) to large outdoor scenes is costly in both data acquisition and computation. Adopting panoramic images with equirectangular projection (ERP) can reduce capture effort via their full $360^{\circ}$ field of view, yet the resulting omnipresent visibility invalidates existing partitioning strategies that rely on local camera frustums, causing block-wise optimization to degenerate into global training. Thus, we propose PanoLOG, a two-stage coarse-to-fine framework equipped with a Geometry and Gradient-based Partitioning Strategy tailored for large-scale panoramic 3DGS reconstruction. In the global coarse stage, PanoLOG leverages sky-sphere modeling and panoramic monocular depth supervision for reliable geometry, while in the refinement stage, G$^2$PS builds adaptive bounding volumes via parallax-driven uncertainty and assigns cameras via gradient-based importance scoring. Furthermore, we construct Pano360, the first benchmark on large-scale panoramic dataset for outdoor scene reconstruction. Extensive experiments demonstrate that G$^2$PS achieves state-of-the-art rendering quality while maintaining scalable, block-parallel training. Our models, training code, and dataset are publicly available.

cs.CV

A Decomposition Lemma in Convex Integration via Classical Algebraic Geometry

In this paper, we prove a decomposition lemma for symmetric matrix fields on bounded domains: $D+\mathrm{Sym}\nablaΦ=\sum_i a_i^2ξ_i\otimesξ_i$ with uniform control on $Φ$ and $a_i^2$, using fewer than the usual $n(n+1)/2$ rank-one symmetric terms. Except possibly in dimensions $n=8,16$, the decomposition is shown to be optimal through algebraic arguments. This reduces the number of steps in convex integration for a nonlinear PDE system, improving Hölder regularity of flexible solutions in dimension $n\ge3$. This PDE is a partial linearization of the codimension-one local isometric embedding equation in the Nash--Kuiper theorem, and also yields improved regularity for very weak solutions of related 2D Monge--Ampére and $2$-Hessian systems. The improved Hölder exponent is any $α<(n^2+1)^{-1}$ for $n=2,4,8,16$ and any $α<(n^2+n-2ρ(n/2)-1)^{-1}$ otherwise, where $ρ$ is the Radon--Hurwitz number, related to Bott periodicity. The proof involves novel applications of algebraic geometry and topology that yield the optimality of decomposition, including Adams' theorem on vector fields on spheres, intersections of projective varieties, and projective duality, combined with an elliptic method that avoids loss of differentiability.

math.AP

Some variational problems for the complex Monge--Amp{è}re operator

We consider the Dirichlet problem for the complex Monge--Ampère equation on strongly pseudoconvex Kähler manifolds when the right-hand side is decreasing in the solution. Using flow-based arguments, we establish existence of smooth solutions in a number of natural circumstances, following work of Chou-Wang.

math.CV

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models

Large vision-language models (VLMs) have demonstrated remarkable capabilities in open-world multimodal understanding, yet their high computational overheads pose great challenges for practical deployment. Some recent works have proposed methods to accelerate VLMs by pruning redundant visual tokens guided by the attention maps of VLM's early layers. Despite the success of these token pruning methods, they still suffer from two major shortcomings: (i) considerable accuracy drop due to insensitive attention signals in early layers, and (ii) limited speedup when generating long responses (e.g., 30 tokens). To address the limitations above, we present TwigVLM -- a simple and general architecture by growing a lightweight module, named twig, upon an early layer of the base VLM. Compared with most existing VLM acceleration methods purely based on visual token pruning, our TwigVLM not only achieves better accuracy retention by employing a twig-guided token pruning (TTP) strategy, but also yields higher generation speed by utilizing a self-speculative decoding (SSD) strategy. Taking LLaVA-1.5-7B as the base VLM, experimental results show that TwigVLM preserves 96% of the original performance after pruning 88.9% of the visual tokens and achieves 154% speedup in generating long responses, delivering significantly better performance in terms of both accuracy and speed over the state-of-the-art VLM acceleration methods. Moreover, we extend TwigVLM to an improved TwigVLM++ variant by introducing a novel multi-head twig architecture with a specialized pruning head. TwigVLM++ improves pruning quality via a two-stage training paradigm combining a distillation learning stage and a pruning-oriented reinforcement learning stage, and further accelerates inference via a tree-based SSD strategy.

cs.CV

Multimodal OCR: Parse Anything from Documents

We present Multimodal OCR (MOCR), a document parsing paradigm that jointly parses text and graphics into unified textual representations. Unlike conventional OCR systems that focus on text recognition and leave graphical regions as cropped pixels, our method, termed dots.mocr, treats visual elements such as charts, diagrams, tables, and icons as first-class parsing targets, enabling systems to parse documents while preserving semantic relationships across elements. It offers several advantages: (1) it reconstructs both text and graphics as structured outputs, enabling more faithful document reconstruction; (2) it supports end-to-end training over heterogeneous document elements, allowing models to exploit semantic relations between textual and visual components; and (3) it converts previously discarded graphics into reusable code-level supervision, unlocking multimodal supervision embedded in existing documents. To make this paradigm practical at scale, we build a comprehensive data engine from PDFs, rendered webpages, and native SVG assets, and train a compact 3B-parameter model through staged pretraining and supervised fine-tuning. We evaluate dots.mocr from two perspectives: document parsing and structured graphics parsing. On document parsing benchmarks, it ranks second only to Gemini 3 Pro on our OCR Arena Elo leaderboard, surpasses existing open-source document parsing systems, and sets a new state of the art of 83.9 on olmOCR Bench. On structured graphics parsing, our model achieves higher reconstruction quality than Gemini 3 Pro across image-to-SVG benchmarks, demonstrating strong performance on charts, UI layouts, scientific figures, and chemical diagrams. These results show a scalable path toward building large-scale image-to-code corpora for multimodal pretraining. Code and models are publicly available at https://github.com/rednote-hilab/dots.mocr.

cs.CV

Access in the Shadow of Ableism: An Autoethnography of a Blind Student's Higher Education Experience in China

The HCI research community has witnessed a growing body of research on accessibility and disability driven by efforts to improve access. Yet, the concept of access reveals its limitations when examined within broader ableist structures. Drawing on an autoethnographic method, this study shares the co-first author Zhang's experiences at two higher-education institutions in China, including a specialized program exclusively for blind and low-vision students and a mainstream university where he was the first blind student admitted. Our analysis revealed tensions around access in both institutions: they either marginalized blind students within society at large or imposed pressures to conform to sighted norms. Both institutions were further constrained by systemic issues, including limited accessible resources, pervasive ableist cultures, and the lack of formalized policies. In response to these tensions, we conceptualize access as a contradictory construct and argue for understanding accessibility as an ongoing, exploratory practice within ableist structures.

cs.HC

Manifold Function Encoder: Identifying Different Functions Defined on Different Manifolds

We propose the Manifold Function Encoder (MFE) for identifying different functions defined on different manifolds. Both a manifold in Euclidean space and a function defined on this manifold can be viewed as bounded linear functionals on a suitable space of continuous functions. From this perspective, we treat manifold functions as elements of the dual space. By expanding them in the dual space based on appropriate approximating sequence of bases, we obtain a corresponding method for encoding manifold functions, that is MFE. Especially, we prove that MFE achieves super-algebraic convergence based on smooth bases commonly used in spectral methods, such as Legendre polynomials and Fourier basis. We further extend MFE to handle more complex cases, including joint manifold functions of different dimensions and manifold functions with different measures. In addition, we show the approximation theory for MFE-based operator learning, in particular learning the solution mappings of PDEs defined on varying domains, together with several numerical experiments including the 2-d Poisson equation and the 3-d elasticity problem on the real-world bearing.

math.NA

AirSim360: A Panoramic Simulation Platform within Drone View

The field of 360-degree omnidirectional understanding has been receiving increasing attention for advancing spatial intelligence. However, the lack of large-scale and diverse data remains a major limitation. In this work, we propose AirSim360, a simulation platform for omnidirectional data from aerial viewpoints, enabling wide-ranging scene sampling with drones. Specifically, AirSim360 focuses on three key aspects: a render-aligned data and labeling paradigm for pixel-level geometric, semantic, and entity-level understanding; an interactive pedestrian-aware system for modeling human behavior; and an automated trajectory generation paradigm to support navigation tasks. Furthermore, we collect more than 60K panoramic samples and conduct extensive experiments across various tasks to demonstrate the effectiveness of our simulator. Unlike existing simulators, our work is the first to systematically model the 4D real world under an omnidirectional setting. The entire platform, including the toolkit, plugins, and collected datasets, will be made publicly available at https://insta360-research-team.github.io/AirSim360-website.

cs.CV

dots.llm1 Technical Report

Mixture of Experts (MoE) models have emerged as a promising paradigm for scaling language models efficiently by activating only a subset of parameters for each input token. In this report, we present dots.llm1, a large-scale MoE model that activates 14B parameters out of a total of 142B parameters, delivering performance on par with state-of-the-art models while reducing training and inference costs. Leveraging our meticulously crafted and efficient data processing pipeline, dots.llm1 achieves performance comparable to Qwen2.5-72B after pretraining on 11.2T high-quality tokens and post-training to fully unlock its capabilities. Notably, no synthetic data is used during pretraining. To foster further research, we open-source intermediate training checkpoints at every one trillion tokens, providing valuable insights into the learning dynamics of large language models.

cs.CL

CFTrack: Enhancing Lightweight Visual Tracking through Contrastive Learning and Feature Matching

Achieving both efficiency and strong discriminative ability in lightweight visual tracking is a challenge, especially on mobile and edge devices with limited computational resources. Conventional lightweight trackers often struggle with robustness under occlusion and interference, while deep trackers, when compressed to meet resource constraints, suffer from performance degradation. To address these issues, we introduce CFTrack, a lightweight tracker that integrates contrastive learning and feature matching to enhance discriminative feature representations. CFTrack dynamically assesses target similarity during prediction through a novel contrastive feature matching module optimized with an adaptive contrastive loss, thereby improving tracking accuracy. Extensive experiments on LaSOT, OTB100, and UAV123 show that CFTrack surpasses many state-of-the-art lightweight trackers, operating at 136 frames per second on the NVIDIA Jetson NX platform. Results on the HOOT dataset further demonstrate CFTrack's strong discriminative ability under heavy occlusion.

cs.CV

Symmetry of Convex Solutions to Fully Nonlinear Elliptic Systems: Unbounded Domains

In this paper, we are concerned with the monotonic and symmetric properties of convex solutions Monge-Ampère systems for instance, considering \begin{equation*} \det(D^2u^i)=f^i(x,{\bf u},\nabla u^i), \ 1\leq i\leq m, \end{equation*} over unbounded domains of various cases, including the whole spaces $\mathbb{R}^n$, the half spaces $\mathbb{R}^n_+$ and the unbounded tube shape domains in $\mathbb{R}^n$. We obtain monotonic and symmetric properties of the solutions to the problem with respect to the geometry of domains and the monotonic and symmetric properties of right-hand side terms. The proof is based on carefully using the moving plane method together with various maximum principles and Hopf's lemmas.

math.AP

Bifurcation on Fully Nonlinear Elliptic Equations and Systems

In this paper, we study the following fully nonlinear elliptic equations \begin{equation*} \left\{\begin{array}{rl} \left(S_{k}(D^{2}u)\right)^{\frac1k}=λf(-u) & in\quadΩ\\ u=0 & on\quad \partialΩ\\ \end{array} \right. \end{equation*} and coupled systems \begin{equation*} \left\{\begin{array}{rl} (S_{k}(D^{2}u))^\frac1k=λg(-u,-v) & in\quadΩ\\ (S_{k}(D^{2}v))^\frac1k=λh(-u,-v) & in\quadΩ\\ u=v=0 & on\quad \partialΩ\\ \end{array} \right. \end{equation*} dominated by $k$-Hessian operators, where $Ω$ is a $(k$-$1)$-convex bounded domain in $\mathbb{R}^{N}$, $λ$ is a non-negative parameter, $f:\left[0,+\infty\right)\rightarrow\left[0,+\infty\right)$ is a continuous function with zeros only at $0$ and $g,h:\left[0,+\infty\right)\times \left[0,+\infty\right)\rightarrow \left[0,+\infty\right)$ are continuous functions with zeros only at $(\cdot,0)$ and $(0,\cdot)$. We determine the interval of $λ$ about the existence, non-existence, uniqueness and multiplicity of $k$-convex solutions to the above problems according to various cases of $f,g,h$, which is a complete supplement to the known results in previous literature. In particular, the above results are also new for Laplacian and Monge-Ampère operators. We mainly use bifurcation theory, a-priori estimates, various maximum principles and technical strategies in the proof.

math.AP

Symmetry of Convex Solutions to Fully Nonlinear Elliptic Systems: Bounded Domains

In this paper, we are concerned with the monotonic and symmetric properties of convex solutions to fully nonlinear elliptic systems. We mainly discuss Monge-Ampère type systems for instance, considering \begin{equation*} \det(D^2u^i)=f^i(x,{\bf u},\nabla u^i), \ 1\leq i\leq m, \end{equation*} over bounded domains of various cases, including the bounded smooth simply connected domains and bounded tube shape domains in $\mathbb{R}^n$. We obtain monotonic and symmetric properties of the solutions to the problem with respect to the geometry of domains and the monotonic and symmetric properties of right-hand side terms. The proof is based on carefully using the moving plane method together with various maximum principles and Hopf's lemmas. The existence and uniqueness to an interesting example of such system is also discussed as an application of our results.

math.AP

Improving photon number resolvability of a superconducting nanowire detector array using a level comparator circuit

Photon number resolving (PNR) capability is very important in many optical applications, including quantum information processing, fluorescence detection, and few-photon-level ranging and imaging. Superconducting nanowire single-photon detectors (SNSPDs) with a multipixel interleaved architecture give the array an excellent spatial PNR capability. However, the signal-to-noise ratio (SNR) of the photon number resolution (SNRPNR) of the array will be degraded with increasing the element number due to the electronic noise in the readout circuit, which limits the PNR resolution as well as the maximum PNR number. In this study, a 16-element interleaved SNSPD array was fabricated, and the PNR capability of the array was investigated and analyzed. By introducing a level comparator circuit (LCC), the SNRPNR of the detector array was improved over a factor of four. In addition, we performed a statistical analysis of the photon number on this SNSPD array with LCC, showing that the LCC method effectively enhances the PNR resolution. Besides, the system timing jitter of the detector was reduced from 90 ps to 72 ps due to the improved electrical SNR.

physics.app-ph

Multi-query Vehicle Re-identification: Viewpoint-conditioned Network, Unified Dataset and New Metric

Existing vehicle re-identification methods mainly rely on the single query, which has limited information for vehicle representation and thus significantly hinders the performance of vehicle Re-ID in complicated surveillance networks. In this paper, we propose a more realistic and easily accessible task, called multi-query vehicle Re-ID, which leverages multiple queries to overcome viewpoint limitation of single one. Based on this task, we make three major contributions. First, we design a novel viewpoint-conditioned network (VCNet), which adaptively combines the complementary information from different vehicle viewpoints, for multi-query vehicle Re-ID. Moreover, to deal with the problem of missing vehicle viewpoints, we propose a cross-view feature recovery module which recovers the features of the missing viewpoints by learnt the correlation between the features of available and missing viewpoints. Second, we create a unified benchmark dataset, taken by 6142 cameras from a real-life transportation surveillance system, with comprehensive viewpoints and large number of crossed scenes of each vehicle for multi-query vehicle Re-ID evaluation. Finally, we design a new evaluation metric, called mean cross-scene precision (mCSP), which measures the ability of cross-scene recognition by suppressing the positive samples with similar viewpoints from same camera. Comprehensive experiments validate the superiority of the proposed method against other methods, as well as the effectiveness of the designed metric in the evaluation of multi-query vehicle Re-ID.

cs.CV

Experimental Side-Channel-Free Quantum Key Distribution

Quantum key distribution can provide unconditionally secure key exchange for remote users in theory. In practice, however, in most quantum key distribution systems, quantum hackers might steal the secure keys by listening to the side channels in the source, such as the photon frequency spectrum, emission time, propagation direction, spatial angular momentum, and so on. It is hard to prevent such kinds of attacks because side channels may exist in any of the encoding space whether the designers take care of or not. Here we report an experimental realization of a side-channel-free quantum key distribution protocol which is not only measurement-device-independent, but also immune to all side-channel attacks in the source. We achieve a secure key rate of 4.80e-7 per pulse through 50 km fiber spools.

quant-ph

Field Test of Twin-Field Quantum Key Distribution through Sending-or-Not-Sending over 428 km

Quantum key distribution endows people with information-theoretical security in communications. Twin-field quantum key distribution (TF-QKD) has attracted considerable attention because of its outstanding key rates over long distances. Recently, several demonstrations of TF-QKD have been realized. Nevertheless, those experiments are implemented in the laboratory, remaining a critical question about whether the TF-QKD is feasible in real-world circumstances. Here, by adopting the sending-or-not-sending twin-field QKD (SNS-TF-QKD) with the method of actively odd parity pairing (AOPP), we demonstrate a field-test QKD over 428~km deployed commercial fiber and two users are physically separated by about 300~km in a straight line. To this end, we explicitly measure the relevant properties of the deployed fiber and develop a carefully designed system with high stability. The secure key rate we achieved breaks the absolute key rate limit of repeater-less QKD. The result provides a new distance record for the field test of both TF-QKD and all types of fiber-based QKD systems. Our work bridges the gap of QKD between laboratory demonstrations and practical applications, and paves the way for intercity QKD network with high-speed and measurement-device-independent security.

quant-ph

Field demonstration of distributed quantum sensing without post-selection

Distributed quantum sensing can provide quantum-enhanced sensitivity beyond the shot-noise limit (SNL) for sensing spatially distributed parameters. To date, distributed quantum sensing experiments have been mostly accomplished in laboratory environments without a real space separation for the sensors. In addition, the post-selection is normally assumed to demonstrate the sensitivity advantage over the SNL. Here, we demonstrate distributed quantum sensing in field and show the unconditional violation (without post-selection) of SNL up to 0.916 dB for the field distance of 240 m. The achievement is based on a loophole free Bell test setup with entangled photon pairs at the averaged heralding efficiency of 73.88%. Moreover, to test quantum sensing in real life, we demonstrate the experiment for long distances (with 10-km fiber) together with the sensing of a completely random and unknown parameter. The results represent an important step towards a practical quantum sensing network for widespread applications.

quant-ph