SearcharxivSearch

arXiv subjects

Nannan Li

Publications and source records attributed to Nannan Li.

At least 19 recordsLinked to original sources

CoDS: Robust Collaborative Perception via Expert-driven Detection and BEV Segmentation

Collaborative perception breaks through single-view limitations via multi-agent information exchange. However, multi-source noise such as pose errors and communication delays degrades fusion feature quality, constraining perception performance. Joint training of detection and BEV segmentation provides a natural remedy, where segmented road regions help constrain target distributions and detection bounding boxes help recover ambiguous segmentation boundaries. To this end, we propose a robust Collaborative perception framework with expert-driven Detection and bev Segmentation (CoDS). To address spatial inconsistency in fusion quality, we first introduce the Collaborative Reliability Map (CoRM) to explicitly quantify feature quality distribution. Based on CoRM, we design the Semantic Mixture-of-Experts (S-MoE) module to extract differentiated features for inconsistent feature demands. Finally, to further mitigate feature noise degradation, the Bidirectional Task Complementary Interaction (BTCI) refines task-aware features through bidirectional injection. Extensive experiments on OPV2V and V2V4Real datasets show that our CoDS surpasses existing baselines on both tasks and maintains stable robustness under multi-source noise. Code: https://github.com/JinlongW128/CoDS and https://openi.pcl.ac.cn/OpenAIDriving/CoDS.

cs.CV

Exact Signature Tail Asymptotics for Pure Rough Paths

We prove~\cite[Conjecture 2.12]{BGS20} on the signature tail asymptotics of pure rough paths and extend it to arbitrary reasonable tensor norms. In more details, let \[ \mathbf X_t=\exp(tl) \,\text{ with }\, l=l_1+\cdots+l_m\,\text{ and }\, l_r\in\mathcal L_r(V), \] be a pure $m$-rough path over a finite dimensional real or complex Banach space, and equip the tensor powers of $V$ with arbitrary reasonable tensor algebra norms. We prove that \[ \limsup_{n\to\infty}\left(\left(\frac{n}{m}\right)!\left\|\pi_n(\exp l)\right\|_n\right)^{m/n}=\|l_m\|_m . \] In particular, this identifies the signature tail with the local $m$-variation of the pure rough path. The upper bound was obtained in~\cite{BGS20}; the main contribution of the paper is the matching lower bound. Its proof is based on finite dimensional developments and a norming cyclic construction. For every top-level tensor $l_m$, we also build a contractive development in which $\|l_m\|_m$ appears as an eigenvalue at degree $m$.

math.PR

A Low-Regularity Semigroup Sewing Lemma via Quotient Structures

We develop a low-regularity Sewing theory for the semigroup coboundary $\hat\delta=\delta-a$ associated with a strongly continuous semigroup $S$. Unlike the ordinary low-regularity Sewing problem, the semigroup setting has an intrinsic algebraic non-uniqueness below the threshold $1$, in the sense that solutions are canonical only modulo semigroup cocycles. Accordingly, the natural target is a quotient space rather than an increment space. We identify this quotient structure and construct the corresponding semigroup Sewing map. The construction uses a frozen terminal-time transform, which rewrites semigroup defects, for each terminal time, as ordinary low-regularity Sewing problems on a frozen simplex. This reduction, however, does not by itself produce a genuine semigroup increment; the main additional step is to prove that the frozen solution classes are compatible as the terminal time varies and hence assemble into a canonical quotient class for $\hat\delta$. This yields canonical classes for $0<\gamma<1$, and at $\gamma=1$ under logarithmic control. We further provide a scale-dependent criterion for selecting genuine representatives, verified for heat semigroups on Sobolev scales through a parabolic Littlewood--Paley tail condition.

math.PR

Jump It\^o-type formula with arbitrary regularity

We establish an It\^o-type formula for finite $p$-variation paths with jumps for arbitrary $p\geq 1$. The formula is stated in a fully pathwise form and separates the reduced rough integral from explicit left- and right-jump correction terms. In the c\`adl\`ag case, only the left-jump correction remains, while in the continuous case, both jump correction terms vanish and the formula recovers the corresponding continuous arbitrary-regularity change-of-variable formula. The proof is based on the reduced rough path framework and a refinement Riemann-Stieltjes convergence criterion adapted to discontinuous paths. This approach allows us to handle the higher-order Taylor expansions required for large values of $p$ and to control the interaction between rough increments and discrete jumps. As applications, we derive It\^o-type formulas for stochastic processes whose sample paths have finite $p$-variation, including pure-jump models and mixed fractional Brownian-jump signals. The latter class includes cases with Hurst parameter $H\leq 1/3$, which fall outside the regime $2\leq p<3$. We also obtain chain-rule identities for nonlinear observables of c\`adl\`ag finite-$p$-variation solutions of random differential equations with jumps, together with a pathwise log-wealth decomposition.

math.PR

USCNet: Transformer-Based Multimodal Fusion with Segmentation Guidance for Urolithiasis Classification

Kidney stone disease ranks among the most prevalent conditions in urology, and understanding the composition of these stones is essential for creating personalized treatment plans and preventing recurrence. Current methods for analyzing kidney stones depend on postoperative specimens, which prevents rapid classification before surgery. To overcome this limitation, we introduce a new approach called the Urinary Stone Segmentation and Classification Network (USCNet). This innovative method allows for precise preoperative classification of kidney stones by integrating Computed Tomography (CT) images with clinical data from Electronic Health Records (EHR). USCNet employs a Transformer-based multimodal fusion framework with CT-EHR attention and segmentation-guided attention modules for accurate classification. Moreover, a dynamic loss function is introduced to effectively balance the dual objectives of segmentation and classification. Experiments on an in-house kidney stone dataset show that USCNet demonstrates outstanding performance across all evaluation metrics, with its classification efficacy significantly surpassing existing mainstream methods. This study presents a promising solution for the precise preoperative classification of kidney stones, offering substantial clinical benefits. The source code has been made publicly available: https://github.com/ZhangSongqi0506/KidneyStone.

cs.CV

JEPA-MSAC: A Joint-Embedding Predictive Architecture for Multimodal Sensing-Assisted Communications

Future wireless systems increasingly require predictive and transferable representations that can support multiple physical-layer (PHY) tasks under dynamic environments. However, most existing supervised learning-based methods are designed for a single task, which leads to high adaptation cost. To address this issue, we propose a joint-embedding predictive architecture for multimodal sensing-assisted communications (JEPA-MSAC), a self-supervised multimodal predictive representation learning framework for wireless environments. The proposed framework first maps multimodal sensing and communication measurements into a unified token space, and then pretrains a shared backbone using temporal block-masked JEPA to learn a predictive latent space that captures environment dynamics and cross-modal dependencies. After pretraining, the backbone is frozen and reused as a general future-feature generator, on top of which lightweight task heads are trained for localization, beam prediction, and received signal strength indicator (RSSI) prediction. Extensive experiments show the latent state supports accurate multi-task prediction with low adaptation cost. Additionally, ablation studies reveal its scaling behavior and the impact of key pretraining setups.

eess.SP

Universal limit theorem for rough differential equations driven by controlled rough paths

We study rough differential equations driven by controlled rough paths in the level-$2$ regime $1/3<\alpha\le 1/2$. Given a reference rough path $\mathbf X=(1,X,\mathbb X)$ and an $\mathbf X$-controlled driver $\mathbf Z=(Z,Z')$, we first give a point-removal construction of the controlled rough integral $ \int_s^t Y_r\,d\mathbf Z_r $ and prove the corresponding remainder estimates. We then establish local and global well-posedness for the controlled-driven rough differential equation $ dY_t=F(Y_t)\,d\mathbf Z_t. $ A key structural result is the canonical lift of the controlled driver: from the controlled data $(\mathbf X,\mathbf Z)$ we construct a level-$2$ rough path \[ \widehat{\mathbf Z}=(1,Z,\mathbb Z), \qquad \mathbb Z_{s,t}:=\int_s^t Z_{s,u}\otimes dZ_u, \] and show that the controlled-driven equation is equivalent to the classical rough differential equation driven by $\widehat{\mathbf Z}$. This equivalence shows compatibility with classical rough path theory, while the controlled formulation keeps track of the dependence of the effective driver $Z$ on the reference rough path $\X$. Finally, we prove a universal limit theorem for the solution map $ (\mathbf X,\mathbf Z,Y_0)\longmapsto Y, $ which gives stability with respect to perturbations of the initial condition, the reference rough path, and the controlled driver. These results provide a natural framework for layered rough systems and equations driven by transformed or previously evolved rough signals.

math.PR

Emergence of Pascal's triangle in cascaded polarization optics: an intuitive framework for field transformation

Nature is imbued with mathematics, manifested through its stunning patterns, symmetries, and structures. Here, we unveil that in a multilayered framework of twisted birefringent optical components, a recursive number pattern of Pascal's triangle is naturally embedded in the structure of the Jones matrix which intuitively provide a generalized solution for pixel-to-pixel field transformation. The resulting standalone solution is universal across the electromagnetic spectrum, unifies N-layered metasurface and conventional bulk waveplates in a single framework, offers comprehensive insights about the bidirectional complex amplitude modulation and wavefront engineering in linear and circular polarization bases, and at the same time substantially reduces the computational cost. In essence, the discovery of number patterns in polarization optics/photonics will have broad impact across quantum optics, theory informed artificial intelligence model trainings, biomedical engineering and imaging, polarization information encryption, and advanced sensing applications.

physics.optics

Rough differential equations and reduced rough paths: a Lie bracket characterization

This paper studies rough differential equations from the viewpoint of reduced rough paths in the H\"older regime \(\frac13<\alpha\le\frac12\). A reduced rough path retains the first level and the symmetric part of the second level, while discarding the antisymmetric L\'evy-area component. We identify the precise obstruction to determining rough differential equation solutions from this reduced information. For an RDE driven by vector fields \(F_1,\ldots,F_d\), we prove that any two rough paths with the same reduced projection produce the same solution for every common initial value if and only if \[ [F_i,F_j]=0, \quad 1\le i,j\le d. \] Thus the antisymmetric L\'evy area is irrelevant exactly in the commuting-vector-field case.

math.PR

Controlled rough paths: a general Hopf-algebraic setting

We set up controlled rough paths for a class of combinatorial Hopf algebras, encompassing shuffle, Butcher-Connes-Kreimer and Munthe-Kaas--Wright Hopf algebras. The class of controls we consider encompasses both H\"older continuous paths and (not necessarily continuous) paths with bounded $p$-variation. We prove existence and uniqueness of the solution of a lifted initial value problem in this general setting by applying the fixed point method in a suitable Banach space of controlled rough paths, and we prove a universal limit theorem addressing the robustness of the solution with respect to the parameters and the initial condition.

math.PR

It\^o formula for reduced rough paths

The It\^o formula, also known as the change-of-variables formula, is a cornerstone of It\^o stochastic calculus. Over time, this formula has been extended to apply to random processes for which classical calculus is insufficient. Since every random process exhibits some degree of regularity, rough path theory provides a natural framework for treating them uniformly. In this paper, we extend the It\^o formula for reduced rough paths, broadening the range of roughness from the previously known case $\frac{1}{3} < \alpha \leq \frac{1}{2}$ to the more singular regime $\frac{1}{4} < \alpha \leq \frac{1}{3}$.

math.PR

High-fidelity 3D Gaussian Inpainting: preserving multi-view consistency and photorealistic details

Recent advancements in multi-view 3D reconstruction and novel-view synthesis, particularly through Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS), have greatly enhanced the fidelity and efficiency of 3D content creation. However, inpainting 3D scenes remains a challenging task due to the inherent irregularity of 3D structures and the critical need for maintaining multi-view consistency. In this work, we propose a novel 3D Gaussian inpainting framework that reconstructs complete 3D scenes by leveraging sparse inpainted views. Our framework incorporates an automatic Mask Refinement Process and region-wise Uncertainty-guided Optimization. Specifically, we refine the inpainting mask using a series of operations, including Gaussian scene filtering and back-projection, enabling more accurate localization of occluded regions and realistic boundary restoration. Furthermore, our Uncertainty-guided Fine-grained Optimization strategy, which estimates the importance of each region across multi-view images during training, alleviates multi-view inconsistencies and enhances the fidelity of fine details in the inpainted results. Comprehensive experiments conducted on diverse datasets demonstrate that our approach outperforms existing state-of-the-art methods in both visual quality and view consistency.

cs.CV

3D-Telepathy: Reconstructing 3D Objects from EEG Signals

Reconstructing 3D visual stimuli from Electroencephalography (EEG) data holds significant potential for applications in Brain-Computer Interfaces (BCIs) and aiding individuals with communication disorders. Traditionally, efforts have focused on converting brain activity into 2D images, neglecting the translation of EEG data into 3D objects. This limitation is noteworthy, as the human brain inherently processes three-dimensional spatial information regardless of whether observing 2D images or the real world. The neural activities captured by EEG contain rich spatial information that is inevitably lost when reconstructing only 2D images, thus limiting its practical applications in BCI. The transition from EEG data to 3D object reconstruction faces considerable obstacles. These include the presence of extensive noise within EEG signals and a scarcity of datasets that include both EEG and 3D information, which complicates the extraction process of 3D visual data. Addressing this challenging task, we propose an innovative EEG encoder architecture that integrates a dual self-attention mechanism. We use a hybrid training strategy to train the EEG Encoder, which includes cross-attention, contrastive learning, and self-supervised learning techniques. Additionally, by employing stable diffusion as a prior distribution and utilizing Variational Score Distillation to train a neural radiation field, we successfully generate 3D objects with similar content and structure from EEG data.

cs.CV

Multiple Object Tracking in Video SAR: A Benchmark and Tracking Baseline

In the context of multi-object tracking using video synthetic aperture radar (Video SAR), Doppler shifts induced by target motion result in artifacts that are easily mistaken for shadows caused by static occlusions. Moreover, appearance changes of the target caused by Doppler mismatch may lead to association failures and disrupt trajectory continuity. A major limitation in this field is the lack of public benchmark datasets for standardized algorithm evaluation. To address the above challenges, we collected and annotated 45 video SAR sequences containing moving targets, and named the Video SAR MOT Benchmark (VSMB). Specifically, to mitigate the effects of trailing and defocusing in moving targets, we introduce a line feature enhancement mechanism that emphasizes the positive role of motion shadows and reduces false alarms induced by static occlusions. In addition, to mitigate the adverse effects of target appearance variations, we propose a motion-aware clue discarding mechanism that substantially improves tracking robustness in Video SAR. The proposed model achieves state-of-the-art performance on the VSMB, and the dataset and model are released at https://github.com/softwarePupil/VSMB.

cs.CV

Rough Burger-like SPDEs

We study a class of nonlinear Burgers-type stochastic partial differential equations driven by additive space-time white noise in one spatial dimension. Building on the rough path framework initiated by Hairer, which provides a pathwise solution theory under spatial regularity $\alpha \in(\frac{1}{3}, \frac{1}{2})$, we extend this approach to the full subcritical regime $\alpha \in(0, \frac{1}{2})$. Our main contribution is the establishment of pathwise existence and uniqueness of mild (equivalently, weak) solutions when the spatial regularity of the solution lies strictly below the classical rough path threshold. This is achieved through refined estimates for controlled rough paths, including a new upper bound for compositions with smooth functions and a scaling analysis for rough integrals against heat kernels. In particular, we extend and sharpen key analytic estimates originating from Hairer's work, incorporating refined scaling arguments that are effective in the low-regularity regime. As a result, our framework significantly enlarges the class of Burgers-type SPDEs that can be treated pathwise using rough path techniques.

math.PR

Towards Better Robustness: Pose-Free 3D Gaussian Splatting for Arbitrarily Long Videos

3D Gaussian Splatting (3DGS) has emerged as a powerful representation due to its efficiency and high-fidelity rendering. 3DGS training requires a known camera pose for each input view, typically obtained by Structure-from-Motion (SfM) pipelines. Pioneering works have attempted to relax this restriction but still face difficulties when handling long sequences with complex camera trajectories. In this paper, we propose Rob-GS, a robust framework to progressively estimate camera poses and optimize 3DGS for arbitrarily long video inputs. In particular, by leveraging the inherent continuity of videos, we design an adjacent pose tracking method to ensure stable pose estimation between consecutive frames. To handle arbitrarily long inputs, we propose a Gaussian visibility retention check strategy to adaptively split the video sequence into several segments and optimize them separately. Extensive experiments on Tanks and Temples, ScanNet, and a self-captured dataset show that Rob-GS outperforms the state-of-the-arts.

cs.CV

It\^o formula for planarly branched rough paths

The It\^o formula, originated by K. It\^o, is focus on the stochastic calculus, where many stochastic processes can be placed under the framework of rough paths. In rough path theory, It\^o formulas have been proved for rough paths with roughness $\frac{1}{3}< \alpha \leq \frac{1}{2}$ and branched rough paths with roughness $0< \alpha \leq 1$. Planarly branched rough paths contain more random processes than rough paths and branched rough paths. In the present paper, we prove the It\^o formula for planarly branched rough paths with roughness $\frac{1}{4}< \alpha \leq \frac{1}{2}$.

math.PR

Enhancing Virtual Try-On with Synthetic Pairs and Error-Aware Noise Scheduling

Given an isolated garment image in a canonical product view and a separate image of a person, the virtual try-on task aims to generate a new image of the person wearing the target garment. Prior virtual try-on works face two major challenges in achieving this goal: a) the paired (human, garment) training data has limited availability; b) generating textures on the human that perfectly match that of the prompted garment is difficult, often resulting in distorted text and faded textures. Our work explores ways to tackle these issues through both synthetic data as well as model refinement. We introduce a garment extraction model that generates (human, synthetic garment) pairs from a single image of a clothed individual. The synthetic pairs can then be used to augment the training of virtual try-on. We also propose an Error-Aware Refinement-based Schr\"odinger Bridge (EARSB) that surgically targets localized generation errors for correcting the output of a base virtual try-on model. To identify likely errors, we propose a weakly-supervised error classifier that localizes regions for refinement, subsequently augmenting the Schr\"odinger Bridge's noise schedule with its confidence heatmap. Experiments on VITON-HD and DressCode-Upper demonstrate that our synthetic data augmentation enhances the performance of prior work, while EARSB improves the overall image quality. In user studies, our model is preferred by the users in an average of 59% of cases.

cs.CV