SearcharxivSearch

arXiv subjects

Jing Ou

Publications and source records attributed to Jing Ou.

5 recordsLinked to original sources

UniD-Shift: Towards Unified Semantic Segmentation via Interpretable Share-Private Multimodal Decomposition

Semantic segmentation of large-scale 3D point clouds is crucial for applications such as autonomous driving and urban digital twins. However, the sparse sampling pattern of LiDAR and the view-dependent geometric distortion in image observations complicate cross-modal alignment and hinder stable fusion. Inspired by the fact that 2D images captured by cameras are representations of the 3D world, we recognize that the features learned from 2D and 3D segmentation share some common semantics, while other aspects remain modality-specific. This insight motivates a unified multimodal framework for joint 2D-3D semantic segmentation. We combine a SAM-based vision encoder with a SPTNet-based geometric encoder to extract complementary semantic and geometric representations. The resulting features from both modalities are explicitly decomposed into shared and private subspaces, where the shared components summarize semantic factors common to both domains, and the private components preserve properties that are unique to each modality. A lightweight attention-based fusion module aggregates the shared features into a consistent cross-modal representation, and a regularized training objective ensures both semantic alignment and subspace independence. Experiments on the SemanticKITTI and nuScenes benchmarks demonstrate consistent improvements in segmentation accuracy over representative multimodal baselines, accompanied by competitive computational efficiency. Cross-domain evaluation on nuScenes USA-Singapore shows stable performance under distribution shifts, demonstrating strong generalization. The implementation code is publicly available at: https://github.com/shuaizhang69/UniD-Shift.

cs.CV

Holo360D: A Large-Scale Real-World Dataset with Continuous Trajectories for Advancing Panoramic 3D Reconstruction and Beyond

While feed-forward 3D reconstruction models have advanced rapidly, they still exhibit degraded performance on panoramas due to spherical distortions. Moreover, existing panoramic 3D datasets are predominantly collected with 360 cameras fixed at discrete locations, resulting in discontinuous trajectories. These limitations critically hinder the development of panoramic feed-forward 3D reconstruction, especially for the multi-view setting. In this paper, we present Holo360D, a comprehensive dataset containing 109,495 panoramas paired with registered point clouds, meshes, and aligned camera poses. To our knowledge, Holo360D is the first large-scale dataset that provides continuous panoramic sequences with accurately aligned high-completeness depth maps. The raw data are initially collected using a 3D laser scanner coupled with a 360 camera. Subsequently, the raw data are processed with both online and offline SLAM systems. Furthermore, to enhance the 3D data quality, a post-processing pipeline tailored for the 360 dataset is proposed, including geometry denoising, mesh hole filling, and region-specific remeshing. Finally, we establish a new benchmark by fine-tuning 3D reconstruction models on Holo360D, providing key insights into effective fine-tuning strategies. Our results demonstrate that Holo360D delivers superior training signals and provides a comprehensive benchmark for advancing panoramic 3D reconstruction models. Datasets and Code will be made publicly available.

cs.CV

EventVGGT: Exploring Cross-Modal Distillation for Consistent Event-based Depth Estimation

Event cameras offer superior sensitivity to high-speed motion and extreme lighting, making event-based monocular depth estimation a promising approach for robust 3D perception in challenging conditions. However, progress is severely hindered by the scarcity of dense depth annotations. While recent annotation-free approaches mitigate this by distilling knowledge from Vision Foundation Models (VFMs), a critical limitation persists: they process event streams as independent frames. By neglecting the inherent temporal continuity of event data, these methods fail to leverage the rich temporal priors encoded in VFMs, ultimately yielding temporally inconsistent and less accurate depth predictions. To address this, we introduce EventVGGT, a novel framework that explicitly models the event stream as a coherent video sequence. To the best of our knowledge, we are the first to distill spatio-temporal and multi-view geometric priors from the Visual Geometry Grounded Transformer (VGGT) into the event domain. We achieve this via a comprehensive tri-level distillation strategy: (i) Cross-Modal Feature Mixture (CMFM) bridges the modality gap at the output level by fusing RGB and event features to generate auxiliary depth predictions; (ii) Spatio-Temporal Feature Distillation (STFD) distills VGGT's powerful spatio-temporal representations at the feature level; and (iii) Temporal Consistency Distillation (TCD) enforces cross-frame coherence at the temporal level by aligning inter-frame depth changes. Extensive experiments demonstrate that EventVGGT consistently outperforms existing methods -- reducing the absolute mean depth error at 30m by over 53\% on EventScape (from 2.30 to 1.06) -- while exhibiting robust zero-shot generalization on the unseen DENSE and MVSEC datasets. The code is available at https://github.com/yinruiRen/EventVGGT.

cs.CV

Comparison of 2D simulation models to estimate the critical current of a coated superconducting coil

Superconductors have been being applied to a variety of large-scale power applications, including magnets, electric machines, and fault current limiters, because they can enable a compact, lightweight and high efficiency design. In applications such those mentioned above, superconducting coils are always a key component. For example, in a superconducting electric machine, the superconducting coils are used to generate the main flux density in the air gap, which is significantly important for the energy conversion. It is the performance of the superconducting coils that plays an essential role in determining the performance of the device. However, the performance of a superconducting coil is limited by its critical current, which is determined by temperature and the magnitude and orientation of the magnetic field inside the superconductors. Hence, in-depth investigations to estimate the critical current of the superconducting coils are necessary before manufacturing. Available transient simulation models to estimate the critical current are through the H- and T-A formulations of Maxwell's equations. Both methods consider the same current ramp-up process occurring in experiments. Besides these transient models, static simulations can also be used: a modified load-line method and the so-called P-model, which is based on the asymptotic limit of Faraday's equation when time approaches infinity. To find the best way to calculate the critical current, the four methods are used to estimate the critical current of a double pancake superconducting coils and results are compared with experiments. As a conclusion, T-A formulation, P-model, and the modified load-line methods are recommended for estimating the critical current of the superconducting coils.

cond-mat.supr-con

Co-current rotation of the bulk ions due to the ion orbit loss at the edge of a tokamak plasma

Flux-surface-averaged momentum loss and parallel rotation of the bulk ions at the edge of a tokamak plasma due to the ion orbit loss are calculated by computing the minimum loss energy of both the trapped and the passing thermal ions. The flux-surface-averaged parallel rotation of the bulk ions is in the co-current direction. The peak of the co-current rotation speed locates inside the last closed flux surface due to the orbit loss of the co-current thermal ions at the very edge of a tokamak plasma. The peaking position moves inward when the ion temperature increases.

physics.plasm-ph