SearcharxivSearch

arXiv subjects

Guoqiang Zhao

Publications and source records attributed to Guoqiang Zhao.

At least 19 recordsLinked to original sources

Spheriverse: 3D Scene Understanding from Spherical Observations in the Wild

Spherical observations provide global visual context for 3D scene understanding. However, visual information is encoded in an angular domain, whereas the physical world is represented in Cartesian coordinates. This cross-space representation gap complicates geometric correspondence and semantic evidence aggregation. To delve into this challenge, we introduce Spheriverse, comprising $64,400$ temporally aligned spherical image-LiDAR pairs organized into 644 sequences. The dataset spans diverse scenes, illumination, and weather conditions, with fine-grained semantic classes. We further establish benchmarks for semantic occupancy prediction, semantic mapping, and 3D object detection, evaluating 30+ methods through overall and scene-wise comparisons. For dense prediction, we propose SphereOcc, an occupancy framework that couples spherical geometry modeling with semantic evidence retrieval. Cartesian-Spherical Representation Remodeling (CSRR) incorporates spherical range-azimuth geometry into Cartesian voxel features through region-wise modulation. Spherical Evidence Re-querying (SER) then conditions queries on voxel content and range-height-azimuth geometry to adaptively retrieve relevant semantic evidence from source spherical image features. SphereOcc achieves 13.91% mIoU and 24.65% GeoIoU, outperforming the respective best-performing methods, TPVFormer and SurroundOcc, by 1.70 and 2.10 percentage points. It also ranks first in both metrics across all five scenes, with consistent advantages across the evaluated spatial partitions and reduced fields of view. The established benchmark and source code will be available at https://feit-feiteng.github.io/Spheriverse.

cs.CV

Tilting pairs and Wakamatsu tilting pairs of subcategories over cleft extensions

Let $(\mathcal{B},\mathcal{A}, i, e, l)$ be a cleft extension of abelian categories. We prove that the functor $l$ preserves and reflects (Wakamatsu) tilting pairs of subcategories under certain conditions, unifying an abundance of known results. Then, we apply our results to the cleft extensions of module categories, and give characterizations of tilting pairs and Wakamatsu tilting pairs over $\theta$-extension of rings and tensor rings, which not only recover the earlier results in this direction, but also obtain some new conclusions.

math.RT

Panoramic Multimodal Semantic Occupancy Prediction for Quadruped Robots

Panoramic imagery provides holistic 360{\deg} visual coverage for environmental perception in quadruped robots. However, existing occupancy prediction methods are primarily designed for wheeled autonomous driving and rely heavily on RGB cues, which limits their robustness in complex, dynamically changing environments. To bridge this gap, we introduce PanoMMOcc, the first real-world panoramic multimodal occupancy dataset for quadruped robots, comprising four sensing modalities collected across diverse scenes. We further propose VoxelHound, a panoramic multimodal occupancy perception framework tailored to legged locomotion and spherical imaging. VoxelHound incorporates a Vertical Jitter Compensation (VJC) module to mitigate severe viewpoint perturbations caused by body pitch and roll during locomotion, enabling more consistent spatial reasoning, and a Multimodal Information Prompt Fusion (MIPF) module to effectively integrate panoramic visual cues with auxiliary modalities for enhanced volumetric occupancy prediction. We also establish a comprehensive benchmark on PanoMMOcc and provide detailed dataset analyses to enable systematic evaluation in challenging embodied perception scenarios. Extensive experiments demonstrate that VoxelHound achieves state-of-the-art performance on PanoMMOcc, with a +4.16 gain in mIoU. The dataset and code will be publicly released to facilitate future research on panoramic multimodal 3D perception for embodied robotic systems at https://github.com/SXDR/PanoMMOcc.

cs.RO

O3N: Omnidirectional Open-Vocabulary Occupancy Prediction for Embodied Intelligent Robotics

The rapid evolution of consumer electronics toward embodied intelligence has accelerated the emergence of Consumer Embodied Intelligent Robotics (CEIRs), where intelligent devices are expected to perceive, understand, and interact with complex real-world environments. Understanding and reconstructing the 3D world through omnidirectional perception is therefore becoming increasingly important for CEIRs operating in complex and dynamic environments. However, existing vision-based 3D occupancy prediction methods are constrained by limited perspective inputs and a predefined training distribution, making them difficult to support embodied intelligent systems that require comprehensive and safe perception of scenes in open-world exploration. To address this, we present O3N, the first framework for open-vocabulary occupancy prediction from a single omnidirectional RGB image. O3N embeds omnidirectional voxels in a polar-spiral topology via the Polar-spiral Mamba (PsM) module, enabling continuous spatial representation and long-range context modeling across 360{\deg}. The Occupancy Cost Aggregation (OCA) module introduces a principled mechanism for unifying geometric and semantic supervision within the voxel space, ensuring consistency between reconstructed geometry and underlying semantic structure. Moreover, Natural Modality Alignment (NMA) establishes a gradient-free alignment pathway that harmonizes visual features, voxel embeddings, and text semantics, forming a consistent pixel-voxel-text representation triad for open-world perception. Extensive experiments on multiple models demonstrate that our method not only achieves state-of-the-art performance on QuadOcc and Human360Occ benchmarks but also exhibits remarkable cross-scene generalization and semantic scalability. The source code will be made publicly available at https://github.com/MengfeiD/O3N.

cs.CV

Spherical-GOF: Geometry-Aware Panoramic Gaussian Opacity Fields for 3D Scene Reconstruction

Omnidirectional images are increasingly used in robotics and vision due to their wide field of view. However, extending 3D Gaussian Splatting (3DGS) to panoramic camera models remains challenging, as existing formulations are designed for perspective projections and naive adaptations often introduce distortion and geometric inconsistencies. We present Spherical-GOF, an omnidirectional Gaussian rendering framework built upon Gaussian Opacity Fields (GOF). Unlike projection-based rasterization, Spherical-GOF performs GOF ray sampling directly on the unit sphere in spherical ray space, enabling consistent ray-Gaussian interactions for panoramic rendering. To make the spherical ray casting efficient and robust, we derive a conservative spherical bounding rule for fast ray-Gaussian culling and introduce a spherical filtering scheme that adapts Gaussian footprints to distortion-varying panoramic pixel sampling. Extensive experiments on standard panoramic benchmarks (OmniBlender and OmniPhotos) demonstrate competitive photometric quality and substantially improved geometric consistency. Compared with the strongest baseline, Spherical-GOF reduces depth reprojection error by 57% and improves cycle inlier ratio by 21%. Qualitative results show cleaner depth and more coherent normal maps, with strong robustness to global panorama rotations. We further validate generalization on OmniRob, a real-world robotic omnidirectional dataset introduced in this work, featuring UAV and quadruped platforms. The source code and the OmniRob dataset will be released at https://github.com/1170632760/Spherical-GOF.

cs.CV

Silting and tilting objects in cleft extensions of abelian categories

We establish connections between silting and tilting objects in an abelian category $\mathcal{B}$ and those in a cleft extension $\mathcal{A}$ of $\mathcal{B}$, which provides a method for constructing more silting and tilting objects. Then we apply our results to the cleft extensions of module categories, and characterize silting and tilting modules over $\theta$-extension of rings. Some known results over trivial extension of rings are extended and strengthened.

math.RT

$\mathcal{X}$-Gorenstein projective and $\mathcal{Y}$-Gorenstein injective modules over tensor rings

Let $T_R(M)$ be a tensor ring and $\mathcal{X}$, $\mathcal{Y}$ be two classes of $R$-modules. Under certain conditions, we prove that a $T_R(M)$-module $(A, u)$ is $Ind(\mathcal{X})$-Gorenstein projective if and only if $u$ is monomorphic and $coker(u)$ is an $\mathcal{X}$-Gorenstein projective $R$-module. $\mathcal{Y}$-Gorenstein injective $T_R(M)$-modules are also explicitly described. As a consequence, the characterizations of Ding projective and Ding injective modules over $T_R(M)$ are obtained. Some applications to trivial ring extensions and Morita context rings are given.

math.RA

Construction of Gorenstein projective modules over tensor rings

For a tensor ring $T_R(M)$, we obtain sufficient and necessary conditions to describe all complete projective resolutions and all Gorenstein projective modules. As a consequence, we provide a method for constructing Gorenstein projective modules over $T_R(M)$ from the ones of $R$. Some applications to trivial ring extensions, Morita context rings and triangular matrix rings are given.

math.AC

Constructions of symmetric separable equivalences and their applications

Let $\Lambda$ and $\Gamma$ be symmetrically separably equivalent Artin algebras. We prove that there exist symmetrical separable equivalences between certain endomorphism algebras of modules. As applications, we provide several methods to construct symmetrical separable equivalences from given ones and discuss when the rigidity dimension is an invariant under symmetrical separable equivalences. Moreover, we show that a symmetrical separable equivalence preserves the Frobenius-finite type, Auslander-type condition, the (strong) Nakayama conjecture, the Auslander-Gorenstein conjecture and so on.

math.RT

Magnetism of kagome metals $\left(\text{Fe}_{1-x} \text{Co}_{x}\right) \text{Sn}$ studied by $\mu$SR

We study the magnetic properties of the metallic kagome system $\left(\mathrm{Fe}_{1-x} \mathrm{Co}_{x}\right) \mathrm{Sn}$ by a combination of Muon Spin Relaxation ($\mu \mathrm{SR}$), magnetic susceptibility and Scanning Tunneling Microscopy (STM) measurements, in single crystal specimens with Co concentrations $\mathrm{x}=0,0.11,0.8$. In the undoped antiferromagnetic compound FeSn, we find possible signatures for a previously unidentified phase that sets in at $T^*\sim 50$ K, well beneath the Neel temperature $T_N \sim 376$ K, as indicated by a peak in the relaxation rate $1/T_1$ observed in zero field (ZF) and longitudinal field (LF) $\mu \mathrm{SR}$ measurements, with a corresponding anomaly in the ac and dc-susceptibility, and an increase in the static width $1/T_2$ in ZF measurements. No signatures of spatial symmetry breaking are found in STM down to $7$ K. In $\mathrm{Fe}_{0.2} \mathrm{Co}_{0.8} \mathrm{Sn}$, we find canonical spin glass behavior with freezing temperature $T_{g} \sim 3.5 \mathrm{~K}$; the ZF and LF time spectra exhibit results similar to those observed in dilute alloy spin glasses CuMn and AuFe, with a critical behavior of $1 / T_{1}$ at $T_{g}$ and $1 / \mathrm{T}_{1}\rightarrow 0$ as $T \rightarrow 0$. The absence of spin dynamics at low temperatures makes a clear contrast to the spin dynamics observed by $\mu \mathrm{SR}$ in many geometrically frustrated spin systems on insulating kagome, pyrochlore, and triangular lattices. The spin glass behavior of CoSn doped with dilute Fe moments is shown to originate primarily from the randomness of doped Fe moments rather than due to geometrical frustration of the underlying lattice.

cond-mat.str-el

$\omega$-left approximation dimensions under Stable equivalence

In this paper, we investigate some transfer properties of $\omega$-left approximation dimensions of modules of stably equivalent Artin algebras having neither nodes nor semisimple direct summands. As applications, we give a one-to-one correspondence between basic (Wakamatsu) tilting modules, and prove that the Wakamatsu tilting conjecture is preserved under those equivalences.

math.RT

Gorenstein categories and separable equivalences

Let $\mathscr{C}$ be an additive subcategory of left $\Lambda$-modules, we establish relations of the orthogonal classes of $\mathscr{C}$ and (co)res $\widetilde{\mathscr{C}}$ under separable equivalences. As applications, we obtain that the (one-sided) Gorenstein category and Wakamatsu tilting module are preserved under separable equivalences. Furthermore, we discuss when $G_{C}$-projective (injective) modules and Auslander (Bass) class with respect to $C$ are invariant under separable equivalences.

math.RT

Unveiling the Potential of Segment Anything Model 2 for RGB-Thermal Semantic Segmentation with Language Guidance

The perception capability of robotic systems relies on the richness of the dataset. Although Segment Anything Model 2 (SAM2), trained on large datasets, demonstrates strong perception potential in perception tasks, its inherent training paradigm prevents it from being suitable for RGB-T tasks. To address these challenges, we propose SHIFNet, a novel SAM2-driven Hybrid Interaction Paradigm that unlocks the potential of SAM2 with linguistic guidance for efficient RGB-Thermal perception. Our framework consists of two key components: (1) Semantic-Aware Cross-modal Fusion (SACF) module that dynamically balances modality contributions through text-guided affinity learning, overcoming SAM2's inherent RGB bias; (2) Heterogeneous Prompting Decoder (HPD) that enhances global semantic information through a semantic enhancement module and then combined with category embeddings to amplify cross-modal semantic consistency. With 32.27M trainable parameters, SHIFNet achieves state-of-the-art segmentation performance on public benchmarks, reaching 89.8% on PST900 and 67.8% on FMB, respectively. The framework facilitates the adaptation of pre-trained large models to RGB-T segmentation tasks, effectively mitigating the high costs associated with data collection while endowing robotic systems with comprehensive perception capabilities. The source code will be made publicly available at https://github.com/iAsakiT3T/SHIFNet.

cs.CV

Deep Learning-based Cross-modal Reconstruction of Vehicle Target from Sparse 3D SAR Image

Three-dimensional synthetic aperture radar (3D SAR) is an advanced active microwave imaging technology widely utilized in remote sensing area. To achieve high-resolution 3D imaging,3D SAR requires observations from multiple aspects and altitude baselines surrounding the target. However, constrained flight trajectories often lead to sparse observations, which degrade imaging quality, particularly for anisotropic man-made small targets, such as vehicles and aircraft. In the past, compressive sensing (CS) was the mainstream approach for sparse 3D SAR image reconstruction. More recently, deep learning (DL) has emerged as a powerful alternative, markedly boosting reconstruction quality and efficiency. However, existing DL-based methods typically rely solely on high-quality 3D SAR images as supervisory signals to train deep neural networks (DNNs). This unimodal learning paradigm prevents the integration of complementary information from other data modalities, which limits reconstruction performance and reduces target discriminability due to the inherent constraints of electromagnetic scattering. In this paper, we introduce cross-modal learning and propose a Cross-Modal 3D-SAR Reconstruction Network (CMAR-Net) for enhancing sparse 3D SAR images of vehicle targets by fusing optical information. Leveraging cross-modal supervision from 2D optical images and error propagation guaranteed by differentiable rendering, CMAR-Net achieves efficient training and reconstructs sparse 3D SAR images, which are derived from highly sparse-aspect observations, into visually structured 3D vehicle images. Trained exclusively on simulated data, CMAR-Net exhibits robust generalization to real-world data, outperforming state-of-the-art CS and DL methods in structural accuracy within a large-scale parking lot experiment involving numerous civilian vehicles, thereby demonstrating its strong practical applicability.

cs.CV

AxonCallosumEM Dataset: Axon Semantic Segmentation of Whole Corpus Callosum cross section from EM Images

The electron microscope (EM) remains the predominant technique for elucidating intricate details of the animal nervous system at the nanometer scale. However, accurately reconstructing the complex morphology of axons and myelin sheaths poses a significant challenge. Furthermore, the absence of publicly available, large-scale EM datasets encompassing complete cross sections of the corpus callosum, with dense ground truth segmentation for axons and myelin sheaths, hinders the advancement and evaluation of holistic corpus callosum reconstructions. To surmount these obstacles, we introduce the AxonCallosumEM dataset, comprising a 1.83 times 5.76mm EM image captured from the corpus callosum of the Rett Syndrome (RTT) mouse model, which entail extensive axon bundles. We meticulously proofread over 600,000 patches at a resolution of 1024 times 1024, thus providing a comprehensive ground truth for myelinated axons and myelin sheaths. Additionally, we extensively annotated three distinct regions within the dataset for the purposes of training, testing, and validation. Utilizing this dataset, we develop a fine-tuning methodology that adapts Segment Anything Model (SAM) to EM images segmentation tasks, called EM-SAM, enabling outperforms other state-of-the-art methods. Furthermore, we present the evaluation results of EM-SAM as a baseline.

eess.IV

Towards Large-scale Single-shot Millimeter-wave Imaging for Low-cost Security Inspection

Millimeter-wave (MMW) imaging is emerging as a promising technique for safe security inspection. It achieves a delicate balance between imaging resolution, penetrability and human safety, resulting in higher resolution compared to low-frequency microwave, stronger penetrability compared to visible light, and stronger safety compared to X ray. Despite of recent advance in the last decades, the high cost of requisite large-scale antenna array hinders widespread adoption of MMW imaging in practice. To tackle this challenge, we report a large-scale single-shot MMW imaging framework using sparse antenna array, achieving low-cost but high-fidelity security inspection under an interpretable learning scheme. We first collected extensive full-sampled MMW echoes to study the statistical ranking of each element in the large-scale array. These elements are then sampled based on the ranking, building the experimentally optimal sparse sampling strategy that reduces the cost of antenna array by up to one order of magnitude. Additionally, we derived an untrained interpretable learning scheme, which realizes robust and accurate image reconstruction from sparsely sampled echoes. Last, we developed a neural network for automatic object detection, and experimentally demonstrated successful detection of concealed centimeter-sized targets using 10% sparse array, whereas all the other contemporary approaches failed at the same sample sampling ratio. The performance of the reported technique presents higher than 50% superiority over the existing MMW imaging schemes on various metrics including precision, recall, and mAP50. With such strong detection ability and order-of-magnitude cost reduction, we anticipate that this technique provides a practical way for large-scale single-shot MMW imaging, and could advocate its further practical applications.

eess.IV

Compressive Sensing Based Sparse MIMO Array Optimization for Wideband Near-Field Imaging

In the area of near-field millimeter-wave imaging, the generalized sparse array synthesis (SAS) method is in great demand. The traditional methods usually employ the greedy algorithms, which may have the convergence problem. This paper proposes a convex optimization model for the multiple-input multiple-output (MIMO) array design based on the compressive sensing (CS) approach. We generate a block shaped reference pattern, to be used as an optimizing target. The pattern occupies the entire imaging area of interest in order to involve the effect of each pixel into the optimization model. In MIMO scenarios, we can fix the transmit subarray and synthesize the receive subarray, and vice versa, or doing the synthesis sequentially. The problems associated with focusing, sidelobes suppression, and grating lobes suppression of the synthesized array are examined in details. Numerical and experimental results demonstrate that the synthesized sparse array can offer better image qualities than the sparse arrays with equally spaced or randomly spaced antennas with the same number of antenna elements.

eess.SP

Near-Field Millimeter-Wave Imaging via Arrays in the Shape of Polyline

This paper proposes a polyline shaped array based system scheme, associated with mechanical scanning along the perpendicular direction of the array, for near-field millimeter-wave (MMW) imaging. Each section of the polyline is a chord of a circle with equal length. The polyline array, which can be realized as a monostatic array or a multistatic one, is capable of providing more observation angles than the linear or planar arrays. Further, we present the related three-dimensional (3-D) imaging algorithms based on a hybrid processing in the time domain and the spatial frequency domain. The nonuniform fast Fourier transform (NUFFT) is utilized to improve the computational efficiency. Simulations and experimental results are provided to demonstrate the efficacy of the proposed method in comparison with the back-projection (BP) algorithm.

eess.SP