SearcharxivSearch

arXiv subjects

Weijia Fan

Publications and source records attributed to Weijia Fan.

10 recordsLinked to original sources

X$^2$Localizer: Cross-grained Alignment for Progressive Cross-view Video Geo-localization

Cross-view Video Geo-localization (CVG) aims to localize ground-view videos by retrieving their corresponding geo-tagged aerial images. However, CVG approaches rely on fixed-length inputs and post-hoc refinement, hindering online-oriented localization under partial or dynamic observations. In this work, we formulate Progressive Cross-view Video Geo-localization (PCVG) as a deployment-oriented extension and evaluation protocol of CVG, enabling localization under varying temporal budgets, prefix-based inference, random-start evaluation, and long-range localization with interruptions. To explore PCVG, we introduce X$^2$Localizer, a cross-grained alignment framework that jointly supervises global prefix-to-aerial retrieval and token-aggregated frame--aerial-tile matching with a budget-dependent asymmetric objective. Furthermore, we introduce a Sliding-Window Re-Localization (SWRL) strategy that dynamically refreshes candidate regions for failure recovery and long-range deployment without full-sequence reprocessing. Extensive experiments show that X$^2$Localizer preserves conventional full-video performance, with marginal gains of +0.1 Recall@1 and +0.3 Recall@10, while substantially improving early localization. In the challenging single-frame setting, X$^2$Localizer improves coarse retrieval by +4.7 Recall@1 and +11.5 Recall@10 over the previous state-of-the-art method. With SWRL, our approach further enables robust progressive localization under random-start and long-distance scenarios, narrowing the gap between benchmark evaluation and real-world deployment.

cs.CV

Faster or Stronger: Towards Flexible Visual Place Recognition via Weighted Aggregation and Token Pruning

Visual Place Recognition (VPR) aims to match a query image to reference images of the same place in a large-scale database. Recent state-of-the-art methods employ Vision Transformers (ViTs) as backbone foundation models to extract patch-level features that are robust to viewpoint, illumination, and seasonal variations, which are then aggregated into a compact global descriptor for retrieval. Most existing aggregation methods uniformly pool patch tokens into learned clusters, despite the fact that different clusters often encode distinct spatial or semantic patterns and contribute unequally to VPR performance. To address this limitation, we propose Weighted Aggregated Descriptor (WeiAD), which assigns weights to clusters during aggregation, producing more discriminative global representations. Beyond accuracy, retrieval latency is a critical concern for large-scale deployments and resource-constrained edge devices. Prior work mainly reduces latency by compressing global descriptors, while overlooking the cost of feature extraction, an issue exacerbated by ViT-based backbones. We therefore introduce WeiToP, a VPR-oriented token pruning framework that reduces feature extraction cost via self-distillation, where aggregation-induced token importance supervises a lightweight pruning module attached to an early transformer layer, enabling inference-time token pruning. After a single joint training phase, WeiToP enables plug-and-play token pruning at inference time, allowing flexible and on-demand control over the accuracy-efficiency trade-off without additional training. Moreover, WeiToP outperforms existing token pruning methods adapted from general vision tasks.

cs.CV

More than the Sum: Panorama-Language Models for Adverse Omni-Scenes

Existing vision-language models (VLMs) are tailored for pinhole imagery, stitching multiple narrow field-of-view inputs to piece together a complete omni-scene understanding. Yet, such multi-view perception overlooks the holistic spatial and contextual relationships that a single panorama inherently preserves. In this work, we introduce the Panorama-Language Modeling (PLM)paradigm, a unified $360^\circ$ vision-language reasoning that is more than the sum of its pinhole counterparts. Besides, we present PanoVQA, a large-scale panoramic VQA dataset that involves adverse omni-scenes, enabling comprehensive reasoning under object occlusions and driving accidents. To establish a foundation for PLM, we develop a plug-and-play panoramic sparse attention module that allows existing pinhole-based VLMs to process equirectangular panoramas without retraining. Extensive experiments demonstrate that our PLM achieves superior robustness and holistic reasoning under challenging omni-scenes, yielding understanding greater than the sum of its narrow parts. Project page: https://github.com/InSAI-Lab/PanoVQA.

cs.CV

SGR3 Model: Scene Graph Retrieval-Reasoning Model in 3D

3D scene graphs provide a structured representation of object entities and their relationships, enabling high-level interpretation and reasoning for robots while remaining intuitively understandable to humans. Existing approaches for 3D scene graph generation typically combine scene reconstruction with graph neural networks (GNNs). However, such pipelines require multi-modal data that may not always be available, and their reliance on heuristic graph construction can constrain the prediction of relationship triplets. In this work, we introduce a Scene Graph Retrieval-Reasoning Model in 3D (SGR3 Model), a training-free framework that leverages multi-modal large language models (MLLMs) with retrieval-augmented generation (RAG) for semantic scene graph generation. SGR3 Model bypasses the need for explicit 3D reconstruction. Instead, it enhances relational reasoning by incorporating semantically aligned scene graphs retrieved via a ColPali-style cross-modal framework. To improve retrieval robustness, we further introduce a weighted patch-level similarity selection mechanism that mitigates the negative impact of blurry or semantically uninformative regions. Experiments demonstrate that SGR3 Model achieves competitive performance compared to training-free baselines and on par with GNN-based expert models. Moreover, an ablation study on the retrieval module and knowledge base scale reveals that retrieved external information is explicitly integrated into the token generation process, rather than being implicitly internalized through abstraction.

cs.CV

BCE3S: Binary Cross-Entropy Based Tripartite Synergistic Learning for Long-tailed Recognition

For long-tailed recognition (LTR) tasks, high intra-class compactness and inter-class separability in both head and tail classes, as well as balanced separability among all the classifier vectors, are preferred. The existing LTR methods based on cross-entropy (CE) loss not only struggle to learn features with desirable properties but also couple imbalanced classifier vectors in the denominator of its Softmax, amplifying the imbalance effects in LTR. In this paper, for the LTR, we propose a binary cross-entropy (BCE)-based tripartite synergistic learning, termed BCE3S, which consists of three components: (1) BCE-based joint learning optimizes both the classifier and sample features, which achieves better compactness and separability among features than the CE-based joint learning, by decoupling the metrics between feature and the imbalanced classifier vectors in multiple Sigmoid; (2) BCE-based contrastive learning further improves the intra-class compactness of features; (3) BCE-based uniform learning balances the separability among classifier vectors and interactively enhances the feature properties by combining with the joint learning. The extensive experiments show that the LTR model trained by BCE3S not only achieves higher compactness and separability among sample features, but also balances the classifier's separability, achieving SOTA performance on various long-tailed datasets such as CIFAR10-LT, CIFAR100-LT, ImageNet-LT, and iNaturalist2018.

cs.CV

Unveiling the nonrelativistic spin current polarization in an altermagnet

Spin current plays a central role in spintronics for driving exotic spin-dependent phenomena and high-performance device applications. Recently, a magnetic spin Hall effect has been discovered in spin-split antiferromagnets including noncollinear antiferromagnets and altermagnets, allowing the efficient generation of unconventional spin currents even in the absence of spin-orbit coupling. However, although such nonrelativistic spin currents are proposed to have a magnetic origin, the direct connection between the Neel vector and the spin current polarization is still missing. Here, using the altermagnetic RuO2 as a representative example, we unveil the spin current polarization by disentangling the conventional and magnetic spin Hall effects using a technique we developed based on the spin Hall magnetoresistance measurement. The results suggest that the nonrelativistic spin current hosts a polarization very close to the Neel vector. Our work offers unambiguous evidence of the magnetic origin of the nonrelativistic spin current in altermagnetic RuO2, and paves a straightforward route to understand the unconventional spin currents that are crucial in spintronics.

cond-mat.mes-hall

EPL: Empirical Prototype Learning for Deep Face Recognition

Prototype learning is widely used in face recognition, which takes the row vectors of coefficient matrix in the last linear layer of the feature extraction model as the prototypes for each class. When the prototypes are updated using the facial sample feature gradients in the model training, they are prone to being pulled away from the class center by the hard samples, resulting in decreased overall model performance. In this paper, we explicitly define prototypes as the expectations of sample features in each class and design the empirical prototypes using the existing samples in the dataset. We then devise a strategy to adaptively update these empirical prototypes during the model training based on the similarity between the sample features and the empirical prototypes. Furthermore, we propose an empirical prototype learning (EPL) method, which utilizes an adaptive margin parameter with respect to sample features. EPL assigns larger margins to the normal samples and smaller margins to the hard samples, allowing the learned empirical prototypes to better reflect the class center dominated by the normal samples and finally pull the hard samples towards the empirical prototypes through the learning. The extensive experiments on MFR, IJB-C, LFW, CFP-FP, AgeDB, and MegaFace demonstrate the effectiveness of EPL. Our code is available at $\href{https://github.com/WakingHours-GitHub/EPL}{https://github.com/WakingHours-GitHub/EPL}$.

cs.CV

Spin-Transfer-Torque Induced Spatially Nonuniform Switching in Ferrimagnets

Ferrimagnet (FiM), (FeCo)1-xGdx, attracts research attention due to its ultrafast magnetic dynamics and finite net magnetization. Incorporating FiM into the magnetic tunnel junction will be beneficial to further improve the writing speed of magnetic random access memory (MRAM). It is commonly assumed that the FeCo and Gd atoms are switched together due to the strong exchange coupling, which remains valid even if one performs the two-sublattice macrospin simulation. Interestingly, using the atomistic model developed by our group, it is clearly seen that different atoms are not switched together. In addition, our study reveals that the nature of switching is spatially nonuniform even in the small sample with the dimension of 20 nm-20 nm. Furthermore, the characteristics of nonuniformity are completely different for samples with different Gd composition (x). When x is close to the magnetization compensation point, successful switching cannot be obtained, but is accompanied by the stable oscillation. The atom type that dominates the oscillation is different from that predicted by the two-sublattice macrospin model. In addition, the size of singular region is a non-monotonic function of current density. All these results can only be understood by considering the spatial nonuniform magnetization dynamics.

cond-mat.mes-hall

Efficient field-free perpendicular magnetization switching by a magnetic spin Hall effect

Current induced spin-orbit torques driven by the conventional spin Hall effect are widely used to manipulate the magnetization. This approach, however, is nondeterministic and inefficient for the switching of magnets with perpendicular magnetic anisotropy that are demanded by the high-density magnetic storage and memory devices. Here, we demonstrate that this limitation can be overcome by exploiting a magnetic spin Hall effect in noncollinear antiferromagnets, such as Mn3Sn. The magnetic group symmetry of Mn3Sn allows generation of the out-of-plane spin current carrying spin polarization induced by an in-plane charge current. This spin current drives an out-of-plane anti-damping torque providing deterministic switching of perpendicular magnetization of an adjacent Ni/Co multilayer. Compared to the conventional spin-orbit torque devices, the observed switching does not need any external magnetic field and requires much lower current density. Our results demonstrate great prospects of exploiting the magnetic spin Hall effect in noncollinear antiferromagnets for low-power spintronics.

cond-mat.mtrl-sci

A multistate approach for mediation analysis in the presence of semi-competing risks with application in cancer survival disparities

We propose a novel methodology to quantify the effect of stochastic interventions on non-terminal time-to-events that lie on the pathway between an exposure and a terminal time-to-event outcome. Investigating these effects is particularly important in health disparities research when we seek to quantify inequities in timely delivery of treatment and its impact on patients survival time. Current approaches fail to account for semi-competing risks arising in this setting. Under the potential outcome framework, we define and provide identifiability conditions for causal estimands for stochastic direct and indirect effects. Causal contrasts are estimated in continuous time within a multistate modeling framework and analytic formulae for the estimators of the causal contrasts are developed. We show via simulations that ignoring censoring in mediator and or outcome time-to-event processes, or ignoring competing risks may give misleading results. This work demonstrates that rigorous definition of the direct and indirect effects and joint estimation of the outcome and mediator time-to-event distributions in the presence of semi-competing risks are crucial for valid investigation of mechanisms in continuous time. We employ this novel methodology to investigate the role of delaying treatment uptake in explaining racial disparities in cancer survival in a cohort study of colon cancer patients.

stat.ME