SearcharxivSearch

arXiv subjects

Shiqi Yu

Publications and source records attributed to Shiqi Yu.

At least 19 recordsLinked to original sources

Galactic Microquasar and Supernova Remnants Imprinting on Diffuse Neutrino and Gamma-Ray Sky

Recent detections of Galactic diffuse neutrinos by IceCube and $\gamma$-rays by LHAASO offer direct probes into the origin of Galactic cosmic rays. Conventional diffuse templates typically assume a single cosmic-ray injection spectrum across a wide energy range, without accounting for independent contributions from distinct accelerator populations. Here, we present a numerical framework that models Galactic diffuse neutrino and $\gamma$-ray emission that incorporates contributions from both microquasar and supernova remnant populations. By anchoring CR injection to local observations and utilizing high-resolution 3D target gas distributions, our model suggests multi-population contributions to the diffuse sky: escaped cosmic rays from supernova remnants dominate below $\sim10\text{ TeV}$, while those from microquasars become the primary driver at higher energies. Our predicted neutrino flux agrees well with the recent 12-year IceCube measurements, establishing a physically motivated baseline for the diffuse hadronic background while leaving room for unresolved point-like sources. This flexible framework provides testable predictions for current and future multi-messenger observatories, accommodating diverse accelerator populations and updated observational constraints.

astro-ph.HE

IJCB-AFMFR 2026: Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data

This paper presents a summary of the Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data (AFMFR), held at the 2026 International Joint Conference on Biometrics (IJCB 2026). The competition received a total of eight valid submissions from four distinct teams across two complementary tracks: a Full Data Track, in which participants adapt the CLIP ViT-L/14 foundation model using large-scale synthetic identity data, and a Limited Data Track, designed to reflect more resource-constrained adaptation regimes. All training data was generated exclusively using IDPERTURB. Submitted solutions are ranked based on verification and identification performance across a diverse suite of benchmarks, including LFW, CFP-FP, AgeDB-30, CALFW, CPLFW, IJB-B, IJB-C, and TinyFace, using the Borda count method. Fairness evaluation is additionally conducted on the RFW dataset across four demographic groups. The results demonstrate that adaptation of the CLIP foundation model with synthetic training data substantially improves over the off-the-shelf model and, in several cases, surpasses the baseline. Notably, full fine-tuning with Sub-Center ArcFace (DMSTI-Neurotechnology) leads the Full Data Track, while rank-stabilized LoRA adaptation (Idiap-BSP) proves most effective under limited-data conditions.

cs.CV

Microquasar Remnants as Pevatrons Illuminating the Galactic Cosmic Ray Knee

Microquasars are primary candidates for Galactic PeVatrons, yet their collective contribution to the cosmic ray (CR) ``knee" remains poorly understood. We investigate this contribution by simulating anisotropic diffusive propagation through the Galactic magnetic field. Our results demonstrate that spatial alignment and magnetic connectivity between source locations and the solar neighborhood govern the local flux: sources aligned with local magnetic field lines yield pronounced flux enhancements, whereas magnetically disconnected locations are strongly suppressed. We find that the characteristic proton bump near PeV is robust across random realizations of the Galactic microquasar remnant population, with as few as four nearby remnants accounting for 50\% of the observed PeV flux. Our findings suggest that the integrated history of microquasar remnants, governed by source temporal and spatial distribution and magnetic transport, naturally populates the observed CR ``knee''.

astro-ph.HE

Human Identification at a Distance: Challenges, Methods and Results on the Competition HID 2025

Human identification at a distance (HID) is challenging because traditional biometric modalities such as face and fingerprints are often difficult to acquire in real-world scenarios. Gait recognition provides a practical alternative, as it can be captured reliably at a distance. To promote progress in gait recognition and provide a fair evaluation platform, the International Competition on Human Identification at a Distance (HID) has been organized annually since 2020. Since 2023, the competition has adopted the challenging SUSTech-Competition dataset, which features substantial variations in clothing, carried objects, and view angles. No dedicated training data are provided, requiring participants to train their models using external datasets. Each year, the competition applies a different random seed to generate distinct evaluation splits, which reduces the risk of overfitting and supports a fair assessment of cross-domain generalization. While HID 2023 and HID 2024 already used this dataset, HID 2025 explicitly examined whether algorithmic advances could surpass the accuracy limits observed previously. Despite the heightened difficulty, participants achieved further improvements, and the best-performing method reached 94.2% accuracy, setting a new benchmark on this dataset. We also analyze key technical trends and outline potential directions for future research in gait recognition.

cs.CV

Multi-Messenger Modeling of Low-Luminosity Gamma-Ray Bursts

Low-luminosity gamma-ray bursts (LL GRBs), a subclass of the most powerful transients in the Universe, remain promising sources of high-energy astrophysical neutrinos, despite strong IceCube constraints on typical long GRBs. In this work, a novel approach is introduced to study a sample of seven LL~GRBs with their multi-wavelength observations to investigate leptohadronic processes during their prompt emission phases. The relative energy densities in magnetic fields, non-thermal electrons, and protons are constrained, with the latter defining the cosmic-ray (CR) loading factor. Our results suggest that LL~GRBs exhibit diverse emission processes, as confirmed by a machine-learning analysis of the fitted parameters. Across the seven LL~GRBs, we find the posterior medians of the CR loading factor in the range of $\xi_p \sim 0.2$--$1.6$. GRB~060218 and GRB~100316D, the lowest-luminosity bursts ($L_{\gamma, \rm iso} \sim 10^{46}$-$10^{47}\rm~erg~s^{-1}$) consistent with the shock-breakout (SBO) scenario, yield the highest CR loading factor and therefore are expected to produce neutrinos more efficiently. Our model predicts the expected number of neutrino signals that are consistent with current limits but would be detectable with next-generation neutrino observatories. These results strengthen the case for LL~GRBs as promising sources of high-energy astrophysical neutrinos and motivate real-time searches for coincident LL~GRB and neutrino events. Next-generation X-ray and MeV facilities will be critical for identifying more LL~GRBs and strengthening their role in multi-messenger astrophysics.

astro-ph.HE

A Unified Framework for 10 TeV to EeV Diffuse Neutrino Sky and KM3-230213A

Establishing a unified framework that simultaneously accounts for the wideband diffuse neutrino flux and the physical origin of individual ultra-high-energy (UHE) neutrino detections, including KM3-230213A, remains a pressing challenge in multi-messenger astrophysics. In intrinsically low-luminosity gamma-ray bursts (LL~GRBs) driven by shock breakouts (SBOs), the evolving physical conditions naturally produce a multicomponent neutrino flux extending from 10 TeV to the EeV scale. By integrating prompt and afterglow phases within a unified framework grounded in multiwavelength observations of representative events, we show that LL GRB population accounts for this broadband neutrino emission through a characteristic two-hump spectrum. In this framework, the prompt emission from GRB~060218-like events accounts for $\gtrsim 10\%$ of the diffuse flux at 100~TeV, while GRB~100316D-like afterglow configuration predicts a distinct flux peak near $10^{-9}\rm~GeV~cm^{-2}~s^{-1}~sr^{-1}$ at 100~PeV. This two-hump spectrum provides a high-energy component flux consistent with the 220 PeV KM3-230213A event, while the low-energy component contributes non-trivially to the observed diffuse neutrinos and supports the lack of individual low-energy counterparts. Furthermore, we utilize Fermi-LAT gamma-ray upper limits to place constraints on the source distance and luminosity of the event, assuming a GRB 100316D-like afterglow configuration. Ultimately, this framework identifies SBO-like LL~GRBs as a unifying origin for these phenomena, providing a physical link across the 10 TeV to EeV neutrino sky that is testable by next-generation observatories, including GRAND, IceCube-Gen2, and RNO-G.

astro-ph.HE

Is Visual Realism Enough? Evaluating Gait Biometric Fidelity in Generative AI Human Animation

Generative AI (GenAI) models have revolutionized animation, enabling the synthesis of humans and motion patterns with remarkable visual fidelity. However, generating truly realistic human animation remains a formidable challenge, where even minor inconsistencies can make a subject appear unnatural. This limitation is particularly critical when AI-generated videos are evaluated for behavioral biometrics, where subtle motion cues that define identity are easily lost or distorted. The present study investigates whether state-of-the-art GenAI human animation models can preserve the subtle spatio-temporal details needed for person identification through gait biometrics. Specifically, we evaluate four different GenAI models across two primary evaluation tasks to assess their ability to i) restore gait patterns from reference videos under varying conditions of complexity, and ii) transfer these gait patterns to different visual identities. Our results show that while visual quality is mostly high, biometric fidelity remains low in tasks focusing on identification, suggesting that current GenAI models struggle to disentangle identity from motion. Furthermore, through an identity transfer task, we expose a fundamental flaw in appearance-based gait recognition: when texture is disentangled from motion, identification collapses, proving current GenAI models rely on visual attributes rather than temporal dynamics.

cs.CV

Prefrontal scaling of reward prediction error readout gates reinforcement-derived adaptive behavior in primates

Reinforcement learning (RL) enables adaptive behavior across species via reward prediction errors (RPEs), but the neural origins of species-specific adaptability remain unknown. Integrating RL modeling, transcriptomics, and neuroimaging during reversal learning, we discovered convergent RPE signatures - shared monoaminergic/synaptic gene upregulation and neuroanatomical representations, yet humans outperformed macaques behaviorally. Single-trial decoding showed RPEs guided choices similarly in both species, but humans disproportionately recruited dorsal anterior cingulate (dACC) and dorsolateral prefrontal cortex (dlPFC). Cross-species alignment uncovered that macaque prefrontal circuits encode human-like optimal RPEs yet fail to translate them into action. Adaptability scaled not with RPE encoding fidelity, but with the areal extent of dACC/dlPFC recruitment governing RPE-to-action transformation. These findings resolve an evolutionary puzzle: behavioral performance gaps arise from executive cortical readout efficiency, not encoding capacity.

q-bio.NC

IntAttention: A Fully Integer Attention Pipeline for Efficient Edge Inference

Deploying Transformer models on edge devices is limited by latency and energy budgets. While INT8 quantization effectively accelerates the primary matrix multiplications, it exposes the softmax-related path as the dominant bottleneck. This stage incurs a costly dequantize -> softmax -> requantize detour, which can account for up to 65% of total attention latency and disrupts the end-to-end integer dataflow critical for edge hardware efficiency. To address this limitation, we present IntAttention, the first fully integer attention pipeline that serves as a training-free drop-in replacement. At the core of our approach lies IndexSoftmax, a hardware-friendly operator that replaces floating-point exponentials entirely within the integer domain. IntAttention integrates sparsity-aware clipping, a 32-entry lookup table approximation, and direct integer normalization, thereby eliminating datatype conversion overhead along the attention path. Experiments on Armv8 CPUs show that our method achieves up to 3.7x speedup and 61% energy reduction over FP16 baselines, and up to 2.0x speedup over conventional INT8 attention pipelines. Across diverse language and vision models, as well as additional reasoning and long-context evaluations, IntAttention maintains strong overall fidelity and demonstrates a more favorable trade-off than existing LUT-based softmax approximations. Code is available at https://github.com/WanliZhong/IntAttention

cs.LG

Pose as Clinical Prior: Learning Dual Representations for Scoliosis Screening

Recent AI-based scoliosis screening methods primarily rely on large-scale silhouette datasets, often neglecting clinically relevant postural asymmetries-key indicators in traditional screening. In contrast, pose data provide an intuitive skeletal representation, enhancing clinical interpretability across various medical applications. However, pose-based scoliosis screening remains underexplored due to two main challenges: (1) the scarcity of large-scale, annotated pose datasets; and (2) the discrete and noise-sensitive nature of raw pose coordinates, which hinders the modeling of subtle asymmetries. To address these limitations, we introduce Scoliosis1K-Pose, a 2D human pose annotation set that extends the original Scoliosis1K dataset, comprising 447,900 frames of 2D keypoints from 1,050 adolescents. Building on this dataset, we introduce the Dual Representation Framework (DRF), which integrates a continuous skeleton map to preserve spatial structure with a discrete Postural Asymmetry Vector (PAV) that encodes clinically relevant asymmetry descriptors. A novel PAV-Guided Attention (PGA) module further uses the PAV as clinical prior to direct feature extraction from the skeleton map, focusing on clinically meaningful asymmetries. Extensive experiments demonstrate that DRF achieves state-of-the-art performance. Visualizations further confirm that the model leverages clinical asymmetry cues to guide feature extraction and promote synergy between its dual representations. The dataset and code are publicly available at https://zhouzi180.github.io/Scoliosis1K/.

cs.CV

On Denoising Walking Videos for Gait Recognition

To capture individual gait patterns, excluding identity-irrelevant cues in walking videos, such as clothing texture and color, remains a persistent challenge for vision-based gait recognition. Traditional silhouette- and pose-based methods, though theoretically effective at removing such distractions, often fall short of high accuracy due to their sparse and less informative inputs. Emerging end-to-end methods address this by directly denoising RGB videos using human priors. Building on this trend, we propose DenoisingGait, a novel gait denoising method. Inspired by the philosophy that "what I cannot create, I do not understand", we turn to generative diffusion models, uncovering how they partially filter out irrelevant factors for gait understanding. Additionally, we introduce a geometry-driven Feature Matching module, which, combined with background removal via human silhouettes, condenses the multi-channel diffusion features at each foreground pixel into a two-channel direction vector. Specifically, the proposed within- and cross-frame matching respectively capture the local vectorized structures of gait appearance and motion, producing a novel flow-like gait representation termed Gait Feature Field, which further reduces residual noise in diffusion features. Experiments on the CCPG, CASIA-B*, and SUSTech1K datasets demonstrate that DenoisingGait achieves a new SoTA performance in most cases for both within- and cross-domain evaluations. Code is available at https://github.com/ShiqiYu/OpenGait.

cs.CV

BiggerGait: Unlocking Gait Recognition with Layer-wise Representations from Large Vision Models

Large vision models (LVM) based gait recognition has achieved impressive performance. However, existing LVM-based approaches may overemphasize gait priors while neglecting the intrinsic value of LVM itself, particularly the rich, distinct representations across its multi-layers. To adequately unlock LVM's potential, this work investigates the impact of layer-wise representations on downstream recognition tasks. Our analysis reveals that LVM's intermediate layers offer complementary properties across tasks, integrating them yields an impressive improvement even without rich well-designed gait priors. Building on this insight, we propose a simple and universal baseline for LVM-based gait recognition, termed BiggerGait. Comprehensive evaluations on CCPG, CAISA-B*, SUSTech1K, and CCGR\_MINI validate the superiority of BiggerGait across both within- and cross-domain tasks, establishing it as a simple yet practical baseline for gait representation learning. All the models and code will be publicly available.

cs.CV

Improving the generalization of gait recognition with limited datasets

Generalized gait recognition remains challenging due to significant domain shifts in viewpoints, appearances, and environments. Mixed-dataset training has recently become a practical route to improve cross-domain robustness, but it introduces underexplored issues: 1) inter-dataset supervision conflicts, which distract identity learning, and 2) redundant or noisy samples, which reduce data efficiency and may reinforce dataset-specific patterns. To address these challenges, we introduce a unified paradigm for cross-dataset gait learning that simultaneously improves motion-signal quality and supervision consistency. We first increase the reliability of training data by suppressing sequences dominated by redundant gait cycles or unstable silhouettes, guided by representation redundancy and prediction uncertainty. This refinement concentrates learning on informative gait dynamics when mixing heterogeneous datasets. In parallel, we stabilize supervision by disentangling metric learning across datasets, forming triplets within each source to prevent destructive cross-domain gradients while preserving transferable identity cues. These components act in synergy to stabilize optimization and strengthen generalization without modifying network architectures or requiring extra annotations. Experiments on CASIA-B, OU-MVLP, Gait3D, and GREW with both GaitBase and DeepGaitV2 backbones consistently show improved cross-domain performance without sacrificing in-domain accuracy. These results demonstrate that data selection and aligning supervision effectively enables scalable mixed-dataset gait learning.

cs.CV

Neutrinos in IceCube and KM3NeT

Neutrino observatories such as IceCube, Cubic Kilometre Neutrino Telescope (KM3NeT), and Super-Kamiokande cover a broad energy range that enables the study of both atmospheric neutrinos and astrophysical neutrinos. IceCube and KM3NeT focus on a similar energy range, from a few GeV to PeV, and have conducted competitive work on the atmospheric neutrino flux, three-flavor oscillation parameter measurements, searches beyond the Standard Model, and investigations of cosmic-ray accelerators using high-energy astrophysical neutrinos. Recent IceCube findings of evidence of neutrino signals from NGC~1068 have triggered a series of follow-up studies. These studies provide evidence that a subset of Seyfert galaxies may produce high-energy neutrinos. The emerging candidates are NGC~4151, NGC~3079, CGCG~420-015, and Circinus Galaxy. Furthermore, a stacking analysis of 13 selected sources in the Southern Hemisphere reported a cumulative neutrino signal at 3.0\,$\sigma$, offering independent evidence that some X-ray-bright Seyfert galaxies could be potential high-energy neutrino sources. KM3NeT, still under construction, continues to accumulate data that will support future studies of astrophysical neutrino sources. However, with its currently deployed detection units, it has detected an ultra-high-energy event of several tens of PeV originating from approximately 1 degree above the horizon. This contribution highlights and summarizes recent findings from IceCube and KM3NeT in both neutrino physics and astrophysics.

astro-ph.HE

Exploring More from Multiple Gait Modalities for Human Identification

The gait, as a kind of soft biometric characteristic, can reflect the distinct walking patterns of individuals at a distance, exhibiting a promising technique for unrestrained human identification. With largely excluding gait-unrelated cues hidden in RGB videos, the silhouette and skeleton, though visually compact, have acted as two of the most prevailing gait modalities for a long time. Recently, several attempts have been made to introduce more informative data forms like human parsing and optical flow images to capture gait characteristics, along with multi-branch architectures. However, due to the inconsistency within model designs and experiment settings, we argue that a comprehensive and fair comparative study among these popular gait modalities, involving the representational capacity and fusion strategy exploration, is still lacking. From the perspectives of fine vs. coarse-grained shape and whole vs. pixel-wise motion modeling, this work presents an in-depth investigation of three popular gait representations, i.e., silhouette, human parsing, and optical flow, with various fusion evaluations, and experimentally exposes their similarities and differences. Based on the obtained insights, we further develop a C$^2$Fusion strategy, consequently building our new framework MultiGait++. C$^2$Fusion preserves commonalities while highlighting differences to enrich the learning of gait features. To verify our findings and conclusions, extensive experiments on Gait3D, GREW, CCPG, and SUSTech1K are conducted. The code is available at https://github.com/ShiqiYu/OpenGait.

cs.CV

Gait Patterns as Biomarkers: A Video-Based Approach for Classifying Scoliosis

Scoliosis presents significant diagnostic challenges, particularly in adolescents, where early detection is crucial for effective treatment. Traditional diagnostic and follow-up methods, which rely on physical examinations and radiography, face limitations due to the need for clinical expertise and the risk of radiation exposure, thus restricting their use for widespread early screening. In response, we introduce a novel video-based, non-invasive method for scoliosis classification using gait analysis, effectively circumventing these limitations. This study presents Scoliosis1K, the first large-scale dataset specifically designed for video-based scoliosis classification, encompassing over one thousand adolescents. Leveraging this dataset, we developed ScoNet, an initial model that faced challenges in handling the complexities of real-world data. This led to the development of ScoNet-MT, an enhanced model incorporating multi-task learning, which demonstrates promising diagnostic accuracy for practical applications. Our findings demonstrate that gait can serve as a non-invasive biomarker for scoliosis, revolutionizing screening practices through deep learning and setting a precedent for non-invasive diagnostic methodologies. The dataset and code are publicly available at https://zhouzi180.github.io/Scoliosis1K/.

cs.CV

OpenGait: A Comprehensive Benchmark Study for Gait Recognition towards Better Practicality

Gait recognition, a rapidly advancing vision technology for person identification from a distance, has made significant strides in indoor settings. However, evidence suggests that existing methods often yield unsatisfactory results when applied to newly released real-world gait datasets. Furthermore, conclusions drawn from indoor gait datasets may not easily generalize to outdoor ones. Therefore, the primary goal of this paper is to present a comprehensive benchmark study aimed at improving practicality rather than solely focusing on enhancing performance. To this end, we developed OpenGait, a flexible and efficient gait recognition platform. Using OpenGait, we conducted in-depth ablation experiments to revisit recent developments in gait recognition. Surprisingly, we detected some imperfect parts of some prior methods and thereby uncovered several critical yet previously neglected insights. These findings led us to develop three structurally simple yet empirically powerful and practically robust baseline models: DeepGaitV2, SkeletonGait, and SkeletonGait++, which represent the appearance-based, model-based, and multi-modal methodologies for gait pattern description, respectively. In addition to achieving state-of-the-art performance, our careful exploration provides new perspectives on the modeling experience of deep gait models and the representational capacity of typical gait modalities. In the end, we discuss the key trends and challenges in current gait recognition, aiming to inspire further advancements towards better practicality. The code is available at https://github.com/ShiqiYu/OpenGait.

cs.CV

Cross-Modality Gait Recognition: Bridging LiDAR and Camera Modalities for Human Identification

Current gait recognition research mainly focuses on identifying pedestrians captured by the same type of sensor, neglecting the fact that individuals may be captured by different sensors in order to adapt to various environments. A more practical approach should involve cross-modality matching across different sensors. Hence, this paper focuses on investigating the problem of cross-modality gait recognition, with the objective of accurately identifying pedestrians across diverse vision sensors. We present CrossGait inspired by the feature alignment strategy, capable of cross retrieving diverse data modalities. Specifically, we investigate the cross-modality recognition task by initially extracting features within each modality and subsequently aligning these features across modalities. To further enhance the cross-modality performance, we propose a Prototypical Modality-shared Attention Module that learns modality-shared features from two modality-specific features. Additionally, we design a Cross-modality Feature Adapter that transforms the learned modality-specific features into a unified feature space. Extensive experiments conducted on the SUSTech1K dataset demonstrate the effectiveness of CrossGait: (1) it exhibits promising cross-modality ability in retrieving pedestrians across various modalities from different sensors in diverse scenes, and (2) CrossGait not only learns modality-shared features for cross-modality gait recognition but also maintains modality-specific features for single-modality recognition.

cs.CV