SearcharxivSearch

arXiv subjects

Ruifeng Wang

Publications and source records attributed to Ruifeng Wang.

13 recordsLinked to original sources

Entity-Aware Sequence Transduction for Player-Centric Ball Action Spotting

Player-centric ball action spotting requires temporally precise event detection together with actor attribution in crowded, partially observed multi-agent sports videos. Existing Denoising Sequence Transduction (DST) baselines treat the player-role dimension as part of a flattened frame-level representation, which weakens the inductive bias for modeling player-specific temporal evolution and inter-player interactions. To address this limitation, we propose Multi-Entity Denoising Sequence Transduction (ME-DST). ME-DST keeps the role-slot dimension throughout encoding. It uses temporal attention to model the history of each role slot, and spatial attention to exchange information across role slots at each frame. This factorized design gives the model a direct structure for separating within-player evolution from inter-player context. We also add learnable role embeddings, tracking-derived tactical features, and fused visual predictions from X3D-L and Swin3D-S. Experiments on the FOOTPASS dataset show that ME-DST reaches a Micro F1 of 0.778. This improves the strongest official TAAD+DST baseline by 10.3 percentage points. Controlled ablations show that preserving the entity axis and encoding role identity are central to this gain. These results suggest that explicit entity modeling is an effective inductive bias for player-centric sports event understanding.

cs.CV

SoccerNet 2026 Challenges Results

The SoccerNet 2026 Challenges constitute the sixth annual edition of the SoccerNet open benchmarking effort, dedicated to advancing computer vision research in sports video understanding. This year's challenges span five vision-based tasks: (1) Ball Action Anticipation, predicting the timing and class of ball-related actions within a short future window from a preceding observation window; (2) Player-Centric Ball Action Spotting, temporally localizing and classifying ball-related actions while assigning each action to the acting player through team affiliation and jersey number; (3) Novel View Synthesis, rendering images from unobserved camera poses in multi-view football scenes; (4) Spiideo SoccerNet Synloc, localizing athletes in real-world pitch coordinates from a single calibrated static-camera image; and (5) Visual Question Answering, answering multiple-choice questions about football broadcasts across text, image, and video inputs. For each task, participants were provided with annotated data, a unified evaluation protocol, and a public baseline. This edition saw broad participation, with 427 teams submitting 1,129 entries across the five tasks and 28 teams contributing reviewed technical reports. This paper describes each task and its evaluation protocol, presents the challenge leaderboards, and summarizes the leading submissions, with the aim of documenting the current state of each task as measured on held-out challenge data.

cs.CV

Nonmonotonic Scaling of the Anomalous Hall Effect in a Bicollinear Antiferromagnet

An anomalous Hall effect (AHE) in antiferromagnetic (AF) systems with no net magnetization is of considerable interest for both fundamental physics and spintronic applications. Of particular interest is the two-dimensional van der Waals antiferromagnet FeTe that has an unusual fully magnetically compensated bicollinear AF structure and exhibits pronounced Kondo interaction leading to strong band renormalization. Here, we investigate the AHE in epitaxial FeTe thin films grown by molecular beam epitaxy. A large anomalous Hall conductivity is exhibited below the Neel temperature (T_N ~ 60 K) and, strikingly, becomes nonlinear at high fields within a narrow temperature window around 49 K, deviating from conventional AHE scaling behavior versus its longitudinal conductivity. Linear fits reveal a pronounced negative peak in the intercept, accompanied by a field-induced canted magnetic moment. The AHE responses are related to the Berry curvature derived from FeTe's topological band structure, highlighting the intricate interplay between topology, magnetism, and electronic transport.

cond-mat.mtrl-sci

FMRFT: Fusion Mamba and DETR for Query Time Sequence Intersection Fish Tracking

Early detection of abnormal fish behavior caused by disease or hunger can be achieved through fish tracking using deep learning techniques, which holds significant value for industrial aquaculture. However, underwater reflections and some reasons with fish, such as the high similarity, rapid swimming caused by stimuli and mutual occlusion bring challenges to multi-target tracking of fish. To address these challenges, this paper establishes a complex multi-scenario sturgeon tracking dataset and introduces the FMRFT model, a real-time end-to-end fish tracking solution. The model incorporates the low video memory consumption Mamba In Mamba (MIM) architecture, which facilitates multi-frame temporal memory and feature extraction, thereby addressing the challenges to track multiple fish across frames. Additionally, the FMRFT model with the Query Time Sequence Intersection (QTSI) module effectively manages occluded objects and reduces redundant tracking frames using the superior feature interaction and prior frame processing capabilities of RT-DETR. This combination significantly enhances the accuracy and stability of fish tracking. Trained and tested on the dataset, the model achieves an IDF1 score of 90.3% and a MOTA accuracy of 94.3%. Experimental results show that the proposed FMRFT model effectively addresses the challenges of high similarity and mutual occlusion in fish populations, enabling accurate tracking in factory farming environments.

cs.CV

FA-YOLO: Research On Efficient Feature Selection YOLO Improved Algorithm Based On FMDS and AGMF Modules

Over the past few years, the YOLO series of models has emerged as one of the dominant methodologies in the realm of object detection. Many studies have advanced these baseline models by modifying their architectures, enhancing data quality, and developing new loss functions. However, current models still exhibit deficiencies in processing feature maps, such as overlooking the fusion of cross-scale features and a static fusion approach that lacks the capability for dynamic feature adjustment. To address these issues, this paper introduces an efficient Fine-grained Multi-scale Dynamic Selection Module (FMDS Module), which applies a more effective dynamic feature selection and fusion method on fine-grained multi-scale feature maps, significantly enhancing the detection accuracy of small, medium, and large-sized targets in complex environments. Furthermore, this paper proposes an Adaptive Gated Multi-branch Focus Fusion Module (AGMF Module), which utilizes multiple parallel branches to perform complementary fusion of various features captured by the gated unit branch, FMDS Module branch, and TripletAttention branch. This approach further enhances the comprehensiveness, diversity, and integrity of feature fusion. This paper has integrated the FMDS Module, AGMF Module, into Yolov9 to develop a novel object detection model named FA-YOLO. Extensive experimental results show that under identical experimental conditions, FA-YOLO achieves an outstanding 66.1% mean Average Precision (mAP) on the PASCAL VOC 2007 dataset, representing 1.0% improvement over YOLOv9's 65.1%. Additionally, the detection accuracies of FA-YOLO for small, medium, and large targets are 44.1%, 54.6%, and 70.8%, respectively, showing improvements of 2.0%, 3.1%, and 0.9% compared to YOLOv9's 42.1%, 51.5%, and 69.9%.

cs.CV

Where to Fetch: Extracting Visual Scene Representation from Large Pre-Trained Models for Robotic Goal Navigation

To complete a complex task where a robot navigates to a goal object and fetches it, the robot needs to have a good understanding of the instructions and the surrounding environment. Large pre-trained models have shown capabilities to interpret tasks defined via language descriptions. However, previous methods attempting to integrate large pre-trained models with daily tasks are not competent in many robotic goal navigation tasks due to poor understanding of the environment. In this work, we present a visual scene representation built with large-scale visual language models to form a feature representation of the environment capable of handling natural language queries. Combined with large language models, this method can parse language instructions into action sequences for a robot to follow, and accomplish goal navigation with querying the scene representation. Experiments demonstrate that our method enables the robot to follow a wide range of instructions and complete complex goal navigation tasks.

cs.RO

Self-supervised transformer-based pre-training method with General Plant Infection dataset

Pest and disease classification is a challenging issue in agriculture. The performance of deep learning models is intricately linked to training data diversity and quantity, posing issues for plant pest and disease datasets that remain underdeveloped. This study addresses these challenges by constructing a comprehensive dataset and proposing an advanced network architecture that combines Contrastive Learning and Masked Image Modeling (MIM). The dataset comprises diverse plant species and pest categories, making it one of the largest and most varied in the field. The proposed network architecture demonstrates effectiveness in addressing plant pest and disease recognition tasks, achieving notable detection accuracy. This approach offers a viable solution for rapid, efficient, and cost-effective plant pest and disease detection, thereby reducing agricultural production costs. Our code and dataset will be publicly available to advance research in plant pest and disease recognition the GitHub repository at https://github.com/WASSER2545/GPID-22

cs.CV

Conformance Testing of Relational DBMS Against SQL Specifications

A Relational Database Management System (RDBMS) is one of the fundamental software that supports a wide range of applications, making it critical to identify bugs within these systems. There has been active research on testing RDBMS, most of which employ crash or use metamorphic relations as the oracle. Although existing approaches can detect bugs in RDBMS, they are far from comprehensively evaluating the RDBMS's correctness (i.e., with respect to the semantics of SQL). In this work, we propose a method to test the semantic conformance of RDBMS i.e., whether its behavior respects the intended semantics of SQL. Specifically, we have formally defined the semantics of SQL and implemented them in Prolog. Then, the Prolog implementation serves as the reference RDBMS, enabling differential testing on existing RDBMS. We applied our approach to four widely-used and thoroughly tested RDBMSs, i.e., MySQL, TiDB, SQLite, and DuckDB. In total, our approach uncovered 19 bugs and 11 inconsistencies, which are all related to violating the SQL specification or missing/unclear specification, thereby demonstrating the effectiveness and applicability of our approach.

cs.DB

Neural Radiance Field-based Visual Rendering: A Comprehensive Review

In recent years, Neural Radiance Fields (NeRF) has made remarkable progress in the field of computer vision and graphics, providing strong technical support for solving key tasks including 3D scene understanding, new perspective synthesis, human body reconstruction, robotics, and so on, the attention of academics to this research result is growing. As a revolutionary neural implicit field representation, NeRF has caused a continuous research boom in the academic community. Therefore, the purpose of this review is to provide an in-depth analysis of the research literature on NeRF within the past two years, to provide a comprehensive academic perspective for budding researchers. In this paper, the core architecture of NeRF is first elaborated in detail, followed by a discussion of various improvement strategies for NeRF, and case studies of NeRF in diverse application scenarios, demonstrating its practical utility in different domains. In terms of datasets and evaluation metrics, This paper details the key resources needed for NeRF model training. Finally, this paper provides a prospective discussion on the future development trends and potential challenges of NeRF, aiming to provide research inspiration for researchers in the field and to promote the further development of related technologies.

cs.CV

Direct observation of nodeless superconductivity and phonon modes in electron-doped copper oxide Sr$_{1-x}$Nd$_x$CuO$_2$

The microscopic understanding of high-temperature superconductivity in cuprates has been hindered by the apparent complexity of crystal structures in these materials. We used scanning tunneling microscopy and spectroscopy to study an electron-doped copper oxide compound Sr$_{1-x}$Nd$_x$CuO$_2$ that has only bare cations separating the CuO$_2$ planes and thus the simplest infinite-layer structure among all cuprate superconductors. Tunneling conductance spectra of the major CuO$_2$ planes in the superconducting state revealed direct evidence for a nodeless pairing gap, regardless of variation of its magnitude with the local doping of trivalent neodymium. Furthermore, three distinct bosonic modes are observed as multiple peak-dip-hump features outside the superconducting gaps and their respective energies depend little on the spatially varying gaps. Along with the bosonic modes with energies identical to those of the external, bending and stretching phonons of copper oxides, our findings indicate their origin from lattice vibrations rather than spin excitations.

cond-mat.supr-con

A Thermoplastic Elastomer Belt Based Robotic Gripper

Novel robotic grippers have captured increasing interests recently because of their abilities to adapt to varieties of circumstances and their powerful functionalities. Differing from traditional gripper with mechanical components-made fingers, novel robotic grippers are typically made of novel structures and materials, using a novel manufacturing process. In this paper, a novel robotic gripper with external frame and internal thermoplastic elastomer belt-made net is proposed. The gripper grasps objects using the friction between the net and objects. It has the ability of adaptive gripping through flexible contact surface. Stress simulation has been used to explore the regularity between the normal stress on the net and the deformation of the net. Experiments are conducted on a variety of objects to measure the force needed to reliably grip and hold the object. Test results show that the gripper can successfully grip objects with varying shape, dimensions, and textures. It is promising that the gripper can be used for grasping fragile objects in the industry or out in the field, and also grasping the marine organisms without hurting them.

cs.RO

Direct visualization of ambipolar Mott transition in cuprate CuO2 planes

Identifying the essence of doped Mott insulators is one of the major outstanding problems in condensed matter physics and the key to understanding the high-temperature superconductivity in cuprates. We report real space visualization of Mott transition in Sr1-xLaxCuO2+y cuprate films that cover the entire electron- and hole-doped regimes. Tunneling conductance measurements directly on the cooper-oxide (CuO2) planes reveal a systematic shift in the Fermi level, while the fundamental Mott-Hubbard band structure remains unchanged. This is further demonstrated by exploring atomic-scale electronic response of CuO2 to substitutional dopants and intrinsic defects in a sister compound Sr0.92Nd0.08CuO2. The results could be better explained in the framework of self-modulation doping, similar to that in semiconductor heterostructures, and form a basis for developing any microscopic theories for cuprate superconductivity.

cond-mat.supr-con

Real-space observation of charge ordering in epitaxial La2-xSrxCuO4 films

The cuprate superconductors exhibit ubiquitous instabilities toward charge-ordered states. These unusual electronic states break the spatial symmetries of the host crystal, and have been widely appreciated as essential ingredients for constructing a theory for high-temperature superconductivity in cuprates. Here we report real-space imaging of the doping-dependent charge orders in the epitaxial thin films of a canonical cuprate compound La2-xSrxCuO4 using scanning tunneling microscopy. As the films are moderately doped, we observe a crossover from incommensurate to commensurate (4a0, where a0 is the Cu-O-Cu distance) stripes. Furthermore, at lower and higher doping levels, the charge orders occur in the form of distorted Wigner crystal and grid phase of crossed vertical and horizontal stripes. We discuss how the charge orders are stabilized, and their interplay with superconductivity.

cond-mat.supr-con