SearcharxivSearch

arXiv subjects

Tengfei Zhang

Publications and source records attributed to Tengfei Zhang.

15 recordsLinked to original sources

A Vision-language Framework for Comparative Reasoning in Radiology

Medical imaging artificial intelligence has achieved strong performance in isolated image interpretation, but remains poorly aligned with radiological practice, where diagnosis and follow-up rely on comparison across prior studies and analogous reference cases. Here we formulate radiological comparison as an entity-aware cross-image reasoning problem and introduce a framework that supports both reference-case retrieval and temporal comparative interpretation. We construct MedReCo-DB, a large-scale comparative imaging resource derived from routine image-report pairs, comprising more than 690,000 images from over 160,000 patients across eight institutions, four countries and seven imaging modalities. Reports are decomposed into anatomical structures, abnormal findings and pathological conditions to provide supervision for entity-conditioned retrieval and comparative visual question answering. Using this resource, we develop MedReCo, an entity-aware visual encoder for controllable retrieval of clinically analogous cases, and MedReCo-VLM, a vision--language extension for generative interpretation of interval change. Across internal, external and cross-center evaluations, MedReCo achieved the highest Recall@1 in all 12 internal retrieval settings and improved external retrieval by a mean of 6.0 percentage points. In clinically confusable differential groups, it consistently outperformed the strongest baselines. MedReCo-VLM achieved the best performance across all comparative generation evaluations and improved longitudinal follow-up accuracy by 14.5-46.5 percentage points on chest radiographs and 13.0-27.9 percentage points on CT. These findings suggest that entity-aware comparative reasoning can be learned from routine clinical data at scale and may provide a more clinically aligned foundation for medical imaging AI.

cs.CV

GeVI-SLAM: Gravity-Enhanced Stereo Visua Inertial SLAM for Underwater Robots

Accurate visual inertial simultaneous localization and mapping (VI SLAM) for underwater robots remains a significant challenge due to frequent visual degeneracy and insufficient inertial measurement unit (IMU) motion excitation. In this paper, we present GeVI-SLAM, a gravity-enhanced stereo VI SLAM system designed to address these issues. By leveraging the stereo camera's direct depth estimation ability, we eliminate the need to estimate scale during IMU initialization, enabling stable operation even under low acceleration dynamics. With precise gravity initialization, we decouple the pitch and roll from the pose estimation and solve a 4 degrees of freedom (DOF) Perspective-n-Point (PnP) problem for pose tracking. This allows the use of a minimal 3-point solver, which significantly reduces computational time to reject outliers within a Random Sample Consensus framework. We further propose a bias-eliminated 4-DOF PnP estimator with provable consistency, ensuring the relative pose converges to the true value as the feature number increases. To handle dynamic motion, we refine the full 6-DOF pose while jointly estimating the IMU covariance, enabling adaptive weighting of the gravity prior. Extensive experiments on simulated and real-world data demonstrate that GeVI-SLAM achieves higher accuracy and greater stability compared to state-of-the-art methods.

cs.RO

Spatiotemporal Calibration of Doppler Velocity Logs for Underwater Robots

The calibration of extrinsic parameters and clock offsets between sensors for high-accuracy performance in underwater SLAM systems remains insufficiently explored. Existing methods for Doppler Velocity Log (DVL) calibration are either constrained to specific sensor configurations or rely on oversimplified assumptions, and none jointly estimate translational extrinsics and time offsets. We propose a Unified Iterative Calibration (UIC) framework for general DVL sensor setups, formulated as a Maximum A Posteriori (MAP) estimation with a Gaussian Process (GP) motion prior for high-fidelity motion interpolation. UIC alternates between efficient GP-based motion state updates and gradient-based calibration variable updates, supported by a provably statistically consistent sequential initialization scheme. The proposed UIC can be applied to IMU, cameras and other modalities as co-sensors. We release an open-source DVL-camera calibration toolbox. Beyond underwater applications, several aspects of UIC-such as the integration of GP priors for MAP-based calibration and the design of provably reliable initialization procedures-are broadly applicable to other multi-sensor calibration problems. Finally, simulations and real-world tests validate our approach.

cs.RO

RadIR: A Scalable Framework for Multi-Grained Medical Image Retrieval via Radiology Report Mining

Developing advanced medical imaging retrieval systems is challenging due to the varying definitions of `similar images' across different medical contexts. This challenge is compounded by the lack of large-scale, high-quality medical imaging retrieval datasets and benchmarks. In this paper, we propose a novel methodology that leverages dense radiology reports to define image-wise similarity ordering at multiple granularities in a scalable and fully automatic manner. Using this approach, we construct two comprehensive medical imaging retrieval datasets: MIMIC-IR for Chest X-rays and CTRATE-IR for CT scans, providing detailed image-image ranking annotations conditioned on diverse anatomical structures. Furthermore, we develop two retrieval systems, RadIR-CXR and model-ChestCT, which demonstrate superior performance in traditional image-image and image-report retrieval tasks. These systems also enable flexible, effective image retrieval conditioned on specific anatomical structures described in text, achieving state-of-the-art results on 77 out of 78 metrics.

cs.CV

First Place Solution to the ECCV 2024 ROAD++ Challenge @ ROAD++ Spatiotemporal Agent Detection 2024

This report presents our team's solutions for the Track 1 of the 2024 ECCV ROAD++ Challenge. The task of Track 1 is spatiotemporal agent detection, which aims to construct an "agent tube" for road agents in consecutive video frames. Our solutions focus on the challenges in this task, including extreme-size objects, low-light scenarios, class imbalance, and fine-grained classification. Firstly, the extreme-size object detection heads are introduced to improve the detection performance of large and small objects. Secondly, we design a dual-stream detection model with a low-light enhancement stream to improve the performance of spatiotemporal agent detection in low-light scenes, and the feature fusion module to integrate features from different branches. Subsequently, we develop a multi-branch detection framework to mitigate the issues of class imbalance and fine-grained classification, and we design a pre-training and fine-tuning approach to optimize the above multi-branch framework. Besides, we employ some common data augmentation techniques, and improve the loss function and upsampling operation. We rank first in the test set of Track 1 for the ROAD++ Challenge 2024, and achieve 30.82% average video-mAP.

cs.CV

First Place Solution to the ECCV 2024 ROAD++ Challenge @ ROAD++ Atomic Activity Recognition 2024

This report presents our team's technical solution for participating in Track 3 of the 2024 ECCV ROAD++ Challenge. The task of Track 3 is atomic activity recognition, which aims to identify 64 types of atomic activities in road scenes based on video content. Our approach primarily addresses the challenges of small objects, discriminating between single object and a group of objects, as well as model overfitting in this task. Firstly, we construct a multi-branch activity recognition framework that not only separates different object categories but also the tasks of single object and object group recognition, thereby enhancing recognition accuracy. Subsequently, we develop various model ensembling strategies, including integrations of multiple frame sampling sequences, different frame sampling sequence lengths, multiple training epochs, and different backbone networks. Furthermore, we propose an atomic activity recognition data augmentation method, which greatly expands the sample space by flipping video frames and road topology, effectively mitigating model overfitting. Our methods rank first in the test set of Track 3 for the ROAD++ Challenge 2024, and achieve 69% mAP.

cs.CV

Solvable non-Hermitian skin effects and real-space exceptional points: Non-Hermitian generalized Bloch theorem

Non-Hermitian systems can exhibit extraordinary boundary behaviors, known as the non-Hermitian skin effects, where all the eigenstates are localized exponentially at one side of lattice model. To give a full understanding and control of non-Hermitian skin effects, we have developed the non-Hermitian generalized Bloch theorem to provide the analytical expression for all solvable eigenvalues and eigenstates, in which translation symmetry is broken due to the open boundary condition. By introducing the Vieta's theorem for any polynomial equation with arbitrary degree, our approach is widely applicable for one-dimensional non-Hermitian tight-binding models. With the non-Hermitian generalized Bloch theorem, we can analyze the condition of existence or non-existence of the non-Hermitian skin effects at a mathematically rigorous level. Additionally, the non-Hermitian generalized Bloch theorem allows us to explore the real-space exceptional points. We also establish the connection between our approach and the generalized Brillouin zone method. To illustrate our main results, we examine two concrete examples including the Su-Schrieffer-Heeger chain model with long-range couplings, and the ladder model with non-reciprocal interaction. Our non-Hermitian generalized Bloch theorem provides an efficient way to analytically study various non-Hermitian phenomena in more general cases.

quant-ph

Synergy between Machine/Deep Learning and Software Engineering: How Far Are We?

Since 2009, the deep learning revolution, which was triggered by the introduction of ImageNet, has stimulated the synergy between Machine Learning (ML)/Deep Learning (DL) and Software Engineering (SE). Meanwhile, critical reviews have emerged that suggest that ML/DL should be used cautiously. To improve the quality (especially the applicability and generalizability) of ML/DL-related SE studies, and to stimulate and enhance future collaborations between SE/AI researchers and industry practitioners, we conducted a 10-year Systematic Literature Review (SLR) on 906 ML/DL-related SE papers published between 2009 and 2018. Our trend analysis demonstrated the mutual impacts that ML/DL and SE have had on each other. At the same time, however, we also observed a paucity of replicable and reproducible ML/DL-related SE studies and identified five factors that influence their replicability and reproducibility. To improve the applicability and generalizability of research results, we analyzed what ingredients in a study would facilitate an understanding of why a ML/DL technique was selected for a specific SE problem. In addition, we identified the unique trends of impacts of DL models on SE tasks, as well as five unique challenges that needed to be met in order to better leverage DL to improve the productivity of SE tasks. Finally, we outlined a road-map that we believe can facilitate the transfer of ML/DL-based SE research results into real-world industry practices.

cs.SE

Comparison Network for One-Shot Conditional Object Detection

The current advances in object detection depend on large-scale datasets to get good performance. However, there may not always be sufficient samples in many scenarios, which leads to the research on few-shot detection as well as its extreme variation one-shot detection. In this paper, the one-shot detection has been formulated as a conditional probability problem. With this insight, a novel one-shot conditional object detection (OSCD) framework, referred as Comparison Network (ComparisonNet), has been proposed. Specifically, query and target image features are extracted through a Siamese network as mapped metrics of marginal probabilities. A two-stage detector for OSCD is introduced to compare the extracted query and target features with the learnable metric to approach the optimized non-linear conditional probability. Once trained, ComparisonNet can detect objects of both seen and unseen classes without further training, which also has the advantages including class-agnostic, training-free for unseen classes, and without catastrophic forgetting. Experiments show that the proposed approach achieves state-of-the-art performance on the proposed datasets of Fashion-MNIST and PASCAL VOC.

cs.CV

SCRDet: Towards More Robust Detection for Small, Cluttered and Rotated Objects

Object detection has been a building block in computer vision. Though considerable progress has been made, there still exist challenges for objects with small size, arbitrary direction, and dense distribution. Apart from natural images, such issues are especially pronounced for aerial images of great importance. This paper presents a novel multi-category rotation detector for small, cluttered and rotated objects, namely SCRDet. Specifically, a sampling fusion network is devised which fuses multi-layer feature with effective anchor sampling, to improve the sensitivity to small objects. Meanwhile, the supervised pixel attention network and the channel attention network are jointly explored for small and cluttered object detection by suppressing the noise and highlighting the objects feature. For more accurate rotation estimation, the IoU constant factor is added to the smooth L1 loss to address the boundary problem for the rotating bounding box. Extensive experiments on two remote sensing public datasets DOTA, NWPU VHR-10 as well as natural image datasets COCO, VOC2007 and scene text data ICDAR2015 show the state-of-the-art performance of our detector. The code and models will be available at https://github.com/DetectionTeamUCAS.

cs.CV

A Training-free, One-shot Detection Framework For Geospatial Objects In Remote Sensing Images

Deep learning based object detection has achieved great success. However, these supervised learning methods are data-hungry and time-consuming. This restriction makes them unsuitable for limited data and urgent tasks, especially in the applications of remote sensing. Inspired by the ability of humans to quickly learn new visual concepts from very few examples, we propose a training-free, one-shot geospatial object detection framework for remote sensing images. It consists of (1) a feature extractor with remote sensing domain knowledge, (2) a multi-level feature fusion method, (3) a novel similarity metric method, and (4) a 2-stage object detection pipeline. Experiments on sewage treatment plant and airport detections show that proposed method has achieved a certain effect. Our method can serve as a baseline for training-free, one-shot geospatial object detection.

cs.CV

High Efficiency and Low Distortion Photoacoustic Effect in 3D Graphene Sponge

The conversion of light in sound plays a crucial role in spectroscopy, applied physics, and technology. In this paper, light sound conversion in 3D graphene sponge through a photothermoacoustic mechanism is reported. It is shown that the unique combination of mechanical, optical, and thermodynamic properties of graphene assembled in a 3D sponge structure allows an unprecedented high efficiency conversion independent of light wavelength from infrared to ultraviolet. As a first application of this effect, a photothermal based graphene sponge loudspeaker is demonstrated, providing a full digital operation for frequencies from acoustic to ultrasound. The present results suggest a new pathway for light generation and control of sound and ultrasound signals potentially usable in a variety of new technological applications from high fidelity loudspeaker and radiation detectors to medical devices.

physics.app-ph

On the 4D Nonlinear Schrödinger equation with combined terms under the energy threshold

In this paper, we consider the longtime dynamics of the solutions to focusing energy-critical Schrödinger equation with a defocusing energy-subcritical perturbation term under a ground state energy threshold in four spatial dimension. This extends the results in Miao et al. (Commun Math Phys 318(3):767-808, 2013, The dynamics of the NLS with the combined terms in five and higher dimensions. Some topics in harmonic analysis and applications, advanced lectures in mathematics, ALM34, Higher Education Press, Beijing, pp 265-298, 2015) to four dimension without radial assumption and the proof of scattering is based on the interaction Morawetz estimates developed in Dodson (Global well-posedness and scattering for the focusing, energy-critical nonlinear Schröinger problem in dimension $d=4$ for initial data below a ground state threshold, arXiv:1409.1950), the main ingredients of which requires us to overcome the logarithmic failure in the double Duhamel argument in four dimensions.

math.AP

Macroscopic and direct light propulsion of bulk graphene material

It has been a great challenge to achieve the direct light manipulation of matter on a bulk scale. In this work, the direct light propulsion of matter was observed on a macroscopic scale for the first time using a bulk graphene based material. The unique structure and properties of graphene and the morphology of the bulk graphene material make it capable of not only absorbing light at various wavelengths but also emitting energetic electrons efficiently enough to drive the bulk material following Newtonian mechanics. Thus, the unique photonic and electronic properties of individual graphene sheets are manifested in the response of the bulk state. These results offer an exciting opportunity to bring about bulk scale light manipulation with the potential to realize long-sought proposals in areas such as the solar sail and space transportation driven directly by sunlight.

physics.optics

Solar LImb Prominence CAtcher and Tracker (SLIPCAT): An Automated System and Its Preliminary Statistical Results

In this paper, we present an automated system, which has the capability to catch and track solar limb prominences based on observations from EUV 304 passband. The characteristic parameters and their evolution, including height, position angle, area, length and brightness, are obtained without manual interventions. By applying the system to the STEREO-B/SECCHI/EUVI 304 data during 2007 April -2009 October, we obtain a total of 9477 well-tracked prominences and a catalog of these events available online at http://space.ustc.edu.cn/dreams/slipcat/. A detailed analysis of these prominences suggests that the system has a rather good performance. We have obtained several interesting statistical results based on the catalog. Most prominences appear below the latitude of 60 degrees and at the height of about 26 Mm above the solar surface. Most of them are quite stable during the period they are tracked. Nevertheless, some prominences have an upward speed of more than 100 km/s, and some others show significant downward and/or azimuthal speeds. There are strong correlations among the brightness, area and height. The expansion of a prominence is probably one major cause of its fading during the rising or erupting process.

astro-ph.SR