SearcharxivSearch

arXiv subjects

Zhenzhen Wang

Publications and source records attributed to Zhenzhen Wang.

10 recordsLinked to original sources

AE-UAV: An Air-to-Air Event-Based UAV Tracking Benchmark and a Real-Time Frequency-Domain Tracker

Air-to-air (A2A) unmanned aerial vehicle (UAV) tracking is fundamental to airborne remote sensing of low-altitude aerial targets. However, the deployment of continuous, real-time tracking systems on UAVs presents significant challenges. In A2A scenarios, traditional frame-based cameras suffer from severe performance degradation under low illumination, overexposure, and high-speed motion owing to their limited dynamic range and fixed temporal sampling. Although event cameras offer a promising alternative with microsecond temporal resolution and a high dynamic range, current research is bottlenecked by two primary issues: 1) the absence of dedicated A2A event-based datasets, and 2) the heavy reliance of existing trackers on GPU acceleration and extensive training data, rendering them impractical for resource-constrained UAVs. To bridge these gaps, we introduce AE-UAV, an air-to-air event-based UAV tracking benchmark. To the best of our knowledge, this is the first airborne-captured event camera dataset for A2A tracking, comprising 178 flight sequences with continuous-time cubic B-spline annotations. Furthermore, we propose the Fast-Slow Frequency-domain Tracking (FSFT) method. This lightweight, training-free framework seamlessly integrates frequency-domain template matching with search region prediction and detection-based drift correction. Extensive experiments demonstrate that FSFT operates at an ultra-fast 420 frames per second (FPS) on CPU-only hardware. It retains 93.97% of the accuracy of state-of-the-art GPU-dependent methods while delivering a 5.32-fold effective speedup and exhibiting superior temporal resolution generalization, thereby providing a highly efficient and robust solution for airborne remote sensing of aerial targets. The dataset and source code are available at https://github.com/MSP-xEN/AE-UAV.

cs.CV

DiffusionQC: Artifact Detection in Histopathology via Diffusion Model

Digital pathology plays a vital role across modern medicine, offering critical insights for disease diagnosis, prognosis, and treatment. However, histopathology images often contain artifacts introduced during slide preparation and digitization. Detecting and excluding them is essential to ensure reliable downstream analysis. Traditional supervised models typically require large annotated datasets, which is resource-intensive and not generalizable to novel artifact types. To address this, we propose DiffusionQC, which detects artifacts as outliers among clean images using a diffusion model. It requires only a set of clean images for training rather than pixel-level artifact annotations and predefined artifact types. Furthermore, we introduce a contrastive learning module to explicitly enlarge the distribution separation between artifact and clean images, yielding an enhanced version of our method. Empirical results demonstrate superior performance to state-of-the-art and offer cross-stain generalization capacity, with significantly less data and annotations.

cs.CV

Direct Preference Optimization for Adaptive Concept-based Explanations

Concept-based explanation methods aim at making machine learning models more transparent by finding the most important semantic features of an input (e.g., colors, patterns, shapes) for a given prediction task. However, these methods generally ignore the communicative context of explanations, such as the preferences of a listener. For example, medical doctors understand explanations in terms of clinical markers, but patients may not, needing a different vocabulary to rationalize the same diagnosis. We address this gap with listener-adaptive explanations grounded in principles of pragmatic reasoning and the rational speech act. We introduce an iterative training procedure based on direct preference optimization where a speaker learns to compose explanations that maximize communicative utility for a listener. Our approach only needs access to pairwise preferences, which can be collected from human feedback, making it particularly relevant in real-world scenarios where a model of the listener may not be available. We demonstrate that our method is able to align speakers with the preferences of simulated listeners on image classification across three datasets, and further validate that pragmatic explanations generated with our method improve the classification accuracy of participants in a user study.

cs.LG

Magnetism and berry phase manipulation in an emergent structure of perovskite ruthenate by (111) strain engineering

The interplay among symmetry of lattices, electronic correlations, and Berry phase of the Bloch states in solids has led to fascinating quantum phases of matter. A prototypical system is the magnetic Weyl candidate SrRuO3, where designing and creating electronic and topological properties on artificial lattice geometry is highly demanded yet remains elusive. Here, we establish an emergent trigonal structure of SrRuO3 by means of heteroepitaxial strain engineering along the [111] crystallographic axis. Distinctive from bulk, the trigonal SrRuO3 exhibits a peculiar XY-type ferromagnetic ground state, with the coexistence of high-mobility holes likely from linear Weyl bands and low-mobility electrons from normal quadratic bands as carriers. The presence of Weyl nodes are further corroborated by capturing intrinsic anomalous Hall effect, acting as momentum-space sources of Berry curvatures. The experimental observations are consistent with our first-principles calculations, shedding light on the detailed band topology of trigonal SrRuO3 with multiple pairs of Weyl nodes near the Fermi level. Our findings signify the essence of magnetism and Berry phase manipulation via lattice design and pave the way towards unveiling nontrivial correlated topological phenomena.

cond-mat.str-el

Model-independent test for the cosmic distance duality relation with Pantheon and eBOSS DR16 quasar sample

In this paper, we carry out a new model-independent cosmological test for the cosmic distance duality relation~(CDDR) by combining the latest five baryon acoustic oscillations (BAO) measurements and the Pantheon type Ia supernova (SNIa) sample. Particularly, the BAO measurement from extended Baryon Oscillation Spectroscopic Survey~(eBOSS) data release~(DR) 16 quasar sample at effective redshift $z=1.48$ is used, and two methods, i.e. a compressed form of Pantheon sample and the Artificial Neural Network~(ANN) combined with the binning SNIa method, are applied to overcome the redshift-matching problem. Our results suggest that the CDDR is compatible with the observations, and the high-redshift BAO and SNIa data can effectively strengthen the constraints on the violation parameters of CDDR with the confidence interval decreasing by more than 20 percent. In addition, we find that the compressed form of observational data can provide a more rigorous constraint on the CDDR, and thus can be generalized to the applications of other actual observational data with limited sample size in the test for CDDR.

astro-ph.CO

Engineered Kondo screening and nonzero Berry phase in SrTiO3/LaTiO3/SrTiO3 heterostructures

Controlling the interplay between localized spins and itinerant electrons at the oxide interfaces can lead to exotic magnetic states. Here we devise SrTiO3/LaTiO3/SrTiO3 heterostructures with varied thickness of the LaTiO3 layer (n monolayers) to investigate the magnetic interactions in the two-dimensional electron gas system. The heterostructures exhibit significant Kondo effect when the LaTiO3 layer is rather thin (n = 2, 10), manifesting the strong interaction between the itinerant electrons and the localized magnetic moments at the interfaces, while the Kondo effect is greatly inhibited when n = 20. Notably, distinct Shubnikov-de Haas oscillations are observed and a nonzero Berry phase of π is extracted when the LaTiO3 layer is rather thin (n = 2, 10), which is absent in the heterostructure with thicker LaTiO3 layer (n = 20). The observed phenomena are consistently interpreted as a result of sub-band splitting and symmetry breaking due to the interplay between the interfacial Rashba spin-orbit coupling and the magnetic orderings in the heterostructures. Our findings provide a route for exploring and manipulating nontrivial electronic band structures at complex oxide interfaces.

cond-mat.str-el

Label Cleaning Multiple Instance Learning: Refining Coarse Annotations on Single Whole-Slide Images

Annotating cancerous regions in whole-slide images (WSIs) of pathology samples plays a critical role in clinical diagnosis, biomedical research, and machine learning algorithms development. However, generating exhaustive and accurate annotations is labor-intensive, challenging, and costly. Drawing only coarse and approximate annotations is a much easier task, less costly, and it alleviates pathologists' workload. In this paper, we study the problem of refining these approximate annotations in digital pathology to obtain more accurate ones. Some previous works have explored obtaining machine learning models from these inaccurate annotations, but few of them tackle the refinement problem where the mislabeled regions should be explicitly identified and corrected, and all of them require a -- often very large -- number of training samples. We present a method, named Label Cleaning Multiple Instance Learning (LC-MIL), to refine coarse annotations on a single WSI without the need of external training data. Patches cropped from a WSI with inaccurate labels are processed jointly within a multiple instance learning framework, mitigating their impact on the predictive model and refining the segmentation. Our experiments on a heterogeneous WSI set with breast cancer lymph node metastasis, liver cancer, and colorectal cancer samples show that LC-MIL significantly refines the coarse annotations, outperforming state-of-the-art alternatives, even while learning from a single slide. Moreover, we demonstrate how real annotations drawn by pathologists can be efficiently refined and improved by the proposed approach. All these results demonstrate that LC-MIL is a promising, light-weight tool to provide fine-grained annotations from coarsely annotated pathology sets.

cs.CV

Attention-Aware Noisy Label Learning for Image Classification

Deep convolutional neural networks (CNNs) learned on large-scale labeled samples have achieved remarkable progress in computer vision, such as image/video classification. The cheapest way to obtain a large body of labeled visual data is to crawl from websites with user-supplied labels, such as Flickr. However, these samples often tend to contain incorrect labels (i.e. noisy labels), which will significantly degrade the network performance. In this paper, the attention-aware noisy label learning approach ($A^2NL$) is proposed to improve the discriminative capability of the network trained on datasets with potential label noise. Specifically, a Noise-Attention model, which contains multiple noise-specific units, is designed to better capture noisy information. Each unit is expected to learn a specific noisy distribution for a subset of images so that different disturbances are more precisely modeled. Furthermore, a recursive learning process is introduced to strengthen the learning ability of the attention network by taking advantage of the learned high-level knowledge. To fully evaluate the proposed method, we conduct experiments from two aspects: manually flipped label noise on large-scale image classification datasets, including CIFAR-10, SVHN; and real-world label noise on an online crawled clothing dataset with multiple attributes. The superior results over state-of-the-art methods validate the effectiveness of our proposed approach.

cs.CV

Deep Reinforcement Learning with Label Embedding Reward for Supervised Image Hashing

Deep hashing has shown promising results in image retrieval and recognition. Despite its success, most existing deep hashing approaches are rather similar: either multi-layer perceptron or CNN is applied to extract image feature, followed by different binarization activation functions such as sigmoid, tanh or autoencoder to generate binary code. In this work, we introduce a novel decision-making approach for deep supervised hashing. We formulate the hashing problem as travelling across the vertices in the binary code space, and learn a deep Q-network with a novel label embedding reward defined by Bose-Chaudhuri-Hocquenghem (BCH) codes to explore the best path. Extensive experiments and analysis on the CIFAR-10 and NUS-WIDE dataset show that our approach outperforms state-of-the-art supervised hashing methods under various code lengths.

cs.CV

Well-balanced finite difference WENO schemes for the blood flow model

The blood flow model maintains the steady state solutions, in which the flux gradients are non-zero but exactly balanced by the source term. In this paper, we design high order finite difference weighted non-oscillatory (WENO) schemes to this model with such well-balanced property and at the same time keeping genuine high order accuracy. Rigorous theoretical analysis as well as extensive numerical results all indicate that the resulting schemes verify high order accuracy, maintain the well-balanced property, and keep good resolution for smooth and discontinuous solutions.

math.NA