SearcharxivSearch

arXiv subjects

Wenhua Zhang

Publications and source records attributed to Wenhua Zhang.

18 recordsLinked to original sources

Exploring the Limits of End-to-End Feature-Affinity Propagation for Single-Point Supervised Infrared Small Target Detection

Single-point supervised infrared small target detection (IRSTD) drastically reduces dense annotation costs. Current state-of-the-art (SOTA) methods achieve high precision by recovering mask supervision through explicit, offline pseudo-label construction, such as multi-stage active learning and physics-driven mask generation. In this paper, we study a minimalist alternative: generating point-to-mask supervision online through in-batch, point-anchored feature-affinity propagation. We instantiate this paradigm as GSACP, an end-to-end testbed that directly supervises the detector using hard-margin feature affinity gated by local image priors, entirely eliminating external label-evolution loops. This compact design, however, exposes an optimization bottleneck. Because the affinity target is generated from the same feature representation being optimized, training forms a self-referential loop. We theoretically formalize this as \emph{Self-Referential Propagation Drift}, a representation-supervision entanglement that can sharpen true boundaries or distort the feature space to satisfy its own targets. To systematically isolate these failure modes, we apply a protocolized single-variable ablation procedure spanning local EMA teacher decoupling, hard-background contrastive separation, and adaptive support geometry. On the SIRST3 dataset, GSACP-Final establishes a new ultra-low false-alarm operating regime, achieving a highly competitive $0.6674$ mIoU while demonstrating a $38\% relative reduction in false-positive artifacts ($\mathrm{Fa}$) compared with PAL. By systematically deconstructing the end-to-end paradigm, we map its performance boundaries and show that in-batch feature propagation provides a compact alternative for deployment scenarios where false-alarm suppression is paramount.

cs.CV

On the role of inertia and self-sustaining mechanism in two-dimensional elasto-inertial turbulence

Elasto-inertial turbulence (EIT) is primarily driven by polymer elasticity, yet the modulating role of fluid inertia is non-negligible and remains largely unexplored. To investigate the effect of inertia, we perform direct numerical simulations of two-dimensional EIT in channel flow over a wide range of Reynolds numbers ($Re$). We show that increasing inertia promotes both the enhancement of dynamic amplitudes and the wallward migration of core structures. Specifically, inertia intensifies the turbulent fluctuations, facilitates the fragmentation of large-scale structures, and amplifies statistical quantities such as the root-mean-square of velocity fluctuations and polymer extension. The peak location of nonlinear elastic shear stress follows a scaling law $y^+ \propto Re_τ^{1/2}$, closely resembling that of Reynolds shear stress in Newtonian turbulence, indicating a change of the momentum transfer mechanism. Meanwhile, the peak location of energy conversion between elastic and turbulent kinetic energies exhibits a $y^+ \propto Re_τ^{0.1}$ scaling law migration, remaining mostly confined to the near-wall region. Remarkably, despite the inertial modulation, the probability density functions (PDFs) of velocity and elastic stress fluctuations extracted at the energy-conversion peak collapse convincingly over the range of $Re$ investigated. This reveals a robust statistical self-similarity across a wide range of inertia magnitude. Furthermore, the PDFs of wall-normal velocity and elastic stress fluctuations exhibit pronounced exponential heavy tails.

physics.flu-dyn

Hierarchical Vision-Language Interaction for Facial Action Unit Detection

Facial Action Unit (AU) detection seeks to recognize subtle facial muscle activations as defined by the Facial Action Coding System (FACS). A primary challenge w.r.t AU detection is the effective learning of discriminative and generalizable AU representations under conditions of limited annotated data. To address this, we propose a Hierarchical Vision-language Interaction for AU Understanding (HiVA) method, which leverages textual AU descriptions as semantic priors to guide and enhance AU detection. Specifically, HiVA employs a large language model to generate diverse and contextually rich AU descriptions to strengthen language-based representation learning. To capture both fine-grained and holistic vision-language associations, HiVA introduces an AU-aware dynamic graph module that facilitates the learning of AU-specific visual representations. These features are further integrated within a hierarchical cross-modal attention architecture comprising two complementary mechanisms: Disentangled Dual Cross-Attention (DDCA), which establishes fine-grained, AU-specific interactions between visual and textual features, and Contextual Dual Cross-Attention (CDCA), which models global inter-AU dependencies. This collaborative, cross-modal learning paradigm enables HiVA to leverage multi-grained vision-based AU features in conjunction with refined language-based AU details, culminating in robust and semantically enriched AU detection capabilities. Extensive experiments show that HiVA consistently surpasses state-of-the-art approaches. Besides, qualitative analyses reveal that HiVA produces semantically meaningful activation patterns, highlighting its efficacy in learning robust and interpretable cross-modal correspondences for comprehensive facial behavior analysis.

cs.CV

Colorectal Cancer Tumor Grade Segmentation in Digital Histopathology Images: From Giga to Mini Challenge

Colorectal cancer (CRC) is the third most diagnosed cancer and the second leading cause of cancer-related death worldwide. Accurate histopathological grading of CRC is essential for prognosis and treatment planning but remains a subjective process prone to observer variability and limited by global shortages of trained pathologists. To promote automated and standardized solutions, we organized the ICIP Grand Challenge on Colorectal Cancer Tumor Grading and Segmentation using the publicly available METU CCTGS dataset. The dataset comprises 103 whole-slide images with expert pixel-level annotations for five tissue classes. Participants submitted segmentation masks via Codalab, evaluated using metrics such as macro F-score and mIoU. Among 39 participating teams, six outperformed the Swin Transformer baseline (62.92 F-score). This paper presents an overview of the challenge, dataset, and the top-performing methods

cs.CV

Anomalous Reynolds stress and dynamic mechanisms in two-dimensional elasto-inertial turbulence of viscoelastic channel flow

Elasto-inertial turbulence (EIT) has been demonstrated to be able to sustain in two-dimensional (2D) channel flow; however the systematic investigations on 2D EIT remain scare. This study addresses this gap by examining the statistical characteristics and dynamic mechanisms of 2D EIT, while exploring its similarities to and differences from three-dimensional (3D) EIT. We demonstrate that the influence of elasticity on the statistical properties of 2D EIT follows distinct trends compared to those observed in 3D EIT and drag-reducing turbulence (DRT). These differences can be attributed to variations in the underlying dynamical processes. As nonlinear elasticity increases, the dominant dynamic evolution in 3D flows involves the gradual suppression of inertial turbulence (IT). In contrast, 2D flows exhibit a progressive enhancement of EIT. More strikingly, we identify an anomalous Reynolds stress in 2D EIT that contributes negatively to flow resistance, a behavior opposite to that of IT. Quadrant analysis of velocity fluctuations reveals the predominance of motions in the first and third quadrants. These motions are closely associated with polymer sheet-like extension structures, which are inclined from the near-wall region toward the channel center along the streamwise direction. Finally, we present the dynamical budget of 2D EIT, which shows significant similarities to that of 3D EIT, thereby providing compelling evidence for the objective existence of the 2D nature of EIT.

physics.flu-dyn

Keep It Accurate and Robust: An Enhanced Nuclei Analysis Framework

Accurate segmentation and classification of nuclei in histology images is critical but challenging due to nuclei heterogeneity, staining variations, and tissue complexity. Existing methods often struggle with limited dataset variability, with patches extracted from similar whole slide images (WSI), making models prone to falling into local optima. Here we propose a new framework to address this limitation and enable robust nuclear analysis. Our method leverages dual-level ensemble modeling to overcome issues stemming from limited dataset variation. Intra-ensembling applies diverse transformations to individual samples, while inter-ensembling combines networks of different scales. We also introduce enhancements to the HoVer-Net architecture, including updated encoders, nested dense decoding and model regularization strategy. We achieve state-of-the-art results on public benchmarks, including 1st place for nuclear composition prediction and 3rd place for segmentation/classification in the 2022 Colon Nuclei Identification and Counting (CoNIC) Challenge. This success validates our approach for accurate histological nuclei analysis. Extensive experiments and ablation studies provide insights into optimal network design choices and training techniques. In conclusion, this work proposes an improved framework advancing the state-of-the-art in nuclei analysis. We release our code and models (https://github.com/WinnieLaugh/CONIC_Pathology_AI) to serve as a toolkit for the community.

eess.IV

Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases

Large Vision-Language Models (LVLMs) have received widespread attention for advancing the interpretable self-driving. Existing evaluations of LVLMs primarily focus on multi-faceted capabilities in natural circumstances, lacking automated and quantifiable assessment for self-driving, let alone the severe road corner cases. In this work, we propose CODA-LM, the very first benchmark for the automatic evaluation of LVLMs for self-driving corner cases. We adopt a hierarchical data structure and prompt powerful LVLMs to analyze complex driving scenes and generate high-quality pre-annotations for the human annotators, while for LVLM evaluation, we show that using the text-only large language models (LLMs) as judges reveals even better alignment with human preferences than the LVLM judges. Moreover, with our CODA-LM, we build CODA-VLM, a new driving LVLM surpassing all open-sourced counterparts on CODA-LM. Our CODA-VLM performs comparably with GPT-4V, even surpassing GPT-4V by +21.42% on the regional perception task. We hope CODA-LM can become the catalyst to promote interpretable self-driving empowered by LVLMs.

cs.CV

MMM-RS: A Multi-modal, Multi-GSD, Multi-scene Remote Sensing Dataset and Benchmark for Text-to-Image Generation

Recently, the diffusion-based generative paradigm has achieved impressive general image generation capabilities with text prompts due to its accurate distribution modeling and stable training process. However, generating diverse remote sensing (RS) images that are tremendously different from general images in terms of scale and perspective remains a formidable challenge due to the lack of a comprehensive remote sensing image generation dataset with various modalities, ground sample distances (GSD), and scenes. In this paper, we propose a Multi-modal, Multi-GSD, Multi-scene Remote Sensing (MMM-RS) dataset and benchmark for text-to-image generation in diverse remote sensing scenarios. Specifically, we first collect nine publicly available RS datasets and conduct standardization for all samples. To bridge RS images to textual semantic information, we utilize a large-scale pretrained vision-language model to automatically output text prompts and perform hand-crafted rectification, resulting in information-rich text-image pairs (including multi-modal images). In particular, we design some methods to obtain the images with different GSD and various environments (e.g., low-light, foggy) in a single sample. With extensive manual screening and refining annotations, we ultimately obtain a MMM-RS dataset that comprises approximately 2.1 million text-image pairs. Extensive experimental results verify that our proposed MMM-RS dataset allows off-the-shelf diffusion models to generate diverse RS images across various modalities, scenes, weather conditions, and GSD. The dataset is available at https://github.com/ljl5261/MMM-RS.

eess.IV

Hybrid-SORT: Weak Cues Matter for Online Multi-Object Tracking

Multi-Object Tracking (MOT) aims to detect and associate all desired objects across frames. Most methods accomplish the task by explicitly or implicitly leveraging strong cues (i.e., spatial and appearance information), which exhibit powerful instance-level discrimination. However, when object occlusion and clustering occur, spatial and appearance information will become ambiguous simultaneously due to the high overlap among objects. In this paper, we demonstrate this long-standing challenge in MOT can be efficiently and effectively resolved by incorporating weak cues to compensate for strong cues. Along with velocity direction, we introduce the confidence and height state as potential weak cues. With superior performance, our method still maintains Simple, Online and Real-Time (SORT) characteristics. Also, our method shows strong generalization for diverse trackers and scenarios in a plug-and-play and training-free manner. Significant and consistent improvements are observed when applying our method to 5 different representative trackers. Further, with both strong and weak cues, our method Hybrid-SORT achieves superior performance on diverse benchmarks, including MOT17, MOT20, and especially DanceTrack where interaction and severe occlusion frequently happen with complex motions. The code and models are available at https://github.com/ymzis69/HybridSORT.

cs.CV

CoNIC Challenge: Pushing the Frontiers of Nuclear Detection, Segmentation, Classification and Counting

Nuclear detection, segmentation and morphometric profiling are essential in helping us further understand the relationship between histology and patient outcome. To drive innovation in this area, we setup a community-wide challenge using the largest available dataset of its kind to assess nuclear segmentation and cellular composition. Our challenge, named CoNIC, stimulated the development of reproducible algorithms for cellular recognition with real-time result inspection on public leaderboards. We conducted an extensive post-challenge analysis based on the top-performing models using 1,658 whole-slide images of colon tissue. With around 700 million detected nuclei per model, associated features were used for dysplasia grading and survival analysis, where we demonstrated that the challenge's improvement over the previous state-of-the-art led to significant boosts in downstream performance. Our findings also suggest that eosinophils and neutrophils play an important role in the tumour microevironment. We release challenge models and WSI-level results to foster the development of further methods for biomarker discovery.

cs.CV

dual unet:a novel siamese network for change detection with cascade differential fusion

Change detection (CD) of remote sensing images is to detect the change region by analyzing the difference between two bitemporal images. It is extensively used in land resource planning, natural hazards monitoring and other fields. In our study, we propose a novel Siamese neural network for change detection task, namely Dual-UNet. In contrast to previous individually encoded the bitemporal images, we design an encoder differential-attention module to focus on the spatial difference relationships of pixels. In order to improve the generalization of networks, it computes the attention weights between any pixels between bitemporal images and uses them to engender more discriminating features. In order to improve the feature fusion and avoid gradient vanishing, multi-scale weighted variance map fusion strategy is proposed in the decoding stage. Experiments demonstrate that the proposed approach consistently outperforms the most advanced methods on popular seasonal change detection datasets.

cs.CV

A Dual-fusion Semantic Segmentation Framework With GAN For SAR Images

Deep learning based semantic segmentation is one of the popular methods in remote sensing image segmentation. In this paper, a network based on the widely used encoderdecoder architecture is proposed to accomplish the synthetic aperture radar (SAR) images segmentation. With the better representation capability of optical images, we propose to enrich SAR images with generated optical images via the generative adversative network (GAN) trained by numerous SAR and optical images. These optical images can be used as expansions of original SAR images, thus ensuring robust result of segmentation. Then the optical images generated by the GAN are stitched together with the corresponding real images. An attention module following the stitched data is used to strengthen the representation of the objects. Experiments indicate that our method is efficient compared to other commonly used methods

eess.IV

Hierarchical Feature-Aware Tracking

In this paper, we propose a hierarchical feature-aware tracking framework for efficient visual tracking. Recent years, ensembled trackers which combine multiple component trackers have achieved impressive performance. In ensembled trackers, the decision of results is usually a post-event process, i.e., tracking result for each tracker is first obtained and then the suitable one is selected according to result ensemble. In this paper, we propose a pre-event method. We construct an expert pool with each expert being one set of features. For each frame, several experts are first selected in the pool according to their past performance and then they are used to predict the object. The selection rate of each expert in the pool is then updated and tracking result is obtained according to result ensemble. We propose a novel pre-known expert-adaptive selection strategy. Since the process is more efficient, more experts can be constructed by fusing more types of features which leads to more robustness. Moreover, with the novel expert selection strategy, overfitting caused by fixed experts for each frame can be mitigated. Experiments on several public available datasets demonstrate the superiority of the proposed method and its state-of-the-art performance among ensembled trackers.

cs.CV

Atomistic Mechanisms of Nonlinear Graphene Growth on Ir Surface

As a two-dimensional material, graphene can be naturally obtained via epitaxial growth on a suitable substrate. Growth condition optimization usually requires an atomistic level understanding of the growth mechanism. In this article, we perform a mechanistic study about graphene growth on Ir(111) surface by combining first principles calculations and kinetic Monte Carlo (kMC) simulations. Small carbon clusters on the Ir surface are checked first. On terraces, arching chain configurations are favorable in energy and they are also of relatively high mobilities. At steps, some magic two-dimensional compact structures are identified, which show clear relevance to the nucleation process. Attachment of carbon species to a graphene edge is then studied. Due to the effect of substrate, at some edge sites, atomic carbon attachment becomes thermodynamically unfavorable. Graphene growth at these difficult sites has to proceed via cluster attachment, which is the growth rate determining step. Based on such an inhomogeneous growth picture, kMC simulations are made possible by successfully separating different timescales, and they well reproduce the experimentally observed nonlinear kinetics. Different growth rates and nonlinear behaviors are predicted for different graphene orientations, which is consistent with available experimental results. Importantly, as a phenomenon originated from lattice mismatch, inhomogeneity revealed in this case is expected to be quite universal and it should also make important roles in many other hetero-epitaxial systems.

cond-mat.mtrl-sci

First-Principles Thermodynamics of Graphene Growth on Cu Surface

Chemical vapor deposition (CVD) is an important method to synthesis grapheme on a substract. Recently, Cu becomes the most popular CVD substrate for graphene growth. Here, we combine electronic structure calculation, molecular dynamics simulation, and thermodynamics analysis to study the graphene growth process on Cu surface. As a fundamentally important but previously overlooked fact, we find that carbon atoms are thermodynamically unfavorable on Cu surface under typical experimental conditions. The active species for graphene growth should thus mainly be CHx instead of atomic carbon. Based on this new picture, the nucleation behavior can be understood, which explains many experimental observations and also provides us a guide to improve graphene sample quality.

cond-mat.mtrl-sci

Chemical Vapor Deposition Growth of Graphene using Other Hydrocarbon Sources

Graphene has attracted a lot of research interests due to its exotic properties and a wide spectrum of potential applications. Chemical vapor deposition (CVD) from gaseous hydrocarbon sources has shown great promises for large-scale graphene growth. However, high growth temperature, typically 1000°C, is required for such growth. Here we demonstrate a revised CVD route to grow graphene on Cu foils at low temperature, adopting solid and liquid hydrocarbon feedstocks. For solid PMMA and polystyrene precursors, centimeter-scale monolayer graphene films are synthesized at a growth temperature down to 400°C. When benzene is used as the hydrocarbon source, monolayer graphene flakes with excellent quality are achieved at a growth temperature as low as 300°C. The successful low-temperature growth can be qualitatively understood from the first principles calculations. Our work might pave a way to undemanding route for economical and convenient graphene growth.

cond-mat.mtrl-sci

Coalescence of Carbon Atoms on Cu (111) Surface: Emergence of a Stable Bridging-Metal Structure Motif

By combining first principles transition state location and molecular dynamics simulation, we unambiguously identify a carbon atom approaching induced bridging metal structure formation on Cu (111) surface, which strongly modify the carbon atom coalescence dynamics. The emergence of this new structural motif turns out to be a result of the subtle balance between Cu-C and Cu-Cu interactions. Based on this picture, a simple theoretical model is proposed, which describes a variety of surface chemistries very well.

cond-mat.mtrl-sci

Oxidation States of Graphene: Insights from Computational Spectroscopy

When it is oxidized, graphite can be easily exfoliated forming graphene oxide (GO). GO is a critical intermediate for massive production of graphene, and it is also an important material with various application potentials. With many different oxidation species randomly distributed on the basal plane, GO has a complicated nonstoichiometric atomic structure that is still not well understood in spite of of intensive studies involving many experimental techniques. Controversies often exist in experimental data interpretation. We report here a first principles study on binding energy of carbon 1s orbital in GO. The calculated results can be well used to interpret experimental X-ray photoelectron spectroscopy (XPS) data and provide a unified spectral assignment. Based on the first principles understanding of XPS, a GO structure model containing new oxidation species epoxy pair and epoxy-hydroxy pair is proposed. Our results demonstrate that first principles computational spectroscopy provides a powerful means to investigate GO structure.

cond-mat.mtrl-sci