SearcharxivSearch

arXiv subjects

Zhidong Yang

Publications and source records attributed to Zhidong Yang.

11 recordsLinked to original sources

Dual-Pathway Circuits of Object Hallucination in Vision-Language Models

Vision-language models (VLMs) have demonstrated remarkable capabilities in bridging visual perception and natural language understanding, enabling a wide range of multimodal reasoning tasks. However, they often produce object hallucinations, describing content absent from the input image, which limits their reliability and interpretability. To address this limitation, we propose Dual-Pathway Circuit Analysis, a framework that identifies and characterizes hallucination-related circuits in VLMs for mechanistic understanding and causal probing. We first apply activation patching across five architecturally diverse VLMs to identify a visual grounding pathway that supports correct predictions and a hallucination pathway that drives erroneous outputs. We then introduce Conditional Pathway Analysis (CPA) to characterize pathway-level interactions, revealing that grounding components remain strongly redundant in both correct and hallucinating samples but undergo a consistent polarity flip, shifting from supporting the ground truth on correct samples to aligning with the hallucinated answer on erroneous ones. We further perform targeted suppression of hallucination-pathway components, showing that scaling these components reduces object hallucination by up to 76% with minimal accuracy cost, and validate that the same circuit selectively transfers to relational but not attribute hallucination. Evaluations on POPE-adversarial and AMBER show that the identified circuits are consistent across architectures, support causal intervention, and transfer selectively across hallucination types.

cs.CV

WISCA: A Lightweight Model Transition Method to Improve LLM Training via Weight Scaling

Transformer architecture gradually dominates the LLM field. Recent advances in training optimization for Transformer-based large language models (LLMs) primarily focus on architectural modifications or optimizer adjustments. However, these approaches lack systematic optimization of weight patterns during training. Weight pattern refers to the distribution and relative magnitudes of weight parameters in a neural network. To address this issue, we propose a Weight Scaling method called WISCA to enhance training efficiency and model quality by strategically improving neural network weight patterns without changing network structures. By rescaling weights while preserving model outputs, WISCA indirectly optimizes the model's training trajectory. Experiments demonstrate that WISCA significantly improves convergence quality (measured by generalization capability and loss reduction), particularly in LLMs with Grouped Query Attention (GQA) architectures and LoRA fine-tuning tasks. Empirical results show 5.6% average improvement on zero-shot validation tasks and 2.12% average reduction in training perplexity across multiple architectures.

cs.LG

QCAgent: An agentic framework for quality-controllable pathology report generation from whole slide image

Recent methods for pathology report generation from whole-slide image (WSI) are capable of producing slide-level diagnostic descriptions but fail to ground fine-grained statements in localized visual evidence. Furthermore, they lack control over which diagnostic details to include and how to verify them. Inspired by emerging agentic analysis paradigms and the diagnostic workflow of pathologists,who selectively examine multiple fields of view, we propose QCAgent, an agentic framework for quality-controllable WSI report generation. The core innovations of this framework are as follows: (i) it incorporates a customized critique mechanism guided by a user-defined checklist specifying required diagnostic details and constraints; (ii) it re-identifies informative regions in the WSI based on the critique feedback and text-patch semantic retrieval, a process that iteratively enriches and reconciles the report. Experiments demonstrate that by making report requirements explicitly prompt-defined, constraint-aware, and verifiable through evidence-grounded refinement, QCAgent enables controllable generation of clinically meaningful and high-coverage pathology reports from WSI.

cs.CV

AFA-LoRA: Enabling Non-Linear Adaptations in LoRA with Activation Function Annealing

Low-Rank Adaptation (LoRA) is a widely adopted parameter-efficient fine-tuning (PEFT) method. However, its linear adaptation process limits its expressive power. This means there is a gap between the expressive power of linear training and non-linear training. To bridge this gap, we propose AFA-LoRA, a novel training strategy that brings non-linear expressivity to LoRA while maintaining its seamless mergeability. Our key innovation is an annealed activation function that transitions from a non-linear to a linear transformation during training, allowing the adapter to initially adopt stronger representational capabilities before converging to a mergeable linear form. We implement our method on supervised fine-tuning, reinforcement learning, and speculative decoding. The results show that AFA-LoRA reduces the performance gap between LoRA and full-parameter training. This work enables a more powerful and practical paradigm of parameter-efficient adaptation.

cs.LG

Fusion of Multi-scale Heterogeneous Pathology Foundation Models for Whole Slide Image Analysis

Whole slide image (WSI) analysis has emerged as an increasingly essential technique in computational pathology. Recent advances in the pathology foundation models (FMs) have demonstrated significant advantages in deriving meaningful patch-level or slide-level multi-scale features from WSIs. However, current pathology FMs have exhibited substantial heterogeneity caused by diverse private training datasets and different network architectures. This heterogeneity introduces performance variability when we utilize the features from different FMs in the downstream tasks. To fully explore the advantages of multiple FMs effectively, in this work, we propose a novel framework for the fusion of multi-scale heterogeneous pathology FMs, called FuseCPath, yielding a model with a superior ensemble performance. The main contributions of our framework can be summarized as follows: (i) To guarantee the representativeness of the training patches, we propose a multi-view clustering-based method to filter out the discriminative patches via multiple FMs' embeddings. (ii) To effectively fuse the patch-level FMs, we devise a cluster-level re-embedding strategy to online capture patch-level local features. (iii) To effectively fuse the slide-level FMs, we devise a collaborative distillation strategy to explore the connections between slide-level FMs. Extensive experiments demonstrate that the proposed FuseCPath achieves state-of-the-art performance across multiple tasks on diverse datasets.

cs.CV

SED-MVS: Segmentation-Driven and Edge-Aligned Deformation Multi-View Stereo with Depth Restoration and Occlusion Constraint

Recently, patch-deformation methods have exhibited significant effectiveness in multi-view stereo owing to the deformable and expandable patches in reconstructing textureless areas. However, such methods primarily emphasize broadening the receptive field in textureless areas, while neglecting deformation instability caused by easily overlooked edge-skipping, potentially leading to matching distortions. To address this, we propose SED-MVS, which adopts panoptic segmentation and multi-trajectory diffusion strategy for segmentation-driven and edge-aligned patch deformation. Specifically, to prevent unanticipated edge-skipping, we first employ SAM2 for panoptic segmentation as depth-edge guidance to guide patch deformation, followed by multi-trajectory diffusion strategy to ensure patches are comprehensively aligned with depth edges. Moreover, to avoid potential inaccuracy of random initialization, we combine both sparse points from LoFTR and monocular depth map from DepthAnything V2 to restore reliable and realistic depth map for initialization and supervised guidance. Finally, we integrate segmentation image with monocular depth map to exploit inter-instance occlusion relationship, then further regard them as occlusion map to implement two distinct edge constraint, thereby facilitating occlusion-aware patch deformation. Extensive results on ETH3D, Tanks & Temples, BlendedMVS and Strecha datasets validate the state-of-the-art performance and robust generalization capability of our proposed method.

cs.CV

Baryon density dependence of viscosities of the quark-gluon plasma at hadronization

The $ϕ$ meson and $Ω$ baryon provide unique probes of the properties of the quark-gluon plasma (QGP) at hadronization in relativistic heavy-ion collisions. Using the quark recombination model with the quark phase-space information parameterized in a viscous blastwave, we perform Bayesian inference of the shear and bulk viscosities of the QGP at hadronization with a temperature of $T\sim 160$ MeV by analyzing the $ϕ$ and $Ω$ data in Au+Au collisions at $\sqrt{s_{\rm NN}}=$ 19.6-200 GeV and Pb+Pb collisions at $\sqrt{s_{\rm NN}}=$ 2.76 TeV, corresponding to a baryon chemical potential variation from $μ_B\approx 0$ (at $\sqrt{s_{\rm NN}}= 2.76$ TeV) to $200$ MeV (at $\sqrt{s_{\rm NN}}= 19.6$ GeV). We find that the shear viscosity to enthalpy ratio $ηT/(ε+P)$ of the QGP at hadronization decreases as $μ_B$ increases, with $ηT/(ε+P)\approx 0.18$ at $μ_B=0$ and $ηT/(ε+P)\approx 0.08$ at $μ_B=200$ MeV, while the corresponding specific bulk viscosity is essentially constant with $ζT/(ε+ P)=0.02\sim 0.04$ for $μ_B<200$ MeV. Our results suggest that the QGP at hadronization ($T\sim 160$ MeV) with finite baryon density is more close to perfect fluid than that with zero baryon density.

nucl-th

Parameterizing Smooth Viscous Fluid Dynamics With a Viscous Blast Wave

Blast wave fits are widely used in high energy nuclear collisions to capture essential features of global properties of systems near kinetic equilibrium. They usually provide temperature fields and collective velocity fields on a given hypersurface. We systematically compare blast wave fits of fluid dynamic simulations for Au+Au collisions at $\sqrt{s_{NN}}=200$ GeV and Pb+Pb collisions at $\sqrt{s_{NN}}=2.76$ TeV with the original simulations. In particular, we investigate how faithful the viscous blast wave introduced in \cite{Yang:2022yxa} can reproduce the given temperature and specific shear viscosity fixed at freeze-out of a viscous fluid dynamic calculation, if the final spectrum and elliptic flow of several particle species are fitted. We find that viscous blast wave fits describe fluid dynamic pseudodata rather well and reproduce the specific shear viscosities to good accuracy. However, extracted temperatures tend to be underpredicted, especially for peripheral collisions. We investigate possible reasons for these deviations. We establish maps from true to fitted values. These maps can be used to improve raw fit results from viscous blast wave fits. Although our work is limited to two specific, albeit important, parameters and two collision systems, the same procedure can be easily generalized to other parameters and collision systems.

nucl-th

Bayesian Inference of the Specific Shear and Bulk Viscosities of the Quark-Gluon Plasma at Crossover from $ϕ$ and $Ω$ Observables

Due to their weak final state interactions, the $ϕ$ meson and $Ω$ baryon provide unique probes of the properties of the quark-gluon plasma (QGP) formed in relativistic heavy-ion collisions. Using the quark recombination model with the quark phase-space information parameterized in a viscous blastwave, we study the transverse-momentum spectra and elliptic flows of $ϕ$ and $Ω$ in Au+Au collisions at $\sqrt{s_{\rm NN}} = 200$ GeV and Pb+Pb collisions at $\sqrt{s_{\rm NN}} = 2.76$ TeV. The viscous blastwave includes non-equilibrium deformations of thermal distributions due to shear and bulk stresses and thus carries information on the specific shear viscosity $η/s$ and the specific bulk viscosity $ζ/s$ of the QGP. We perform a model-to-data comparison with Bayesian inference and simultaneously obtain $η/s=(2.08^{+1.10}_{-1.09})/4π$ and $ζ/s= 0.06^{+0.04}_{-0.04}$ at $90\%$ C.L. for the baryon-free QGP at crossover temperature of about $160$ MeV. Our work provides a novel approach to simultaneously determine the $η/s$ and $ζ/s$ of the QGP at hadronization.

nucl-th

Extraction of the Specific Shear Viscosity of Hot Hadron Gas

We extract the specific shear viscosity $η/s$ of nuclear matter for various temperatures and chemical potentials in the hadronic phase using data taken in high energy nuclear collisions. We use a blastwave parameterization of the final state of nuclear collisions, including non-equilibrium deformations of particle distributions due to shear stress in the Navier-Stokes approximation. We fit spectra and elliptic flow of identified hadrons for a variety of collision energies and impact parameters at the Relativistic Heavy Ion Collider (RHIC) and the Large Hadron Collider (LHC). The systems analyzed cover a temperature range from about 110 to 140 MeV and vary in their chemical potentials for stable hadrons. We attempt to assign meaningful systematic uncertainties to our results. This work is complementary to efforts using viscous fluid dynamics to extract the specific shear viscosity of quark gluon plasma at higher temperatures. We put our work in context with existing theoretical calculations of the specific shear viscosity.

nucl-th

A Blast Wave Model With Viscous Corrections

Hadronic observables in the final stage of heavy ion collision can be described well by fluid dynamics or blast wave parameterizations. We improve existing blast wave models by adding shear viscous corrections to the particle distributions in the Navier-Stokes approximation. The specific shear viscosity $η/s$ of a hadron gas at the freeze-out temperature is a new parameter in this model. We extract the blast wave parameters with viscous corrections from experimental data which leads to constraints on the specific shear viscosity at kinetic freeze-out. Preliminary results show $η/s$ is rather small.

nucl-th