SearcharxivSearch

arXiv subjects

Zhisong Wang

Publications and source records attributed to Zhisong Wang.

15 recordsLinked to original sources

Volumetric Radiology AI in the Era of Multimodal Large Language Models

Advances in multimodal large language models (MLLMs) are extending radiological artificial intelligence (AI) beyond task-specific image analysis toward multimodal understanding and reasoning. Volumetric radiology, however, presents a fundamental representational mismatch: clinical interpretation often requires full-volume spatial context and acquisition-dependent quantitative information, whereas current MLLMs are commonly conditioned on selected two-dimensional (2D) images, compressed visual representations, or report-derived text. Reliable volumetric radiology AI therefore requires representations that preserve task-relevant three-dimensional (3D) information and systems that can access, verify, and integrate this information across clinical workflows. In this Review, we examine more than 200 publications through July 2026. We organize the literature around volumetric representation and multimodal understanding at the model level, agentic orchestration at the system level, and their links to clinical applications and evaluation. We review volumetric foundation models, language alignment and compression strategies, and agentic systems that extend MLLMs through planning, tools, memory, and workflow interaction. We distinguish settings in which selected 2D views or report-mediated reasoning may suffice from those that warrant native volumetric modeling. We also introduce a Claim-Design-Validation framework to assess whether technical, workflow, and clinical claims are matched by appropriate design and validation. Across the literature, native volumetric modeling and agentic capabilities depend on the spatial, quantitative, contextual, and workflow requirements of the intended task. Clinical credibility requires faithful volumetric representation, traceable system behavior, claim-aligned validation, and clearly defined human oversight in realistic workflows.

cs.AI

InViC: Intent-aware Visual Cues for Medical Visual Question Answering

Medical visual question answering (Med-VQA) aims to answer clinically relevant questions grounded in medical images. However, existing multimodal large language models (MLLMs) often exhibit shortcut answering, producing plausible responses by exploiting language priors or dataset biases while insufficiently attending to visual evidence. This behavior undermines clinical reliability, especially when subtle imaging findings are decisive. We propose a lightweight plug-in framework, termed Intent-aware Visual Cues (InViC), to explicitly enhance image-based answer generation in medical VQA. InViC introduces a Cue Tokens Extraction (CTE) module that distills dense visual tokens into a compact set of K question-conditioned cue tokens, which serve as structured visual intermediaries injected into the LLM decoder to promote intent-aligned visual evidence. To discourage bypassing of visual information, we further design a two-stage fine-tuning strategy with a cue-bottleneck attention mask. In Stage I, we employ an attention mask to block the LLM's direct view of raw visual features, thereby funneling all visual evidence through the cue pathway. In Stage II, standard causal attention is restored to train the LLM to jointly exploit the visual and cue tokens. We evaluate InViC on three public Med-VQA benchmarks (VQA-RAD, SLAKE, and ImageCLEF VQA-Med 2019) across multiple representative MLLMs. InViC consistently improves over zero-shot inference and standard LoRA fine-tuning, demonstrating that intent-aware visual cues with bottlenecked training is a practical and effective strategy for improving trustworthy Med-VQA.

cs.CV

Topology-Guided Biomechanical Profiling: A White-Box Framework for Opportunistic Screening of Spinal Instability on Routine CT

Routine oncologic computed tomography (CT) presents an ideal opportunity for screening spinal instability, yet prophylactic stabilization windows are frequently missed due to the complex geometric reasoning required by the Spinal Instability Neoplastic Score (SINS). Automating SINS is fundamentally hindered by metastatic osteolysis, which induces topological ambiguity that confounds standard segmentation and black-box AI. We propose Topology-Guided Biomechanical Profiling (TGBP), an auditable white-box framework decoupling anatomical perception from structural reasoning. TGBP anchors SINS assessment on two deterministic geometric innovations: (i) canal-referenced partitioning to resolve posterolateral boundary ambiguity, and (ii) context-aware morphometric normalization via covariance-based oriented bounding boxes (OBB) to quantify vertebral collapse. Integrated with auxiliary radiomic and large language model (LLM) modules, TGBP provides an end-to-end, interpretable SINS evaluation. Validated on a multi-center, multi-cancer cohort ($N=482$), TGBP achieved 90.2\% accuracy in 3-tier stability triage. In a blinded reader study ($N=30$), TGBP significantly outperformed medical oncologists on complex structural features ($κ=0.857$ vs.\ $0.570$) and prevented compounding errors in Total Score estimation ($κ=0.625$ vs.\ $0.207$), democratizing expert-level opportunistic screening.

q-bio.QM

From Few to More: Scribble-based Medical Image Segmentation via Masked Context Modeling and Continuous Pseudo Labels

Scribble-based weakly supervised segmentation methods have shown promising results in medical image segmentation, significantly reducing annotation costs. However, existing approaches often rely on auxiliary tasks to enforce semantic consistency and use hard pseudo labels for supervision, overlooking the unique challenges faced by models trained with sparse annotations. These models must predict pixel-wise segmentation maps from limited data, making it crucial to handle varying levels of annotation richness effectively. In this paper, we propose MaCo, a weakly supervised model designed for medical image segmentation, based on the principle of "from few to more." MaCo leverages Masked Context Modeling (MCM) and Continuous Pseudo Labels (CPL). MCM employs an attention-based masking strategy to perturb the input image, ensuring that the model's predictions align with those of the original image. CPL converts scribble annotations into continuous pixel-wise labels by applying an exponential decay function to distance maps, producing confidence maps that represent the likelihood of each pixel belonging to a specific category, rather than relying on hard pseudo labels. We evaluate MaCo on three public datasets, comparing it with other weakly supervised methods. Our results show that MaCo outperforms competing methods across all datasets, establishing a new record in weakly supervised medical image segmentation.

cs.CV

Pre-training Everywhere: Parameter-Efficient Fine-Tuning for Medical Image Analysis via Target Parameter Pre-training

Parameter-efficient fine-tuning (PEFT) techniques have emerged to address overfitting and high computational costs associated with fully fine-tuning in self-supervised learning. Mainstream PEFT methods add a few trainable parameters while keeping the pre-trained backbone parameters fixed. These methods achieve comparative, and often superior, performance to fully fine-tuning, demonstrating the powerful representation ability of the pre-trained backbone. Despite this success, these methods typically ignore the initialization of the new parameters, often relying solely on random initialization. We argue that if pre-training is significantly beneficial, it should be applied to all parameters requiring representational capacity. Motivated by this, we propose Target Parameter Pre-training (TPP), a simple yet effective fine-tuning framework. TPP pre-trains target parameters, i.e., the new parameters introduced during fine-tuning, in an additional stage before PEFT. During this stage, the pre-trained backbone parameters are frozen, and only the new parameters are trainable. A defined pretext task encourages the new parameters to learn specific representations of downstream data. Subsequently, when PEFT is employed, the pre-trained new parameters are loaded to enhance fine-tuning efficiency. The proposed TPP framework is versatile, allowing integration with various pre-trained backbones, pretext tasks, and PEFT methods. We evaluated the fine-tuning performance of our method on seven public datasets, covering four modalities and two task types. The results demonstrate that TPP can be easily integrated into existing PEFT methods, significantly improving performance.

cs.CV

Enjoying Information Dividend: Gaze Track-based Medical Weakly Supervised Segmentation

Weakly supervised semantic segmentation (WSSS) in medical imaging struggles with effectively using sparse annotations. One promising direction for WSSS leverages gaze annotations, captured via eye trackers that record regions of interest during diagnostic procedures. However, existing gaze-based methods, such as GazeMedSeg, do not fully exploit the rich information embedded in gaze data. In this paper, we propose GradTrack, a framework that utilizes physicians' gaze track, including fixation points, durations, and temporal order, to enhance WSSS performance. GradTrack comprises two key components: Gaze Track Map Generation and Track Attention, which collaboratively enable progressive feature refinement through multi-level gaze supervision during the decoding process. Experiments on the Kvasir-SEG and NCI-ISBI datasets demonstrate that GradTrack consistently outperforms existing gaze-based methods, achieving Dice score improvements of 3.21\% and 2.61\%, respectively. Moreover, GradTrack significantly narrows the performance gap with fully supervised models such as nnUNet.

cs.CV

CoSAM: Self-Correcting SAM for Domain Generalization in 2D Medical Image Segmentation

Medical images often exhibit distribution shifts due to variations in imaging protocols and scanners across different medical centers. Domain Generalization (DG) methods aim to train models on source domains that can generalize to unseen target domains. Recently, the segment anything model (SAM) has demonstrated strong generalization capabilities due to its prompt-based design, and has gained significant attention in image segmentation tasks. Existing SAM-based approaches attempt to address the need for manual prompts by introducing prompt generators that automatically generate these prompts. However, we argue that auto-generated prompts may not be sufficiently accurate under distribution shifts, potentially leading to incorrect predictions that still require manual verification and correction by clinicians. To address this challenge, we propose a method for 2D medical image segmentation called Self-Correcting SAM (CoSAM). Our approach begins by generating coarse masks using SAM in a prompt-free manner, providing prior prompts for the subsequent stages, and eliminating the need for prompt generators. To automatically refine these coarse masks, we introduce a generalized error decoder that simulates the correction process typically performed by clinicians. Furthermore, we generate diverse prompts as feedback based on the corrected masks, which are used to iteratively refine the predictions within a self-correcting loop, enhancing the generalization performance of our model. Extensive experiments on two medical image segmentation benchmarks across multiple scenarios demonstrate the superiority of CoSAM over state-of-the-art SAM-based methods.

cs.CV

Directional fidelity of nanoscale motors and particles is limited by the second law of thermodynamics via a universal equality

Directional motion of nanoscale motors and driven particles in an isothermal environment costs a finite amount of energy despite zero work as decreed by the 2nd law, but quantifying this general limit remains difficult. Here we derive a universal equality linking directional fidelity of an arbitrary nanoscale object to the least possible energy driving it. The fidelity-energy equality depends on the environmental temperature alone; any lower energy would violate the 2nd law in a thought experiment. Real experimental proof for the equality comes from force-induced motion of biological nanomotors by three independent groups for translational as well as rotational motion. Interestingly, the natural self-propelled motion of a biological nanomotor (F1-ATPase) known to have nearly 100% energy efficiency evidently pays the 2nd-law decreed least energy cost for direction production.

cond-mat.stat-mech

Role of directional fidelity in multiple extreme performance of F1-ATPase motor

Quantitative understanding of the best possible performance of nanomotors allowed by physical laws pertains to study of nanomotors from biology as well as nanotechnology. Biological nanomotor F1-ATPase is the best available model system as it is the only nanomotor known for extreme energy conversion near the limit of energy conservation. Using a unified theoretical framework centred on a concept called directional fidelity, we analyze recent experiments in which F1-motor's performance was measured for controlled chemical potentials, and expose from the experiments quantitative evidence for the motor's multiple extreme performance in directional fidelity, speed and catalytic capability close to physical limits. Specifically, the motor nearly exhausts available energy from the fuel to retain the highest possible directional fidelity for arbitrary load, encompassing the motor's extreme energy conversion and beyond. The theory-experiment comparison implies a tight chemomechanical coupling up to stalemate as futile steps occur but unlikely involve fuel consumption. The F1-motor data also helps clarify relation between directional fidelity and experimentally measured stepping ratio.

physics.bio-ph

Bipedal nanowalker by pure physical mechanisms

Artificial nanowalkers are inspired by biomolecular counterparts from living cells, but remain far from comparable to the latter in design principles. The walkers reported to date mostly rely on chemical mechanisms to gain a direction; they all produce chemical wastes. Here we report a light-powered DNA bipedal walker based on a design principle derived from cellular walkers. The walker has two identical feet and the track has equal binding sites; yet the walker gains a direction by pure physical mechanisms that autonomously amplify an intra-site asymmetry into a ratchet effect. The nanowalker is free of any chemical waste. It has a distinct thermodynamic feature that it possesses the same equilibrium before and after operation, but generates a truly non-equilibrium distribution during operation. The demonstrated design principle exploits mechanical effects and is adaptable for use in other nanomachines.

physics.bio-ph

Synergic mechanism and fabrication target for bipedal nanomotors

Inspired by dimeric motor proteins capable of undergoing transportation in living cells, significant efforts have been expended to the fabrication of track-walking nanomotors possessing two foot-like components that each can bind or detach from an array of anchorage groups on the track in response to local events of reagent consumption. The central problem in fabricating bipedal nanomotors is how the motor as a whole can gain the synergic capacity of directional track-walking, given the fact that each pedal component alone often is incapable of any directional drift. Implemented bipedal motors to date solve this thermodynamically intricate problem by an intuitive strategy that requires a hetero-pedal motor, multiple anchorage species for the track, and multiple reagent species for motor operation. Here we presented a detailed molecular mechanism by which motor-level directionality arises from a homo-pedal motor along a minimally heterogeneous track. Optimally, the operation may be reduced to a random supply of a single species of reagents to allow the motor's autonomous functioning. The mechanism suggests a distinct class of fabrication targets of drastically reduced system requirements. Intriguingly, a defective form of the mechanism falls into the realm of the well known Brownian motor mechanism, yet distinct features emerge from the normal working of the mechanism.

cond-mat.mes-hall

Modeling Motility of the Kinesin Dimer from Molecular Properties of Individual Monomers

Conventional kinesin is a homodimeric motor protein that unidirectionally transports organelles along filamentous microtubule (MT) by hydrolyzing ATP molecules. This study shows that the load modulations of ATP turnover and head diffusion are both essential in determining the performance of the dimer under loads. It is found that the consecutive run length of the dimer critically depends upon a few pathways, leading to the detachment of individual heads from MT. Modifying rates for these detachment pathways changes the run length but not the velocity of the dimer, consistent with mutant experiments. The run length may increase with or without the ATP concentration, depending upon a single rate for pure mechanical detachment. This finding provides an explanation to a previous controversy concerning ATP dependence of the run length, and related quantitative predictions of this study can be tested by a future experiment. This study also finds that the experimental observations for assisting loads can be quantitatively explained by load-biased head diffusion. We thus conclude that the dimer motility under resisting as well as assisting loads is governed by essentially the same mechanisms.

physics.bio-ph

Influence of Local and Residual Structures on the Scaling Behavior and Dimensions of Unfolded Proteins

Although recent spectroscopic studies of chemically denatured proteins hint at significant nonrandom residual structure, the results of extensive small angle X-ray scattering studies suggest random coil behavior, calling for a coherent understanding of these seemingly contradicting observations. Here, we report the results of a Monte Carlo study of the effects of two types of local structures, a helix and Polyproline II (PPII) helix, on the dimensions of random coil polyalanine chains viewed as a model of highly denatured proteins. With an alpha helix content of 20%, corresponding to the Ramachandran probability of being in the helical basin, experimentally observed radii of gyration are recovered. Experimental radii are similarly recovered at an a helix content of 87%, providing an explanation for the previously puzzling experimental finding that the dimensions of the highly helical methanol-induced unfolded state are experimentally indistinguishable from those of the helix-poor urea-unfolded state. In contrast, the radius of gyration increases monotonically with increasing PPII content, and is always more expanded than the dimensions observed experimentally. These results suggest that PPII is unlikely the sole, dominant preferred conformation for unfolded proteins.

physics.bio-ph

Nonadiabatic simulation study of photoisomerization of azobenzene: Detailed mechanism and load-resisting capacity

Nonadiabatic dynamical simulations were carried out to study cis-to-trans isomerization of azobenzene under laser irradiation and/or external mechanical loads. We used a semiclassical electron-radiation-ion dynamics method that is able to describe the coevolution of the structural dynamics and the underlying electronic dynamics in a real-time manner. It is found that azobenzene photoisomerization occurs predominantly by an out-of-plane rotation mechanism even under a nontrivial resisting force of several tens of piconewtons. We have repeated the simulations systematically for a broad range of parameters for laser pulses, but could not find any photoisomerization event by a previously suggested in-plane inversion mechanism. The simulations found that the photoisomerization process can be held back by an external resisting force of 90 - 200 pN depending on the frequency and intensity of the lasers. This study also found that a pure mechanical isomerization is possible from the cis state if the azobenzene molecule is stretched by an external force of 1250 -1650 pN. Remarkably, the mechanical isomerization first proceeds through a mechanically activated inversion, and then is diverted to an ultrafast downhill rotation that accomplishes the isomerization. Implications of these findings to azobenzene-based nanomechanical devices are discussed.

physics.chem-ph

Kinesin Is an Evolutionarily Fine-Tuned Molecular Ratchet-and-Pawl Device of Decisively Locked Direction

Conventional kinesin is a dimeric motor protein that transports membranous organelles toward the plus-end of microtubules (MTs). Individual kinesin dimers show steadfast directionality and hundreds of consecutive steps, yetthe detailed physical mechanism remains unclear. Here we compute free energies for the entire dimer-MT system for all possible interacting configurations by taking full account of molecular details. Employing merely first principles and several measured binding and barrier energies, the system-level analysis reveals insurmountable energy gaps between configurations, asymmetric ground state caused by mechanically lifted configurational degeneracy, and forbidden transitions ensuring coordination between both motor domains for alternating catalysis. This wealth of physical effects converts a kinesin dimer into a molecular ratchet-and-pawl device, which determinedly locks the dimer's movement into the MT plus-end and ensures consecutive steps in hand-over-hand gait.Under a certain range of extreme loads, however, the ratchet-and-pawl device becomes defective but not entirely abolished to allow consecutive back-steps. This study yielded quantitative evidence that kinesin's multiple molecular properties have been evolutionarily adapted to fine-tune the ratchet-and-pawl device so as to ensure the motor's distinguished performance.

physics.bio-ph