SearcharxivSearch

arXiv subjects

Zhonghao Zhang

Publications and source records attributed to Zhonghao Zhang.

11 recordsLinked to original sources

Transferable Tool-Tissue Contact Detection from Stereo Depth in Robot-Assisted Surgery

Reliable tool--tissue contact detection can support interaction-aware control and downstream force estimation in robot-assisted surgery. Most existing methods learn a contact classifier from RGB appearance, which is hard to generalize. In this work, we use the depth image generated from a stereo pair to give more information about tool--tissue contact. For each depth frame, we localize a spatially supported minimum-distance patch around the tool boundary and reduce it to a single scalar, $-\log_{10}|d|$; this signal rises and falls in step with ground-truth contact. We formalize this observation with a fully supervised two-state hidden Markov model. We fit this model as a six-fold leave-one-session-out (LOSO) ensemble on six palpation sessions against a single silicone cup-like phantom, with the decision threshold selected from the pooled out-of-fold predictions. It is evaluated on four held-out sessions of three categories: 1. same task on same phantom; 2. same task on different phantom; 3. different task on different phantom. This model reaches held-out macro F1 $0.927$ and AUPRC $0.980$. We further compare against a reproduction of an RGB-based contact classifier from prior work. This RGB-based model achieves high performance on the first category (F1 $0.965$), but substantially lower performance on the other two, resulting in macro F1 $0.320$ across all four sessions. These results indicate that the tool--tissue distance is a strong, transferable cue for contact detection in robot-assisted surgery.

cs.RO

Ecosystem service demand relationship and trade-off patterns in urban parks across China

Understanding public demand for urban ecosystem services (ES) is crucial for effective green space management, yet the intricate relationships and potential trade-offs among these diverse demands remain poorly understood. Previous studies have yielded inconsistent findings, often limited by small samples or reliance on indirect proxies. Here, we provide the first national-scale, direct assessment of the relationship among demands for nine urban park ES using a survey dataset comprising 20,075 responses across China and a point-allotment experiment that directly quantifies the trade-off patterns among service demands. We found particularly strong preferences among urban residents in China for air purification and recreation services, at the expense of other services. These preferences were further reflected in three distinct demand bundles: air purification-dominated, recreation-dominated, and balanced demands, each delineating a typical group of people with distinct representative characteristics. Socio-economic and environmental factors, such as age, environmental interest, and mean annual precipitation, significantly influence the trade-off intensity among service demands. Our study pioneers the direct, quantitative analysis of relationships among ecosystem service demands, and the results underscore the need for tailored urban park designs that address diverse service demands to sustainably enhance the quality of city life in China and beyond.

econ.GN

HELLoRA: Hot Experts Layer-Level Low-Rank Adaptation for Mixture-of-Experts Models

Low-Rank Adaptation (LoRA) dominates parameter-efficient fine-tuning of large language models, yet most variants target dense architectures. Mixture-of-Experts (MoE) models scale parameters at near-constant per-token compute, and their sparse activation patterns create untapped opportunities for more efficient adaptation. We propose Hot-Experts Layer-level Low-Rank Adaptation (HELLoRA), which attaches LoRA modules only to the most frequently activated experts at each layer. This simple mechanism reduces trainable parameters and adapter-induced FLOPs while improving downstream performance, an effect we attribute to a form of structured regularization that preserves pretrained expert specialization. To stress-test HELLoRA under extreme parameter budgets, we further compose it with LoRI to form HELLoRI, which freezes the up-projection and sparsifies the down-projection. Across three MoE backbones, namely OlMoE-1B-7B, Mixtral-8x7B, and DeepSeekMoE, and three task families covering mathematical reasoning, code generation, and safety alignment, HELLoRA consistently outperforms strong PEFT baselines. Relative to vanilla LoRA on OlMoE, HELLoRA uses 15.7% of the trainable parameters, reduces adapter FLOPs by 38.7%, achieves 1.9x the training throughput, and improves accuracy by 9.2%. On DeepSeekMoE, HELLoRA outperforms LoRA while using only 23.2% of its trainable parameters. These results demonstrate that activation-aware adapter placement is an effective and practical route to scaling PEFT for MoE language models.

cs.LG

Ivy-Fake: A Unified Explainable Framework and Benchmark for Image and Video AIGC Detection

The rapid development of Artificial Intelligence Generated Content (AIGC) techniques has enabled the creation of high-quality synthetic content, but it also raises significant security concerns. Current detection methods face two major limitations: (1) the lack of multidimensional explainable datasets for generated images and videos. Existing open-source datasets (e.g., WildFake, GenVideo) rely on oversimplified binary annotations, which restrict the explainability and trustworthiness of trained detectors. (2) Prior MLLM-based forgery detectors (e.g., FakeVLM) exhibit insufficiently fine-grained interpretability in their step-by-step reasoning, which hinders reliable localization and explanation. To address these challenges, we introduce Ivy-Fake, the first large-scale multimodal benchmark for explainable AIGC detection. It consists of over 106K richly annotated training samples (images and videos) and 5,000 manually verified evaluation examples, sourced from multiple generative models and real world datasets through a carefully designed pipeline to ensure both diversity and quality. Furthermore, we propose Ivy-xDetector, a reinforcement learning model based on Group Relative Policy Optimization (GRPO), capable of producing explainable reasoning chains and achieving robust performance across multiple synthetic content detection benchmarks. Extensive experiments demonstrate the superiority of our dataset and confirm the effectiveness of our approach. Notably, our method improves performance on GenImage from 86.88% to 96.32%, surpassing prior state-of-the-art methods by a clear margin.

cs.CV

SpineBench: A Clinically Salient, Level-Aware Benchmark Powered by the SpineMed-450k Corpus

Spine disorders affect 619 million people globally and are a leading cause of disability, yet AI-assisted diagnosis remains limited by the lack of level-aware, multimodal datasets. Clinical decision-making for spine disorders requires sophisticated reasoning across X-ray, CT, and MRI at specific vertebral levels. However, progress has been constrained by the absence of traceable, clinically-grounded instruction data and standardized, spine-specific benchmarks. To address this, we introduce SpineMed, an ecosystem co-designed with practicing spine surgeons. It features SpineMed-450k, the first large-scale dataset explicitly designed for vertebral-level reasoning across imaging modalities with over 450,000 instruction instances, and SpineBench, a clinically-grounded evaluation framework. SpineMed-450k is curated from diverse sources, including textbooks, guidelines, open datasets, and ~1,000 de-identified hospital cases, using a clinician-in-the-loop pipeline with a two-stage LLM generation method (draft and revision) to ensure high-quality, traceable data for question-answering, multi-turn consultations, and report generation. SpineBench evaluates models on clinically salient axes, including level identification, pathology assessment, and surgical planning. Our comprehensive evaluation of several recently advanced large vision-language models (LVLMs) on SpineBench reveals systematic weaknesses in fine-grained, level-specific reasoning. In contrast, our model fine-tuned on SpineMed-450k demonstrates consistent and significant improvements across all tasks. Clinician assessments confirm the diagnostic clarity and practical utility of our model's outputs.

cs.CV

Fast Electromagnetic and RF Circuit Co-Simulation for Passive Resonator Field Calculation and Optimization in MRI

Passive resonators have been widely used in MRI to manipulate RF field distributions. However, optimizing these structures using full-wave electromagnetic simulations is computationally prohibitive, particularly for large passive resonator arrays with many degrees of freedom. This work presents a co-simulation framework tailored specifically for the analysis and optimization of passive resonators. The framework performs a single full-wave electromagnetic simulation in which the resonator's lumped components are replaced by ports, followed by circuit-level computations to evaluate arbitrary capacitor and inductor configurations. This allows integration with a genetic algorithm to rapidly optimize resonator parameters and enhance the B1 field in a targeted region of interest (ROI). We validated the method in three scenarios: (1) a single-loop passive resonator on a spherical phantom, (2) a two-loop array on a cylindrical phantom, and (3) a two-loop array on a human head model. In all cases, the co-simulation results showed excellent agreement with full-wave simulations, with relative errors below 1%. The genetic-algorithm-driven optimization, involving tens of thousands of capacitor combinations, completed in under 5 minutes, whereas equivalent full-wave EM sweeps would require an impractically long computation time. This work extends co-simulation methodology to passive resonator design for first time, enabling the fast, accurate, and scalable optimization. The approach significantly reduces computational burden while preserving full-wave accuracy, making it a powerful tool for passive RF structure development in MRI.

physics.med-ph

Integrated user scheduling and beam steering in over-the-air federated learning for mobile IoT

The rising popularity of Internet of things (IoT) has spurred technological advancements in mobile internet and interconnected systems. While offering flexible connectivity and intelligent applications across various domains, IoT service providers must gather vast amounts of sensitive data from users, which nonetheless concomitantly raises concerns about privacy breaches. Federated learning (FL) has emerged as a promising decentralized training paradigm to tackle this challenge. This work focuses on enhancing the aggregation efficiency of distributed local models by introducing over-the-air computation into the FL framework. Due to radio resource scarcity in large-scale networks, only a subset of users can participate in each training round. This highlights the need for effective user scheduling and model transmission strategies to optimize communication efficiency and inference accuracy. To address this, we propose an integrated approach to user scheduling and receive beam steering, subject to constraints on the number of selected users and transmit power. Leveraging the difference-of-convex technique, we decompose the primal non-convex optimization problem into two sub-problems, yielding an iterative solution. While effective, the computational load of the iterative method hampers its practical implementation. To overcome this, we further propose a low-complexity user scheduling policy based on characteristic analysis of the wireless channel to directly determine the user subset without iteration. Extensive experiments validate the superiority of the proposed method in terms of aggregation error and learning performance over existing approaches.

cs.DC

Efficient cascaded wavelength conversion under two-peak Stark-chirped rapid adiabatic passage via grating structures

In this paper, we demonstrate a domain inversion crystal structure to study the cascaded three-wave mixing process in a two peak Stark-chirped rapid adiabatic passage. We have achieved efficient wavelength conversion, which can be performed in intuitive order and counterintuitive order. The requirement of crystal is reduced and the flexibility of structure design is improved. When the conversion wavelength is fixed, increasing the coupling coefficient between the two peaks can reduce the intensity of the intermediate wavelength while maintaining high conversion efficiency. Compared with the cascaded wavelength conversion process based on stimulated Raman rapid adiabatic passage, the two-peak Stark-chirped rapid adiabatic passage has a larger convertible wavelength bandwidth. This scheme provides a theoretical basis for obtaining mid-infrared laser source via flexible crystal structure.

physics.optics

Scalable Deep Compressive Sensing

Deep learning has been used to image compressive sensing (CS) for enhanced reconstruction performance. However, most existing deep learning methods train different models for different subsampling ratios, which brings additional hardware burden. In this paper, we develop a general framework named scalable deep compressive sensing (SDCS) for the scalable sampling and reconstruction (SSR) of all existing end-to-end-trained models. In the proposed way, images are measured and initialized linearly. Two sampling masks are introduced to flexibly control the subsampling ratios used in sampling and reconstruction, respectively. To make the reconstruction model adapt to any subsampling ratio, a training strategy dubbed scalable training is developed. In scalable training, the model is trained with the sampling matrix and the initialization matrix at various subsampling ratios by integrating different sampling matrix masks. Experimental results show that models with SDCS can achieve SSR without changing their structure while maintaining good performance, and SDCS outperforms other SSR methods.

cs.CV

AMP-Net: Denoising based Deep Unfolding for Compressive Image Sensing

Most compressive sensing (CS) reconstruction methods can be divided into two categories, i.e. model-based methods and classical deep network methods. By unfolding the iterative optimization algorithm for model-based methods onto networks, deep unfolding methods have the good interpretation of model-based methods and the high speed of classical deep network methods. In this paper, to solve the visual image CS problem, we propose a deep unfolding model dubbed AMP-Net. Rather than learning regularization terms, it is established by unfolding the iterative denoising process of the well-known approximate message passing algorithm. Furthermore, AMP-Net integrates deblocking modules in order to eliminate the blocking artifacts that usually appear in CS of visual images. In addition, the sampling matrix is jointly trained with other network parameters to enhance the reconstruction performance. Experimental results show that the proposed AMP-Net has better reconstruction accuracy than other state-of-the-art methods with high reconstruction speed and a small number of network parameters.

eess.IV

An Entropy-based Proof of Threshold Saturation for Nonbinary SC-LDPC Ensembles on the BEC

In this paper we are concerned with the asymptotic analysis of nonbinary spatially-coupled low-density parity-check (SC-LDPC) ensembles defined over GL$\left(2^{m}\right)$ (the general linear group of degree $m$ over GF$\left(2\right)$). Our purpose is to prove threshold saturation when the transmission takes place on the binary erasure channel (BEC). To this end, we establish the duality rule for entropy for nonbinary variable-node (VN) and check-node (CN) convolutional operators to accommodate the nonbinary density evolution (DE) analysis. Based on this, we construct the explicit forms of the potential functions for uncoupled and coupled DE recursions. In addition, we show that these functions exhibit similar monotonicity properties as those for binary LDPC and SC-LDPC ensembles over general binary memoryless symmetric (BMS) channels. This leads to the threshold saturation theorem and its converse for nonbinary SC-LDPC ensembles on the BEC, following the proof technique developed by S. Kumar et al.

cs.IT