SearcharxivSearch

arXiv subjects

Mingyue Zhao

Publications and source records attributed to Mingyue Zhao.

11 recordsLinked to original sources

Thinking Like a Clinician: A Cognitive AI Agent for Clinical Diagnosis via Panoramic Profiling and Adversarial Debate

The application of large language models (LLMs) in clinical decision support faces significant challenges of "tunnel vision" and diagnostic hallucinations present in their processing unstructured electronic health records (EHRs). To address these challenges, we propose a novel chain-based clinical reasoning framework, called DxChain, which transforms the diagnostic workflow into an iterative process by mirroring a clinician's cognitive trajectory that consists of "Memory Anchoring", "Navigation" and "Verification" phases. DxChain introduces three key methodological innovations to elicit the potential of LLM: (i) a Profile-Then-Plan paradigm to mitigate cold-start hallucinations by establishing a panoramic patient baseline, (ii) a Medical Tree-of-Thoughts (Med-ToT) algorithm for strategic look ahead planning and resource aware navigation, and (iii) a Dialectical Diagnostic Verification procedure utilizing "Angel-Devil" adversarial debates to resolve complex evidence conflicts. Evaluated on two real world benchmarks, MIMIC-IV-Ext Cardiac Disease and MIMIC-IV-Ext CDM, DxChain achieves state-of-the-art performances in both diagnostic accuracy and logical consistency, offering a modular and reliable architecture for next-generation clinical AI. The code is at https://anonymous.4open.science/r/Dx-Chain.

cs.AI

Scalar Spin Chiral Order via Bond Selectivity in Strained Collinear Ferrimagnets

Scalar spin chirality (SSC) drives a series of topological transports in noncoplanar magnets. However, the ordering temperature of magnet hosting intrinsic SSC order is typically below 100 K. Current approaches to achieve near room temperature SSC order largely rely on external fields or chemical doping in noncollinear magnets. A significant challenge persists in generating and controlling SSC order in high temperature collinear magnets. Here, using the collinear ferrimagnet Mn4N with Neel temperature ~740 K as a platform, we demonstrate that isotropic strain acts as a clean and continuous tuning parameter to induce long range SSC order by first principles calculations. As strain increases from to, the magnetic ground state evolves continuously from a collinear to a noncoplanar configuration, activating the SSC order and enhancing its magnitude from 0 to ~2.32. Our quantitative orbital-resolved bonding analysis reveals that strain selectively suppresses the bond between Mn 3d orbitals and N 2p orbitals, driving dual prerequisites for the SSC order. Specifically, the decreased covalent spin-pairing activates Mn3c moments within the plane, simultaneously the suppressed N-mediated ferromagnetic superexchange interaction shifts the balance of the nearest-neighbor Mn3c sites toward antiferromagnetic exchange interaction. Our findings establish a powerful strain mediated route to construct the SSC order in high temperature collinear magnets.

cond-mat.mtrl-sci

GreenRFM: Learning a resource-efficient radiology vision-language foundation model via supervision-centric pre-training

Radiology foundation models (RFMs) have largely inherited the scale-first recipe of natural-image vision--language pre-training. This recipe is difficult to deploy in 3D radiology, where training corpora are smaller, reports vary across institutions, and receiving hospitals often need local adaptation under privacy and compute constraints. We ask whether routine radiology reports can instead be converted into auditable diagnostic supervision that shapes the image encoder, text encoder, aligned space, and local-adaptation procedure. We develop GreenRFM, a supervision-centric pre-training framework organized around four empirical principles: More distilled, Ubiquitous, Semantic-enforcing, and Task-aligning (MUST) supervision. These principles convert noisy reports into structured diagnostic signals and use them to learn discriminative unimodal encoders plus an aligned image--text space for diagnosis-centered multimodal use. GreenRFM requires 24 GPU-hours on a single 24GB GPU (lightweight variant: 6GB VRAM, 4~hours) and reaches a zero-shot CT-RATE AUC of 84.8. Evaluations using more than 200,000 volumes from six institutions and two modalities show transfer to private clinical cohorts and to musculoskeletal MRI. On a local institutional cohort, computationally feasible retraining raises macro-AUC from 70.5 to 82.1. The aligned space also improves hepatocellular-carcinoma microvascular-invasion prediction and trans-arterial chemoembolization response analysis over established clinical scores. These results support supervision-centric pre-training as a practical route to resource-efficient, locally adaptable, diagnosis-centered radiology vision--language representations.

cs.CV

MAC-Gaze: Motion-Aware Continual Calibration for Mobile Gaze Tracking

Mobile gaze tracking faces a fundamental challenge: maintaining accuracy as users naturally change their postures and device orientations. Traditional calibration approaches, like one-off, fail to adapt to these dynamic conditions, leading to degraded performance over time. We present MAC-Gaze, a Motion-Aware continual Calibration approach that leverages smartphone Inertial measurement unit (IMU) sensors and continual learning techniques to automatically detect changes in user motion states and update the gaze tracking model accordingly. Our system integrates a pre-trained visual gaze estimator and an IMU-based activity recognition model with a clustering-based hybrid decision-making mechanism that triggers recalibration when motion patterns deviate significantly from previously encountered states. To enable accumulative learning of new motion conditions while mitigating catastrophic forgetting, we employ replay-based continual learning, allowing the model to maintain performance across previously encountered motion conditions. We evaluate our system through extensive experiments on the publicly available RGBDGaze dataset and our own 10-hour multimodal MotionGaze dataset (481K+ images, 800K+ IMU readings), encompassing a wide range of postures under various motion conditions including sitting, standing, lying, and walking. Results demonstrate that our method reduces gaze estimation error by 19.9% on RGBDGaze (from 1.73 cm to 1.41 cm) and by 31.7% on MotionGaze (from 2.81 cm to 1.92 cm) compared to traditional calibration approaches. Our framework provides a robust solution for maintaining gaze estimation accuracy in mobile scenarios.

cs.HC

Quantifying the Impact of Motion on 2D Gaze Estimation in Real-World Mobile Interactions

Mobile gaze tracking involves inferring a user's gaze point or direction on a mobile device's screen from facial images captured by the device's front camera. While this technology inspires an increasing number of gaze-interaction applications, achieving consistent accuracy remains challenging due to dynamic user-device spatial relationships and varied motion conditions inherent in mobile contexts. This paper provides empirical evidence on how user mobility and behaviour affect mobile gaze tracking accuracy. We conduct two user studies collecting behaviour and gaze data under various motion conditions - from lying to maze navigation - and during different interaction tasks. Quantitative analysis has revealed behavioural regularities among daily tasks and identified head distance, head pose, and device orientation as key factors affecting accuracy, with errors increasing by up to 48.91% in dynamic conditions compared to static ones. These findings highlight the need for more robust, adaptive eye-tracking systems that account for head movements and device deflection to maintain accuracy across diverse mobile contexts.

cs.HC

3DGR-CAR: Coronary artery reconstruction from ultra-sparse 2D X-ray views with a 3D Gaussians representation

Reconstructing 3D coronary arteries is important for coronary artery disease diagnosis, treatment planning and operation navigation. Traditional reconstruction techniques often require many projections, while reconstruction from sparse-view X-ray projections is a potential way of reducing radiation dose. However, the extreme sparsity of coronary arteries in a 3D volume and ultra-limited number of projections pose significant challenges for efficient and accurate 3D reconstruction. To this end, we propose 3DGR-CAR, a 3D Gaussian Representation for Coronary Artery Reconstruction from ultra-sparse X-ray projections. We leverage 3D Gaussian representation to avoid the inefficiency caused by the extreme sparsity of coronary artery data and propose a Gaussian center predictor to overcome the noisy Gaussian initialization from ultra-sparse view projections. The proposed scheme enables fast and accurate 3D coronary artery reconstruction with only 2 views. Experimental results on two datasets indicate that the proposed approach significantly outperforms other methods in terms of voxel accuracy and visual quality of coronary arteries. The code will be available in https://github.com/windrise/3DGR-CAR.

eess.IV

Addressing Fairness Issues in Deep Learning-Based Medical Image Analysis: A Systematic Review

Deep learning algorithms have demonstrated remarkable efficacy in various medical image analysis (MedIA) applications. However, recent research highlights a performance disparity in these algorithms when applied to specific subgroups, such as exhibiting poorer predictive performance in elderly females. Addressing this fairness issue has become a collaborative effort involving AI scientists and clinicians seeking to understand its origins and develop solutions for mitigation within MedIA. In this survey, we thoroughly examine the current advancements in addressing fairness issues in MedIA, focusing on methodological approaches. We introduce the basics of group fairness and subsequently categorize studies on fair MedIA into fairness evaluation and unfairness mitigation. Detailed methods employed in these studies are presented too. Our survey concludes with a discussion of existing challenges and opportunities in establishing a fair MedIA and healthcare system. By offering this comprehensive review, we aim to foster a shared understanding of fairness among AI researchers and clinicians, enhance the development of unfairness mitigation methods, and contribute to the creation of an equitable MedIA society.

cs.CV

Large topological Hall effect arising from spin reorientation in kagome magnet Fe3Ge

Materials systems with spin chirality can provide ultra-high-density, ultra-fast, and ultralow-power information carriers for digital transformation. These material systems include magnetic skyrmions, chiral domain walls, spin reorientation,and so on. The topological Hall effect (THE) has been identified as the most convenient and effective tool for detecting the presence of spin chirality in these systems. The research on the THE that may arise from spin reorientation and specifically in Fe3Ge with spin reorientation remains an unexplored area, so we study the THE in Fe3Ge Conduct systematic research. X-Ray Diffraction (XRD) results indicate that our Fe3Ge ribbon sample has a D019 structure. First-principles calculations and magnetic and electrical testing confirm spin reorientation in the Fe3Ge ribbon sample at 350 K.The Hall resistivity test results are consistent with our expectations, indicating the presence of the THE in the Fe3Ge ribbon sample. The topological Hall resistivity reaches a maximum value of 0.69 mΩ cm at 400 K. For the first time, a detailed experimental study of the THE in Fe3Ge with spin reorientation has been conducted, introducing a new member to the family of THE.

cond-mat.mtrl-sci

Skeleton Supervised Airway Segmentation

Fully-supervised airway segmentation has accomplished significant triumphs over the years in aiding pre-operative diagnosis and intra-operative navigation. However, full voxel-level annotation constitutes a labor-intensive and time-consuming task, often plagued by issues such as missing branches, branch annotation discontinuity, or erroneous edge delineation. label-efficient solutions for airway extraction are rarely explored yet primarily demanding in medical practice. To this end, we introduce a novel skeleton-level annotation (SkA) tailored to the airway, which simplifies the annotation workflow while enhancing annotation consistency and accuracy, preserving the complete topology. Furthermore, we propose a skeleton-supervised learning framework to achieve accurate airway segmentation. Firstly, a dual-stream buffer inference is introduced to realize initial label propagation from SkA, avoiding the collapse of direct learning from SkA. Then, we construct a geometry-aware dual-path propagation framework (GDP) to further promote complementary propagation learning, composed of hard geometry-aware propagation learning and soft geometry-aware propagation guidance. Experiments reveal that our proposed framework outperforms the competing methods with SKA, which amounts to only 1.96% airways, and achieves comparable performance with the baseline model that is fully supervised with 100% airways, demonstrating its significant potential in achieving label-efficient segmentation for other tubular structures, such as vessels.

cs.CV

Landauer-QFLPS model for mixed Schottky-Ohmic contact two-dimensional transistors

Two-dimensional material-based field effect transistors (2DM-FETs) are playing a revolutionary role in electronic devices. However, after years of development, no device model can match the Pao-Sah model for standard silicon-based transistors in terms of physical accuracy and computational efficiency to support large-scale integrated circuit design. One remaining critical obstacle is the contacts of 2DM-FETs. In order to self-consistently include the contact effect in the current model, it is necessary to perform self-consistent calculations, which is a fatal flaw for applications that prioritize efficiency. Here, we report that the Landauer-QFLPS model effectively overcomes the above contradiction, where QFLPS means quasi-Fermi-level phase space theory. By connecting the physical pictures of the contact and the intrinsic channel part, we have successfully derived a drain-source current formula including the contact effect. To verify the model, we prepared transistors based on two typical 2DMs, black phosphorus (BP) and molybdenum disulfide (MoS2), the former having ambipolar transport and the latter showing electron-dominant unipolar transport. The proposed new formula could describe both 2DM-FETs with Schottky or Ohmic contacts. Moreover, compared with traditional methods, the proposed model has the advantages of accuracy and efficiency, especially in describing non-monotonic drain conductance characteristics, because the contact effect is self-consistently and compactly packaged as an exponential term. More importantly, we also examined the model at the circuit level. Here, we fabricated a three-bit threshold inverter quantizer circuit based on ambipolar-BP process and experimentally demonstrated that the model can accurately predict the circuit performance. This industry-benign 2DM-FET model is supposed to be very useful for the development of 2DM-FET-based integrated circuits.

physics.app-ph

GDDS: Pulmonary Bronchioles Segmentation with Group Deep Dense Supervision

Airway segmentation, especially bronchioles segmentation, is an important but challenging task because distal bronchus are sparsely distributed and of a fine scale. Existing neural networks usually exploit sparse topology to learn the connectivity of bronchioles and inefficient shallow features to capture such high-frequency information, leading to the breakage or missed detection of individual thin branches. To address these problems, we contribute a new bronchial segmentation method based on Group Deep Dense Supervision (GDDS) that emphasizes fine-scale bronchioles segmentation in a simple-but-effective manner. First, Deep Dense Supervision (DDS) is proposed by constructing local dense topology skillfully and implementing dense topological learning on a specific shallow feature layer. GDDS further empowers the shallow features with better perception ability to detect bronchioles, even the ones that are not easily discernible to the naked eye. Extensive experiments on the BAS benchmark dataset have shown that our method promotes the network to have a high sensitivity in capturing fine-scale branches and outperforms state-of-the-art methods by a large margin (+12.8 % in BD and +8.8 % in TD) while only introducing a small number of extra parameters.

cs.CV