SearcharxivSearch

arXiv subjects

Zimu Yuan

Publications and source records attributed to Zimu Yuan.

11 recordsLinked to original sources

OrthoDiffusion: A Generalizable Multi-Task Diffusion Foundation Model for Musculoskeletal MRI Interpretation

Musculoskeletal disorders represent a significant global health burden and are a leading cause of disability worldwide. While MRI is essential for accurate diagnosis, its interpretation remains exceptionally challenging. Radiologists must identify multiple potential abnormalities within complex anatomical structures across different imaging planes, a process that requires significant expertise and is prone to variability. We developed OrthoDiffusion, a unified diffusion-based foundation model designed for multi-task musculoskeletal MRI interpretation. The framework utilizes three orientation-specific 3D diffusion models, pre-trained in a self-supervised manner on 15,948 unlabeled knee MRI scans, to learn robust anatomical features from sagittal, coronal, and axial views. These view-specific representations are integrated to support diverse clinical tasks, including anatomical segmentation and multi-label diagnosis. Our evaluation demonstrates that OrthoDiffusion achieves excellent performance in the segmentation of 11 knee structures and the detection of 8 knee abnormalities. The model exhibited remarkable robustness across different clinical centers and MRI field strengths, consistently outperforming traditional supervised models. Notably, in settings where labeled data was scarce, OrthoDiffusion maintained high diagnostic precision using only 10\% of training labels. Furthermore, the anatomical representations learned from knee imaging proved highly transferable to other joints, achieving strong diagnostic performance across 11 diseases of the ankle and shoulder. These findings suggest that diffusion-based foundation models can serve as a unified platform for multi-disease diagnosis and anatomical segmentation, potentially improving the efficiency and accuracy of musculoskeletal MRI interpretation in real-world clinical workflows.

cs.CV

A multimodal vision foundation model for generalizable knee pathology

Musculoskeletal disorders represent a leading cause of global disability, creating an urgent demand for precise interpretation of medical imaging. Current artificial intelligence (AI) approaches in orthopedics predominantly rely on task-specific, supervised learning paradigms. These methods are inherently fragmented, require extensive annotated datasets, and often lack generalizability across different modalities and clinical scenarios. The development of foundation models in this field has been constrained by the scarcity of large-scale, curated, and open-source musculoskeletal datasets. To address these challenges, we introduce OrthoFoundation, a multimodal vision foundation model optimized for musculoskeletal pathology. We constructed a pre-training dataset of 1.2 million unlabeled knee X-ray and MRI images from internal and public databases. Utilizing a Dinov3 backbone, the model was trained via self-supervised contrastive learning to capture robust radiological representations. OrthoFoundation achieves state-of-the-art (SOTA) performance across 14 downstream tasks. It attained superior accuracy in X-ray osteoarthritis diagnosis and ranked first in MRI structural injury detection. The model demonstrated remarkable label efficiency, matching supervised baselines using only 50% of labeled data. Furthermore, despite being pre-trained on knee images, OrthoFoundation exhibited exceptional cross-anatomy generalization to the hip, shoulder, and ankle. OrthoFoundation represents a significant advancement toward general-purpose AI for musculoskeletal imaging. By learning fundamental, joint-agnostic radiological semantics from large-scale multimodal data, it overcomes the limitations of conventional models, which provides a robust framework for reducing annotation burdens and enhancing diagnostic accuracy in clinical practice.

cs.CV

The Illusion of Clinical Reasoning: A Benchmark Reveals the Pervasive Gap in Vision-Language Models for Clinical Competency

Background: The rapid integration of foundation models into clinical practice and public health necessitates a rigorous evaluation of their true clinical reasoning capabilities beyond narrow examination success. Current benchmarks, typically based on medical licensing exams or curated vignettes, fail to capture the integrated, multimodal reasoning essential for real-world patient care. Methods: We developed the Bones and Joints (B&J) Benchmark, a comprehensive evaluation framework comprising 1,245 questions derived from real-world patient cases in orthopedics and sports medicine. This benchmark assesses models across 7 tasks that mirror the clinical reasoning pathway, including knowledge recall, text and image interpretation, diagnosis generation, treatment planning, and rationale provision. We evaluated eleven vision-language models (VLMs) and six large language models (LLMs), comparing their performance against expert-derived ground truth. Results: Our results demonstrate a pronounced performance gap between task types. While state-of-the-art models achieved high accuracy, exceeding 90%, on structured multiple-choice questions, their performance markedly declined on open-ended tasks requiring multimodal integration, with accuracy scarcely reaching 60%. VLMs demonstrated substantial limitations in interpreting medical images and frequently exhibited severe text-driven hallucinations, often ignoring contradictory visual evidence. Notably, models specifically fine-tuned for medical applications showed no consistent advantage over general-purpose counterparts. Conclusions: Current artificial intelligence models are not yet clinically competent for complex, multimodal reasoning. Their safe deployment should currently be limited to supportive, text-based roles. Future advancement in core clinical tasks awaits fundamental breakthroughs in multimodal integration and visual understanding.

cs.CV

Closed-Form Error Analysis on RSS-based Indoor Localization Method

Received Signal Strength (RSS) is considered as a promising measurement for indoor positioning. Lots of RSS-based localization methods have been proposed by its convenience and low cost. This paper focuses on two challenging issues in RSS-based localization schemes: finding optimal localization algorithms and knowing what affects the accuracy of these algorithms. Through theoretical and experimental analysis, we present three important results: 1) we prove that the Non-linear Least Square (NLS) method is efficient to solve the RSS-based localization problem, i.e., its localization error is minimal in theory; 2) we provide a closed-form expression of the localization error for the NLS method. Such an expression reveals the key factors that affect the accuracy of the NLS method; 3) we further conduct a new lower bound of localization error, i.e., the best accuracy achieved possibly. The study of this paper shows the inherent limitation of RSS-based localization schemes and provides guidance for achieving optimal accuracy.

cs.NI

Flow Demands Oriented Node Placement in Multi-Hop Wireless Networks

In multi-hop wireless networks, flow demands mean that some nodes have routing demands of transmitting their data to other nodes with a certain level of transmission rate. When a set of nodes have been deployed with flow demands, it is worth to know how to construct paths to satisfy these flow demands with nodes placed as few as possible. In this paper, we study this flow demands oriented node placement problem that has not been addressed before. In particular, we divide and conquer the problem by three steps: calculating the maximal flow for single routing demand, calculating the maximal flow for multiple routing demands, and finding the minimal number of nodes for multiple routing demands with flow requirement. During the above solving procedure, we prove that the second and third step are NP-hard and propose two algorithms that have polynomial-time complexity. The proposed algorithms are evaluated under practical scenarios. The experiments show that the proposed algorithms can achieve satisfactory results on both flow demands and total number of wireless nodes.

cs.NI

A Cyber-Human Interaction Based System on Mobile Phone for Indoor Localization

In this article, we study the Cyber-Human Interaction (CHI) based approach that the "Human" part sets a list of location-based objectives and makes the pathway decision whereas the "Cyber" part provides the pathway suggestion, infer heuristics from the environment along the pathway and incrementally resolve the location-based objectives with new heuristics for indoor localization. For this study, we implement a CHI-based system on mobile phone. The CHI-based system offers the pathway suggestion and the solution of the location-based objectives based on its trajectory management. Without any priori knowledge on the area of interest and any aid from other equipments, a laborer can achieve his location-based objectives by walking through the area of interest and simultaneously online interacting with the CHI-based system installed in his phone. In evaluation, we conduct the experiments and show the advantage CHI in reducing the time cost and the expense cost for the laborer.

cs.HC

Online Query Scheduling on Source Permutation for Big Data Integration

Big data integration could involve a large number of sources with unpredictable redundancy information between them. The approach of building a central warehousing to integrate big data from all sources then becomes infeasible because of so large number of sources and continuous updates happening. A practical approach is to apply online query scheduling that inquires data from sources at runtime upon receiving a query. In this paper, we address the Time-Cost Minimization Problem for online query scheduling, and tackle the challenges of source permutation and statistics estimation to minimize the time cost of retrieving answers for the real-time receiving query. We propose the online scheduling strategy that enables the improvement of statistics, the construction of source permutation and the execution of query working in parallel. Experimental results show high efficiency and scalability of our scheduling strategy.

cs.DB

Beacon Node Placement for Minimal Localization Error

Beacon node placement, node-to-node measurement, and target node positioning are the three key steps for a localization process. However, compared with the other two steps, beacon node placement still lacks a comprehensive, systematic study in research literatures. To fill this gap, we address the Beacon Node Placment (BNP) problem that deploys beacon nodes for minimal localization error in this paper. BNP is difficult in that the localization error is determined by a complicated combination of factors, i.e., the localization error differing greatly under a different environment, with a different algorithm applied, or with a different type of beacon node used. In view of the hardness of BNP, we propose an approximate function to reduce time cost in localization error calculation, and also prove its time complexity and error bound. By approximation, a sub-optimal distribution of beacon nodes could be found within acceptable time cost for placement. In the experiment, we test our method and compare it with other node placement methods under various settings and environments. The experimental results show feasibility and effectiveness of our method in practice.

cs.NI

CIUV: Collaborating Information Against Unreliable Views

In many real world applications, the information of an object can be obtained from multiple sources. The sources may provide different point of views based on their own origin. As a consequence, conflicting pieces of information are inevitable, which gives rise to a crucial problem: how to find the truth from these conflicts. Many truth-finding methods have been proposed to resolve conflicts based on information trustworthy (i.e. more appearance means more trustworthy) as well as source reliability. However, the factor of men's involvement, i.e., information may be falsified by men with malicious intension, is more or less ignored in existing methods. Collaborating the possible relationship between information's origins and men's participation are still not studied in research. To deal with this challenge, we propose a method -- Collaborating Information against Unreliable Views (CIUV) --- in dealing with men's involvement for finding the truth. CIUV contains 3 stages for interactively mitigating the impact of unreliable views, and calculate the truth by weighting possible biases between sources. We theoretically analyze the error bound of CIUV, and conduct intensive experiments on real dataset for evaluation. The experimental results show that CIUV is feasible and has the smallest error compared with other methods.

cs.DB

Founding Digital Currency on Imprecise Commodity

Current digital currency schemes provide instantaneous exchange on precise commodity, in which "precise" means a buyer can possibly verify the function of the commodity without error. However, imprecise commodities, e.g. statistical data, with error existing are abundant in digital world. Existing digital currency schemes do not offer a mechanism to help the buyer for payment decision on precision of commodity, which may lead the buyer to a dilemma between having to buy and being unconfident. In this paper, we design a currency schemes IDCS for imprecise digital commodity. IDCS completes a trade in three stages of handshake between a buyer and providers. We present an IDCS prototype implementation that assigns weights on the trustworthy of the providers, and calculates a confidence level for the buyer to decide the quality of a imprecise commodity. In experiment, we characterize the performance of IDCS prototype under varying impact factors.

cs.CY

Minimum Latency Broadcast Scheduling in Single-Radio Multi-Channel Wireless Ad-Hoc Networks

We study the minimum latency broadcast scheduling (MLBS) problem in Single-Radio Multi-Channel (SR-MC) wireless ad-hoc networks (WANETs), which are modeled by Unit Disk Graphs. Nodes with this capability have their fixed reception channels, but can switch their transmission channels to communicate with their neighbors. The single-radio and multi-channel model prevents existing algorithms for single-channel networks achieving good performance. First, the common assumption that one transmission reaches all the neighboring nodes does not hold naturally. Second, the multi-channel dimension provides new opportunities to schedule the broadcast transmissions in parallel. We show MLBS problem in SR-MC WANETs is NP-hard, and present a benchmark algorithm: Basic Transmission Scheduling (BTS), which has approximation ratio of 4k + 12. Here k is the number of orthogonal channels in SR-MC WANETs. Then we propose an Enhanced Transmission Scheduling (ETS) algorithm, improving the approximation ratio to k + 23. Simulation results show that ETS achieves better performance over BTS, and the performance of ETS approaches the lower bound.

cs.NI