Searcharxiv⌕ Search

arXiv · 2609.36525

Reliability Testing of Medical Model Performance under Distributed Deployment

Abstract

Distributed inference has become an indispensable part of deploying medical models under practical latency, memory, and throughput constraints. Although modern frameworks improve serving efficiency through tensor parallelism, mixed precision, kernel fusion, and multi-device communication, they are generally assumed to preserve the behavior observed during centralized HuggingFace evaluation. This assumption creates an evaluation-deployment mismatch: a model may pass offline evaluation but produce a different output after the execution stack changes. To address this mismatch, we propose a testing framework and an improved, distributed-execution-sensitive medical-model benchmark that evaluates the same checkpoint and input under a centralized HuggingFace reference and matched distributed deployments. Extensive experiments across language, vision, and multimodal medical models show that execution changes can produce measurable output disagreements. Across supported visual settings, the test success rate ranges from 0.21 to 0.43 for single-modality models and from 0.32 to 0.98 for multimodal models. The benchmark is aimed at extending medical-model evaluation from capability and security to evaluation-deployment consistency.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yifei Wang, Xiaohan Zhang, Youtao Ding, Tianlin Li, Xiaoyu Zhang, Yida Yang, Li Pan. 2026-09-29. Reliability Testing of Medical Model Performance under Distributed Deployment. https://arxiv.org/abs/2609.36525

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Gen2-IC: Bridging Generative Models and Image Codecs through Latent Transport

Diffusion-based image compression has achieved strong perceptual quality at ultra-low bitrates. However, existing codecs are often tied to specific backbones and specialized components, making diverse, rapidly evolving generative models difficult to reuse. This raises a natural question: Can modern generative foundation models be connected to image compression through a simple and extensible interface? Two insights guide our design: stronger generative priors make a simpler codec interface viable, and generation and compression can be intrinsically linked through latent transport. We therefore propose Gen2-IC with two stages: (1) Latent Compression maps clean image latents to entropy-constrained latents; and (2) Latent Transport refines them with one near-terminal update based on the pretrained model. Gen2-IC requires neither auxiliary conditioning signals nor task-specific backbone modifications. With lightweight adaptation and no distillation, it supports fast encoding and one-step decoding across multiple bitrates. We validate Gen2-IC on SD-2.1, SANA-1.5, FLUX.1-dev, and Qwen-Image-2512, spanning U-Net and Transformer architectures as well as diffusion and flow-matching formulations. With stronger priors, Gen2-IC delivers gains below 0.05 bpp: the Qwen variant leads diffusion-based generative codecs in reconstruction fidelity (PSNR), perceptual similarity (LPIPS and DISTS), and recognizer-based semantic fidelity (OCR CER/WER and face-ROI similarity) across four benchmarks.

eess.IV↗

MedForge-RSI: Medical Deepfake Detection via Recursive Self-Improvement

Text-guided image editors can generate high-fidelity medical deepfakes, challenging the reliability of clinical imagery. Although reasoning-based detectors perform strongly in distribution, they degrade substantially under deployment shift. MedForge-Reasoner, an 8B vision-language model trained with supervised fine-tuning and reinforcement learning, achieves 99.2% accuracy on its target distribution, yet misclassifies 40% of authentic scans, reaches only 77% accuracy on unseen generators, and falls to 59% under transmission distortion. Adapting such models through conventional retraining is costly, requiring large-scale supervision and expert-designed guidelines. We introduce MedForge-RSI, a recursive self-improvement framework that enables a deployed detector to adapt while keeping its model weights frozen. Over 20 rounds, the detector analyzes verified errors, accumulates reusable experience, and autonomously develops image-analysis tools, while an independent acceptance test retains only validated improvements. Across 49 registered configurations, MedForge-RSI increases average accuracy over four test sets from 75.0% to 84.4%. On a held-out 4,000-image evaluation, it improves clean accuracy from 76.5% to 87.9% and transmission-distorted accuracy from 59.0% to 70.5%, with the largest gains in authentic-image recall. Controlled analysis across all 49 trajectories identifies which self-improvement mechanisms replicate across seeds, shows acceptance testing to be the largest individual contributor, and reveals a taxonomy of failed adaptations. We release the complete trajectories, including all rejected and rolled-back changes.

eess.IV↗

Rigid Motion Estimation using Accelerated Iterative Coordinate Descent (REACT) for MR Imaging

Purpose: To develop a computationally viable autofocus method for estimating 3D rigid motion in MR imaging. Theory and Methods: The proposed method, REACT, assumes a piecewise-constant motion trajectory and estimates the rigid motion parameters of individual temporal segments by optimizing an image-quality metric. Coordinate descent is adopted to decompose the high-dimensional optimization problem into a series of subproblems, each updating the motion parameters of a single temporal segment. The cost function of each subproblem is assumed to be approximately locally convex under suitable acquisition conditions. Each subproblem is then solved using a derivative-free solver, thereby avoiding an exhaustive grid search. Numerical simulations investigated the local convexity assumption and data acquisition requirements. REACT was evaluated for respiratory motion correction on in vivo free-breathing coronary MR angiography datasets. Coronary artery sharpness was quantified using unbounded image edge profile acutance (u-IEPA). Results: In numerical simulations, the objective surfaces of the subproblems were approximately locally convex when the current motion estimate was sufficiently close to the desired solution, and REACT required the data collected within each temporal segment to be sufficiently distributed across k-space. In the in vivo study, REACT yielded higher u-IEPA for both the left anterior descending artery (LAD) and the right coronary artery than did a conventional translational motion-estimation method using image-based navigators. REACT also yielded higher u-IEPA for the LAD than did a conventional autofocus nonrigid motion correction method. Conclusion: This study demonstrates the feasibility of coordinate descent for autofocus motion correction in MR imaging.

eess.IV↗