SearcharxivSearch

arXiv subjects

Tianyu Han

Publications and source records attributed to Tianyu Han.

At least 19 recordsLinked to original sources

Auditable CT Phenotyping Through Report-derived Radiological Observations

Medical image foundation models can predict clinical phenotypes from computed tomography (CT), but strong performance leaves open whether they read disease-specific findings or shortcuts that correlate with the diagnosis. We tested this in 221 electronic-health-record (EHR) phenotypes using Auditable CT phenotyping (ACT), built on report-derived radiological observations. We trained ACT on 38,317 patients, mined 376,194 observations and evaluated it in 25,183 held-out patients. ACT exceeded five vision-language baselines on zero-shot annotation, and CT-CLIP across 221 phenotypes from unseen CT pulmonary angiography, both under zero-shot scoring (0.651 versus 0.572) and under linear probing (0.709 versus 0.662). Reading each probe exposes what accuracy conceals: only 97 observations occupy the 221 rank-1 positions, and one phrase describing aortic and coronary calcification ranks first for 20 phenotypes, including osteoporosis, urinary tract infection and major depressive disorder. Restricting the bank to clinician-specified evidence redirects those probes onto phenotype-related observations in 86 phenotypes at no accuracy cost (0.751 versus 0.741). Accurate CT-based EHR phenotyping can therefore rest on observations that are not valid evidence for the coded phenotype and that ACT can identify and intervene on.

cs.CV

Scaling-Based Reciprocal Control Barrier Functions for Nonholonomic Mobile Robots

This paper studies the construction of control barrier functions (CBFs) for force-controlled nonholonomic mobile robots subject to relative-degree-two safety constraints arising from position-level obstacle avoidance. A scaling-based reciprocal barrier construction is proposed, in which a positive motion-dependent scaling factor is placed in the numerator of a reciprocal barrier associated with the original physical safety function. The resulting barrier is defined exactly on the interior of the physical safe set and becomes singular on its boundary, thereby preserving the certified interior domain of the original safety constraint while recovering first-order control authority. For a force-controlled nonholonomic robot model, sufficient conditions are derived under which the proposed construction defines a reciprocal CBF, and the interior of the physical safe set is forward invariant under controllers satisfying the induced reciprocal-CBF condition. A scalar strict-feedback system is further used to provide a structural interpretation of the underlying higher-relative-degree cascade under explicit structural assumptions. Numerical simulations demonstrate the induced safe-set geometry and its integration with an optimization-based control framework for obstacle avoidance.

eess.SY

Semi-Blind Fluid Antenna System: Port Selection via Statistical Analysis

The fluid antenna system (FAS) enables position reconfigurability. A potential drawback of real-time FAS, however, is that it requires complete channel state information (CSI) for each FAS port at every communication time slot, an approach referred to as ideal-FAS. Recognizing the difficulties of achieving ideal-FAS, we propose a FAS scheme based on incomplete CSI, referred to as semi-blind FAS. This paper first introduces the spatial-temporal framework of FAS, upon which the proposed semi-blind FAS is developed. The proposed semi-blind FAS is lightweight and computationally efficient, scalable to an arbitrary number of ports and time slots, and operates without pre-training or deep learning structures. The scheme effectively exploits incomplete historical CSI to estimate the conditional distribution across all FAS ports at the desired time slot, thereby identifying the statistical optimal port for signal reception. Generally, the key idea of semi-blind FAS is to select the optimal port through conditional distribution analysis, from a statistical perspective, with optimality defined according to the scenario of interest. Inspired by information-theoretic entropy, we further develop the residual entropy power ratio to characterize how physical parameters influence the performance gap between semi-blind FAS and ideal-FAS. Our analysis reveals that estimation performance depends not only on the number of sampled ports and time slots, but also on the specific indices of ports with given CSI at each time slot, i.e., the port sampling strategy. This critical factor has been largely overlooked in existing port estimation studies. Numerical results demonstrate that the proposed semi-blind FAS achieves performance comparable to, and in some cases indistinguishable from, that of ideal-FAS, while requiring significantly fewer port CSI measurements and lower port switching speeds.

eess.SP

Fluid Antenna Enabled Compact Ultra Massive Antenna Array for Satellite Communications

Satellites provide seamless coverage and are critical for emergency communications during natural disasters. However, their performance is constrained by limited spectrum and high deployment cost. To address these issues, we propose a fluid antenna system (FAS)-based solution that enables dynamic signal adaptation. Building on this concept, a compact ultra-massive antenna array (CUMA) is introduced, where multiple ports are simultaneously activated to coherently combine signal components. This design mitigates interference while reducing cost, as each fluid antenna requires only a single RF chain yet achieves significant improvement in the received signal-to-interference-plus-noise ratio (SINR). We consider a satellite CUMA network where all ground users share the same satellite for uplink transmission, and CUMA is employed to suppress inter-user interference. Closed-form expressions for the received signal power, interference power, and their distributions are derived. Based on these results, the outage probability is obtained in a unified form along with an accurate approximation, and the ergodic rate is characterized. Our analysis identifies the conditions under which CUMA outperforms maximum ratio combining in satellite systems. Notably, with sufficiently compact fluid antenna configurations, the received signal becomes deterministic, indicating that system performance is dominated by interference statistics. Moreover, increasing the number of ports yields a linear beamforming gain. Numerical results further compare orthogonal and non-orthogonal multiple access CUMA, showing that the latter achieves superior performance under wideband conditions.

eess.SP

Finite Boundary-Layer Residence Certificates for Non-Strict Control Barrier Functions

Non-strict control barrier function (CBF) conditions guarantee safety through forward invariance, but they do not preclude trajectories from remaining near the safe-set boundary for extended continuous time intervals. This paper develops a finite boundary-layer residence certificate for such settings. The certificate preserves the standard non-strict CBF safety condition and uses a bounded auxiliary function whose derivative is bounded away from zero in a prescribed boundary layer, yielding an explicit upper bound on every uninterrupted residence interval. For control-affine systems, a selected auxiliary branch is implemented as an additional affine constraint in a CBF-QP, and a tangential-input compatibility condition is given to ensure simultaneous feasibility with the hard CBF constraint for unconstrained inputs. A local-chart version handles angular or multi-valued auxiliary functions such as $\operatorname{atan2}$. Single-integrator, double-integrator, and nonholonomic unicycle examples illustrate the resulting radial--tangential construction and its local-chart and feasibility limitations.

eess.SY

Medical SAM3: A Foundation Model for Universal Prompt-Driven Medical Image Segmentation

Promptable segmentation foundation models such as SAM3 have demonstrated strong generalization capabilities through interactive and concept-based prompting. However, their direct applicability to medical image segmentation remains limited by severe domain shifts, the absence of privileged spatial prompts, and the need to reason over complex anatomical and volumetric structures. Here we present Medical SAM3, a foundation model for universal prompt-driven medical image segmentation, obtained by fully fine-tuning SAM3 on large-scale, heterogeneous 2D and 3D medical imaging datasets with paired segmentation masks and text prompts. Through a systematic analysis of vanilla SAM3, we observe that its performance degrades substantially on medical data, with its apparent competitiveness largely relying on strong geometric priors such as ground-truth-derived bounding boxes. These findings motivate full model adaptation beyond prompt engineering alone. By fine-tuning SAM3's model parameters on 33 datasets spanning 10 medical imaging modalities, Medical SAM3 acquires robust domain-specific representations while preserving prompt-driven flexibility. Extensive experiments across organs, imaging modalities, and dimensionalities demonstrate consistent and significant performance gains, particularly in challenging scenarios characterized by semantic ambiguity, complex morphology, and long-range 3D context. Our results establish Medical SAM3 as a universal, text-guided segmentation foundation model for medical imaging and highlight the importance of holistic model adaptation for achieving robust prompt-driven segmentation under severe domain shift. Code and model will be made available at https://github.com/AIM-Research-Lab/Medical-SAM3.

cs.CV

Further Results on Safety-Critical Stabilization of Force-Controlled Nonholonomic Mobile Robots

In this paper, we address the stabilization problem for force-controlled nonholonomic mobile robots under safety-critical constraints. We propose a continuous, time-invariant control law based on the gamma m-quadratic programming (gamma m-QP) framework, which unifies control Lyapunov functions (CLFs) and control barrier functions (CBFs) to enforce both stability and safety in the closed-loop system. For the first time, we construct a global, time-invariant, strict Lyapunov function for the closed-loop nonholonomic mobile robot full-dynamic system with a nominal stabilization controller in polar coordinates; this strict Lyapunov function then serves as the CLF in the QP design. Next, by exploiting the inherent cascaded structure of the vehicle dynamics, we develop a CBF for the mobile robot via an integrator backstepping procedure. Our main results guarantee both asymptotic stability and safety for the closed-loop system. Both the simulation and experimental results are presented to illustrate the effectiveness and performance of our approach.

eess.SY

Cell-free Fluid Antenna Multiple Access Networks

Fluid antenna enables position reconfigurability that gives transceiver access to a high-resolution spatial signal and the ability to avoid interference through the ups and downs of fading channels. Previous studies investigated this fluid antenna multiple access (FAMA) approach in a single-cell setup only. In this paper, we consider a cell-free network architecture in which users are associated with the nearest base stations (BSs) and all users share the same physical channel. Each BS has multiple fixed antennas that employ maximum ratio transmission (MRT) to beam to its associated users while each user relies on its fluid antenna system (FAS) on one radio frequency (RF) chain to overcome the inter-user interference. Our aim is to analyze the outage probability performance of such cell-free FAMA network when both large- and small-scale fading effects are considered. To do so, we derive the distribution of the received \textcolor{black}{magnitude} for a typical user and then the interference distribution under both fast and slow port switching techniques. The outage probability is finally obtained in integral form in each case. Numerical results demonstrate that in an interference-limited situation, although fast port switching is typically understood as the superior method for FAMA, slow port switching emerges as a more effective solution when there is a large antenna array at the BS. Moreover, it is revealed that FAS at each user can serve to greatly reduce the burden of BS in terms of both antenna costs and CSI estimation overhead, thereby enhancing the scalability of cell-free networks.

eess.SP

Medical Slice Transformer: Improved Diagnosis and Explainability on 3D Medical Images with DINOv2

MRI and CT are essential clinical cross-sectional imaging techniques for diagnosing complex conditions. However, large 3D datasets with annotations for deep learning are scarce. While methods like DINOv2 are encouraging for 2D image analysis, these methods have not been applied to 3D medical images. Furthermore, deep learning models often lack explainability due to their "black-box" nature. This study aims to extend 2D self-supervised models, specifically DINOv2, to 3D medical imaging while evaluating their potential for explainable outcomes. We introduce the Medical Slice Transformer (MST) framework to adapt 2D self-supervised models for 3D medical image analysis. MST combines a Transformer architecture with a 2D feature extractor, i.e., DINOv2. We evaluate its diagnostic performance against a 3D convolutional neural network (3D ResNet) across three clinical datasets: breast MRI (651 patients), chest CT (722 patients), and knee MRI (1199 patients). Both methods were tested for diagnosing breast cancer, predicting lung nodule dignity, and detecting meniscus tears. Diagnostic performance was assessed by calculating the Area Under the Receiver Operating Characteristic Curve (AUC). Explainability was evaluated through a radiologist's qualitative comparison of saliency maps based on slice and lesion correctness. P-values were calculated using Delong's test. MST achieved higher AUC values compared to ResNet across all three datasets: breast (0.94$\pm$0.01 vs. 0.91$\pm$0.02, P=0.02), chest (0.95$\pm$0.01 vs. 0.92$\pm$0.02, P=0.13), and knee (0.85$\pm$0.04 vs. 0.69$\pm$0.05, P=0.001). Saliency maps were consistently more precise and anatomically correct for MST than for ResNet. Self-supervised 2D models like DINOv2 can be effectively adapted for 3D medical imaging using MST, offering enhanced diagnostic accuracy and explainability compared to convolutional neural networks.

eess.IV

Biomedical Large Languages Models Seem not to be Superior to Generalist Models on Unseen Medical Data

Large language models (LLMs) have shown potential in biomedical applications, leading to efforts to fine-tune them on domain-specific data. However, the effectiveness of this approach remains unclear. This study evaluates the performance of biomedically fine-tuned LLMs against their general-purpose counterparts on a variety of clinical tasks. We evaluated their performance on clinical case challenges from the New England Journal of Medicine (NEJM) and the Journal of the American Medical Association (JAMA) and on several clinical tasks (e.g., information extraction, document summarization, and clinical coding). Using benchmarks specifically chosen to be likely outside the fine-tuning datasets of biomedical models, we found that biomedical LLMs mostly perform inferior to their general-purpose counterparts, especially on tasks not focused on medical knowledge. While larger models showed similar performance on case tasks (e.g., OpenBioLLM-70B: 66.4% vs. Llama-3-70B-Instruct: 65% on JAMA cases), smaller biomedical models showed more pronounced underperformance (e.g., OpenBioLLM-8B: 30% vs. Llama-3-8B-Instruct: 64.3% on NEJM cases). Similar trends were observed across the CLUE (Clinical Language Understanding Evaluation) benchmark tasks, with general-purpose models often performing better on text generation, question answering, and coding tasks. Our results suggest that fine-tuning LLMs to biomedical data may not provide the expected benefits and may potentially lead to reduced performance, challenging prevailing assumptions about domain-specific adaptation of LLMs and highlighting the need for more rigorous evaluation frameworks in healthcare AI. Alternative approaches, such as retrieval-augmented generation, may be more effective in enhancing the biomedical capabilities of LLMs without compromising their general knowledge.

cs.CL

Safety-Critical Stabilization of Force-Controlled Nonholonomic Mobile Robots

We present a safety-critical controller for the problem of stabilization for force-controlled nonholonomic mobile robots. The proposed control law is based on the constructions of control Lyapunov functions (CLFs) and control barrier functions (CBFs) for cascaded systems. To address nonholonomicity, we design the nominal controller that guarantees global asymptotic stability and local exponential stability for the closed-loop system in polar coordinates and construct a strict Lyapunov function valid on any compact sets. Furthermore, we present a procedure for constructing CBFs for cascaded systems, utilizing the CBF of the kinematic model through integrator backstepping. Quadratic programming is employed to combine CLFs and CBFs to integrate both stability and safety in the closed loop. The proposed control law is time-invariant, continuous along trajectories, and easy to implement. Our main results guarantee both safety and local asymptotic stability for the closed-loop system.

eess.SY

On Instabilities of Unsupervised Denoising Diffusion Models in Magnetic Resonance Imaging Reconstruction

Denoising diffusion models offer a promising approach to accelerating magnetic resonance imaging (MRI) and producing diagnostic-level images in an unsupervised manner. However, our study demonstrates that even tiny worst-case potential perturbations transferred from a surrogate model can cause these models to generate fake tissue structures that may mislead clinicians. The transferability of such worst-case perturbations indicates that the robustness of image reconstruction may be compromised due to MR system imperfections or other sources of noise. Moreover, at larger perturbation strengths, diffusion models exhibit Gaussian noise-like artifacts that are distinct from those observed in supervised models and are more challenging to detect. Our results highlight the vulnerability of current state-of-the-art diffusion-based reconstruction models to possible worst-case perturbations and underscore the need for further research to improve their robustness and reliability in clinical settings.

eess.IV

Compute-Efficient Medical Image Classification with Softmax-Free Transformers and Sequence Normalization

The Transformer model has been pivotal in advancing fields such as natural language processing, speech recognition, and computer vision. However, a critical limitation of this model is its quadratic computational and memory complexity relative to the sequence length, which constrains its application to longer sequences. This is especially crucial in medical imaging where high-resolution images can reach gigapixel scale. Efforts to address this issue have predominantely focused on complex techniques, such as decomposing the softmax operation integral to the Transformer's architecture. This paper addresses this quadratic computational complexity of Transformer models and introduces a remarkably simple and effective method that circumvents this issue by eliminating the softmax function from the attention mechanism and adopting a sequence normalization technique for the key, query, and value tokens. Coupled with a reordering of matrix multiplications this approach reduces the memory- and compute complexity to a linear scale. We evaluate this approach across various medical imaging datasets comprising fundoscopic, dermascopic, radiologic and histologic imaging data. Our findings highlight that these models exhibit a comparable performance to traditional transformer models, while efficiently handling longer sequences.

cs.CV

LongHealth: A Question Answering Benchmark with Long Clinical Documents

Background: Recent advancements in large language models (LLMs) offer potential benefits in healthcare, particularly in processing extensive patient records. However, existing benchmarks do not fully assess LLMs' capability in handling real-world, lengthy clinical data. Methods: We present the LongHealth benchmark, comprising 20 detailed fictional patient cases across various diseases, with each case containing 5,090 to 6,754 words. The benchmark challenges LLMs with 400 multiple-choice questions in three categories: information extraction, negation, and sorting, challenging LLMs to extract and interpret information from large clinical documents. Results: We evaluated nine open-source LLMs with a minimum of 16,000 tokens and also included OpenAI's proprietary and cost-efficient GPT-3.5 Turbo for comparison. The highest accuracy was observed for Mixtral-8x7B-Instruct-v0.1, particularly in tasks focused on information retrieval from single and multiple patient documents. However, all models struggled significantly in tasks requiring the identification of missing information, highlighting a critical area for improvement in clinical data interpretation. Conclusion: While LLMs show considerable potential for processing long clinical documents, their current accuracy levels are insufficient for reliable clinical use, especially in scenarios requiring the identification of missing information. The LongHealth benchmark provides a more realistic assessment of LLMs in a healthcare setting and highlights the need for further model refinement for safe and effective clinical application. We make the benchmark and evaluation code publicly available.

cs.CL

Medical Foundation Models are Susceptible to Targeted Misinformation Attacks

Large language models (LLMs) have broad medical knowledge and can reason about medical information across many domains, holding promising potential for diverse medical applications in the near future. In this study, we demonstrate a concerning vulnerability of LLMs in medicine. Through targeted manipulation of just 1.1% of the model's weights, we can deliberately inject an incorrect biomedical fact. The erroneous information is then propagated in the model's output, whilst its performance on other biomedical tasks remains intact. We validate our findings in a set of 1,038 incorrect biomedical facts. This peculiar susceptibility raises serious security and trustworthiness concerns for the application of LLMs in healthcare settings. It accentuates the need for robust protective measures, thorough verification mechanisms, and stringent management of access to these models, ensuring their reliable and safe use in medical practice.

cs.LG

Reconstruction of Patient-Specific Confounders in AI-based Radiologic Image Interpretation using Generative Pretraining

Detecting misleading patterns in automated diagnostic assistance systems, such as those powered by Artificial Intelligence, is critical to ensuring their reliability, particularly in healthcare. Current techniques for evaluating deep learning models cannot visualize confounding factors at a diagnostic level. Here, we propose a self-conditioned diffusion model termed DiffChest and train it on a dataset of 515,704 chest radiographs from 194,956 patients from multiple healthcare centers in the United States and Europe. DiffChest explains classifications on a patient-specific level and visualizes the confounding factors that may mislead the model. We found high inter-reader agreement when evaluating DiffChest's capability to identify treatment-related confounders, with Fleiss' Kappa values of 0.8 or higher across most imaging findings. Confounders were accurately captured with 11.1% to 100% prevalence rates. Furthermore, our pretraining process optimized the model to capture the most relevant information from the input radiographs. DiffChest achieved excellent diagnostic accuracy when diagnosing 11 chest conditions, such as pleural effusion and cardiac insufficiency, and at least sufficient diagnostic accuracy for the remaining conditions. Our findings highlight the potential of pretraining based on diffusion models in medical image classification, specifically in providing insights into confounding factors and model robustness.

cs.CV

Large Language Models Streamline Automated Machine Learning for Clinical Studies

A knowledge gap persists between machine learning (ML) developers (e.g., data scientists) and practitioners (e.g., clinicians), hampering the full utilization of ML for clinical data analysis. We investigated the potential of the ChatGPT Advanced Data Analysis (ADA), an extension of GPT-4, to bridge this gap and perform ML analyses efficiently. Real-world clinical datasets and study details from large trials across various medical specialties were presented to ChatGPT ADA without specific guidance. ChatGPT ADA autonomously developed state-of-the-art ML models based on the original study's training data to predict clinical outcomes such as cancer development, cancer progression, disease complications, or biomarkers such as pathogenic gene sequences. Following the re-implementation and optimization of the published models, the head-to-head comparison of the ChatGPT ADA-crafted ML models and their respective manually crafted counterparts revealed no significant differences in traditional performance metrics (P>0.071). Strikingly, the ChatGPT ADA-crafted ML models often outperformed their counterparts. In conclusion, ChatGPT ADA offers a promising avenue to democratize ML in medicine by simplifying complex data analyses, yet should enhance, not replace, specialized training and resources, to promote broader applications in medical research and practice.

cs.LG