SearcharxivSearch

arXiv subjects

Xiaoyan Song

Publications and source records attributed to Xiaoyan Song.

3 recordsLinked to original sources

Improving O-RADS Risk Stratification from Ultrasound Reports: A Comparative Evaluation of Hybrid versus End-to-End LLM Reasoning Strategies

Background: Automating clinical guideline-based decision-making with large language models (LLMs) remains challenging because of reliability, hallucination, and limited interpretability. We compared the performance of LLMs and reasoning strategies for automated Ovarian-Adnexal Reporting and Data System (O-RADS) classification from free-text pelvic ultrasound reports. Methods: In this retrospective study, consecutive patients with ovarian masses who underwent pelvic ultrasound were included. Eight LLMs were tested with three reasoning strategies: implicit-knowledge end-to-end, rule-informed end-to-end, and a feature-based hybrid architecture that decoupled feature extraction from rule-based classification. The reference standard was O-RADS categorization established by expert consensus. Results: A total of 310 women with 390 ovarian masses were evaluated. The feature-based hybrid architecture using Gemini 3.6 Flash demonstrated the best performance, achieving an accuracy of 99.2% (387 of 390) and almost perfect agreement with the reference standard (weighted kappa = 1.00; 95% CI: 0.99-1.00). Its performance surpassed that of original clinical reports (accuracy, 87.7% [342 of 390]; weighted kappa = 0.94; 95% CI: 0.91-0.96) and end-to-end LLM strategies (accuracy range, 65.6% [256 of 390] to 95.9% [374 of 390]). For structured feature extraction, Gemini 3.6 Flash demonstrated higher overall accuracy than Claude Fable 5 (98.9% vs 97.8%; P < 0.001). The hybrid architecture reduced misclassification errors and mitigated the overstaging tendency observed in original reports. Conclusion: The feature-based hybrid LLM architecture that separates clinical feature extraction from deterministic guideline execution enables highly accurate, reliable, and interpretable automated O-RADS classification, providing a promising approach for standardized, guideline-based clinical decision-making.

cs.AI

An improved implicit sampling for Bayesian inverse problems of multi-term time fractional multiscale diffusion models

This paper presents an improved implicit sampling method for hierarchical Bayesian inverse problems. A widely used approach for sampling posterior distribution is based on Markov chain Monte Carlo (MCMC). However, the samples generated by MCMC are usually strongly correlated. This may lead to a small size of effective samples from a long Markov chain and the resultant posterior estimate may be inaccurate. An implicit sampling method proposed in [11] can generate independent samples and capture some inherent non-Gaussian features of the posterior based on the weights of samples. In the implicit sampling method, the posterior samples are generated by constructing a map and distribute around the MAP point. However, the weights of implicit sampling in previous works may cause excessive concentration of samples and lead to ensemble collapse. To overcome this issue, we propose a new weight formulation and make resampling based on the new weights. In practice, some parameters in prior density are often unknown and a hierarchical Bayesian inference is necessary for posterior exploration. To this end, the hierarchical Bayesian formulation is used to estimate the MAP point and integrated in the implicit sampling framework. Compared to conventional implicit sampling, the proposed implicit sampling method can significantly improve the posterior estimator and the applicability for high dimensional inverse problems. The improved implicit sampling method is applied to the Bayesian inverse problems of multi-term time fractional diffusion models in heterogeneous media. To effectively capture the heterogeneity effect, we present a mixed generalized multiscale finite element method (mixed GMsFEM) to solve the time fractional diffusion models in a coarse grid, which can substantially speed up the Bayesian inversion.

math.NA

Recovering the reaction coefficient for two dimensional time fractional diffusion equations

In this paper, we present an inverse problem of identifying the reaction coefficient for time fractional diffusion equations in two dimensional spaces by using boundary Neumann data. It is proved that the forward operator is continuous with respect to the unknown parameter. Because the inverse problem is often ill-posed, regularization strategies are imposed on the least fit-to-data functional to overcome the stability issue. There may exist various kinds of functions to reconstruct. It is crucial to choose a suitable regularization method. We present a multi-parameter regularization $L^{2}+BV$ method for the inverse problem. This can extend the applicability for reconstructing the unknown functions. Rigorous analysis is carried out for the inverse problem. In particular, we analyze the existence and stability of regularized variational problem and the convergence. To reduce the dimension in the inversion for numerical simulation, the unknown coefficient is represented by a suitable set of basis functions based on a priori information. A few numerical examples are presented for the inverse problem in time fractional diffusion equations to confirm the theoretic analysis and the efficacy of the different regularization methods.

math.NA