SearcharxivSearch

arXiv subjects

Grazziela Figueredo

Publications and source records attributed to Grazziela Figueredo.

14 recordsLinked to original sources

Standardizing Longitudinal Radiology Report Evaluation via Large Language Model Annotation

Longitudinal information in radiology reports refers to the sequential tracking of findings across multiple examinations over time, which is crucial for monitoring disease progression and guiding clinical decisions. Many recent automated radiology report generation methods are designed to capture longitudinal information; however, validating their performance is challenging. There is no proper tool to consistently label temporal changes in both ground-truth and model-generated texts for meaningful comparisons. Large language models (LLMs) offer a promising annotation alternative, as they are capable of capturing nuanced linguistic patterns and semantic similarities without extensive manual intervention. They also adapt well to new contexts. In this study, we therefore propose an LLM-based pipeline to automatically annotate longitudinal information in radiology reports. The pipeline first identifies sentences containing relevant information and then extracts the progression of diseases. We evaluate and compare five mainstream LLMs on these two tasks using 500 manually annotated reports. Considering both efficiency and performance, Qwen2.5-32B was subsequently selected and used to annotate another 95,169 reports from the public MIMIC-CXR dataset. Our Qwen2.5-32B-annotated dataset provided us with a standardized benchmark for evaluating report generation models. Using this new benchmark, we assessed seven state-of-the-art report generation models. Our LLM-based annotation method outperforms existing annotation solutions, achieving 11.3\% and 5.3\% higher F1-scores for longitudinal information detection and disease tracking, respectively. The source code is available at https://github.com/wxinyi1996/Standardizing-Longitudinal-Chest-X-ray-Report-Evaluation-via-Large-Language-Model-Annotation.git.

cs.CL

T$_2$* and Susceptibility Mapping as Indicators of Placental Health

Objective(s): T$_2$* and susceptibility ($χ$) MRI mapping provide complimentary measures of the haemodynamic environment in the placenta. The aims of this work were to use these simultaneously obtained measures to investigate the role of oxygen distribution on the well-established reduction of T$_2$* with gestational age found in healthy pregnancies and explore differences in both measures in compromised placentas. Methods: T$_2$* and $χ$ were measured simultaneously from a double echo, echo planar scan of the whole placenta, across a range of gestational ages and pregnancy complications. Regional variations across the placenta were investigated. Results: Whole placental mean T$_2$* was more correlated with standard deviation of $χ$ than mean $χ$ indicating it is more driven by increasing local inhomogeneities rather than bulk deoxygenation with healthy gestation. Compromised placentas also showed increased standard deviation of $χ$ as well as lower mean T$_2$* suggesting flow/uptake mismatch and reduced oxygenation. Regionally, the susceptibility was lowest (most oxygenated) and least variable in the central region of the placenta indicating good mixing and refreshment of blood in this area. The susceptibility was highest (most deoxygenated) and most variable at the fetal side, suggesting less effective perfusion in this region. Compromised cases showed the greatest difference on the fetal side for both mean and standard deviation of $χ$. T$_2$* was lowest at the fetal side for healthy and compromised cases but the maternal and central regions better distinguished between the two groups. Conclusion(s): T$_2$* and susceptibility can be mapped simultaneously from a single MRI scan and provide complimentary information about the function of the placenta across healthy gestational development, and as a potential indicator of placental compromise.

physics.med-ph

Llettuce: An Open Source Natural Language Processing Tool for the Translation of Medical Terms into Uniform Clinical Encoding

This paper introduces Llettuce, an open-source tool designed to address the complexities of converting medical terms into OMOP standard concepts. Unlike existing solutions such as the Athena database search and Usagi, which struggle with semantic nuances and require substantial manual input, Llettuce leverages advanced natural language processing, including large language models and fuzzy matching, to automate and enhance the mapping process. Developed with a focus on GDPR compliance, Llettuce can be deployed locally, ensuring data protection while maintaining high performance in converting informal medical terms to standardised concepts.

cs.CL

Placental contractions in uncomplicated pregnancies

In 2020 we first described placental contractions, and we have now undertaken a study to characterise them and seek features that might automatically separate them from uterine contractions. We recruited 36 healthy pregnant women to undergo magnetic resonance imaging (MRI) between 29 and 42 weeks of pregnancy in a single-centre, prospective, observational study. Participants had fetal ultrasound to confirm normal growth. Dynamic MRI was acquired for between 15 and 32 minutes using respiratory triggered, multi-slice, single shot, gradient echo, echo planar imaging covering the whole uterus. All participants had a live birth of a healthy baby weighing over the 10th centile for gestational age and none developed any associated conditions of placental dysfunction e.g. pre-eclampsia, or severe maternal or fetal villous malperfusion on placental histopathology. Any visible contractions were recorded for all participants who completed their MRI scan and placental contractions occurred in at least 60% of our healthy pregnant population with a median frequency of approximately 2 per hour, and a median duration of 2.4 minutes. Contractions involving a decrease in placental volume of >10% were classified as either placental or uterine by visual observation. Placental contractions occurred more frequently than uterine contractions (p=0.0061), were associated with a larger increase in the surface area of the uterine wall not covered by the placenta (p=0.0015), placental sphericity (p<0.0001) and longer duration (p=0.0151). All contractions led to an increase in the MRI parameter R2* in the placenta. There was large variation both between participants and between contractions from the same individual, in terms of time course and contractions features, with no apparent change across the gestational age range studied, although the largest fractional volume changes were detected at early gestation.

physics.med-ph

DF-ACBlurGAN: Structure-Aware Conditional Generation of Internally Repeated Patterns for Biomaterial Microtopography Design

Learning to generate images with internally repeated and periodic structures poses a fundamental challenge for machine learning and computer vision models, which are typically optimised for local texture statistics and semantic realism rather than global structural consistency. This limitation is particularly pronounced in applications requiring strict control over repetition scale, spacing, and boundary coherence, such as microtopographical biomaterial surfaces. In this work, biomaterial design serves as a use case to study conditional generation of repeated patterns under weak supervision and class imbalance. We propose DF-ACBlurGAN, a structure-aware conditional generative adversarial network that explicitly reasons about long-range repetition during training. The approach integrates frequency-domain repetition scale estimation, scale-adaptive Gaussian blurring, and unit-cell reconstruction to balance sharp local features with stable global periodicity. Conditioning on experimentally derived biological response labels, the model synthesises designs aligned with target functional outcomes. Evaluation across multiple biomaterial datasets demonstrates improved repetition consistency and controllable structural variation compared to conventional generative approaches.

cs.CV

Active Learning Driven Materials Discovery for Low Thermal Conductivity Rare-Earth Pyrochlore for Thermal Barrier Coatings

High-Entropy/multicomponent rare-earth oxides (HECs and MCCs) show promise as alternative materials for thermal barrier coatings (TBC) with the ability to tailor properties based on the combination of rare-earth elements present. By enabling the substitution of scarce or supply-risk rare-earths with more readily available alternatives while maintaining comparable material performance, HECs and MCCs offer a valuable path towards alternative TBC material design. However, navigating this search space of compositionally complex materials is both time and resource intensive. In this study, an active learning (AL) framework was employed to identify HEC/MCC materials with a pyrochlore structure, with acceptable thermal conductivity (TC) for TBC applications. The AL framework was applied through a Bayesian optimisation (BO) strategy, coupled with a random forest surrogate model. TC was selected as the optimisation criterion as that is the most basic requirement of TBC materials. Over two iterations of the AL cycle, four compositions were generated and synthesized in the lab for experimental evaluation. The first iteration yielded two single-phase pyrochlores, $(La_{0.29}Nd_{0.36}Gd_{0.36})_2Zr_2O_7$ and $(La_{0.333}Nd_{0.26}Gd_{0.15}Ho_{0.15}Yb_{0.111})_2Zr_2O_7$, with measured thermal conductivities of 2.03 and 1.90 $W/mK$, respectively. The surrogate model predicted a TC of 2.009 $W/mK$ for both compositions, demonstrating it's accuracy for completely new compositions. The second iteration compositions showed dual-phase when synthesized, highlighting the need to take into account phase formation in the AL framework.

cond-mat.mtrl-sci

A Methodological Study on Data Representation for Machine Learning Modelling of Thermal Conductivity of Rare-Earth Oxides

Quantitative structure-activity relationship (QSAR) modelling is widely employed in materials science to predict properties of interest and extract useful descriptors for measured properties. In thermal barrier coatings (TBC), QSAR can significantly shorten the experimental discovery cycle, which can take years. Although machine learning methods are commonly employed for QSAR, their performance depends on the data quality and how instances are represented. Traditional, hand-crafted descriptors based on known material properties are limited to represent materials that share the same basic crystal structure, limited the size of the dataset. By contrast, graph neural networks offer a more expressive representation, encoding atomic positions and bonds in the crystal lattice. In this study, we compare Random Forest (RF) and Gaussian Process (GP) models trained on hand-crafted descriptors from the literature with graph-based representations for high-entropy, rare-earth pyrochlore oxides using the Crystal Graph Convolutional Neural Network (CGCNN). Two different types of augmentation methods are also explored to account for the limited data size, one of which is only applicable to graph-based representations. Our findings show that the CGCNN model substantially outperforms the RF and GP models, underscoring the potential of graph-based representations for enhanced QSAR modelling in TBC research.

cond-mat.mtrl-sci

Helix 1.0: An Open-Source Framework for Reproducible and Interpretable Machine Learning on Tabular Scientific Data

Helix is an open-source, extensible, Python-based software framework to facilitate reproducible and interpretable machine learning workflows for tabular data. It addresses the growing need for transparent experimental data analytics provenance, ensuring that the entire analytical process -- including decisions around data transformation and methodological choices -- is documented, accessible, reproducible, and comprehensible to relevant stakeholders. The platform comprises modules for standardised data preprocessing, visualisation, machine learning model training, evaluation, interpretation, results inspection, and model prediction for unseen data. To further empower researchers without formal training in data science to derive meaningful and actionable insights, Helix features a user-friendly interface that enables the design of computational experiments, inspection of outcomes, including a novel interpretation approach to machine learning decisions using linguistic terms all within an integrated environment. Released under the MIT licence, Helix is accessible via GitHub and PyPI, supporting community-driven development and promoting adherence to the FAIR principles.

cs.LG

A Survey of Deep Learning-based Radiology Report Generation Using Multimodal Data

Automatic radiology report generation can alleviate the workload for physicians and minimize regional disparities in medical resources, therefore becoming an important topic in the medical image analysis field. It is a challenging task, as the computational model needs to mimic physicians to obtain information from multi-modal input data (i.e., medical images, clinical information, medical knowledge, etc.), and produce comprehensive and accurate reports. Recently, numerous works have emerged to address this issue using deep-learning-based methods, such as transformers, contrastive learning, and knowledge-base construction. This survey summarizes the key techniques developed in the most recent works and proposes a general workflow for deep-learning-based report generation with five main components, including multi-modality data acquisition, data preparation, feature learning, feature fusion and interaction, and report generation. The state-of-the-art methods for each of these components are highlighted. Additionally, we summarize the latest developments in large model-based methods and model explainability, along with public datasets, evaluation methods, current challenges, and future directions in this field. We have also conducted a quantitative comparison between different methods in the same experimental setting. This is the most up-to-date survey that focuses on multi-modality inputs and data fusion for radiology report generation. The aim is to provide comprehensive and rich information for researchers interested in automatic clinical report generation and medical image analysis, especially when using multimodal inputs, and to assist them in developing new algorithms to advance the field.

cs.CV

A Unified Framework for Semi-Supervised Image Segmentation and Registration

Semi-supervised learning, which leverages both annotated and unannotated data, is an efficient approach for medical image segmentation, where obtaining annotations for the whole dataset is time-consuming and costly. Traditional semi-supervised methods primarily focus on extracting features and learning data distributions from unannotated data to enhance model training. In this paper, we introduce a novel approach incorporating an image registration model to generate pseudo-labels for the unannotated data, producing more geometrically correct pseudo-labels to improve the model training. Our method was evaluated on a 2D brain data set, showing excellent performance even using only 1\% of the annotated data. The results show that our approach outperforms conventional semi-supervised segmentation methods (e.g. teacher-student model), particularly in a low percentage of annotation scenario. GitHub: https://github.com/ruizhe-l/UniSegReg.

cs.CV

KemenkeuGPT: Leveraging a Large Language Model on Indonesia's Government Financial Data and Regulations to Enhance Decision Making

Data is crucial for evidence-based policymaking and enhancing public services, including those at the Ministry of Finance of the Republic of Indonesia. However, the complexity and dynamic nature of governmental financial data and regulations can hinder decision-making. This study investigates the potential of Large Language Models (LLMs) to address these challenges, focusing on Indonesia's financial data and regulations. While LLMs are effective in the financial sector, their use in the public sector in Indonesia is unexplored. This study undertakes an iterative process to develop KemenkeuGPT using the LangChain with Retrieval-Augmented Generation (RAG), prompt engineering and fine-tuning. The dataset from 2003 to 2023 was collected from the Ministry of Finance, Statistics Indonesia and the International Monetary Fund (IMF). Surveys and interviews with Ministry officials informed, enhanced and fine-tuned the model. We evaluated the model using human feedback, LLM-based evaluation and benchmarking. The model's accuracy improved from 35% to 61%, with correctness increasing from 48% to 64%. The Retrieval-Augmented Generation Assessment (RAGAS) framework showed that KemenkeuGPT achieved 44% correctness with 73% faithfulness, 40% precision and 60% recall, outperforming several other base models. An interview with an expert from the Ministry of Finance indicated that KemenkeuGPT has the potential to become an essential tool for decision-making. These results are expected to improve with continuous human feedback.

cs.AI

MrRegNet: Multi-resolution Mask Guided Convolutional Neural Network for Medical Image Registration with Large Deformations

Deformable image registration (alignment) is highly sought after in numerous clinical applications, such as computer aided diagnosis and disease progression analysis. Deep Convolutional Neural Network (DCNN)-based image registration methods have demonstrated advantages in terms of registration accuracy and computational speed. However, while most methods excel at global alignment, they often perform worse in aligning local regions. To address this challenge, this paper proposes a mask-guided encoder-decoder DCNN-based image registration method, named as MrRegNet. This approach employs a multi-resolution encoder for feature extraction and subsequently estimates multi-resolution displacement fields in the decoder to handle the substantial deformation of images. Furthermore, segmentation masks are employed to direct the model's attention toward aligning local regions. The results show that the proposed method outperforms traditional methods like Demons and a well-known deep learning method, VoxelMorph, on a public 3D brain MRI dataset (OASIS) and a local 2D brain MRI dataset with large deformations. Importantly, the image alignment accuracies are significantly improved at local regions guided by segmentation masks. Github link:https://github.com/ruizhe-l/MrRegNet.

eess.IV

Active learning-driven uncertainty reduction for in-flight particle characteristics of atmospheric plasma spraying of silicon

In this study, the first-of-its-kind use of active learning (AL) framework in thermal spray is adapted to improve the prediction accuracy of the in-flight particle characteristics and uses Gaussian Process (GP) ML model as a surrogate that generalises a global solution without necessarily involving physical mechanisms. The AL framework via the Bayesian Optimisation was utilised to: (a) reduce the maximum uncertainty in the given database and (b) reduce local uncertainty around a contrived test point. The initial dataset consists of 26 atmospheric plasma spray (APS) parameters of silicon, aimed at ceramic matrix composites (CMCs) for the next generation of aerospace applications. The maximum uncertainty in the initial dataset was reduced by AL-driven identification of search spaces and conducting six guided spray trails in the identified search spaces. On average, a 52.9% improvement (error reduction) of RMSE and an R2 increase of 8.5% were reported on the predicted in-flight particle velocities and temperatures after the AL-driven optimisation. Furthermore, the Bayesian Optimisation around a contrived test point to predict the best possible characteristics resulted in a three-fold increase in prediction accuracy as compared to the non-optimised prediction. These AL-guided experimental validations not only increase the informativeness of the limited dataset but is adaptable for other thermal spraying methods without necessarily involving physical mechanisms and underlying mechanisms. The use of AL-driven optimisation may drive the thermal spraying towards resource-efficiency and may serve as the first step towards fully digital thermal spraying environments.

physics.plasm-ph

Towards a More Reliable Interpretation of Machine Learning Outputs for Safety-Critical Systems using Feature Importance Fusion

When machine learning supports decision-making in safety-critical systems, it is important to verify and understand the reasons why a particular output is produced. Although feature importance calculation approaches assist in interpretation, there is a lack of consensus regarding how features' importance is quantified, which makes the explanations offered for the outcomes mostly unreliable. A possible solution to address the lack of agreement is to combine the results from multiple feature importance quantifiers to reduce the variance of estimates. Our hypothesis is that this will lead to more robust and trustworthy interpretations of the contribution of each feature to machine learning predictions. To assist test this hypothesis, we propose an extensible Framework divided in four main parts: (i) traditional data pre-processing and preparation for predictive machine learning models; (ii) predictive machine learning; (iii) feature importance quantification and (iv) feature importance decision fusion using an ensemble strategy. We also introduce a novel fusion metric and compare it to the state-of-the-art. Our approach is tested on synthetic data, where the ground truth is known. We compare different fusion approaches and their results for both training and test sets. We also investigate how different characteristics within the datasets affect the feature importance ensembles studied. Results show that our feature importance ensemble Framework overall produces 15% less feature importance error compared to existing methods. Additionally, results reveal that different levels of noise in the datasets do not affect the feature importance ensembles' ability to accurately quantify feature importance, whereas the feature importance quantification error increases with the number of features and number of orthogonal informative features.

stat.ML