SearcharxivSearch

arXiv subjects

Marcel Haas

Publications and source records attributed to Marcel Haas.

8 recordsLinked to original sources

Translation as a Computationally Efficient Bridge: Feasibility of English BERT for Low-Resource Languages

BERT models have revolutionised Natural Language Processing (NLP) through their ability to process unstructured text across diverse domains. However, developing high-quality BERT models for non-English languages remains challenging due to limited annotated data and high computational demands. Translating non-English data into English and fine-tuning existing English BERT models offers a resource-efficient alternative, yet few studies have structurally compared translation-based fine-tuning with native-language BERT performance across tasks and languages. This study provides such a comparison, evaluating the feasibility of translation-based fine-tuning across six NLP tasks: Sentiment Analysis, Hate Speech Detection, Question Answering, Named Entity Recognition, Part-of-Speech Tagging, and Natural Language Inference, using datasets translated from Bulgarian, Chinese, Dutch, Italian, and Russian. Across all settings, the translation-based approach was comparable or superior in 53.3 percent of cases. Gains were most frequent in Question Answering, Part-of-Speech Tagging, and Natural Language Inference, while performance declines were common in Named Entity Recognition and Hate Speech Detection. The results show that translation-based fine-tuning is most effective for tasks relying on syntactic or structural patterns and for languages typologically close to English, such as Dutch, but less effective for token-level or culturally nuanced tasks, particularly in Chinese. Overall, this study demonstrates that translation-based fine-tuning offers a scalable, resource-efficient, and empirically validated path for extending NLP to low-resource languages while advancing linguistic inclusivity and sustainability in artificial intelligence.

cs.CL

XGenBoost: Synthesizing Small and Large Tabular Datasets with XGBoost

Tree ensembles such as XGBoost are often preferred for discriminative tasks in mixed-type tabular data, due to their inductive biases, minimal hyperparameter tuning, and training efficiency. We argue that these qualities, when leveraged correctly, can make for better generative models as well. As such, we present XGenBoost, a pair of generative models based on XGBoost: i) a Denoising Diffusion Implicit Model (DDIM) with XGBoost as score-estimator suited for smaller datasets, and ii) a hierarchical autoregressive model whose conditionals are learned via XGBoost classifiers, suited for large-scale tabular synthesis. The architectures follow from the natural constraints imposed by tree-based learners, e.g., in the diffusion model, combining Gaussian and multinomial diffusion to leverage native categorical splits and avoid one-hot encoding while accurately modelling mixed data types. In the autoregressive model, we use a fixed-order factorization, a hierarchical classifier to impose ordinal inductive biases when modelling numerical features, and de-quantization based on empirical quantile functions to model the non-continuous nature of most real-world tabular datasets. Through two benchmarks, one containing smaller and the other larger datasets, we show that our proposed architectures outperform previous neural- and tree-based generative models for mixed-type tabular synthesis at lower training cost.

cs.LG

OpenExtract: Automated Data Extraction for Systematic Reviews in Health

This study presents OpenExtract, an open-source pipeline for automated data extraction in large-scale systematic literature reviews. The pipeline queries large language models (LLMs) to predict data entries based on relevant sections of scientific articles. To test the efficacy of OpenExtract, we apply it to a systematic literature review in digital health and compare its outputs with those of human researchers. OpenExtract achieves precision and recall scores of > 0.8 in this task, indicating that it can be effective at extracting data automatically and efficiently. OpenExtract: https://github.com/JimAchterbergLUMC/OpenExtract.

cs.IR

The Data Sharing Paradox of Synthetic Data in Healthcare

Synthetic data offers a promising solution to privacy concerns in healthcare by generating useful datasets in a privacy-aware manner. However, although synthetic data is typically developed with the intention of sharing said data, ambiguous reidentification risk assessments often prevent synthetic data from seeing the light of day. One of the main causes is that privacy metrics for synthetic data, which inform on reidentification risks, are not well-aligned with practical requirements and regulations regarding data sharing in healthcare. This article discusses the paradoxical situation where synthetic data is designed for data sharing but is often still restricted. We also discuss how the field should move forward to mitigate this issue.

cs.DB

A Novel Taxonomy for Navigating and Classifying Synthetic Data in Healthcare Applications

Data-driven technologies have improved the efficiency, reliability and effectiveness of healthcare services, but come with an increasing demand for data, which is challenging due to privacy-related constraints on sharing data in healthcare contexts. Synthetic data has recently gained popularity as potential solution, but in the flurry of current research it can be hard to oversee its potential. This paper proposes a novel taxonomy of synthetic data in healthcare to navigate the landscape in terms of three main varieties. Data Proportion comprises different ratios of synthetic data in a dataset and associated pros and cons. Data Modality refers to the different data formats amenable to synthesis and format-specific challenges. Data Transformation concerns improving specific aspects of a dataset like its utility or privacy with synthetic data. Our taxonomy aims to help researchers in the healthcare domain interested in synthetic data to grasp what types of datasets, data modalities, and transformations are possible with synthetic data, and where the challenges and overlaps between the varieties lie.

cs.CY

The Astropy Problem

The Astropy Project (http://astropy.org) is, in its own words, "a community effort to develop a single core package for Astronomy in Python and foster interoperability between Python astronomy packages." For five years this project has been managed, written, and operated as a grassroots, self-organized, almost entirely volunteer effort while the software is used by the majority of the astronomical community. Despite this, the project has always been and remains to this day effectively unfunded. Further, contributors receive little or no formal recognition for creating and supporting what is now critical software. This paper explores the problem in detail, outlines possible solutions to correct this, and presents a few suggestions on how to address the sustainability of general purpose astronomical software.

astro-ph.IM

Stellar Signatures of AGN Jet Triggered Star Formation

To investigate feedback between relativistic jets emanating from Active Galactic Nuclei (AGN) and the stellar population of the host galaxy, we analyze the long-term evolution of the galaxy-scale simulations by Gaibler et al. (2012) of jets in massive, gas-rich galaxies at z ~ 2 - 3 and of stars formed in the host galaxies. We find strong, jet-induced differences in the resulting stellar populations of galaxies that host relativistic jets and galaxies that do not, including correlations in stellar locations, velocities, and ages. Jets are found to generate distributions of increased radial and vertical velocities that persist long enough to effectively extend the stellar structure of the host. The jets cause the formation of bow shocks that move out through the disk, generating rings of star formation within the disk. The bow shock often accelerates pockets of gas in which stars form, yielding populations of stars with significant radial and vertical velocities, some of which have large enough velocities to escape the galaxy. These stellar population signatures can serve to identify past jet activity as well as jet-induced star formation.

astro-ph.GA

Observational evidence for a truncation of the star cluster initial mass function at the high mass end

We present the luminosity function (LF) of star clusters in M51 based on HST/ACS observations taken as part of the Hubble Heritage project. The clusters are selected based on their size and with the resulting 5990 clusters we present one of the largest cluster samples of a single galaxy. We find that the LF can be approximated with a double power-law distribution with a break around M_V = -8.9. On the bright side the index of the power-law distribution is steeper (a = 2.75) than on the faint-side (a = 1.93), similar to what was found earlier for the ``Antennae'' galaxies. The location of the bend, however, occurs about 1.6 mag fainter in M51. We confront the observed LF with the model for the evolution of integrated properties of cluster populations of Gieles et al., which predicts that a truncated cluster initial mass function would result in a bend in, and a double power-law behaviour of, the integrated LF. The combination of the large field-of view and the high star cluster formation rate of M51 make it possible to detect such a bend in the LF. Hence, we conclude that there exists a fundamental upper limit to the mass of star clusters in M51. Assuming a power-law cluster initial mass function with exponentional cut-off of the form NdM ~ M^-b * exp(-M/M_C)dM, we find that M_C = 10^5 M_sun. A direct comparison with the LF of the ``Antennae'' suggests that there M_C = 4*10^5 M_sun.

astro-ph