SearcharxivSearch

arXiv subjects

Eeva Vilkkumaa

Publications and source records attributed to Eeva Vilkkumaa.

3 recordsLinked to original sources

Convex Regularization and Convergence of Policy Gradient Flows under Safety Constraints

This paper examines reinforcement learning (RL) in infinite-horizon decision processes with almost-sure safety constraints, crucial for applications like autonomous systems, finance, and resource management. We propose a doubly-regularized RL framework combining reward and parameter regularization to address safety constraints in continuous state-action spaces. The problem is formulated as a convex regularized objective with parametrized policies in the mean-field regime. Leveraging mean-field theory and Wasserstein gradient flows, policies are modeled on an infinite-dimensional statistical manifold, with updates governed by parameter distribution gradient flows. Key contributions include solvability conditions for safety-constrained problems, smooth bounded approximations for gradient flows, and exponential convergence guarantees under sufficient regularization. General regularization conditions, including entropy regularization, support practical particle method implementations. This framework provides robust theoretical insights and guarantees for safe RL in complex, high-dimensional settings.

cs.LG

Predicting Visit Cost of Obstructive Sleep Apnea using Electronic Healthcare Records with Transformer

Background: Obstructive sleep apnea (OSA) is growing increasingly prevalent in many countries as obesity rises. Sufficient, effective treatment of OSA entails high social and financial costs for healthcare. Objective: For treatment purposes, predicting OSA patients' visit expenses for the coming year is crucial. Reliable estimates enable healthcare decision-makers to perform careful fiscal management and budget well for effective distribution of resources to hospitals. The challenges created by scarcity of high-quality patient data are exacerbated by the fact that just a third of those data from OSA patients can be used to train analytics models: only OSA patients with more than 365 days of follow-up are relevant for predicting a year's expenditures. Methods and procedures: The authors propose a method applying two Transformer models, one for augmenting the input via data from shorter visit histories and the other predicting the costs by considering both the material thus enriched and cases with more than a year's follow-up. Results: The two-model solution permits putting the limited body of OSA patient data to productive use. Relative to a single-Transformer solution using only a third of the high-quality patient data, the solution with two models improved the prediction performance's $R^{2}$ from 88.8% to 97.5%. Even using baseline models with the model-augmented data improved the $R^{2}$ considerably, from 61.6% to 81.9%. Conclusion: The proposed method makes prediction with the most of the available high-quality data by carefully exploiting details, which are not directly relevant for answering the question of the next year's likely expenditure.

cs.AI

Decision Programming for optimizing the Finnish Colorectal Cancer Screening Program

In Finland colorectal cancer (CRC) incidence rates have steadily increased over the last decades and as of 2017, CRC is the sixth most common cause of death. CRC is a crucial concern for the public health of Finland. We optimize the faecal immunochemical test (FIT) cut-off level for specified target populations with regard to minimization of cancer prevalence with Decision Programming, which is a novel approach to solving discrete multi-stage decision problems under uncertainty. The results present optimal cut-off levels for Finnish target groups with different colonoscopy capacity constraints. Finally, the estimated resulting cancer prevalence, the amount of required colonoscopies and third-party payer costs resulting from found strategies to the newly implemented Finnish CRC Screening Programme, which have not previously been published, are presented.

math.OC