SearcharxivSearch

arXiv subjects

Geert Molenberghs

Publications and source records attributed to Geert Molenberghs.

18 recordsLinked to original sources

Probability distributions of initial rotation velocities and core-boundary mixing efficiencies of γ Doradus stars

The theory the rotational and chemical evolution is incomplete, thereby limiting the accuracy of model-dependent stellar mass and age determinations. The $γ$ Doradus pulsators are excellent points of calibration for the current state-of-the-art stellar evolution models, as their gravity modes probe the physical conditions in the deep stellar interior. Yet, individual asteroseismic modelling of these stars is not always possible because of insufficient observed oscillation modes. This paper presents a novel method to derive distributions of the stellar mass, age, core-boundary mixing efficiency and initial rotation rates for $γ$ Dor stars. We compute a grid of rotating stellar evolution models covering the entire $γ$ Dor instability strip. We then use the observed distributions of the luminosity, effective temperature, buoyancy travel time and near-core rotation frequency of a sample of 539 stars to assign a statistical weight to each of our models. This weight is a measure of how likely the combination of a specific model is. We then compute weighted histograms to derive the most likely distributions of the fundamental stellar properties. We find that the rotation frequency at zero-age main sequence follows a normal distribution, peaking around 25% of the critical Keplerian rotation frequency. The probability-density function for extent of the core-boundary mixing zone, given by a factor $f_{\rm CBM}$ times the local pressure scale height (assuming an exponentially decaying parameterisation) decreases linearly with increasing $f_{\rm CBM}$. Converting the distribution of fractions of critical rotation at the zero-age main sequence to units of d$^{-1}$, we find most F-type stars start the main sequence with a rotation frequency between 0.5 and 2 d$^{-1}$. Regarding the core-boundary mixing efficiency, we find that it is generally weak in this mass regime.

astro-ph.SR

Astrophysical properties of 15062 Gaia DR3 gravity-mode pulsators: pulsation amplitudes, rotation, and spectral line broadening

Gravito-inertial asteroseismology saw its birth thanks to high-precision CoRoT and Kepler space photometric light curves. So far, it gave rise to the internal rotation frequency of a few hundred intermediate-mass stars, yet only several tens of these have been weighed, sized, and age-dated with high precision from asteroseismic modelling. We aim to increase the sample of optimal targets for future gravito-inertial asteroseismology by assessing the properties of 15062 newly found Gaia DR3 gravity-mode pulsators. We also wish to investigate if there is any connection between their fundamental parameters and dominant mode on the one hand, and their spectral line broadening measured by Gaia on the other hand. After re-classifying about 22% of the F-type gravity-mode pulsators as B-type according to their effective temperature, we construct histograms of the fundamental parameters and mode properties of the 15062 new Gaia DR3 pulsators. We compare these histograms with those of 63 Kepler bona fide class members. We fit errors-in-variables regression models to couple the effective temperature, luminosity, gravity, and oscillation properties to the two Gaia DR3 parameters capturing spectral line broadening for a fraction of the pulsators. We find that the selected 15062 gravity-mode pulsators have properties fully in line with those of their well-known Kepler analogues, revealing that Gaia has a role to play in asteroseismology. The dominant g-mode frequency is a significant predictor of the spectral line broadening for the class members having this quantity measured. We show that the Gaia vbroad parameter captures the joint effect of time-independent intrinsic and rotational line broadening and time-dependent tangential pulsational broadening. Gaia was not desiged to detect non-radial oscillations, yet its homogeneous data treatment allow us to identify many new gravity-mode pulsators.

astro-ph.SR

Identifying predictive biomarkers of CIMAvaxEGF success in advanced Lung Cancer Patients

Objectives: To identify predictive biomarkers of CIMAvaxEGF success in the treatment of Non-Small Cell Lung Cancer Patients. Methods: Data from a clinical trial evaluating the effect on survival time of CIMAvax-EGF versus best supportive care were analyzed retrospectively following the causal inference approach. Pre-treatment potential predictive biomarkers included basal serum EGF concentration, peripheral blood parameters and immunosenescence biomarkers (The proportion of CD8 + CD28- T cells, CD4+ and CD8+ T cells, CD4 CD8 ratio and CD19+ B cells. The 33 patients with complete information were included. The predictive causal information (PCI) was calculated for all possible models. The model with a minimum number of predictors, but with high prediction accuracy (PCI>0.7) was selected. Good, rare and poor responder patients were identified using the predictive probability of treatment success. Results: The mean of PCI increased from 0.486, when only one predictor is considered, to 0.98 using the multivariate approach with all predictors. The model considering the proportion of CD4+ T cell, basal EGF concentration, NLR, Monocytes, and Neutrophils as predictors were selected (PCI>0.74). Patients predicted as good responders according to the pre-treatment biomarkers values treated with CIMAvax-EGF had a significant higher observed survival compared with the control group (p=0.03). No difference was observed for bad responders. Conclusions: Peripheral blood parameters and immunosenescence biomarkers together with basal EGF concentration in serum resulted in good predictors of the CIMAvax-EGF success in advanced NSCLC. The study illustrates the application of a new methodology, based on causal inference, to evaluate multivariate pre-treatment predictors

stat.AP

Asteroseismic masses, ages, and core properties of $γ$ Doradus stars using gravito-inertial dipole modes and spectroscopy

The asteroseismic modelling of period spacing patterns from gravito-inertial modes in stars with a convective core is a high-dimensional problem. We utilise the measured period spacing pattern of prograde dipole gravity modes (acquiring $Π_0$), in combination with the effective temperature ($T_{\rm eff}$) and surface gravity ($\log g$) derived from spectroscopy, to estimate the fundamental stellar parameters and core properties of 37 $γ~$Doradus ($γ~$Dor) stars whose rotation frequency has been derived from $\textit{Kepler}$ photometry. We make use of two 6D grids of stellar models, one with step core overshooting and one with exponential core overshooting, to evaluate correlations between the three observables $Π_0$, $T_{\rm eff}$, and $\log g$ and the mass, age, core overshooting, metallicity, initial hydrogen mass fraction and envelope mixing. We provide multivariate linear model recipes relating the stellar parameters to be estimated to the three observables ($Π_0$, $T_{\rm eff}$, $\log g$). We estimate the (core) mass, age, core overshooting and metallicity of $γ~$Dor stars from an ensemble analysis and achieve relative uncertainties of $\sim\!10$ per cent for the parameters. The asteroseismic age determination allows us to conclude that efficient angular momentum transport occurs already early on during the main sequence. We find that the nine stars with observed Rossby modes occur across almost the entire main-sequence phase, except close to core-hydrogen exhaustion. Future improvements of our work will come from the inclusion of more types of detected modes per star, larger samples, and modelling of individual mode frequencies.

astro-ph.SR

On asymptotic normality in estimation after a group sequential trial

We prove that in many realistic cases, the ordinary sample mean after a group sequential trial is asymptotically normal if the maximal number of observations increases. We derive that it is often safe to use naive confidence intervals for the mean of the collected observations, based on the ordinary sample mean. Our theoretical findings are confirmed by a simulation study.

math.ST

On the sample mean after a group sequential trial

A popular setting in medical statistics is a group sequential trial with independent and identically distributed normal outcomes, in which interim analyses of the sum of the outcomes are performed. Based on a prescribed stopping rule, one decides after each interim analysis whether the trial is stopped or continued. Consequently, the actual length of the study is a random variable. It is reported in the literature that the interim analyses may cause bias if one uses the ordinary sample mean to estimate the location parameter. For a generic stopping rule, which contains many classical stopping rules as a special case, explicit formulas for the expected length of the trial, the bias, and the mean squared error (MSE) are provided. It is deduced that, for a fixed number of interim analyses, the bias and the MSE converge to zero if the first interim analysis is performed not too early. In addition, optimal rates for this convergence are provided. Furthermore, under a regularity condition, asymptotic normality in total variation distance for the sample mean is established. A conclusion for naive confidence intervals based on the sample mean is derived. It is also shown how the developed theory naturally fits in the broader framework of likelihood theory in a group sequential trial setting. A simulation study underpins the theoretical findings.

math.ST

On the asymptotic normality and the construction of confidence intervals for estimators after sampling with probabilistic and deterministic stopping rules

A key feature of a sequential study is that the actual sample size is a random variable that typically depends on the outcomes collected. While hypothesis testing theory for sequential designs is well established, parameter and precision estimation is less well understood. Even though earlier work has established a number of ad hoc estimators to overcome alleged bias in the ordinary sample average, recent work has shown the sample average to be consistent. Building upon these results, by providing a rate of convergence for the total variation distance, it is established that the asympotic distribution of the sample average is normal, in almost all cases, except in a very specific one where the stopping rule is deterministic and the true population mean coincides with the cut-off between stopping and continuing. For this pathological case, the Kolmogorov distance with the normal is found to equal 0.125. While noticeable in the asymptotic distribution, simulations show that there fortunately are no consequences for the coverage of normally-based confidence intervals.

math.ST

On the asymptotic behavior of the contaminated sample mean

An observation of a cumulative distribution function $F$ with finite variance is said to be contaminated according to the inflated variance model if it has a large probability of coming from the original target distribution $F$, but a small probability of coming from a contaminating distribution that has the same mean and shape as $F$, though a larger variance. It is well known that in the presence of data contamination, the ordinary sample mean looses many of its good properties, making it preferable to use more robust estimators. From a didactical point of view, it is insightful to see to what extent an intuitive estimator such as the sample mean becomes less favorable in a contaminated setting. In this paper, we investigate under which conditions the sample mean, based on a finite number of independent observations of $F$ which are contaminated according to the inflated variance model, is a valid estimator for the mean of $F$. In particular, we examine to what extent this estimator is weakly consistent for the mean of $F$ and asymptotically normal. As classical central limit theory is generally inaccurate to cope with the asymptotic normality in this setting, we invoke more general approximate central limit theory as developed by Berckmoes, Lowen, and Van Casteren (2013). Our theoretical results are illustrated by a specific example and a simulation study.

math.ST

Approximate central limit theorems

We refine the classical Lindeberg-Feller central limit theorem by obtaining asymptotic bounds on the Kolmogorov distance, the Wasserstein distance, and the parametrized Prokhorov distances in terms of a Lindeberg index. We thus obtain more general approximate central limit theorems, which roughly state that the row-wise sums of a triangular array are approximately asymptotically normal if the array approximately satisfies Lindeberg's condition. This allows us to continue to provide information in non-standard settings in which the classical central limit theorem fails to hold. Stein's method plays a key role in the development of this theory.

math.PR

A permutational-splitting sample procedure to quantify expert opinion on clusters of chemical compounds using high-dimensional data

Expert opinion plays an important role when selecting promising clusters of chemical compounds in the drug discovery process. We propose a method to quantify these qualitative assessments using hierarchical models. However, with the most commonly available computing resources, the high dimensionality of the vectors of fixed effects and correlated responses renders maximum likelihood unfeasible in this scenario. We devise a reliable procedure to tackle this problem and show, using theoretical arguments and simulations, that the new methodology compares favorably with maximum likelihood, when the latter option is available. The approach was motivated by a case study, which we present and analyze.

stat.AP

The surface nitrogen abundance of a massive star in relation to its oscillations, rotation, and magnetic field

We have composed a sample of 68 massive stars in our galaxy whose projected rotational velocity, effective temperature and gravity are available from high-precision spectroscopic measurements. The additional seven observed variables considered here are their surface nitrogen abundance, rotational frequency, magnetic field strength, and the amplitude and frequency of their dominant acoustic and gravity mode of oscillation. Multiple linear regression to estimate the nitrogen abundance combined with principal components analysis, after addressing the incomplete and truncated nature of the data, reveals that the effective temperature and the frequency of the dominant acoustic oscillation mode are the only two significant predictors for the nitrogen abundance, while the projected rotational velocity and the rotational frequency have no predictive power. The dominant gravity mode and the magnetic field strength are correlated with the effective temperature but have no predictive power for the nitrogen abundance. Our findings are completely based on observations and their proper statistical treatment and call for a new strategy in evaluating the outcome of stellar evolution computations.

astro-ph.SR

Dynamic Predictions with Time-Dependent Covariates in Survival Analysis using Joint Modeling and Landmarking

A key question in clinical practice is accurate prediction of patient prognosis. To this end, nowadays, physicians have at their disposal a variety of tests and biomarkers to aid them in optimizing medical care. These tests are often performed on a regular basis in order to closely follow the progression of the disease. In this setting it is of medical interest to optimally utilize the recorded information and provide medically-relevant summary measures, such as survival probabilities, that will aid in decision making. In this work we present and compare two statistical techniques that provide dynamically-updated estimates of survival probabilities, namely landmark analysis and joint models for longitudinal and time-to-event data. Special attention is given to the functional form linking the longitudinal and event time processes, and to measures of discrimination and calibration in the context of dynamic prediction.

stat.AP

A Family of Generalized Linear Models for Repeated Measures with Normal and Conjugate Random Effects

Non-Gaussian outcomes are often modeled using members of the so-called exponential family. Notorious members are the Bernoulli model for binary data, leading to logistic regression, and the Poisson model for count data, leading to Poisson regression. Two of the main reasons for extending this family are (1) the occurrence of overdispersion, meaning that the variability in the data is not adequately described by the models, which often exhibit a prescribed mean--variance link, and (2) the accommodation of hierarchical structure in the data, stemming from clustering in the data which, in turn, may result from repeatedly measuring the outcome, for various members of the same family, etc. The first issue is dealt with through a variety of overdispersion models, such as, for example, the beta-binomial model for grouped binary data and the negative-binomial model for counts. Clustering is often accommodated through the inclusion of random subject-specific effects. Though not always, one conventionally assumes such random effects to be normally distributed. While both of these phenomena may occur simultaneously, models combining them are uncommon. This paper proposes a broad class of generalized linear models accommodating overdispersion and clustering through two separate sets of random effects. We place particular emphasis on so-called conjugate random effects at the level of the mean for the first aspect and normal random effects embedded within the linear predictor for the second aspect, even though our family is more general. The binary, count and time-to-event cases are given particular emphasis. Apart from model formulation, we present an overview of estimation methods, and then settle for maximum likelihood estimation with analytic--numerical integration. Implications for the derivation of marginal correlations functions are discussed. The methodology is applied to data from a study in epileptic seizures, a clinical trial in toenail infection named onychomycosis and survival data in children with asthma.

stat.ME

Formal and Informal Model Selection with Incomplete Data

Model selection and assessment with incomplete data pose challenges in addition to the ones encountered with complete data. There are two main reasons for this. First, many models describe characteristics of the complete data, in spite of the fact that only an incomplete subset is observed. Direct comparison between model and data is then less than straightforward. Second, many commonly used models are more sensitive to assumptions than in the complete-data situation and some of their properties vanish when they are fitted to incomplete, unbalanced data. These and other issues are brought forward using two key examples, one of a continuous and one of a categorical nature. We argue that model assessment ought to consist of two parts: (i) assessment of a model's fit to the observed data and (ii) assessment of the sensitivity of inferences to unverifiable assumptions, that is, to how a model described the unobserved data given the observed ones.

stat.ME

Analyzing Incomplete Discrete Longitudinal Clinical Trial Data

Commonly used methods to analyze incomplete longitudinal clinical trial data include complete case analysis (CC) and last observation carried forward (LOCF). However, such methods rest on strong assumptions, including missing completely at random (MCAR) for CC and unchanging profile after dropout for LOCF. Such assumptions are too strong to generally hold. Over the last decades, a number of full longitudinal data analysis methods have become available, such as the linear mixed model for Gaussian outcomes, that are valid under the much weaker missing at random (MAR) assumption. Such a method is useful, even if the scientific question is in terms of a single time point, for example, the last planned measurement occasion, and it is generally consistent with the intention-to-treat principle. The validity of such a method rests on the use of maximum likelihood, under which the missing data mechanism is ignorable as soon as it is MAR. In this paper, we will focus on non-Gaussian outcomes, such as binary, categorical or count data. This setting is less straightforward since there is no unambiguous counterpart to the linear mixed model. We first provide an overview of the various modeling frameworks for non-Gaussian longitudinal data, and subsequently focus on generalized linear mixed-effects models, on the one hand, of which the parameters can be estimated using full likelihood, and on generalized estimating equations, on the other hand, which is a nonlikelihood method and hence requires a modification to be valid under MAR. We briefly comment on the position of models that assume missingness not at random and argue they are most useful to perform sensitivity analysis. Our developments are underscored using data from two studies. While the case studies feature binary outcomes, the methodology applies equally well to other discrete-data settings, hence the qualifier ``discrete'' in the title.

math.ST

Estimating stellar oscillation-related parameters and their uncertainties with the moment method

The moment method is a well known mode identification technique in asteroseismology (where `mode' is to be understood in an astronomical rather than in a statistical sense), which uses a time series of the first 3 moments of a spectral line to estimate the discrete oscillation mode parameters l and m. The method, contrary to many other mode identification techniques, also provides estimates of other important continuous parameters such as the inclination angle alpha, and the rotational velocity v_e. We developed a statistical formalism for the moment method based on so-called generalized estimating equations (GEE). This formalism allows the estimation of the uncertainty of the continuous parameters taking into account that the different moments of a line profile are correlated and that the uncertainty of the observed moments also depends on the model parameters. Furthermore, we set up a procedure to take into account the mode uncertainty, i.e., the fact that often several modes (l,m) can adequately describe the data. We also introduce a new lack of fit function which works at least as well as a previous discriminant function, and which in addition allows us to identify the sign of the azimuthal order m. We applied our method to the star HD181558, using several numerical methods, from which we learned that numerically solving the estimating equations is an intensive task. We report on the numerical results, from which we gain insight in the statistical uncertainties of the physical parameters involved in the moment method.

astro-ph