SearcharxivSearch

arXiv subjects

Jaroslaw Harezlak

Publications and source records attributed to Jaroslaw Harezlak.

17 recordsLinked to original sources

Joint Nuclear and $\ell_1$ Regularization for Logistic Matrix Regression with Applications to Brain Imaging

We introduce a new convex optimization framework for logistic scalar-on-matrix regression which incorporates nuclear and $\ell_1$ norm penalties to enforce simultaneously low-rank and sparse structures in the estimated coefficient matrix. The proposed method enables interpretable modeling of high-dimensional matrix-valued predictors in the presence of binary responses. We derive a custom algorithm based on the Alternating Direction Method of Multipliers (ADMM) to efficiently solve the resulting convex optimization problem and establish the theoretical properties of the obtained solution. Numerical experiments clearly demonstrate the effectiveness of our method in recovering meaningful predictive patterns. Finally, we apply our method to the brain imaging data to identify structures in functional brain connectivity matrices that are characteristic of subjects with a family history of alcohol use disorders (AUDs).

stat.ME

The Genetic and Environmental Architecture of the Human Functional Connectome

Functional connectivity varies across individuals due to genetic and environmental factors, yet classical twin models typically confound non-shared environment with measurement error and are largely limited to resting-state analyses. We hypothesized that: i) explicitly modeling measurement error from repeated fMRI sessions enables more accurate application of classical twin models (ACE/ADE) to functional connectivity; ii) model applicability depends on scan-length and parcellation granularity; iii) genetic and environmental effects on functional connectomes show differentiated functional modules across conditions. We extended ACE/ADE models to include a repeated-scan derived error term by analyzing monozygotic and dizygotic twins from the Young-Adult Human Connectome Project dataset. Genetic and environment variance components were estimated for all functional couplings across resting-state and task conditions, integrated across conditions using a minimum-error criterion, and analyzed using multilayer community detection across resolution scales. Functional couplings segregated into distinct categories characterized by shared environmental, additive, dominant, or epistatic influences, with a substantial fraction not meeting twin-model assumptions. Integrating across conditions revealed hierarchical community structure in genetic and environmental components observed across community resolution scales. Incorporating measurement error into twin models improves interpretability and applicability at the functional connectome level, revealing that genetic and environmental influences are structured into coherent, multiscale brain networks.

q-bio.NC

A Novel Testing Approach for Differences Among Brain Connectomes

Statistical analysis on non-Euclidean spaces typically relies on distances as the primary tool for constructing likelihoods. However, manifold-valued data admits richer structures in addition to Riemannian distances. We demonstrate that simple, tractable models that do not rely exclusively on distances can be constructed on the manifold of symmetric positive definite (SPD) matrices, which naturally arises in brain connectivity analysis. Specifically, we highlight the manifold-valued Mahalanobis distribution, a parametric family that extends classical multivariate concepts to the SPD manifold. We develop estimators for this distribution and establish their asymptotic properties. Building on this framework, we propose a novel ANOVA test that leverages the manifold structure to obtain a test statistic that better captures the dimensionality of the data. We theoretically demonstrate that our test achieves superior statistical power compared to distance-based Fréchet ANOVA methods.

math.ST

Resting State Functional Connectivity Patterns Associate with Alcohol Use Disorder Characteristics: Insights from the Triple Network Model

Prolonged alcohol use results in neuroadaptations that mark more severe and treatment-resistant alcohol use. The goal of this study was to identify functional connectivity brain patterns underlying Alcohol Use Disorder (AUD)-related characteristics in fifty-five adults (31 female) who endorsed heavy alcohol use. We hypothesized that resting-state functional connectivity (rsFC) of the Salience (SN), Frontoparietal (FPN), and Default Mode (DMN) networks would reflect self-reported recent and lifetime alcohol use, laboratory-based alcohol seeking, urgency, and sociodemographic characteristics related to AUD. To test our hypothesis, we combined the triple network model (TNM) of psychopathology with a multivariate data-driven approach, regularized partial least squares (rPLS), to unfold concurrent functional connectivity (FC) patterns and their association with AUD-related characteristics. We observed three concurrent associations of interest: i) drinking and age-related cross communication between the SN and both the FPN and DMN; ii) family history density of AUD and urgency anticorrelations between the SN and FPN; and iii) alcohol seeking and sex-associated SN and DMN interactions. These findings provide an integrative interpretation for many individual findings reported in the literature relating functional connectivity signatures and AUD factors. Moreover, we identified a set of neural mechanisms and brain regions concomitant with AUD-related characteristics that can serve as potential treatment targets across clinical and preclinical models.

q-bio.NC

Tangent space functional reconfigurations in individuals at risk for alcohol use disorder

Human brain function dynamically adjusts to ever-changing stimuli from the external environment. Studies characterizing brain functional reconfiguration are nevertheless scarce. Here we present a principled mathematical framework to quantify brain functional reconfiguration when engaging and disengaging from a stop signal task (SST). We apply tangent space projection (a Riemannian geometry mapping technique) to transform functional connectomes (FCs) of 54 participants and quantify functional reconfiguration using the correlation distance of the resulting tangent-FCs. Our goal was to compare functional reconfigurations in individuals at risk for alcohol use disorder (AUD). We hypothesized that functional reconfigurations when transitioning to/from a task would be influenced by family history of alcohol use disorder (FHA) and other AUD risk factors. Multilinear regression models showed that engaging and disengaging functional reconfiguration were associated with FHA and recent drinking. When engaging in the SST after a rest condition, functional reconfiguration was negatively associated with recent drinking, while functional reconfiguration when disengaging from the SST was negatively associated with FHA. In both models, several other factors contributed to the functional reconfiguration. This study 1demonstrates that tangent-FCs can characterize task-induced functional reconfiguration, and that it is related to AUD risk.

q-bio.NC

Tangent functional connectomes uncover more unique phenotypic traits

Functional connectomes (FCs) contain pairwise estimations of functional couplings based on pairs of brain regions activity. FCs are commonly represented as correlation matrices that are symmetric positive definite (SPD) lying on or inside the SPD manifold. Since the geometry on the SPD manifold is non-Euclidean, the inter-related entries of FCs undermine the use of Euclidean-based distances. By projecting FCs into a tangent space, we can obtain tangent functional connectomes (tangent-FCs). Tangent-FCs have shown a higher predictive power of behavior and cognition, but no studies have evaluated the effect of such projections with respect to fingerprinting. We hypothesize that tangent-FCs have a higher fingerprint than regular FCs. Fingerprinting was measured by identification rates (ID rates) on test-retest FCs as well as on monozygotic and dizygotic twins. Our results showed that identification rates are systematically higher when using tangent-FCs. Specifically, we found: (i) Riemann and log-Euclidean matrix references systematically led to higher ID rates. (ii) In tangent-FCs, Main-diagonal regularization prior to tangent space projection was critical for ID rate when using Euclidean distance, whereas barely affected ID rates when using correlation distance. (iii) ID rates were dependent on condition and fMRI scan length. (iv) Parcellation granularity was key for ID rates in FCs, as well as in tangent-FCs with fixed regularization, whereas optimal regularization of tangent-FCs mostly removed this effect. (v) Correlation distance in tangent-FCs outperformed any other configuration of distance on FCs or on tangent-FCs across the fingerprint gradient (here sampled by assessing test-retest, Monozygotic and Dizygotic twins). (vi)ID rates tended to be higher in task scans compared to resting-state scans when accounting for fMRI scan length.

q-bio.NC

Geodesic distance on optimally regularized functional connectomes uncovers individual fingerprints

Background: Functional connectomes (FCs), have been shown to provide a reproducible individual fingerprint, which has opened the possibility of personalized medicine for neuro/psychiatric disorders. Thus, developing accurate ways to compare FCs is essential to establish associations with behavior and/or cognition at the individual-level. Methods: Canonically, FCs are compared using Pearson's correlation coefficient of the entire functional connectivity profiles. Recently, it has been proposed that the use of geodesic distance is a more accurate way of comparing functional connectomes, one which reflects the underlying non-Euclidean geometry of the data. Computing geodesic distance requires FCs to be positive-definite and hence invertible matrices. As this requirement depends on the fMRI scanning length and the parcellation used, it is not always attainable and sometimes a regularization procedure is required. Results: In the present work, we show that regularization is not only an algebraic operation for making FCs invertible, but also that an optimal magnitude of regularization leads to systematically higher fingerprints. We also show evidence that optimal regularization is dataset-dependent, and varies as a function of condition, parcellation, scanning length, and the number of frames used to compute the FCs. Discussion: We demonstrate that a universally fixed regularization does not fully uncover the potential of geodesic distance on individual fingerprinting, and indeed could severely diminish it. Thus, an optimal regularization must be estimated on each dataset to uncover the most differentiable across-subject and reproducible within-subject geodesic distances between FCs. The resulting pairwise geodesic distances at the optimal regularization level constitute a very reliable quantification of differences between subjects.

q-bio.NC

Towards an information theoretical description of communication in brain networks

Modeling communication dynamics in the brain is a key challenge in network neuroscience. We present here a framework that combines two measurements for any system where different communication processes are taking place on top of a fixed structural topology: Path Processing Score (PPS) estimates how much the brain signal has changed or has been transformed between any two brain regions (source and target); Path Broadcasting Strength (PBS) estimates the propagation of the signal through edges adjacent to the path being assessed. We use PPS and PBS to explore communication dynamics in large-scale brain networks. We show that brain communication dynamics can be divided into three main 'communication regimes' of information transfer: absent communication (no communication happening); relay communication (information is being transferred almost intact); transducted communication (the information is being transformed). We use PBS to categorize brain regions based on the way they broadcast information. Subcortical regions are mainly direct broadcasters to multiple receivers; Temporal and frontal nodes mainly operate as broadcast relay brain stations; Visual and somato-motor cortices act as multi-channel transducted broadcasters. This work paves the way towards the field of brain network information theory by providing a principled methodology to explore communication dynamics in large-scale brain networks.

q-bio.NC

A Sparsity Inducing Nuclear-Norm Estimator (SpINNEr) for Matrix-Variate Regression in Brain Connectivity Analysis

Classical scalar-response regression methods treat covariates as a vector and estimate a corresponding vector of regression coefficients. In medical applications, however, regressors are often in a form of multi-dimensional arrays. For example, one may be interested in using MRI imaging to identify which brain regions are associated with a health outcome. Vectorizing the two-dimensional image arrays is an unsatisfactory approach since it destroys the inherent spatial structure of the images and can be computationally challenging. We present an alternative approach - regularized matrix regression - where the matrix of regression coefficients is defined as a solution to the specific optimization problem. The method, called SParsity Inducing Nuclear Norm EstimatoR (SpINNEr), simultaneously imposes two penalty types on the regression coefficient matrix---the nuclear norm and the lasso norm---to encourage a low rank matrix solution that also has entry-wise sparsity. A specific implementation of the alternating direction method of multipliers (ADMM) is used to build a fast and efficient numerical solver. Our simulations show that SpINNEr outperforms other methods in estimation accuracy when the response-related entries (representing the brain's functional connectivity) are arranged in well-connected communities. SpINNEr is applied to investigate associations between HIV-related outcomes and functional connectivity in the human brain.

stat.ME

GEFF: Graph Embedding for Functional Fingerprinting

It has been well established that Functional Connectomes (FCs), as estimated from functional MRI (fMRI) data, have an individual fingerprint that can be used to identify an individual from a population (subject-identification). Although identification rate is high when using resting-state FCs, other tasks show moderate to low values. Furthermore, identification rate is task-dependent, and is low when distinct cognitive states, as captured by different fMRI tasks, are compared. Here we propose an embedding framework, GEFF (Graph Embedding for Functional Fingerprinting), based on group-level decomposition of FCs into eigenvectors. GEFF creates an eigenspace representation of a group of subjects using one or more task FCs (Learning Stage). In the Identification Stage, we compare new instances of FCs from the Learning subjects within this eigenspace (validation dataset). The validation dataset contains FCs either from the same tasks as the Learning dataset or from the remaining tasks that were not included in Learning. Assessment of validation FCs within the eigenspace results in significantly increased subject-identification rates for all fMRI tasks tested and potentially task-independent fingerprinting process. It is noteworthy that combining resting-state with one fMRI task for GEFF Learning Stage covers most of the cognitive space for subject identification. In addition to subject-identification, GEFF was also used for identification of cognitive states, i.e. to identify the task associated to a given FC, regardless of the subject being already in the Learning dataset or not (subject-independent task-identification). In addition, we also show that eigenvectors from the Learning Stage can be characterized as task-dominant, subject dominant or neither, providing a deeper insight into the extent of variance in functional connectivity across individuals and cognitive states.

q-bio.NC

Brain perfusion mediates the relationship between miRNA levels and postural control

Transcriptomics, regional cerebral blood flow (rCBF), and a spatial motor virtual reality task were integrated using mediation analysis in a novel demonstration of "imaging omics". Data collected in NCAA Division I football athletes cleared for play before in-season training showed significant relationships in a) elevated levels of miR-30d and miR-92a to elevated putamen rCBF, (b) elevated putamen rCBF to compromised balance scores, and (c) compromised balance scores to elevated miRNA levels. rCBF acted as a mediator variable (minimum 70% mediation, significant Sobel's test) between abnormal miRNA levels and compromised balance scores. Given the involvement of these miRNAs in inflammation and immune function, and that vascular perfusion is a component of the inflammatory response, these findings support a chronic inflammatory model of repetitive head acceleration events (HAEs). rCBF, a systems biology measure, was necessary for miRNA to affect behavior. These results suggest miRNA as a potential diagnostic biomarker for repetitive HAEs.

q-bio.NC

Prediction of sustained harmonic walking in the free-living environment using raw accelerometry data

Objective. Using raw, sub-second level, accelerometry data, we propose and validate a method for identifying and characterizing walking in the free-living environment. We focus on the sustained harmonic walking (SHW), which we define as walking for at least 10 seconds with low variability of step frequency. Approach. We utilize the harmonic nature of SHW and quantify local periodicity of the tri-axial raw accelerometry data. We also estimate fundamental frequency of observed signals and link it to the instantaneous walking (step-to-step) frequency (IWF). Next, we report total time spent in SHW, number and durations of SHW bouts, time of the day when SHW occurred and IWF for 49 healthy, elderly individuals. Main results. Sensitivity of the proposed classification method was found to be 97%, while specificity ranged between 87% and 97% and prediction accuracy between 94% and 97%. We report total time in SHW between 140 and 10 minutes-per-day distributed between 340 and 50 bouts. We estimate the average IWF to be 1.7 steps-per-second. Significance. We propose a simple approach for detection of SHW and estimation of IWF, based on Fourier decomposition. The resulting approach is fast and allows processing of a week-long raw accelerometry data (approx. 150 million measurements) in relatively short time (~half an hour) on a common laptop computer (2.8 GHz Intel Core i7, 16 GB DDR3 RAM).

stat.AP

Validation of gait characteristics extracted from raw accelerometry during walking against measures of physical function, mobility, fatigability, and fitness

Background. Wearable accelerometry devices allow collection of high-density activity data in large epidemiological studies both in-the-lab as well as in-the-wild (free-living). Such data can be used to detect and identify periods of sustained harmonic walking. This report aims to establish whether the micro- and macro-features of walking identified in the laboratory and free-living environments are associated with measures of physical function, mobility, fatigability, and fitness. Methods. Fifty-one older adults (median age 77.5) enrolled in the Developmental Epidemiologic Cohort Study in Pittsburgh, Pennsylvania were included in the analyses. The study included an in-the-lab component as well as 7 days of monitoring in-the-wild. Participants were equipped with hip-worn Actigraph GT3X+ activity monitors, which collect high-density raw accelerometry data. We applied a walking identification algorithm to the data and defined features of walking, such as participant-specific walking acceleration and cadence. The association between these walking features and physical function, mobility, fatigability, and fitness was quantified using linear regression analysis. Results. Micro-scale features of walking (acceleration and cadence) estimated from in-the-lab and in-the-wild data were associated with measures of physical function, mobility, fatigability, and fitness. In-the-lab median walking acceleration was strongly inversely associated with physical function, mobility, fatigability and fitness. Additionally, in-the-wild daily walking time was inversely associated with usual- and fast-paced 400m walking time. Conclusions. The proposed accelerometry-derived walking features are significantly associated with measures of physical function, mobility, fatigability, and fitness, which provides evidence of convergent validity.

q-bio.QM

Stride variability measures derived from wrist- and hip-worn accelerometers

Many epidemiological and clinical studies use accelerometry to objectively measure physical activity using the activity counts, vector magnitude, or number of steps. These measures use just a fraction of the information in the raw accelerometry data as they are typically summarized at the minute level. To address this problem we define and estimate two gait measures of temporal stride-to-stride variability based on raw accelerometry data: Amplitude Deviation (AD) and Phase Deviation (PD). We explore the sensitivity of our approach to on-body placement of the accelerometer by comparing hip, left and right wrist placements. We illustrate the approach by estimating AD and PD in 46 elderly participants in the Developmental Epidemiologic Cohort Study (DECOS) who worn accelerometers during a 400 meter walk test. We also show that AD and PD have a statistically significant association with the gait speed and sit-to-stand test performance

q-bio.NC

Regression Trees for Longitudinal Data

While studying response trajectory, often the population of interest may be diverse enough to exist distinct subgroups within it and the longitudinal change in response may not be uniform in these subgroups. That is, the timeslope and/or influence of covariates in longitudinal profile may vary among these different subgroups. For example, Raudenbush (2001) used depression as an example to argue that it is incorrect to assume that all the people in a given population would be experiencing either increasing or decreasing levels of depression. In such cases, traditional linear mixed effects model (assuming common parametric form for covariates and time) is not directly applicable for the entire population as a group-averaged trajectory can mask important subgroup differences. Our aim is to identify and characterize longitudinally homogeneous subgroups based on the combination of baseline covariates in the most parsimonious way. This goal can be achieved via constructing regression tree for longitudinal data using baseline covariates as partitioning variables. We have proposed LongCART algorithm to construct regression tree for the longitudinal data. In each node, the proposed LongCART algorithm determines the need for further splitting (i.e. whether parameter(s) of longitudinal profile is influenced by any baseline attributes) via parameter instability tests and thus the decision of further splitting is type-I error controlled. We have obtained the asymptotic results for the proposed instability test and examined finite sample behavior of the whole algorithm through simulation studies. Finally, we have applied the LongCART algorithm to study the longitudinal changes in choline level among HIV patients.

stat.ME

Longitudinal Functional Models with Structured Penalties

This paper addresses estimation in a longitudinal regression model for association between a scalar outcome and a set of longitudinally-collected functional covariates or predictor curves. The framework consists of estimating a time-varying coefficient function that is modeled as a linear combination of time-invariant functions but having time-varying coefficients. The estimation procedure exploits the equivalence between penalized least squares estimation and a linear mixed model representation. The process is empirically evaluated with several simulations and it is applied to analyze the neurocognitive impairment of HIV patients and its association with longitudinally-collected magnetic resonance spectroscopy curves.

stat.AP

Structured penalties for functional linear models---partially empirical eigenvectors for regression

One of the challenges with functional data is incorporating spatial structure, or local correlation, into the analysis. This structure is inherent in the output from an increasing number of biomedical technologies, and a functional linear model is often used to estimate the relationship between the predictor functions and scalar responses. Common approaches to the ill-posed problem of estimating a coefficient function typically involve two stages: regularization and estimation. Regularization is usually done via dimension reduction, projecting onto a predefined span of basis functions or a reduced set of eigenvectors (principal components). In contrast, we present a unified approach that directly incorporates spatial structure into the estimation process by exploiting the joint eigenproperties of the predictors and a linear penalty operator. In this sense, the components in the regression are `partially empirical' and the framework is provided by the generalized singular value decomposition (GSVD). The GSVD clarifies the penalized estimation process and informs the choice of penalty by making explicit the joint influence of the penalty and predictors on the bias, variance, and performance of the estimated coefficient function. Laboratory spectroscopy data and simulations are used to illustrate the concepts.

stat.AP