SearcharxivSearch

arXiv subjects

Rutvik H. Desai

Publications and source records attributed to Rutvik H. Desai.

8 recordsLinked to original sources

Lesioned Multimodal Language Models Reproduce Aphasic Picture-Naming Patterns

Aphasia following stroke commonly produces systematic naming errors with characteristic profiles, but whether general-purpose language models not designed for clinical simulation can reproduce these patterns remains untested. We investigated (1) whether lesions or controlled perturbations to a multimodal language model can reproduce different types of errors in picture naming, and (2) whether the framework can reproduce the complete error profile of individual persons with aphasia (PWAs). Using LLaVA 1.6, we evaluated perturbation configurations that varied the layer, proportion, and amount of noise applied to model units. We examined 278 PWAs on the Philadelphia Naming Test, classifying responses into seven categories using a validated neural classifier. Six of seven response categories (correct, semantic, mixed, unrelated, neologism, no response errors) emerged at clinically-comparable proportions across distinct parameter space regions, with formal paraphasia being the exception. Searching the perturbation space revealed configurations that reproduced the individual error profile in at least six of seven categories for 97.8% of PWAs and in all seven categories for 79.5% of PWAs. Monte Carlo baselines confirmed that this matching reflects joint inter-category structure rather than marginal overlap. These results establish a quantitative framework for reproducing individual aphasic error patterns in picture naming. They suggest the potential for language models to serve as digital twins of individuals with post-stroke aphasia.

cs.AI

Recovering Lesion Parameters from Aphasic Picture Naming Error Profiles in Large Language Models

Interpretability methods for large language models (LLMs) describe internal state but do not directly test whether that state is causally sufficient to produce the observed behavior. In earlier work, we lesioned LLMs to produce error profiles in picture naming, a central task for assessing aphasia, and found that specific lesions produced errors resembling those of individual stroke survivors. Here we ask the inverse question: given an error profile, can the lesion parameters that produced it be recovered, and what does this inverse problem reveal about transformer computation? Lesions in LLaVA-Vicuna 13B were parameterized by layer index, modification percentage, and noise sigma across 4,840 configurations, and error profiles were characterized by a seven-category clinical taxonomy (correct, semantic, unrelated, formal, mixed, neologism, no-response). We trained a multi-task neural network to map error profiles back to perturbation parameters. The problem admitted a partial solution: across 10 independently trained inverse models, modification percentage and noise sigma were recoverable, whereas layer index was recoverable only within a neighborhood. In counterfactual validation, a fresh model instance perturbed with the recovered parameters reproduced the target behavior in 81.4% of cases. This dissociation between low layer recovery and high counterfactual fidelity is consistent with functional redundancy across transformer layers, a property not captured by standard interpretability methods. As an out-of-distribution test, we applied the trained model to picture-naming error profiles from 278 stroke survivors; recovered parameters were syndrome-discriminative, most strongly for perturbation intensity, indicating generalization beyond the training distribution. Counterfactual validation provides a general framework for LLM interpretability claims beyond inverse mapping.

cs.CL

Model Selection with Regression and Representational Similarity Analysis for Linear and Nonlinear Data

In cognitive psychology and neuroscience, adjudicating between competing theoretical models is a common methodological challenge. Researchers often rely on either first-order direct mapping approaches (e.g., linear regression) or second-order abstraction methods (e.g., Representational Similarity Analysis [RSA]). However, it remains unclear whether or how the nature of the underlying data and feature characteristics affect the performance of these methods. Here, we systematically evaluated regression, RSA, and Pattern Component Modeling (PCM) across distinct data-generating schemes, including first-order linear mappings and geometry-to-first-order transformations with either linear or nonlinear sigmoid readouts, using both univariate behavioral and multivariate fMRI spatial-pattern simulations. Our results suggest that the relative performance of these methods depends on the underlying generative mechanism. Under linear generative assumptions, regression and PCM showed higher model-selection accuracy than RSA. Under nonlinear but order-preserving transformations, rank-based RSA showed an advantage over regression and PCM. We also found that feature multicollinearity affected these methods differently across generative schemes, and that orthogonalizing the predictor space via principal component analysis (PCA) reduced several collinearity-related differences. Finally, analyses of empirical datasets were consistent with the simulation results under approximately linear conditions, with regression showing clearer model discrimination than RSA. Overall, these findings suggest that the relative performance of regression, RSA, and PCM depends on the form of the mapping between features and responses, as well as on the structure of the feature space.

stat.ME

Topological inference on brain networks with application to lesion symptom mapping

Persistent homology (PH) characterizes the shape of brain networks through persistence features. Group comparison of persistence features from brain networks can be challenging as they are inherently heterogeneous. A recent scale-space representation of persistence diagrams (PDs) through heat diffusion reparameterizes them using a finite number of Fourier coefficients with respect to the Laplace--Beltrami (LB) eigenfunction expansion of the domain, providing a powerful vectorized algebraic representation for group comparisons. In this study, we develop a transposition-based permutation test for comparing multiple groups of PDs using heat-diffusion estimates. We evaluate the empirical performance of the spectral transposition test in capturing within- and between-group similarity and dissimilarity under varying levels of topological noise and cycle location variability. In application, we propose a topological lesion symptom mapping (TLSM) method based on the proposed framework. The method is applied to resting-state functional brain networks of individuals with post-stroke aphasia to identify characteristic cycles associated with varying levels of speech-language impairment.

stat.ME

Multifaceted neural representation of words in naturalistic language

Understanding how the brain represents the multifaceted properties of words in context is essential for explaining the neural architecture of human language. Here, we combine large-scale psycholinguistic modeling with naturalistic fMRI to uncover the latent structure of word properties and their neural representations during narrative comprehension. By analyzing 106 psycholinguistic variables across 13,850 English words, we identified eight interpretable latent dimensions spanning lexical usage, word form, phonology orthography mapping, sublexical regularity, and semantic organization. These factors robustly predicted behavioral performance across lexical decision, naming, recognition, and semantic judgment tasks, demonstrating their cognitive relevance. Parcel-based and multivariate fMRI analyses of narrative listening revealed that these latent dimensions are encoded in overlapping yet functionally differentiated cortical systems. Multidimensional scaling and hierarchical clustering analyses further identified four interacting subsystems supporting sensorimotor grounding, controlled semantic retrieval, resolution of lexical competition, and contextual episodic integration. Together, these findings provide a unified neurocognitive framework linking fundamental lexical psycholinguistic dimensions to distributed cortical systems engaged during naturalistic language comprehension.

q-bio.NC

Topological inference on brain networks across subtypes of post-stroke aphasia

Persistent homology (PH) characterizes the shape of brain networks through the persistence features. Group comparison of persistence features from brain networks can be challenging as they are inherently heterogeneous. A recent scale-space representation of persistence diagram (PD) through heat diffusion reparameterizes using the finite number of Fourier coefficients with respect to the Laplace-Beltrami (LB) eigenfunction expansion of the domain, which provides a powerful vectorized algebraic representation for group comparisons of PDs. In this study, we advance a transposition-based permutation test for comparing multiple groups of PDs through the heat-diffusion estimates of the PDs. We evaluate the empirical performance of the spectral transposition test in capturing within- and between-group similarity and dissimilarity with respect to statistical variation of topological noise and hole location. We also illustrate how the method extends naturally into a clustering scheme by subtyping individuals with post-stroke aphasia through the PDs of their resting-state functional brain networks.

stat.ME

The Two Word Test: A Semantic Benchmark for Large Language Models

Large Language Models (LLMs) have shown remarkable abilities recently, including passing advanced professional exams and demanding benchmark tests. This performance has led many to suggest that they are close to achieving humanlike or 'true' understanding of language, and even Artificial General Intelligence (AGI). Here, we provide a new open-source benchmark that can assess semantic abilities of LLMs using two-word phrases using a task that can be performed relatively easily by humans without advanced training. Combining multiple words into a single concept is a fundamental aspect of human language and intelligence. The test requires meaningfulness judgments of 1768 noun-noun combinations that have been rated as meaningful (e.g., baby boy) or not meaningful (e.g., goat sky). by 150 human raters. We provide versions of the task that probe meaningfulness ratings on a 0-4 scale as well as binary judgments. We conducted a series of experiments using the TWT on GPT-4, GPT-3.5, and Bard, with both versions. Results demonstrated that, compared to humans, all models perform poorly at rating meaningfulness of these phrases. GPT-3.5 and Bard are also unable to make binary discriminations between sensible and nonsense phrases as making sense. GPT-4 makes a substantial improvement in binary discrimination of combinatorial phrases but is still significantly worse than human performance. The TWT can be used to understand the limitations and weaknesses of current LLMs, and potentially improve them. The test also reminds us that caution is warranted in attributing 'true understanding' or AGI to LLMs. TWT is available at: https://github.com/NickRiccardi/two-word-test

cs.CL

Network-based Statistics Distinguish Anomic and Broca Aphasia

Aphasia is a speech-language impairment commonly caused by damage to the left hemisphere. Due to the complexity of speech-language processing, the neural mechanisms that underpin various symptoms between different types of aphasia are still not fully understood. We used the network-based statistic method to identify distinct subnetwork(s) of connections differentiating the resting-state functional networks of the anomic and Broca groups. We identified one such subnetwork that mainly involved the brain regions in the premotor, primary motor, primary auditory, and primary sensory cortices in both hemispheres. The majority of connections in the subnetwork were weaker in the Broca group than the anomic group. The network properties of the subnetwork were examined through complex network measures, which indicated that the regions in the superior temporal gyrus and auditory cortex bilaterally exhibit intensive interaction, and primary motor, premotor and primary sensory cortices in the left hemisphere play an important role in information flow and overall communication efficiency. These findings underlied articulatory difficulties and reduced repetition performance in Broca aphasia, which are rarely observed in anomic aphasia. This research provides novel findings into the resting-state brain network differences between groups of individuals with anomic and Broca aphasia. We identified a subnetwork of, rather than isolated, connections that statistically differentiate the resting-state brain networks of the two groups, in comparison with standard lesion symptom mapping results that yield isolated connections.

q-bio.NC