SearcharxivSearch

arXiv subjects

Felippe Alves

Publications and source records attributed to Felippe Alves.

5 recordsLinked to original sources

Kolmogorov--Arnold Networks for Small Language Models

Kolmogorov--Arnold Networks (KANs) replace fixed node activations with learned one-dimensional edge functions, offering an explicit interface for interpretation and a possible alternative to transformer feed-forward networks. We test these claims separately. In a six-layer, 10M-parameter B-spline KAN, we reconstruct all 884,736 feed-forward edges: 87.8\% exceed (NLS>0.1) and 0.4\% are inactive. Pruning the lowest-activity 20--25\% causes negligible loss increase, although structured MLP neuron pruning tolerates comparable sparsity. The audit replicates on BabyLM, but grid-size sweeps show that near-total fPCA compression and high closed-form-fit coverage are properties of the low-capacity grid-2 basis, not universal KAN behavior. For replacement, we evaluate MLP, SwiGLU, grouped Chebyshev, and rational GR-KAN networks on BabyLM. The KAN-family and gated variants improve validation loss over the GELU MLP, but this ordering does not transfer to standardized benchmarks: across ten seeds and 59,875 BLiMP pairs, accuracies span 62.4--63.1\%, EWoK remains at chance, and a (+0.7)-point GR-KAN effect on BLiMP reverses on the supplement. Larger tests are also cautionary: parameter-matched MLPEdge underperforms the MLP on Wikitext-103, and 286M-parameter GR-KAN remains below a SwiGLU ClimbMix baseline after stabilization. Thus, small-basis KANs provide a practical, corpus-transferable interface for auditing learned scalar transformations, but the tested replacements show no consistent benchmark, quality, or latency advantage over strong MLP baselines.

cs.LG

Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models

Large language models (LLMs) increasingly operate in social contexts, motivating analysis of how they express and shift moral judgments. In this work, we investigate the moral response of LLMs to persona role-play, prompting a LLM to assume a specific character. Using the Moral Foundations Questionnaire (MFQ), we introduce a benchmark that quantifies two properties: moral susceptibility and moral robustness, defined from the variability of MFQ scores across- and within-personas. We estimate these quantities with two complementary procedures, repeated sampling and a logit-based method that directly estimates the rating distributions and enables temperature analysis. We evaluate 15 models across six families: Claude, DeepSeek, Gemini, GPT, Grok, and Llama. The two metrics show qualitatively different patterns. Moral robustness varies by more than an order of magnitude, with a coefficient of variation of about $152\%$, and is explained almost entirely by model family. The Claude family is, by a significant margin, the most robust, about 30 times more so than the lower-performing families (DeepSeek, Grok, and Llama), while Gemini and GPT occupy an intermediate tier. This strong family dependence suggests that robustness is primarily shaped by post-training. Moral susceptibility, by contrast, spans a much narrower range, with a coefficient of variation of about $13\%$, and the most susceptible model is only 1.6 times more susceptible than the least. Unlike robustness, susceptibility shows no clear family dependence, suggesting that it is primarily determined by pre-training. Additionally, we present moral foundation profiles for models without persona role-play and for personas averaged across models. Together, these analyses provide a systematic view of how persona conditioning shapes moral behavior in LLMs and a window into the internal machinery they use to instantiate personas.

cs.CL

The futility of being selfish in vaccine distribution

We study vaccine budget-sharing strategies in the SIR (Susceptible-Infected-Recovered) model given a structured community network to investigate the benefit of sharing vaccine across communities. The network studied comprises two communities, one of which controls vaccine budget and may share it with the other. Different scenarios are considered regarding the connectivity between communities, infection rates and the unvaccinated fraction of the population. Properties of the SIR model facilitates the use of Dynamic Message Passing (DMP) and optimal control methods to investigate preventive and reactive budget-sharing scenarios. Our results show a large set of budget-sharing strategies in which the sharing community benefits from the reduced global infection rates with no detrimental impact on its local infection rate.

physics.soc-ph

Homo Entropicus, the emotional agent and societies of Neural Networks

A neural network with a learning algorithm optimized by information theory entropic dynamics is used to build an agent dubbed Homo Entropicus. The algorithm can be described at a macroscopic level in terms of aggregate variables interpretable as quantitative markers of proto-emotions. We use systems of such interacting neural networks to construct a framework for modeling societies that show complex emergent behavior. A few applications are presented to investigate the role the interactions of opinions about multidimensional issues and trust on the information source play on the state of the agent society. These include the case of a class of $N$ agents learning from a fixed teacher; two dynamical agents; panels of three agents modeling the interactions that occur in decisions of the US Court of Appeals, where we quantify how politically biased are the agents, how trustful of other agents-judges of other parties, how much the agents follow a common understanding of the law. Finally we address under which conditions ideological polarization follows or precedes affective polarization in large societies and how simpler versions of the learning algorithm may change these relations.

physics.soc-ph

Frustration, glassy behavior and dynamical annealing in societies of Neural Networks

We study maximum entropy mechanisms of information exchange between agents modeled by neural networks and the macroscopic states of a society of such agents in a few situations. Mathematical quantification of surprise, distrust of other agents and confidence about its opinion emerge as essential ingredients in the entropy based learning dynamics. Learning is shown to be driven by surprises, i.e. the receptor agent is confronted with the concurring opinion of a distrusted agent or with a trusted agent's disagreeing opinion. Attribution of blame for the surprise derives from measures of distrust of the receiver towards the emitter agent and the receiver's confidence about its own opinion. The dynamics proceeds by changes of mainly one or the other: the receptor opinion about the issue or the distrust about the emitter. A society with $N$ agents exchanging binary opinions about a set of issues show rich behavior which depend on the complexity of the agenda. For small sets the society reaches a steady state polarized into antagonistic factions, where balanced norms such as "the friend of an enemy is an enemy" are strictly satisfied. For larger sets of issues, societies can persist for a long time in spin-glass like states. There are two types of frustration: ideological and affective, with dynamical annealing properties depending on the complexity of the set of questions under discussion, leading to the lack of sharply defined parties for long transients.

physics.soc-ph