Searcharxiv⌕ Search

arXiv subjects

Bernardo L. Sabatini

Publications and source records attributed to Bernardo L. Sabatini.

9 recordsLinked to original sources

Conditioned Direct Feedback Alignment via Activity and Error Geometry

Direct feedback alignment (DFA) trains hidden layers with fixed random projections of the output error. Even when this feedback provides useful credit, unequal scales across activity or error directions can distort the local update. We study conditioned DFA (nDFA), a family that adapts established inverse-moment preconditioning to either side of this update. An aligned linear analysis describes how activity conditioning changes spectral learning rates and early-stopping risk. Synthetic experiments show gains from both stabilization and changes in update direction. Confirmatory CIFAR-10 experiments show that activity conditioning, including its FOOF formulation, improves tuned DFA at matched measured training time. Controls support a role for centered correlations beyond the tested mean and diagonal alternatives, while early-only conditioning retains most of the benefit with less training work. Conditioning also benefits backpropagation, and timing interventions do not establish a mechanism specific to random feedback. Error conditioning improves short-budget models and changes update directions, but its small additional improvement under stable training does not survive correction for multiple comparisons. Predicted benefits on image-background benchmarks are not confirmed. These findings establish practical benefits and important limits of conditioning learning with fixed random feedback.

cs.LG↗

Common-Mode Collapse and Recovery in Direct Feedback Alignment

Direct feedback alignment (DFA) trains hidden layers through fixed random projections of output error. With tanh hidden units and independent sigmoid outputs, plain stochastic gradient descent can stall near the loss of a constant predictor of class frequencies. We trace this stall to the error's common mode, the component shared across inputs. An exact mean-covariance decomposition separates a rank-one update formed by the mean teaching signal and mean presynaptic activity. Its leading component drives tanh units toward saturation. At initialization, random feedback provides no systematic correction of the shared error on average; readout learning limits its duration. A reduced model initialized from the network, without fitted parameters, predicts the concentration of activation sensitivity across 48 settings. On MNIST, class decodability largely survives collapse, but readout learning remains slow at a fixed learning rate. Adam learns faster despite deeper collapse. Calibrating the baseline readout to the class prior suppresses collapse and speeds learning; weaker feedback trades less collapse for slower learning. Replacing errors by their signs sustains collapse; subtracting the signal's batch mean prevents sustained collapse and improves learning in the tested setting. Related effects occur in deeper and convolutional networks and on CIFAR-10, with severity and cost depending on the readout, optimizer and input statistics.

cs.LG↗

Shunting Inhibition and Dendritic Branching Shape Local Credit Assignment

Biological neurons assign credit across branching dendrites, where synaptic drive, conductance, local voltage, and somatic teaching signals interact to shape plasticity. We study conductance-based dendritic networks with excitatory and inhibitory synapses, shunting inhibition, and tree-structured branch-to-soma coupling, asking when restricted somatic feedback can approximate compartment-specific backpropagated errors. Exact gradients factor into a synapse-local eligibility term, set by presynaptic activity, driving force, and input resistance, and a path-specific compartment error obtained by transporting a somatic error through dendritic gains. This turns local learning into a credit-signal approximation problem. We test whether shunting improves learning when its effect on dendritic gain makes compartment errors more compatible with restricted feedback. Exact-gradient reconstruction verifies the factorization, while path-gain, feedback-fidelity, inhibition-intervention, and transported-error controls probe the mechanism and its limits. With nonnegative conductances and a five-factor rule using matched-width feedback with scalar fallback, shunting LocalCA remains 5 to 6 percentage points below matched backpropagation on MNIST, Fashion-MNIST, and figure-ground MNIST, showing that feedback fidelity remains a major bottleneck. A three-factor rule approaches matched backpropagation with exact transported feedback in the shunting model and with neuron-wise feedback in both architectures, but shunting has no general advantage under matched initialization. These results show how conductance and dendritic branching enter the exact credit equation and identify restricted feedback as a principal limit.

q-bio.NC↗

When Branch-Local Shunting Helps: A Gain-Load-Alignment Principle for Dendritic E/I Networks

Biological neurons combine excitatory and inhibitory (E/I) activity on branched dendrites through shunting, in which inhibition divisively attenuates excitation. Whether this improves population readout over additive E/I integration of the same nonnegative inputs remains unclear. We introduce DendriNet, a trainable framework that varies integration rule, morphology, synaptic allocation, divisor locality, and dendritic nonlinearities. For population codes with multiplicative gain, a local linearization of any realizable shunting readout yields a decision direction within the positive additive E/I cone; matching the additive optimum requires a positive self-consistent shunting realization. Every scalar shunting threshold also has an exact affine additive realization. Beyond this local limit, performance follows a gain-load-alignment principle: branch-local shunting helps when a reliable divisor suppresses signal-aligned gain more than it attenuates signal or adds denominator variability. Passive additive trees flatten to linear readouts, whereas shunting trees compose local divisors. In a designed hierarchy, deep shunting outperforms tangent and fitted-linear controls, but flexible nonlinear predictors overtake it with enough labels. Support shuffling reverses the linear comparisons, sensor corruption reverses the fitted-linear comparison, and resource-matched activated training shows no consistent depth benefit. The same support and reliability interaction appears in frozen-feature normalization. Across three mouse V1 sessions, the shunting-over-additive decoder gap is largest for narrow readouts, reverses under strong private noise at the widest readout, and varies across running states. Morphology can determine where reliable nuisance estimates meet task-relevant signals, but neither depth nor shunting is intrinsically advantageous.

q-bio.NC↗

Why Depth Matters in Parallelizable Sequence Models: A Lie Algebraic View

Scalable sequence models, such as Transformer variants and structured state-space models, often trade expressivity power for sequence-level parallelism, which enables efficient training. Here we examine the bounds on error and how error scales when models operate outside of their expressivity regimes using a Lie-algebraic control perspective. Our theory formulates a correspondence between the depth of a sequence model and the tower of Lie algebra extensions. Echoing recent theoretical studies, we characterize the Lie-algebraic class of constant-depth sequence models and their corresponding expressivity bounds. Furthermore, we analytically derive an approximation error bound and show that error diminishes exponentially as the depth increases, consistent with the strong empirical performance of these models. We validate our theoretical predictions using experiments on symbolic word and continuous-valued state-tracking problems.

cs.LG↗

Feature leakage and the identifiability of direct-dependency entropy models of neural activity

Biological neurons receive thousands of synaptic inputs on branching, electrically excitable dendrites, yet population activity is often modeled with direct input-output rules in which each input contributes independently to a scalar drive. We study what successful prediction by such models does, and does not, reveal about neural computation. For conditional maximum-entropy models that match output rates and pairwise output-input coactivities, the entropy explained by a direct model is a prediction measure under the sampled input distribution, not a mechanism-identification test. A restricted MaxEnt fit is an information projection: omitted interaction, temporal, or hidden-state terms can be absorbed into fitted first-order parameters whenever they are correlated with the included sufficient statistics. For sparse correlated binary inputs, this absorption has an explicit coskewness form. We introduce diagnostics that separate in-distribution prediction from recovery of the response rule: state reweighting that holds P(y|x) fixed while changing P(x), conditional log-odds contrasts for local additivity, and temporal leakage controls. In ground-truth simulations, purely higher-order responses can pass first-order entropy and raw coactivity tests under leakage-prone sampling, but are correctly classified after reweighting. Applied to selected, leakage-enriched local tables from CA1 hippocampal recordings, approximately half of tables that appear first-order under empirical weights become distribution-sensitive under balanced reweighting, far above a matched additive-surrogate null. Thus direct entropy-explained fractions and raw coactivity predictions should be interpreted as predictions under the observed state distribution, not as evidence that mechanisms outside the direct model are absent or small.

q-bio.NC↗

Task Relevance Is Not Local Replaceability: A Two-Axis View of Channel Information

Channel importance in vision networks is usually summarized by a single score. That summary hides two different questions: how much a channel is related to the task, and whether its function can be supplied by same-layer peers when the channel is removed. We call the second property local replaceability. We introduce a two-axis view that separates these questions. The local axis measures input capture and peer overlap, while the target axis measures task information and target-excess information. Across ResNet-18, VGG-16, and MobileNetV2 trained on CIFAR-100, the two axes are weakly aligned, induce different channel groupings, and separate rapidly during training despite being strongly coupled at random initialization. A Gaussian linear analysis accounts for how this separation can arise through residualized gradient directions, and lesion plus peer-replacement experiments show that peer support refines removability beyond input capture and task relevance alone. Under the fixed FLOPs-matched pruning protocol, local-axis metrics are more reliable predictors of removability than target-axis metrics across the three CIFAR-100 backbones, with the same direction preserved in stress tests on CIFAR-10, Tiny-ImageNet, ImageNet-100, and a ConvNeXt-T/ImageNet-100 pilot. These findings identify an axis-level distinction rather than a universal ranking of pruning scores: local replaceability is a more reliable guide to removability than target relevance, while norm-based baselines remain competitive in architectures such as VGG-16. Relevance-based scores ask what a channel says about the task; pruning asks whether the network still needs that channel when its peers remain available.

cs.CV↗

ALIGN: Adversarial Learning for Generalizable Speech Neuroprosthesis

Intracortical brain-computer interfaces (BCIs) can decode speech from neural activity with high accuracy when trained on data pooled across recording sessions. In realistic deployment, however, models must generalize to new sessions without labeled data, and performance often degrades due to cross-session nonstationarities (e.g., electrode shifts, neural turnover, and changes in user strategy). In this paper, we propose ALIGN, a session-invariant learning framework based on multi-domain adversarial neural networks for semi-supervised cross-session adaptation. ALIGN trains a feature encoder jointly with a phoneme classifier and a domain classifier operating on the latent representation. Through adversarial optimization, the encoder is encouraged to preserve task-relevant information while suppressing session-specific cues. We evaluate ALIGN on intracortical speech decoding and find that it generalizes consistently better to previously unseen sessions, improving both phoneme error rate and word error rate relative to baselines. These results indicate that adversarial domain alignment is an effective approach for mitigating session-level distribution shift and enabling robust longitudinal BCI decoding.

cs.LG↗

Efficient and accurate extraction of in vivo calcium signals from microendoscopic video data

In vivo calcium imaging through microscopes has enabled deep brain imaging of previously inaccessible neuronal populations within the brains of freely moving subjects. However, microendoscopic data suffer from high levels of background fluorescence as well as an increased potential for overlapping neuronal signals. Previous methods fail in identifying neurons and demixing their temporal activity because the cellular signals are often submerged in the large fluctuating background. Here we develop an efficient method to extract cellular signals with minimal influence from the background. We model the background with two realistic components: (1) one models the constant baseline and slow trends of each pixel, and (2) the other models the fast fluctuations from out-of-focus signals and is therefore constrained to have low spatial-frequency structure. This decomposition avoids cellular signals being absorbed into the background term. After subtracting the background approximated with this model, we use Constrained Nonnegative Matrix Factorization (CNMF, Pnevmatikakis et al. (2016)) to better demix neural signals and get their denoised and deconvolved temporal activity. We validate our method on simulated and experimental data, where it shows fast, reliable, and high quality signal extraction under a wide variety of imaging parameters.

q-bio.NC↗