SearcharxivSearch

arXiv subjects

Houman Safaai

Publications and source records attributed to Houman Safaai.

13 recordsLinked to original sources

Shunting Inhibition and Dendritic Branching Shape Local Credit Assignment

Biological neurons assign credit across branching dendrites, where synaptic drive, conductance, local voltage, and somatic teaching signals interact to shape plasticity. We study conductance-based dendritic networks with excitatory and inhibitory synapses, shunting inhibition, and tree-structured branch-to-soma coupling, asking when restricted somatic feedback can approximate compartment-specific backpropagated errors. Exact gradients factor into a synapse-local eligibility term, set by presynaptic activity, driving force, and input resistance, and a path-specific compartment error obtained by transporting a somatic error through dendritic gains. This turns local learning into a credit-signal approximation problem. We test whether shunting improves learning when its effect on dendritic gain makes compartment errors more compatible with restricted feedback. Exact-gradient reconstruction verifies the factorization, while path-gain, feedback-fidelity, inhibition-intervention, and transported-error controls probe the mechanism and its limits. With nonnegative conductances and a five-factor rule using matched-width feedback with scalar fallback, shunting LocalCA remains 5 to 6 percentage points below matched backpropagation on MNIST, Fashion-MNIST, and figure-ground MNIST, showing that feedback fidelity remains a major bottleneck. A three-factor rule approaches matched backpropagation with exact transported feedback in the shunting model and with neuron-wise feedback in both architectures, but shunting has no general advantage under matched initialization. These results show how conductance and dendritic branching enter the exact credit equation and identify restricted feedback as a principal limit.

q-bio.NC

Synaptic clustering emerges from learning and supports covariance discrimination

Functional synapse clusters (FSCs) are synapses with correlated presynaptic activity that are colocalized on the same neuronal dendritic branch. FSCs have been observed after learning in cortical and hippocampal pyramidal neurons. However, previous efforts to ablate FSCs by pharmacologically blocking dendritic nonlinearities to establish causal necessity may have confounded effects. Therefore, whether FSCs are causally necessary for computation is unknown. Here, we attempt to isolate FSCs from this potential confounder in silico. We train Dendrinet, an artificial neural network architecture with hierarchical dendritic segments and sparse conductance-based synapses, on a Permuted-Covariance Classification (PCC) task. This task cannot be solved by single-layer linear-nonlinear artificial neural networks. We find that neurons with dendrites can be trained to solve the task and develop excitatory and inhibitory FSCs if both dendritic nonlinearities and synaptic structural plasticity are active. Turning off dendritic nonlinearities reduces excitatory FSCs, which replicates experimental findings, and reduces performance while unexpectedly increasing inhibitory FSCs. Furthermore, shuffling learned synaptic connectivity while keeping the nonlinearities fixed reduces performance. This shows sensitivity to learned connectivity, but the shuffle does not change only FSCs. Shuffling inhibitory synapse properties reduces performance more than the corresponding excitatory shuffle, showing higher sensitivity to inhibitory organization. This work suggests that dendritic compartmentalization and learned synaptic organization can support computation of covariance structure.

q-bio.NC

When Branch-Local Shunting Helps: A Gain-Load-Alignment Principle for Dendritic E/I Networks

Biological neurons combine excitatory and inhibitory (E/I) activity on branched dendrites through shunting, in which inhibition divisively attenuates excitation. Whether this improves population readout over additive E/I integration of the same nonnegative inputs remains unclear. We introduce DendriNet, a trainable framework that varies integration rule, morphology, synaptic allocation, divisor locality, and dendritic nonlinearities. For population codes with multiplicative gain, a local linearization of any realizable shunting readout yields a decision direction within the positive additive E/I cone; matching the additive optimum requires a positive self-consistent shunting realization. Every scalar shunting threshold also has an exact affine additive realization. Beyond this local limit, performance follows a gain-load-alignment principle: branch-local shunting helps when a reliable divisor suppresses signal-aligned gain more than it attenuates signal or adds denominator variability. Passive additive trees flatten to linear readouts, whereas shunting trees compose local divisors. In a designed hierarchy, deep shunting outperforms tangent and fitted-linear controls, but flexible nonlinear predictors overtake it with enough labels. Support shuffling reverses the linear comparisons, sensor corruption reverses the fitted-linear comparison, and resource-matched activated training shows no consistent depth benefit. The same support and reliability interaction appears in frozen-feature normalization. Across three mouse V1 sessions, the shunting-over-additive decoder gap is largest for narrow readouts, reverses under strong private noise at the widest readout, and varies across running states. Morphology can determine where reliable nuisance estimates meet task-relevant signals, but neither depth nor shunting is intrinsically advantageous.

q-bio.NC

Conditioned Direct Feedback Alignment via Activity and Error Geometry

Direct feedback alignment (DFA) trains hidden layers with fixed random projections of the output error, avoiding the transposed-weight backward pass of backpropagation (BP). We study a failure mode of DFA training that is distinct from feedback quality: the local weight update is calculated by an outer product, so anisotropy can enter through either its presynaptic-activity factor or its local-error factor. Our analyses with controlled synthetic regimes isolate the first failure mode and show an approximately 40-percentage-point activity-conditioning gain when high-variance directions contain task-irrelevant nuisance. Three clean confirmations isolate a different regime: error conditioning improves raw DFA by 1.77--7.53 percentage points, and combining independently selected activity and error factors adds 0.40--0.90 points over activity conditioning. The signs hold for tanh/one-vs-rest MNIST and preregistered Fashion-MNIST, and replicate on eight fresh seeds in a ReLU/softmax MNIST model. This factorization yields a symmetric block-local family of normalized DFA (nDFA): activity nDFA right-preconditions by an inverse activity second moment, error nDFA left-preconditions by an inverse local-error second moment, and K-nDFA applies both factors with separately tuned damping. A linearized post-alignment calculation gives an exact input-side spectral identity and a Kronecker-factor motivation for the two-sided rule, whereas norm matching rules out a scalar step-size explanation. The error factor is fragile when under-damped, BatchNorm is a strong activity-side alternative, and convnet gains remain partial. We therefore frame conditioned DFA as a factor-level study of when local outer-product rules fail, not as a general replacement for BP or a solution to all-layer convolutional credit assignment.

cs.LG

Feature leakage and the identifiability of direct-dependency entropy models of neural activity

Biological neurons receive thousands of synaptic inputs on branching, electrically excitable dendrites, yet population activity is often modeled with direct input-output rules in which each input contributes independently to a scalar drive. We study what successful prediction by such models does, and does not, reveal about neural computation. For conditional maximum-entropy models that match output rates and pairwise output-input coactivities, the entropy explained by a direct model is a prediction measure under the sampled input distribution, not a mechanism-identification test. A restricted MaxEnt fit is an information projection: omitted interaction, temporal, or hidden-state terms can be absorbed into fitted first-order parameters whenever they are correlated with the included sufficient statistics. For sparse correlated binary inputs, this absorption has an explicit coskewness form. We introduce diagnostics that separate in-distribution prediction from recovery of the response rule: state reweighting that holds P(y|x) fixed while changing P(x), conditional log-odds contrasts for local additivity, and temporal leakage controls. In ground-truth simulations, purely higher-order responses can pass first-order entropy and raw coactivity tests under leakage-prone sampling, but are correctly classified after reweighting. Applied to selected, leakage-enriched local tables from CA1 hippocampal recordings, approximately half of tables that appear first-order under empirical weights become distribution-sensitive under balanced reweighting, far above a matched additive-surrogate null. Thus direct entropy-explained fractions and raw coactivity predictions should be interpreted as predictions under the observed state distribution, not as evidence that mechanisms outside the direct model are absent or small.

q-bio.NC

Task Relevance Is Not Local Replaceability: A Two-Axis View of Channel Information

Channel importance in vision networks is usually summarized by a single score. That summary hides two different questions: how much a channel is related to the task, and whether its function can be supplied by same-layer peers when the channel is removed. We call the second property local replaceability. We introduce a two-axis view that separates these questions. The local axis measures input capture and peer overlap, while the target axis measures task information and target-excess information. Across ResNet-18, VGG-16, and MobileNetV2 trained on CIFAR-100, the two axes are weakly aligned, induce different channel groupings, and separate rapidly during training despite being strongly coupled at random initialization. A Gaussian linear analysis accounts for how this separation can arise through residualized gradient directions, and lesion plus peer-replacement experiments show that peer support refines removability beyond input capture and task relevance alone. Under the fixed FLOPs-matched pruning protocol, local-axis metrics are more reliable predictors of removability than target-axis metrics across the three CIFAR-100 backbones, with the same direction preserved in stress tests on CIFAR-10, Tiny-ImageNet, ImageNet-100, and a ConvNeXt-T/ImageNet-100 pilot. These findings identify an axis-level distinction rather than a universal ranking of pruning scores: local replaceability is a more reliable guide to removability than target relevance, while norm-based baselines remain competitive in architectures such as VGG-16. Relevance-based scores ask what a channel says about the task; pruning asks whether the network still needs that channel when its peers remain available.

cs.CV

Dynamic Vine Copulas: Detecting and Quantifying Time-Varying Higher-Order Interactions

Time-varying dependence is often modeled with dynamic correlations or Gaussian graphical models, but multivariate systems can change through tail behavior, asymmetry, or conditional structure even when correlations are nearly stable. We introduce Dynamic Vine Copulas (DVC), a temporal vine-copula framework for estimating and diagnosing sequence-wide non-Gaussian dependence. DVC fixes a chosen vine factorization for comparability; the framework applies to C-, D-, and R-vines, and our experiments use fixed-root-order C-vines. Pair-copula states evolve through smooth parameter trajectories or temporally regularized family-switching paths. The main diagnostic is a held-out comparison between a full vine and its matched 1-truncated version, which separates flexible first-tree pairwise dependence from evidence contributed by higher-tree conditional terms. At the population level, under a correct fixed vine and the simplifying assumption, this contrast equals the higher-tree component of a vine total-correlation decomposition; in finite samples, it is a predictive diagnostic. In controlled benchmarks, DVC detects Student-t degrees-of-freedom changes, Clayton-to-Gumbel switches, and recurrent conditional-interaction episodes missed or conflated by Gaussian dynamic baselines. The higher-tree score remains near zero in pairwise-only regimes and rises during conditional-interaction regimes. On Allen Visual Behavior Neuropixels data, DVC identifies a reproducible time-indexed higher-tree signal that is positive across held-out splits and vanishes under a decorrelated null, indicating simultaneous cross-area dependence. DVC therefore provides a flexible temporal copula model and an interpretable test of whether temporal dependence changes are pairwise or conditional.

stat.ML

Amortized Vine Copulas for High-Dimensional Density and Information Estimation

Modeling high-dimensional dependencies while keeping likelihoods tractable remains challenging. Classical vine-copula pipelines are interpretable but can be expensive, while many neural estimators are flexible but less structured. In this work, we propose Vine Denoising Copula (VDC), an amortized vine-copula pipeline for continuous-data, simplified-vine dependence modeling. VDC trains a single bivariate denoising model and reuses it across all vine edges. For each edge, given pseudo-observations, the model predicts a piecewise-constant density grid. We then apply an IPFP/Sinkhorn projection that normalizes mass and drives the marginals to uniformity. This preserves the tractable vine-likelihood structure and the usual copula interpretation while replacing repeated per-edge optimization with GPU inference. Across synthetic and real-data benchmarks, VDC delivers strong bivariate density accuracy, competitive MI/TC estimation, and faster high-dimensional vine fitting. These gains make explicit information estimation and dependence decomposition feasible when repeated vine fitting would otherwise be costly, while conditional downstream tasks remain a limitation.

cs.LG

Supernodes and Halos: Loss-Critical Hubs in LLM Feed-Forward Layers

We study the organization of channel-level importance in transformer feed-forward networks (FFNs). Using a Fisher-style loss proxy (LP) based on activation-gradient second moments, we show that loss sensitivity is concentrated in a small set of channels within each layer. In Llama-3.1-8B, the top 1% of channels per layer accounts for a median of 58.7% of LP mass, with a range of 33.0% to 86.1%. We call these loss-critical channels supernodes. Although FFN layers also contain strong activation outliers, LP-defined supernodes overlap only weakly with activation-defined outliers and are not explained by activation power or weight norms alone. Around this core, we find a weaker but consistent halo structure: some non-supernode channels share the supernodes' write support and show stronger redundancy with the protected core. We use one-shot structured FFN pruning as a diagnostic test of this organization. At 50% FFN sparsity, baselines that prune many supernodes degrade sharply, whereas our SCAR variants explicitly protect the supernode core; the strongest variant, SCAR-Prot, reaches perplexity 54.8 compared with 989.2 for Wanda-channel. The LP-concentration pattern appears across Mistral-7B, Llama-2-7B, and Qwen2-7B, remains visible in targeted Llama-3.1-70B experiments, and increases during OLMo-2-7B pretraining. These results suggest that LLM FFNs develop a small learned core of loss-critical channels, and that preserving this core is important for reliable structured pruning.

cs.LG

Information estimation using nonparametric copulas

Estimation of mutual information between random variables has become crucial in a range of fields, from physics to neuroscience to finance. Estimating information accurately over a wide range of conditions relies on the development of flexible methods to describe statistical dependencies among variables, without imposing potentially invalid assumptions on the data. Such methods are needed in cases that lack prior knowledge of their statistical properties and that have limited sample numbers. Here we propose a powerful and generally applicable information estimator based on non-parametric copulas. This estimator, called the non-parametric copula-based estimator (NPC), is tailored to take into account detailed stochastic relationships in the data independently of the data's marginal distributions. The NPC estimator can be used both for continuous and discrete numerical variables and thus provides a single framework for the mutual information estimation of both continuous and discrete data. By extensive validation on artificial samples drawn from various statistical distributions, we found that the NPC estimator compares well against commonly used alternatives. Unlike methods not based on copulas, it allows an estimation of information that is robust to changes of the details of the marginal distributions. Unlike parametric copula methods, it remains accurate regardless of the precise form of the interactions between the variables. In addition, the NPC estimator had accurate information estimates even at low sample numbers, in comparison to alternative estimators. The NPC estimator therefore provides a good balance between general applicability to arbitrarily shaped statistical dependencies in the data and shows accurate and robust performance when working with small sample sizes.

stat.ME

Quantifying how much sensory information in a neural code is relevant for behavior

Determining how much of the sensory information carried by a neural code contributes to behavioral performance is key to understand sensory function and neural information flow. However, there are as yet no analytical tools to compute this information that lies at the intersection between sensory coding and behavioral readout. Here we develop a novel measure, termed the information-theoretic intersection information $I_{II}(S;R;C)$, that quantifies how much of the sensory information carried by a neural response R is used for behavior during perceptual discrimination tasks. Building on the Partial Information Decomposition framework, we define $I_{II}(S;R;C)$ as the part of the mutual information between the stimulus S and the response R that also informs the consequent behavioral choice C. We compute $I_{II}(S;R;C)$ in the analysis of two experimental cortical datasets, to show how this measure can be used to compare quantitatively the contributions of spike timing and spike rates to task performance, and to identify brain areas or neural populations that specifically transform sensory information into choice.

q-bio.NC

Exploring Pure Spinor String Theory on AdS_4 x CP^3

In this paper we formulate the pure spinor superstring theory on AdS_4 x CP^3. By recasting the pure spinor action as a topological A-model on the fermionic supercoset Osp(6|4)/SO(6)xSp(4) plus a BRST exact term, we prove the exactness of the sigma-model. We then give a gauged linear sigma-model which reduces to the superstring in the limit of large volume and we study its branch geometry in different phases. Moreover, we discuss possible D-brane boundary conditions and the principal chiral model for the fermionic supercoset.

hep-th

On gauge/string correspondence and mirror symmetry

We consider a mirror dual of the Berkovits-Vafa A-model for the BPS superstring on $AdS_5\times S^5$ in the form of a deformed superconifold. Via geometric transition, the theory has a dual description as the hermitian gaussian one-matrix model. We show that the A-model amplitudes of generic $AdS_2\times S^4$ branes, breaking the superconformal symmetry as $U(2,2|4)\to OSp(4^*|4)$, are evaluated in terms of observables in the matrix model. As such, upon the usual identification $g_{YM}^2=g_s$, these can be expanded as Drukker-Gross circular 1/2-BPS Wilson loops in the perturbative regime of ${\cal N}=4$ SYM.

hep-th