Searcharxiv⌕ Search

arXiv subjects

Seyed Morteza Emadi

Publications and source records attributed to Seyed Morteza Emadi.

6 recordsLinked to original sources

When does a scaling result justify a different allocation? A critical review of resource-allocation evidence for AI systems

AI scaling studies increasingly evaluate systems that combine a pretrained model with retrieval, search, verification, tools, and interaction. Yet a higher score under a larger budget does not by itself show where additional resources are best spent. This critical integrative review asks when a reported scaling result supports a resource-allocation decision. It compares evidence across pretraining, test-time computation, retrieval, and agent evaluation, distinguishing the performance of a tested procedure from the best performance achievable under a resource limit. The synthesis shows that three mismatches recur across this evidence: success counted before an answer is chosen, information a deployed system will not have, and costs left out of the comparison. A capability surface expresses performance as a function of budgets, mechanisms, and available information. Worked analytical examples show how the evaluation metric, deployment volume, selection rule, and stopping policy can alter an allocation conclusion. A resource envelope provides a structured record of the task, development and run-time resources, information access, and procedure behind a reported score. Its application to a published comparison illustrates which conclusions the evidence supports and which deployment questions remain unresolved. The resulting framework specifies the comparisons needed to choose among feasible systems and motivates experiments on the transfer of allocation rules across tasks and operating conditions. It does not propose a universal scaling law or infer general intelligence from benchmark gains.

cs.AI↗

Exact Attention Sensitivity and the Geometry of Transformer Stability

We develop a sensitivity analysis for transformer attention in a geometry aligned with tokenwise computation. Our main result is the exact identity $\|J_τ(u)\|_{\infty\to1}=θ(p)/τ$ for the Jacobian $J_τ(u)$ of the tempered softmax $u\mapsto\mathrm{softmax}(u/τ)$, where $θ(p)=4\max_{S\subseteq[L]}p(S)(1-p(S))$ measures how evenly the attention distribution can be bisected rather than how concentrated it is. We combine this identity with a block-$\infty$/RMS norm under which row-stochastic attention mixing is nonexpansive. This yields a distribution-aware local Jacobian bound for multi-head attention and a sequence-length-independent Lipschitz bound on bounded input sets, with explicit dependence on width, input magnitude, temperature, and projection norms. We also identify a structural distinction between normalization placements: a pre-LN residual-sublayer Jacobian contains an additive identity term, whereas a post-LN residual-sublayer Jacobian does not. A LayerNorm projection lemma gives a sufficient condition under which the LayerNorm-only term in the post-LN expansion contracts geometrically; the condition is not tested by our experiments. Across three Pre-LN early-training runs of $774$M-parameter models, attention becomes substantially more concentrated while the median lower-bound certificate for $θ(p)$ remains near one at every sampled layer and checkpoint. This certifies near-maximal exact sensitivity for at least half of the sampled rows within each layer. A minority of rows enters a dominant-atom regime with lower exact sensitivity, consistent with the deterministic relationship between $p_{\max}$ and $θ(p)$.

cs.LG↗

The Causal Description Gap: Information-Theoretic Separations Across Pearl's Hierarchy

Pearl's causal hierarchy shows that observational, interventional, and counterfactual queries are qualitatively distinct. We ask a quantitative version of this question: how many additional bits are needed to specify higher-rung causal answers once lower-rung answers are known? We formalize this via query-class description length, the Kolmogorov complexity of the answer oracle induced by an SCM for a class of queries. Our main construction gives binary acyclic SCMs whose observational distribution has constant description length, while the single-variable interventional answer oracle has description length $Θ(n^2)$. A degree-sensitive upper bound shows that finite-gate-schema SCMs of indegree $d$ have observational-interventional gap at most $O(nd \log(en/d) + n \log n)$, making the quadratic construction order-optimal in the dense regime and a rooted-tree construction order-optimal for bounded indegree. The quadratic separation persists under $\varepsilon$-accurate total-variation descriptions for every fixed $\varepsilon < 1/4$. At the next rung, the full hard-do interventional oracle can still leave a $Θ(n)$ counterfactual description gap. A general ambiguity-to-bits theorem and Shannon analogue show that these gaps equal the logarithm of residual higher-rung ambiguity up to lower-order terms.

stat.ML↗

Rank-Aware Spectral Bounds on Attention Logits for Stable Low-Precision Training

Attention scores in transformers are bilinear forms $S_{ij} = x_i^\top M x_j / \sqrt{d_h}$ whose maximum magnitude governs overflow risk in low-precision training. We derive a \emph{rank-aware concentration inequality}: when the interaction matrix $M = W^Q W^{K\top}$ has rank $r \ll d$, tail probabilities for $\max_{i,j}|S_{ij}|$ decay as $\exp(-d^{2}α^{2}/(γr))$ rather than $\exp(-dα^{2})$, where $γ> 1$ is a typicality parameter. For transformer attention where $r = d_h$, this yields $8$--$28\times$ tighter concentration than rank-agnostic bounds in modern architectures. We apply this result to FP8 training, deriving \emph{geometry-aware scale factors} that provide principled overflow guarantees without observing activations. The method computes per-layer scales from the spectral norm $\|W^Q W^{K\top}\|_2$ via implicit power iteration, includes a grouped query attention formulation that avoids key expansion, and remains compatible with fused attention kernels. Across GPT-2 XL to Llama-2-70B, geometry-aware scaling eliminates overflows in transient scenarios where delayed scaling fails, while achieving comparable downstream MMLU accuracy.

cs.LG↗

The Critical Horizon: Inspection Design Principles for Multi-Stage Operations and Deep Reasoning

Manufacturing lines, service journeys, supply chains, and AI reasoning chains share a common challenge: attributing a terminal outcome to the intermediate stage that caused it. We establish an information-theoretic barrier to this credit assignment problem: the signal connecting early steps to final outcomes decays exponentially with depth, creating a critical horizon beyond which reliable learning from endpoint data alone requires exponentially many samples. We prove four results. First, a Signal Decay Bound: sample complexity for attributing outcomes to early stages grows exponentially in the number of intervening steps. Second, Width Limits: parallel rollouts provide only logarithmic relief, with correlation capping the effective number of independent samples. Third, an Objective Mismatch: additive reward aggregation optimizes the wrong quantity when sequential validity requires all steps to be correct. Fourth, Optimal Inspection Design: uniform checkpoint spacing is minimax-optimal under homogeneous signal attenuation, while a greedy algorithm yields optimal non-uniform schedules under heterogeneous attenuation. Together, these results provide a common analytical foundation for inspection design in operations and supervision design in AI.

stat.ML↗

Testing the Exogeneity of Instrumental Variables and Regressors in Linear Regression Models Using Copulas

We provide a Copula-based approach to test the exogeneity of instrumental variables in linear regression models. We show that the exogeneity of instrumental variables is equivalent to the exogeneity of their standard normal transformations with the same CDF value. Then, we establish a Wald test for the exogeneity of the instrumental variables. We demonstrate the performance of our test using simulation studies. Our simulations show that if the instruments are actually endogenous, our test rejects the exogeneity hypothesis approximately 93% of the time at the 5% significance level. Conversely, when instruments are truly exogenous, it dismisses the exogeneity assumption less than 30% of the time on average for data with 200 observations and less than 2% of the time for data with 1,000 observations. Our results demonstrate our test's effectiveness, offering significant value to applied econometricians.

stat.ME↗