SearcharxivSearch

arXiv subjects

Suraj Yadav

Publications and source records attributed to Suraj Yadav.

9 recordsLinked to original sources

Self-Routed Tensor Adapters for Parameter-Efficient Universal Visual Adaptation

Universal visual representations require adaptation mechanisms that adapt across heterogeneous domains without fragmenting knowledge into domain-specific modules. Parameter-efficient fine-tuning adapts frozen visual foundation models efficiently, but standard low-rank adapters use a fixed subspace for all inputs, which can be restrictive when domains differ in style, background, and semantic context. MoE-based adapters improve specialization through multiple expert pathways, but often rely on external routers and large expert banks, adding parameters and separating routing from adaptation. We propose \textbf{Self-Routed Tensor Adapters}, a compact framework for multi-domain visual adaptation. SRTA projects each input into a low-rank space, computes routing weights from this representation using a learnable domain matrix, and uses these weights to blend slices of a shared Tucker core. This produces a sample-specific adaptation matrix without an external gating network, allowing shared visual factors to be reused while supporting domain-aware specialization. To strengthen pathway learning, we introduce a progressive depth-weighted routing objective that supervises routing decisions across adapter layers. Across five heterogeneous multi-domain visual classification benchmarks, SRTA achieves competitive or slightly stronger average accuracy than MoE-style PEFT baselines while using substantially fewer trainable parameters. At rank 64, SRTA uses 2.77M parameters in the 4-domain setting compared with 9.52M for MoLoRA, and 3.00M in the 6-domain setting compared with 14.31M. Overall, SRTA offers an effective accuracy-parameter trade-off for adapting visual foundation models toward universal multi-domain representations. \href{https://github.com/surajyadav-research/SRTA}{GitHub}

cs.CV

Breaking Spurious Correlations via Generative Randomization and Cross-Variant Self-Supervised Learning

Deep neural networks trained with Empirical Risk Minimization (ERM) often fail under distribution shifts because they exploit spurious correlations between object labels and background context. Recent generative approaches address this issue by creating counterfactual images with altered contexts, but typically use these samples as standard data augmentation, leaving the model free to retain background-sensitive representations. We propose a two-stage framework that uses generative intervention to explicitly learn background-invariant visual representations. First, we isolate the foreground object using zero-shot segmentation and generate context-shifted variants with a structure-preserving diffusion model, preserving object identity while varying the surrounding environment. We then introduce Cross-Variant Self-Supervised Learning, where variants of the same object under different backgrounds form positive pairs in a contrastive objective. This encourages the encoder to align object-centric representations while suppressing background-specific cues. Then, we fine-tune the pretrained encoder using an ERM warm-up followed by GroupDRO with layer-wise learning rates. Experiments on distribution-shift benchmarks demonstrate best worst-group performance, achieving 92.5% on Waterbirds, 81.7% on MetaShift, and 87.4% on NICO++. Code: https://github.com/surajyadav-research/GRSSL

cs.CV

Limits of Difficulty Scaling: Hard Samples Yield Diminishing Returns in GRPO-Tuned SLMs

Recent alignment work on Large Language Models (LLMs) suggests preference optimization can improve reasoning by shifting probability mass toward better solutions. We test this claim in a resource-constrained setting by applying GRPO with LoRA to SLMs (up to 3B) for math reasoning on GSM8K and MATH datasets with difficulty-stratified analyses. As problem difficulty increases, accuracy plateaus, revealing a capacity boundary: GRPO primarily reshapes output preferences without reliably improving hardest-tier solving. Consistent with this, training GRPO only on lower-difficulty problems matches full-dataset accuracy across difficulty tiers while using only ~45% training steps, indicating diminishing returns from harder samples in this regime. We also find a cross-dataset generalization effect: GSM8K-trained GRPO achieves higher accuracy on the numeric subset of MATH than MATH-trained GRPO, exceeding it by ~5% at 1.5B and by ~3% at 3B. We show that the best achievable gains depend strongly on the base model's prior reasoning competence and the dataset's difficulty profile.

cs.LG

$\mathbb{A}^1$-connected stacky curves and the Brauer group of moduli of elliptic curves

Given a smooth scheme X with an action by an affine algebraic group G, we give a formula to compute the Nisnevich sheaf of the motivic connected components of the quotient stack [X/G] in the case of an orbifold. We apply it to identify all the $\mathbb{A}^1$-connected stacky curves and prove the $\mathbb{A}^1$ -connectedness of the moduli stack of elliptic curves. We also prove homotopy purity for the smooth algebraic stacks and recover a computation of the Picard group of stacky orbifold curves by computing their mixed motive.

math.AG

$\mathbb{A}^1$-connectivity of moduli of vector bundles on a curve

In this note we prove that the moduli stack of vector bundles on a curve, with a fixed determinant is $\mathbb{A}^1$-connected. We obtain this result by classifying vector bundles on a curve upto $\mathbb{A}^1$-concordance. Consequently we classify$\mathbb{P}^n$- bundles on a curve upto $\mathbb{A}^1$-weak equivalence, extending a result of Asok-Morel. We also give an explicit example of a variety which is $\mathbb{A}^1$-h-cobordant to a projective bundle over $\mathbb{P}^2$ but does not have the structure of a projective bundle over $\mathbb{P}^2$, thus answering a question of Asok-Kebekus-Wendt

math.AG

Nisnevich local Good compactifications

For a local complete intersection morphism, we establish fiberwise denseness in the $n$-dimensional irreducible components of the compactification Nisnevich locally.

math.AG

Grothendieck-Serre Conjecture for Quasi-split Reductive Groups

We prove the Grothendieck-Serre conjecture for quasi-split reductive groups schemes. Our method involves reducing to the Borel subgroup in order to conclude the result from purity for tori and the structure theorem for unipotent radicals of parabolic subgroups in a reductive group.

math.AG