SearcharxivSearch

arXiv subjects

Sunny Gupta

Publications and source records attributed to Sunny Gupta.

At least 19 recordsLinked to original sources

Amortizing Federated Adaptation: Hypernetwork Driven LoRA for Personalized Foundation Models

Federated fine-tuning of foundation models using Low-Rank Adaptation (LoRA) offers a communication efficient solution for distributed learning. However, existing federated LoRA methods suffer from two fundamental limitations: (1) structural aggregation bias, where independently averaging low rank factors fails to approximate the true combined update, and (2) client side initialization lag, as clients repeatedly reinitialize LoRA parameters across communication rounds, slowing convergence. We propose HyperLoRA, a unified framework that addresses both issues through amortized federated adaptation through hypernetwork-driven LoRA generation and product space aggregation. Instead of iterative per-client optimization, HyperLoRA employs a learned generator that maps client distribution signatures to LoRA initializations, effectively amortizing per client adaptation. On the server side, we introduce a learned aggregation module that directly synthesizes updates in the low-rank product space, eliminating the inconsistencies of factor-wise averaging. A lightweight residual correction module further improves stability under heterogenous (non-IID) client distributions.By replacing iterative optimization and heuristic averaging with learned operators, HyperLoRA jointly enables efficient personalization, unbiased aggregation, and faster convergence. Experiments on federated vision and vision-language benchmarks show that HyperLoRA achieves improved convergence speed, greater robustness to distribution shift, and stronger personalization performance compared to prior federated LoRA methods.

cs.AI

BiPrompt: Bilateral Prompt Optimization for Visual and Textual Debiasing in Vision-Language Models

Vision language foundation models such as CLIP exhibit impressive zero-shot generalization yet remain vulnerable to spurious correlations across visual and textual modalities. Existing debiasing approaches often address a single modality either visual or textual leading to partial robustness and unstable adaptation under distribution shifts. We propose a bilateral prompt optimization framework (BiPrompt) that simultaneously mitigates non-causal feature reliance in both modalities during test-time adaptation. On the visual side, it employs structured attention-guided erasure to suppress background activations and enforce orthogonal prediction consistency between causal and spurious regions. On the textual side, it introduces balanced prompt normalization, a learnable re-centering mechanism that aligns class embeddings toward an isotropic semantic space. Together, these modules jointly minimize conditional mutual information between spurious cues and predictions, steering the model toward causal, domain invariant reasoning without retraining or domain supervision. Extensive evaluations on real-world and synthetic bias benchmarks demonstrate consistent improvements in both average and worst-group accuracies over prior test-time debiasing methods, establishing a lightweight yet effective path toward trustworthy and causally grounded vision-language adaptation.

cs.CV

FedHypeVAE: Federated Learning with Hypernetwork Generated Conditional VAEs for Differentially Private Embedding Sharing

Federated data sharing promises utility without centralizing raw data, yet existing embedding-level generators struggle under non-IID client heterogeneity and provide limited formal protection against gradient leakage. We propose FedHypeVAE, a differentially private, hypernetwork-driven framework for synthesizing embedding-level data across decentralized clients. Building on a conditional VAE backbone, we replace the single global decoder and fixed latent prior with client-aware decoders and class-conditional priors generated by a shared hypernetwork from private, trainable client codes. This bi-level design personalizes the generative layerrather than the downstream modelwhile decoupling local data from communicated parameters. The shared hypernetwork is optimized under differential privacy, ensuring that only noise-perturbed, clipped gradients are aggregated across clients. A local MMD alignment between real and synthetic embeddings and a Lipschitz regularizer on hypernetwork outputs further enhance stability and distributional coherence under non-IID conditions. After training, a neutral meta-code enables domain agnostic synthesis, while mixtures of meta-codes provide controllable multi-domain coverage. FedHypeVAE unifies personalization, privacy, and distribution alignment at the generator level, establishing a principled foundation for privacy-preserving data synthesis in federated settings. Code: github.com/sunnyinAI/FedHypeVAE

cs.LG

Federated Cross-Modal Style-Aware Prompt Generation

Prompt learning has propelled vision-language models like CLIP to excel in diverse tasks, making them ideal for federated learning due to computational efficiency. However, conventional approaches that rely solely on final-layer features miss out on rich multi-scale visual cues and domain-specific style variations in decentralized client data. To bridge this gap, we introduce FedCSAP (Federated Cross-Modal Style-Aware Prompt Generation). Our framework harnesses low, mid, and high-level features from CLIP's vision encoder alongside client-specific style indicators derived from batch-level statistics. By merging intricate visual details with textual context, FedCSAP produces robust, context-aware prompt tokens that are both distinct and non-redundant, thereby boosting generalization across seen and unseen classes. Operating within a federated learning paradigm, our approach ensures data privacy through local training and global aggregation, adeptly handling non-IID class distributions and diverse domain-specific styles. Comprehensive experiments on multiple image classification datasets confirm that FedCSAP outperforms existing federated prompt learning methods in both accuracy and overall generalization.

cs.CV

FEDTAIL: Federated Long-Tailed Domain Generalization with Sharpness-Guided Gradient Matching

Domain Generalization (DG) seeks to train models that perform reliably on unseen target domains without access to target data during training. While recent progress in smoothing the loss landscape has improved generalization, existing methods often falter under long-tailed class distributions and conflicting optimization objectives. We introduce FedTAIL, a federated domain generalization framework that explicitly addresses these challenges through sharpness-guided, gradient-aligned optimization. Our method incorporates a gradient coherence regularizer to mitigate conflicts between classification and adversarial objectives, leading to more stable convergence. To combat class imbalance, we perform class-wise sharpness minimization and propose a curvature-aware dynamic weighting scheme that adaptively emphasizes underrepresented tail classes. Furthermore, we enhance conditional distribution alignment by integrating sharpness-aware perturbations into entropy regularization, improving robustness under domain shift. FedTAIL unifies optimization harmonization, class-aware regularization, and conditional alignment into a scalable, federated-compatible framework. Extensive evaluations across standard domain generalization benchmarks demonstrate that FedTAIL achieves state-of-the-art performance, particularly in the presence of domain shifts and label imbalance, validating its effectiveness in both centralized and federated settings. Code: https://github.com/sunnyinAI/FedTail

cs.AI

UniVarFL: Uniformity and Variance Regularized Federated Learning for Heterogeneous Data

Federated Learning (FL) often suffers from severe performance degradation when faced with non-IID data, largely due to local classifier bias. Traditional remedies such as global model regularization or layer freezing either incur high computational costs or struggle to adapt to feature shifts. In this work, we propose UniVarFL, a novel FL framework that emulates IID-like training dynamics directly at the client level, eliminating the need for global model dependency. UniVarFL leverages two complementary regularization strategies during local training: Classifier Variance Regularization, which aligns class-wise probability distributions with those expected under IID conditions, effectively mitigating local classifier bias; and Hyperspherical Uniformity Regularization, which encourages a uniform distribution of feature representations across the hypersphere, thereby enhancing the model's ability to generalize under diverse data distributions. Extensive experiments on multiple benchmark datasets demonstrate that UniVarFL outperforms existing methods in accuracy, highlighting its potential as a highly scalable and efficient solution for real-world FL deployments, especially in resource-constrained settings. Code: https://github.com/sunnyinAI/UniVarFL

cs.LG

Colossal anomalous Stark shift in defect emission of undulated 2D materials

We report a strikingly new physical phenomenon that mirror symmetry breaking in undulated two-dimensional (2D) materials induces a colossal Stark shift in defect emissions, occurring without external electric field F, termed anomalous Stark effect. First-principles calculations of multiple defects in bent 2D hBN uncover the fundamental physical reasonings for this anomalous effect and reveal this arises due to strong coupling between flexoelectric polarization and defect dipole moment. This flexo-dipole interaction, similar to that in traditional Stark effect due to F, results in zero-phonon line (ZPL) shifts >500 meV for defects like NBVN and CBVN at $\kappa$ = 1/nm, exceeding typical Stark shifts by 2-3 orders of magnitude. The large ZPL shifts variations with curvature and bending direction offers a method to identify nanotube chirality and explain the large variability in single photon emitters' wavelength in 2D materials, with additional implications for designing nano-electro-mechanical and photonic devices.

cond-mat.mtrl-sci

FedAlign: Federated Domain Generalization with Cross-Client Feature Alignment

Federated Learning (FL) offers a decentralized paradigm for collaborative model training without direct data sharing, yet it poses unique challenges for Domain Generalization (DG), including strict privacy constraints, non-i.i.d. local data, and limited domain diversity. We introduce FedAlign, a lightweight, privacy-preserving framework designed to enhance DG in federated settings by simultaneously increasing feature diversity and promoting domain invariance. First, a cross-client feature extension module broadens local domain representations through domain-invariant feature perturbation and selective cross-client feature transfer, allowing each client to safely access a richer domain space. Second, a dual-stage alignment module refines global feature learning by aligning both feature embeddings and predictions across clients, thereby distilling robust, domain-invariant features. By integrating these modules, our method achieves superior generalization to unseen domains while maintaining data privacy and operating with minimal computational and communication overhead.

cs.LG

Sequential Compression Layers for Efficient Federated Learning in Foundational Models

Federated Learning (FL) has gained popularity for fine-tuning large language models (LLMs) across multiple nodes, each with its own private data. While LoRA has been widely adopted for parameter efficient federated fine-tuning, recent theoretical and empirical studies highlight its suboptimal performance in the federated learning context. In response, we propose a novel, simple, and more effective parameter-efficient fine-tuning method that does not rely on LoRA. Our approach introduces a small multi-layer perceptron (MLP) layer between two existing MLP layers the up proj (the FFN projection layer following the self-attention module) and down proj within the feed forward network of the transformer block. This solution addresses the bottlenecks associated with LoRA in federated fine tuning and outperforms recent LoRA-based approaches, demonstrating superior performance for both language models and vision encoders.

cs.LG

Undulation-induced moir\'e superlattices with 1D polarization domains and 1D flat bands in 2D bilayer semiconductors

Two-dimensional (2D) materials have a high F\"oppl-von K\'arm\'an number and can be easily bent, much like a paper, making undulations a novel way to design distinct electronic phases. Through first-principles calculations, we reveal the formation of 1D polarization domains and 1D flat electronic bands by 1D bending modulation to a 2D bilayer semiconductor. Using 1D sinusoidal undulation of a hexagonal boron nitride (hBN) bilayer as an example, we demonstrate how undulation induces nonuniform shear patterns, creating regions with unique local stacking and vertical polarization akin to sliding-induced ferroelectrics observed in twisted moir\'e systems. This sliding-induced polarization is also observed in double-wall BN nanotubes due to curvature differences between inner and outer tubes. Furthermore, undulation generates a shear-induced 1D moir\'e pattern that perturbs electronic states, confining them into 1D quantum-well-like bands with kinetic energy quenched in modulation direction while dispersive in other directions (1D flat bands). This electronic confinement is attributed to modulated shear deformation potential resulting from tangential polarization due to the moir\'e pattern. Thus, bending modulation and interlayer shear offer an alternative avenue, termed "curvytronics", to induce exotic phenomena in 2D bilayer materials.

cond-mat.mes-hall

Undulated 2D materials as a platform for large Rashba spin-splitting and persistent spin-helix states

Materials with large unidirectional Rashba spin-orbit coupling (SOC), resulting in persistent-spin helix states with small spin-precession length, are critical for advancing spintronics. We demonstrate a design principle achieving it through specific undulations of 2D materials. Analytical model and first-principles calculations reveal that bending-induced asymmetric hybridization brings about and even enhances Rashba SOC. Its strength $\alpha_R \propto \kappa$ (curvature) and shifting electronic levels $\Delta \propto \kappa^2$. Despite the vanishing integral curvature of typical topographies, implying a net-zero Rashba effect, our two-band analysis and electronic structure calculation of a bent 2D MoTe$_2$ show that only an interplay of $\alpha_R$ and $\Delta$ modulations results in large unidirectional Rashba SOC with well-isolated states. Their high spin-splitting $\sim 0.16$ eV, and attractively small spin-precession length $\sim 1$ nm, are among the best known. Our work uncovers major physical effects of undulations on Rashba SOC in 2D materials, opening new avenues for using their topographical deformation for spintronics and quantum computing.

cond-mat.mtrl-sci

Taming the Tail: Leveraging Asymmetric Loss and Pade Approximation to Overcome Medical Image Long-Tailed Class Imbalance

Long-tailed problems in healthcare emerge from data imbalance due to variability in the prevalence and representation of different medical conditions, warranting the requirement of precise and dependable classification methods. Traditional loss functions such as cross-entropy and binary cross-entropy are often inadequate due to their inability to address the imbalances between the classes with high representation and the classes with low representation found in medical image datasets. We introduce a novel polynomial loss function based on Pade approximation, designed specifically to overcome the challenges associated with long-tailed classification. This approach incorporates asymmetric sampling techniques to better classify under-represented classes. We conducted extensive evaluations on three publicly available medical datasets and a proprietary medical dataset. Our implementation of the proposed loss function is open-sourced in the public repository:https://github.com/ipankhi/ALPA.

cs.CV

FedStein: Enhancing Multi-Domain Federated Learning Through James-Stein Estimator

Federated Learning (FL) facilitates data privacy by enabling collaborative in-situ training across decentralized clients. Despite its inherent advantages, FL faces significant challenges of performance and convergence when dealing with data that is not independently and identically distributed (non-i.i.d.). While previous research has primarily addressed the issue of skewed label distribution across clients, this study focuses on the less explored challenge of multi-domain FL, where client data originates from distinct domains with varying feature distributions. We introduce a novel method designed to address these challenges FedStein: Enhancing Multi-Domain Federated Learning Through the James-Stein Estimator. FedStein uniquely shares only the James-Stein (JS) estimates of batch normalization (BN) statistics across clients, while maintaining local BN parameters. The non-BN layer parameters are exchanged via standard FL techniques. Extensive experiments conducted across three datasets and multiple models demonstrate that FedStein surpasses existing methods such as FedAvg and FedBN, with accuracy improvements exceeding 14% in certain domains leading to enhanced domain generalization. The code is available at https://github.com/sunnyinAI/FedStein

cs.LG

FLeNS: Federated Learning with Enhanced Nesterov-Newton Sketch

Federated learning faces a critical challenge in balancing communication efficiency with rapid convergence, especially for second-order methods. While Newton-type algorithms achieve linear convergence in communication rounds, transmitting full Hessian matrices is often impractical due to quadratic complexity. We introduce Federated Learning with Enhanced Nesterov-Newton Sketch (FLeNS), a novel method that harnesses both the acceleration capabilities of Nesterov's method and the dimensionality reduction benefits of Hessian sketching. FLeNS approximates the centralized Newton's method without relying on the exact Hessian, significantly reducing communication overhead. By combining Nesterov's acceleration with adaptive Hessian sketching, FLeNS preserves crucial second-order information while preserving the rapid convergence characteristics. Our theoretical analysis, grounded in statistical learning, demonstrates that FLeNS achieves super-linear convergence rates in communication rounds - a notable advancement in federated optimization. We provide rigorous convergence guarantees and characterize tradeoffs between acceleration, sketch size, and convergence speed. Extensive empirical evaluation validates our theoretical findings, showcasing FLeNS's state-of-the-art performance with reduced communication requirements, particularly in privacy-sensitive and edge-computing scenarios. The code is available at https://github.com/sunnyinAI/FLeNS

cs.LG

CCVA-FL: Cross-Client Variations Adaptive Federated Learning for Medical Imaging

Federated Learning (FL) offers a privacy-preserving approach to train models on decentralized data. Its potential in healthcare is significant, but challenges arise due to cross-client variations in medical image data, exacerbated by limited annotations. This paper introduces Cross-Client Variations Adaptive Federated Learning (CCVA-FL) to address these issues. CCVA-FL aims to minimize cross-client variations by transforming images into a common feature space. It involves expert annotation of a subset of images from each client, followed by the selection of a client with the least data complexity as the target. Synthetic medical images are then generated using Scalable Diffusion Models with Transformers (DiT) based on the target client's annotated images. These synthetic images, capturing diversity and representing the original data, are shared with other clients. Each client then translates its local images into the target image space using image-to-image translation. The translated images are subsequently used in a federated learning setting to develop a server model. Our results demonstrate that CCVA-FL outperforms Vanilla Federated Averaging by effectively addressing data distribution differences across clients without compromising privacy.

cs.CV

Effect of cation-disorder on lithium transport in halide superionic conductors

Li$_2$ZrCl$_6$ (LZC) is a promising solid-state electrolyte due to its affordability, moisture stability, and high ionic conductivity. We computationally investigate the role of cation disorder in LZC and its effect on Li-ion transport by integrating thermodynamic and kinetic modeling. The results demonstrate that fast Li-ion conductivity requires Li/vacancy disorder, which is dependent on the degree of Zr disorder. The high temperature required to form equilibrium Zr-disorder precludes any equilibrium synthesis processes for achieving fast Li-ion conductivity, rationalizing why only non-equilibrium synthesis methods, such as ball milling, lead to good conductivity. Our simulations show that Zr disorder lowers the Li/vacancy order-disorder transition temperature, which is necessary for creating high Li diffusivity at room temperature. These insights raise a challenge for the large-scale production of these materials and the potential for the long-term stability of their properties.

cond-mat.mtrl-sci

What dictates soft clay-like Lithium superionic conductor formation from rigid-salts mixture

Soft clay-like Li-superionic conductors have been recently synthesized by mixing rigid-salts. Through computational and experimental analysis, we clarify how a soft clay-like material can be created from a mixture of rigid-salts. Using molecular dynamics simulations with a deep learning-based interatomic potential energy model, we uncover the microscopic features responsible for soft clay-formation from ionic solid mixtures. We find that salt mixtures capable of forming molecular solid units on anion exchange, along with the slow kinetics of such reactions, are key to soft-clay formation. Molecular solid units serve as sites for shear transformation zones, and their inherent softness enables plasticity at low stress. Extended X-ray absorption fine structure spectroscopy confirms the formation of molecular solid units. A general strategy for creating soft clay-like materials from ionic solid mixtures is formulated.

cond-mat.mtrl-sci

Designing 1D correlated-electron states by non-Euclidean topography of 2D monolayers

Two-dimensional (2D) bilayers, twisted to particular angles to display electronic flat bands, are being extensively explored for physics of strongly correlated 2D systems. However, the similar rich physics of one-dimensional (1D) strongly correlated systems remains elusive as it is largely inaccessible by twists. Here, a distinctive way to create 1D flat bands is proposed, by either stamping or growing a 2D monolayer on a non-Euclidean topography-patterned surface. Using boron nitride (hBN) as an example, our analysis employing elastic plate theory, density-functional and coarse-grained tight-binding method reveals that hBN's bi-periodic sinusoidal deformation creates pseudo-electric and magnetic fields with unexpected spatial dependence. A combination of these fields leads to anisotropic confinement and 1D flat bands. Moreover, changing the periodic undulations can tune the bandwidth, to drive the system to different strongly correlated regimes such as density waves, Luttinger liquid, and Mott insulator. The 1D nature of these states differs from those obtained in twisted materials and can be exploited to study the exciting physics of 1D quantum systems.

cond-mat.mtrl-sci