SearcharxivSearch

arXiv subjects

Igor Sokolov

Publications and source records attributed to Igor Sokolov.

At least 19 recordsLinked to original sources

SILAGE: Memory-Efficient, Full-Gradient-Free Nonconvex Optimization for Nested Finite Sums

Empirical risk minimization on massive datasets naturally exhibits a nested double finite-sum structure, where $N=nm$ total samples are logically or physically partitioned into $n$ blocks of size $m$ (e.g., in pooled data silos, out-of-core learning, or deliberate stratification). While variance-reduced methods achieve optimal oracle complexities for nonconvex objectives, they suffer from severe scaling bottlenecks in this centralized regime. Recursive estimators, such as PAGE, require periodic global full-gradient refreshes over all $nm$ samples, which are computationally expensive. Conversely, single-loop methods, such as SILVER, avoid such refreshes but require an impractical $\mathcal{O}(nm)$ memory footprint to store a control variate for every sample. In this paper, we propose SILAGE, a variance-reduced algorithm that addresses this trade-off. By actively exploiting the double-sum structure, SILAGE eliminates periodic global full-gradient refreshes over all $nm$ components (evaluating at most one local group gradient per iteration) while requiring only $\mathcal{O}(n)$ memory. Furthermore, we provide a tight convergence analysis that avoids pessimistic worst-case Lipschitz constants. Instead, SILAGE's complexity natively adapts to the underlying data geometry via nested functional similarities: across-group ($\delta_1$) and within-group ($\delta_2$) heterogeneity. Our results improve existing state-of-the-art bounds in several practically relevant regimes.

cs.LG

Improved Convergence in Parameter-Agnostic Error Feedback through Momentum

Communication compression is essential for scalable distributed training of modern machine learning models, but it often degrades convergence due to the noise it introduces. Error Feedback (EF) mechanisms are widely adopted to mitigate this issue of distributed compression algorithms. Despite their popularity and training efficiency, existing distributed EF algorithms often require prior knowledge of problem parameters (e.g., smoothness constants) to fine-tune stepsizes. This limits their practical applicability especially in large-scale neural network training. In this paper, we study normalized error feedback algorithms that combine EF with normalized updates, various momentum variants, and parameter-agnostic, time-varying stepsizes, thus eliminating the need for problem-dependent tuning. We analyze the convergence of these algorithms for minimizing smooth functions, and establish parameter-agnostic complexity bounds that are close to the best-known bounds with carefully-tuned problem-dependent stepsizes. Specifically, we show that normalized EF21 achieve the convergence rate of near ${O}(1/T^{1/4})$ for Polyak's heavy-ball momentum, ${O}(1/T^{2/7})$ for Iterative Gradient Transport (IGT), and ${O}(1/T^{1/3})$ for STORM and Hessian-corrected momentum. Our results hold with decreasing stepsizes and small mini-batches. Finally, our empirical experiments confirm our theoretical insights.

math.OC

Bernoulli-LoRA: A Theoretical Framework for Randomized Low-Rank Adaptation

Parameter-efficient fine-tuning (PEFT) has emerged as a crucial approach for adapting large foundational models to specific tasks, particularly as model sizes continue to grow exponentially. Among PEFT methods, Low-Rank Adaptation (LoRA) (arXiv:2106.09685) stands out for its effectiveness and simplicity, expressing adaptations as a product of two low-rank matrices. While extensive empirical studies demonstrate LoRA's practical utility, theoretical understanding of such methods remains limited. Recent work on RAC-LoRA (arXiv:2410.08305) took initial steps toward rigorous analysis. In this work, we introduce Bernoulli-LoRA, a novel theoretical framework that unifies and extends existing LoRA approaches. Our method introduces a probabilistic Bernoulli mechanism for selecting which matrix to update. This approach encompasses and generalizes various existing update strategies while maintaining theoretical tractability. Under standard assumptions from non-convex optimization literature, we analyze several variants of our framework: Bernoulli-LoRA-GD, Bernoulli-LoRA-SGD, Bernoulli-LoRA-PAGE, Bernoulli-LoRA-MVR, Bernoulli-LoRA-QGD, Bernoulli-LoRA-MARINA, and Bernoulli-LoRA-EF21, establishing convergence guarantees for each variant. Additionally, we extend our analysis to convex non-smooth functions, providing convergence rates for both constant and adaptive (Polyak-type) stepsizes. Through extensive experiments on various tasks, we validate our theoretical findings and demonstrate the practical efficacy of our approach. This work is a step toward developing theoretically grounded yet practically effective PEFT methods.

cs.LG

Evidence of Time-Dependent Diffusive Shock Acceleration in the 2022 September 5 Solar Energetic Particle Event

On 2022 September 5, a large solar energetic particle (SEP) event was detected by Parker Solar Probe (PSP) and Solar Orbiter (SolO), at heliocentric distances of 0.07 and 0.71 au, respectively. PSP observed an unusual velocity-dispersion signature: particles below $\sim$1 MeV exhibited a normal velocity dispersion, while higher-energy particles displayed an inverse velocity arrival feature, with the most energetic particles arriving later than those at lower energies. The maximum energy increased from about 20-30 MeV upstream to over 60 MeV downstream of the shock. The arrival of SEPs at PSP was significantly delayed relative to the expected onset of the eruption. In contrast, SolO detected a typical large SEP event characterized by a regular velocity dispersion at all energies up to 100 MeV. To understand these features, we simulate particle acceleration and transport from the shock to the observers with our newly developed SEP model - Particle ARizona and MIchigan Solver on Advected Nodes (PARMISAN). Our results reveal that the inverse velocity arrival and delayed particle onset detected by PSP originate from the time-dependent diffusive shock acceleration processes. After shock passage, PSP's magnetic connectivity gradually shifted due to its high velocity near perihelion, detecting high-energy SEPs streaming sunward. Conversely, SolO maintained a stable magnetic connection to the strong shock region where efficient acceleration was achieved. These results underscore the importance of spatial and temporal dependence in SEP acceleration at interplanetary shocks, and provide new insights to understand SEP variations in the inner heliosphere.

astro-ph.SR

EF21 with Bells & Whistles: Six Algorithmic Extensions of Modern Error Feedback

First proposed by Seide (2014) as a heuristic, error feedback (EF) is a very popular mechanism for enforcing convergence of distributed gradient-based optimization methods enhanced with communication compression strategies based on the application of contractive compression operators. However, existing theory of EF relies on very strong assumptions (e.g., bounded gradients), and provides pessimistic convergence rates (e.g., while the best known rate for EF in the smooth nonconvex regime, and when full gradients are compressed, is $O(1/T^{2/3})$, the rate of gradient descent in the same regime is $O(1/T)$). Recently, Richtárik et al. (2021) proposed a new error feedback mechanism, EF21, based on the construction of a Markov compressor induced by a contractive compressor. EF21 removes the aforementioned theoretical deficiencies of EF and at the same time works better in practice. In this work we propose six practical extensions of EF21, all supported by strong convergence theory: partial participation, stochastic approximation, variance reduction, proximal setting, momentum, and bidirectional compression. To the best of our knowledge, several of these techniques have not been previously analyzed in combination with EF, and in cases where prior analysis exists -- such as for bidirectional compression -- our theoretical convergence guarantees significantly improve upon existing results.

cs.LG

Single Ultrabright Fluorescent Silica Nanoparticles Can Be Used as Individual Fast Real-Time Nanothermometers

Optical-based nanothermometry represents a transformative approach for precise temperature measurements at the nanoscale, which finds versatile applications across biology, medicine, and electronics. The assembly of ratiometric fluorescent 40 nm nanoparticles designed to serve as individual nanothermometers is introduced here. These nanoparticles exhibit unprecedented sensitivity (11% /K) and temperature resolution 128 \, \mathrm{K} \cdot \mathrm{Hz}^{-1/2} \cdot \mathrm{W} \cdot \mathrm{cm}^{-2}, outperforming existing optical nanothermometers by factors of 2-6 and 455, respectively. The enhanced performance is attributed to the encapsulation of fluorescent molecules with high density inside the mesoporous matrix. It becomes possible after incorporating hydrophobic groups into the silica matrix, which effectively prevents water ingress and dye leaking. A practical application of these nanothermometers is demonstrated using confocal microscopy, showcasing their ability to map temperature distributions accurately. This methodology is compatible with any fluorescent microscope capable of recording dual fluorescent channels in any transparent medium or on a sample surface. This work not only sets a new benchmark for optical nano-thermometry but also provides a relatively simple yet powerful tool for exploring thermal phenomena at the nanoscale across various scientific domains.

physics.optics

MARINA-P: Superior Performance in Non-smooth Federated Optimization with Adaptive Stepsizes

Non-smooth communication-efficient federated optimization is crucial for many machine learning applications, yet remains largely unexplored theoretically. Recent advancements have primarily focused on smooth convex and non-convex regimes, leaving a significant gap in understanding the non-smooth convex setting. Additionally, existing literature often overlooks efficient server-to-worker communication (downlink), focusing primarily on worker-to-server communication (uplink). We consider a setup where uplink costs are negligible and focus on optimizing downlink communication by improving state-of-the-art schemes like EF21-P (arXiv:2209.15218) and MARINA-P (arXiv:2402.06412) in the non-smooth convex setting. We extend the non-smooth convex theory of EF21-P [Anonymous, 2024], originally developed for single-node scenarios, to the distributed setting, and extend MARINA-P to the non-smooth convex setting. For both algorithms, we prove an optimal $O(1/\sqrt{T})$ convergence rate and establish communication complexity bounds matching classical subgradient methods. We provide theoretical guarantees under constant, decreasing, and adaptive (Polyak-type) stepsizes. Our experiments demonstrate that MARINA-P with correlated compressors outperforms other methods in both smooth non-convex and non-smooth convex settings. This work presents the first theoretical results for distributed non-smooth optimization with server-to-worker compression, along with comprehensive analysis for various stepsize schemes.

cs.LG

Cohort Squeeze: Beyond a Single Communication Round per Cohort in Cross-Device Federated Learning

Virtually all federated learning (FL) methods, including FedAvg, operate in the following manner: i) an orchestrating server sends the current model parameters to a cohort of clients selected via certain rule, ii) these clients then independently perform a local training procedure (e.g., via SGD or Adam) using their own training data, and iii) the resulting models are shipped to the server for aggregation. This process is repeated until a model of suitable quality is found. A notable feature of these methods is that each cohort is involved in a single communication round with the server only. In this work we challenge this algorithmic design primitive and investigate whether it is possible to ``squeeze more juice" out of each cohort than what is possible in a single communication round. Surprisingly, we find that this is indeed the case, and our approach leads to up to 74% reduction in the total communication cost needed to train a FL model in the cross-device setting. Our method is based on a novel variant of the stochastic proximal point method (SPPM-AS) which supports a large collection of client sampling procedures some of which lead to further gains when compared to classical client selection approaches.

cs.LG

Quantum Positional Encodings for Graph Neural Networks

In this work, we propose novel families of positional encodings tailored to graph neural networks obtained with quantum computers. These encodings leverage the long-range correlations inherent in quantum systems that arise from mapping the topology of a graph onto interactions between qubits in a quantum computer. Our inspiration stems from the recent advancements in quantum processing units, which offer computational capabilities beyond the reach of classical hardware. We prove that some of these quantum features are theoretically more expressive for certain graphs than the commonly used relative random walk probabilities. Empirically, we show that the performance of state-of-the-art models can be improved on standard benchmarks and large-scale datasets by computing tractable versions of quantum features. Our findings highlight the potential of leveraging quantum computing capabilities to enhance the performance of transformers in handling graph data.

quant-ph

On machine learning analysis of atomic force microscopy images for image classification, sample surface recognition

Atomic force microscopy (AFM or SPM) imaging is one of the best matches with machine learning (ML) analysis among microscopy techniques. The digital format of AFM images allows for direct utilization in ML algorithms without the need for additional processing. Additionally, AFM enables the simultaneous imaging of distributions of over a dozen different physicochemical properties of sample surfaces, a process known as multidimensional imaging. While this wealth of information can be challenging to analyze using traditional methods, ML provides a seamless approach to this task. However, the relatively slow speed of AFM imaging poses a challenge in applying deep learning methods broadly used in image recognition. This Prospective is focused on ML recognition/classification when using a relatively small number of AFM images, small database. We discuss ML methods other than popular deep-learning neural networks. The described approach has already been successfully used to analyze and classify the surfaces of biological cells. It can be applied to recognize medical images, specific material processing, in forensic studies, even to identify the authenticity of arts. A general template for ML analysis specific to AFM is suggested, with a specific example of the identification of cell phenotype. Special attention is given to the analysis of the statistical significance of the obtained results, an important feature that is often overlooked in papers dealing with machine learning. A simple method for finding statistical significance is also described.

physics.bio-ph

Enhancing Graph Neural Networks with Quantum Computed Encodings

Transformers are increasingly employed for graph data, demonstrating competitive performance in diverse tasks. To incorporate graph information into these models, it is essential to enhance node and edge features with positional encodings. In this work, we propose novel families of positional encodings tailored for graph transformers. These encodings leverage the long-range correlations inherent in quantum systems, which arise from mapping the topology of a graph onto interactions between qubits in a quantum computer. Our inspiration stems from the recent advancements in quantum processing units, which offer computational capabilities beyond the reach of classical hardware. We prove that some of these quantum features are theoretically more expressive for certain graphs than the commonly used relative random walk probabilities. Empirically, we show that the performance of state-of-the-art models can be improved on standard benchmarks and large-scale datasets by computing tractable versions of quantum features. Our findings highlight the potential of leveraging quantum computing capabilities to potentially enhance the performance of transformers in handling graph data.

quant-ph

Solar Wind with Field Lines and Energetic Particles (SOFIE) Model: Application to Historical Solar Energetic Particle Events

In this paper, we demonstrate the applicability of the data-driven and self-consistent solar energetic particle model, Solar-wind with FIeld-lines and Energetic-particles (SOFIE), to simulate acceleration and transport processes of solar energetic particles. SOFIE model is built upon the Space Weather Modeling Framework (SWMF) developed at the University of Michigan. In SOFIE, the background solar wind plasma in the solar corona and interplanetary space is calculated by the Aflvén Wave Solar-atmosphere Model(-Realtime) (AWSoM-R) driven by the near-real-time hourly updated Global Oscillation Network Group (GONG) solar magnetograms. In the background solar wind, coronal mass ejections (CMEs) are launched by placing an imbalanced magnetic flux rope on top of the parent active region, using the Eruptive Event Generator using Gibson-Low model (EEGGL). The acceleration and transport processes are modeled by the Multiple-Field-Line Advection Model for Particle Acceleration (M-FLAMPA). In this work, nine solar energetic particle events (Solar Heliospheric and INterplanetary Environment (SHINE) challenge/campaign events) are modeled. The three modules in SOFIE are validated and evaluated by comparing with observations, including the steady-state background solar wind properties, the white-light image of the CME, and the flux of solar energetic protons, at energies of > 10 MeV.

astro-ph.SR

A Guide Through the Zoo of Biased SGD

Stochastic Gradient Descent (SGD) is arguably the most important single algorithm in modern machine learning. Although SGD with unbiased gradient estimators has been studied extensively over at least half a century, SGD variants relying on biased estimators are rare. Nevertheless, there has been an increased interest in this topic in recent years. However, existing literature on SGD with biased estimators (BiasedSGD) lacks coherence since each new paper relies on a different set of assumptions, without any clear understanding of how they are connected, which may lead to confusion. We address this gap by establishing connections among the existing assumptions, and presenting a comprehensive map of the underlying relationships. Additionally, we introduce a new set of assumptions that is provably weaker than all previous assumptions, and use it to present a thorough analysis of BiasedSGD in both convex and non-convex settings, offering advantages over previous results. We also provide examples where biased estimators outperform their unbiased counterparts or where unbiased versions are simply not available. Finally, we demonstrate the effectiveness of our framework through experimental results that validate our theoretical findings.

cs.LG

Modeling the Solar Wind During Different Phases of the Last Solar Cycle

We describe our first attempt to systematically simulate the solar wind during different phases of the last solar cycle with the Alfvén Wave Solar atmosphere Model (AWSoM) developed at the University of Michigan. Key to this study is the determination of the optimal values of one of the most important input parameters of the model, the Poynting flux parameter, which prescribes the energy flux passing through the chromospheric boundary of the model in the form of Alfvén wave turbulence. It is found that the optimal value of the Poynting flux parameter is correlated with the area of the open magnetic field regions with the Spearman's correlation coefficient of 0.96 and anti-correlated with the average unsigned radial component of the magnetic field with the Spearman's correlation coefficient of -0.91. Moreover, the Poynting flux in the open field regions is approximately constant in the last solar cycle, which needs to be validated with observations and can shed light on how Alfvén wave turbulence accelerates the solar wind during different phases of the solar cycle. Our results can also be used to set the Poynting flux parameter for real-time solar wind simulations with AWSoM.

astro-ph.SR

Three-dimensional, Time-dependent MHD Simulation of Disk-Magnetosphere-Stellar Wind Interaction in a T Tauri, Protoplanetary System

We present a three-dimensional, time-dependent, MHD simulation of the short-term interaction between a protoplanetary disk and the stellar corona in a T Tauri system. The simulation includes the stellar magnetic field, self-consistent coronal heating and stellar wind acceleration, and a disk rotating at sub-Keplerian velocity to induce accretion. We find that initially, as the system relaxes from the assumed initial conditions, the inner part of the disk winds around and moves inward and close to the star as expected. However, the self-consistent coronal heating and stellar wind acceleration build up the original state after some time, significantly pushing the disk out beyond $10R_\star$. After this initial relaxation period, we do not find clear evidence of a strong, steady accretion flow funneled along coronal field lines, but only weak, sporadic accretion. We produce synthetic coronal X-ray line emission light curves which show flare-like increases that are not correlated with accretion events nor with heating events. These variations in the line emission flux are the result of compression and expansion due to disk-corona pressure variations. Vertical disk evaporation evolves above and below the disk. However, the disk - stellar wind boundary stays quite stable, and any disk material that reaches the stellar wind region is advected out by the stellar wind.

astro-ph.SR

Solar wind modeling with the Alfven Wave Solar atmosphere Model driven by HMI-based Near-Real-Time maps by the National Solar Observatory

We explore model performance for the Alfven Wave Solar atmosphere Model (AWSoM) with near-real-time (NRT) synoptic maps of the photospheric vector magnetic field. These maps, produced by assimilating data from the Helioseismic Magnetic Imager (HMI) onboard the Solar Dynamics Observatory (SDO), use a different method developed at the National Solar Observatory (NSO) to provide a near contemporaneous source of data to drive numerical models. Here, we apply these NSO-HMI-NRT maps to simulate three Carrington rotations (CRs): 2107-2108 (centered on 2011/03/07 20:12 CME event), 2123 (integer CR) and 2218--2219 (centered on 2019/07/2 solar eclipse), which together cover a wide range of activity level for solar cycle 24. We show simulation results, which reproduce both extreme ultraviolet emission (EUV) from the low corona while simultaneously matching in situ observations at 1 au as well as quantify the total unsigned open magnetic flux from these maps.

astro-ph.SR

Federated Optimization Algorithms with Random Reshuffling and Gradient Compression

Gradient compression is a popular technique for improving communication complexity of stochastic first-order methods in distributed training of machine learning models. However, the existing works consider only with-replacement sampling of stochastic gradients. In contrast, it is well-known in practice and recently confirmed in theory that stochastic methods based on without-replacement sampling, e.g., Random Reshuffling (RR) method, perform better than ones that sample the gradients with-replacement. In this work, we close this gap in the literature and provide the first analysis of methods with gradient compression and without-replacement sampling. We first develop a naïve combination of random reshuffling with gradient compression (Q-RR). Perhaps surprisingly, but the theoretical analysis of Q-RR does not show any benefits of using RR. Our extensive numerical experiments confirm this phenomenon. This happens due to the additional compression variance. To reveal the true advantages of RR in the distributed learning with compression, we propose a new method called DIANA-RR that reduces the compression variance and has provably better convergence rates than existing counterparts with with-replacement sampling of stochastic gradients. Next, to have a better fit to Federated Learning applications, we incorporate local computation, i.e., we propose and analyze the variants of Q-RR and DIANA-RR -- Q-NASTYA and DIANA-NASTYA that use local gradient steps and different local and global stepsizes. Finally, we conducted several numerical experiments to illustrate our theoretical results.

cs.LG

Application of the Monte Carlo Method in Modeling Transport and Acceleration of Solar Energetic Particles

The need for quantitative characterization of the solar energetic particle (SEP) dynamics goes beyond being an academic discipline only. It has numerous practical implications related to human activity in space. The terrestrial magnetic field shields the International Space Station (ISS) and most uncrewed missions from exposure to SEP radiation. However, extreme SEP events with hard energy spectra are particularly rich in hundreds of MeV to several GeV protons that can reach the altitudes of the Low Earth Orbit (LEO). These protons have a high penetrating capability, thus producing significant radiation hazards for human spaceflight. SEPs also have a significant effect on the atmosphere. Sudden ionization of the upper atmosphere at high latitudes that occurs during polar cap absorption (PCA) events can block high frequency (HF) communication for hours, affecting communication with aircraft on intercontinental high-altitude flights. Another effect of SEPs in the atmosphere is creating NOx molecules in the upper atmosphere that can deplete the atmospheric ozone population. The paper also presents an analysis of (1) how various pitch angle diffusion coefficient approximations affect the properties of the simulated SEPs population and (2) discusses how pitch angle scattering when SEPs are beyond 1 AU affects a SEP event decay phase at the Earth's orbit.

physics.space-ph