SearcharxivSearch

arXiv subjects

Mikhail Ivanov

Publications and source records attributed to Mikhail Ivanov.

At least 19 recordsLinked to original sources

CORDEX-ML-Bench: A Benchmark for Data-Driven Regional Climate Downscaling -Experiment Design and Overview

Machine learning (ML) has emerged as a cost-effective approach to complement dynamical downscaling for producing high-resolution regional climate projections. However, the absence of standardised training and evaluation protocols, applied consistently across multiple domains, continues to hinder meaningful model intercomparison. We introduce CORDEX-ML-Bench, a benchmark aligned with CORDEX, which constitutes the first phase of a community initiative to advance data-driven downscaling toward operational readiness, and complement future dynamical downscaling efforts under CMIP7. The framework targets downscaled daily maximum temperature and precipitation to ~10 km resolution (20x increase) across three pilot regions; European Alps, New Zealand, and Southern Africa. Using a perfect-model experimental design, we evaluate 40 ML configurations developed independently, spanning traditional ML, convolutional U-Nets, vision transformers, graph neural networks, and generative models based on diffusion, flow matching, and generative adversarial networks. Models are trained under two experimental periods, an empirical-statistical downscaling pseudo-reality (historical period only) and Emulator (historical and future periods) -and are evaluated against a core set of metrics developed specifically for assessing downscaling skill. Generative models consistently outperform deterministic approaches for precipitation, better capturing fine-scale variability and extremes. For temperature, the generative advantage narrows and deterministic architectures remain competitive. Models trained solely on the historical period systematically underestimate future climate-change signals while those additionally trained on a future period perform better. These findings raise concerns about historically trained models widely used in an operational setting, underscoring the need for rigorous extrapolation testing.

physics.ao-ph

Extraction of informative statistical features in the problem of forecasting time series generated by It{\^{o}}-type processes

In this paper, we consider the problem of extraction of most informative features from time series that are regarded as observed values of stochastic processes satisfying the It{\^{o}} stochastic differential equations with unknown random drift and diffusion coefficients. We do not attract any additional information and use only the information contained in the time series as it is. Therefore, as additional features, we use the parameters of statistically adjusted mixture-type models of the observed regularities of the behavior of the time series. Several algorithms of construction of these parameters are discussed. These algorithms are based on statistical reconstruction of the coefficients which, in turn, is based on statistical separation of normal mixtures. We obtain two types of parameters by the techniques of the uniform and non-uniform statistical reconstruction of the coefficients of the underlying It{\^{o}} process. The reconstructed coefficients obtained by uniform techniques do not depend on the current value of the process, while the non-uniform techniques reconstruct the coefficients with the account of their dependence on the value of the process. Actually, the non-uniform techniques used in this paper represent a stochastic analog of the Taylor expansion for the time series. The efficiency of the obtained additional features is compared by using them in the autoregressive algorithms of prediction of time series. In order to obtain pure conclusion that is not affected by unwanted factors, say, related to a special choice of the architecture of the neural network prediction methods, we used only simple autoregressive algorithms. We show that the use of additional statistical features improves the prediction.

stat.ML

Climate Downscaling with Stochastic Interpolants (CDSI)

Global climate projections rely on computationally demanding Earth System Models (ESMs), which are typically limited to coarse spatial resolutions due to their high cost. To obtain high-resolution projections for regions of interest, it is common to use Regional Climate Models (RCMs), which are driven by data produced by ESMs as boundary conditions. While more efficient than running ESMs at fine resolution, RCMs remain expensive and restrict the size of ensemble simulations. Inspired by recent advances in probabilistic machine learning for weather and climate, we introduce a data-driven climate downscaling method based on stochastic interpolants. Our approach efficiently transforms coarse ESM output into high-resolution regional climate projections at a fraction of the computational cost of traditional RCMs. Through extensive validation, we demonstrate that our method generates accurate regional ensembles, enabling both improved uncertainty quantification and broader use of high-resolution climate information.

physics.ao-ph

SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks

The rapid advancement of Large Language Models (LLMs) in software engineering has revealed critical limitations in existing benchmarks, particularly the widely used SWE-bench dataset. Recent studies have uncovered severe data contamination issues, e.g., SWE-bench reports 32.67% of successful patches involve direct solution leakage and 31.08% pass due to inadequate test cases. We introduce SWE-MERA, a dynamic, continuously updated benchmark designed to address these fundamental challenges through an automated collection of real-world GitHub issues and rigorous quality validation. Our approach implements a reliable pipeline that ensures quality while minimizing contamination risks, resulting in approximately 10,000 potential tasks with 728 samples currently available. Evaluation using the Aider coding agent demonstrates strong discriminative power in state-of-the-art models. We report performance across a dozen recent LLMs evaluated on tasks collected between September 2024 and June 2025.

cs.SE

In Search Of Lost Tunneling Time

The measurement of tunneling times in strong-field ionization has been the topic of much controversy in recent years, with the attoclock and Larmor clock being two of the main contenders for correctly reproducing these times. By expressing the attoclock as the weak value of temporal delay, we extend its meaning beyond the traditional setup. This allows us to calculate the attoclock time for a static one-dimensional tunneling model consisting of a binding delta potential and a constant electric field. We apply the Steinberg weak-value interpretation of the Larmor clock. Using this definition, we obtain the position-resolved time density during tunnel ionization, yielding a non-zero Larmor tunneling time. Our model allows us to derive the analogue of the position-resolved attoclock tunneling time. While non-zero at the tunnel exit, it vanishes at the detector, far away from the atom. Formally, this means that the attoclock does not measure the "local" Larmor time, but instead a "non-local" time closely related to the phase time.

quant-ph

The Spectroscopic Stage-5 Experiment

The existence, properties, and dynamics of the dark sectors of our universe pose fundamental challenges to our current model of physics, and large-scale astronomical surveys may be our only hope to unravel these long-standing mysteries. In this white paper, we describe the science motivation, instrumentation, and survey plan for the next-generation spectroscopic observatory, the Stage-5 Spectroscopic Experiment (Spec-S5). Spec-S5 is a new all-sky spectroscopic instrument optimized to efficiently carry out cosmological surveys of unprecedented scale and precision. The baseline plan for Spec-S5 involves upgrading two existing 4-m telescopes to new 6-m wide-field facilities, each with a highly multiplexed spectroscopic instrument capable of simultaneously measuring the spectra of 13,000 astronomical targets. Spec-S5, which builds and improves on the hardware used for previous cosmology experiments, represents a cost-effective and rapid approach to realizing a more than 10$\times$ gain in spectroscopic capability compared to the current state-of-the-art represented by the Dark Energy Spectroscopic Instrument project (DESI). Spec-S5 will provide a critical scientific capability in the post-Rubin and post-DESI era for advancing cosmology, fundamental physics, and astrophysics in the 2030s.

astro-ph.CO

Snowmass Theory Frontier: Effective Field Theory

We summarize recent progress in the development, application, and understanding of effective field theories and highlight promising directions for future research. This Report is prepared as the TF02 "Effective Field Theory" topical group summary for the Theory Frontier as part of the Snowmass 2021 process.

hep-ph

The lock-on effect and collapsing bipolar Gunn domains in high-voltage GaAs avalanche p-n junction diode

We present experimental evidence and physics-based simulations of the lock-on effect in high-voltage GaAs avalanche diodes. The avalanche triggering is initiated by steep voltage ramp applied to the diode and in-series 50 Ohm load. After subnanosecond avalanche switching the reversely biased GaAs diode remains in the conducting state for the whole duration of the applied pulse (dozens of nanoseconds). There is no indication of the p-n junction recovery that is commonly expected to develop on the nanosecond scale due to the drift extraction of non-equilibrium carriers. The diode voltage keeps a constant value of ~70 V much lower than the stationary breakdown voltage of 400 V. Numerical simulations reveal that the conducting state is supported by impact ionization in narrow high-field collapsing Gunn domains as well as in quasi-stationary cathode and anode ionizing domains. Collapsing Gunn domains spontaneously appear in the dense electron-hole plasma due to the negative differential mobility of electrons in GaAs. The effect resembles the lock-on effect of GaAs bulk photoconductive switches but is observed in reversely biased p-n junction diode switched by a non-optical method.

cond-mat.mtrl-sci

Characterization of laser-induced ionization dynamics in solid dielectrics

The formation of an electron-hole plasma during the interaction of intense femtosecond laser pulses with transparent solids lies at the heart of femtosecond laser processing. Advanced micro- and nanomachining applications require improved control over the excitation characteristics. Here, we relate the emission of low-order harmonics to the strong laser-field-induced plasma formation. Together with a measurement of the total plasma density we identify the contribution of two competing ionization mechanisms - strong-field and electron-impact ionization.

physics.optics

Unequal Error Protection in Coded Slotted ALOHA

We analyze the performance of coded slotted ALOHA systems for a scenario where users have different error protection requirements and correspondingly can be divided into user classes. The main goal is to design the system so that the requirements for each class are satisfied. To that end, we derive analytical error floor approximations of the packet loss rate for each class in the finite frame length regime, as well as the density evolution in the asymptotic case. Based on this analysis, we propose a heuristic approach for the optimization of the degree distributions to provide the required unequal error protection. In addition, we analyze the decoding delay for users in different classes and show that better protected users experience a smaller average decoding delay.

cs.IT

Broadcast Coded Slotted ALOHA: A Finite Frame Length Analysis

We propose an uncoordinated medium access control (MAC) protocol, called all-to-all broadcast coded slotted ALOHA (B-CSA) for reliable all-to-all broadcast with strict latency constraints. In B-CSA, each user acts as both transmitter and receiver in a half-duplex mode. The half-duplex mode gives rise to a double unequal error protection (DUEP) phenomenon: the more a user repeats its packet, the higher the probability that this packet is decoded by other users, but the lower the probability for this user to decode packets from others. We analyze the performance of B-CSA over the packet erasure channel for a finite frame length. In particular, we provide a general analysis of stopping sets for B-CSA and derive an analytical approximation of the performance in the error floor (EF) region, which captures the DUEP feature of B-CSA. Simulation results reveal that the proposed approximation predicts very well the performance of B-CSA in the EF region. Finally, we consider the application of B-CSA to vehicular communications and compare its performance with that of carrier sense multiple access (CSMA), the current MAC protocol in vehicular networks. The results show that B-CSA is able to support a much larger number of users than CSMA with the same reliability.

cs.IT

On the Information Loss of the Max-Log Approximation in BICM Systems

We present a comprehensive study of the information rate loss of the max-log approximation for $M$-ary pulse-amplitude modulation (PAM) in a bit-interleaved coded modulation (BICM) system. It is widely assumed that the calculation of L-values using the max-log approximation leads to an information loss. We prove that this assumption is correct for all $M$-PAM constellations and labelings with the exception of a symmetric 4-PAM constellation labeled with a Gray code. We also show that for max-log L-values, the BICM generalized mutual information (GMI), which is an achievable rate for a standard BICM decoder, is too pessimistic. In particular, it is proved that the so-called "harmonized" GMI, which can be seen as the sum of bit-level GMIs, is achievable without any modifications to the decoder. We then study how bit-level channel symmetrization and mixing affect the mutual information (MI) and the GMI for max-log L-values. Our results show that these operations, which are often used when analyzing BICM systems, preserve the GMI. However, this is not necessarily the case when the MI is considered. Necessary and sufficient conditions under which these operations preserve the MI are provided.

cs.IT

Probabilistic Handshake in All-to-all Broadcast Coded Slotted ALOHA

We propose a probabilistic handshake mechanism for all-to-all broadcast coded slotted ALOHA. We consider a fully connected network where each user acts as both transmitter and receiver in a half-duplex mode. Users attempt to exchange messages with each other and to establish one-to-one handshakes, in the sense that each user decides whether its packet was successfully received by the other users: After performing decoding, each user estimates in which slots the resolved users transmitted their packets and, based on that, decides if these users successfully received its packet. The simulation results show that the proposed handshake algorithm allows the users to reliably perform the handshake. The paper also provides some analytical bounds on the performance of the proposed algorithm which are in good agreement with the simulation results.

cs.IT

All-to-all Broadcast for Vehicular Networks Based on Coded Slotted ALOHA

We propose an uncoordinated all-to-all broadcast protocol for periodic messages in vehicular networks based on coded slotted ALOHA (CSA). Unlike classical CSA, each user acts as both transmitter and receiver in a half-duplex mode. As in CSA, each user transmits its packet several times. The half-duplex mode gives rise to an interesting design trade-off: the more the user repeats its packet, the higher the probability that this packet is decoded by other users, but the lower the probability for this user to decode packets from others. We compare the proposed protocol with carrier sense multiple access with collision avoidance, currently adopted as a multiple access protocol for vehicular networks. The results show that the proposed protocol greatly increases the number of users in the network that reliably communicate with each other. We also provide analytical tools to predict the performance of the proposed protocol.

cs.IT

Error Floor Analysis of Coded Slotted ALOHA over Packet Erasure Channels

We present a framework for the analysis of the error floor of coded slotted ALOHA (CSA) for finite frame lengths over the packet erasure channel. The error floor is caused by stopping sets in the corresponding bipartite graph, whose enumeration is, in general, not a trivial problem. We therefore identify the most dominant stopping sets for the distributions of practical interest. The derived analytical expressions allow us to accurately predict the error floor at low to moderate channel loads and characterize the unequal error protection inherent in CSA.

cs.IT

On the Asymptotic Performance of Bit-Wise Decoders for Coded Modulation

Two decoder structures for coded modulation over the Gaussian and flat fading channels are studied: the maximum likelihood symbol-wise decoder, and the (suboptimal) bit-wise decoder based on the bit-interleaved coded modulation paradigm. We consider a 16-ary quadrature amplitude constellation labeled by a Gray labeling. It is shown that the asymptotic loss in terms of pairwise error probability, for any two codewords caused by the bit-wise decoder, is bounded by 1.25 dB. The analysis also shows that for the Gaussian channel the asymptotic loss is zero for a wide range of linear codes, including all rate-1/2 convolutional codes.

cs.IT

General BER Expression for One-Dimensional Constellations

A novel general ready-to-use bit-error rate (BER) expression for one-dimensional constellations is developed. The BER analysis is performed for bit patterns that form a labeling. The number of patterns for equally spaced M-PAM constellations with different BER is analyzed.

cs.IT