SearcharxivSearch

arXiv subjects

Andrey Goncharov

Publications and source records attributed to Andrey Goncharov.

8 recordsLinked to original sources

PRAGMA: Revolut Foundation Model

Modern financial systems generate vast quantities of transactional and event-level data that encode rich economic signals. This paper presents PRAGMA, a family of foundation models for banking event sequences. Our approach pre-trains a Transformer-based architecture with masked modelling on a large-scale, heterogeneous banking event corpus using a self-supervised objective tailored to the discrete, variable-length nature of financial records. The resulting model supports a wide range of downstream tasks such as credit scoring, fraud detection, and lifetime value prediction: strong performance can be achieved by training a simple linear model on top of the extracted embeddings and can be further improved with lightweight fine-tuning. Through extensive evaluation on downstream tasks, we demonstrate that PRAGMA achieves superior performance across multiple domains directly from raw event sequences, providing a general-purpose representation layer for financial applications.

cs.LG

Language steering in latent space to mitigate unintended code-switching

Multilingual Large Language Models (LLMs) often exhibit hallucinations such as unintended code-switching, reducing reliability in downstream tasks. We propose latent-space language steering, a lightweight inference-time method that identifies language directions via Principal Component Analysis (PCA) on parallel translations and steers token embeddings along these axes to control language identity. Our approach mitigates code-switching while preserving semantics with negligible computational overhead and requires only minimal parallel data for calibration. Empirically, we achieve 95-99\% language classification accuracy using a single principal component and reduce next-token distributional divergence by up to 55\% across multiple language pairs on Qwen2.5 and Llama-3.2 models. Generation-based evaluation on Llama-3.2 further demonstrates 63--99\% reduction in Code-Switching Index across four language pairs ($p < 0.001$). We further analyze the layer-wise evolution of language representations, revealing that language identity concentrates in final layers with near-perfect linear separability.

cs.CL

Enhancing the Reliability of Medical AI through Expert-guided Uncertainty Modeling

Artificial intelligence (AI) systems accelerate medical workflows and improve diagnostic accuracy in healthcare, serving as second-opinion systems. However, the unpredictability of AI errors poses a significant challenge, particularly in healthcare contexts, where mistakes can have severe consequences. A widely adopted safeguard is to pair predictions with uncertainty estimation, enabling human experts to focus on high-risk cases while streamlining routine verification. Current uncertainty estimation methods, however, remain limited, particularly in quantifying aleatoric uncertainty, which arises from data ambiguity and noise. To address this, we propose a novel approach that leverages disagreement in expert responses to generate targets for training machine learning models. These targets are used in conjunction with standard data labels to estimate two components of uncertainty separately, as given by the law of total variance, via a two-ensemble approach, as well as its lightweight variant. We validate our method on binary image classification, binary and multi-class image segmentation, and multiple-choice question answering. Our experiments demonstrate that incorporating expert knowledge can enhance uncertainty estimation quality by $9\%$ to $50\%$ depending on the task, making this source of information invaluable for the construction of risk-aware AI systems in healthcare applications.

cs.LG

Complexity-aware fine-tuning

General-purpose Large Language Models (LLMs) are frequently fine-tuned through supervised fine-tuning (SFT) to enhance performance in specific domains. Better results can be achieved by distilling the chain-of-thought of a larger model at the cost of numerous expensive calls and a much greater amount of data. We propose a novel blueprint for efficient fine-tuning that uses reasoning only for complex data identified by entropy. Specifically, across three small open models ($\approx 3B$) we split the training data into complexity categories by a single token answer entropy (ROC AUC $0.73$), fine-tune large language models (LLMs) via SFT and distillation, and show that our pipeline significantly outperforms the standard SFT approach ($0.58$ vs $0.45$ average accuracy) and outperforms the distillation approach ($0.58$ vs $0.56$ average accuracy) while using $81\%$ less data.

cs.LG

When an LLM is apprehensive about its answers -- and when its uncertainty is justified

Uncertainty estimation is crucial for evaluating Large Language Models (LLMs), particularly in high-stakes domains where incorrect answers result in significant consequences. Numerous approaches consider this problem, while focusing on a specific type of uncertainty, ignoring others. We investigate what estimates, specifically token-wise entropy and model-as-judge (MASJ), would work for multiple-choice question-answering tasks for different question topics. Our experiments consider three LLMs: Phi-4, Mistral, and Qwen of different sizes from 1.5B to 72B and $14$ topics. While MASJ performs similarly to a random error predictor, the response entropy predicts model error in knowledge-dependent domains and serves as an effective indicator of question difficulty: for biology ROC AUC is $0.73$. This correlation vanishes for the reasoning-dependent domain: for math questions ROC-AUC is $0.55$. More principally, we found out that the entropy measure required a reasoning amount. Thus, data-uncertainty related entropy should be integrated within uncertainty estimates frameworks, while MASJ requires refinement. Moreover, existing MMLU-Pro samples are biased, and should balance required amount of reasoning for different subdomains to provide a more fair assessment of LLMs performance.

cs.CL

Extending frequency metrology to increasingly complex molecules: SI-traceable sub-Doppler mid-IR spectroscopy of trioxane

Bringing increasingly complex polyatomic molecules within reach of precision measurement experiments offers fascinating and far-reaching prospects ranging from Earth sciences and astrophysics, to metrology and quantum sciences. Here, we demonstrate sub-Doppler spectroscopic measurements in the mid-IR fingerprint region of, to our knowledge, the largest molecule to date. To this end, we use a high-resolution ~10.3 $μ$m spectrometer based on a sub-Hz quantum cascade laser remotely calibrated against state-of-the-art primary frequency standards via a metrology-grade fibre link. We perform saturated absorption spectroscopy in the v5 CO stretching mode of 1,3,5-trioxane, (H2CO)3, at a resolution of ~100 kHz, allowing us to measure the absolute frequency of hundreds of rovibrational transitions at unprecedented uncertainties for such a complex species, as low as ~5 kHz. Our work demonstrates the extension of frequency metrology methodologies to ever larger molecular system, confirming the potential of the technologies we develop for bringing increasingly complex species within reach of ultra-precise measurement experiments.

physics.atom-ph

All-optical atomic magnetometry using an elliptically polarized amplitude-modulated light wave

We study a resonant interaction of an elliptically polarized light wave with $^{87}$Rb vapor (D$_1$ line) exposed to a transverse magnetic field. A $5$$\times$$5$$\times$$5$~mm$^3$ glass vapor cell is used for the experiments. The wave intensity is modulated at the frequency $Ω_m$. By scanning $Ω_m$ near the Larmor frequency $Ω_L$, a magnetic resonance (MR) can be observed as a change in the ellipticity parameter of the wave polarization. This method for observing MR allows to significantly improve the signal-to-noise ratio compared to a classical Bell-Bloom scheme using a circularly polarized wave. The sensitivity of the magnetic field sensor is estimated to be $\approx\,$$130$~fT/$\surd$Hz in a $2$~kHz bandwidth, confidently competing with widely used Faraday-rotation Bell-Bloom schemes. The results can be used to develop a miniature all-optical magnetic field sensor for medicine and geophysics.

physics.atom-ph

A widely tunable 10-$μ$m quantum cascade laser phase-locked to a state-of-the-art mid-infrared reference for precision molecular spectroscopy

We report the coherent phase-locking of a quantum cascade laser (QCL) at 10-$μ$m to the secondary frequency standard of this spectral region, a CO2 laser stabilized on a saturated absorption line of OsO4. The stability and accuracy of the standard are transferred to the QCL resulting in a line width of the order of 10 Hz, and leading to our knowledge to the narrowest QCL to date. The locked QCL is then used to perform absorption spectroscopy spanning 6 GHz of NH3 and methyltrioxorhenium, two species of interest for applications in precision measurements.

physics.optics