Searcharxiv⌕ Search

arXiv subjects

Daniil Volkov

Publications and source records attributed to Daniil Volkov.

6 recordsLinked to original sources

BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models

While Multimodal Large Language Models (MLLMs) have made significant strides in visual comprehension, their ability to reason about text-dense, professional documents remains incompletely evaluated. Existing benchmarks emphasize information extraction, require external domain knowledge, or cover professional documents only as one of many settings. They are also largely English- or Chinese-centric, leaving other languages and Russian, in particular, substantially underrepresented. To address these limitations, we introduce BEAR-Bench (Bilingual Enterprise and Academic Reasoning), a self-contained, complex English-and-Russian benchmark comprising 1000 human-annotated questions based on text-rich business and scientific documents. We evaluate 16 proprietary and open-weight MLLMs, including Gemini 3.1 Pro and Qwen3.5-397B, on BEAR-Bench and observe clear headroom even for the strongest systems. Finally, we use the resulting model outputs to compare existing hallucination detection methods, evaluating not only how often models fail on BEAR-Bench but also how reliably those failures can be identified.

cs.CL↗

Batch Size or Negatives? A Selection Rule for Memory-Constrained Recommender Training

Large-scale neural recommender systems are typically trained with a softmax cross-entropy objective over the full item vocabulary. For a typical large number of possible items $K$, the final classification layer dominates memory, requiring $O(nK)$ logits and gradients to materialize for a batch of $n$ examples. Sampled softmax reduces this cost by restricting the objective to only $k \ll K$ candidate negative items, resulting in an $O(nk)$ memory. However, for a fixed budget $B = n k$, it remains unclear whether one should prioritize larger batches or the inclusion of more negative items. We address this question by analyzing sampled-softmax training under a fixed memory constraint. Under standard smoothness and variance assumptions, our theoretical evidence suggests that the fastest convergence arises from an $ n \sim B, k \sim 1$ allocation. So, an actionable rule is to include as many objects as possible given computational constraints. Our theory is supported by controlled synthetic and synthetic and four real sequential recommendation benchmarks, including MovieLens-20M. The suggested configuration achieve faster convergence and better final recommendation quality than imbalanced alternatives within the same memory constraint. These findings provide a theoretical and empirical foundation for configuring memory during the training of recommender systems. Code, reproducibility materials, and all scripts for generating figures are available at https://anonymous.4open.science/r/LimitedMemoryRule-BBFB

cs.LG↗

Faster and Memory-Efficient Training of Sequential Recommendation Models for Large Catalogs

Sequential recommendations (SR) with transformer-based architectures are widely adopted in real-world applications, where SR models require frequent retraining to adapt to ever-changing user preferences. However, training transformer-based SR models often encounters a high computational cost associated with scoring extensive item catalogs, often exceeding thousands of items. This occurs mainly due to the use of cross-entropy loss, where peak memory scales proportionally to catalog size, batch size, and sequence length. Recognizing this, practitioners in the field of recommendation systems typically address memory consumption by integrating the cross-entropy (CE) loss with negative sampling, thereby reducing the explicit memory demands of the final layer. However, a small number of negative samples would degrade model performance, and as we demonstrate in our work, increasing the number of negative samples and the batch size further improves the model's performance, but rapidly starts to exceed industrial GPUs' size (~40Gb). In this work, we introduce the CCE- method, which offers a GPU-efficient implementation of the CE loss with negative sampling. Our method accelerates training by up to two times while reducing memory consumption by more than 10 times. Leveraging the memory savings afforded by using CCE- for model training, it becomes feasible to enhance its accuracy on datasets with a large item catalog compared to those trained with original PyTorch-implemented loss functions. Finally, we perform an analysis of key memory-related hyperparameters and highlight the necessity of a delicate balance among these factors. We demonstrate that scaling both the number of negative samples and batch size leads to better results rather than maximizing only one of them. To facilitate further adoption of CCE-, we release a Triton kernel that efficiently implements the proposed method.

cs.IR↗

Electron EDM and $Γ(μ\to e γ)$ in the 2HDM

We present the first complete two-loop calculation of the electric dipole moment of the electron, as well as the rates of the lepton-flavor violating decays $μ\to e + γ$ and $τ\to e/μ+ γ$, in the unconstrained two-Higgs doublet model. We include the most general Yukawa interactions of the Higgs doublets with the Standard Model fermions up to quadratic order, and allow for generic phases in the Higgs potential. A python implementation of our results is provided via a public git repository.

hep-ph↗

Hadronic CP Violation in the 2HDM

We present the first complete two-loop calculation of the electric and chromo-electric dipole moments of the light quarks and the gluon, as well as contributions to CP-violating lepton-quark interactions, in the unconstrained two-Higgs doublet model. We include the most general Yukawa interactions of the Higgs doublets with the Standard Model fermions up to quadratic order, and allow for generic phases in the Higgs potential. We pay particular attention to a consistent treatment of all fermionic contributions in the low-energy effective theory, including a consistent renormalization-group summation of all leading-logarithmic effects. This latter part of the work is independent of the specific UV model and can generally be applied to a large class of models that do not introduce new light degrees of freedom. A python implementation of our results is provided via a public git repository.

hep-ph↗

Detection states of ions in a Paul trap via conventional and quantum machine learning algorithms

Trapped ions are among the leading platforms for quantum technologies, particularly in the field of quantum computing. Detecting states of trapped ions is essential for ensuring high-fidelity readouts of quantum states. In this work, we develop and benchmark a set of methods for ion quantum state detection using images obtained by a highly sensitive camera. By transforming the images from the camera and applying conventional and quantum machine learning methods, including convolution, support vector machine (classical and quantum), and quantum annealing, we demonstrate a possibility to detect the positions and quantum states of ytterbium ions in a Paul trap. Quantum state detection is performed with an electron shelving technique: depending on the quantum state of the ion its fluorescence under the influence of a 369.5 nm laser beam is either suppressed or not. We estimate fidelities for conventional and quantum detection techniques. In particular, conventional algorithms for detecting $^{171}$Yb$^{+}$, such as the support vector machine and photon statistics-based method,as well as our quantum annealing-based approach, have achieved perfect fidelity, which is beneficial compared to standard techniques. This result may pave the way for ultrahigh-fidelity detection of trapped ions via conventional and quantum machine learning techniques.

quant-ph↗