SearcharxivSearch

arXiv subjects

Alexander Sergeev

Publications and source records attributed to Alexander Sergeev.

At least 19 recordsLinked to original sources

Optimizing Multimodal Language Models through Attention-based Interpretability

Modern large language models become multimodal, analyzing various data formats like text and images. While fine-tuning is effective for adapting these multimodal language models (MLMs) to downstream tasks, full fine-tuning is computationally expensive. Parameter-Efficient Fine-Tuning (PEFT) methods address this by training only a small portion of model weights. However, MLMs are difficult to interpret, making it challenging to identify which components are most effective for training to balance efficiency and performance. We propose an attention-based interpretability method for MLMs by analyzing attention scores relative to image tokens. The core idea is to identify attention heads that focus on image key objects. We utilize this information to select optimal model components for PEFT in multimodal models. Our contributions include a method for identifying attention heads associated with image key objects, its application to PEFT for image captioning, and the creation of a new dataset containing images, key object masks, and their textual descriptions. We conducted experiments on MLMs with 2-3 billion parameters to validate the method's effectiveness. By calculating Head Impact (HI) scores we quantify an attention head's focus on key objects, indicating its significance in image understanding. Our fine-tuning experiments demonstrate that adapting layers with the highest HI scores leads to the most significant shifts in metrics compared to pre-trained, randomly selected, or lowest-HI-score layers. This indicates that fine-tuning a small percentage (around 0.01%) of parameters in these crucial layers can substantially influence image understanding capabilities.

cs.CL

Do LLMs Understand Why We Write Diaries? A Method for Purpose Extraction and Clustering

Diary analysis presents challenges, particularly in extracting meaningful information from large corpora, where traditional methods often fail to deliver satisfactory results. This study introduces a novel method based on Large Language Models (LLMs) to identify and cluster the various purposes of diary writing. By "purposes," we refer to the intentions behind diary writing, such as documenting life events, self-reflection, or practicing language skills. Our approach is applied to Soviet-era diaries (1922-1929) from the Prozhito digital archive, a rich collection of personal narratives. We evaluate different proprietary and open-source LLMs, finding that GPT-4o and o1-mini achieve the best performance, while a template-based baseline is significantly less effective. Additionally, we analyze the retrieved purposes based on gender, age of the authors, and the year of writing. Furthermore, we examine the types of errors made by the models, providing a deeper understanding of their limitations and potential areas for improvement in future research.

cs.CL

Talking to Data: Designing Smart Assistants for Humanities Databases

Access to humanities research databases is often hindered by the limitations of traditional interaction formats, particularly in the methods of searching and response generation. This study introduces an LLM-based smart assistant designed to facilitate natural language communication with digital humanities data. The assistant, developed in a chatbot format, leverages the RAG approach and integrates state-of-the-art technologies such as hybrid search, automatic query generation, text-to-SQL filtering, semantic database search, and hyperlink insertion. To evaluate the effectiveness of the system, experiments were conducted to assess the response quality of various language models. The testing was based on the Prozhito digital archive, which contains diary entries from predominantly Russian-speaking individuals who lived in the 20th century. The chatbot is tailored to support anthropology and history researchers, as well as non-specialist users with an interest in the field, without requiring prior technical training. By enabling researchers to query complex databases with natural language, this tool aims to enhance accessibility and efficiency in humanities research. The study highlights the potential of Large Language Models to transform the way researchers and the public interact with digital archives, making them more intuitive and inclusive. Additional materials are presented in GitHub repository: https://github.com/alekosus/talking-to-data-intersys2025.

cs.CL

Mg$_2$Si is the new black: introducing a black silicide with $>$95% average absorption at 200-1800 nm wavelengths

Textured silicon surface structures, in particular black silicon (b-Si), open up possibilities for Si-based solar cells and photodetectors to be extremely thin and highly sensitive owing to perfect light-trapping and anti-reflection properties. However, near-infrared (NIR) performance of bare b-Si is limited by Si band gap of 1.12 eV or 1100 nm. This work reports a simple method to increase NIR absorption of b-Si by $in$ $vacuo$ silicidation with magnesium. Obtained Mg$_2$Si/b-Si heterostructure has a complex geometry where b-Si nanocones are covered by Mg$_2$Si shells and crowned with flake-like Mg$_2$Si hexagons. Mg$_2$Si formation atop b-Si resulted in 5-fold lower reflectivity and optical absorption to be no lower than 88\% over 200-1800 nm spectral range. More importantly, Mg$_2$Si/b-Si heterostructure is more adjusted to match AM-1.5 solar spectrum with theoretically higher photogenerated current density. The maximal advantage is demonstrated in the NIR region compared to bare b-Si in full accordance with one's expectations about NIR sensitive narrow band gap ($\sim$0.75 eV) semiconductor with high absorption coefficient, which is Mg$_2$Si. Results of optical simulation confirmed the superiority of Mg$_2$Si/b-Si NIR performance. Therefore, this new wide-band optical absorber called black silicide proved rather competitive alongside state-of-the-art approaches to extend b-Si spectral blackness.

physics.optics

Densifying Assumed-sparse Tensors: Improving Memory Efficiency and MPI Collective Performance during Tensor Accumulation for Parallelized Training of Neural Machine Translation Models

Neural machine translation - using neural networks to translate human language - is an area of active research exploring new neuron types and network topologies with the goal of dramatically improving machine translation performance. Current state-of-the-art approaches, such as the multi-head attention-based transformer, require very large translation corpuses and many epochs to produce models of reasonable quality. Recent attempts to parallelize the official TensorFlow "Transformer" model across multiple nodes have hit roadblocks due to excessive memory use and resulting out of memory errors when performing MPI collectives. This paper describes modifications made to the Horovod MPI-based distributed training framework to reduce memory usage for transformer models by converting assumed-sparse tensors to dense tensors, and subsequently replacing sparse gradient gather with dense gradient reduction. The result is a dramatic increase in scale-out capability, with CPU-only scaling tests achieving 91% weak scaling efficiency up to 1200 MPI processes (300 nodes), and up to 65% strong scaling efficiency up to 400 MPI processes (200 nodes) using the Stampede2 supercomputer.

cs.LG

Horovod: fast and easy distributed deep learning in TensorFlow

Training modern deep learning models requires large amounts of computation, often provided by GPUs. Scaling computation from one GPU to many can enable much faster training and research progress but entails two complications. First, the training library must support inter-GPU communication. Depending on the particular methods employed, this communication may entail anywhere from negligible to significant overhead. Second, the user must modify his or her training code to take advantage of inter-GPU communication. Depending on the training library's API, the modification required may be either significant or minimal. Existing methods for enabling multi-GPU training under the TensorFlow library entail non-negligible communication overhead and require users to heavily modify their model-building code, leading many researchers to avoid the whole mess and stick with slower single-GPU training. In this paper we introduce Horovod, an open source library that improves on both obstructions to scaling: it employs efficient inter-GPU communication via ring reduction and requires only a few lines of modification to user code, enabling faster, easier distributed training in TensorFlow. Horovod is available under the Apache 2.0 license at https://github.com/uber/horovod

cs.LG

Orthogonal polynomials of discrete variable and Lie algebras of complex size matrices

We give a uniform interpretation of the classical continuous Chebyshev's and Hahn's orthogonal polynomials of discrete variable in terms of Feigin's Lie algebra gl(N), where N is any complex number. One can similarly interpret Chebyshev's and Hahn's q-polynomials and introduce orthogonal polynomials corresponding to Lie superlagebras. We also describe the real forms of gl(N), quasi-finite modules over gl(N), and conditions for unitarity of the quasi-finite modules. Analogs of tensors over gl(N) are also introduced.

math.RT

Centralizer construction of the Yangian of the queer Lie superalgebra

Consider the complex matrix Lie superalgebra $gl_{N|N}$ with the standard generators $E_{ij}$ where $i,j=-N,...,-1,1,...,N$. Define an involutive automorphism $η$ of $gl_{N|N}$ by $η(E_{ij})=E_{-i,-j}$. The queer Lie superalgebra $q_N$ is the fixed point subalgebra in $gl_{N|N}$ relative to the automorphism $η$. Consider the twisted polynomial current Lie superalgebra $g=\{X(t)\in\gl_{N|N}[t]:η(X(t))=X(-t)\}$. The enveloping algebra $U(g)$ of the Lie superalgebra $g$ has a deformation, called the Yangian of $q_N$. For each $M=1,2,...$ denote by $A_N^M$ the centralizer of $q_M\subset q_{N+M}$ in the superalgebra $U(q_{N+M})$. We describe the projective limit of the sequence of centralizer algebras $A_N^1,A_N^2,...$ in terms of the Yangian of $q_N$.

math.RT

Casimir operators for Lie superalgebras

Casimir operators -- the generators of the center of the enveloping algebra -- are described for simple or close to them ``classical'' finite dimensional Lie superalgebras with nondegenerate symmetric even bilinear form in Sergeev A., The invariant polynomials on simple Lie superalgebras. Represent. Theory 3 (1999), 250--280; math-RT/9810111 and for the ``queer'' series in Sergeev A., The centre of enveloping algebra for Lie superalgebra Q(n, C). Lett. Math. Phys. 7, no. 3, 1983, 177--179. Here we consider the remaining cases, and state conjectures proved for small values of parameter. Under deformation (quantization) the Poisson Lie superalgebra po(0|2n) on purely odd superspace turns into gl(2^{n-1}|2^{n-1}) and, conjecturally, the lowest terms of the Taylor series expansion with respect to the deformation parameter (Planck's constant) of the Casimir operators for gl(2^{n-1}|2^{n-1}) are the Casimir operators for po(0|2n). Similarly, quantization sends po(0|2n-1) into q(2^{n-1}) and the above procedure makes Casimir operators for q(2^{n-1}) into same for po(0|2n-1). Casimir operators for the Lie superalgebra vect(0|m) of vector fields on purely odd superspace are only constants for m>2. Conjecturally, same is true for the Lie superalgebra svect(0|m) of divergence free vector fields, and its deform, for m>3. Invariant polynomials on po(0|2n-1) are also described. They do not correspond to Casimir operators.

math.RT

Enveloping algebra U(gl(3)) and orthogonal polynomials in several discrete indeterminates

Let A be an associative complex algebra and L an invariant linear functional on it (trace). Let i be an involutive antiautomorphism of A such that L(i(a))=L(a) for any a in A. Then A admits a symmetric invariant bilinear form (a, b)=L(a i(b)). For A=U(sl(2))/m, where m is any maximal ideal of U(sl(2)), Leites and I have constructed orthogonal basis whose elements turned out to be, essentially, Chebyshev and Hahn polynomials in one discrete variable. Here I take A=U(gl(3))/m for the maximal ideals m which annihilate irreducible highest weight gl(3)-modules of particular form (generalizations of symmetric powers of the identity representation). In this way we obtain multivariable analogs of Hahn polynomials.

math.RT

Projective Schur functions as a bispherical functions on certain homogeneous superspaces

I show that the projective Schur functions may be interpreted as bispherical functions of either the triple (q(n),gl(n,n),q(n)), where q(n) is the "odd" (queer) analog of the general liner Lie algebra, or the triple (p(n),gl(n,n),p(n)), where p(n) is the periplectic Lie superalgebra which preserves the nondegenerate odd, bilinear form (either symmetric or skew symmetric). Making use of this interpretation I characterize projective Schur functions as common eigenfunctions of an algebra differential operators.

math.RT

Enveloping superalgebra U(osp(1|2)) and orthogonal polynomials in discrete indeterminate

Let $A$ be an associative simple (central) superalgebra over ${\mathbb C}$ and $L$ an invariant linear functional on it (trace). Let $a\mapsto a^t$ be an antiautomorphism of $A$ such that $(a^t)^ t=(-1)^{p(a)}a$, where $p(a)$ is the parity of $a$, and let $L(a^t)=L(a)$. Then $A$ admits a nondegenerate supersymmetric invariant bilinear form $\langle a, b\rangle=L(ab^t)$. For $A=U({\mathfrak{sl}}(2))/{\mathfrak{m}}$, where ${\mathfrak{m}}$ is any maximal ideal of $U({\mathfrak{sl}}(2))$, Leites and I have constructed orthogonal basis in $A$ whose elements turned out to be, essentially, Chebyshev (Hahn) polynomials in one discrete variable. Here I take $A=U({\mathfrak{osp}}(1|2))/{\mathfrak{m}}$ for any maximal ideal ${\mathfrak{m}}$ and apply a similar procedure. As a result we obtain either Hahn polynomials over ${\mathbb C}[τ]$, where $τ^2\in{\mathbb C}$, or a particular case of Meixner polynomials, or --- when $A=\mbox{Mat}(n+1|n)$ --- dual Hahn polynomials of even degree, or their (hopefully, new) analogs of odd degree. Observe that the nondegenerate bilinear forms we consider for orthogonality are, as a rule, not sign definite.

math.RT

Superanalogs of the Calogero operators and Jack polynomials

A depending on a complex parameter $k$ superanalog ${\mathcal S}{\mathcal L}$ of Calogero operator is constructed; it is related with the root system of the Lie superalgebra ${\mathfrak{gl}}(n|m)$. For $m=0$ we obtain the usual Calogero operator; for $m=1$ we obtain, up to a change of indeterminates and parameter $k$ the operator constructed by Veselov, Chalykh and Feigin [2,3]. For $k=1, \frac12$ the operator ${\mathcal S}{\mathcal L}$ is the radial part of the 2nd order Laplace operator for the symmetric superspaces corresponding to pairs $(GL(V)\times GL(V), GL(V))$ and $(GL(V), OSp(V))$, respectively. We will show that for the generic $m$ and $n$ the superanalogs of the Jack polynomials constructed by Kerov, Okunkov and Olshanskii [5] are eigenfunctions of ${\mathcal S}{\mathcal L}$; for $k=1, \frac12$ they coinside with the spherical functions corresponding to the above mentioned symmetric superspaces. We also study the inner product induced by Berezin's integral on these superspaces.

math.RT

An analog of the classical invariant theory for Lie superlagebras. II

Let V be a finite dimensional complex superspace and G a simple (or a ``close'' to simple) Lie superalgebra of matrix type, i.e., a Lie subsuperalgebra in GL(V). Under the classical invariant theory for G we mean the description of G-invariant elements of the tensor algebra of V. In math.RT/9810113 the invariants are described up to polarization operators. Here we give a complete description of the generators in the algebra of invariants and describe the relations between the invariants of the scalar product type.

math.RT

The Howe duality and the Projective Representations of Symmetric Groups

The symmetric group S_n possesses a nontrivial central extension, whose irreducible representations, different from the irreducible representations of S_n itself, coincide with the irreducible representations of a certain algebra A_n. Recently M.~Nazarov realized irreducible representations of A_n and Young symmetrizers by means of the Howe duality between the Lie superalgebra q(n) and the Hecke algebra H_n, the semidirect product of S_n with the Clifford algebra C_n on n indeterminates. Here I construct one more analog of Young symmetrizers in H_n as well as the analogs of Specht modules for A_n and H_n.

math.RT

Irreducible representations of solvable Lie superalgebras

The description of irreducible finite dimensional representations of finite dimensional solvable Lie superalgebras over complex numbers given by V.~Kac is refined. In reality these representations are not just induced from a polarization but twisted, as infinite dimensional representations of solvable Lie algebras. Various cases of irreducibility (general and of type Q) are classified.

math.RT

Orthogonal polynomials and Lie superalgebras

For the orthogonal Lie algebra O(2n+1), in addition to the conventional set of orthogonal polynomials, another set is produced with the help of the Lie superalgebra OSP(1|2n). Difficulties related with expression of Dyson's constant for the Lie superalgebras are discussed.

math.RT