SearcharxivSearch

arXiv subjects

Eliot Brenner

Publications and source records attributed to Eliot Brenner.

11 recordsLinked to original sources

Fin-RATE: A Real-world Financial Analytics and Tracking Evaluation Benchmark for LLMs on SEC Filings

With the increasing deployment of Large Language Models (LLMs) in the finance domain, LLMs are increasingly expected to parse complex regulatory disclosures. However, existing benchmarks often focus on isolated details, failing to reflect the complexity of professional analysis that requires synthesizing information across multiple documents, reporting periods, and corporate entities. Furthermore, these benchmarks do not disentangle whether errors arise from retrieval failures, generation inaccuracies, domain-specific reasoning mistakes, or misinterpretation of the query or context, making it difficult to precisely diagnose performance bottlenecks. To bridge these gaps, we introduce Fin-RATE, a benchmark built on U.S. Securities and Exchange Commission (SEC) filings and mirroring financial analyst workflows through three pathways: detail-oriented reasoning within individual disclosures, cross-entity comparison under shared topics, and longitudinal tracking of the same firm across reporting periods. We benchmark 17 leading LLMs, spanning open-source, closed-source, and finance-specialized models, under both ground-truth context and retrieval-augmented settings. Results show substantial performance degradation, with accuracy dropping by 18.60% and 14.35% as tasks shift from single-document reasoning to longitudinal and cross-entity analysis. This degradation is associated with increased comparison hallucinations, temporal and entity mismatches, and is further reflected in declines in reasoning quality and factual consistency--limitations that existing benchmarks have yet to formally categorize or quantify.

cs.CE

Adaptation of Embedding Models to Financial Filings via LLM Distillation

Despite advances in generative large language models (LLMs), practical application of specialized conversational AI agents remains constrained by computation costs, latency requirements, and the need for precise domain-specific relevance measures. While existing embedding models address the first two constraints, they underperform on information retrieval in specialized domains like finance. This paper introduces a scalable pipeline that trains specialized models from an unlabeled corpus using a general purpose retrieval embedding model as foundation. Our method yields an average of 27.7% improvement in MRR$\texttt{@}$5, 44.6% improvement in mean DCG$\texttt{@}$5 across 14 financial filing types measured over 21,800 query-document pairs, and improved NDCG on 3 of 4 document classes in FinanceBench. We adapt retrieval embeddings (bi-encoder) for RAG, not LLM generators, using LLM-judged relevance to distill domain knowledge into a compact retriever. There are prior works which pair synthetically generated queries with real passages to directly fine-tune the retrieval model. Our pipeline differs from these by introducing interaction between student and teacher models that interleaves retrieval-based mining of hard positive/negative examples from the unlabeled corpus with iterative retraining of the student model's weights using these examples. Each retrieval iteration uses the refined student model to mine the corpus for progressively harder training examples for the subsequent training iteration. The methodology provides a cost-effective solution to bridging the gap between general-purpose models and specialized domains without requiring labor-intensive human annotation.

cs.CL

Long Document Summarization in a Low Resource Setting using Pretrained Language Models

Abstractive summarization is the task of compressing a long document into a coherent short document while retaining salient information. Modern abstractive summarization methods are based on deep neural networks which often require large training datasets. Since collecting summarization datasets is an expensive and time-consuming task, practical industrial settings are usually low-resource. In this paper, we study a challenging low-resource setting of summarizing long legal briefs with an average source document length of 4268 words and only 120 available (document, summary) pairs. To account for data scarcity, we used a modern pretrained abstractive summarizer BART (Lewis et al., 2020), which only achieves 17.9 ROUGE-L as it struggles with long documents. We thus attempt to compress these long documents by identifying salient sentences in the source which best ground the summary, using a novel algorithm based on GPT-2 (Radford et al., 2019) language model perplexity scores, that operates within the low resource regime. On feeding the compressed documents to BART, we observe a 6.0 ROUGE-L improvement. Our method also beats several competitive salience detection baselines. Furthermore, the identified salient sentences tend to agree with an independent human labeling by domain experts.

cs.CL

End-to-End Neural Ranking for eCommerce Product Search: an application of task models and textual embeddings

We consider the problem of retrieving and ranking items in an eCommerce catalog, often called SKUs, in order of relevance to a user-issued query. The input data for the ranking are the texts of the queries and textual fields of the SKUs indexed in the catalog. We review the ways in which this problem both resembles and differs from the problems of IR in the context of web search. The differences between the product-search problem and the IR problem of web search necessitate a different approach in terms of both models and datasets. We first review the recent state-of-the-art models for web search IR, distinguishing between two distinct types of model which we call the distributed type and the local-interaction type. The different types of relevance models developed for IR have complementary advantages and disadvantages when applied to eCommerce product search. Further, we explain why the conventional methods for dataset construction employed in the IR literature fail to produce data which suffices for training or evaluation of models for eCommerce product search. We explain how our own approach, applying task modeling techniques to the click-through logs of an eCommerce site, enables the construction of a large-scale dataset for training and robust benchmarking of relevance models. Our experiments consist of applying several of the models from the IR literature to our own dataset. Empirically, we have established that, when applied to our dataset, certain models of local-interaction type reduce ranking errors by one-third compared to the baseline tf-idf. Applied to our dataset, the distributed models fail to outperform the baseline. As a basis for a deployed system, the distributed models have several advantages, computationally, over the local-interaction models. This motivates an ongoing program of work, which we outline at the conclusion of the paper.

cs.IR

Incorporating Type II Error Probabilities from Independence Tests into Score-Based Learning of Bayesian Network Structure

We give a new consistent scoring function for structure learning of Bayesian networks. In contrast to traditional approaches to score-based structure learning, such as BDeu or MDL, the complexity penalty that we propose is data-dependent and is given by the probability that a conditional independence test correctly shows that an edge cannot exist. What really distinguishes this new scoring function from earlier work is that it has the property of becoming computationally easier to maximize as the amount of data increases. We prove a polynomial sample complexity result, showing that maximizing this score is guaranteed to correctly learn a structure with no false edges and a distribution close to the generating distribution, whenever there exists a Bayesian network which is a perfect map for the data generating distribution. Although the new score can be used with any search algorithm, in our related UAI 2013 paper [BS13], we have given empirical results showing that it is particularly effective when used together with a linear programming relaxation approach to Bayesian network structure learning. The present paper contains all details of the proofs of the finite-sample complexity results in [BS13] as well as detailed explanation of the computation of the certain error probabilities called beta-values, whose precomputation and tabulation is necessary for the implementation of the algorithm in [BS13].

cs.LG

SparsityBoost: A New Scoring Function for Learning Bayesian Network Structure

We give a new consistent scoring function for structure learning of Bayesian networks. In contrast to traditional approaches to scorebased structure learning, such as BDeu or MDL, the complexity penalty that we propose is data-dependent and is given by the probability that a conditional independence test correctly shows that an edge cannot exist. What really distinguishes this new scoring function from earlier work is that it has the property of becoming computationally easier to maximize as the amount of data increases. We prove a polynomial sample complexity result, showing that maximizing this score is guaranteed to correctly learn a structure with no false edges and a distribution close to the generating distribution, whenever there exists a Bayesian network which is a perfect map for the data generating distribution. Although the new score can be used with any search algorithm, we give empirical results showing that it is particularly effective when used together with a linear programming relaxation approach to Bayesian network structure learning.

cs.LG

Notes on Analytic Properties of Residual Eisenstein Series, I

We partially generalize the results of Kudla and Rallis on the poles of degenerate, Siegel-parabolic Eisenstein series to residual-data Eisenstein series. In particular, for $a,b$ integers greater than 1, we show that poles of the Eisenstein series induced from the Speh representation $Δ(τ,b)$ on the Levi $\mathrm{GL}_{ab}$ of $\mathrm{Sp}_{2ab}$ are located in the "segment" of half integers $X_{b}$ between a "right endpoint" and its negative, inclusive of endpoints. The right endpoint is $\pm b/2$, or $(b-1)/2$, depending on the analytic properties of the automorphic $L$-functions attached to $τ$. We study the automorphic forms $Φ_{i}^{(b)}$ obtained as residues at the points $s_i^{(b)}$ (defined precisely in the paper) by calculating their cuspidal exponents in certain cases. In the case of the "endpoint" $s_0^{(b)}$ and `first interior point' $s_1^{(b)}$ in the segment of singularity points, we are able to determine a set containing \textit{all possible} cuspidal exponents of $Φ_0^{(b)}$ and $Φ_1^{(b)}$ precisely for all $a$ and $b$. In these cases, we use the result of the calculation to deduce that the residual automorphic forms lie in $L^2(G(k)\backslash G(\mathbf{A}))$. In a more precise sense, our result establishes a relationship between, on the one hand, the actually occurring cuspidal exponents of $Φ_i^{(b)}$, residues at interior points which lie to the right of the origin, and, on the other hand, the "analytic properties" of the original residual-data Eisenstein series at the origin. This preprint is a longer version of the paper "Analytic Properties of Residual Eisenstein Series, I", with the details of some proofs added and some additional examples adduced in support of the main conjecture.

math.NT

Artin formalism for Selberg zeta functions of co-finite Kleinian groups

Let $Γ\backslash\mathbb H^3$ be a finite-volume quotient of the upper-half space, where $Γ\subset {\rm SL}(2,\mathbb C)$ is a discrete subgroup. To a finite dimensional unitary representation $χ$ of $Γ$ one associates the Selberg zeta function $Z(s;Γ;χ)$. In this paper we prove the Artin formalism for the Selberg zeta function. Namely, if $\tildeΓ$ is a finite index group extension of $Γ$ in ${\rm SL}(2,\mathbb C)$, and $π={\rm Ind}_Γ^{\tildeΓ}χ$ is the induced representation, then $Z(s;Γ;χ)=Z(s;\tildeΓ;π)$. In the second part of the paper we prove by a direct method the analogous identity for the scattering function, namely $ϕ(s;Γ;χ)=ϕ(s;\tildeΓ;π)$, for an appropriate normalization of the Eisenstein series.

math.NT

A fundamental domain of Ford type for $SO(3,Z[i])\backslash SO(3,C)/SO(3)$, and for $SO(2,1)_Z\backslash SO(2,1)/SO(2)$

Let $G=SO(3,C)$, $Γ=SO(3,Z[i])$, $K=SO(3)$, and let $X$ be the locally symmetric space $Γ\backslash G/K$. In this paper, we write down explicit equations defining a fundamental domain for the action of $Γ$ on $G/K$. The fundamental domain is well-adapted for studying the theory of $Γ$-invariant functions on $G/K$. We write down equations defining a fundamental domain for the subgroup $Γ_Z=SO(2,1)_Z$ of $Γ$ acting on the symmetric space $G_{R}/K_R$, where $G_R$ is the split real form SO(2,1) of $G$ and $K_R$ is its maximal compact subgroup SO(2). We formulate a simple geometric relation between the fundamental domain of $Γ$ and $Γ_Z$ so described. These fundamental domains are geared towards the detailed study of the spectral theory of $X$ and the embedded subspace $X_R=Γ_Z\backslash G_R/K_R$.}

math.NT

Stability of the local gamma factor in the unitary case

Rallis and Soudry have proven the stability under twists by highly ramified characters of the local gamma factor arising from the doubling method, in the case of a symplectic group or orthogonal group G over a local non-archimedean field F of characteristic zero, and a representation of G, which is not necessarily generic. This paper extends their arguments to show the stability in the case when G is a unitary group over a quadratic extension E of F, thereby completing the proof of the stability for classical groups. This stability property is important in Cogdell, Piatetski-Shapiro, and Shahidi's use of the converse theorem to prove the existence of a weak lift from automorphic, cuspidal, generic representations of G(A) to automorphic representations of GL(n,A) for appropriate n, to which references are given in the paper of Rallis and Soudry.

math.NT

A fundamental domain of Ford type for some subgroups of the orthogonal group

We initiate a study of the spectral theory of the locally symmetric space $X=Γ\backslash G/K$, where $G=SO(3,Complex)$, $Γ=SO(3,Z[i])$, $K=SO{3}$. We write down explicit equations defining a fundamental domain for the action of $Γ$ on $G/K$. The fundamental domain is well-adapted for studying the theory of $Γ$-invariant functions on $G/K$. We write down equations defining a fundamental domain for the subgroup $Γ_Z=\SO(2,1)_Z$ of $Γ$ acting on the symmetric space $G_R/K_R$, where $G_R$ is the split real form $\SO(2,1)$ of $G$ and $K_R$ is its maximal compact subgroup $\SO(2)$. We formulate a simple geometric relation between the fundamental domains of $Γ$ and $Γ_Z$ so described. We then use the previous results compute the covolumes of of the lattices $Γ$ and $Γ_Z$ in $G$ and $G_R$.

math.NT