SearcharxivSearch

arXiv subjects

Sebastian Frank

Publications and source records attributed to Sebastian Frank.

6 recordsLinked to original sources

What Makes a Peer? Valuation-Anchored Similarity in Private Markets

As more investors contemplate private markets and contend with limited transparency, sparse disclosures, and infrequent transactions, identifying economically meaningful peer companies for comparison is a fundamental challenge for valuation, due diligence, portfolio construction, and risk management. We propose an ensemble tree-based supervised similarity learning framework that defines company similarity through the lens of market valuation rather than static feature matching or semantic descriptions. Specifically, we train a CatBoost gradient-boosted decision tree model on observed private company valuations and derive a valuation-aware similarity metric from importance-weighted leaf-node co-occurrences across the ensemble. The similarity metric captures shared valuation drivers while accommodating nonlinear relationships, mixed data types, and pervasive missing data common in private markets. Using a global private-market universe of approximately 270,000 companies, including more than 53,000 firms with observed or derivable post-money valuations spanning multiple industries, geographies, and deal stages, we demonstrate that the proposed similarity framework improves upon traditional distance-based and text-embedding-based approaches in downstream k-nearest-neighbor valuation tasks in the evaluated industry groups, while retaining case-based explainability.

q-fin.ST

How to Choose a Threshold for an Evaluation Metric for Large Language Models

To ensure and monitor large language models (LLMs) reliably, various evaluation metrics have been proposed in the literature. However, there is little research on prescribing a methodology to identify a robust threshold on these metrics even though there are many serious implications of an incorrect choice of the thresholds during deployment of the LLMs. Translating the traditional model risk management (MRM) guidelines within regulated industries such as the financial industry, we propose a step-by-step recipe for picking a threshold for a given LLM evaluation metric. We emphasize that such a methodology should start with identifying the risks of the LLM application under consideration and risk tolerance of the stakeholders. We then propose concrete and statistically rigorous procedures to determine a threshold for the given LLM evaluation metric using available ground-truth data. As a concrete example to demonstrate the proposed methodology at work, we employ it on the Faithfulness metric, as implemented in various publicly available libraries, using the publicly available HaluBench dataset. We also lay a foundation for creating systematic approaches to select thresholds, not only for LLMs but for any GenAI applications.

stat.ML

Can an unsupervised clustering algorithm reproduce a categorization system?

Peer analysis is a critical component of investment management, often relying on expert-provided categorization systems. These systems' consistency is questioned when they do not align with cohorts from unsupervised clustering algorithms optimized for various metrics. We investigate whether unsupervised clustering can reproduce ground truth classes in a labeled dataset, showing that success depends on feature selection and the chosen distance metric. Using toy datasets and fund categorization as real-world examples we demonstrate that accurately reproducing ground truth classes is challenging. We also highlight the limitations of standard clustering evaluation metrics in identifying the optimal number of clusters relative to the ground truth classes. We then show that if appropriate features are available in the dataset, and a proper distance metric is known (e.g., using a supervised Random Forest-based distance metric learning method), then an unsupervised clustering can indeed reproduce the ground truth classes as distinct clusters.

stat.ML

Introducing Interactions in Multi-Objective Optimization of Software Architectures

Software architecture optimization aims to enhance non-functional attributes like performance and reliability while meeting functional requirements. Multi-objective optimization employs metaheuristic search techniques, such as genetic algorithms, to explore feasible architectural changes and propose alternatives to designers. However, this resource-intensive process may not always align with practical constraints. This study investigates the impact of designer interactions on multi-objective software architecture optimization. Designers can intervene at intermediate points in the fully automated optimization process, making choices that guide exploration towards more desirable solutions. Through several controlled experiments as well as an initial user study (14 subjects), we compare this interactive approach with a fully automated optimization process, which serves as a baseline. The findings demonstrate that designer interactions lead to a more focused solution space, resulting in improved architectural quality. By directing the search towards regions of interest, the interaction uncovers architectures that remain unexplored in the fully automated process. In the user study, participants found that our interactive approach provides a better trade-off between sufficient exploration of the solution space and the required computation time.

cs.SE

Orbital signatures of Fano-Kondo line shapes in STM adatom spectroscopy

We investigate the orbital origin of the Fano-Kondo line shapes measured in STM spectroscopy of magnetic adatoms on metal substrates. To this end we calculate the low-bias tunnel spectra of a Co adatom on the (001) and (111) Cu surfaces with our density functional theory-based ab initio transport scheme augmented by local correlations. In order to associate different $d$-orbitals with different Fano line shapes we only correlate individual $3d$-orbitals instead of the full Co $3d$-shell. We find that Kondo peaks arising in different $d$-levels indeed give rise to different Fano features in the conductance spectra. Hence the shape of measured Fano features allows to draw some conclusions about the orbital responsible for the Kondo resonance, although the actual shape is also influenced by temperature, effective interaction and charge fluctuations. Comparison with a simplified model shows that line shapes are mostly the result of interference between tunneling paths through the correlated $d$-orbital and the $sp$-type orbitals on the Co atom. Very importantly, the amplitudes of the Fano features vary strongly among orbitals, with the $3z^2$-orbital featuring by far the largest amplitude due to its strong direct coupling to the $s$-type conduction electrons.

cond-mat.str-el

Spectroscopy of high-energy states of lanthanide ions

We discuss recent progress and future prospects for the analysis of the 4f${N-1}$5d excited states of lanthanide ions in host materials. Ab-initio calculations for Ce$^{3+}$ in LiYF$_4$ are used to estimate crystal-field and spin-orbit parameters for the 4f$^1$ and 5d$^1$ configurations. We discuss the possibility of using excited-state absorption to probe the electronic and geometric structure of the 4f$^{N-1}$5d excited states in more detail and we illustrate these ideas with calculations for Yb$^{2+}$ ions in SrCl$_2$.

cond-mat.mtrl-sci