SearcharxivSearch

arXiv subjects

Peter Sarlin

Publications and source records attributed to Peter Sarlin.

17 recordsLinked to original sources

Poro 34B and the Blessing of Multilinguality

The pretraining of state-of-the-art large language models now requires trillions of words of text, which is orders of magnitude more than available for the vast majority of languages. While including text in more than one language is an obvious way to acquire more pretraining data, multilinguality is often seen as a curse, and most model training efforts continue to focus near-exclusively on individual large languages. We believe that multilinguality can be a blessing: when the lack of training data is a constraint for effectively training larger models for a target language, augmenting the dataset with other languages can offer a way to improve over the capabilities of monolingual models for that language. In this study, we introduce Poro 34B, a 34 billion parameter model trained for 1 trillion tokens of Finnish, English, and programming languages, and demonstrate that a multilingual training approach can produce a model that substantially advances over the capabilities of existing models for Finnish and excels in translation, while also achieving competitive performance in its class for English and programming languages. We release the model parameters, scripts, and data under open licenses at https://huggingface.co/LumiOpen/Poro-34B.

cs.CL

Deep learning bank distress from news and numerical financial data

In this paper we focus our attention on the exploitation of the information contained in financial news to enhance the performance of a classifier of bank distress. Such information should be analyzed and inserted into the predictive model in the most efficient way and this task deals with all the issues related to text analysis and specifically analysis of news media. Among the different models proposed for such purpose, we investigate one of the possible deep learning approaches, based on a doc2vec representation of the textual data, a kind of neural network able to map the sequential and symbolic text input onto a reduced latent semantic space. Afterwards, a second supervised neural network is trained combining news data with standard financial figures to classify banks whether in distressed or tranquil states, based on a small set of known distress events. Then the final aim is not only the improvement of the predictive performance of the classifier but also to assess the importance of news data in the classification process. Does news data really bring more useful information not contained in standard financial variables? Our results seem to confirm such hypothesis.

stat.ML

News-sentiment networks as a risk indicator

To understand the relationship between news sentiment and company stock price movements, and to better understand connectivity among companies, we define an algorithm for measuring sentiment-based network risk. The algorithm ranks companies in networks of co-occurrences, and measures sentiment-based risk, by calculating both individual risks and aggregated network risks. We extract relative sentiment for companies to get a measure of individual company risk, and input it into our risk model together with co-occurrences of companies extracted from news on a quarterly basis. We can show that the highest quarterly risk value outputted by our risk model, is correlated to a higher chance of stock price decline, up to 70 days after a risk measurement. Our results show that the highest difference in the probability of stock price decline, compared to the benchmark containing all risk values for the same period, is during the interval from 21 to 30 days after a quarterly measurement. The highest average probability of company stock price decline, is found at a delay of 28 days, after a company has reached its maximum risk value. The highest probability differences for a daily decline were calculated to be 13 percentage points.

q-fin.RM

Bank distress in the news: Describing events through deep learning

While many models are purposed for detecting the occurrence of significant events in financial systems, the task of providing qualitative detail on the developments is not usually as well automated. We present a deep learning approach for detecting relevant discussion in text and extracting natural language descriptions of events. Supervised by only a small set of event information, comprising entity names and dates, the model is leveraged by unsupervised learning of semantic vector representations on extensive text data. We demonstrate applicability to the study of financial risk based on news (6.6M articles), particularly bank distress and government interventions (243 events), where indices can signal the level of bank-stress-related reporting at the entity level, or aggregated at national or European level, while being coupled with explanations. Thus, we exemplify how text, as timely, widely available and descriptive data, can serve as a useful complementary source of information for financial and systemic risk analytics.

cs.CL

RiskRank: Measuring interconnected risk

This paper proposes RiskRank as a joint measure of cyclical and cross-sectional systemic risk. RiskRank is a general-purpose aggregation operator that concurrently accounts for risk levels for individual entities and their interconnectedness. The measure relies on the decomposition of systemic risk into sub-components that are in turn assessed using a set of risk measures and their relationships. For this purpose, motivated by the development of the Choquet integral, we employ the RiskRank function to aggregate risk measures, allowing for the integration of the interrelation of different factors in the aggregation process. The use of RiskRank is illustrated through a real-world case in a European setting, in which we show that it performs well in out-of-sample analysis. In the example, we provide an estimation of systemic risk from country-level risk and cross-border linkages.

q-fin.RM

Bank Networks from Text: Interrelations, Centrality and Determinants

In the wake of the still ongoing global financial crisis, bank interdependencies have come into focus in trying to assess linkages among banks and systemic risk. To date, such analysis has largely been based on numerical data. By contrast, this study attempts to gain further insight into bank interconnections by tapping into financial discourse. We present a text-to-network process, which has its basis in co-occurrences of bank names and can be analyzed quantitatively and visualized. To quantify bank importance, we propose an information centrality measure to rank and assess trends of bank centrality in discussion. For qualitative assessment of bank networks, we put forward a visual, interactive interface for better illustrating network structures. We illustrate the text-based approach on European Large and Complex Banking Groups (LCBGs) during the ongoing financial crisis by quantifying bank interrelations and centrality from discussion in 3M news articles, spanning 2007Q1 to 2014Q3.

q-fin.CP

Detect & Describe: Deep learning of bank stress in the news

News is a pertinent source of information on financial risks and stress factors, which nevertheless is challenging to harness due to the sparse and unstructured nature of natural text. We propose an approach based on distributional semantics and deep learning with neural networks to model and link text to a scarce set of bank distress events. Through unsupervised training, we learn semantic vector representations of news articles as predictors of distress events. The predictive model that we learn can signal coinciding stress with an aggregated index at bank or European level, while crucially allowing for automatic extraction of text descriptions of the events, based on passages with high stress levels. The method offers insight that models based on other types of data cannot provide, while offering a general means for interpreting this type of semantic-predictive model. We model bank distress with data on 243 events and 6.6M news articles for 101 large European banks.

q-fin.CP

Toward robust early-warning models: A horse race, ensembles and model uncertainty

This paper presents first steps toward robust models for crisis prediction. We conduct a horse race of conventional statistical methods and more recent machine learning methods as early-warning models. As individual models are in the literature most often built in isolation of other methods, the exercise is of high relevance for assessing the relative performance of a wide variety of methods. Further, we test various ensemble approaches to aggregating the information products of the built models, providing a more robust basis for measuring country-level vulnerabilities. Finally, we provide approaches to estimating model uncertainty in early-warning exercises, particularly model performance uncertainty and model output uncertainty. The approaches put forward in this paper are shown with Europe as a playground. Generally, our results show that the conventional statistical approaches are outperformed by more advanced machine learning methods, such as k-nearest neighbors and neural networks, and particularly by model aggregation approaches through ensemble learning.

q-fin.ST

Aggregation operators for the measurement of systemic risk

The policy objective of safeguarding financial stability has stimulated a wave of research on systemic risk analytics, yet it still faces challenges in measurability. This paper models systemic risk by tapping into expert knowledge of financial supervisors. We decompose systemic risk into a number of interconnected segments, for which the level of vulnerability is measured. The system is modeled in the form of a Fuzzy Cognitive Map (FCM), in which nodes represent vulnerability in segments and links their interconnectedness. A main problem tackled in this paper is the aggregation of values in different interrelated nodes of the network to obtain an estimate systemic risk. To this end, the Choquet integral is employed for aggregating expert evaluations of measures, as it allows for the integration of interrelations among factors in the aggregation process. The approach is illustrated through two applications in a European setting. First, we provide an estimation of systemic risk with a of pan-European set-up. Second, we estimate country-level risks, allowing for a more granular decomposition. This sets a starting point for the use of the rich, oftentimes tacit, knowledge in policy organizations.

q-fin.GN

Interactive Visual Exploration of Topic Models using Graphs

Probabilistic topic modeling is a popular and powerful family of tools for uncovering thematic structure in large sets of unstructured text documents. While much attention has been directed towards the modeling algorithms and their various extensions, comparatively few studies have concerned how to present or visualize topic models in meaningful ways. In this paper, we present a novel design that uses graphs to visually communicate topic structure and meaning. By connecting topic nodes via descriptive keyterms, the graph representation reveals topic similarities, topic meaning and shared, ambiguous keyterms. At the same time, the graph can be used for information retrieval purposes, to find documents by topic or topic subsets. To exemplify the utility of the design, we illustrate its use for organizing and exploring corpora of financial patents.

cs.IR

The process of macroprudential oversight in Europe

The 2007--2008 financial crisis has paved the way for the use of macroprudential policies in supervising the financial system as a whole. This paper views macroprudential oversight in Europe as a process, a sequence of activities with the ultimate aim of safeguarding financial stability. To conceptualize a process in this context, we introduce the notion of a public collaborative process (PCP). PCPs involve multiple organizations with a common objective, where a number of dispersed organizations cooperate under various unstructured forms and take a collaborative approach to reaching the final goal. We argue that PCPs can and should essentially be managed using the tools and practices common for business processes. To this end, we conduct an assessment of process readiness for macroprudential oversight in Europe. Based upon interviews with key European policymakers and supervisors, we provide an analysis model to assess the maturity of five process enablers for macroprudential oversight. With the results of our analysis, we give clear recommendations on the areas that need further attention when macroprudential oversight is being developed, in addition to providing a general purpose framework for monitoring the impact of improvement efforts.

q-fin.GN

Macroprudential oversight, risk communication and visualization

This paper discusses the role of risk communication in macroprudential oversight and of visualization in risk communication. Beyond the soar in data availability and precision, the transition from firm-centric to system-wide supervision imposes vast data needs. Moreover, except for internal communication as in any organization, broad and effective external communication of timely information related to systemic risks is a key mandate of macroprudential supervisors, further stressing the importance of simple representations of complex data. This paper focuses on the background and theory of information visualization and visual analytics, as well as techniques within these fields, as potential means for risk communication. We define the task of visualization in risk communication, discuss the structure of macroprudential data, and review visualization techniques applied to systemic risk. We conclude that two essential, yet rare, features for supporting the analysis of big data and communication of risks are analytical visualizations and interactive interfaces. For visualizing the so-called macroprudential data cube, we provide the VisRisk platform with three modules: plots, maps and networks. While VisRisk is herein illustrated with five web-based interactive visualizations of systemic risk indicators and models, the platform enables and is open to the visualization of any data from the macroprudential data cube.

q-fin.CP

Automated and Weighted Self-Organizing Time Maps

This paper proposes schemes for automated and weighted Self-Organizing Time Maps (SOTMs). The SOTM provides means for a visual approach to evolutionary clustering, which aims at producing a sequence of clustering solutions. This task we denote as visual dynamic clustering. The implication of an automated SOTM is not only a data-driven parametrization of the SOTM, but also the feature of adjusting the training to the characteristics of the data at each time step. The aim of the weighted SOTM is to improve learning from more trustworthy or important data with an instance-varying weight. The schemes for automated and weighted SOTMs are illustrated on two real-world datasets: (i) country-level risk indicators to measure the evolution of global imbalances, and (ii) credit applicant data to measure the evolution of firm-level credit risks.

cs.NE

From Bits to Atoms: 3D Printing in the Context of Supply Chain Strategies

A lot of attention in supply chain management has been devoted to understanding customer requirements. What are customer priorities in terms of price and service level, and how can companies go about fulfilling these requirements in an optimal way? New manufacturing technology in the form of 3D printing is about to change some of the underlying assumptions for different supply chain set-ups. This paper explores opportunities and barriers of 3D printing technology, specifically in a supply chain context. We are proposing a set of principles that can act to bridge existing research on different supply chain strategies and 3D printing. With these principles, researchers and practitioners alike can better understand the opportunities and limitations of 3D printing in a supply chain management context.

cs.CY

From Text to Bank Interrelation Maps

In the wake of the ongoing global financial crisis, interdependencies among banks have come into focus in trying to assess systemic risk. To date, such analysis has largely been based on numerical data. By contrast, this study attempts to gain further insight into bank interconnections by tapping into financial discussion. Co-mentions of bank names are turned into a network, which can be visualized and analyzed quantitatively, in order to illustrate characteristics of individual banks and the network as a whole. The approach allows for the study of temporal dynamics of the network, to highlight changing patterns of discussion that reflect real-world events, the current financial crisis in particular. For instance, it depicts how connections from distressed banks to other banks and supervisory authorities have emerged and faded over time, as well as how global shifts in network structure coincide with severe crisis episodes. The usage of textual data holds an additional advantage in the possibility of gaining a more qualitative understanding of an observed interrelation, through its context. We illustrate our approach using a case study on Finnish banks and financial institutions. The data set comprises 3.9M posts from online, financial and business-related discussion, during the years 2004 to 2012. Future research includes analyzing European news articles with a broader perspective, and a focus on improving semantic description of relations.

q-fin.RM

Cluster coloring of the Self-Organizing Map: An information visualization perspective

This paper takes an information visualization perspective to visual representations in the general SOM paradigm. This involves viewing SOM-based visualizations through the eyes of Bertin's and Tufte's theories on data graphics. The regular grid shape of the Self-Organizing Map (SOM), while being a virtue for linking visualizations to it, restricts representation of cluster structures. From the viewpoint of information visualization, this paper provides a general, yet simple, solution to projection-based coloring of the SOM that reveals structures. First, the proposed color space is easy to construct and customize to the purpose of use, while aiming at being perceptually correct and informative through two separable dimensions. Second, the coloring method is not dependent on any specific method of projection, but is rather modular to fit any objective function suitable for the task at hand. The cluster coloring is illustrated on two datasets: the iris data, and welfare and poverty indicators.

cs.LG

Self-Organizing Time Map: An Abstraction of Temporal Multivariate Patterns

This paper adopts and adapts Kohonen's standard Self-Organizing Map (SOM) for exploratory temporal structure analysis. The Self-Organizing Time Map (SOTM) implements SOM-type learning to one-dimensional arrays for individual time units, preserves the orientation with short-term memory and arranges the arrays in an ascending order of time. The two-dimensional representation of the SOTM attempts thus twofold topology preservation, where the horizontal direction preserves time topology and the vertical direction data topology. This enables discovering the occurrence and exploring the properties of temporal structural changes in data. For representing qualities and properties of SOTMs, we adapt measures and visualizations from the standard SOM paradigm, as well as introduce a measure of temporal structural changes. The functioning of the SOTM, and its visualizations and quality and property measures, are illustrated on artificial toy data. The usefulness of the SOTM in a real-world setting is shown on poverty, welfare and development indicators.

cs.LG