SearcharxivSearch

arXiv subjects

Stanislav S. Borysov

Publications and source records attributed to Stanislav S. Borysov.

15 recordsLinked to original sources

Estimating Causal Effects with the Neural Autoregressive Density Estimator

Estimation of causal effects is fundamental in situations were the underlying system will be subject to active interventions. Part of building a causal inference engine is defining how variables relate to each other, that is, defining the functional relationship between variables given conditional dependencies. In this paper, we deviate from the common assumption of linear relationships in causal models by making use of neural autoregressive density estimators and use them to estimate causal effects within the Pearl's do-calculus framework. Using synthetic data, we show that the approach can retrieve causal effects from non-linear systems without explicitly modeling the interactions between the variables.

stat.ME

Prediction of rare feature combinations in population synthesis: Application of deep generative modelling

In population synthesis applications, when considering populations with many attributes, a fundamental problem is the estimation of rare combinations of feature attributes. Unsurprisingly, it is notably more difficult to reliably representthe sparser regions of such multivariate distributions and in particular combinations of attributes which are absent from the original sample. In the literature this is commonly known as sampling zeros for which no systematic solution has been proposed so far. In this paper, two machine learning algorithms, from the family of deep generative models,are proposed for the problem of population synthesis and with particular attention to the problem of sampling zeros. Specifically, we introduce the Wasserstein Generative Adversarial Network (WGAN) and the Variational Autoencoder(VAE), and adapt these algorithms for a large-scale population synthesis application. The models are implemented on a Danish travel survey with a feature-space of more than 60 variables. The models are validated in a cross-validation scheme and a set of new metrics for the evaluation of the sampling-zero problem is proposed. Results show how these models are able to recover sampling zeros while keeping the estimation of truly impossible combinations, the structural zeros, at a comparatively low level. Particularly, for a low dimensional experiment, the VAE, the marginal sampler and the fully random sampler generate 5%, 21% and 26%, respectively, more structural zeros per sampling zero generated by the WGAN, while for a high dimensional case, these figures escalate to 44%, 2217% and 170440%, respectively. This research directly supports the development of agent-based systems and in particular cases where detailed socio-economic or geographical representations are required.

stat.ML

Band gap prediction for large organic crystal structures with machine learning

Machine-learning models are capable of capturing the structure-property relationship from a dataset of computationally demanding ab initio calculations. Over the past two years, the Organic Materials Database (OMDB) has hosted a growing number of calculated electronic properties of previously synthesized organic crystal structures. The complexity of the organic crystals contained within the OMDB, which have on average 82 atoms per unit cell, makes this database a challenging platform for machine learning applications. In this paper, the focus is on predicting the band gap which represents one of the basic properties of a crystalline materials. With this aim, a consistent dataset of 12 500 crystal structures and their corresponding DFT band gap are released, freely available for download at https://omdb.mathub.io/dataset. An ensemble of two state-of-the-art models reach a mean absolute error (MAE) of 0.388 eV, which corresponds to a percentage error of 13% for an average band gap of 3.05 eV. Finally, the trained models are employed to predict the band gap for 260 092 materials contained within the Crystallography Open Database (COD) and made available online so that the predictions can be obtained for any arbitrary crystal structure uploaded by a user.

cond-mat.mtrl-sci

Scalable Population Synthesis with Deep Generative Modeling

Population synthesis is concerned with the generation of synthetic yet realistic representations of populations. It is a fundamental problem in the modeling of transport where the synthetic populations of micro-agents represent a key input to most agent-based models. In this paper, a new methodological framework for how to 'grow' pools of micro-agents is presented. The model framework adopts a deep generative modeling approach from machine learning based on a Variational Autoencoder (VAE). Compared to the previous population synthesis approaches, including Iterative Proportional Fitting (IPF), Gibbs sampling and traditional generative models such as Bayesian Networks or Hidden Markov Models, the proposed method allows fitting the full joint distribution for high dimensions. The proposed methodology is compared with a conventional Gibbs sampler and a Bayesian Network by using a large-scale Danish trip diary. It is shown that, while these two methods outperform the VAE in the low-dimensional case, they both suffer from scalability issues when the number of modeled attributes increases. It is also shown that the Gibbs sampler essentially replicates the agents from the original sample when the required conditional distributions are estimated as frequency tables. In contrast, the VAE allows addressing the problem of sampling zeros by generating agents that are virtually different from those in the original data but have similar statistical properties. The presented approach can support agent-based modeling at all levels by enabling richer synthetic populations with smaller zones and more detailed individual characteristics.

stat.ML

Introducing Super Pseudo Panels: Application to Transport Preference Dynamics

We propose a new approach for constructing synthetic pseudo-panel data from cross-sectional data. The pseudo panel and the preferences it intends to describe is constructed at the individual level and is not affected by aggregation bias across cohorts. This is accomplished by creating a high-dimensional probabilistic model representation of the entire data set, which allows sampling from the probabilistic model in such a way that all of the intrinsic correlation properties of the original data are preserved. The key to this is the use of deep learning algorithms based on the Conditional Variational Autoencoder (CVAE) framework. From a modelling perspective, the concept of a model-based resampling creates a number of opportunities in that data can be organized and constructed to serve very specific needs of which the forming of heterogeneous pseudo panels represents one. The advantage, in that respect, is the ability to trade a serious aggregation bias (when aggregating into cohorts) for an unsystematic noise disturbance. Moreover, the approach makes it possible to explore high-dimensional sparse preference distributions and their linkage to individual specific characteristics, which is not possible if applying traditional pseudo-panel methods. We use the presented approach to reveal the dynamics of transport preferences for a fixed pseudo panel of individuals based on a large Danish cross-sectional data set covering the period from 2006 to 2016. The model is also utilized to classify individuals into 'slow' and 'fast' movers with respect to the speed at which their preferences change over time. It is found that the prototypical fast mover is a young woman who lives as a single in a large city whereas the typical slow mover is a middle-aged man with high income from a nuclear family who lives in a detached house outside a city.

stat.ML

A Bayesian Additive Model for Understanding Public Transport Usage in Special Events

Public special events, like sports games, concerts and festivals are well known to create disruptions in transportation systems, often catching the operators by surprise. Although these are usually planned well in advance, their impact is difficult to predict, even when organisers and transportation operators coordinate. The problem highly increases when several events happen concurrently. To solve these problems, costly processes, heavily reliant on manual search and personal experience, are usual practice in large cities like Singapore, London or Tokyo. This paper presents a Bayesian additive model with Gaussian process components that combines smart card records from public transport with context information about events that is continuously mined from the Web. We develop an efficient approximate inference algorithm using expectation propagation, which allows us to predict the total number of public transportation trips to the special event areas, thereby contributing to a more adaptive transportation system. Furthermore, for multiple concurrent event scenarios, the proposed algorithm is able to disaggregate gross trip counts into their most likely components related to specific events and routine behavior. Using real data from Singapore, we show that the presented model outperforms the best baseline model by up to 26% in R2 and also has explanatory power for its individual components.

stat.ML

Online Search Tool for Graphical Patterns in Electronic Band Structures

We present an online graphical pattern search tool for electronic band structure data contained within the Organic Materials Database (OMDB) available at https://omdb.diracmaterials.org/search/pattern. The tool is capable of finding user-specified graphical patterns in the collection of thousands of band structures from high-throughput ab initio calculations in the online regime. Using this tool, it only takes a few seconds to find an arbitrary graphical pattern within the ten electronic bands near the Fermi level for 26,739 organic crystals. The tool can be used to find realizations of functional materials characterized by a specific pattern in their electronic structure, for example, Dirac materials, characterized by a linear crossing of bands; topological insulators, characterized by a "Mexican hat" pattern or an effectively free electron gas, characterized by a parabolic dispersion. The source code of the developed tool is freely available at https://github.com/OrganicMaterialsDatabase/EBS-search and can be transferred to any other electronic band structure database. The approach allows for an automatic online analysis of a large collection of band structures where the amount of data makes its manual inspection impracticable.

cond-mat.mtrl-sci

Towards Novel Organic High-$T_\mathrm{c}$ Superconductors: Data Mining using Density of States Similarity Search

Identifying novel functional materials with desired key properties is an important part of bridging the gap between fundamental research and technological advancement. In this context, high-throughput calculations combined with data-mining techniques highly accelerated this process in different areas of research during the past years. The strength of a data-driven approach for materials prediction lies in narrowing down the search space of thousands of materials to a subset of prospective candidates. Recently, the open-access organic materials database OMDB was released providing electronic structure data for thousands of previously synthesized three-dimensional organic crystals. Based on the OMDB, we report about the implementation of a novel density of states similarity search tool which is capable of retrieving materials with similar density of states to a reference material. The tool is based on the approximate nearest neighbor algorithm as implemented in the ANNOY library and can be applied via the OMDB web interface. The approach presented here is wide-ranging and can be applied to various problems where the density of states is responsible for certain key properties of a material. As the first application, we report about materials exhibiting electronic structure similarities to the aromatic hydrocarbon p-terphenyl which was recently discussed as a potential organic high-temperature superconductor exhibiting a transition temperature in the order of 120~K under strong potassium doping. Although the mechanism driving the remarkable transition temperature remains under debate, we argue that the density of states, reflecting the electronic structure of a material, might serve as a crucial ingredient for the observed high-$T_\mathrm{c}$. To provide candidates which might exhibit comparable properties, we present 15 purely organic materials with similar features to p-terphenyl...

cond-mat.supr-con

Data Mining for Three-Dimensional Organic Dirac Materials: Focus on Space Group 19

We combined the group theory and data mining approach within the Organic Materials Database that leads to the prediction of stable Dirac-point nodes within the electronic band structure of 3-dimensional organic crystals. We find a particular space group $P2_12_12_1$ ($\#19$) that is conducive to the Dirac nodes formation. We prove that nodes are a consequence of the orthorhombic crystal structure. Within the electronic band structure, two different kinds of nodes can be distinguished: 8-fold degenerate Dirac nodes protected by the crystalline symmetry and 4-fold degenerate Dirac nodes protected by band topology. Mining the Organic Materials Database, we present band structure calculations and symmetry analysis for 6 previously synthesized organic materials. In all these materials, the Dirac nodes are well separated within the energy and located near the Fermi surface, which opens up a possibility for their direct experimental observation.

cond-mat.mtrl-sci

Three-dimensional organic Dirac-line materials due to nonsymmorphic symmetry: a data mining approach

A data mining study of electronic Kohn-Sham band structures was performed to identify Dirac materials within the Organic Materials Database (OMDB). Out of that, the 3-dimensional organic crystal 5,6-bis(trifluoromethyl)-2-methoxy-1$H$-1,3-diazepine was found to host different Dirac line-nodes within the band structure. From a group theoretical analysis, it is possible to distinguish between Dirac line-nodes occurring due to 2-fold degenerate energy levels protected by the monoclinic crystalline symmetry and 2-fold degenerate accidental crossings protected by the topology of the electronic band structure. The obtained results can be generalized to all materials having the space group $P2_1/c$ (No. 14, $C^5_{2h}$) by introducing three distinct topological classes.

cond-mat.mtrl-sci

U.S. stock market interaction network as learned by the Boltzmann Machine

We study historical dynamics of joint equilibrium distribution of stock returns in the U.S. stock market using the Boltzmann distribution model being parametrized by external fields and pairwise couplings. Within Boltzmann learning framework for statistical inference, we analyze historical behavior of the parameters inferred using exact and approximate learning algorithms. Since the model and inference methods require use of binary variables, effect of this mapping of continuous returns to the discrete domain is studied. The presented analysis shows that binarization preserves market correlation structure. Properties of distributions of external fields and couplings as well as industry sector clustering structure are studied for different historical dates and moving window sizes. We found that a heavy positive tail in the distribution of couplings is responsible for the sparse market clustering structure. We also show that discrepancies between the model parameters might be used as a precursor of financial instabilities.

q-fin.ST

Determining surface properties with bimodal and multimodal AFM

Conventional dynamic atomic force microscopy (AFM) can be extended to bimodal and multimodal AFM in which the cantilever is simultaneously excited at two ore more resonance frequencies. Such excitation schemes result in one additional amplitude and phase images for each driven resonance, and potentially convey more information about the surface under investigation. Here we present a theoretical basis for using this information to approximate the parameters of a tip-surface interaction model. The theory is verified by simulations with added noise corresponding to room-temperature measurements.

cond-mat.mes-hall

Cross-correlation asymmetries and causal relationships between stock and market risk

We study historical correlations and lead-lag relationships between individual stock risk (volatility of daily stock returns) and market risk (volatility of daily returns of a market-representative portfolio) in the US stock market. We consider the cross-correlation functions averaged over all stocks, using 71 stock prices from the Standard \& Poor's 500 index for 1994--2013. We focus on the behavior of the cross-correlations at the times of financial crises with significant jumps of market volatility. The observed historical dynamics showed that the dependence between the risks was almost linear during the US stock market downturn of 2002 and after the US housing bubble in 2007, remaining on that level until 2013. Moreover, the averaged cross-correlation function often had an asymmetric shape with respect to zero lag in the periods of high correlation. We develop the analysis by the application of the linear response formalism to study underlying causal relations. The calculated response functions suggest the presence of characteristic regimes near financial crashes, when the volatility of an individual stock follows the market volatility and vice versa.

q-fin.ST

Dynamic Calibration of Higher Eigenmode Parameters of a Cantilever in Atomic Force Microscopy Using Tip-Surface Interactions

We present a theoretical framework for the dynamic calibration of the higher eigenmode parameters (stiffness and optical lever responsivity) of a cantilever. The method is based on the tip-surface force reconstruction technique and does not require any prior knowledge of the eigenmode shape or the particular form of the tip-surface interaction. The calibration method proposed requires a single-point force measurement using a multimodal drive and its accuracy is independent of the unknown physical amplitude of a higher eigenmode.

physics.ins-det

Reconstruction of Tip-Surface Interactions with Multimodal Intermodulation Atomic Force Microscopy

We propose a theoretical framework for reconstructing tip-surface interactions using the intermodulation technique when more than one eigenmode is required to describe the cantilever motion. Two particular cases of bimodal motion are studied numerically: one bending and one torsional mode, and two bending modes. We demonstrate the possibility of accurate reconstruction of a two-dimensional conservative force field for the former case, while dissipative forces are studied for the latter.

cond-mat.mes-hall