Searcharxiv⌕ Search

arXiv subjects

Tomislav Šmuc

Publications and source records attributed to Tomislav Šmuc.

6 recordsLinked to original sources

Approaches For Multi-View Redescription Mining

The task of redescription mining explores ways to re-describe different subsets of entities contained in a dataset and to reveal non-trivial associations between different subsets of attributes, called views. This interesting and challenging task is encountered in different scientific fields, and is addressed by a number of approaches that obtain redescriptions and allow for the exploration and analyses of attribute associations. The main limitation of existing approaches to this task is their inability to use more than two views. Our work alleviates this drawback. We present a memory efficient, extensible multi-view redescription mining framework that can be used to relate multiple, i.e. more than two views, disjoint sets of attributes describing one set of entities. The framework can use any multi-target regression or multi-label classification algorithm, with models that can be represented as sets of rules, to generate redescriptions. Multi-view redescriptions are built using incremental view-extending heuristic from initially created two-view redescriptions. In this work, we use different types of Predictive Clustering trees algorithms (regular, extra, with random output selection) and the Random Forest thereof in order to improve the quality of final redescription sets and/or execution time needed to generate them. We provide multiple performance analyses of the proposed framework and compare it against the naive approach to multi-view redescription mining. We demonstrate the usefulness of the proposed multi-view extension on several datasets, including a use-case on understanding of machine learning models - a topic of growing importance in machine learning and artificial intelligence in general.

cs.LG↗

Disentangling sources of influence in online social networks

Information propagation in online social networks is facilitated by two types of influence - endogenous (peer) influence that acts between users of the social network and exogenous (external) that corresponds to various external mediators such as online news media. However, inference of these influences from data remains a challenge, especially when data on the activation of users is scarce. In this paper we propose a methodology that yields estimates of both endogenous and exogenous influence using only a social network structure and a single activation cascade. Our method exploits the statistical differences between the two types of influence - endogenous is dependent on the social network structure and current state of each user while exogenous is independent of these. We evaluate our methodology on simulated activation cascades as well as on cascades obtained from several large Facebook political survey applications. We show that our methodology is able to provide estimates of endogenous and exogenous influence in online social networks, characterize activation of each individual user as being endogenously or exogenously driven, and identify most influential groups of users.

cs.SI↗

Using Redescription Mining to Relate Clinical and Biological Characteristics of Cognitively Impaired and Alzheimer's Disease Patients

We used redescription mining to find interpretable rules revealing associations between those determinants that provide insights about the Alzheimer's disease (AD). We extended the CLUS-RM redescription mining algorithm to a constraint-based redescription mining (CBRM) setting, which enables several modes of targeted exploration of specific, user-constrained associations. Redescription mining enabled finding specific constructs of clinical and biological attributes that describe many groups of subjects of different size, homogeneity and levels of cognitive impairment. We confirmed some previously known findings. However, in some instances, as with the attributes: testosterone, the imaging attribute Spatial Pattern of Abnormalities for Recognition of Early AD, as well as the levels of leptin and angiopoietin-2 in plasma, we corroborated previously debatable findings or provided additional information about these variables and their association with AD pathogenesis. Applying redescription mining on ADNI data resulted with the discovery of one largely unknown attribute: the Pregnancy-Associated Protein-A (PAPP-A), which we found highly associated with cognitive impairment in AD. Statistically significant correlations (p <= 0.01) were found between PAPP-A and various different clinical tests. The high importance of this finding lies in the fact that PAPP-A is a metalloproteinase, known to cleave insulin-like growth factor binding proteins. Since it also shares similar substrates with A Disintegrin and the Metalloproteinase family of enzymes that act as α-secretase to physiologically cleave amyloid precursor protein (APP) in the non-amyloidogenic pathway, it could be directly involved in the metabolism of APP very early during the disease course. Therefore, further studies should investigate the role of PAPP-A in the development of AD more thoroughly.

q-bio.QM↗

A framework for redescription set construction

Redescription mining is a field of knowledge discovery that aims at finding different descriptions of similar subsets of instances in the data. These descriptions are represented as rules inferred from one or more disjoint sets of attributes, called views. As such, they support knowledge discovery process and help domain experts in formulating new hypotheses or constructing new knowledge bases and decision support systems. In contrast to previous approaches that typically create one smaller set of redescriptions satisfying a pre-defined set of constraints, we introduce a framework that creates large and heterogeneous redescription set from which user/expert can extract compact sets of differing properties, according to its own preferences. Construction of large and heterogeneous redescription set relies on CLUS-RM algorithm and a novel, conjunctive refinement procedure that facilitates generation of larger and more accurate redescription sets. The work also introduces the variability of redescription accuracy when missing values are present in the data, which significantly extends applicability of the method. Crucial part of the framework is the redescription set extraction based on heuristic multi-objective optimization procedure that allows user to define importance levels towards one or more redescription quality criteria. We provide both theoretical and empirical comparison of the novel framework against current state of the art redescription mining algorithms and show that it represents more efficient and versatile approach for mining redescriptions from data.

cs.AI↗

Modeling peer and external influence in online social networks

Opinion polls mediated through a social network can give us, in addition to usual demographics data like age, gender and geographic location, a friendship structure between voters and the temporal dynamics of their activity during the voting process. Using a Facebook application we collected friendship relationships, demographics and votes of over ten thousand users on the referendum on the definition of marriage in Croatia held on 1st of December 2013. We also collected data on online news articles mentioning our application. Publication of these articles align closely with large peaks of voting activity, indicating that these external events have a crucial influence in engaging the voters. Also, existence of strongly connected friendship communities where majority of users vote during short time period, and the fact that majority of users in general tend to friend users that voted the same suggest that peer influence also has its role in engaging the voters. As we are not able to track activity of our users at all times, and we do not know their motivations for expressing their votes through our application, the question is whether we can infer peer and external influence using friendship network of users and the times of their voting. We propose a new method for estimation of magnitude of peer and external influence in friendship network and demonstrate its validity on both simulated and actual data.

cs.SI↗

News Cohesiveness: an Indicator of Systemic Risk in Financial Markets

Motivated by recent financial crises significant research efforts have been put into studying contagion effects and herding behaviour in financial markets. Much less has been said about influence of financial news on financial markets. We propose a novel measure of collective behaviour in financial news on the Web, News Cohesiveness Index (NCI), and show that it can be used as a systemic risk indicator. We evaluate the NCI on financial documents from large Web news sources on a daily basis from October 2011 to July 2013 and analyse the interplay between financial markets and financially related news. We hypothesized that strong cohesion in financial news reflects movements in the financial markets. Cohesiveness is more general and robust measure of systemic risk expressed in news, than measures based on simple occurrences of specific terms. Our results indicate that cohesiveness in the financial news is highly correlated with and driven by volatility on the financial markets.

cs.SI↗