SearcharxivSearch

arXiv subjects

Mirian Halfeld-Ferrari

Publications and source records attributed to Mirian Halfeld-Ferrari.

3 recordsLinked to original sources

Extracting node comparison insights for the interactive exploration of property graphs

While scoring nodes in graphs to understand their importance (e.g., in terms of centrality) has been investigated for decades, comparing nodes in property graphs based on their properties has not, to our knowledge, yet been addressed. In this paper, we propose an approach to automatically extract comparison of nodes in property graphs, to support the interactive exploratory analysis of said graphs. We first present a way of devising comparison indicators using the context of nodes to be compared. Then, we formally define the problem of using these indicators to group the nodes so that the comparisons extracted are both significant and not straightforward. We propose various heuristics for solving this problem. Our tests on real property graph databases show that simple heuristics can be used to obtain insights within minutes while slower heuristics are needed to obtain insights of higher quality.

cs.DB

From Text to Databases: attribute grammar as database meta-model

We present a general methodology for structuring textual data, represented as syntax trees enriched with semantic information, guided by a meta-model G defined as an attribute grammar. The method involves an evolution process where both the instance and its grammar evolve, with instance transformations guided by rewriting rules and a similarity measure. Each new instance generates a corresponding grammar, culminating in a target grammar GT that satisfies G. This methodology is applied to build a database populated from textual data. The process generates both a database schema and its instance, independent of specific database models. We demonstrate the approach using clinical medical cases, where trees represent database instances and grammars act as database schemas. Key contributions include the proposal of a general attribute grammar G, a formalization of grammar evolution, and a proof-of-concept implementation for database structuring.

cs.DB

Querying Linked Data: how to ensure user's quality requirements

In the distributed and dynamic framework of the Web, data quality is a big challenge. The Linked Open Data (LOD) provides an enormous amount of data, the quality of which is difficult to control. Quality is intrinsically a matter of usage, so consumers need ways to specify quality rules that make sense for their use, in order to get only data conforming to these rules. We propose a user-side query framework equipped with a checker of constraints and confidence levels on data resulting from LOD providers\' query evaluations. We detail its theoretical foundations and we provide experimental results showing that the check additional cost is reasonable and that integrating the constraints in the queries further improves it significantly.

cs.DB