SearcharxivSearch

arXiv subjects

Behshid Behkamal

Publications and source records attributed to Behshid Behkamal.

5 recordsLinked to original sources

Heterogeneity in Entity Matching: A Survey and Experimental Analysis

Entity matching (EM) is a fundamental task in data integration and analytics, essential for identifying records that refer to the same real-world entity across diverse sources. In practice, datasets often differ widely in structure, format, schema, and semantics, creating substantial challenges for EM. We refer to this setting as Heterogeneous EM (HEM). This survey offers a unified perspective on HEM by introducing a taxonomy, grounded in prior work, that distinguishes two primary categories -- representation and semantic heterogeneity -- and their subtypes. The taxonomy provides a systematic lens for understanding how variations in data form and meaning shape the complexity of matching tasks. We then connect this framework to the FAIR principles -- Findability, Accessibility, Interoperability, and Reusability -- demonstrating how they both reveal the challenges of HEM and suggest strategies for mitigating them. Building on this foundation, we critically review recent EM methods, examining their ability to address different heterogeneity types, and conduct targeted experiments on state-of-the-art models to evaluate their robustness and adaptability under semantic heterogeneity. Our analysis uncovers persistent limitations in current approaches and points to promising directions for future research, including multimodal matching, human-in-the-loop workflows, deeper integration with large language models and knowledge graphs, and fairness-aware evaluation in heterogeneous settings.

cs.DB

Geo-Enabled Business Process Modeling

Recent advancements in location-aware analytics have created novel opportunities in different domains. In the area of process mining, enriching process models with geolocation helps to gain a better understanding of how the process activities are executed in practice. In this paper, we introduce our idea of geo-enabled process modeling and report on our industrial experience. To this end, we present a real-world case study to describe the importance of considering the location in process mining. Then we discuss the shortcomings of currently available process mining tools and propose our novel approach for modeling geo-enabled processes focusing on 1) increasing process interpretability through geo-visualization, 2) incorporating location-related metadata into process analysis, and 3) using location-based measures for the assessment of process performance. Finally, we conclude the paper by future research directions.

cs.SE

Big Data Quality: A systematic literature review and future research directions

One of the most significant problems of Big Data is to extract knowledge through the huge amount of data. The usefulness of the extracted information depends strongly on data quality. In addition to the importance, data quality has recently been taken into consideration by the big data community and there is not any comprehensive review conducted in this area. Therefore, the purpose of this study is to review and present the state of the art on the quality of big data research through a hierarchical framework. The dimensions of the proposed framework cover various aspects in the quality assessment of Big Data including 1) the processing types of big data, i.e. stream, batch, and hybrid, 2) the main task, and 3) the method used to conduct the task. We compare and critically review all of the studies reported during the last ten years through our proposed framework to identify which of the available data quality assessment methods have been successfully adopted by the big data community. Finally, we provide a critical discussion on the limitations of existing methods and offer suggestions on potential valuable research directions that can be taken in future research in this domain.

cs.DB

A metric Suite for Systematic Quality Assessment of Linked Open Data

Abstract- The vision of the Linked Open Data (LOD) initiative is to provide a distributed model for publishing and meaningfully interlinking open data. The realization of this goal depends strongly on the quality of the data that is published as a part of the LOD. This paper focuses on the systematic quality assessment of datasets prior to publication on the LOD cloud. To this end, we identify important quality deficiencies that need to be avoided and/or resolved prior to the publication of a dataset. We then propose a set of metrics to measure these quality deficiencies in a dataset. This way, we enable the assessment and identification of undesirable quality characteristics of a dataset through our proposed metrics. This will help publishers to filter out low-quality data based on the quality assessment results, which in turn enables data consumers to make better and more informed decisions when using the open datasets.

cs.DB

Contextualization of Big Data Quality: A framework for comparison

With the advent of big data applications and the increasing amount of data being produced in these applications, the importance of efficient methods for big data analysis has become highly evident. However, the success of any such method will be hindered should the data lacks the required quality. Big data quality assessment is therefore a major requirement for any organization or business that use big data analytics for its decision making. On the other hand, using contextual information is advantageous in many analysis tasks in various domains, e.g. user behavior analysis in the social networks. However, the big data quality assessment has benefited less from this potential. There is a vast variety of data sources in the big data domain that can be utilized to improve the quality evaluation of big data. Including contextual information provided by these sources into the big data quality assessment process is an emerging trend towards more advanced techniques aimed at enhancing the performance and accuracy of quality assessment. This paper presents a context classification framework for big data quality, categorizing the context features into four primary dimensions: 1) context category, 2) data source type that contextual features come from, 3) discovery and extraction method of context, and 4) the quality factors affected by the contextual data. The proposed model introduces new context features and dimensions that need to be taken into consideration in quality assessment of big data. The initial evaluation demonstrates that the model is more understandable, more comprehensive, richer, and more useful compared to existing models.

cs.CY