SearcharxivSearch

arXiv subjects

Mattia Fumagalli

Publications and source records attributed to Mattia Fumagalli.

At least 19 recordsLinked to original sources

Time and Relations into Focus: Ontological Foundations of Object-Centric Event Data

Object-centric process mining is a new branch of process mining where events are associated with multiple objects, and where object-to-object interactions are essential to understand the process dynamics. Traditional event data models, also called case-centric, are unable to cope with the complexity introduced by these more refined relationships. Several models have been made to move from case-centric to Object-Centric Event Data (OCED), trying to retain simplicity as much as possible. Still, these suffer from inherent ambiguities, and lack a comprehensive support of essential dimensions related to time and (dynamic) relations. In this work, we propose to fill this gap by leveraging a well-founded ontology of events and bringing ontological foundations to OCED, with a three-step approach. First, we start from key open issues reported in the literature regarding current OCED metamodels, and witness their ambiguity and expressiveness limitations on illustrative and representative examples proposed therein. Second, we consider the OCED Core Model, currently proposed as the basis for defining a new standard for object-centric event data, and we enhance it by grounding it on a lightweight version of UFO-B called gUFO, a well-known foundational ontology tailored to the representation of objects, events, time, and their (dynamic) relations. This results in a new metamodel, which we call gOCED. The third contribution then shows how gOCED at once covers the features of existing metamodels preserving their simplicity, and extends them with the essential features needed to overcome the ambiguity and expressiveness issues reported in the literature.

cs.DB

An ontological lens on attack trees: Toward adequacy and interoperability

Attack Trees (AT) are a popular formalism for security analysis. They are meant to display an attacker's goal decomposed into attack steps needed to achieve it and compute certain security metrics (e.g., attack cost, probability, and damage). ATs offer three important services: (a) conceptual modeling capabilities for representing security risk management scenarios, (b) a qualitative assessment to find root causes and minimal conditions of successful attacks, and (c) quantitative analyses via security metrics computation under formal semantics, such as minimal time and cost among all attacks. Still, the AT language presents limitations due to its lack of ontological foundations, thus compromising associated services. Via an ontological analysis grounded in the Common Ontology of Value and Risk (COVER) -- a reference core ontology based on the Unified Foundational Ontology (UFO) -- we investigate the ontological adequacy of AT and reveal four significant shortcomings: (1) ambiguous syntactical terms that can be interpreted in various ways; (2) ontological deficit concerning crucial domain-specific concepts; (3) lacking modeling guidance to construct ATs decomposing a goal; (4) lack of semantic interoperability, resulting in ad hoc stand-alone tools. We also discuss existing incremental solutions and how our analysis paves the way for overcoming those issues through a broader approach to risk management modeling.

cs.CR

WATCHDOG: an ontology-aWare risk AssessmenT approaCH via object-oriented DisruptiOn Graphs

When considering risky events or actions, we must not downplay the role of involved objects: a charged battery in our phone averts the risk of being stranded in the desert after a flat tyre, and a functional firewall mitigates the risk of a hacker intruding the network. The Common Ontology of Value and Risk (COVER) highlights how the role of objects and their relationships remains pivotal to performing transparent, complete and accountable risk assessment. In this paper, we operationalize some of the notions proposed by COVER -- such as parthood between objects and participation of objects in events/actions -- by presenting a new framework for risk assessment: WATCHDOG. WATCHDOG enriches the expressivity of vetted formal models for risk -- i.e., fault trees and attack trees -- by bridging the disciplines of ontology and formal methods into an ontology-aware formal framework composed by a more expressive modelling formalism, Object-Oriented Disruption Graphs (DOGs), logic (DOGLog) and an intermediate query language (DOGLang). With these, WATCHDOG allows risk assessors to pose questions about disruption propagation, disruption likelihood and risk levels, keeping the fundamental role of objects at risk always in sight.

cs.AI

Mining Frequent Structures in Conceptual Models

The problem of using structured methods to represent knowledge is well-known in conceptual modeling and has been studied for many years. It has been proven that adopting modeling patterns represents an effective structural method. Patterns are, indeed, generalizable recurrent structures that can be exploited as solutions to design problems. They aid in understanding and improving the process of creating models. The undeniable value of using patterns in conceptual modeling was demonstrated in several experimental studies. However, discovering patterns in conceptual models is widely recognized as a highly complex task and a systematic solution to pattern identification is currently lacking. In this paper, we propose a general approach to the problem of discovering frequent structures, as they occur in conceptual modeling languages. As proof of concept, we implement our approach by focusing on two widely-used conceptual modeling languages. This implementation includes an exploratory tool that integrates a frequent subgraph mining algorithm with graph manipulation techniques. The tool processes multiple conceptual models and identifies recurrent structures based on various criteria. We validate the tool using two state-of-the-art curated datasets: one consisting of models encoded in OntoUML and the other in ArchiMate. The primary objective of our approach is to provide a support tool for language engineers. This tool can be used to identify both effective and ineffective modeling practices, enabling the refinement and evolution of conceptual modeling languages. Furthermore, it facilitates the reuse of accumulated expertise, ultimately supporting the creation of higher-quality models in a given language.

cs.AI

Towards a Gateway for Knowledge Graph Schemas Collection, Analysis, and Embedding

One of the significant barriers to the training of statistical models on knowledge graphs is the difficulty that scientists have in finding the best input data to address their prediction goal. In addition to this, a key challenge is to determine how to manipulate these relational data, which are often in the form of particular triples (i.e., subject, predicate, object), to enable the learning process. Currently, many high-quality catalogs of knowledge graphs, are available. However, their primary goal is the re-usability of these resources, and their interconnection, in the context of the Semantic Web. This paper describes the LiveSchema initiative, namely, a first version of a gateway that has the main scope of leveraging the gold mine of data collected by many existing catalogs collecting relational data like ontologies and knowledge graphs. At the current state, LiveSchema contains - 1000 datasets from 4 main sources and offers some key facilities, which allow to: i) evolving LiveSchema, by aggregating other source catalogs and repositories as input sources; ii) querying all the collected resources; iii) transforming each given dataset into formal concept analysis matrices that enable analysis and visualization services; iv) generating models and tensors from each given dataset.

cs.AI

Integrating 3D City Data through Knowledge Graphs

CityGML is a widely adopted standard by the Open Geospatial Consortium (OGC) for representing and exchanging 3D city models. The representation of semantic and topological properties in CityGML makes it possible to query such 3D city data to perform analysis in various applications, e.g., security management and emergency response, energy consumption and estimation, and occupancy measurement. However, the potential of querying CityGML data has not been fully exploited. The official GML/XML encoding of CityGML is only intended as an exchange format but is not suitable for query answering. The most common way of dealing with CityGML data is to store them in the 3DCityDB system as relational tables and then query them with the standard SQL query language. Nevertheless, for end users, it remains a challenging task to formulate queries over 3DCityDB directly for their ad-hoc analytical tasks, because there is a gap between the conceptual semantics of CityGML and the relational schema adopted in 3DCityDB. In fact, the semantics of CityGML itself can be modeled as a suitable ontology. The technology of Knowledge Graphs (KGs), where an ontology is at the core, is a good solution to bridge such a gap. Moreover, embracing KGs makes it easier to integrate with other spatial data sources, e.g., OpenStreetMap and existing (Geo)KGs (e.g., Wikidata, DBPedia, and GeoNames), and to perform queries combining information from multiple data sources. In this work, we describe a CityGML KG framework to populate the concepts in the CityGML ontology using declarative mappings to 3DCityDB, thus exposing the CityGML data therein as a KG. To demonstrate the feasibility of our approach, we use CityGML data from the city of Munich as test data and integrate OpenStreeMap data in the same area.

cs.DB

Towards Ranking Schemas by Focus

The main goal of this paper is to evaluate knowledge base schemas, modeled as a set of entity types, each such type being associated with a set of properties, according to their focus. We intuitively model the notion of focus as ''the state or quality of being relevant in storing and retrieving information''. This definition of focus is adapted from the notion of ''categorization purpose'', as first defined in cognitive psychology, thus giving us a high level of understandability on the side of users. In turn, this notion is formalized based on a set of knowledge metrics that, for any given focus, rank knowledge base schemas according to their quality. We apply the proposed methodology to more than 200 state-of-the-art knowledge base schemas. The experimental results show the utility of our approach

cs.AI

Towards an Ontology-Driven Approach for Process-Aware Risk Propagation

The rapid development of cyber-physical systems creates an increasing demand for a general approach to risk, especially considering how physical and digital components affect the processes of the system itself. In risk analytics and management, risk propagation is a central technique, which allows the calculation of the cascading effect of risk within a system and supports risk mitigation activities. However, one open challenge is to devise a process-aware risk propagation solution that can be used to assess the impact of risk at different levels of abstraction, accounting for actors, processes, physical-digital objects, and their interrelations. To address this challenge, we propose a process-aware risk propagation approach that builds on two main components: i. an ontology, which supports functionalities typical of Semantic Web technologies (SWT), and semantics-based intelligent systems, representing a system with processes and objects having different levels of abstraction, and ii. a method to calculate the propagation of risk within the given system. We implemented our approach in a proof-of-concept tool, which was validated and demonstrated in the cybersecurity domain.

cs.IT

Popularity Driven Data Integration

More and more, with the growing focus on large scale analytics, we are confronted with the need of integrating data from multiple sources. The problem is that these data are impossible to reuse as-is. The net result is high cost, with the further drawback that the resulting integrated data will again be hardly reusable as-is. iTelos is a general purpose methodology aiming at minimizing the effects of this process. The intuition is that data will be treated differently based on their popularity: the more a certain set of data have been reused, the more they will be reused and the less they will be changed across reuses, thus decreasing the overall data preprocessing costs, while increasing backward compatibility and future sharing

cs.AI

LiveSchema: A Gateway Towards Learning on Knowledge Graph Schemas

One of the major barriers to the training of algorithms on knowledge graph schemas, such as vocabularies or ontologies, is the difficulty that scientists have in finding the best input resource to address the target prediction tasks. In addition to this, a key challenge is to determine how to manipulate (and embed) these data, which are often in the form of particular triples (i.e., subject, predicate, object), to enable the learning process. In this paper, we describe the LiveSchema initiative, namely a gateway that offers a family of services to easily access, analyze, transform and exploit knowledge graph schemas, with the main goal of facilitating the reuse of these resources in machine learning use cases. As an early implementation of the initiative, we also advance an online catalog, which relies on more than 800 resources, with the first set of example services.

cs.AI

iTelos -- Purpose Driven Knowledge Graph Generation

When building a new application we are more and more confronted with the need of reusing and integrating pre-existing knowledge, e.g., ontologies, schemas, data of any kind, from multiple sources. Nevertheless, it is a fact that this prior knowledge is virtually impossible to reuse as-is. This difficulty is the cause of high costs, with the further drawback that the resulting application will again be hardly reusable. It is a negative loop which consistently reinforces itself. iTelos is a general purpose methodology aiming at minimizing as much as possible the effects of this loop. iTelos is based on the intuition that the data level and the schema level of an application should be developed independently, thus allowing for maximum flexibility in the reuse of the prior knowledge, but under the overall guidance of the needs to be satisfied, formalized as competence queries. This intuition is implemented by codifying all the requirements, including those concerning reuse, as part of an a-priori defined purpose, which is then used to drive a middle-out development process where the application schema and data are continuously aligned.

cs.DB

Ages of massive galaxies at $0.5 < z < 2.0$ from 3D-HST rest-frame optical spectroscopy

We present low-resolution near-infrared stacked spectra from the 3D-HST survey up to $z=2.0$ and fit them with commonly used stellar population synthesis models: BC03 (Bruzual & Charlot, 2003), FSPS10 (Flexible Stellar Population Synthesis, Conroy & Gunn 2010), and FSPS-C3K (Conroy, Kurucz, Cargile, Castelli, in prep). The accuracy of the grism redshifts allows the unambiguous detection of many emission and absorption features, and thus a first systematic exploration of the rest-frame optical spectra of galaxies up to $z=2$. We select massive galaxies ($\rm log(M_{*} / M_{\odot}) > 10.8$), we divide them into quiescent and star-forming via a rest-frame color-color technique, and we median-stack the samples in 3 redshift bins between $z=0.5$ and $z=2.0$. We find that stellar population models fit the observations well at wavelengths below $\rm 6500 Å$ rest-frame, but show systematic residuals at redder wavelengths. The FSPS-C3K model generally provides the best fits (evaluated with a $χ^2_{red}$ statistics) for quiescent galaxies, while BC03 performs the best for star-forming galaxies. The stellar ages of quiescent galaxies implied by the models, assuming solar metallicity, vary from 4 Gyr at $z \sim 0.75$ to 1.5 Gyr at $z \sim 1.75$, with an uncertainty of a factor of 2 caused by the unknown metallicity. On average the stellar ages are half the age of the Universe at these redshifts. We show that the inferred evolution of ages of quiescent galaxies is in agreement with fundamental plane measurements, assuming an 8 Gyr age for local galaxies. For star-forming galaxies the inferred ages depend strongly on the stellar population model and the shape of the assumed star-formation history.

astro-ph.GA

The 3D-HST Survey: Hubble Space Telescope WFC3/G141 grism spectra, redshifts, and emission line measurements for $\sim 100,000$ galaxies

We present reduced data and data products from the 3D-HST survey, a 248-orbit HST Treasury program. The survey obtained WFC3 G141 grism spectroscopy in four of the five CANDELS fields: AEGIS, COSMOS, GOODS-S, and UDS, along with WFC3 $H_{140}$ imaging, parallel ACS G800L spectroscopy, and parallel $I_{814}$ imaging. In a previous paper (Skelton et al. 2014) we presented photometric catalogs in these four fields and in GOODS-N, the fifth CANDELS field. Here we describe and present the WFC3 G141 spectroscopic data, again augmented with data from GO-1600 in GOODS-N. The data analysis is complicated by the fact that no slits are used: all objects in the WFC3 field are dispersed, and many spectra overlap. We developed software to automatically and optimally extract interlaced 2D and 1D spectra for all objects in the Skelton et al. (2014) photometric catalogs. The 2D spectra and the multi-band photometry were fit simultaneously to determine redshifts and emission line strengths, taking the morphology of the galaxies explicitly into account. The resulting catalog has 98,663 measured redshifts and line strengths down to $JH_{IR}\leq 26$ and 22,548 with $JH_{IR}\leq 24$, where we comfortably detect continuum emission. Of this sample 5,459 galaxies are at $z>1.5$ and 9,621 are at $0.7<z<1.5$, where H$α$ falls in the G141 wavelength coverage. Based on comparisons with ground-based spectroscopic redshifts, and on analyses of paired galaxies and repeat observations, the typical redshift error for $JH_{IR}\leq 24$ galaxies in our catalog is $σ_z \approx 0.003 \times (1+z)$, i.e., one native WFC3 pixel. The $3σ$ limit for emission line fluxes of point sources is $1.5\times10^{-17}$ ergs s$^{-1}$ cm$^{-2}$. We show various representations of the full dataset, as well as individual examples that highlight the range of spectra that we find in the survey.

astro-ph.GA

Forming Compact Massive Galaxies

In this paper we study a key phase in the formation of massive galaxies: the transition of star forming galaxies into massive (M_stars~10^11 Msun), compact (r_e~1 kpc) quiescent galaxies, which takes place from z~3 to z~1.5. We use HST grism redshifts and extensive photometry in all five 3D-HST/CANDELS fields, more than doubling the area used previously for such studies, and combine these data with Keck MOSFIRE and NIRSPEC spectroscopy. We first confirm that a population of massive, compact, star forming galaxies exists at z~2, using K-band spectroscopy of 25 of these objects at 2.0<z<2.5. They have a median NII/Halpha ratio of 0.6, are highly obscured with SFR(tot)/SFR(Halpha)~10, and have a large range of observed line widths. We infer from the kinematics and spatial distribution of Halpha that the galaxies have rotating disks of ionized gas that are a factor of ~2 more extended than the stellar distribution. By combining measurements of individual galaxies, we find that the kinematics are consistent with a nearly Keplerian fall-off from V_rot~500 km/s at 1 kpc to V_rot~250 km/s at 7 kpc, and that the total mass out to this radius is dominated by the dense stellar component. Next, we study the size and mass evolution of the progenitors of compact massive galaxies. Even though individual galaxies may have had complex histories with periods of compaction and mergers, we show that the population of progenitors likely followed a simple inside-out growth track in the size-mass plane of d(log r_e) ~ 0.3 d(log M_stars). This mode of growth gradually increases the stellar mass within a fixed physical radius, and galaxies quench when they reach a stellar density or velocity dispersion threshold. As shown in other studies, the mode of growth changes after quenching, as dry mergers take the galaxies on a relatively steep track in the size-mass plane.

astro-ph.GA

Where stars form: inside-out growth and coherent star formation from HST Halpha maps of 2676 galaxies across the main sequence at z~1

We present Ha maps at 1kpc spatial resolution for star-forming galaxies at z~1, made possible by the WFC3 grism on HST. Employing this capability over all five 3D-HST/CANDELS fields provides a sample of 2676 galaxies. By creating deep stacked Halpha (Ha) images, we reach surface brightness limits of 1x10^-18\erg\s\cm^2\arcsec^2, allowing us to map the distribution of ionized gas out to >10kpc for typical L* galaxies at this epoch. We find that the spatial extent of the Ha distribution increases with stellar mass as r(Ha)[kpc]=1.5(Mstars/10^10Msun)^0.23. Furthermore, the Ha emission is more extended than the stellar continuum emission, consistent with inside-out assembly of galactic disks. This effect, however, is mass dependent with r(Ha)/r(stars)=1.1(M/10^10Msun)^0.054, such that at low masses r(Ha)~r(stars). We map the Ha distribution as a function of SFR(IR+UV) and find evidence for `coherent star formation' across the SFR-M plane: above the main sequence, Ha is enhanced at all radii; below the main sequence, Ha is depressed at all radii. This suggests that at all masses the physical processes driving the enhancement or suppression of star formation act throughout the disks of galaxies. It also confirms that the scatter in the star forming main sequence is real and caused by variations in the star formation rate at fixed mass. At high masses (10^10.5<M/Msun<10^11), above the main sequence, Ha is particularly enhanced in the center, plausibly building bulges and/or supermassive black holes. Below the main sequence, the star forming disks are more compact and a strong central dip in the EW(Ha), and the inferred specific star formation rate, appears. Importantly though, across the entirety of the SFR-M plane, the absolute star formation rate as traced by Ha is always centrally peaked, even in galaxies below the main sequence.

astro-ph.GA

On the importance of using appropriate spectral models to derive physical properties of galaxies at 0.7<z<2.8

Interpreting observations of distant galaxies in terms of constraints on physical parameters - such as stellar mass, star-formation rate (SFR) and dust optical depth - requires spectral synthesis modelling. We analyse the reliability of these physical parameters as determined under commonly adopted `classical' assumptions: star-formation histories assumed to be exponentially declining functions of time, a simple dust law and no emission-line contribution. Improved modelling techniques and data quality now allow us to use a more sophisticated approach, including realistic star-formation histories, combined with modern prescriptions for dust attenuation and nebular emission (Pacifici et al. 2012). We present a Bayesian analysis of the spectra and multi-wavelength photometry of 1048 galaxies from the 3D-HST survey in the redshift range 0.7<z<2.8 and in the stellar mass range 9<log(M/Mo)<12. We find that, using the classical spectral library, stellar masses are systematically overestimated (~0.1 dex) and SFRs are systematically underestimated (~0.6 dex) relative to our more sophisticated approach. We also find that the simultaneous fit of photometric fluxes and emission-line equivalent widths helps break a degeneracy between SFR and optical depth of the dust, reducing the uncertainties on these parameters. Finally, we show how the biases of classical approaches can affect the correlation between stellar mass and SFR for star-forming galaxies (the `Star-Formation Main Sequence'). We conclude that the normalization, slope and scatter of this relation strongly depend on the adopted approach and demonstrate that the classical, oversimplified approach cannot recover the true distribution of stellar mass and SFR.

astro-ph.GA

How dead are dead galaxies? Mid-Infrared fluxes of quiescent galaxies at redshift 0.3 < z < 2.5: implications for star formation rates and dust heating

We investigate the star formation rates of quiescent galaxies at high redshift (0.3 < z < 2.5) using 3D-HST WFC3 grism spectroscopy and Spitzer mid-infrared data. We select quiescent galaxies on the basis of the widely used UVJ color-color criteria. Spectral energy distribution fitting (rest frame optical and near-IR) indicates very low star formation rates for quiescent galaxies (sSFR ~ 10^-12 yr^-1). However, SED fitting can miss star formation if it is hidden behind high dust obscuration and ionizing radiation is re-emitted in the mid-infrared. It is therefore fundamental to measure the dust-obscured SFRs with a mid-IR indicator. We stack the MIPS-24um images of quiescent objects in five redshift bins centered on z = 0.5, 0.9, 1.2, 1.7, 2.2 and perform aperture photometry. Including direct 24um detections, we find sSFR ~ 10^-11.9 * (1+z)^4 yr^-1. These values are higher than those indicated by SED fitting, but at each redshift they are 20-40 times lower than those of typical star forming galaxies. The true SFRs of quiescent galaxies might be even lower, as we show that the mid-IR fluxes can be due to processes unrelated to ongoing star formation, such as cirrus dust heated by old stellar populations and circumstellar dust. Our measurements show that star formation quenching is very efficient at every redshift. The measured SFR values are at z > 1.5 marginally consistent with the ones expected from gas recycling (assuming that mass loss from evolved stars refuels star formation) and well above that at lower redshifts.

astro-ph.CO

Constraining the Low-Mass Slope of the Star Formation Sequence at 0.5<z<2.5

We constrain the slope of the star formation rate ($\logΨ$) to stellar mass ($\log\mathrm{M_{\star}}$) relation down to $\log(\mathrm{M_{\star}/M_{\odot}})=8.4$ ($\log(\mathrm{M_{\star}/M_{\odot}})=9.2$) at $z=0.5$ ($z=2.5$) with a mass-complete sample of 39,106 star-forming galaxies selected from the 3D-HST photometric catalogs, using deep photometry in the CANDELS fields. For the first time, we find that the slope is dependent on stellar mass, such that it is steeper at low masses ($\log\mathrmΨ\propto\log\mathrm{M_{\star}}$) than at high masses ($\log\mathrmΨ\propto(0.3-0.6)\log\mathrm{M_{\star}}$). These steeper low mass slopes are found for three different star formation indicators: the combination of the ultraviolet (UV) and infrared (IR), calibrated from a stacking analysis of Spitzer/MIPS 24$μ$m imaging; $β$-corrected UV SFRs; and H$α$ SFRs. The normalization of the sequence evolves differently in distinct mass regimes as well: for galaxies less massive than $\log(\mathrm{M_{\star}/M_{\odot}})<10$ the specific SFR ($Ψ/\mathrm{M_{\star}}$) is observed to be roughly self-similar with $Ψ/\mathrm{M_{\star}}\propto(1+z)^{1.9}$, whereas more massive galaxies show a stronger evolution with $Ψ/\mathrm{M_{\star}}\propto(1+z)^{2.2-3.5}$ for $\log(\mathrm{M_{\star}/M_{\odot}})=10.2-11.2$. The fact that we find a steep slope of the star formation sequence for the lower mass galaxies will help reconcile theoretical galaxy formation models with the observations. The results of this study support the analytical conclusions of Leja et al. (2014).

astro-ph.GA