Searcharxiv⌕ Search

arXiv subjects

Manuel Parra-Royón

Publications and source records attributed to Manuel Parra-Royón.

8 recordsLinked to original sources

Towards the exploration of astrophysical datacubes through volumetric rendering within CARTA

Data products from astrophysical observations may contain three or more dimensions: most commonly two spatial dimensions and a third spectral dimension, and possibly additional axes such as polarisation or time. Visualising three-dimensional image data is particularly useful for analysing the kinematics and evolution of various objects; however, the large scale of data produced by next-generation observatories and the complexity of displaying 3D data on a 2D screen make visualisation challenging. In this article, we present a scalable and interactive 3D visualisation approach within the astronomy software CARTA. We implemented a prototype 3D rendering widget with ThreeJS and known techniques to improve efficiency (e.g. compression, downsampling) and found that large datacubes up to 128 GB can be visualised locally with a loading time of a few minutes. We analysed the performance in the frontend and in the backend, finding that the bottleneck of our implementation lies in the frontend. This is due to storing the data in a single array in the frontend, unlike in the backend, where it is accessed and sent in blocks. This study paves the way for making 3D visualisation of large datacubes available to the community by implementing a 3D rendering widget in a future release of CARTA.

astro-ph.IM↗

Bringing Computation to the data: Interoperable serverless function execution for astrophysical data analysis in the SRCNet

Serverless computing is a paradigm in which the underlying infrastructure is fully managed by the provider, enabling applications and services to be executed with elastic resource provisioning and minimal operational overhead. A core model within this paradigm is Function-as-a-Service (FaaS), where lightweight functions are deployed and triggered on demand, scaling seamlessly with workload. FaaS offers flexibility, cost-effectiveness, and fine-grained scalability, qualities particularly relevant for large-scale scientific infrastructures where data volumes are too large to centralise and computation must increasingly occur close to the data. The Square Kilometre Array Observatory (SKAO) exemplifies this challenge. Once operational, it will generate about 700~PB of data products annually, distributed across the SKA Regional Centre Network (SRCNet), a federation of international centres providing storage, computing, and analysis services. In such a context, FaaS offers a mechanism to bring computation to the data. We studied the principles of serverless and FaaS computing and explored their application to radio astronomy workflows. Representative functions for astrophysical data analysis were developed and deployed, including micro-functions derived from existing libraries and wrappers around domain-specific applications. In particular, a Gaussian convolution function was implemented and integrated within the SRCNet ecosystem. The use case demonstrates that FaaS can be embedded into the existing SRCNet ecosystem of services, allowing functions to run directly at sites where data replicas are stored. This reduces latency, minimises transfers, and improves efficiency, aligning with federated, data-proximate computation. The results show that serverless models provide a scalable and efficient pathway to address the data volumes of the SKA era.

cs.DC↗

Bringing computation to the data: A MOEA-driven approach for optimising data processing in the context of the SKA and SRCNet

The Square Kilometre Array (SKA) will generate unprecedented data volumes, making efficient data processing a critical challenge. Within this context, the SKA Regional Centres Network (SRCNet) must operate in a near-exascale environment where traditional data-centric computing models based on moving large datasets to centralised resources are no longer viable due to network and storage bottlenecks. To address this limitation, this work proposes a shift towards distributed and in-situ computing, where computation is moved closer to the data. We explore the integration of Function-as-a-Service (FaaS) with an intelligent decision-making entity based on Evolutionary Algorithms (EAs) to optimise data-intensive workflows within SRCNet. FaaS enables lightweight and modular function execution near data sources while abstracting infrastructure management. The proposed decision-making entity employs Multi-Objective Evolutionary Algorithms (MOEAs) to explore near-optimal execution plans considering execution time and energy consumption, together with constraints related to data location and transfer costs. This work establishes a baseline framework for efficient and cost-aware computation-to-data strategies within the SRCNet architecture.

cs.DC↗

Semantic Model for the SKA Regional Centre Network

The unprecedented volume of data from the Square Kilometre Array (SKA) telescopes will require the implementation of robust and solid strategies for efficient data processing and management. In this context, the SKA Regional Centre Network (SRCNet) -- a collaborative global infrastructure comprising multiple regional centres distributed across various geographical regions around the globe -- is poised to play a critical role. This network will be instrumental in facilitating the effective handling and analysis of extensive data streams generated by the telescopes, thereby enabling significant advancements in astronomical research and exploration. This paper introduces a semantic model implemented with JSON-LD designed specifically for the SRCNet, detailing its architecture, data distribution, and computing service. By explicitly defining nodes, resources, relationships, and workflows, this model lays a foundation for interoperability and efficient resource management within the distributed network. The model presented in this text supports two possible configurations: centralized and decentralized -- depending where data reside -- enabling a future service broker to efficiently plan workflows by querying nodes for real-time system availability. Consistency tests conducted using SPARQL queries were made on the model in order to validate and test its integrity. Therefore, this research contributes to the advancement of semantic modeling in astronomy by addressing the semantic model for the SRCNet, a topic that has not been previously explored. This semantic model serves as a precursor to the development of a precise mathematical representation of the network and establishes a foundational framework for a future service broker.

astro-ph.IM↗

Photometry and kinematics of dwarf galaxies from the Apertif HI survey

Context. Understanding the dwarf galaxy population in low density environments is crucial for testing the LCDM cosmological model. The increase in diversity towards low mass galaxies is seen as an increase in the scatter of scaling relations such as the stellar mass-size and the baryonic Tully-Fisher relation (BTFR), and is also demonstrated by recent in-depth studies of an extreme subclass of dwarf galaxies of low surface brightness, but large physical sizes, called ultra-diffuse galaxies (UDGs). Aims. We select galaxies from the Apertif HI survey, and apply a constraint on their i-band absolute magnitude to exclude high mass systems. The sample consists of 24 galaxies, and span HI mass ranges of 8.6 < log ($M_{HI}/M_{Sun}$) < 9.7 and stellar mass range of 8.0 < log ($M_*/M_{Sun}$) < 9.7 (with only three galaxies having log ($M_*/M_{Sun}$) > 9). Methods. We determine the geometrical parameters of the HI and stellar discs, build kinematic models from the HI data using 3DBarolo, and extract surface brightness profiles in g-, r- and i-band from the Pan-STARRS 1 photometric survey. Results. We find that, at fixed stellar mass, our HI selected dwarfs have larger optical effective radii than isolated, optically-selected dwarfs from the literature. We find misalignments between the optical and HI morphologies for some of our sample. For most of our galaxies, we use the HI morphology to determine their kinematics, and we stress that deep optical observations are needed to trace the underlying stellar discs. Standard dwarfs in our sample follow the same BTFR of high-mass galaxies, whereas UDGs are slightly offset towards lower rotational velocities, in qualitative agreement with results from previous studies. Finally, our sample features a fraction (25%) of dwarf galaxies in pairs that is significantly larger with respect to previous estimates based on optical spectroscopic data.

astro-ph.GA↗

Integration of storage endpoints into a Rucio data lake, as an activity to prototype a SKA Regional Centres Network

The Square Kilometre Array (SKA) infrastructure will consist of two radio telescopes that will be the most sensitive telescopes on Earth. The SKA community will have to process and manage near exascale data, which will be a technical challenge for the coming years. In this respect, the SKA Global Network of Regional Centres plays a key role in data distribution and management. The SRCNet will provide distributed computing and data storage capacity, as well as other important services for the network. Within the SRCNet, several teams have been set up for the research, design and development of 5 prototypes. One of these prototypes is related to data management and distribution, where a data lake has been deployed using Rucio. In this paper we focus on the tasks performed by several of the teams to deploy new storage endpoints within the SKAO data lake. In particular, we will describe the steps and deployment instructions for the services required to provide the Rucio data lake with a new Rucio Storage Element based on StoRM and WebDAV within the Spanish SRC prototype.

astro-ph.IM↗

An approach to provide serverless scientific pipelines within the context of SKA

Function-as-a-Service (FaaS) is a type of serverless computing that allows developers to write and deploy code as individual functions, which can be triggered by specific events or requests. FaaS platforms automatically manage the underlying infrastructure, scaling it up or down as needed, being highly scalable, cost-effective and offering a high level of abstraction. Prototypes being developed within the SKA Regional Center Network (SRCNet) are exploring models for data distribution, software delivery and distributed computing with the goal of moving and executing computation to where the data is. Since SKA will be the largest data producer on the planet, it will be necessary to distribute this massive volume of data to the SRCNet nodes that will serve as a hub for computing and analysis operations on the closest data. Within this context, in this work we want to validate the feasibility of designing and deploying functions and applications commonly used in radio interferometry workflows within a FaaS platform to demonstrate the value of this computing model as an alternative to explore for data processing in the distributed nodes of the SRCNet. We have analyzed several FaaS platforms and successfully deployed one of them, where we have imported several functions using two different methods: microfunctions from the CASA framework, which are written in Python code, and highly specific native applications like wsclean. Therefore, we have designed a simple catalogue that can be easily scaled to provide all the key features of FaaS in highly distributed environments using orchestrators, as well as having the ability to integrate them with workflows or APIs. This paper contributes to the ongoing discussion of the potential of FaaS models for scientific data processing, particularly in the context of large-scale, distributed projects such as SKA.

cs.DC↗

Semantic of Cloud Computing services for Time Series workflows

Time series (TS) are present in many fields of knowledge, research, and engineering. The processing and analysis of TS are essential in order to extract knowledge from the data and to tackle forecasting or predictive maintenance tasks among others The modeling of TS is a challenging task, requiring high statistical expertise as well as outstanding knowledge about the application of Data Mining(DM) and Machine Learning (ML) methods. The overall work with TS is not limited to the linear application of several techniques, but is composed of an open workflow of methods and tests. These workflow, developed mainly on programming languages, are complicated to execute and run effectively on different systems, including Cloud Computing (CC) environments. The adoption of CC can facilitate the integration and portability of services allowing to adopt solutions towards services Internet Technologies (IT) industrialization. The definition and description of workflow services for TS open up a new set of possibilities regarding the reduction of complexity in the deployment of this type of issues in CC environments. In this sense, we have designed an effective proposal based on semantic modeling (or vocabulary) that provides the full description of workflow for Time Series modeling as a CC service. Our proposal includes a broad spectrum of the most extended operations, accommodating any workflow applied to classification, regression, or clustering problems for Time Series, as well as including evaluation measures, information, tests, or machine learning algorithms among others.

cs.CL↗