SearcharxivSearch

arXiv subjects

Kevin Vinsen

Publications and source records attributed to Kevin Vinsen.

15 recordsLinked to original sources

DeepPySR -- A Symbolic Regression Framework with Dynamic Pruning, Pareto Selection, and Hierarchical Composition for Real-World Scientific Discovery

Symbolic regression (SR) discovers analytical equations from data, yielding glass-box models with directly interpretable formulas, unlike black-box methods that rely on unstable post-hoc tools such as SHAP or LIME. This transparency is crucial in clinical medicine and social science, but SR faces three challenges: high-dimensional inputs, principled selection of Pareto-front formulae, and data irregularities such as multicollinearity and class imbalance. We introduce DeepPySR, which addresses these issues with a dynamic variable-pruning schedule to remove irrelevant features during search, an exponential Pareto selection criterion that eliminates trade-offs between accuracy and complexity, and a multi-layer architecture for hierarchical symbolic composition. On four Feynman physics benchmarks and seven biomedical and social-science datasets, DeepPySR outperforms PySR and baselines on body fat (R$^2$: 0.794 vs.\ 0.702), heart disease (F1: 0.898 vs.\ 0.787), student performance (R$^2$: 0.964 vs.\ 0.948), and Raine BMI (R$^2$: 0.525 vs.\ 0.370), producing interpretable formulas aligned with domain risk factors.

cs.LG

SM-Net: Learning a Continuous Spectral Manifold from Multiple Stellar Libraries

We present SM-Net, a machine-learning model that learns a continuous spectral manifold from multiple high-resolution stellar libraries. SM-Net generates stellar spectra directly from the fundamental stellar parameters effective temperature (Teff), surface gravity (log g), and metallicity (log Z). It is trained on a combined grid derived from the PHOENIX-Husser, C3K-Conroy, OB-PoWR, and TMAP-Werner libraries. By combining their parameter spaces, we construct a composite dataset that spans a broader and more continuous region of stellar parameter space than any individual library. The unified grid covers Teff = 2,000-190,000 K, log g = -1 to 9, and log Z = -4 to 1, with spectra spanning 3,000-100,000 Angstrom. Within this domain, SM-Net provides smooth interpolation across heterogeneous library boundaries. Outside the sampled region, it can produce numerically smooth exploratory predictions, although these extrapolations are not directly validated against reference models. Zero or masked flux values are treated as unknowns rather than physical zeros, allowing the network to infer missing regions using correlations learned from neighbouring grid points. Across 3,538 training and 11,530 test spectra, SM-Net achieves mean squared errors of 1.47 x 10^-5 on the training set and 2.34 x 10^-5 on the test set in the transformed log1p-scaled flux representation. Inference throughput exceeds 14,000 spectra per second on a single GPU. We also release the model together with an interactive web dashboard for real-time spectral generation and visualisation. SM-Net provides a fast, robust, and flexible data-driven complement to traditional stellar population synthesis libraries.

astro-ph.IM

Spatial Temporal Approach for High-Resolution Gridded Wind Forecasting across Southwest Western Australia

Accurate wind speed and direction forecasting is paramount across many sectors, spanning agriculture, renewable energy generation, and bushfire management. However, conventional forecasting models encounter significant challenges in precisely predicting wind conditions at high spatial resolutions for individual locations or small geographical areas (< 20 km2) and capturing medium to long-range temporal trends and comprehensive spatio-temporal patterns. This study focuses on a spatial temporal approach for high-resolution gridded wind forecasting at the height of 3 and 10 metres across large areas of the Southwest of Western Australia to overcome these challenges. The model utilises the data that covers a broad geographic area and harnesses a diverse array of meteorological factors, including terrain characteristics, air pressure, 10-metre wind forecasts from the European Centre for Medium-Range Weather Forecasts, and limited observation data from sparsely distributed weather stations (such as 3-metre wind profiles, humidity, and temperature), the model demonstrates promising advancements in wind forecasting accuracy and reliability across the entire region of interest. This paper shows the potential of our machine learning model for wind forecasts across various prediction horizons and spatial coverage. It can help facilitate more informed decision-making and enhance resilience across critical sectors.

cs.LG

Rapid localization of gravitational wave sources from compact binary coalescences using deep learning

The mergers of neutron star-neutron star and neutron star-black hole binaries are the most promising gravitational wave events with electromagnetic counterparts. The rapid detection, localization and simultaneous multi-messenger follow-up of these sources is of primary importance in the upcoming science runs of the LIGO-Virgo-KAGRA Collaboration. While prompt electromagnetic counterparts during binary mergers can last less than two seconds, the time scales of existing localization methods that use Bayesian techniques, varies from seconds to days. In this paper, we propose the first deep learning-based approach for rapid and accurate sky localization of all types of binary coalescences, including neutron star-neutron star and neutron star-black hole binaries for the first time. Specifically, we train and test a normalizing flow model on matched-filtering output from gravitational wave searches. Our model produces sky direction posteriors in milliseconds using a single P100 GPU, which is three to six orders of magnitude faster than Bayesian techniques.

gr-qc

Extraction of Binary Black Hole Gravitational Wave Signals from Detector Data Using Deep Learning

Accurate extractions of the detected gravitational wave (GW) signal waveforms are essential to validate a detection and to probe the astrophysics behind the sources producing the GWs. This however could be difficult in realistic scenarios where the signals detected by existing GW detectors could be contaminated with non-stationary and non-Gaussian noise. While the performance of existing waveform extraction methods are optimal, they are not fast enough for online application, which is important for multi-messenger astronomy. In this paper, we demonstrate that a deep learning architecture consisting of Convolutional Neural Network and bidirectional Long Short-Term Memory components can be used to extract binary black hole (BBH) GW waveforms from realistic noise in a few milli-seconds. We have tested our network systematically on injected GW signals, with component masses uniformly distributed in the range of 10 to 80 solar masses, on Gaussian noise and LIGO detector noise. We find that our model can extract GW waveforms with overlaps of more than 0.95 with pure Numerical Relativity templates for signals with signal-to-noise ratio (SNR) greater than six, and is also robust against interfering glitches. We then apply our model to all ten detected BBH events from the first (O1) and second (O2) observation runs, obtaining greater than 0.97 overlaps for all ten extracted BBH waveforms with the corresponding pure templates. We discuss the implication of our result and its future applications to GW localization and mass estimation.

gr-qc

Using Deep Learning to Localize Gravitational Wave Sources

In this paper, we report on the construction of a deep Artificial Neural Network (ANN) to localize simulated gravitational wave signals in the sky with high accuracy. We have modelled the sky as a sphere and have considered cases where the sphere is divided into 18, 50, 128, 1024, 2048 and 4096 sectors. The sky direction of the gravitational wave source is estimated by classifying the signal into one of these sectors based on it's right ascension and declination values for each of these cases. In order to do this, we have injected simulated binary black hole gravitational wave signals of component masses sampled uniformly between 30-80 solar mass into Gaussian noise and used the whitened strain values to obtain the input features for training our ANN. We input features such as the delays in arrival times, phase differences and amplitude ratios at each of the three detectors Hanford, Livingston and Virgo, from the raw time-domain strain values as well as from analytical versions of these signals, obtained through Hilbert transformation. We show that our model is able to classify gravitational wave samples, not used in the training process, into their correct sectors with very high accuracy (>90%) for coarse angular resolution using 18, 50 and 128 sectors. We also test our localization on test samples with injection parameters of the published LIGO binary black hole merger events GW150914, GW170818 and GW170823 for 1024, 2048 and 4096 sectors and compare the result with that from BAYESTAR and Parameter Estimation (PE). In addition, we report that the time taken by our model to localize one GW signal is around 0.018 secs on 14 Intel Xeon CPU cores.

astro-ph.IM

CHILES: HI morphology and galaxy environment at z=0.12 and z=0.17

We present a study of 16 HI-detected galaxies found in 178 hours of observations from Epoch 1 of the COSMOS HI Large Extragalactic Survey (CHILES). We focus on two redshift ranges between 0.108 <= z <= 0.127 and 0.162 <= z <= 0.183 which are among the worst affected by radio frequency interference (RFI). While this represents only 10% of the total frequency coverage and 18% of the total expected time on source compared to what will be the full CHILES survey, we demonstrate that our data reduction pipeline recovers high quality data even in regions severely impacted by RFI. We report on our in-depth testing of an automated spectral line source finder to produce HI total intensity maps which we present side-by-side with significance maps to evaluate the reliability of the morphology recovered by the source finder. We recommend that this become a common place manner of presenting data from upcoming HI surveys of resolved objects. We use the COSMOS 20k group catalogue, and we extract filamentary structure using the topological DisPerSE algorithm to evaluate the \hi\ morphology in the context of both local and large-scale environments and we discuss the shortcomings of both methods. Many of the detections show disturbed HI morphologies suggesting they have undergone a recent interaction which is not evident from deep optical imaging alone. Overall, the sample showcases the broad range of ways in which galaxies interact with their environment. This is a first look at the population of galaxies and their local and large-scale environments observed in HI by CHILES at redshifts beyond the z=0.1 Universe.

astro-ph.GA

GAMA/G10-COSMOS/3D-HST: The 0<z<5 cosmic star-formation history, stellar- and dust-mass densities

We use the energy-balance code MAGPHYS to determine stellar and dust masses, and dust corrected star-formation rates for over 200,000 GAMA galaxies, 170,000 G10-COSMOS galaxies and 200,000 3D-HST galaxies. Our values agree well with previously reported measurements and constitute a representative and homogeneous dataset spanning a broad range in stellar mass (10^8---10^12 Msol), dust mass (10^6---10^9 Msol), and star-formation rates (0.01---100 Msol per yr), and over a broad redshift range (0.0 < z < 5.0). We combine these data to measure the cosmic star-formation history (CSFH), the stellar-mass density (SMD), and the dust-mass density (DMD) over a 12 Gyr timeline. The data mostly agree with previous estimates, where they exist, and provide a quasi-homogeneous dataset using consistent mass and star-formation estimators with consistent underlying assumptions over the full time range. As a consequence our formal errors are significantly reduced when compared to the historic literature. Integrating our cosmic star-formation history we precisely reproduce the stellar-mass density with an ISM replenishment factor of 0.50 +/- 0.07, consistent with our choice of Chabrier IMF plus some modest amount of stripped stellar mass. Exploring the cosmic dust density evolution, we find a gradual increase in dust density with lookback time. We build a simple phenomenological model from the CSFH to account for the dust mass evolution, and infer two key conclusions: (1) For every unit of stellar mass which is formed 0.0065---0.004 units of dust mass is also formed; (2) Over the history of the Universe approximately 90 to 95 per cent of all dust formed has been destroyed and/or ejected.

astro-ph.GA

Galaxy And Mass Assembly: the evolution of the cosmic spectral energy distribution from z = 1 to z = 0

We present the evolution of the Cosmic Spectral Energy Distribution (CSED) from $z = 1 - 0$. Our CSEDs originate from stacking individual spectral energy distribution fits based on panchromatic photometry from the Galaxy and Mass Assembly (GAMA) and COSMOS datasets in ten redshift intervals with completeness corrections applied. Below $z = 0.45$, we have credible SED fits from 100 nm to 1 mm. Due to the relatively low sensitivity of the far-infrared data, our far-infrared CSEDs contain a mix of predicted and measured fluxes above $z = 0.45$. Our results include appropriate errors to highlight the impact of these corrections. We show that the bolometric energy output of the Universe has declined by a factor of roughly four -- from $5.1 \pm 1.0$ at $z \sim 1$ to $1.3 \pm 0.3 \times 10^{35}~h_{70}$~W~Mpc$^{-3}$ at the current epoch. We show that this decrease is robust to cosmic variance, SED modelling and other various types of error. Our CSEDs are also consistent with an increase in the mean age of stellar populations. We also show that dust attenuation has decreased over the same period, with the photon escape fraction at 150~nm increasing from $16 \pm 3$ at $z \sim 1$ to $24 \pm 5$ per cent at the current epoch, equivalent to a decrease in $A_\mathrm{FUV}$ of 0.4~mag. Our CSEDs account for $68 \pm 12$ and $61 \pm 13$ per cent of the cosmic optical and infrared backgrounds respectively as defined from integrated galaxy counts and are consistent with previous estimates of the cosmic infrared background with redshift.

astro-ph.GA

DALiuGE: A Graph Execution Framework for Harnessing the Astronomical Data Deluge

The Data Activated Liu Graph Engine - DALiuGE - is an execution framework for processing large astronomical datasets at a scale required by the Square Kilometre Array Phase 1 (SKA1). It includes an interface for expressing complex data reduction pipelines consisting of both data sets and algorithmic components and an implementation run-time to execute such pipelines on distributed resources. By mapping the logical view of a pipeline to its physical realisation, DALiuGE separates the concerns of multiple stakeholders, allowing them to collectively optimise large-scale data processing solutions in a coherent manner. The execution in DALiuGE is data-activated, where each individual data item autonomously triggers the processing on itself. Such decentralisation also makes the execution framework very scalable and flexible, supporting pipeline sizes ranging from less than ten tasks running on a laptop to tens of millions of concurrent tasks on the second fastest supercomputer in the world. DALiuGE has been used in production for reducing interferometry data sets from the Karl E. Jansky Very Large Array and the Mingantu Ultrawide Spectral Radioheliograph; and is being developed as the execution framework prototype for the Science Data Processor (SDP) consortium of the Square Kilometre Array (SKA) telescope. This paper presents a technical overview of DALiuGE and discusses case studies from the CHILES and MUSER projects that use DALiuGE to execute production pipelines. In a companion paper, we provide in-depth analysis of DALiuGE's scalability to very large numbers of tasks on two supercomputing facilities.

cs.DC

Operations in the era of large distributed telescopes

The previous generation of astronomical instruments tended to consist of single receivers in the focal point of one or more physical reflectors. Because of this, most astronomical data sets were small enough that the raw data could easily be downloaded and processed on a single machine. In the last decade, several large, complex Radio Astronomy instruments have been built and the SKA is currently being designed. Many of these instruments have been designed by international teams, and, in the case of LOFAR span an area larger than a single country. Such systems are ICT telescopes and consist mainly of complex software. This causes the main operational issues to be related to the ICT systems and not the telescope hardware. However, it is important that the operations of the ICT systems are coordinated with the traditional operational work. Managing the operations of such telescopes therefore requires an approach that significantly differs from classical telescope operations. The goal of this session is to bring together members of operational teams responsible for such large-scale ICT telescopes. This gathering will be used to exchange experiences and knowledge between those teams. Also, we consider such a meeting as very valuable input for future instrumentation, especially the SKA and its regional centres.

astro-ph.IM

Galaxy And Mass Assembly (GAMA): Panchromatic Data Release (far-UV --- far-IR) and the low-z energy budget

We present the GAMA Panchromatic Data Release (PDR) constituting over 230deg$^2$ of imaging with photometry in 21 bands extending from the far-UV to the far-IR. These data complement our spectroscopic campaign of over 300k galaxies, and are compiled from observations with a variety of facilities including: GALEX, SDSS, VISTA, WISE, and Herschel, with the GAMA regions currently being surveyed by VST and scheduled for observations by ASKAP. These data are processed to a common astrometric solution, from which photometry is derived for 221,373 galaxies with r<19.8 mag. Online tools are provided to access and download data cutouts, or the full mosaics of the GAMA regions in each band. We focus, in particular, on the reduction and analysis of the VISTA VIKING data, and compare to earlier datasets (i.e., 2MASS and UKIDSS) before combining the data and examining its integrity. Having derived the 21-band photometric catalogue we proceed to fit the data using the energy balance code MAGPHYS. These measurements are then used to obtain the first fully empirical measurement of the 0.1-500$μ$m energy output of the Universe. Exploring the Cosmic Spectral Energy Distribution (CSED) across three time-intervals (0.3-1.1Gyr, 1.1-1.8~Gyr and 1.8---2.4~Gyr), we find that the Universe is currently generating $(1.5 \pm 0.3) \times 10^{35}$ h$_{70}$ W Mpc$^{-3}$, down from $(2.5 \pm 0.2) \times 10^{35}$ h$_{70}$ W Mpc$^{-3}$ 2.3~Gyr ago. More importantly, we identify significant and smooth evolution in the integrated photon escape fraction at all wavelengths, with the UV escape fraction increasing from 27(18)% at z=0.18 in NUV(FUV) to 34(23)% at z=0.06. The GAMA PDR will allow for detailed studies of the energy production and outputs of individual systems, sub-populations, and representative galaxy samples at $z<0.5$. The GAMA PDR can be found at: http://gama-psi.icrar.org/

astro-ph.GA

Imaging SKA-Scale data in three different computing environments

We present the results of our investigations into options for the computing platform for the imaging pipeline in the CHILES project, an ultra-deep HI pathfinder for the era of the Square Kilometre Array. CHILES pushes the current computing infrastructure to its limits and understanding how to deliver the images from this project is clarifying the Science Data Processing requirements for the SKA. We have tested three platforms: a moderately sized cluster, a massive High Performance Computing (HPC) system, and the Amazon Web Services (AWS) cloud computing platform. We have used well-established tools for data reduction and performance measurement to investigate the behaviour of these platforms for the complicated access patterns of real-life Radio Astronomy data reduction. All of these platforms have strengths and weaknesses and the system tools allow us to identify and evaluate them in a quantitative manner. With the insights from these tests we are able to complete the imaging pipeline processing on both the HPC platform and also on the cloud computing platform, which paves the way for meeting big data challenges in the era of SKA in the field of Radio Astronomy. We discuss the implications that all similar projects will have to consider, in both performance and costs, to make recommendations for the planning of Radio Astronomy imaging workflows.

astro-ph.IM

A BOINC based, citizen-science project for pixel Spectral Energy Distribution fitting of resolved galaxies in multi-wavelength surveys

In this work we present our experience from the first year of theSkyNet Pan-STARRS1 Optical Galaxy Survey (POGS) project. This citizen-scientist driven research project uses the Berkeley Open Infrastructure for Network Computing (BOINC) middleware and thousands of Internet-connected computers to measure the resolved galactic structural properties of ~100,000 low redshift galaxies. We are combining the spectral coverage of GALEX, Pan-STARRS1, SDSS, and WISE to generate a value-added, multi-wavelength UV-optical-NIR galaxy atlas for the nearby Universe. Specifically, we are measuring physical parameters (such as local stellar mass, star formation rate, and first-order star formation history) on a resolved pixel-by-pixel basis using spectral energy distribution (SED) fitting techniques in a distributed computing mode.

astro-ph.IM

SkuareView: Client-Server Framework for Accessing Extremely Large Radio Astronomy Image Data

The new wide-field radio telescopes, such as: ASKAP, MWA, and SKA; will produce spectral-imaging data-cubes (SIDC) of unprecedented volume. This requires new approaches to managing and servicing the data to the end-user. We present a new integrated framework based on the JPEG2000/ISO/IEC 15444 standard to address the challenges of working with extremely large SIDC. We also present the developed j2k software, that converts and encodes FITS image cubes into JPEG2000 images, paving the way to implementing the pre- sented framework.

astro-ph.IM