SearcharxivSearch

arXiv subjects

Duarte Santos

Publications and source records attributed to Duarte Santos.

2 recordsLinked to original sources

The 2026 Algorithmic Information Theory Data Compression Challenge

Lossless data compression remains central to computer science, with direct impact on storage, communication bandwidth, computational cost, and energy consumption. It is also closely related to Algorithmic Information Theory, where compressibility provides an operational measure of structure and non-randomness. This paper presents the 2026 Algorithmic Information Theory Data Compression Challenge, a benchmark for evaluating general-purpose lossless compressors under realistic constraints. Submissions were encouraged to use arithmetic or range coding, limited to at most 8 GB of memory, and required to include a decompressor no larger than 1 MB. The benchmark comprised sixteen heterogeneous files, split into public training and hidden testing datasets. In total, 117 valid submitted compressors were evaluated alongside established reference compressors using compression ratio, compression and decompression time, Weissman score, and Pareto-frontier analysis. The results show that performance depends strongly on the optimization criterion: fast compressors achieved the best speed-oriented scores, whereas modelling-intensive compressors produced smaller outputs at higher computational cost. A Normalized Compression Distance analysis further revealed clusters of related submissions and distinguished incremental variants from more independent implementations. Selected submissions were described for their methodological novelty or competitive performance and further tested on four large external datasets, where several achieved competitive or superior results relative to established compressors. Overall, the challenge confirms the importance of probabilistic modelling, hidden testing, and external datasets for assessing compression performance and generalization. Benchmark resources, leaderboard data, binaries, and selected source code are publicly available at https://aitdcc.github.io.

cs.IT

An investigation of the star-forming main sequence considering the nebular continuum emission at low-z

The code FADO is the first publicly available population spectral synthesis tool that treats the contribution from ionised gas to the observed emission self-consistently. We study the impact of the nebular contribution on the determination of the star formation rate (SFR), stellar mass, and consequent effect on the star-forming main sequence (SFMS) at low redshift. We applied FADO to the spectral database of the SDSS to derive the physical properties of galaxies. As a comparison, we used the data in the MPA-JHU catalogue, which contains the properties of SDSS galaxies derived without the nebular contribution. We selected a sample of SF galaxies with H$\alpha$ and H$\beta$ flux measurements, and we corrected the fluxes for the nebular extinction through the Balmer decrement. We then calculated the H$\alpha$ luminosity to estimate the SFR. Then, by combining the stellar mass and SFR estimates from FADO and MPA-JHU, the SFMS was obtained. The H$\alpha$ flux estimates are similar between FADO and MPA-JHU. Because the H$\alpha$ flux was used as tracer of the SFR, FADO and MPA-JHU agree in their SFR. The stellar mass estimates are slightly higher for FADO than for MPA-JHU on average. However, considering the uncertainties, the differences are negligible. With similar SFR and stellar mass estimates, the derived SFMS is also similar between FADO and MPA-JHU. Our results show that for SDSS normal SF galaxies, the additional modelling of the nebular contribution does not affect the retrieved fluxes and consequentially also does not influence SFR estimators based on the extinction-corrected H$\alpha$ luminosity. For the stellar masses, the results point to the same conclusion. These results are a consequence of the fact that the vast majority of normal SF galaxies in the SDSS have a low nebular contribution.

astro-ph.GA