SearcharxivSearch

arXiv subjects

Kornel Howil

Publications and source records attributed to Kornel Howil.

8 recordsLinked to original sources

TOM-GS: Editable Video Representation via Temporal Opacity Modulation of Static 3D Gaussians

While Implicit Neural Representations (INRs) and dynamic 3D Gaussian Splatting (3DGS) achieve impressive results in video processing, they often fall short of producing representations that are easily editable. Recent methods address this by introducing complex spatial deformations or folded distributions, which constrain optimization and reduce flexibility for downstream editing. In this paper, we introduce TOM-GS, an editable video representation that forgoes complex deformations in favor of regular 3D Gaussians equipped with a continuous temporal opacity formulation. By assigning a learnable temporal mean and scale to the opacity of each Gaussian, our model enables static 3D spatial components to fade smoothly in and out of the scene. Grounded by robust, off-the-shelf pose estimation, our approach maintains a static spatial geometry that naturally supports a wide range of manual and physics-based edits. TOM-GS outperforms prior editable video representations in visual fidelity, while its reliance on standard 3D Gaussians ensures seamless compatibility with established 3D editing tools.

cs.CV

OmniStyle-INR: Universal and Multimodal Style Transfer for INRs

Style transfer remains a fundamental and highly important task across various data modalities, enabling creative manipulation conditioned by both reference images and textual descriptions. Recently, methods utilizing Gaussian Splatting have emerged as a unified representation for 2D images, video, 3D scenes, and 4D dynamics. However, representing videos and 2D images with Gaussian Splatting is structurally sub-optimal for dense continuous domains. The number of required Gaussians often approaches the total number of pixels, raising questions about the actual utility of such a representation for these specific modalities. In contrast, Implicit Neural Representations have established themselves as a much more popular and natural choice across all these data domains. Implicit Neural Representations naturally provide significant advantages, including data compression, inherent capabilities for super resolution, and seamless integration with deep generative models. To this end, we introduce OmniStyle-INR, a novel framework that leverages network-based continuous representations as a truly universal domain. Our approach successfully performs high-quality style transfer across all visual modalities, guided seamlessly by both text prompts and visual exemplars.

cs.CV

APEX: Audio Prototype EXplanations for Classification Tasks

Explainable AI (XAI) has achieved remarkable success in image classification, yet the audio domain lacks equally mature solutions. Current methods apply vision-based attribution techniques to spectrograms, overlooking fundamental differences between visual and acoustic signals. While prototype reasoning is promising, acoustic similarity remains multidimensional. We introduce APEX (Audio Prototype EXplanations), a post-hoc framework for interpreting pre-trained audio classifiers. Crucially, APEX requires no fine-tuning of the original backbone and strictly preserves output invariance. APEX disentangles explanations into four perspectives: Square-based prototypes to localize transient events, Time-based for temporal patterns, Frequency-based highlighting spectral bands, and Time-Frequency-based integrating both. This yields intuitive, example-based explanations that respect acoustic properties, providing greater semantic clarity than standard gradient-based methods.

cs.SD

CLIPGaussian: Universal and Multimodal Style Transfer Based on Gaussian Splatting

Gaussian Splatting (GS) has recently emerged as an efficient representation for rendering 3D scenes from 2D images and has been extended to images, videos, and dynamic 4D content. However, applying style transfer to GS-based representations, especially beyond simple color changes, remains challenging. In this work, we introduce CLIPGaussian, the first unified style transfer framework that supports text- and image-guided stylization across multiple modalities: 2D images, videos, 3D objects, and 4D scenes. Our method operates directly on Gaussian primitives and integrates into existing GS pipelines as a plug-in module, without requiring large generative models or retraining from scratch. The CLIPGaussian approach enables joint optimization of color and geometry in 3D and 4D settings, and achieves temporal coherence in videos, while preserving the model size. We demonstrate superior style fidelity and consistency across all tasks, validating CLIPGaussian as a universal and efficient solution for multimodal style transfer.

cs.CV

VeGaS: Video Gaussian Splatting

Implicit Neural Representations (INRs) employ neural networks to approximate discrete data as continuous functions. In the context of video data, such models can be utilized to transform the coordinates of pixel locations along with frame occurrence times (or indices) into RGB color values. Although INRs facilitate effective compression, they are unsuitable for editing purposes. One potential solution is to use a 3D Gaussian Splatting (3DGS) based model, such as the Video Gaussian Representation (VGR), which is capable of encoding video as a multitude of 3D Gaussians and is applicable for numerous video processing operations, including editing. Nevertheless, in this case, the capacity for modification is constrained to a limited set of basic transformations. To address this issue, we introduce the Video Gaussian Splatting (VeGaS) model, which enables realistic modifications of video data. To construct VeGaS, we propose a novel family of Folded-Gaussian distributions designed to capture nonlinear dynamics in a video stream and model consecutive frames by 2D Gaussians obtained as respective conditional distributions. Our experiments demonstrate that VeGaS outperforms state-of-the-art solutions in frame reconstruction tasks and allows realistic modifications of video data. The code is available at: https://github.com/gmum/VeGaS.

cs.CV

Analysis of the full Spitzer microlensing sample I: Dark remnant candidates and Gaia predictions

In the pursuit of understanding the population of stellar remnants within the Milky Way, we analyze the sample of $\sim 950$ microlensing events observed by the Spitzer Space Telescope between 2014 and 2019. In this study we focus on a sub-sample of nine microlensing events, selected based on their long timescales, small microlensing parallaxes and joint observations by the Gaia mission, to increase the probability that the chosen lenses are massive and the mass is measurable. Among the selected events we identify lensing black holes and neutron star candidates, with potential confirmation through forthcoming release of the Gaia time-series astrometry in 2026. Utilizing Bayesian analysis and Galactic models, along with the Gaia Data Release 3 proper motion data, four good candidates for dark remnants were identified: OGLE-2016-BLG-0293, OGLE-2018-BLG-0483, OGLE-2018-BLG-0662, and OGLE-2015-BLG-0149, with lens masses of $2.98^{+1.75}_{-1.28}~M_{\odot}$, $4.65^{+3.12}_{-2.08}~M_{\odot}$, $3.15^{+0.66}_{-0.64}~M_{\odot}$ and $1.4^{+0.75}_{-0.55}~M_{\odot}$, respectively. Notably, the first two candidates are expected to exhibit astrometric microlensing signals detectable by Gaia, offering the prospect of validating the lens masses. The methodologies developed in this work will be applied to the full Spitzer microlensing sample, populating and analyzing the time-scale ($t_{\rm E}$) vs. parallax ($\pi_{\rm E}$) diagram to derive constraints on the population of lenses in general and massive remnants in particular.

astro-ph.GA

Dark lenses through the dust: parallax microlensing events in the VVV

We use near-infrared photometry and astrometry from the VISTA Variables in the Via Lactea (VVV) survey to analyse microlensing events containing annual microlensing parallax information. These events are located in highly extincted and low-latitude regions of the Galactic bulge typically off-limits to optical microlensing surveys. We fit a catalog of $1959$ events previously found in the VVV and extract $21$ microlensing parallax candidates. The fitting is done using nested sampling to automatically characterise the multi-modal and degenerate posterior distributions of the annual microlensing parallax signal. We compute the probability density in lens mass-distance using the source proper motion and a Galactic model of disc and bulge deflectors. By comparing the expected flux from a main sequence lens to the baseline magnitude and blending parameter, we identify 4 candidates which have probability $> 50$% that the lens is dark. The strongest candidate corresponds to a nearby ($\approx0.78$ kpc), medium-mass ($1.46^{+1.13}_{-0.71} \ M_{\odot}$) dark remnant as lens. In the next strongest, the lens is located at heliocentric distance $\approx5.3$ kpc. It is a dark remnant with a mass of $1.63^{+1.15}_{-0.70} \ M_{\odot}$. Both of those candidates are most likely neutron stars, though possibly high-mass white dwarfs. The last two events may also be caused by dark remnants, though we are unable to rule out other possibilities because of limitations in the data.

astro-ph.GA

Using the Carnot cycle to determine changes of the phase transition temperature

The Clausius-Clapeyron relation and its analogs in other first-order phase transitions, such as type-I superconductors, are derived using very elementary methods, without appealing to the more advanced concepts of entropy or Gibbs free energy. The reasoning is based on Kelvin's formulation of the second law of thermodynamics, and should be accessible to high school students. After recalling some basic facts about the Carnot cycle, we present two very different systems that undergo discontinuous phase transitions (ice/water and normal/superconductor), and construct engines that exploit the properties of these systems to produce work. In each case, we show that if the transition temperature $T_tr$ were independent of other parameters, such as pressure or magnetic field, it would be possible to violate Kelvin's principle, i.e., to construct a perpetuum mobile of the second kind. Since the proposed cyclic processes can be realized reversibly in the limit of infinitesimal changes in temperature, their efficiencies must be equal to that of an ordinary Carnot cycle. We immediately obtain an equation of the form $dT /dX = f(T, X)$, which governs how the transition temperature changes with the parameter $X$.

physics.pop-ph