SearcharxivSearch

arXiv subjects

Georgios Arampatzis

Publications and source records attributed to Georgios Arampatzis.

16 recordsLinked to original sources

Multilingual Humour-Aware Retrieval with Dense and Re-Ranking Models

Humour-aware information retrieval poses unique challenges beyond standard semantic retrieval, as systems must account not only for topical relevance but also for humour-specific linguistic phenomena such as wordplay, phonetic ambiguity, and polysemy. In this paper, Team DUTH studies multilingual humour-aware information retrieval using the CLEF 2025 JOKER Task 1 benchmark, which evaluates humour retrieval in English and Portuguese. Our approach combines multilingual XLM-RoBERTa-based dense retrieval with additional system variants, including neural re-ranking, in order to assess the extent to which general-purpose Transformer models can capture humour-specific relevance. The results reveal substantial cross-lingual variation. While the Portuguese runs demonstrate comparatively strong performance across MAP, MRR, and early precision metrics, the English runs perform significantly worse, with relevant humorous documents frequently appearing at lower ranks. These findings highlight the limitations of purely semantic dense representations for humour retrieval, particularly when humour depends on surface-level cues that are not explicitly modelled by multilingual encoders. We further analyse contributing factors to this discrepancy, including dataset characteristics, query-document alignment, and variation in humour mechanisms. Overall, the Team DUTH experiments establish multilingual dense-retrieval and re-ranking baselines and provide insights into the challenges of modelling humour-aware relevance within the JOKER framework.

cs.IR

Neural Measures for learning distributions of Random PDEs

The integration of Scientific Machine Learning (SciML) techniques with uncertainty quantification (UQ) represents a rapidly evolving frontier in computational science. This work advances Physics-Informed Neural Networks (PINNs) by incorporating probabilistic frameworks to effectively model uncertainty in complex systems. Our approach enhances the representation of uncertainty in forward problems by combining generative modeling techniques with PINNs. This integration enables in a systematic fashion uncertainty control while maintaining the predictive accuracy of the model. We demonstrate the utility of this method through applications to random differential equations and random partial differential equations (PDEs).

stat.ML

Adaptive learning of effective dynamics: Adaptive real-time, online modeling for complex systems

Predictive simulations are essential for applications ranging from weather forecasting to material design. The veracity of these simulations hinges on their capacity to capture the effective system dynamics. Massively parallel simulations predict the systems dynamics by resolving all spatiotemporal scales, often at a cost that prevents experimentation. On the other hand, reduced order models are fast but often limited by the linearization of the system dynamics and the adopted heuristic closures. We propose a novel systematic framework that bridges large scale simulations and reduced order models to extract and forecast adaptively the effective dynamics (AdaLED) of multiscale systems. AdaLED employs an autoencoder to identify reduced-order representations of the system dynamics and an ensemble of probabilistic recurrent neural networks (RNNs) as the latent time-stepper. The framework alternates between the computational solver and the surrogate, accelerating learned dynamics while leaving yet-to-be-learned dynamics regimes to the original solver. AdaLED continuously adapts the surrogate to the new dynamics through online training. The transitions between the surrogate and the computational solver are determined by monitoring the prediction accuracy and uncertainty of the surrogate. The effectiveness of AdaLED is demonstrated on three different systems - a Van der Pol oscillator, a 2D reaction-diffusion equation, and a 2D Navier-Stokes flow past a cylinder for varying Reynolds numbers (400 up to 1200), showcasing its ability to learn effective dynamics online, detect unseen dynamics regimes, and provide net speed-ups. To the best of our knowledge, AdaLED is the first framework that couples a surrogate model with a computational solver to achieve online adaptive learning of effective dynamics. It constitutes a potent tool for applications requiring many expensive simulations.

physics.comp-ph

The stress-free state of human erythrocytes: data driven inference of a transferable RBC model

The stress-free state (SFS) of red blood cells (RBCs) is a fundamental reference configuration for the calibration of computational models, yet it remains unknown. Current experimental methods cannot measure the SFS of cells without affecting their mechanical properties while computational postulates are the subject of controversial discussions. Here, we introduce data driven estimates of the SFS shape and the visco-elastic properties of RBCs. We employ data from single-cell experiments that include measurements of the equilibrium shape, of stretched cells, and relaxation times of initially stretched RBCs. A hierarchical Bayesian model accounts for these experimental and data heterogeneities. We quantify, for the first time, the SFS of RBCs and use it to introduce a transferable RBC (t-RBC) model. The effectiveness of the proposed model is shown on predictions of unseen experimental conditions during the inference, including the critical stress of transitions between tumbling and tank-treading cells in shear flow. Our findings demonstrate that the proposed t-RBC model provides predictions of blood flows with unprecedented accuracy and quantified uncertainties.

cs.CE

Multiscale Simulations of Complex Systems by Learning their Effective Dynamics

Predictive simulations of complex systems are essential for applications ranging from weather forecasting to drug design. The veracity of these predictions hinges on their capacity to capture the effective system dynamics. Massively parallel simulations predict the system dynamics by resolving all spatiotemporal scales, often at a cost that prevents experimentation while their findings may not allow for generalisation. On the other hand reduced order models are fast but limited by the frequently adopted linearization of the system dynamics and/or the utilization of heuristic closures. Here we present a novel systematic framework that bridges large scale simulations and reduced order models to Learn the Effective Dynamics (LED) of diverse complex systems. The framework forms algorithmic alloys between non-linear machine learning algorithms and the Equation-Free approach for modeling complex systems. LED deploys autoencoders to formulate a mapping between fine and coarse-grained representations and evolves the latent space dynamics using recurrent neural networks. The algorithm is validated on benchmark problems and we find that it outperforms state of the art reduced order models in terms of predictability and large scale simulations in terms of cost. LED is applicable to systems ranging from chemistry to fluid mechanics and reduces the computational effort by up to two orders of magnitude while maintaining the prediction accuracy of the full system dynamics. We argue that LED provides a novel potent modality for the accurate prediction of complex systems.

physics.comp-ph

Korali: Efficient and Scalable Software Framework for Bayesian Uncertainty Quantification and Stochastic Optimization

We present Korali, an open-source framework for large-scale Bayesian uncertainty quantification and stochastic optimization. The framework relies on non-intrusive sampling of complex multiphysics models and enables their exploitation for optimization and decision-making. In addition, its distributed sampling engine makes efficient use of massively-parallel architectures while introducing novel fault tolerance and load balancing mechanisms. We demonstrate these features by interfacing Korali with existing high-performance software such as Aphros, Lammps (CPU-based), and Mirheo (GPU-based) and show efficient scaling for up to 512 nodes of the CSCS Piz Daint supercomputer. Finally, we present benchmarks demonstrating that Korali outperforms related state-of-the-art software frameworks.

cs.DC

Optimal sensing for fish school identification

Fish schooling implies an awareness of the swimmers for their companions. In flow mediated environments, in addition to visual cues, pressure and shear sensors on the fish body are critical for providing quantitative information that assists the quantification of proximity to other swimmers. Here we examine the distribution of sensors on the surface of an artificial swimmer so that it can optimally identify a leading group of swimmers. We employ Bayesian experimental design coupled with two-dimensional Navier Stokes equations for multiple self-propelled swimmers. The follower tracks the school using information from its own surface pressure and shear stress. We demonstrate that the optimal sensor distribution of the follower is qualitatively similar to the distribution of neuromasts on fish. Our results show that it is possible to identify accurately the center of mass and even the number of the leading swimmers using surface only information.

physics.flu-dyn

Optimal sensor placement for artificial swimmers

Natural swimmers rely for their survival on sensors that gather information from the environment and guide their actions. The spatial organization of these sensors, such as the visual fish system and lateral line, suggests evolutionary selection, but their optimality remains an open question. Here, we identify sensor configurations that enable swimmers to maximize the information gathered from their surrounding flow field. We examine two-dimensional, self-propelled and stationary swimmers that are exposed to disturbances generated by oscillating, rotating and D-shaped cylinders. We combine simulations of the Navier-Stokes equations with Bayesian experimental design to determine the optimal arrangements of shear and pressure sensors that best identify the locations of the disturbance-generating sources. We find a marked tendency for shear stress sensors to be located in the head and the tail of the swimmer, while they are absent from the midsection. In turn, we find a high density of pressure sensors in the head along with a uniform distribution along the entire body. The resulting optimal sensor arrangements resemble neuromast distributions observed in fish and provide evidence for optimality in sensor distribution for natural swimmers.

physics.flu-dyn

Bayesian selection for coarse-grained models of liquid water

The necessity for accurate and computationally efficient representations of water in atomistic simulations that can span biologically relevant timescales has born the necessity of coarse-grained (CG) modeling. Despite numerous advances, CG water models rely mostly on a-priori specified assumptions. How these assumptions affect the model accuracy, efficiency, and in particular transferability, has not been systematically investigated. Here we propose a data driven, comparison and selection for CG water models through a Hierarchical Bayesian framework. We examine CG water models that differ in their level of coarse-graining, structure, and number of interaction sites. We find that the importance of electrostatic interactions for the physical system under consideration is a dominant criterion for the model selection. Multi-site models are favored, unless the effects of water in electrostatic screening are not relevant, in which case the single site model is preferred due to its computational savings. The charge distribution is found to play an important role in the multi-site model's accuracy while the flexibility of the bonds/angles may only slightly improve the models. Furthermore, we find significant variations in the computational cost of these models. We present a data informed rationale for the selection of CG water models and provide guidance for future water model designs.

physics.comp-ph

$S$-Leaping: An adaptive, accelerated stochastic simulation algorithm, bridging $τ$-leaping and $R$-leaping

We propose the $S$-leaping algorithm for the acceleration of Gillespie's stochastic simulation algorithm that combines the advantages of the two main accelerated methods; the $τ$-leaping and $R$-leaping algorithms. These algorithms are known to be efficient under different conditions; the $τ$-leaping is efficient for non-stiff systems or systems with partial equilibrium, while the $R$-leaping performs better in stiff system thanks to an efficient sampling procedure. However, even a small change in a system's set up can critically affect the nature of the simulated system and thus reduce the efficiency of an accelerated algorithm. The proposed algorithm combines the efficient time step selection from the $τ$-leaping with the effective sampling procedure from the $R$-leaping algorithm. The $S$-leaping is shown to maintain its efficiency under different conditions and in the case of large and stiff systems or systems with fast dynamics, the $S$-leaping outperforms both methods. We demonstrate the performance and the accuracy of the $S$-leaping in comparison with the $τ$-leaping and $R$-leaping on a number of benchmark systems involving biological reaction networks.

math.PR

Langevin Diffusion for Population Based Sampling with an Application in Bayesian Inference for Pharmacodynamics

We propose an algorithm for the efficient and robust sampling of the posterior probability distribution in Bayesian inference problems. The algorithm combines the local search capabilities of the Manifold Metropolis Adjusted Langevin transition kernels with the advantages of global exploration by a population based sampling algorithm, the Transitional Markov Chain Monte Carlo (TMCMC). The Langevin diffusion process is determined by either the Hessian or the Fisher Information of the target distribution with appropriate modifications for non positive definiteness. The present methods is shown to be superior over other population based algorithms, in sampling probability distributions for which gradients are available and is shown to handle otherwise unidentifiable models. We demonstrate the capabilities and advantages of the method in computing the posterior distribution of the parameters in a Pharmacodynamics model, for glioma growth and its drug induced inhibition, using clinical data.

stat.CO

Experimental data over quantum mechanics simulations for inferring the repulsive exponent of the Lennard-Jones potential in Molecular Dynamics

The Lennard-Jones (LJ) potential is a cornerstone of Molecular Dynamics (MD) simulations and among the most widely used computational kernels in science. The potential models atomistic attraction and repulsion with century old prescribed parameters ($q=6, \; p=12$, respectively), originally related by a factor of two for simplicity of calculations. We re-examine the value of the repulsion exponent through data driven uncertainty quantification. We perform Hierarchical Bayesian inference on MD simulations of argon using experimental data of the radial distribution function (RDF) for a range of thermodynamic conditions, as well as dimer interaction energies from quantum mechanics simulations. The experimental data suggest a repulsion exponent ($p \approx 6.5$), in contrast to the quantum simulations data that support values closer to the original ($p=12$) exponent. Most notably, we find that predictions of RDF, diffusion coefficient and density of argon are more accurate and robust in producing the correct argon phase around its triple point, when using the values inferred from experimental data over those from quantum mechanics simulations. The present results suggest the need for data driven recalibration of the LJ potential across MD simulations.

physics.chem-ph

Efficient estimators for likelihood ratio sensitivity indices of complex stochastic dynamics

We demonstrate that centered likelihood ratio estimators for the sensitivity indices of complex stochastic dynamics are highly efficient with low, constant in time variance and consequently they are suitable for sensitivity analysis in long-time and steady-state regimes. These estimators rely on a new covariance formulation of the likelihood ratio that includes as a submatrix a Fisher Information Matrix for stochastic dynamics and can also be used for fast screening of insensitive parameters and parameter combinations. The proposed methods are applicable to broad classes of stochastic dynamics such as chemical reaction networks, Langevin-type equations and stochastic models in finance, including systems with a high dimensional parameter space and/or disparate decorrelation times between different observables. Furthermore, they are simple to implement as a standard observable in any existing simulation algorithms without additional modifications.

math.NA

Pathwise Sensitivity Analysis in Transient Regimes

The instantaneous relative entropy (IRE) and the corresponding instanta- neous Fisher information matrix (IFIM) for transient stochastic processes are pre- sented in this paper. These novel tools for sensitivity analysis of stochastic models serve as an extension of the well known relative entropy rate (RER) and the corre- sponding Fisher information matrix (FIM) that apply to stationary processes. Three cases are studied here, discrete-time Markov chains, continuous-time Markov chains and stochastic differential equations. A biological reaction network is presented as a demonstration numerical example.

math.PR

Accelerated Sensitivity Analysis in High-Dimensional Stochastic Reaction Networks

In this paper, a two-step strategy for parametric sensitivity analysis for such systems is proposed, exploiting advantages and synergies between two recently proposed sensitivity analysis methodologies for stochastic dynamics. The first method performs sensitivity analysis of the stochastic dynamics by means of the Fisher Information Matrix on the underlying distribution of the trajectories; the second method is a reduced-variance, finite-difference, gradient-type sensitivity approach relying on stochastic coupling techniques for variance reduction. Here we demonstrate that these two methods can be combined and deployed together by means of a new sensitivity bound which incorporates the variance of the quantity of interest as well as the Fisher Information Matrix estimated from the first method. The first step of the proposed strategy labels sensitivities using the bound and screens out the insensitive parameters in a controlled manner based also on the new sensitivity bound. In the second step of the proposed strategy, the finite-difference method is applied only for the sensitivity estimation of the (potentially) sensitive parameters that have not been screened out in the first step. Results on an epidermal growth factor network with fifty parameters and on a protein homeostasis with eighty parameters demonstrate that the proposed strategy is able to quickly discover and discard the insensitive parameters and in the remaining potentially sensitive parameters it accurately estimates the sensitivities. The new sensitivity strategy can be several times faster than current state-of-the-art approaches that test all parameters, especially in "sloppy" systems. In particular, the computational acceleration is quantified by the ratio between the total number of parameters over the number of the sensitive parameters.

q-bio.MN

Goal-oriented sensitivity analysis for lattice kinetic Monte Carlo simulations

In this paper we propose a new class of coupling methods for the sensitivity analysis of high dimensional stochastic systems and in particular for lattice Kinetic Monte Carlo. Sensitivity analysis for stochastic systems is typically based on approximating continuous derivatives with respect to model parameters by the mean value of samples from a finite difference scheme. Instead of using independent samples the proposed algorithm reduces the variance of the estimator by developing a strongly correlated-"coupled"- stochastic process for both the perturbed and unperturbed stochastic processes, defined in a common state space. The novelty of our construction is that the new coupled process depends on the targeted observables, e.g. coverage, Hamiltonian, spatial correlations, surface roughness, etc., hence we refer to the proposed method as em goal-oriented sensitivity analysis. In particular, the rates of the coupled Continuous Time Markov Chain are obtained as solutions to a goal-oriented optimization problem, depending on the observable of interest, by considering the minimization functional of the corresponding variance. We show that this functional can be used as a diagnostic tool for the design and evaluation of different classes of couplings. Furthermore the resulting KMC sensitivity algorithm has an easy implementation that is based on the Bortz-Kalos-Lebowitz algorithm's philosophy, where here events are divided in classes depending on level sets of the observable of interest. Finally, we demonstrate in several examples including adsorption, desorption and diffusion Kinetic Monte Carlo that for the same confidence interval and observable, the proposed goal-oriented algorithm can be two orders of magnitude faster than existing coupling algorithms for spatial KMC such as the Common Random Number approach.

math.NA