SearcharxivSearch

arXiv subjects

George Papadakis

Publications and source records attributed to George Papadakis.

At least 19 recordsLinked to original sources

AgenticER: the next frontier in Entity Resolution

Entity Resolution (ER) is a fundamental problem in data management, playing a critical role in tasks like data cleaning and knowledge graph construction. The existing ER approaches range from traditional rule-based to deep learning techniques and LLM-based methods, but typically operate under a ``passive paradigm'', as duplicates are detected through static, one-shot similarity computations. Such approaches fail to capture the inherently uncertain and context-dependent nature of real-world ER tasks, especially in data lakes with streaming content in heterogeneous formats such as CSV files, JSON files, RDF dumps, and free text. In such settings, resolving ambiguity often requires iterative evidence gathering, reasoning across multiple sources, even selective human involvement. To cover this gap, we advocate a paradigm shift from passive to Agentic ER, which frames ER as a sequential decision-making process that is performed by autonomous agents. These agents actively plan ER strategies, acquire external evidence, decide when to query additional sources or humans, and optimize trade-offs between accuracy, cost, and latency. We formalize Agentic ER as a decision-theoretic problem, we propose a reference architecture, we identify core research challenges, and outline new evaluation dimensions tailored to agentic behavior. By introducing Agentic ER, we aim to establish a new research direction at the intersection of data management and intelligent agents.

cs.DB

Numerical Investigation of Elastically-Mounted tandem Cylinders using an ALE Runge-Kutta Discontinuous Galerkin method

This work presents a high-order Arbitrary-Lagrangian-Eulerian (ALE) Discontinuous Galerkin framework for simulating multi-body Vortex-Induced Vibrations. The ALE formulation extends a Runge-Kutta Interior-Penalty nodal DG solver with minimal additional computational overhead, incorporating discrete enforcement of the Geometric Conservation Law (GCL) to ensure free-stream preservation and Radial Basis Function (RBF) mesh deformation to handle large structural displacements. The framework is applied to elastically-mounted tandem cylinder configurations: a two-cylinder arrangement with cross-flow oscillations at Re=200, and a three-cylinder arrangement with two degrees of freedom at Re=150. In the three-cylinder case, the trajectories exhibit highly irregular behavior driven by complex wake interference, including a periodic attract-and-release mechanism governing the trailing cylinder's stream-wise response. Results are verified against established benchmarks through Lissajous curves, Poincar\'{e} phase maps, power spectra, and vortex shedding mode classification. An hp-refinement comparison demonstrates that increasing the polynomial order is more effective and computationally efficient than mesh refinement for capturing multi-body wake dynamics, as the low numerical diffusion of the high-order method preserves vortical structures over long distances on relatively coarse meshes. These findings highlight the importance of high-order methods for CFD-FSI applications where wake interactions drive the structural response.

physics.flu-dyn

DistillER: Knowledge Distillation in Entity Resolution with Large Language Models

Recent advances in Entity Resolution (ER) have leveraged Large Language Models (LLMs), achieving strong performance but at the cost of substantial computational resources or high financial overhead. Existing LLM-based ER approaches operate either in unsupervised settings and rely on very large and costly models, or in supervised settings and require ground-truth annotations, leaving a critical gap between time efficiency and effectiveness. To make LLM-powered ER more practical, we investigate Knowledge Distillation (KD) as a means to transfer knowledge from large, effective models (Teachers) to smaller, more efficient models (Students) without requiring gold labels. We introduce DistillER, the first framework that systematically bridges this gap across three dimensions: (i) Data Selection, where we study strategies for identifying informative subsets of data; (ii) Knowledge Elicitation, where we compare single- and multi-teacher settings across LLMs and smaller language models (SLMs); and (iii) Distillation Algorithms, where we evaluate supervised fine-tuning and reinforcement learning approaches. Our experiments reveal that supervised fine-tuning of Students on noisy labels generated by LLM Teachers consistently outperforms alternative KD strategies, while also enabling high-quality explanation generation. Finally, we benchmark DistillER against established supervised and unsupervised ER methods based on LLMs and SLMs, demonstrating significant improvements in both effectiveness and efficiency.

cs.DB

SPER: Accelerating Progressive Entity Resolution via Stochastic Bipartite Maximization

Entity Resolution (ER) is a critical data cleaning task for identifying records that refer to the same real-world entity. In the era of Big Data, traditional batch ER is often infeasible due to volume and velocity constraints, necessitating Progressive ER methods that maximize recall within a limited computational budget. However, existing progressive approaches fail to scale to high-velocity streams because they rely on deterministic sorting to prioritize candidate pairs, a process that incurs prohibitive super-linear complexity and heavy initialization costs. To address this scalability wall, we introduce SPER (Stochastic Progressive ER), a novel framework that redefines prioritization as a sampling problem rather than a ranking problem. By replacing global sorting with a continuous stochastic bipartite maximization strategy, SPER acts as a probabilistic high-pass filter that selects high-utility pairs in strictly linear time. Extensive experiments on eight real-world datasets demonstrate that SPER achieves significant speedups (3x to >6x) over state-of-the-art baselines while maintaining comparable recall and precision.

cs.DB

Design, Testing and Numerical Modelling of a Low-Speed Wind Tunnel Gust Generator

Understanding and accurately reproducing gust-induced unsteady aerodynamics is essential for improving load prediction, aeroelastic analysis, and control strategies in aircraft, uninhabited aerial vehicles, and wind turbines, particularly in regimes where nonlinear flow phenomena dominate. In this work, a low-speed wind tunnel gust generator based on oscillating vanes is designed, manufactured, and characterised through a combined experimental and numerical investigation. The system is intended to reproduce deterministic gust profiles relevant to aircraft, uninhabited aerial vehicles, and wind-turbine applications, operating in highly unsteady aerodynamic regimes. Experimental measurements using hot-wire anemometry are performed to quantify the generated gust field under a range of free-stream velocities, amplitudes, and forcing frequencies. In parallel, time-accurate CFD simulations are conducted using a deforming-mesh approach to validate the measurements and to analyse the flow physics associated with gust formation and propagation. Particular attention is given to the negative velocity peaks inherent to classical '1-cos' gust profiles. A modified vane motion protocol is proposed and shown to significantly reduce the negative peak factor while maintaining a substantial gust ratio. Numerical results reveal that secondary flow-angle variations arise from nonlinear interactions between vortices shed by adjacent vanes.

physics.flu-dyn

Computational study of airfoil stall flutter Limit Cycle Oscillations

This paper presents a comprehensive numerical investigation of a NACA0012 undergoing Stall Flutter Limit Cycle Oscillations (LCO) across distinct fluid dynamics regimes. It accurately models Small Amplitude Oscillations (SAO) in the transitional Reynolds regime and Large Amplitude Oscillations (LAO) in the moderate regime, observed in different experimental campaigns. The SAO analysis serves as a verification of the computational framework against established numerical benchmarks. Crucially, the LAO simulations represent the first documented prediction across the full experimental velocity range correlated against available measured data, addressing a significant literature gap. The predictions fidelity relies on rigorous computational criteria defined through a detailed sensitivity analysis. This demonstrated numerical requirements significantly more demanding than those typically employed for computing static polars or simulating dynamic pitching motion of rigid airfoils, underscoring the severity of the aeroelastic problem. Quantitatively the simulation systematically over-predicts the critical onset velocity and under-predicts the LCO amplitudes.However, the results show strong qualitative agreement with experimental observations, successfully reproducing key dynamic stall mechanics and bifurcation phenomena.

physics.flu-dyn

An ALE approach to reduce spurious numerical mixing through variational minimizers: application to internal waves

Spurious numerical mixing is a frequent phenomenon in ocean models. In this paper, we present an efficient and robust methodology that defines the vertical grid motion so that this mixing is reduced. This motion is defined as the solution of an optimization problem that -- using the ideas of the calculus of variations -- results in an elliptic partial differential equation, which is straightforward to analyze and discretize. This framework is generally applicable to any ocean model that uses an Arbitrary Lagrangian-Eulerian (ALE) vertical coordinate and can be tuned to fit the modeler's specific needs based on the guidelines presented herein. The method is applied to the nonhydrostatic solver presented by the authors in [Alexandris-Galanopoulos et al., 2024]. While the majority of spurious numerical mixing studies focus on large-scale processes, herein the proposed method is applied and tested in small-scale nonhydrostatic phenomena. Specifically, the effectiveness of the method in capturing fully nonlinear internal waves is investigated for the test cases of wave propagation, breaking and overturning. Overturning serves as a demanding test for the proposed scheme as it induces rapid vertical accelerations and thus the mesh-moving algorithm must incorporate this motion with the goal of reducing numerical mixing, while not suppressing physically relevant vertical mass transfer. These numerical benchmarks show the ability of the method to reduce spurious mixing, while attaining the physical relevance of the results.

physics.comp-ph

Parallel Nodal Interior-Penalty Discontinuous Galerkin Methods for the Subsonic Compressible Navier-Stokes Equations: Applications to Vortical Flows and VIV Problems

We present a Discontinuous Galerkin (DG) solver for the compressible Navier-Stokes system, designed for applications of technological and industrial interest in the subsonic region. More precisely, this work aims to exploit the DG-discretised Navier-Stokes for two dimensional vortex-induced vibration (VIV) problems allowing for high-order of accuracy. The numerical discretisation comprises a nodal DG method on triangular grids, that includes two types of numerical fluxes: 1) the Roe approximate Riemann solver flux for non-linear advection terms, and 2) an Interior-Penalty numerical flux for non-linear diffusion terms. The nodal formulation permits the use of high order polynomial approximations without compromising computational robustness. The spatially-discrete form is integrated in time using a low-storage strong-stability preserving explicit Runge-Kutta scheme, and is coupled weakly with an implicit rigid body dynamics algorithm. The proposed algorithm successfully implements polynomial orders of $p\ge 4$ for the laminar compressible Navier-Stokes equations in massively parallel architectures. The resulting framework is firstly tested in terms of its convergence properties. Then, numerical solutions are validated with experimental and numerical data for the case of a circular cylinder at low Reynolds number, and lastly, the methodology is employed to simulate the problem of an elastically-mounted cylinder, a known configuration characterised by significant computational challenges. The above results showcase that the DG framework can be employed as an accurate and efficient, arbitrary order numerical methodology, for high fidelity fluid-structure-interaction (FSI) problems.

math.NA

Lyapunov stability analysis of the chaotic flow past two square cylinders

We investigate the stability of the flow past two side-by-side square cylinders (at Reynolds number 200 and gap ratio 1) using tools from dynamical systems theory. The flow is highly irregular due to the complex interaction between the flapping jet emanating from the gap and the vortices shed in the wake. We first perform Spectral Proper Orthogonal Decomposition (SPOD) to understand the flow characteristics. We then conduct Lyapunov stability analysis by linearizing the Navier-Stokes equations around the irregular base flow and find that it has two positive Lyapunov exponents. The Covariant Lyapunov Vectors (CLVs) are also computed. Contours of the time-averaged CLVs reveal that the footprint of the leading CLV is in the near-wake, whereas the other CLVs peak further downstream, indicating distinct regions of instability. SPOD of the two unstable CLVs is then employed to extract the dominant coherent structures and oscillation frequencies in the tangent space. For the leading CLV, the two dominant frequencies match closely with the prevalent frequencies in the drag coefficient spectrum, and correspond to instabilities due to vortex shedding and jet-flapping. The second unstable CLV captures the subharmonic instability of the shedding frequency. Global linear stability analysis (GLSA) of the time-averaged flow identifies a neutral eigenmode that resembles the leading SPOD mode of the first CLV, with a very similar structure and frequency. However, while GLSA predicts neutrality, Lyapunov analysis reveals that this direction is unstable, exposing the inherent limitations of the GLSA when applied to chaotic flows.

physics.flu-dyn

Forecasting the evolution of three-dimensional turbulent recirculating flows from sparse sensor data

A data-driven algorithm is proposed that employs sparse data from velocity and/or scalar sensors to forecast the future evolution of three dimensional turbulent flows. The algorithm combines time-delayed embedding together with Koopman theory and linear optimal estimation theory. It consists of 3 steps; dimensionality reduction (currently POD), construction of a linear dynamical system for current and future POD coefficients and system closure using sparse sensor measurements. In essence, the algorithm establishes a mapping from current sparse data to the future state of the dominant structures of the flow over a specified time window. The method is scalable (i.e.\ applicable to very large systems), physically interpretable, and provides sequential forecasting on a sliding time window of prespecified length. It is applied to the turbulent recirculating flow over a surface-mounted cube (with more than $10^8$ degrees of freedom) and is able to forecast accurately the future evolution of the most dominant structures over a time window at least two orders of magnitude larger that the (estimated) Lyapunov time scale of the flow. Most importantly, increasing the size of the forecasting window only slightly reduces the accuracy of the estimated future states. Extensions of the method to include convolutional neural networks for more efficient dimensionality reduction and moving sensors are also discussed.

physics.flu-dyn

Auto-Configuring Entity Resolution Pipelines

The same real-world entity (e.g., a movie, a restaurant, a person) may be described in various ways on different datasets. Entity Resolution (ER) aims to find such different descriptions of the same entity, this way improving data quality and, therefore, data value. However, an ER pipeline typically involves several steps (e.g., blocking, similarity estimation, clustering), with each step requiring its own configurations and tuning. The choice of the best configuration, among a vast number of possible combinations, is a dataset-specific and labor-intensive task both for novice and expert users, while it often requires some ground truth knowledge of real matches. In this work, we examine ways of automatically configuring a state of-the-art end-to-end ER pipeline based on pre-trained language models under two settings: (i) When ground truth is available. In this case, sampling strategies that are typically used for hyperparameter optimization can significantly restrict the search of the configuration space. We experimentally compare their relative effectiveness and time efficiency, applying them to ER pipelines for the first time. (ii) When no ground truth is available. In this case, labelled data extracted from other datasets with available ground truth can be used to train a regression model that predicts the relative effectiveness of parameter configurations. Experimenting with 11 ER benchmark datasets, we evaluate the relative performance of existing techniques that address each problem, but have not been applied to ER before.

cs.DB

Progressive Entity Resolution: A Design Space Exploration

Entity Resolution (ER) is typically implemented as a batch task that processes all available data before identifying duplicate records. However, applications with time or computational constraints, e.g., those running in the cloud, require a progressive approach that produces results in a pay-as-you-go fashion. Numerous algorithms have been proposed for Progressive ER in the literature. In this work, we propose a novel framework for Progressive Entity Resolution that organizes relevant techniques into four consecutive steps: (i) filtering, which reduces the search space to the most likely candidate matches, (ii) weighting, which associates every pair of candidate matches with a similarity score, (iii) scheduling, which prioritizes the execution of the candidate matches so that the real duplicates precede the non-matching pairs, and (iv) matching, which applies a complex, matching function to the pairs in the order defined by the previous step. We associate each step with existing and novel techniques, illustrating that our framework overall generates a superset of the main existing works in the field. We select the most representative combinations resulting from our framework and fine-tune them over 10 established datasets for Record Linkage and 8 for Deduplication, with our results indicating that our taxonomy yields a wide range of high performing progressive techniques both in terms of effectiveness and time efficiency.

cs.DB

On the low drag regime of flatback airfoils

Flatback airfoils, characterized by a blunt trailing edge, are used at the root of large wind turbine blades. A low-drag pocket has recently been identified in the flow past these airfoils at high angles of attack, potentially offering opportunities for enhanced energy extraction. This study uses three-dimensional Detached Eddy Simulations (DES) combined with statistical and data-driven modal analysis techniques to explore the aerodynamics and coherent structures of a flatback airfoil in these conditions. Two angles of attack - one inside $\left(12^{\circ}\right)$ and one outside $\left(0^{\circ}\right)$ of the low-drag pocket - are examined more thoroughly. The spanwise correlation length of secondary instability is analyzed in terms of autocorrelation of the $\Gamma_{1}$ vortex identification criterion, while coherent structures were extracted via the multiscale Proper Orthogonal Decomposition (mPOD). The results show increased base pressure, BL thickness, vortex formation length, and more organized wake structures inside the low-drag regime. While the primary instability (B\'enard-von K\'arm\'an vortex street) dominates in both cases, the secondary instability is distinguishable only for the $12^{\circ}$ case and is identified as a Mode S$^{\prime}$ instability.

physics.flu-dyn

A Critical Re-evaluation of Benchmark Datasets for (Deep) Learning-Based Matching Algorithms

Entity resolution (ER) is the process of identifying records that refer to the same entities within one or across multiple databases. Numerous techniques have been developed to tackle ER challenges over the years, with recent emphasis placed on machine and deep learning methods for the matching phase. However, the quality of the benchmark datasets typically used in the experimental evaluations of learning-based matching algorithms has not been examined in the literature. To cover this gap, we propose four different approaches to assessing the difficulty and appropriateness of 13 established datasets: two theoretical approaches, which involve new measures of linearity and existing measures of complexity, and two practical approaches: the difference between the best non-linear and linear matchers, as well as the difference between the best learning-based matcher and the perfect oracle. Our analysis demonstrates that most of the popular datasets pose rather easy classification tasks. As a result, they are not suitable for properly evaluating learning-based matching algorithms. To address this issue, we propose a new methodology for yielding benchmark datasets. We put it into practice by creating four new matching tasks, and we verify that these new benchmarks are more challenging and therefore more suitable for further advancements in the field.

cs.DB

Uncertainty quantification of time-average quantities of chaotic systems using sensitivity-enhanced polynomial chaos expansion

We consider the effect of multiple stochastic parameters on the time-average quantities of chaotic systems. We employ the recently proposed \cite{Kantarakias_Papadakis_2023} sensitivity-enhanced generalized polynomial chaos expansion, se-gPC, to compute efficiently this effect. se-gPC is an extension of gPC expansion, enriched with the sensitivity of the time-averaged quantities with respect to the stochastic variables. To compute these sensitivities, the adjoint of the shadowing operator is derived in the frequency domain. Coupling the adjoint operator with gPC provides an efficient uncertainty quantification (UQ) algorithm which, in its simplest form, has computational cost that is independent of the number of random variables. The method is applied to the Kuramoto-Sivashinsky equation and is found to produce results that match very well with Monte-Carlo simulations. The efficiency of the proposed method significantly outperforms sparse-grid approaches, like Smolyak Quadrature. These properties make the method suitable for application to other dynamical systems with many stochastic parameters.

nlin.CD

An Overset Algorithm for Multiphase Flows using 3D Multiblock Polyhedral Meshes

In this study, we present a parallel topology algorithm with a suitable interpolation method for chimera simulations in CFD. The implementation is done in the unstructured Finite Volume (FV) framework and special attention is given to the numerical algorithm. The aim of the proposed algorithm is to approximate fields with discontinuities with application to two-phase incompressible flows. First, the overset topology problem in partitioned polyhedral meshes is addressed, and then a new interpolation algorithm for generally discontinuous fields is introduced. We describe how the properties of FV are used in favor of the interpolation algorithm and how this intuitive process helps to achieve high-resolution results. The performance of the proposed algorithm is quantified and tested in various test cases, together with a comparison with an already existing interpolation scheme. Finally, scalability tests are presented to prove computational efficiency. The method suggests to be highly accurate in propagation cases and performs well in unsteady two-phase problems executed in parallel architectures.

physics.flu-dyn

Pre-trained Embeddings for Entity Resolution: An Experimental Analysis [Experiment, Analysis & Benchmark]

Many recent works on Entity Resolution (ER) leverage Deep Learning techniques involving language models to improve effectiveness. This is applied to both main steps of ER, i.e., blocking and matching. Several pre-trained embeddings have been tested, with the most popular ones being fastText and variants of the BERT model. However, there is no detailed analysis of their pros and cons. To cover this gap, we perform a thorough experimental analysis of 12 popular language models over 17 established benchmark datasets. First, we assess their vectorization overhead for converting all input entities into dense embeddings vectors. Second, we investigate their blocking performance, performing a detailed scalability analysis, and comparing them with the state-of-the-art deep learning-based blocking method. Third, we conclude with their relative performance for both supervised and unsupervised matching. Our experimental results provide novel insights into the strengths and weaknesses of the main language models, facilitating researchers and practitioners to select the most suitable ones in practice.

cs.DB

Investigation of Submergence Depth and Wave-Induced Effects on the Performance of a Fully Passive Energy Harvesting Flapping Foil Operating Beneath the Free Surface

This paper investigates the performance of a fully passive flapping foil device for energy harvesting in a free surface flow. The study uses numerical simulations to examine the effects of varying submergence depths and the impact of monochromatic waves on the foil's performance. The results show that the fully passive flapping foil device can achieve high efficiency for submergence depths between 4 and 9 chords, with an "optimum" submergence depth where the flapping foil performance is maximised. The performance was found to be correlated with the resonant frequency of the heaving motion and its proximity to the damped natural frequency. The effects of regular waves on the foil's performance were also investigated, showing that waves with a frequency close to that of the natural frequency of the flapping foil aided energy harvesting. Overall, this study provides insights that could be useful for future design improvements for fully passive flapping foil devices for energy harvesting operating near the free surface.

physics.flu-dyn