SearcharxivSearch

arXiv subjects

Eric Simon

Publications and source records attributed to Eric Simon.

5 recordsLinked to original sources

Sequential and Generative Models for Vehicular Distributed MIMO Channel Prediction

Vehicular communication is a key 6G use case requiring reliable and high-capacity connectivity under fast mobility and highly time-varying propagation conditions. However, large-scale vehicular channel estimation is costly and limited, impacting system-level performance of vehicular communications, and realistic channel prediction models are needed. This paper proposes a vehicular channel prediction framework based on real measured urban channels collected through a dedicated measurement campaign using the MaMIMOSA channel sounder. The framework enables the training and systematic benchmarking of sequential and generative models for both single-step and multi-horizon vehicular channel state information (CSI) prediction to assess prediction robustness across different forecasting horizons, including LSTM, TCN, a CNN-enhanced Transformer, and ChannelGPT, with the goal of accurately predicting channel evolution while preserving spatiotemporal dynamics and non-stationarity. In addition, a system-level evaluation framework is introduced to assess the impact of channel prediction on the performance of vehicular distributed MIMO communications. Using predicted channels, spectral efficiency (SE) is evaluated against true CSI. Results show that ChannelGPT achieves over 94% normalized mean squared error (NMSE) reduction compared to LSTM and significant improvements over other baselines, while reducing FLOPs by 28% and inference latency by 39% relative to the CNN + Transformer. Moreover, ChannelGPT-predicted channels yield SE distributions nearly indistinguishable from those obtained with real measurements, demonstrating its effectiveness for reliable performance evaluation in high-mobility 6G vehicular networks.

eess.SP

LOG.io: Unified Rollback Recovery and Data Lineage Capture for Distributed Data Pipelines

This paper introduces LOG.io, a comprehensive solution designed for correct rollback recovery and fine-grain data lineage capture in distributed data pipelines. It is tailored for serverless scalable architectures and uses a log-based rollback recovery protocol. LOG.io supports a general programming model, accommodating non-deterministic operators, interactions with external systems, and arbitrary custom code. It is non-blocking, allowing failed operators to recover independently without interrupting other active operators, thereby leveraging data parallelization, and it facilitates dynamic scaling of operators during pipeline execution. Performance evaluations, conducted within the SAP Data Intelligence system, compare LOG.io with the Asynchronous Barrier Snapshotting (ABS) protocol, originally implemented in Flink. Our experiments show that when there are straggler operators in a data pipeline and the throughput of events is moderate (e.g., 1 event every 100 ms), LOG.io performs as well as ABS during normal processing and outperforms ABS during recovery. Otherwise, ABS performs better than LOG.io for both normal processing and recovery. However, we show that in these cases, data parallelization can largely reduce the overhead of LOG.io while ABS does not improve. Finally, we show that the overhead of data lineage capture, at the granularity of the event and between any two operators in a pipeline, is marginal, with less than 1.5% in all our experiments.

cs.DC

Automated Planning for Optimal Data Pipeline Instantiation

Data pipeline frameworks provide abstractions for implementing sequences of data-intensive transformation operators, automating the deployment and execution of such transformations in a cluster. Deploying a data pipeline, however, requires computing resources to be allocated in a data center, ideally minimizing the overhead for communicating data and executing operators in the pipeline while considering each operator's execution requirements. In this paper, we model the problem of optimal data pipeline deployment as planning with action costs, where we propose heuristics aiming to minimize total execution time. Experimental results indicate that the heuristics can outperform the baseline deployment and that a heuristic based on connections outperforms other strategies.

cs.AI

A Meta-level Analysis of Online Anomaly Detectors

Real-time detection of anomalies in streaming data is receiving increasing attention as it allows us to raise alerts, predict faults, and detect intrusions or threats across industries. Yet, little attention has been given to compare the effectiveness and efficiency of anomaly detectors for streaming data (i.e., of online algorithms). In this paper, we present a qualitative, synthetic overview of major online detectors from different algorithmic families (i.e., distance, density, tree or projection-based) and highlight their main ideas for constructing, updating and testing detection models. Then, we provide a thorough analysis of the results of a quantitative experimental evaluation of online detection algorithms along with their offline counterparts. The behavior of the detectors is correlated with the characteristics of different datasets (i.e., meta-features), thereby providing a meta-level analysis of their performance. Our study addresses several missing insights from the literature such as (a) how reliable are detectors against a random classifier and what dataset characteristics make them perform randomly; (b) to what extent online detectors approximate the performance of offline counterparts; (c) which sketch strategy and update primitives of detectors are best to detect anomalies visible only within a feature subspace of a dataset; (d) what are the tradeoffs between the effectiveness and the efficiency of detectors belonging to different algorithmic families; (e) which specific characteristics of datasets yield an online algorithm to outperform all others.

cs.LG

Controlling the Correctness of Aggregation Operations During Sessions of Interactive Analytic Queries

We present a comprehensive set of conditions and rules to control the correctness of aggregation queries within an interactive data analysis session. The goal is to extend self-service data preparation and BI tools to automatically detect semantically incorrect aggregate queries on analytic tables and views built by using the common analytic operations including filter, project, join, aggregate, union, difference, and pivot. We introduce aggregable properties to describe for any attribute of an analytic table which aggregation functions correctly aggregates the attribute along which sets of dimension attributes. These properties can also be used to formally identify attributes which are summarizable with respect to some aggregation function along a given set of dimension attributes. This is particularly helpful to detect incorrect aggregations of measures obtained through the use of non-distributive aggregation functions like average and count. We extend the notion of summarizability by introducing a new generalized summarizability condition to control the aggregation of attributes after any analytic operation. Finally, we define propagation rules which transform aggregable properties of the query input tables into new aggregable properties for the result tables, preserving summarizability and generalized summarizability.

cs.DB