SearcharxivSearch

arXiv subjects

Alexandros Sopasakis

Publications and source records attributed to Alexandros Sopasakis.

11 recordsLinked to original sources

Testing the limits of past-adapted explanations by post-endpoint randomisation: anticipatory EEG as a worked case

A predictive model can fit its data even when its information set is insufficient; fit alone cannot establish sufficiency. This Perspective introduces Level II-A, a new design-based inference framework to test this distinction, illustrated in anticipatory EEG using contingent negative variation. A pre-event endpoint is committed before the delay to the imperative event is randomised. That later-assigned delay thereby becomes a negative-control probe of whether past-adapted information was sufficient for an already committed result. Under the past-adapted factorisation, accounts using only pre-commitment information cannot systematically order the endpoint by that delay. Leakage-safe preprocessing, a frozen label-blind comparator and retained-sample qualifications carry the exclusion to the confirmatory residual. A qualified material negative ordering supports conditional insufficiency without identifying a mechanism; an adequately sensitive null supports a bounded affirmative conclusion calibrated by the pipeline's false-adequacy rate. A non-compensatory rule separates these from diagnostic failure, selection-limited, opposite-direction and inconclusive outcomes. No human EEG data are analysed. In the synthetic benchmark, grid-based false-adequacy boundaries are $15\,μ\mathrm{V\,s^{-1}}$ for assignment isolation and $30\,μ\mathrm{V\,s^{-1}}$ for the sequential e-value route, in both directions. The design transfers wherever endpoint commitment precedes an exogenous label, probing sufficiency only where a declared alternative predicts ordering by it. It turns "the past explains it" from a working explanatory assumption into a magnitude-qualified, testable claim.

stat.ME

SMARTHEP: training PhD students in real-time analysis at the LHC and in industry

In this invited Editorial for Software and Computing for Big Science, we describe the SMARTHEP Innovative Training Network funded via the Marie Skłodowska-Curie Actions between 2021 and 2025. SMARTHEP trained 12 PhD students to advance machine learning and real-time analysis in high-energy physics experiments and industrial applications. We present the perspective of students, supervisors, and external observers of the network, concerning the work done within the network, the added value compared to ``typical'' PhD positions, and the emerging themes and directions from our experiences in the past four years.

physics.ed-ph

Using Text-Based Life Trajectories from Swedish Register Data to Predict Residential Mobility with Pretrained Transformers

We transform large-scale Swedish register data into textual life trajectories to address two long-standing challenges in data analysis: high cardinality of categorical variables and inconsistencies in coding schemes over time. Leveraging this uniquely comprehensive population register, we convert register data from 6.9 million individuals (2001-2013) into semantically rich texts and predict individuals' residential mobility in later years (2013-2017). These life trajectories combine demographic information with annual changes in residence, work, education, income, and family circumstances, allowing us to assess how effectively such sequences support longitudinal prediction. We compare multiple NLP architectures (including LSTM, DistilBERT, BERT, and Qwen) and find that sequential and transformer-based models capture temporal and semantic structure more effectively than baseline models. The results show that textualized register data preserves meaningful information about individual pathways and supports complex, scalable modeling. Because few countries maintain longitudinal microdata with comparable coverage and precision, this dataset enables analyses and methodological tests that would be difficult or impossible elsewhere, offering a rigorous testbed for developing and evaluating new sequence-modeling approaches. Overall, our findings demonstrate that combining semantically rich register data with modern language models can substantially advance longitudinal analysis in social sciences.

cs.LG

Learning Traffic Anomalies from Generative Models on Real-Time Observations

Accurate detection of traffic anomalies is crucial for effective urban traffic management and congestion mitigation. We use the Spatiotemporal Generative Adversarial Network (STGAN) framework combining Graph Neural Networks and Long Short-Term Memory networks to capture complex spatial and temporal dependencies in traffic data. We apply STGAN to real-time, minute-by-minute observations from 42 traffic cameras across Gothenburg, Sweden, collected over several months in 2020. The images are processed to compute a flow metric representing vehicle density, which serves as input for the model. Training is conducted on data from April to November 2020, and validation is performed on a separate dataset from November 14 to 23, 2020. Our results demonstrate that the model effectively detects traffic anomalies with high precision and low false positive rates. The detected anomalies include camera signal interruptions, visual artifacts, and extreme weather conditions affecting traffic flow.

cs.LG

Early-Scheduled Handover Preparation in 5G NR Millimeter-Wave Systems

The handover (HO) procedure is one of the most critical functions in a cellular network driven by measurements of the user channel of the serving and neighboring cells. The success rate of the entire HO procedure is significantly affected by the preparation stage. As massive Multiple-Input Multiple-Output (MIMO) systems with large antenna arrays allow resolving finer details of channel behavior, we investigate how machine learning can be applied to time series data of beam measurements in the Fifth Generation (5G) New Radio (NR) system to improve the HO procedure. This paper introduces the Early-Scheduled Handover Preparation scheme designed to enhance the robustness and efficiency of the HO procedure, particularly in scenarios involving high mobility and dense small cell deployments. Early-Scheduled Handover Preparation focuses on optimizing the timing of the HO preparation phase by leveraging machine learning techniques to predict the earliest possible trigger points for HO events. We identify a new early trigger for HO preparation and demonstrate how it can beneficially reduce the required time for HO execution reducing channel quality degradation. These insights enable a new HO preparation scheme that offers a novel, user-aware, and proactive HO decision making in MIMO scenarios incorporating mobility.

cs.LG

Improved Anomaly Detection through Conditional Latent Space VAE Ensembles

We propose a novel Conditional Latent space Variational Autoencoder (CL-VAE) to perform improved pre-processing for anomaly detection on data with known inlier classes and unknown outlier classes. This proposed variational autoencoder (VAE) improves latent space separation by conditioning on information within the data. The method fits a unique prior distribution to each class in the dataset, effectively expanding the classic prior distribution for VAEs to include a Gaussian mixture model. An ensemble of these VAEs are merged in the latent spaces to form a group consensus that greatly improves the accuracy of anomaly detection across data sets. Our approach is compared against the capabilities of a typical VAE, a CNN, and a PCA, with regards AUC for anomaly detection. The proposed model shows increased accuracy in anomaly detection, achieving an AUC of 97.4% on the MNIST dataset compared to 95.7% for the second best model. In addition, the CL-VAE shows increased benefits from ensembling, a more interpretable latent space, and an increased ability to learn patterns in complex data with limited model sizes.

cs.LG

Enhancing Carbon Emission Reduction Strategies using OCO and ICOS data

We propose a methodology to enhance local CO2 monitoring by integrating satellite data from the Orbiting Carbon Observatories (OCO-2 and OCO-3) with ground level observations from the Integrated Carbon Observation System (ICOS) and weather data from the ECMWF Reanalysis v5 (ERA5). Unlike traditional methods that downsample national data, our approach uses multimodal data fusion for high-resolution CO2 estimations. We employ weighted K-nearest neighbor (KNN) interpolation with machine learning models to predict ground level CO2 from satellite measurements, achieving a Root Mean Squared Error of 3.92 ppm. Our results show the effectiveness of integrating diverse data sources in capturing local emission patterns, highlighting the value of high-resolution atmospheric transport models. The developed model improves the granularity of CO2 monitoring, providing precise insights for targeted carbon mitigation strategies, and represents a novel application of neural networks and KNN in environmental monitoring, adaptable to various regions and temporal scales.

cs.LG

Summary of the trigger systems of the Large Hadron Collider experiments ALICE, ATLAS, CMS and LHCb

In modern High Energy Physics (HEP) experiments, triggers perform the important task of selecting, in real time, the data to be recorded and saved for physics analyses. As a result, trigger strategies play a key role in extracting relevant information from the vast streams of data produced at facilities like the Large Hadron Collider (LHC). As the energy and luminosity of the collisions increase, these strategies must be upgraded and maintained to suit the experimental needs. This whitepaper compiled by the SMARTHEP Early Stage Researchers presents a high-level overview and reviews recent developments of triggering practices employed at the LHC. The general trigger principles applied at modern HEP experiments are highlighted, with specific reference to the current trigger state-of-the-art within the ALICE, ATLAS, CMS and LHCb collaborations. Furthermore, a brief synopsis of the new trigger paradigm required by the upcoming high-luminosity upgrade of the LHC is provided.

physics.ins-det

Learning-Based UE Classification in Millimeter-Wave Cellular Systems With Mobility

Millimeter-wave cellular communication requires beamforming procedures that enable alignment of the transmitter and receiver beams as the user equipment (UE) moves. For efficient beam tracking it is advantageous to classify users according to their traffic and mobility patterns. Research to date has demonstrated efficient ways of machine learning based UE classification. Although different machine learning approaches have shown success, most of them are based on physical layer attributes of the received signal. This, however, imposes additional complexity and requires access to those lower layer signals. In this paper, we show that traditional supervised and even unsupervised machine learning methods can successfully be applied on higher layer channel measurement reports in order to perform UE classification, thereby reducing the complexity of the classification process.

cs.IT

Latent space conditioning for improved classification and anomaly detection

We propose a new type of variational autoencoder to perform improved pre-processing for clustering and anomaly detection on data with a given label. Anomalies however are not known or labeled. We call our method conditional latent space variational autonencoder since it separates the latent space by conditioning on information within the data. The method fits one prior distribution to each class in the dataset, effectively expanding the prior distribution to include a Gaussian mixture model. Our approach is compared against the capabilities of a typical variational autoencoder by measuring their V-score during cluster formation with respect to the k-means and EM algorithms. For anomaly detection, we use a new metric composed of the mass-volume and excess-mass curves which can work in an unsupervised setting. We compare the results between established methods such as as isolation forest, local outlier factor and one-class support vector machine.

cs.LG

Error analysis of coarse-grained kinetic Monte Carlo method

In this paper we investigate the approximation properties of the coarse-graining procedure applied to kinetic Monte Carlo simulations of lattice stochastic dynamics. We provide both analytical and numerical evidence that the hierarchy of the coarse models is built in a systematic way that allows for error control in both transient and long-time simulations. We demonstrate that the numerical accuracy of the CGMC algorithm as an approximation of stochastic lattice spin flip dynamics is of order two in terms of the coarse-graining ratio and that the natural small parameter is the coarse-graining ratio over the range of particle/particle interactions. The error estimate is shown to hold in the weak convergence sense. We employ the derived analytical results to guide CGMC algorithms and we demonstrate a CPU speed-up in demanding computational regimes that involve nucleation, phase transitions and metastability.

math.NA