Searcharxiv⌕ Search

arXiv subjects

Kees Jan de Vries

Publications and source records attributed to Kees Jan de Vries.

3 recordsLinked to original sources

SOHET: Sequence Of Heterogeneous Events Transformer with Self-Supervised Pre-Training

Many machine learning applications rely on heterogeneous event streams to make predictions, either causally as events arrive or bidirectionally over complete sequences. We propose SOHET (Sequence Of Heterogeneous Events Transformer), a hierarchical architecture combining event-type-specific tabular encoders with temporal and type embeddings, processed by a causal or bidirectional transformer. We introduce three self-supervised pre-training objectives for the causal setting. On a proprietary large-scale real-world Booking.com fraud detection task with 17 event types, SOHET outperforms FlexTPP, NAPPT, and CIPPT by 5.8%. Pre-training yields an additional 2.6% gain and 2.4% faster convergence. On the EBES benchmark, bidirectional SOHET matches or exceeds the published best on 6 out of 8 tasks.

cs.LG↗

Machine Learning for Fraud Detection in E-Commerce: A Research Agenda

Fraud detection and prevention play an important part in ensuring the sustained operation of any e-commerce business. Machine learning (ML) often plays an important role in these anti-fraud operations, but the organizational context in which these ML models operate cannot be ignored. In this paper, we take an organization-centric view on the topic of fraud detection by formulating an operational model of the anti-fraud departments in e-commerce organizations. We derive 6 research topics and 12 practical challenges for fraud detection from this operational model. We summarize the state of the literature for each research topic, discuss potential solutions to the practical challenges, and identify 22 open research challenges.

cs.LG↗

SUSY fits with full LHC Run I data

We present the latest results from the MasterCode Collaboration on supersymmetric models, in particular on the CMSSM, the NUHM1, the NUHM2 and the pMSSM. We combine the data from LHC Run I with astrophysical observables, flavor and electroweak precision observables. We determine the best fit regions of these models and analyze the discovery potential of squarks and gluinos at LHC Run II and direct detection experiments.

hep-ph↗