SearcharxivSearch

arXiv subjects

F. W. Takes

Publications and source records attributed to F. W. Takes.

2 recordsLinked to original sources

Probabilistic Salary Prediction with Graph Attention Networks and a Mixture Density Network

Accurate salary prediction is critical for bridging the information gap between employers and job seekers in modern labor markets. Existing approaches predominantly yield a single point estimate and treat job attributes such as location, occupation, and industry as independent categorical features, ignoring both the inherent uncertainty and multi-modality of real-world compensation data and the rich hierarchical and semantic-similarity relationships that govern pay norms. In this paper we propose GAT-MDN, a unified framework that addresses both limitations simultaneously. For each of the three attribute domains we construct a domain-specific graph whose edges encode (i) hierarchical parent-child containment and (ii) weighted similarity links derived from a pre-trained Sentence-Transformer. Parallel Graph Attention Networks (GATs) with edge-feature-aware attention learn rich, context-sensitive node representations from these multi-relational graphs. A priority-based hierarchical selection module then assembles a composite feature vector that gracefully handles missing or coarse attributes, and a Mixture Density Network (MDN) head maps this vector to the parameters of a Gaussian Mixture Model (GMM), yielding a full conditional salary distribution. Extensive experiments on a real-world Dutch job-posting dataset of over 1 million records demonstrate that GAT-MDN significantly outperforms a non-graph MLP-MDN baseline in both Negative Log-Likelihood (NLL) and Mean Squared Error (MSE).

cs.SI

A Data-Driven Supply-Side Approach for Measuring Cross-Border Internet Purchases

The digital economy is a highly relevant item on the European Union's policy agenda. Cross-border internet purchases are part of the digital economy, but their total value can currently not be accurately measured or estimated. Traditional approaches based on consumer surveys or business surveys are shown to be inadequate for this purpose, due to language bias and sampling issues, respectively. We address both problems by proposing a novel approach based on supply-side data, namely tax returns. The proposed data-driven record-linkage techniques and machine learning algorithms utilize two additional open data sources: European business registers and internet data. Our main finding is that the value of total cross-border internet purchases within the European Union by Dutch consumers was over EUR 1.3 billion in 2016. This is more than 6 times as high as current estimates. Our finding motivates the implementation of the proposed methodology in other EU member states. Ultimately, it could lead to more accurate estimates of cross-border internet purchases within the entire European Union.

stat.AP