arXiv · 2609.13844
Pre-training with Graph Transformers
Abstract
This article investigates pre-training strategies for graph transformers in the biochemistry domain. By conducting comprehensive experiments, the study reveals that supervised pre-training using computed properties as labels provides the highest performance gain on downstream tasks. The results also highlight the importance of constraining model capacity to mitigate overfitting in graph transformers.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jiaming Wang, Thomas Laurent, Xavier Bresson. 2026-09-12. Pre-training with Graph Transformers. https://arxiv.org/abs/2609.13844
Cite the original work for its findings. Save a collection to share your selection of sources.