arXiv · 2512.06694
TopiCLEAR: Adaptive embedding clustering for interpretable topic discovery from short texts
Abstract
Topic discovery is a fundamental technique for text mining that identifies abstract topics within large document collections. A recent approach to topic discovery is to cluster document or sentence embeddings, typically obtained from pre-trained language models, and represent each cluster as a topic. Despite their strong empirical performance, the design principles linking the geometry of embedding spaces to human-interpretable topic organization remain unclear. Clarifying the relationship between these geometric structures and topic interpretability is therefore a key challenge in topic discovery. In this study, we propose TopiCLEAR (Topic discovery by CLustering Embeddings with Adaptive dimensionality Reduction), a simple framework that integrates document embeddings with iterative clustering based on adaptive dimensionality reduction. TopiCLEAR is guided by the hypothesis that human-interpretable topics correspond to low-dimensional geometric structures in embedding spaces and leverages adaptive dimensionality reduction to identify them. We evaluate topic quality using a document-level approach that combines quantitative evaluation based on human-labeled data with qualitative assessment in the absence of ground-truth labels. Experiments on four benchmark datasets, covering both formal and informal texts, show that TopiCLEAR consistently achieves strong agreement with human annotations, particularly for short and informal texts. Furthermore, a case study on Twitter data demonstrates that TopiCLEAR produces more interpretable topics than Latent Dirichlet Allocation (LDA), recovering both human-annotated topic structure and coherent sub-topic structure. These results highlight the effectiveness of clustering documents in a low-dimensional topic space for topic discovery.
Explore related subjects
Keep this discovery
Aoi Fujita, Taichi Yamamoto, Yuri Nakayama, Ryota Kobayashi. 2025-12-07. TopiCLEAR: Adaptive embedding clustering for interpretable topic discovery from short texts. https://doi.org/10.1016/j.eswa.2026.133996
Cite the original work for its findings. Save a collection to share your selection of sources.