SearcharxivSearch

arXiv · 2601.12491

VASTU: Language Models Struggle to Recognize Online Community Values

Abstract

Online communities develop distinct norms for content they collectively value, yet it remains unclear whether current language models can recognize locally valued contributions in context. We formalize this as \textbf{community-conditioned preference prediction} and introduce \textsc{Vastu} (\underline{V}alue-\underline{A}ware \underline{S}ocial \underline{Tu}ning), a benchmark of 75,000 Reddit comments from 15 communities spanning Gaming, Science, Q\&A, Advice, and Politics. We evaluate four model families---prompted LLMs, LoRA-adapted SLMs, supervised encoders, and feature-based classifiers---across global, local, and context-conditioned settings. Our central finding is that parametric adaptation consistently outperforms prompting: supervised encoders reach 0.74 AUROC and fine-tuned SLMs 0.64--0.71, while the best prompted result is only 0.62. This gap is not merely quantitative---vanilla prompting yields over 80\% false-negative rates, systematically discarding content communities actually value. Conversational context narrows but does not close this divide. Together, these results suggest that local preference recognition requires community-specific training signal, not just better prompting. Our work supports future research on community-aware reward modeling, feed curation, and positive moderation.

Explore related subjects

Keep this discovery

BibTeXRIS

Agam Goyal, Xianyang Zhan, Charlotte Lambert, Koustuv Saha, Eshwar Chandrasekharan. 2026-08-28. VASTU: Language Models Struggle to Recognize Online Community Values. https://arxiv.org/abs/2601.12491

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related discoveries

GRAND-HC: Graph-Refined Author Name Disambiguation

From-Scratch Name Disambiguation (SND) groups papers sharing an ambiguous name into clusters of distinct real-world authors. Existing methods suffer from two critical limitations: (1) inherent long-tailed author distribution biases representation learning, causing over-merging of tail authors; (2) existing cluster number estimation methods are unreliable for long paper sequences, hindering large-scale deployment. We propose \textbf{GRAND-HC}, a complete end-to-end SND framework. We construct a heterogeneous paper graph via co-author, co-organization, and co-venue relations, using a graph attention network as the embedding backbone. \textbf{Harmony Contrastive Learning (HCL)} dynamically reweights training loss to suppress overfitting to prolific authors, learning discriminative embeddings. A \textbf{Graph-Refined Distance Matrix (GRDM)} leverages graph topology to optimize pairwise distances, further preventing tail author over-merging. Meanwhile, a lightweight \textbf{Paper Compression Module (PCM)} achieves accurate cluster number estimation across varying scales. Finally, Hierarchical Agglomerative Clustering outputs the final clusters. Extensive experiments demonstrate state-of-the-art macro F1 performance. GRAND-HC has been deployed in a billion-scale academic database. Source code: https://github.com/baokou-fw2/GRAND-HC.

cs.IR

FocusAdapt: Context-aware Adaptive Focus Assistance in Diminished Reality

Diminished Reality (DR) can reduce visual clutter by removing irrelevant objects. However, removing all task-irrelevant objects may eliminate useful contextual information and reduce situational awareness. We present FocusAdapt, a context-aware DR system that predicts object-level distraction by integrating visual saliency, semantic relevance, and gaze behavior. Based on findings from a formative study, FocusAdapt selectively diminishes highly distracting objects while preserving useful context, enabling adaptive focus assistance during procedural tasks.

cs.HC

TSExplorer: An interactive data annotation and exploration tool for time-series data

We present TSExplorer, a cross-platform tool for interactive annotation and exploration of time-series data. The tool enables users to inspect high-dimensional datasets through multiple complementary 2D visualizations derived from high-dimensional feature representations. TSExplorer is designed as a general-purpose research tool supporting a wide range of workflows, including exploratory data analysis, annotation of unlabeled or partially-labeled datasets, comparison of feature representations, and post-hoc inspection and refinement of existing labels with interactive visual feedback.

cs.HC