arXiv · 2609.37472
Do Evidence-Reading Diagnostics Improve Interface Selection in Small LLM Recommenders?
Abstract
Behavioral tests measure how a language model reads evidence. We ask whether those measurements help choose a recommendation interface. We evaluate six small instruction-tuned checkpoints across four recommendation domains with chronological evaluation and 3,426 evaluation users. Each request ranks eight candidates. A baseline selector chooses among history-only prompting, prompting with collaborative evidence, and score fusion. It uses observable features and six stability prompts that vary wording and candidate order. An augmented selector adds features from six evidence-reading prompts that ask the model to compare support counts. An interface chosen once on development (validation) data for each domain and checkpoint scores 0.5524 NDCG@5, compared with 0.5447 for the baseline selector and 0.5428 for the augmented selector. Adding the diagnostic features changes NDCG@5 by -0.0019 (95% interval [-0.0046, 0.0004]). The interval includes zero, and its upper bound is below the analysis plan's 0.005 improvement target. Matching the selectors' hyperparameters also leaves the interval upper bound below that target. Evidence from retrieved similar users improves prompting by 0.0999 NDCG@5 over a control using randomly selected users matched for activity. The evidence-reading tests also reveal answer-position and tie-response biases. These results concern the tested selectors and candidate sets. They illustrate why diagnostic measurements should be evaluated by whether they improve recommendation choices beyond existing features and a fixed interface.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Han Chen, Yingrui Li. 2026-09-25. Do Evidence-Reading Diagnostics Improve Interface Selection in Small LLM Recommenders?. https://arxiv.org/abs/2609.37472
Cite the original work for its findings. Save a collection to share your selection of sources.