Searcharxiv⌕ Search

arXiv · 2609.40241

Decision-Oriented Recommendation Reranking: An Empirical Study of Jev

Abstract

Large language models (LLMs) have shown promise for recommendation reranking, but their use introduces an important tradeoff between recommendation quality and serving efficiency. We investigate whether a decision-oriented model provides a useful alternative when the reranking task is fundamentally a structured choice among predefined candidate items. Specifically, we conduct a controlled empirical study of Jev, described by TypeSafe AI as a ``System One Model,'' for personalized recommendation reranking and compare it with recommendation-specific models and pointwise and listwise Qwen rerankers across multiple Amazon Reviews domains and candidate-set sizes, evaluating both recommendation effectiveness and observed serving latency. Our results show that Jev maintains strong recommendation effectiveness relative to the evaluated baselines while exhibiting substantially more gradual latency growth than the pointwise Qwen rerankers, although its observed serving latency remains substantially higher than that of recommendation-specific models. Together, these characteristics place Jev in a distinct quality--latency operating regime across candidate sizes and domains. These findings motivate further investigation of decision-oriented models for recommendation and other ranking tasks with structured output spaces.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Hanjia Lyu, Yinglong Xia. 2026-09-30. Decision-Oriented Recommendation Reranking: An Empirical Study of Jev. https://arxiv.org/abs/2609.40241

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Agent-Facing Information Design in LLM Tool Registries: A Preregistered Test of Rhetoric, Position and Structure

AI agents often pick tools from registries, where each tool's provider writes its description. We ask whether sales language in those descriptions changes which tool an agent picks. We built pairs of listings differing in one controlled way (added praise, a verifiable specification, or list order) and asked two OpenAI models to call one tool. In a preregistered study, stacked praise (four kinds combined) raised a tool's pick rate by about 43 percentage points, matching or beating a verifiable specification. Praise also pulled some picks toward tools that could not do the task, but rarely toward tools asking for unneeded data access. With identical listings, the first-listed tool was picked about 72 points more often. On tasks with numeric limits, structured fields helped agents pick the capable tool; adding the provider's sales text beside the fields reduced or erased that gain. Registries could list limits as fields, hide sales text from agents, and randomize order. Stacked praise, but no single kind, replicated on held-out domains. Results are provisional until blind phrase ratings are complete, and cover two small models.

cs.IR↗

omni-macos: On-Device Omni-Modal Search on Apple Silicon

We present omni-macos, a search engine that embeds text, code, documents, images, audio and video into one representation space and runs its encoder, index and store on the Mac that already holds the files, so no indexed file, no typed query and no vector ever leaves the machine. It keeps a background indexer and an interactive search box inside one memory budget the user sets: it embeds and stores each distinct chunk once, re-encodes only the chunks an edit changes, hands the GPU smaller units while the user is typing, answers queries from a one-bit replica of the index with exact rescoring, and propagates that budget to the allocators that draw on unified memory. We measure on five Macs spanning an eightfold range of accelerator width and a thirty-twofold range of memory, each indexing the files it already holds.

cs.IR↗

Two-Sided State-Space Models for Sequential Recommendation with Non-Random Multimodal Review Feedback

Two-sided digital platforms are inherently dynamic: user preferences shift, item popularity evolves, and reviews both reflect and drive these changes. Yet most sequential recommendation systems treat reviews as passive signals for updating user states, leaving two aspects underexplored. First, review generation is nonrandom, depending on evolving latent states of both users and items. Second, reviews can reshape item states, induce spillover across related items, and influence future user decisions. To address these gaps, we propose a two-sided state-space model (TS-SSM) for event-conditioned sequential recommendation. TS-SSM consists of three components: (1) a modality-missing-not-at-random fusion module that encodes review content and informative observation patterns; (2) user-state evolution with temporal variation and local graph message passing that uses related item states to refine user preferences; and (3) item-state evolution with asymmetric carryover of positive and negative review feedback. In experiments across six Amazon categories, TS-SSM increases Recall@20 over BSARec by 14.8%--18.8% and exceeds HM4SR by 11.7% on average. On Goodreads Fantasy, Recall@20 improves HM4SR from .5191 to .5847. Ablations highlight distinct contributions of observation patterns, local propagation, and item dynamics.

cs.IR↗