SearcharxivSearch

arXiv subjects

Teng Tu

Publications and source records attributed to Teng Tu.

8 recordsLinked to original sources

EchoRec: Multi-Item Prediction-Empowered Generative Recommendation via Cycle-Consistent Preference Alignment

Generative recommendation autoregressively generates the semantic IDs of the target item, unifying preference modeling and index retrieval within the shared token space. Recent attempts have introduced Multi-Token Prediction (MTP) into this field, yet they primarily inherit its efficiency merit, leaving its potential as dense supervision unexplored. Unlocking this potential hinges on whether future behaviors qualify as informative supervision. Our analysis reveals that future behaviors carry a semantic echo of the current one far above that of random pairs, which nevertheless decays along horizons under intent transitions, making them informative yet order-dependent signals. Motivated by this, we propose EchoRec, which empowers MTP with cycle-consistent holistic preference alignment across multi-horizon for generative recommendation. It comprises two synergistic modules. Horizon-aware Preference Generation (HPG) sequentially chains lightweight auxiliary branches upon the base recommender, where each branch conditions on its predecessor to respect preference evolution. Verifiable Holistic-Preference Alignment (VHA) further consolidates them into the holistic preference and echoes it back through cycle-consistent projectors to suppress spurious alignment, with theoretical guarantees that exclude the rank-collapse form of spurious alignment under an invertible transport, enabling the holistic preference to be retained in the decoding representation. All auxiliary components serve as disposable scaffolding discarded at inference, introducing negligible online serving overhead. Extensive experiments on three datasets demonstrate the superiority of our EchoRec, together with its naturally acquired multi-item generation ability. Our code and datasets will be available upon acceptance.

cs.IR

Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond

As AI systems move from generating text to accomplishing goals through sustained interaction, the ability to model environment dynamics becomes a central bottleneck. Agents that manipulate objects, navigate software, coordinate with others, or design experiments require predictive environment models, yet the term world model carries different meanings across research communities. We introduce a "levels x laws" taxonomy organized along two axes. The first defines three capability levels: L1 Predictor, which learns one-step local transition operators; L2 Simulator, which composes them into multi-step, action-conditioned rollouts that respect domain laws; and L3 Evolver, which autonomously revises its own model when predictions fail against new evidence. The second identifies four governing-law regimes: physical, digital, social, and scientific. These regimes determine what constraints a world model must satisfy and where it is most likely to fail. Using this framework, we synthesize over 400 works and summarize more than 100 representative systems spanning model-based reinforcement learning, video generation, web and GUI agents, multi-agent social simulation, and AI-driven scientific discovery. We analyze methods, failure modes, and evaluation practices across level-regime pairs, propose decision-centric evaluation principles and a minimal reproducible evaluation package, and outline architectural guidance, open problems, and governance challenges. The resulting roadmap connects previously isolated communities and charts a path from passive next-step prediction toward world models that can simulate, and ultimately reshape, the environments in which agents operate. Code and resources are available at: https://github.com/matrix-agent/awesome-agentic-world-modeling.

cs.AI

MusicSem: A Semantically Rich Language--Audio Dataset of Natural Music Descriptions

Music representation learning is central to music information retrieval and generation. While recent advances in multimodal learning have improved alignment between text and audio for tasks such as cross-modal music retrieval, text-to-music generation, and music-to-text generation, existing models often struggle to capture users' expressed intent in natural language descriptions of music. This observation suggests that the datasets used to train and evaluate these models do not fully reflect the broader and more natural forms of human discourse through which music is described. In this paper, we introduce MusicSem, a dataset of 32,493 language-audio pairs derived from organic music-related discussions on the social media platform Reddit. Compared to existing datasets, MusicSem captures a broader spectrum of musical semantics, reflecting how listeners naturally describe music in nuanced and human-centered ways. To structure these expressions, we propose a taxonomy of five semantic categories: descriptive, atmospheric, situational, metadata-related, and contextual. In addition to the construction, analysis, and release of MusicSem, we use the dataset to evaluate a wide range of multimodal models for retrieval and generation, highlighting the importance of modeling fine-grained semantics. Overall, MusicSem serves as a novel semantics-aware resource to support future research on human-aligned multimodal music representation learning.

cs.MM

Extending Visual Dynamics for Video-to-Music Generation

Music profoundly enhances video production by improving quality, engagement, and emotional resonance, sparking growing interest in video-to-music generation. Despite recent advances, existing approaches remain limited in specific scenarios or undervalue the visual dynamics. To address these limitations, we focus on tackling the complexity of dynamics and resolving temporal misalignment between video and music representations. To this end, we propose DyViM, a novel framework to enhance dynamics modeling for video-to-music generation. Specifically, we extract frame-wise dynamics features via a simplified motion encoder inherited from optical flow methods, followed by a self-attention module for aggregation within frames. These dynamic features are then incorporated to extend existing music tokens for temporal alignment. Additionally, high-level semantics are conveyed through a cross-attention mechanism, and an annealing tuning strategy benefits to fine-tune well-trained music decoders efficiently, therefore facilitating seamless adaptation. Extensive experiments demonstrate DyViM's superiority over state-of-the-art (SOTA) methods.

cs.MM

GOAL: A Challenging Knowledge-grounded Video Captioning Benchmark for Real-time Soccer Commentary Generation

Despite the recent emergence of video captioning models, how to generate vivid, fine-grained video descriptions based on the background knowledge (i.e., long and informative commentary about the domain-specific scenes with appropriate reasoning) is still far from being solved, which however has great applications such as automatic sports narrative. In this paper, we present GOAL, a benchmark of over 8.9k soccer video clips, 22k sentences, and 42k knowledge triples for proposing a challenging new task setting as Knowledge-grounded Video Captioning (KGVC). Moreover, we conduct experimental adaption of existing methods to show the difficulty and potential directions for solving this valuable and applicable task. Our data and code are available at https://github.com/THU-KEG/goal.

cs.CV

Reply to: Low-frequency quantum oscillations in LaRhIn$_5$: Dirac point or nodal line?

We thank G.P. Mikitik and Yu.V. Sharlai for contributing this note and the cordial exchange about it. First and foremost, we note that the aim of our paper is to report a methodology to diagnose topological (semi)metals using magnetic quantum oscillations. Thus far, such diagnosis has been based on the phase offset of quantum oscillations, which is extracted from a "Landau fan plot". A thorough analysis of the Onsager-Lifshitz-Roth quantization rules has shown that the famous $\pi$-phase shift can equally well arise from orbital- or spin magnetic moments in topologically trivial systems with strong spin-orbit coupling or small effective masses. Therefore, the "Landau fan plot" does not by itself constitute a proof of a topologically nontrivial Fermi surface. In the paper at hand, we report an improved analysis method that exploits the strong energy-dependence of the effective mass in linearly dispersing bands. This leads to a characteristic temperature dependence of the oscillation frequency which is a strong indicator of nontrivial topology, even for multi-band metals with complex Fermi surfaces. Three materials, Cd$_3$As$_2$, Bi$_2$O$_2$Se and LaRhIn$_5$ served as test cases for this method. Linear band dispersions were detected for Cd$_3$As$_2$, as well as the $F$ $\approx$ 7 T pocket in LaRhIn$_5$.

cond-mat.str-el

Temperature dependence of quantum oscillations from non-parabolic dispersions

The phase offset of quantum oscillations is commonly used to experimentally diagnose topologically non-trivial Fermi surfaces. This methodology, however, is inconclusive for spin-orbit-coupled metals where $\pi$-phase-shifts can also arise from non-topological origins. Here, we show that the linear dispersion in topological metals leads to a $T^2$-temperature correction to the oscillation frequency that is absent for parabolic dispersions. We confirm this effect experimentally in the Dirac semi-metal Cd$_3$As$_2$ and the multiband Dirac metal LaRhIn$_5$. Both materials match a tuning-parameter-free theoretical prediction, emphasizing their unified origin. For topologically trivial Bi$_2$O$_2$Se, no frequency shift associated to linear bands is observed as expected. However, the $\pi$-phase shift in Bi$_2$O$_2$Se would lead to a false positive in a Landau-fan plot analysis. Our frequency-focused methodology does not require any input from ab-initio calculations, and hence is promising for identifying correlated topological materials.

cond-mat.str-el

Raman Spectra and Strain Effects in Bismuth Oxychalcogenides

A new type of two-dimensional layered semiconductor with weak electrostatic but not van der Waals interlayer interactions, Bi2O2Se, has been recently synthesized, which shown excellent air stability and ultrahigh carrier mobility. Herein, we combined theoretical and experimental approaches to study the Raman spectra of Bi2O2Se and related bismuth oxychalcogenides (Bi2O2Te and Bi2O2S). The experimental peaks were fully consistent with the calculated results, and were successfully assigned. Bi2O2S was predicted to have more Raman-active modes due to its lower symmetry. The shift of the predicted frequencies of Raman active modes was also found to get softened as the interlayer interaction decreases from bulk to monolayer Bi2O2Se and Bi2O2Te. To reveal the strain effects on the Raman shifts, a universal theoretical equation was established based on the symmetry of Bi2O2Se and Bi2O2Te. It was predicted that the doubly degenerate modes split under in-plane uniaxial/shear strains. Under a rotated uniaxial strain, the changes of Raman shifts are anisotropic for degenerate modes although Bi2O2Se and Bi2O2Te were usually regarded as isotropic systems similar to graphene. This implies a novel method to identify the crystallographic orientation from Raman spectra under strain. These results have important consequences for the incorporation of 2D Bismuth oxychalcogenides into nanoelectronic devices.

cond-mat.mes-hall