SearcharxivSearch

arXiv subjects

Wenpeng Zhu

Publications and source records attributed to Wenpeng Zhu.

3 recordsLinked to original sources

Reconsidering Overthinking: Penalizing Internal and External Redundancy in CoT Reasoning

Large reasoning models (LRMs) often exhibit overthinking, producing verbose Chain-of-Thought (CoT) traces that increase inference cost and obscure the underlying reasoning process. Existing CoT compression methods mainly rely on global length rewards, which conflate necessary intermediate reasoning with redundant text and may therefore compromise reasoning fidelity. This paper revisits overthinking from a semantic-efficiency perspective and decomposes CoT redundancy into two distinct forms: internal redundancy, defined as informational stagnation before the first correct answer, and external redundancy, defined as superfluous continuation after the first correct answer. Based on this decomposition, we propose a dual-penalty reinforcement learning framework that separately optimizes reasoning progress and termination behavior. Specifically, a sliding-window semantic similarity metric penalizes low-progress reasoning segments, while a normalized external-redundancy metric discourages post-answer continuation. Experiments on GSM8K, MATH500, and AIME24 across different model scales show that our method reduces average reasoning length by 41.3% on the 1.5B model and 40.1% on the 7B model, while preserving competitive accuracy and achieving the best overall accuracy-efficiency score among evaluated baselines. The learned compression behavior further transfers to out-of-domain reasoning tasks, including GPQA and LiveCodeBench. More importantly, our analysis reveals a clear asymmetry between the two redundancy types: external redundancy can be largely removed with little performance loss, whereas internal redundancy compression follows a sensitive accuracy-efficiency trade-off. These results suggest that effective CoT compression should optimize semantic efficiency rather than sequence length alone, offering a principled route toward more concise, efficient, and interpretable LRMs.

cs.AI

MemoryRewardBench: Benchmarking Reward Models for Long-Term Memory Management in Large Language Models

Existing works increasingly adopt memory-centric mechanisms to process long contexts in a segment manner, and effective memory management is one of the key capabilities that enables large language models to effectively propagate information across the entire sequence. Therefore, leveraging reward models (RMs) to automatically and reliably evaluate memory quality is critical. In this work, we introduce MemoryRewardBench, the first benchmark to systematically study the ability of RMs to evaluate long-term memory management processes. MemoryRewardBench covers both long-context comprehension and long-form generation tasks, featuring 10 distinct settings with different memory management patterns, with context length ranging from 8K to 128K tokens. Evaluations on 13 cutting-edge RMs indicate a diminishing performance gap between open-source and proprietary models, with newer-generation models consistently outperforming their predecessors regardless of parameter count. We further expose the capabilities and fundamental limitations of current RMs in evaluating LLM memory management across diverse settings.

cs.CL

Creation of independently controllable and long lifetime polar skyrmion textures in ferroelectric-metallic heterostructures

Topological textures like vortices, labyrinths and skyrmions formed in ferroic materials have attracted extensive interests during the past decade for their fundamental physics, intriguing topology, and technological prospects. So far, polar skyrmions remain scarce in ferroelectrics as they require a delicate balance between various dipolar interactions. Here, we report that PbTiO3 thin films in a metallic contact undergo a topological phase transition and stabilize a broad family of skyrmion-like textures (e.g., skyrmion bubbles, multiple π-twist target skyrmions, and skyrmion bags) with independent controllability, analogous to those reported in magnetic systems. Weakly-interacted skyrmion arrays with a density over 300 Gb/inch2 are successfully written, erased and read-out by local electrical and mechanical stimuli of a scanning probe. Interestingly, in contrast to the relatively short lifetime <20 hours of the skyrmion bubbles, the multiple π-twist target skyrmions and skyrmion bags show topology-enhanced stability with lifetime over two weeks. Experimental and theoretical analysis implies the heterostructures carry electric Dzyaloshinskii-Moriya interaction mediated by oxygen octahedral tiltings. Our results demonstrate ferroelectric-metallic heterostructures as fertile playground for topological states and emergent phenomena.

cond-mat.mtrl-sci