Searcharxiv⌕ Search

arXiv subjects

Fei Sun

Publications and source records attributed to Fei Sun.

At least 55 records · Page 3Linked to original sources

The Evolution of Thought: Tracking LLM Overthinking via Reasoning Dynamics Analysis

Test-time scaling via explicit reasoning trajectories significantly boosts large language model (LLM) performance but often triggers overthinking. To explore this, we analyze reasoning through two lenses: Reasoning Length Dynamics, which reveals a compensatory trade-off between thinking and answer content length that eventually leads to thinking redundancy, and Reasoning Semantic Dynamics, which identifies semantic convergence and repetitive oscillations. These dynamics uncover an instance-specific Reasoning Completion Point (RCP), beyond which computation continues without further performance gain. Since the RCP varies across instances, we propose a Reasoning Completion Point Detector (RCPD), an inference-time early-exit method that identifies the RCP by monitoring the rank dynamics of termination tokens (e.g., ). Across AIME and GPQA benchmarks using Qwen3 and DeepSeek-R1, RCPD reduces token usage by up to 44% while preserving accuracy, offering a principled approach to efficient test-time scaling.

cs.CL↗

Soft-hard factorization of heavy-quark transport in QCD matter at finite chemical potential

We calculate the collisional energy loss and momentum diffusion coefficients of heavy quarks traversing a hot and dense QCD medium at finite quark chemical potential, $μ\neq0$. The analysis is performed within an extended soft-hard factorization model (SHFM) that consistently incorporates the $μ$-dependence of the Debye screening mass $M_D(μ)$ and of the fermionic thermal distribution functions. Both the energy loss and the diffusion coefficients are found to increase with $μ$, with the enhancement being most pronounced at low temperatures where the chemical potential effects dominate the medium response. To elucidate the origin of this dependence, we derive analytic high-energy approximations in which the leading $μ$-corrections appear as logarithmic terms: a soft logarithm $\simμ^{2}\ln(|t^{*}|/M_{D}^{2})$ from $t$-channel scattering off thermal gluonic excitations, and a hard logarithm $\simμ^{2}\ln(E_{1}T/|t^{*}|)$ from scattering off thermal quarks. In the complete result the dependence on the intermediate separation scale $t^{\ast}$ cancels, as required. We also confirm the expected mass hierarchy $-dE/dz(charm)<-dE/dz(bottom)$ at fixed velocity. Our findings demonstrate that finite chemical potential plays a significant role in heavy-quark transport and must be included in theoretical descriptions of heavy-flavor dynamics in baryon-rich environments, such as those probed in the RHIC Beam Energy Scan, and at FAIR and NICA.

hep-ph↗

The 2nd Workshop on Human-Centered Recommender Systems

Recommender systems shape how people discover information, form opinions, and connect with society. Yet, as their influence grows, traditional metrics, e.g., accuracy, clicks, and engagement, no longer capture what truly matters to humans. The workshop on Human-Centered Recommender Systems (HCRS) calls for a paradigm shift from optimizing engagement toward designing systems that truly understand, involve, and benefit people. It brings together researchers in recommender systems, human-computer interaction, AI safety, and social computing to explore how human values, e.g., trust, safety, fairness, transparency, and well-being, can be integrated into recommendation processes. Centered around three thematic axes-Human Understanding, Human Involvement, and Human Impact-HCRS features keynotes, panels, and papers covering topics from LLM-based interactive recommenders to societal welfare optimization. By fostering interdisciplinary collaboration, HCRS aims to shape the next decade of responsible and human-aligned recommendation research.

cs.IR↗

A Survey on Unlearning in Large Language Models

Large Language Models (LLMs) demonstrate remarkable capabilities, but their training on massive corpora poses significant risks from memorized sensitive information. To mitigate these issues and align with legal standards, unlearning has emerged as a critical technique to selectively erase specific knowledge from LLMs without compromising their overall performance. This survey provides a systematic review of over 180 papers on LLM unlearning published since 2021. First, it introduces a novel taxonomy that categorizes unlearning methods based on the phase in the LLM pipeline of the intervention. This framework further distinguishes between parameter modification and parameter selection strategies, thus enabling deeper insights and more informed comparative analysis. Second, it offers a multidimensional analysis of evaluation paradigms. For datasets, we compare 18 existing benchmarks from the perspectives of task format, content, and experimental paradigms to offer actionable guidance. For metrics, we move beyond mere enumeration by dividing knowledge memorization metrics into 10 categories to analyze their advantages and applicability, while also reviewing metrics for model utility, robustness, and efficiency. By discussing current challenges and future directions, this survey aims to advance the field of LLM unlearning and the development of secure AI systems.

cs.CL↗

Detecting Stealthy Backdoor Samples based on Intra-class Distance for Large Language Models

Stealthy data poisoning during fine-tuning can backdoor large language models (LLMs), threatening downstream safety. Existing detectors either use classifier-style probability signals--ill-suited to generation--or rely on rewriting, which can degrade quality and even introduce new triggers. We address the practical need to efficiently remove poisoned examples before or during fine-tuning. We observe a robust signal in the response space: after applying TF-IDF to model responses, poisoned examples form compact clusters (driven by consistent malicious outputs), while clean examples remain dispersed. We leverage this with RFTC--Reference-Filtration + TF-IDF Clustering. RFTC first compares each example's response with that of a reference model and flags those with large deviations as suspicious; it then performs TF-IDF clustering on the suspicious set and identifies true poisoned examples using intra-class distance. On two machine translation datasets and one QA dataset, RFTC outperforms prior detectors in both detection accuracy and the downstream performance of the fine-tuned models. Ablations with different reference models further validate the effectiveness and robustness of Reference-Filtration.

cs.CL↗

Interactive Recommendation Agent with Active User Commands

Traditional recommender systems rely on passive feedback mechanisms that limit users to simple choices such as like and dislike. However, these coarse-grained signals fail to capture users' nuanced behavior motivations and intentions. In turn, current systems cannot also distinguish which specific item attributes drive user satisfaction or dissatisfaction, resulting in inaccurate preference modeling. These fundamental limitations create a persistent gap between user intentions and system interpretations, ultimately undermining user satisfaction and harming system effectiveness. To address these limitations, we introduce the Interactive Recommendation Feed (IRF), a pioneering paradigm that enables natural language commands within mainstream recommendation feeds. Unlike traditional systems that confine users to passive implicit behavioral influence, IRF empowers active explicit control over recommendation policies through real-time linguistic commands. To support this paradigm, we develop RecBot, a dual-agent architecture where a Parser Agent transforms linguistic expressions into structured preferences and a Planner Agent dynamically orchestrates adaptive tool chains for on-the-fly policy adjustment. To enable practical deployment, we employ simulation-augmented knowledge distillation to achieve efficient performance while maintaining strong reasoning capabilities. Through extensive offline and long-term online experiments, RecBot shows significant improvements in both user satisfaction and business outcomes.

cs.IR↗

Too Consistent to Detect: A Study of Self-Consistent Errors in LLMs

As large language models (LLMs) often generate plausible but incorrect content, error detection has become increasingly critical to ensure truthfulness. However, existing detection methods often overlook a critical problem we term as self-consistent error, where LLMs repeatedly generate the same incorrect response across multiple stochastic samples. This work formally defines self-consistent errors and evaluates mainstream detection methods on them. Our investigation reveals two key findings: (1) Unlike inconsistent errors, whose frequency diminishes significantly as the LLM scale increases, the frequency of self-consistent errors remains stable or even increases. (2) All four types of detection methods significantly struggle to detect self-consistent errors. These findings reveal critical limitations in current detection methods and underscore the need for improvement. Motivated by the observation that self-consistent errors often differ across LLMs, we propose a simple but effective cross-model probe method that fuses hidden state evidence from an external verifier LLM. Our method significantly enhances performance on self-consistent errors across three LLM families.

cs.CL↗

Reinforced Lifelong Editing for Language Models

Large language models (LLMs) acquire information from pre-training corpora, but their stored knowledge can become inaccurate or outdated over time. Model editing addresses this challenge by modifying model parameters without retraining, and prevalent approaches leverage hypernetworks to generate these parameter updates. However, they face significant challenges in lifelong editing due to their incompatibility with LLM parameters that dynamically change during the editing process. To address this, we observed that hypernetwork-based lifelong editing aligns with reinforcement learning modeling and proposed RLEdit, an RL-based editing method. By treating editing losses as rewards and optimizing hypernetwork parameters at the full knowledge sequence level, we enable it to precisely capture LLM changes and generate appropriate parameter updates. Our extensive empirical evaluation across several LLMs demonstrates that RLEdit outperforms existing methods in lifelong editing with superior effectiveness and efficiency, achieving a 59.24% improvement while requiring only 2.11% of the time compared to most approaches. Our code is available at: https://github.com/zhrli324/RLEdit.

cs.CL↗

Broadband Simultaneous Beam Steering and Compressing Device Based on Subwavelength Protrusion Metallic Tunnels

Beam steering and beamwidth compressing play a role in steering the beam and narrowing its half-power beamwidth, respectively, which are both widely applied in extending the effective operational range of 6G communications, IoT devices, and antenna systems. However, research on wave manipulation devices capable of simultaneously achieving both functionalities remains limited, despite their great potential for system miniaturization and functional integration. In this study, we design and realize a broadband device capable of simultaneously steering and compressing the TM-polarized EM waves using subwavelength protrusion metallic tunnels. The underlying physical mechanisms are quantitatively explained through wave optics and optical surface transformation, indicating the size ratio between the incident and output surface governs both the steering angle and the compression ratio. Numerical simulations demonstrate its outstanding performance, achieving a maximum steering angle of 40° and a compression ratio of 0.4 across 3 to 12 GHz, with averaged energy transmittance above 80%. The experiments further validate its effectiveness by measuring the magnetic field distributions of the output beam at various frequencies. The excellent beam steering and compressing effects make the proposed device highly promising for next-generation multifunctional wave manipulation in advanced communication systems.

physics.optics↗

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions

Speculative decoding is a standard method for accelerating the inference speed of large language models. However, scaling it for production environments poses several engineering challenges, including efficiently implementing different operations (e.g., tree attention and multi-round speculative decoding) on GPU. In this paper, we detail the training and inference optimization techniques that we have implemented to enable EAGLE-based speculative decoding at a production scale for Llama models. With these changes, we achieve a new state-of-the-art inference latency for Llama models. For example, Llama4 Maverick decodes at a speed of about 4 ms per token (with a batch size of one) on 8 NVIDIA H100 GPUs, which is 10% faster than the previously best known method. Furthermore, for EAGLE-based speculative decoding, our optimizations enable us to achieve a speed-up for large batch sizes between 1.4x and 2.0x at production scale.

cs.CL↗

From Generation to Consumption: Personalized List Value Estimation for Re-ranking

Re-ranking is critical in recommender systems for optimizing the order of recommendation lists, thus improving user satisfaction and platform revenue. Most existing methods follow a generator-evaluator paradigm, where the evaluator estimates the overall value of each candidate list. However, they often ignore the fact that users may exit before consuming the full list, leading to a mismatch between estimated generation value and actual consumption value. To bridge this gap, we propose CAVE, a personalized Consumption-Aware list Value Estimation framework. CAVE formulates the list value as the expectation over sub-list values, weighted by user-specific exit probabilities at each position. The exit probability is decomposed into an interest-driven component and a stochastic component, the latter modeled via a Weibull distribution to capture random external factors such as fatigue. By jointly modeling sub-list values and user exit behavior, CAVE yields a more faithful estimate of actual list consumption value. We further contribute three large-scale real-world list-wise benchmarks from the Kuaishou platform, varying in size and user activity patterns. Extensive experiments on these benchmarks, two Amazon datasets, and online A/B testing on Kuaishou show that CAVE consistently outperforms strong baselines, highlighting the benefit of explicitly modeling user exits in re-ranking.

cs.IR↗

A Survey on AgentOps: Categorization, Challenges, and Future Directions

As the reasoning capabilities of Large Language Models (LLMs) continue to advance, LLM-based agent systems offer advantages in flexibility and interpretability over traditional systems, garnering increasing attention. However, despite the widespread research interest and industrial application of agent systems, these systems, like their traditional counterparts, frequently encounter anomalies. These anomalies lead to instability and insecurity, hindering their further development. Therefore, a comprehensive and systematic approach to the operation and maintenance of agent systems is urgently needed. Unfortunately, current research on the operations of agent systems is sparse. To address this gap, we have undertaken a survey on agent system operations with the aim of establishing a clear framework for the field, defining the challenges, and facilitating further development. Specifically, this paper begins by systematically defining anomalies within agent systems, categorizing them into intra-agent anomalies and inter-agent anomalies. Next, we introduce a novel and comprehensive operational framework for agent systems, dubbed Agent System Operations (AgentOps). We provide detailed definitions and explanations of its four key stages: monitoring, anomaly detection, root cause analysis, and resolution.

cs.AI↗

LLM4MEA: Data-free Model Extraction Attacks on Sequential Recommenders via Large Language Models

Recent studies have demonstrated the vulnerability of sequential recommender systems to Model Extraction Attacks (MEAs). MEAs collect responses from recommender systems to replicate their functionality, enabling unauthorized deployments and posing critical privacy and security risks. Black-box attacks in prior MEAs are ineffective at exposing recommender system vulnerabilities due to random sampling in data selection, which leads to misaligned synthetic and real-world distributions. To overcome this limitation, we propose LLM4MEA, a novel model extraction method that leverages Large Language Models (LLMs) as human-like rankers to generate data. It generates data through interactions between the LLM ranker and target recommender system. In each interaction, the LLM ranker analyzes historical interactions to understand user behavior, and selects items from recommendations with consistent preferences to extend the interaction history, which serves as training data for MEA. Extensive experiments demonstrate that LLM4MEA significantly outperforms existing approaches in data quality and attack performance, reducing the divergence between synthetic and real-world data by up to 64.98% and improving MEA performance by 44.82% on average. From a defensive perspective, we propose a simple yet effective defense strategy and identify key hyperparameters of recommender systems that can mitigate the risk of MEAs.

cs.IR↗

The Mirage of Model Editing: Revisiting Evaluation in the Wild

Despite near-perfect results reported in the literature, the effectiveness of model editing in real-world applications remains unclear. To bridge this gap, we introduce QAEdit, a new benchmark aligned with widely used question answering (QA) datasets, and WILD, a task-agnostic evaluation framework designed to better reflect real-world usage of model editing. Our single editing experiments show that current editing methods perform substantially worse than previously reported (38.5% vs. 96.8%). We demonstrate that it stems from issues in the synthetic evaluation practices of prior work. Among them, the most severe is the use of teacher forcing during testing, which leaks both content and length of the ground truth, leading to overestimated performance. Furthermore, we simulate practical deployment by sequential editing, revealing that current approaches fail drastically with only 1000 edits. This work calls for a shift in model editing research toward rigorous evaluation and the development of robust, scalable methods that can reliably update knowledge in LLMs for real-world use.

cs.CL↗

Thermal superscatterer: amplification of thermal scattering signatures for arbitrarily shaped thermal materials

The concept of superscattering is extended to the thermal field through the design of a thermal superscatterer based on transformation thermodynamics. A small thermal scatterer of arbitrary shape and conductivity is encapsulated with an engineered negative-conductivity shell, creating a composite that mimics the scattering signature of a significantly larger scatterer. The amplified signature can match either a conformal larger scatterer (preserving conductivity) or a geometry-transformed one (modified conductivity). The implementation employs a positive-conductivity shell integrated with active thermal metasurfaces, demonstrated through three representative examples: super-insulating thermal scattering, super-conducting thermal scattering, and equivalent thermally transparent effects. Experimental validation shows the fabricated superscatterer amplifies the thermal scattering signature of a small insulated circular region by nine times, effectively mimicking the scattering signature of a circular region with ninefold radius. This approach enables thermal signature manipulation beyond physical size constraints, with potential applications in thermal superabsorbers/supersources, thermal camouflage, and energy management.

physics.app-ph↗

Implementation of ultra-broadband optical null media via space-folding

Optical null medium (ONM) has garnered significant attention in electromagnetic wave manipulation. However, existing ONM implementations suffer from either narrow operational bandwidths or low efficiency. Here, we demonstrate an ultra-broadband ONM design that simultaneously addresses both challenges - achieving broad bandwidth while preserving perfect impedance matching with air for near-unity transmittance. The proposed space-folding ONM is realized by introducing precisely engineered folds into a metal channel array, creating an effective dispersion-free medium that enables independent phase control in each channel. The design incorporates optimized boundary layers implemented through gradually tapered folding structures, achieving perfect impedance matching with the surrounding medium. Beam bending effect and broadband beam focusing effect are experimentally verified using the proposed space-folding ONM. Due to its simple material requirements, broadband characteristics, and high transmittance, the proposed space-folding ONM shows potential for applications in electromagnetic camouflage, beam steering devices and ultra-compact microwave components.

physics.optics↗

Broadband source-surrounded cloak for on-chip antenna radiation pattern protection

As the frequency range of electromagnetic wave communication continues to expand and the integration of integrated circuits increases, electromagnetic waves emitted by on-chip antennas are prone to scattering from electronic components, which limits further improvements in integration and the protection of radiation patterns. Cloaks can be used to reduce electromagnetic scattering; however, they cannot achieve both broadband and omnidirectional effectiveness simultaneously. Moreover, their operating modes are typically designed for scenarios where the source is located outside the cloak, making it difficult to address this problem. In this work, we propose a dispersionless air-impedance-matched metamaterial over the 2-8 GHz bandwidth that achieves an adjustable effective refractive index ranging from 1.1 to 1.5, with transmittance maintained above 93%. Based on this metamaterial, we introduce a broadband source-surrounded cloak that can guide electromagnetic waves from a broadband source surrounded by the cloak in any propagation direction to bypass obstacles and reproduce the original wavefronts outside the cloak. Thereby protecting the radiation pattern from distortion due to scattering caused by obstacles. Our work demonstrates significant potential for enhancing the integration density of integrated circuits and improving the operational stability of communication systems.

physics.optics↗

Robust Recommender System: A Survey and Future Directions

With the rapid growth of information, recommender systems have become integral for providing personalized suggestions and overcoming information overload. However, their practical deployment often encounters ``dirty'' data, where noise or malicious information can lead to abnormal recommendations. Research on improving recommender systems' robustness against such dirty data has thus gained significant attention. This survey provides a comprehensive review of recent work on recommender systems' robustness. We first present a taxonomy to organize current techniques for withstanding malicious attacks and natural noise. We then explore state-of-the-art methods in each category, including fraudster detection, adversarial training, certifiable robust training for defending against malicious attacks, and regularization, purification, self-supervised learning for defending against malicious attacks. Additionally, we summarize evaluation metrics and commonly used datasets for assessing robustness. We discuss robustness across varying recommendation scenarios and its interplay with other properties like accuracy, interpretability, privacy, and fairness. Finally, we delve into open issues and future research directions in this emerging field. Our goal is to provide readers with a comprehensive understanding of robust recommender systems and to identify key pathways for future research and development. To facilitate ongoing exploration, we maintain a continuously updated GitHub repository with related research: https://github.com/Kaike-Zhang/Robust-Recommender-System.

cs.IR↗