SearcharxivSearch

arXiv subjects

Yafei Li

Publications and source records attributed to Yafei Li.

At least 19 recordsLinked to original sources

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF

Reinforcement learning from human feedback (RLHF) with preference-based reward models often exhibits unstable training dynamics. A key contributing factor is that standard RLHF relies on a single sequence-level scalar reward, which is propagated to token-level policy updates and leaves credit assignment within a response inherently ambiguous. Recent work has attempted to address this issue by refining rewards into denser token-level supervision, often relying on the implicit assumption that finer-grained credit assignment improves optimization. We argue that this assumption is incomplete: when preference signals are noisy and only defined at the response level, overly fine-grained reward refinement can amplify reward uncertainty and destabilize learning. To address this problem, we propose a granularity-aware principle for hierarchical credit assignment, emphasizing stability-oriented reward design rather than maximal allocation precision. Under this principle, sentences serve as a natural intermediate granularity, balancing semantic coherence with robustness to token-level noise. Guided by this view, we introduce S2T-RLHF. This sentence-to-token reward decomposition framework first allocates sequence-level preference rewards across sentences and then applies bounded token-level refinement within each sentence, without reward-model retraining or token-level supervision. Experiments across multiple datasets and optimization settings show that S2T-RLHF improves training stability and robustness while maintaining competitive preference alignment.

cs.AI

Local Truncation Error-Guided Neural ODEs for Large Scale Traffic Forecasting

Spatiotemporal forecasting in physical systems, such as large-scale traffic networks, requires modeling a dual dynamic: continuous macroscopic rhythms and discrete, unpredictable microscopic shocks. While Neural Ordinary Differential Equations (ODEs) excel at capturing smooth evolution, their inherent Lipschitz continuity constraints inevitably cause severe over-smoothing when confronting abrupt anomalies. Recent physics-informed methods attempt to bypass this by penalizing numerical integration errors to enforce manifold smoothness. However, we mathematically reveal that such rigid regularization inherently triggers gradient conflicts and ``attention collapse,'' stripping the model of its sensitivity to anomalies. To resolve this continuity-shock dilemma, we propose Local Truncation Error-Guided Neural ODEs (LTE-ODE). Rather than treating numerical error as a nuisance to be eliminated, we innovatively repurpose the Local Truncation Error (LTE) as an unsupervised forward inductive bias. By mapping the LTE into a dynamic spatial attention mask, our architecture gracefully preserves high-precision continuous ODE evolution in stable regions, while adaptively triggering a discrete compensation branch exclusively at shock points. Trained purely end-to-end without manifold penalties, LTE-ODE achieves state-of-the-art performance on multiple large-scale benchmarks, exhibiting exceptional robustness against highly non-linear fluctuations. Furthermore, our ablation on integration steps demonstrates high deployment flexibility, allowing the model to seamlessly adapt to varying hardware memory constraints in real-world applications.

cs.LG

Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering

Despite the success of Vision-Language Models (VLMs), misleading charts remain a significant challenge due to their deceptive visual structures and distorted data representations. We present ChartCynics, an agentic dual-path framework designed to unmask visual deception via a "skeptical" reasoning paradigm. Unlike holistic models, ChartCynics decouples perception from verification: a Diagnostic Vision Path captures structural anomalies (e.g., inverted axes) through strategic ROI cropping, while an OCR-Driven Data Path ensures numerical grounding. To resolve cross-modal conflicts, we introduce an Agentic Summarizer optimized via a two-stage protocol: Oracle-Informed SFT for reasoning distillation and Deception-Aware GRPO for adversarial alignment. This pipeline effectively penalizes visual traps and enforces logical consistency. Evaluations on two benchmarks show that ChartCynics achieves 74.43% and 64.55% accuracy, providing an absolute performance boost of ~29% over the Qwen3-VL-8B backbone, outperforming state-of-the-art proprietary models. Our results demonstrate that specialized agentic workflows can grant smaller open-source models superior robustness, establishing a new foundation for trustworthy chart interpretation.

cs.CV

MedKGI: Iterative Differential Diagnosis with Medical Knowledge Graphs and Information-Guided Inquiring

Recent advancements in Large Language Models (LLMs) have demonstrated significant promise in clinical diagnosis. However, current models struggle to emulate the iterative, diagnostic hypothesis-driven reasoning of real clinical scenarios. Specifically, current LLMs suffer from three critical limitations: (1) generating hallucinated medical content due to weak grounding in verified knowledge, (2) asking redundant or inefficient questions rather than discriminative ones that hinder diagnostic progress, and (3) losing coherence over multi-turn dialogues, leading to contradictory or inconsistent conclusions. To address these challenges, we propose MedKGI, a diagnostic framework grounded in clinical practices. MedKGI integrates a medical knowledge graph (KG) to constrain reasoning to validated medical ontologies, selects questions based on information gain to maximize diagnostic efficiency, and adopts an OSCE-format structured state to maintain consistent evidence tracking across turns. Experiments on clinical benchmarks show that MedKGI outperforms strong LLM baselines in both diagnostic accuracy and inquiry efficiency, improving dialogue efficiency by 30% on average while maintaining state-of-the-art accuracy.

cs.CL

Semantic Encryption: Secure and Effective Interaction with Cloud-based Large Language Models via Semantic Transformation

The increasing adoption of Cloud-based Large Language Models (CLLMs) has raised significant concerns regarding data privacy during user interactions. While existing approaches primarily focus on encrypting sensitive information, they often overlook the logical structure of user inputs. This oversight can lead to reduced data utility and degraded performance of CLLMs. To address these limitations and enable secure yet effective interactions, we propose Semantic Encryption (SE)-a plug-and-play framework designed to preserve both privacy and utility. SE consists of two key components: Semantic Encoding and Semantic Decoding. In the encoding phase, a lightweight local model transforms the original user input into an alternative semantic context that maintains the original intent and logical structure while obfuscating sensitive information. This transformed input is then processed by the CLLM, which generates a response based on the transformed semantic context. To maintain a seamless user experience, the decoding phase will reconstruct the CLLM's response back into the original semantic context by referencing the locally stored user input. Extensive experimental evaluations demonstrate that SE effectively protects data privacy without compromising data utility or user experience, offering a practical solution for secure interaction with CLLMs. Particularly, the proposed SE demonstrates a significant improvement over the state-of-the-art InferDPT, surpassing it across various evaluated metrics and datasets.

cs.CR

UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text Synthesis

Text-to-image generation has greatly advanced content creation, yet accurately rendering visual text remains a key challenge due to blurred glyphs, semantic drift, and limited style control. Existing methods often rely on pre-rendered glyph images as conditions, but these struggle to retain original font styles and color cues, necessitating complex multi-branch designs that increase model overhead and reduce flexibility. To address these issues, we propose a segmentation-guided framework that uses pixel-level visual text masks -- rich in glyph shape, color, and spatial detail -- as unified conditional inputs. Our method introduces two core components: (1) a fine-tuned bilingual segmentation model for precise text mask extraction, and (2) a streamlined diffusion model augmented with adaptive glyph conditioning and a region-specific loss to preserve textual fidelity in both content and style. Our approach achieves state-of-the-art performance on the AnyText benchmark, significantly surpassing prior methods in both Chinese and English settings. To enable more rigorous evaluation, we also introduce two new benchmarks: GlyphMM-benchmark for testing layout and glyph consistency in complex typesetting, and MiniText-benchmark for assessing generation quality in small-scale text regions. Experimental results show that our model outperforms existing methods by a large margin in both scenarios, particularly excelling at small text rendering and complex layout preservation, validating its strong generalization and deployment readiness.

cs.CV

A Line Graph-Based Framework for Identifying Optimal Routing Paths in Decentralized Exchanges

Decentralized exchanges, such as those employing constant product market makers (CPMMs) like Uniswap V2, play a crucial role in the blockchain ecosystem by enabling peer-to-peer token swaps without intermediaries. Despite the increasing volume of transactions, there remains limited research on identifying optimal trading paths across multiple DEXs. This paper presents a novel line-graph-based algorithm (LG) designed to efficiently discover profitable trading routes within DEX environments. We benchmark LG against the widely adopted Depth-First Search (DFS) algorithm under a linear routing scenario, encompassing platforms such as Uniswap, SushiSwap, and PancakeSwap. Experimental results demonstrate that LG consistently identifies trading paths that are as profitable as, or more profitable than, those found by DFS, while incurring comparable gas costs. Evaluations on Uniswap V2 token graphs across two temporal snapshots further validate LG's performance. Although LG exhibits exponential runtime growth with respect to graph size in empirical tests, it remains viable for practical, real-world use cases. Our findings underscore the potential of the LG algorithm for industrial adoption, offering tangible benefits to traders and market participants in the DeFi space.

q-fin.CP

KKA: Improving Vision Anomaly Detection through Anomaly-related Knowledge from Large Language Models

Vision anomaly detection, particularly in unsupervised settings, often struggles to distinguish between normal samples and anomalies due to the wide variability in anomalies. Recently, an increasing number of studies have focused on generating anomalies to help detectors learn more effective boundaries between normal samples and anomalies. However, as the generated anomalies are often derived from random factors, they frequently lack realism. Additionally, randomly generated anomalies typically offer limited support in constructing effective boundaries, as most differ substantially from normal samples and lie far from the boundary. To address these challenges, we propose Key Knowledge Augmentation (KKA), a method that extracts anomaly-related knowledge from large language models (LLMs). More specifically, KKA leverages the extensive prior knowledge of LLMs to generate meaningful anomalies based on normal samples. Then, KKA classifies the generated anomalies as easy anomalies and hard anomalies according to their similarity to normal samples. Easy anomalies exhibit significant differences from normal samples, whereas hard anomalies closely resemble normal samples. KKA iteratively updates the generated anomalies, and gradually increasing the proportion of hard anomalies to enable the detector to learn a more effective boundary. Experimental results show that the proposed method significantly improves the performance of various vision anomaly detectors while maintaining low generation costs. The code for CMG can be found at https://github.com/Anfeather/KKA.

cs.CV

Radial symmetry and sharp asymptotic behaviors of nonnegative solutions to weighted doubly $D^{1,p}$-critical quasi-linear nonlocal elliptic equations with Hardy potential

In this paper, we mainly consider nonnegative weak solutions $u\in D^{1,p}(\R^{N})$ to the doubly $D^{1,p}(\R^{N})$-critical nonlocal quasi-linear Schr\"{o}dinger-Hartree equation: \begin{align*} -\Delta_p u- \mu \frac{u^{p-1}}{|x|^p}=\left(|x|^{-2p}\ast |u|^{p}\right)|u|^{p-2}u \qquad &\mbox{in} \,\, \mathbb{R}^N, \end{align*} where $N\geq3$, $0\leq\mu< \bar{\mu}:=\left( (N-p)/p \right)^p$ and $1 0$, due to appearance of the Hardy potential, the equation has singularity at $0\in\mathbb{R}^{N}$ and hence is not translation invariant, so sharp asymptotic estimates near the origin must be involved. First, we establish regularity and the sharp estimates on asymptotic behaviors near the origin and the infinity for any positive solution $u\in D^{1,p}(\R^{N})$ (and $|\nabla u|$) to more general equation $-\triangle_p u - \mu \frac{1}{|x|^p}u^{p-1}=V(x)\frac{1}{|x|^s}u^{p-1}$ with $N\geq2$, $0\leq\mu< \bar{\mu}$, $1<p<N$, $0\leq s < p$ and $0\leq V(x)\in L^\frac{N}{p-s}(\R^N)$. Then, as a consequence, we can apply the method of moving planes to prove that all the nontrivial nonnegative solutions in $D^{1,p}(\R^{N})$ are radially symmetric and strictly radially decreasing about the origin $0\in\mathbb{R}^{N}$. The sharp asymptotic estimates and radial symmetry for more general weighted doubly $D^{1,p}$-critical nonlocal quasi-linear equations were also derived. Our results extend the results in \cite{DLL} from the special case $\mu=0$ to general cases $0\leq\mu<\bar{\mu}$.

math.AP

Neighborhood and Global Perturbations Supported SAM in Federated Learning: From Local Tweaks To Global Awareness

Federated Learning (FL) can be coordinated under the orchestration of a central server to collaboratively build a privacy-preserving model without the need for data exchange. However, participant data heterogeneity leads to local optima divergence, subsequently affecting convergence outcomes. Recent research has focused on global sharpness-aware minimization (SAM) and dynamic regularization techniques to enhance consistency between global and local generalization and optimization objectives. Nonetheless, the estimation of global SAM introduces additional computational and memory overhead, while dynamic regularization suffers from bias in the local and global dual variables due to training isolation. In this paper, we propose a novel FL algorithm, FedTOGA, designed to consider optimization and generalization objectives while maintaining minimal uplink communication overhead. By linking local perturbations to global updates, global generalization consistency is improved. Additionally, global updates are used to correct local dynamic regularizers, reducing dual variables bias and enhancing optimization consistency. Global updates are passively received by clients, reducing overhead. We also propose neighborhood perturbation to approximate local perturbation, analyzing its strengths and limitations. Theoretical analysis shows FedTOGA achieves faster convergence $O(1/T)$ under non-convex functions. Empirical studies demonstrate that FedTOGA outperforms state-of-the-art algorithms, with a 1\% accuracy increase and 30\% faster convergence, achieving state-of-the-art.

cs.LG

Radial symmetry and sharp asymptotic behaviors of nonnegative solutions to $D^{1,p}$-critical quasi-linear static Schrödinger-Hartree equation involving $p$-Laplacian $-Δ_{p}$

In this paper, we mainly consider nonnegative weak solution to the $D^{1,p}(\R^{N})$-critical quasi-linear static Schrödinger-Hartree equation with $p$-Laplacian $-Δ_{p}$ and nonlocal nonlinearity: \begin{align*} -Δ_p u =\left(|x|^{-2p}\ast |u|^{p}\right)|u|^{p-2}u \qquad &\mbox{in} \,\, \mathbb{R}^N, \end{align*} where $1<p<\frac{N}{2}$, $N\geq3$ and $u\in D^{1,p}(\R^N)$. Being different to the $D^{1,p}(\R^{N})$-critical local nonlinear term $u^{p^{\star}-1}$ with $p^{\star}:=\frac{Np}{N-p}$ investigated in \cite{CFR,LDSMLMSB,GV,Ou,BS16,VJ16} etc., since the nonlocal convolution $|x|^{-2p}*u^p$ appears in the Hartree type nonlinearity, it is impossible for us to use the scaling arguments and the Doubling Lemma as in \cite{VJ16} to get preliminary estimates on upper bounds of asymptotic behaviors for any positive solutions $u$. Moreover, it is also quite difficult to obtain the boundedness of the quasi-norm $\|u \|_{L^{s,\infty}(\R^N)}$ and hence derive the sharp estimates on upper bounds of asymptotic behaviors from the preliminary estimates as in \cite{VJ16}. Fortunately, by showing a better preliminary estimates on upper bounds of asymptotic behaviors through the De Giorgi-Moser-Nash iteration method and combining the result from \cite{XCL}, we are able to overcome these difficulties and establish regularity and the sharp estimates on both upper and lower bounds of asymptotic behaviors for any positive solution $u$ to more general equation $-Δ_p u=V(x)u^{p-1}$ with $V\in L^{\frac{N}{p}}(\mathbb{R}^{N})$. Then, by using the arguments from \cite{BS16,VJ16}, we can deduce the sharp estimates on both upper and lower bounds for the decay rate of $|\nabla u|$. Finally, as a consequence, we can apply the method of moving planes to prove that all the nontrivial nonnegative solutions are radially symmetric and strictly decreasing about some point $x_0\in\R^N$.

math.AP

ChatGraph: Chat with Your Graphs

Graph analysis is fundamental in real-world applications. Traditional approaches rely on SPARQL-like languages or clicking-and-dragging interfaces to interact with graph data. However, these methods either require users to possess high programming skills or support only a limited range of graph analysis functionalities. To address the limitations, we propose a large language model (LLM)-based framework called ChatGraph. With ChatGraph, users can interact with graphs through natural language, making it easier to use and more flexible than traditional approaches. The core of ChatGraph lies in generating chains of graph analysis APIs based on the understanding of the texts and graphs inputted in the user prompts. To achieve this, ChatGraph consists of three main modules: an API retrieval module that searches for relevant APIs, a graph-aware LLM module that enables the LLM to comprehend graphs, and an API chain-oriented finetuning module that guides the LLM in generating API chains.

cs.AI

Room-Temperature Ferromagnetism in Fe-doped SnSe Bulk Single Crystalline Semiconductor

The quest for pragmatic room-temperature (RT) magnetic semiconductors (MSs) with a suitable bandgap constitutes one of the contemporary opportunities to be exploited. This may provide a materials platform for to bring new-generation ideal information device technologies into real-world applications where the otherwise conventionally separately utilized charge and spin are simultaneously exploited. Here we present RT ferromagnetism in an Fe-doped SnSe (Fe:SnSe) van der Waals (vdW) single crystalline ferromagnetic semiconductor (FMS) with a semiconducting bandgap of ~1.19 eV (comparable to those of Si and GaAs). The synthesized Fe:SnSe single crystals feature a dilute Fe content of less than 1.0 at%, a Curie temperature of ~310 K, a layered vdW structure identical to that of pristine SnSe, and the absence of in-gap defect states. The Fe:SnSe vdW diluted magnetic semiconductor (DMS) single crystals are grown using a simple temperature-gradient melt-growth process, in which the magnetic Fe atom doping is realized uniquely using FeI2 as the dopant precursor whose melting point is low with respect to crystal growth, and which in principle possesses industrially unlimited scalability. Our work adds a new member in the family of long-searching RT magnetic semiconductors, and may establish a generalized strategy for large-volume production of related DMSs.

cond-mat.mtrl-sci

Ground state solutions to a coupled nonlinear logarithmic Hartree system

In this paper, we study the following coupled nonlinear logarithmic Hartree system \begin{align*} \left\{ \displaystyle \begin{array}{ll} \displaystyle -Δu+ λ_1 u =μ_1\left( -\frac{1}{2π}\ln(|x|) \ast u^2 \right)u+β\left( -\frac{1}{2π}\ln(|x|) \ast v^2 \right)u, & x \in ~ \mathbb R^2, \vspace{.4cm}\\ -Δv+ λ_2 v =μ_2\left( -\frac{1}{2π}\ln(|x|) \ast v^2 \right)v +β\left( -\frac{1}{2π}\ln(|x|) \ast u^2 \right)v, & x \in ~ \mathbb R^2, \end{array} \right.\hspace{1cm} \end{align*} where $β, μ_i, λ_i \ (i=1,2)$ are positive constants, $\ast$ denotes the convolution in $\mathbb R^2$. By considering the constraint minimum problem on the Nehari manifold, we prove the existence of ground state solutions for $β>0$ large enough. Moreover, we also show that every positive solution is radially symmetric and decays exponentially.

math.AP

Realization of one-dimensional electronic flat bands in an untwisted moire superlattice

Two-dimensional electronic flat bands and their induced correlated electronic interactions have been discovered, probed, and tuned in interlayer regions of hexagonally shaped van der Waals moire superlattices. Fabrication of anisotropic one-dimensional correlated bands by moire interference of 2D, however, remains a challenge. Here, we report an experimental discovery of 1D electronic flat bands near the Fermi level in an anisotropic rectangular moire superlattice composed of in situ grown, vdW stacked two-atomic-layer thick Bi(110) well-aligned on a SnSe(001) substrate. The epitaxial lattice mismatch between the aligned Bi and SnSe zigzag atomic chains causes strong three-dimensional anisotropic atomic relaxations with associated one-dimensional out-of- and in-plane strain distributions that are expressed in electronic bands of the Bi(110) layer, which are characterized jointly by scanning probe microscopy and density functional theory. At the regions of the strongest out-of-plane shear strain, a series of 1D flat bands near the Fermi level are experimentally observed and defined in our calculations. We establish that 1D flat bands can arise in moiré superlattices in absence of the relative layer twist, but solely through the lattice strain. We generalize the strategy of utilizing strain in lattice mismatched rectangular hetero-bilayers for engineering correlated anisotropic electronic bands.

cond-mat.mes-hall

Two-dimensional Dirac nodal-line semimetal protected by symmetry

Dirac nodal line semimetals (DNLSs) host relativistic quasiparticles in their one-dimensional (1D) Dirac nodal line (DNL) bands that are protected by certain crystalline symmetries. Their novel low-energy fermion quasiparticle excitations and transport properties invite studies of relativistic physics in the solid state where their linearly dispersing Dirac bands cross at continuous lines with four-fold degeneracy. In materials studied up to now, the four-fold degeneracy, however, has been vulnerable to suppression by the ubiquitous spin-orbit coupling (SOC). Despite the current effort to discover 3D DNLSs that are robust to SOC by theory, positive experimental evidence is yet to emerge. In 2D DNLSs, because of the decreased total density of states as compared with their 3D counterparts, it is anticipated that their physical properties would be dominated by the electronic states defined by the DNL. It has been even more challenging, however, to discover robust 2D DNLSs against SOC because of their lowered symmetry; no such materials have yet been predicted by theory. By combining molecular beam epitaxy growth, STM, nc-AFM characterisation, with DFT calculations and space group theory analysis, here we reveal a novel class of 2D crystalline DNLSs that host the exact symmetry that protects them against SOC. The discovered quantum material is a brick phase 3-AL Bi(110), whose symmetry protection and thermal stability are imparted by the compressive vdW epitaxial growth on black phosphorus substrates. The BP substrate templates the growth of 3-AL Bi(110) nano-islands in a non-symmorphic space group structure. This crystalline symmetry protects the DNL electronic phase against SOC independent of any orbital or elemental factors. We theoretically establish that this intrinsic symmetry imparts a general, robust protection of DNL in a series of isostructural 2D quantum materials.

cond-mat.mtrl-sci

ACSEE: Antagonistic Crowd Simulation Model with Emotional Contagion and Evolutionary Game Theory

Antagonistic crowd behaviors are often observed in cases of serious conflict. Antagonistic emotions, which is the typical psychological state of agents in different roles (i.e. cops, activists, and civilians) in crowd violent scenes, and the way they spread through contagion in a crowd are important causes of crowd antagonistic behaviors. Moreover, games, which refers to the interaction between opposing groups adopting different strategies to obtain higher benefits and less casualties, determine the level of crowd violence. We present an antagonistic crowd simulation model, ACSEE, which is integrated with antagonistic emotional contagion and evolutionary game theories. Our approach models the antagonistic emotions between agents in different roles using two components: mental emotion and external emotion. We combine enhanced susceptible-infectious-susceptible (SIS) and game approaches to evaluate the role of antagonistic emotional contagion in crowd violence. Our evolutionary game theoretic approach incorporates antagonistic emotional contagion through deterrent force, which is modelled by a mixture of emotional forces and physical forces defeating the opponents. Antagonistic emotional contagion and evolutionary game theories influence each other to determine antagonistic crowd behaviors. We evaluate our approach on real-world scenarios consisting of different kinds of agents. We also compare the simulated crowd behaviors with real-world crowd videos and use our approach to predict the trends of crowd movements in violence incidents. We investigate the impact of various factors (number of agents, emotion, strategy, etc.) on the outcome of crowd violence. We present results from user studies suggesting that our model can simulate antagonistic crowd behaviors similar to those seen in real-world scenarios.

cs.MA

Stylize Aesthetic QR Code

With the continued proliferation of smart mobile devices, Quick Response (QR) code has become one of the most-used types of two-dimensional code in the world. Aiming at beautifying the visual-unpleasant appearance of QR codes, existing works have developed a series of techniques. However, these works still leave much to be desired, such as personalization, artistry, and robustness. To address these issues, in this paper, we propose a novel type of aesthetic QR codes, SEE (Stylize aEsthEtic) QR code, and a three-stage approach to automatically produce such robust style-oriented codes. Specifically, in the first stage, we propose a method to generate an optimized baseline aesthetic QR code, which reduces the visual contrast between the noise-like black/white modules and the blended image. In the second stage, to obtain art style QR code, we tailor an appropriate neural style transformation network to endow the baseline aesthetic QR code with artistic elements. In the third stage, we design an error-correction mechanism by balancing two competing terms, visual quality and readability, to ensure the performance robust. Extensive experiments demonstrate that SEE QR code has high quality in terms of both visual appearance and robustness, and also offers a greater variety of personalized choices to users.

cs.MM