Searcharxiv⌕ Search

arXiv subjects

Long Li

Publications and source records attributed to Long Li.

At least 37 records · Page 2Linked to original sources

DyJR: Preserving Diversity in Reinforcement Learning with Verifiable Rewards via Dynamic Jensen-Shannon Replay

While Reinforcement Learning (RL) enhances Large Language Model reasoning, on-policy algorithms like GRPO are sample-inefficient as they discard past rollouts. Existing experience replay methods address this by reusing accurate samples for direct policy updates, but this often incurs high computational costs and causes mode collapse via overfitting. We argue that historical data should prioritize sustaining diversity rather than simply reinforcing accuracy. To this end, we propose Dynamic Jensen-Shannon Replay (DyJR), a simple yet effective regularization framework using a dynamic reference distribution from recent trajectories. DyJR introduces two innovations: (1) A Time-Sensitive Dynamic Buffer that uses FIFO and adaptive sizing to retain only temporally proximal samples, synchronizing with model evolution; and (2) Jensen-Shannon Divergence Regularization, which replaces direct gradient updates with a distributional constraint to prevent diversity collapse. Experiments on mathematical reasoning and Text-to-SQL benchmarks demonstrate that DyJR significantly outperforms GRPO as well as baselines such as RLEP and Ex-GRPO, while maintaining training efficiency comparable to the original GRPO. Furthermore, from the perspective of Rank-$k$ token probability evolution, we show that DyJR enhances diversity and mitigates over-reliance on Rank-1 tokens, elucidating how specific sub-modules of DyJR influence the training dynamics.

cs.LG↗

SQL-ASTRA: Alleviating Sparse Feedback in Agentic SQL via Column-Set Matching and Trajectory Aggregation

Agentic Reinforcement Learning (RL) shows promise for complex tasks, but Text-to-SQL remains mostly restricted to single-turn paradigms. A primary bottleneck is the credit assignment problem. In traditional paradigms, rewards are determined solely by the final-turn feedback, which ignores the intermediate process and leads to ambiguous credit evaluation. To address this, we propose Agentic SQL, a framework featuring a universal two-tiered reward mechanism designed to provide effective trajectory-level evaluation and dense step-level signals. First, we introduce Aggregated Trajectory Reward (ATR) to resolve multi-turn credit assignment. Using an asymmetric transition matrix, ATR aggregates process-oriented scores to incentivize continuous improvement. Leveraging Lyapunov stability theory, we prove ATR acts as an energy dissipation operator, guaranteeing a cycle-free policy and monotonic convergence. Second, Column-Set Matching Reward (CSMR) provides immediate step-level rewards to mitigate sparsity. By executing queries at each turn, CSMR converts binary (0/1) feedback into dense [0, 1] signals based on partial correctness. Evaluations on BIRD show a 5% gain over binary-reward GRPO. Notably, our approach outperforms SOTA Arctic-Text2SQL-R1-7B on BIRD and Spider 2.0 using identical models, propelling Text-to-SQL toward a robust multi-turn agent paradigm.

cs.AI↗

The Choice of Divergence: A Neglected Key to Mitigating Diversity Collapse in Reinforcement Learning with Verifiable Reward

A central paradox in fine-tuning Large Language Models (LLMs) with Reinforcement Learning with Verifiable Reward (RLVR) is the frequent degradation of multi-attempt performance (Pass@k) despite improvements in single-attempt accuracy (Pass@1). This is often accompanied by catastrophic forgetting, where models lose previously acquired skills. While various methods have been proposed, the choice and function of the divergence term have been surprisingly unexamined as a proactive solution. We argue that standard RLVR objectives -- both those using the mode-seeking reverse KL-divergence and those forgoing a divergence term entirely -- lack a crucial mechanism for knowledge retention. The reverse-KL actively accelerates this decay by narrowing the policy, while its absence provides no safeguard against the model drifting from its diverse knowledge base. We propose a fundamental shift in perspective: using the divergence term itself as the solution. Our framework, Diversity-Preserving Hybrid RL (DPH-RL), leverages mass-covering f-divergences (like forward-KL and JS-divergence) to function as a rehearsal mechanism. By continuously referencing the initial policy, this approach forces the model to maintain broad solution coverage. Extensive experiments on math and SQL generation demonstrate that DPH-RL not only resolves the Pass@k degradation but improves both Pass@1 and Pass@k in- and out-of-domain. Additionally, DPH-RL is more training-efficient because it computes f-divergence using generator functions, requiring only sampling from the initial policy and no online reference model. Our work highlights a crucial, overlooked axis for improving RLVR, demonstrating that the proper selection of a divergence measure is a powerful tool for building more general and diverse reasoning models.

cs.LG↗

AT 2024wpp: the most luminous fast-evolving optical transient linked to the merger explosion of a black-hole binary

Fast blue optical transients (FBOTs) represent one of the most exotic astrophysical transients, exhibiting unusually strong emission across X-ray, optical, and radio wavelengths. Their physical origins remain highly debated, with proposed explanations ranging from stellar explosion to tidal disruption event (TDE). Here we report observations of the most luminous FBOT, AT 2024wpp whose post-peak luminosity rebrightens in X ray and becomes flattening in optical in a manner follows the decay rate characteristic of TDEs ($L_{\rm bol} \propto t^{-5/3}$). This invokes energy contribution of accretion by a central compact object, getting further corroborations from hardening of X-ray spectral index and detection of outflow inferred from the emission lines at similar phase. Detailed modeling of luminsoity evolution favors a coalesce explosion of a 34 M$_{\odot}$ Wolf-Rayet star with a 15 M$_{\odot}$ black hole (BH), demonstrating that some FBOTs may be associated with TDE of a stellar blackhole.

astro-ph.HE↗

Anchored Policy Optimization: Mitigating Exploration Collapse Via Support-Constrained Rectification

Reinforcement Learning with Verifiable Rewards (RLVR) is increasingly viewed as a tree pruning mechanism. However, we identify a systemic pathology termed Recursive Space Contraction (RSC), an irreversible collapse driven by the combined dynamics of positive sharpening and negative squeezing, where the sampling probability of valid alternatives vanishes. While Kullback-Leibler (KL) regularization aims to mitigate this, it imposes a rigid Shape Matching constraint that forces the policy to mimic the reference model's full density, creating a gradient conflict with the sharpening required for correctness. We propose Anchored Policy Optimization (APO), shifting the paradigm from global Shape Matching to Support Coverage. By defining a Safe Manifold based on the reference model's high-confidence support, APO permits aggressive sharpening for efficiency while selectively invoking a restorative force during error correction to prevent collapse. We theoretically derive that APO serves as a gradient-aligned mechanism to maximize support coverage, enabling an Elastic Recovery that re-inflates valid branches. Empirical evaluations on mathematical benchmarks demonstrate that APO breaks the accuracy-diversity trade-off, significantly improving Pass@1 while restoring the Pass@K diversity typically lost by standard policy gradient methods.

cs.AI↗

Spectral Transitions and Singular Continuous Spectrum in A New Family of Quasi-periodic Quantum Walks

This paper introduces and rigorously analyzes a new class of one-dimensional discrete-time quantum walks whose dynamics are governed by a parametrized family of extended CMV matrices. The model generalizes the unitary almost Mathieu operator (UAMO) and exhibits a richer spectral phase diagram, closely resembling the extended Harper's model. It provides the first example of a solvable quasi-periodic quantum walk that exhibits a stable region of purely singular continuous spectrum.

quant-ph↗

Electric-current-assisted nucleation of zero-field hopfion rings

Magnetic hopfions are three-dimensional topological solitons -- knotted, vortex-like spin configurations. In chiral magnets, hopfions can appear as isolated structures or they can be linked to skyrmion strings. Previous studies employed a sophisticated protocol and a special sample geometry to nucleate such hopfions linked to one or a few skyrmion strings. Here, we introduce an electric-current-assisted nucleation protocol that is simple and independent of the sample shape and size. The resulting hopfions exhibit extraordinary stability in the presence of both positive and negative magnetic fields, in perfect agreement with micromagnetic simulations. We also present a comprehensive framework for classifying hopfions, skyrmions, and merons by deriving the corresponding homotopy group.

cond-mat.mes-hall↗

Upper and Lower Bounds for The Quantum Dynamics of One-Dimensional Divergence-Type Random Jacobi Operators

We study quantum transport for the discrete one-dimensional random Jacobi operator of divergence-gradient type. For strictly positive and bounded random variables, we analyze the q-moments of the position operator and establish both upper and lower power-law bounds on their growth. Our approach relies on the asymptotic behavior of the integrated density of states and the Lyapunov exponent near the critical energy 0, previously obtained by Pastur and Figotin. A key ingredient in our analysis is the large deviation-type estimates explored via the phase formalism, which play a central role in deriving bounds on the growth of the transfer matrices.

math-ph↗

Uniform Resolvent Estimates for Subwavelength Resonators: The Minnaert Bubble Case

Subwavelength resonators are small scaled objects that exhibit contrasting medium properties (eigher in intensity or sign) while compared to the ones of a uniform background. Such contrasts allow them to resonate at specific frequencies. There are two ways to mathematically define these resonances. First, as the frequencies for which the related system of integral equations is not injective. Second, as the frequencies for which the related resolvent operator of the natural Hamiltonian, given by the wave-operator, has a pole. In this work, we consider, as the subwavelength resonator, the Minneart bubble. We show that these two mentioned definitions are equivalent. Most importantly, 1. we derive the related resolvent estimates which are uniform in terms of the size/contrast of the resonators. As a by product, we show that the resolvent operators have no scattering resonances in the upper half complex plane while they exhibit two scattering resonances in the lower half plane which converge to the real axis, as the size of the bubble tends to zero. As these resonances are poles of the natural Hamiltonian, given by the wave-operator, and have the Minnaert frequency as their dominating real part, this justifies calling them Minnaert resonances. 2. we derive the asymptotic estimates of the generated scattered fields which are uniform in terms of the incident frequency and which are valid everywhere in space (i.e. inside or outside the bubble). The dominating parts, for both the resolvent operator and the scattered fields, are given by the ones of the point-scatterer supported at the location of the bubble. In particular, these dominant parts are non trivial (not the same as those of the background medium) if and only if the used incident frequency identifies with the Minnaert one.

math.AP↗

Uniform Space and Time Behavior for Acoustic Resonators

We deal with the time-domain acoustic wave propagation in the presence of a subwavelength resonator given by a Minneart bubble. This bubble is small scaled and enjoys high contrasting mass density and bulk modulus. It is well known that, under certain regimes between these scales, such a bubble generates a single low-frequency (or subwavelength) resonance called Minnaert resonance. In this paper, we study the wave propagation governed by Minnaert resonance effects in time domain. We derive the point-approximation expansion of the wave field. The dominant part is a sum of two terms. 1. The first one, which we call the primary wave, is the wave field generated in the absence of the bubble. 2. The second one, which we call the resonant wave, is generated by the interaction between the bubble and the background. It is related to a Dirac-source, in space, that is modulated, in time, with a coefficient which is a solution of a $1$D Cauchy problem, for a second order differential equation, having as propagation and attenuation parameters the real and the imaginary parts, respectively, of the Minnaert resonance. We show that the evolution of the resonant wave remains valid for a large time of the order $ε^{-1}$, where $ε$ is the radius of the bubble, after which it collapses by exponentially decaying. Precisely, we confirm that such resonant wave have life-time inversely proportional to the imaginary part of the related subwavelength resonances, which is in our case given by the Minnaert one. In addition, the real part of this resonance fixes the period of the wave.

math.AP↗

High-Contrast Transmission Resonances for the Lamé System

We consider the Lamé transmission problem in $\mathbb{R}^3$ with a bounded isotropic elastic inclusion in a high-contrast setting, where the interior-to-exterior Lamé moduli and densities scale like $1/τ$ as $τ\to0$. We study the scattering resonances of the associated self-adjoint Hamiltonian, defined as the poles of the meromorphic continuation of its resolvent. We obtain a sharp asymptotic description of resonances near the real axis as $τ\to0$. Near each nonzero Neumann eigenvalue of the interior Lamé operator there is a cluster of resonances lying just below it in the complex plane; in this wavelength-scale regime the imaginary parts are of order $τ$ with non-vanishing leading coefficients. In addition, near zero (a subwavelength regime), we identify resonances with real parts of order $\sqrtτ$ and prove a lifetime dichotomy: their imaginary parts are of order $τ$ generically, but of order $τ^2$ for an explicit admissible set $\mathcal E$. This yields a classification of long-lived elastic resonances in the high-contrast limit. We also establish resolvent asymptotics for both fixed-size resonators and microresonators. We derive explicit expansions with a finite-rank leading term and quantitative remainder bounds, valid near both wavelength-scale and subwavelength resonances. For microresonators, at the wavelength scale the dominant contribution is an anisotropic elastic point scatterer. Near the zero eigenvalue, the leading-order behaviour is of monopole or dipole type, and we give a rigorous criterion distinguishing the two cases.

math.AP↗

IIB-LPO: Latent Policy Optimization via Iterative Information Bottleneck

Recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) for Large Language Model (LLM) reasoning have been hindered by a persistent challenge: exploration collapse. The semantic homogeneity of random rollouts often traps models in narrow, over-optimized behaviors. While existing methods leverage policy entropy to encourage exploration, they face inherent limitations. Global entropy regularization is susceptible to reward hacking, which can induce meaningless verbosity, whereas local token-selective updates struggle with the strong inductive bias of pre-trained models. To address this, we propose Latent Policy Optimization via Iterative Information Bottleneck (IIB-LPO), a novel approach that shifts exploration from statistical perturbation of token distributions to topological branching of reasoning trajectories. IIB-LPO triggers latent branching at high-entropy states to diversify reasoning paths and employs the Information Bottleneck principle both as a trajectory filter and a self-reward mechanism, ensuring concise and informative exploration. Empirical results across four mathematical reasoning benchmarks demonstrate that IIB-LPO achieves state-of-the-art performance, surpassing prior methods by margins of up to 5.3% in accuracy and 7.4% in diversity metrics.

cs.LG↗

Constraining the Properties of GRB Accreting Magnetar with $R/I$ Evolutionary Effects Using \emph{Swift}/XRT Data

A newly born millisecond magnetar has been proposed as one possible central engine of some long gamma-ray bursts (LGRBs) with X-ray plateau. In this work, we used a universal correlation between initial spin period ($P_0$) and surface magnetic field ($B_p$) of newborn magnetar based on an LGRB sample in \cite{Lan2025} to explore the propeller properties of accreting magnetars with $R/I$ evolutionary effects. We found that $B_p-P_0$ relation is approximately consistent with $B_p\propto P_{\rm eq}^{7/6}$. Here $P_{\rm eq}$ is equilibrium spin period in magnetic propeller model. The $B_p-P_0$ relation indicates that $P_0$ may not be true initial spin period of newborn magnetar but had reached an equilibrium spin period via fallback accretion in propeller model. The magnetar accretion rate in our LGRBs is in range of $\dot{M}\sim10^{-5}-10^{-2} M_{\odot} \rm s^{-1}$ by incorporating $R/I$ evolutionary effects and using the transition relation between gravitational mass $M_g$ and baryonic mass $M_b$ in different equations of state. Such accretion rates ensure that the accreting magnetars in our sample survive until reaching the equilibrium spin period, and the accretion rate is one order of magnitude lower compared to the statistical results in \cite{Stratta2018} and \cite{Linweili2020}, which used constant $R/I/M_g$ scenario. We suggested that adopting a constant $R/I/M_g$ scenario for modeling propeller regime in accreting magnetar results in a higher mass accretion rate, which may impair our understanding of the physical nature and its surroundings of accreting magnetar, and low-metallicity progenitors can provide enough material to satisfy the accretion requirements of newborn accreting magnetar in LGRBs.

astro-ph.HE↗

Characterisation of the first wafer-scale prototype for the ALICE ITS3 upgrade: the monolithic stitched sensor (MOSS)

This paper presents the characterisation and testing of the first wafer-scale monolithic stitched sensor (MOSS) prototype developed for the ALICE ITS3 upgrade that is to be installed during the LHC Long Shutdown 3 (2026-2030). The MOSS chip design is driven by the truly cylindrical detector geometry that imposes that each layer is built out of two wafer-sized, bent silicon chips. The stitching technique is employed to fabricate sensors with dimensions of 1.4 $\times$ 25.9 cm, thinned to 50 $μ$m. The chip architecture, in-pixel front-end, laboratory and in-beam characterisation, susceptibility to single-event effects, and series testing are discussed. The testing campaign validates the design of a wafer-scale stitched sensor and the performance of the pixel matrix to be within the ITS3 requirements. The MOSS chip demonstrates the feasibility of the ITS3 detector concept and provides insights for further optimisation and development.

physics.ins-det↗

Saliency-R1: Incentivizing Unified Saliency Reasoning Capability in MLLM with Confidence-Guided Reinforcement Learning

Although multimodal large language models (MLLMs) excel in high-level vision-language reasoning, they lack inherent awareness of visual saliency, making it difficult to identify key visual elements. To bridge this gap, we propose Saliency-R1, the first unified MLLM framework that jointly tackles three representative and heterogeneous saliency tasks: Salient Object Detection (SOD), Salient Instance Segmentation (SIS), and Co-salient Object Detection (CoSOD), enhancing the model's capacity for saliency reasoning. We introduce a textual interface with structured tags ( , ) to encode region- and instance-level referring expressions, enabling a single referring segmenter to produce task-appropriate masks. To train the MLLM efficiently, we propose Confidence-Guided Policy Optimization (CGPO), a novel single-sample reinforcement learning algorithm. CGPO improves on GRPO by replacing group-normalized advantages with a per-sample signal based on reward-confidence discrepancy, thereby reducing computational waste, mitigating signal dilution, and lowering training overhead. Our model exceeds or matches the performance of robust open/closed-source MLLMs and specialized state-of-the-art methods across all three tasks, demonstrating the efficacy of our framework in saliency reasoning.

cs.CV↗

FedReplay: A Feature Replay Assisted Federated Transfer Learning Framework for Efficient and Privacy-Preserving Smart Agriculture

Accurate classification plays a pivotal role in smart agriculture, enabling applications such as crop monitoring, fruit recognition, and pest detection. However, conventional centralized training often requires large-scale data collection, which raises privacy concerns, while standard federated learning struggles with non-independent and identically distributed (non-IID) data and incurs high communication costs. To address these challenges, we propose a federated learning framework that integrates a frozen Contrastive Language-Image Pre-training (CLIP) vision transformer (ViT) with a lightweight transformer classifier. By leveraging the strong feature extraction capability of the pre-trained CLIP ViT, the framework avoids training large-scale models from scratch and restricts federated updates to a compact classifier, thereby reducing transmission overhead significantly. Furthermore, to mitigate performance degradation caused by non-IID data distribution, a small subset (1%) of CLIP-extracted feature representations from all classes is shared across clients. These shared features are non-reversible to raw images, ensuring privacy preservation while aligning class representation across participants. Experimental results on agricultural classification tasks show that the proposed method achieve 86.6% accuracy, which is more than 4 times higher compared to baseline federated learning approaches. This demonstrates the effectiveness and efficiency of combining vision-language model features with federated learning for privacy-preserving and scalable agricultural intelligence.

cs.CV↗

Count Counts: Motivating Exploration in LLM Reasoning with Count-based Intrinsic Rewards

Reinforcement Learning (RL) has become a compelling way to strengthen the multi step reasoning ability of Large Language Models (LLMs). However, prevalent RL paradigms still lean on sparse outcome-based rewards and limited exploration, which often drives LLMs toward repetitive and suboptimal reasoning patterns. In this paper, we study the central question of how to design exploration for LLM reasoning and introduce MERCI (Motivating Exploration in LLM Reasoning with Count-based Intrinsic Rewards), a novel RL algorithm that augments policy optimization with a principled intrinsic reward. Building on the idea of count-based exploration, MERCI leverages a lightweight Coin Flipping Network (CFN) to estimate the pseudo count and further epistemic uncertainty over reasoning trajectories, and converts them into an intrinsic reward that values novelty while preserving the learning signal from task rewards. We integrate MERCI into some advanced RL frameworks like Group Relative Policy Optimization (GRPO). Experiments on complex reasoning benchmarks demonstrate that MERCI encourages richer and more varied chains of thought, significantly improves performance over strong baselines, and helps the policy escape local routines to discover better solutions. It indicates that our targeted intrinsic motivation can make exploration reliable for language model reasoning.

cs.AI↗

High Contrast Transmission and Fabry-Pérot-type Resonances

It is well known, in the acoustic model, that highly contrasting transmission leads to the so-called Minnaert subwavelength resonance. In this work, we show that such highly contrasting transmissions create not only one resonance but a family of infinite resonances located near the real axis where the first one (i.e. the smallest) is indeed the Minnaert one. This family of resonances are the shifts (in the lower complex plan) of the Neumann eigenvalues of the Laplacian. The well known Minneart resonance is nothing but the shift of the trivial (zero) Neumann eigenvalue of the bubble. These resonances, other than the Minnaert ones, are Fabry-Pérot-type resonances as the generated total fields, in the bubble, are dominated by a linear combination of the Neumann eigenfunctions which, in particular, might create interferences. In addition, we establish the following properties. 1. We derive the asymptotic expansions, at the second order, of this family of resonances in terms of the contrasting coefficient. 2. In the time-harmonic regime, we derive the resolvent estimates of the related Hamiltonian and the asymptotics of scattered fields that are uniform in the whole space, highlighting the contributions from this sequence of resonances. 3. In the time domain regime, we derive the time behavior of the acoustic microresonator at large time-scales inversely proportional to powers of microresonator's radius. 4. The analysis shows that near Fabry-Pérot resonances, the mircoresonator exhibits pronounced anisotropy. We believe that such a feature may pave the way for designing anisotropic metamaterials from simple configurations of a single microresonator.

math.AP↗