Searcharxiv⌕ Search

arXiv subjects

Long Li

Publications and source records attributed to Long Li.

At least 55 records · Page 3Linked to original sources

FedReplay: A Feature Replay Assisted Federated Transfer Learning Framework for Efficient and Privacy-Preserving Smart Agriculture

Accurate classification plays a pivotal role in smart agriculture, enabling applications such as crop monitoring, fruit recognition, and pest detection. However, conventional centralized training often requires large-scale data collection, which raises privacy concerns, while standard federated learning struggles with non-independent and identically distributed (non-IID) data and incurs high communication costs. To address these challenges, we propose a federated learning framework that integrates a frozen Contrastive Language-Image Pre-training (CLIP) vision transformer (ViT) with a lightweight transformer classifier. By leveraging the strong feature extraction capability of the pre-trained CLIP ViT, the framework avoids training large-scale models from scratch and restricts federated updates to a compact classifier, thereby reducing transmission overhead significantly. Furthermore, to mitigate performance degradation caused by non-IID data distribution, a small subset (1%) of CLIP-extracted feature representations from all classes is shared across clients. These shared features are non-reversible to raw images, ensuring privacy preservation while aligning class representation across participants. Experimental results on agricultural classification tasks show that the proposed method achieve 86.6% accuracy, which is more than 4 times higher compared to baseline federated learning approaches. This demonstrates the effectiveness and efficiency of combining vision-language model features with federated learning for privacy-preserving and scalable agricultural intelligence.

cs.CV↗

Count Counts: Motivating Exploration in LLM Reasoning with Count-based Intrinsic Rewards

Reinforcement Learning (RL) has become a compelling way to strengthen the multi step reasoning ability of Large Language Models (LLMs). However, prevalent RL paradigms still lean on sparse outcome-based rewards and limited exploration, which often drives LLMs toward repetitive and suboptimal reasoning patterns. In this paper, we study the central question of how to design exploration for LLM reasoning and introduce MERCI (Motivating Exploration in LLM Reasoning with Count-based Intrinsic Rewards), a novel RL algorithm that augments policy optimization with a principled intrinsic reward. Building on the idea of count-based exploration, MERCI leverages a lightweight Coin Flipping Network (CFN) to estimate the pseudo count and further epistemic uncertainty over reasoning trajectories, and converts them into an intrinsic reward that values novelty while preserving the learning signal from task rewards. We integrate MERCI into some advanced RL frameworks like Group Relative Policy Optimization (GRPO). Experiments on complex reasoning benchmarks demonstrate that MERCI encourages richer and more varied chains of thought, significantly improves performance over strong baselines, and helps the policy escape local routines to discover better solutions. It indicates that our targeted intrinsic motivation can make exploration reliable for language model reasoning.

cs.AI↗

High Contrast Transmission and Fabry-Pérot-type Resonances

It is well known, in the acoustic model, that highly contrasting transmission leads to the so-called Minnaert subwavelength resonance. In this work, we show that such highly contrasting transmissions create not only one resonance but a family of infinite resonances located near the real axis where the first one (i.e. the smallest) is indeed the Minnaert one. This family of resonances are the shifts (in the lower complex plan) of the Neumann eigenvalues of the Laplacian. The well known Minneart resonance is nothing but the shift of the trivial (zero) Neumann eigenvalue of the bubble. These resonances, other than the Minnaert ones, are Fabry-Pérot-type resonances as the generated total fields, in the bubble, are dominated by a linear combination of the Neumann eigenfunctions which, in particular, might create interferences. In addition, we establish the following properties. 1. We derive the asymptotic expansions, at the second order, of this family of resonances in terms of the contrasting coefficient. 2. In the time-harmonic regime, we derive the resolvent estimates of the related Hamiltonian and the asymptotics of scattered fields that are uniform in the whole space, highlighting the contributions from this sequence of resonances. 3. In the time domain regime, we derive the time behavior of the acoustic microresonator at large time-scales inversely proportional to powers of microresonator's radius. 4. The analysis shows that near Fabry-Pérot resonances, the mircoresonator exhibits pronounced anisotropy. We believe that such a feature may pave the way for designing anisotropic metamaterials from simple configurations of a single microresonator.

math.AP↗

VLLFL: A Vision-Language Model Based Lightweight Federated Learning Framework for Smart Agriculture

In modern smart agriculture, object detection plays a crucial role by enabling automation, precision farming, and monitoring of resources. From identifying crop health and pest infestations to optimizing harvesting processes, accurate object detection enhances both productivity and sustainability. However, training object detection models often requires large-scale data collection and raises privacy concerns, particularly when sensitive agricultural data is distributed across farms. To address these challenges, we propose VLLFL, a vision-language model-based lightweight federated learning framework (VLLFL). It harnesses the generalization and context-aware detection capabilities of the vision-language model (VLM) and leverages the privacy-preserving nature of federated learning. By training a compact prompt generator to boost the performance of the VLM deployed across different farms, VLLFL preserves privacy while reducing communication overhead. Experimental results demonstrate that VLLFL achieves 14.53% improvement in the performance of VLM while reducing 99.3% communication overhead. Spanning tasks from identifying a wide variety of fruits to detecting harmful animals in agriculture, the proposed framework offers an efficient, scalable, and privacy-preserving solution specifically tailored to agricultural applications.

cs.CV↗

Perception Before Reasoning: Two-Stage Reinforcement Learning for Visual Reasoning in Vision-Language Models

Reinforcement learning (RL) has proven highly effective in eliciting the reasoning capabilities of large language models (LLMs). Inspired by this success, recent studies have explored applying similar techniques to vision-language models (VLMs), aiming to enhance their reasoning performance. However, directly transplanting RL methods from LLMs to VLMs is suboptimal, as the tasks faced by VLMs are inherently more complex. Specifically, VLMs must first accurately perceive and understand visual inputs before reasoning can be effectively performed. To address this challenge, we propose a two-stage reinforcement learning framework designed to jointly enhance both the perceptual and reasoning capabilities of VLMs. To mitigate the vanishing advantage issue commonly observed in RL training, we first perform dataset-level sampling to selectively strengthen specific capabilities using distinct data sources. During training, the first stage focuses on improving the model's visual perception through coarse- and fine-grained visual understanding, while the second stage targets the enhancement of reasoning abilities. After the proposed two-stage reinforcement learning process, we obtain PeBR-R1, a vision-language model with significantly enhanced perceptual and reasoning capabilities. Experimental results on seven benchmark datasets demonstrate the effectiveness of our approach and validate the superior performance of PeBR-R1 across diverse visual reasoning tasks.

cs.CV↗

ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning

Reasoning-based large language models have excelled in mathematics and programming, yet their potential in knowledge-intensive medical question answering remains underexplored and insufficiently validated in clinical contexts. To bridge this gap, we introduce ReasonMed, the largest medical reasoning dataset to date, comprising 370k high-quality examples distilled from 1.75 million initial reasoning paths generated by complementary LLMs and curated through a cost-efficient easy-medium-difficult (EMD) pipeline. ReasonMed is built through a multi-agent generation, verification, and refinement process, in which an Error Refiner improves reasoning paths by correcting error-prone steps identified by a verifier. Using ReasonMed, we investigate effective strategies for training medical reasoning models and find that integrating detailed CoT reasoning with concise answer summaries yields the most robust fine-tuning results. Models trained on ReasonMed set a new benchmark: ReasonMed-7B surpasses the prior best sub-10B models by 4.17% and even exceeds LLaMA3.1-70B on PubMedQA by 4.60%. When scaled to ReasonMed-14B, it remains highly competitive, underscoring consistent scaling potential. The codes and datasets are available at https://github.com/YuSun-Work/ReasonMed.

cs.CL↗

MMA-ASIA: A Multilingual and Multimodal Alignment Framework for Culturally-Grounded Evaluation

Large language models (LLMs) are now used worldwide, yet their multimodal understanding and reasoning often degrade outside Western, high-resource settings. We propose MMA-ASIA, a comprehensive framework to evaluate LLMs' cultural awareness with a focus on Asian contexts. MMA-ASIA centers on a human-curated, multilingual, and multimodally aligned multiple-choice benchmark covering 8 Asian countries and 10 languages, comprising 27,000 questions; over 79 percent require multi-step reasoning grounded in cultural context, moving beyond simple memorization. To our knowledge, this is the first dataset aligned at the input level across three modalities: text, image (visual question answering), and speech. This enables direct tests of cross-modal transfer. Building on this benchmark, we propose a five-dimensional evaluation protocol that measures: (i) cultural-awareness disparities across countries, (ii) cross-lingual consistency, (iii) cross-modal consistency, (iv) cultural knowledge generalization, and (v) grounding validity. To ensure rigorous assessment, a Cultural Awareness Grounding Validation Module detects "shortcut learning" by checking whether the requisite cultural knowledge supports correct answers. Finally, through comparative model analysis, attention tracing, and an innovative Vision-ablated Prefix Replay (VPR) method, we probe why models diverge across languages and modalities, offering actionable insights for building culturally reliable multimodal LLMs.

cs.CL↗

On the Kotani-Last Conjecture for the Dirac Operator

We prove a dichotomy of almost periodicity for reflectionless one-dimensional Dirac operators whose spectra satisfy certain geometric conditions, extending work of Volberg--Yuditskii. We also construct a weakly mixing Dirac operator with a non-constant continuous potential whose spectrum is purely absolutely continuous, adapting Avila's argument for continuous Schrödinger operators. In particular, we disprove the Kotani--Last conjecture in the setting of one-dimensional Dirac operators.

math.SP↗

A Stochastic Ekman-Stokes Model for Coupled Ocean-Wave-Atmosphere Dynamics

Accurate representation of atmosphere-ocean boundary layers, including the interplay of turbulence, surface waves, and air-sea fluxes, remains a challenge in geophysical fluid dynamics, particularly for climate simulations. This study introduces a stochastic coupled Ekman-Stokes model (SCESM) developed within the physically consistent Location Uncertainty framework, explicitly incorporating random turbulent fluctuations and surface wave effects. The SCESM integrates established parameterizations for air-sea fluxes, turbulent viscosity, and Stokes drift, and its performance is rigorously assessed through ensemble simulations compared against observations from the LOTUS field experiment. A performance ranking analysis quantifies the impact of different model components, highlighting the critical role of explicit uncertainty representation in both oceanic and atmospheric dynamics for accurately capturing system variability. Among the tested configurations, the full model version -- including both Stokes drift and wave-induced mixing -- shows the best agreement with observations. Wave-induced mixing terms improve model performance, while wave-dependent surface roughness enhances air-sea fluxes but reduces the relative influence of wave-driven mixing. This fully coupled stochastic framework provides a foundation for advancing boundary layer parameterizations in large-scale climate models.

physics.ao-ph↗

Constrain magnetar parameters by taking into account the evolutionary effects of radius and moment of inertia with \emph{Swift}/XRT data

A newly born millisecond magnetar has been proposed as one possible central engine of some GRBs with X-ray plateau emission. In this work, we systematically analyzed the Swift/XRT data of long GRBs with plateau emission that were detected before 2023 December, and estimated the physical parameters by considering the $R/I$ evolutionary effects. We found that neglecting the $R/I$ evolutionary effects can lead to systematic overestimation or underestimation of magnetar parameters such as $B_p$, $P_0$, and $ε$ from 20\% to 50\%. We also found that some tight correlations, which can be approximately expressed as $ε\propto P_0^{1.57\pm0.22}$, $ε\propto B_p^{0.97\pm0.13}$, $B_p\propto P_0^{1.30\pm0.16}$, $E_{\rm wind}\propto E_{\rm jet,iso}^{0.83\pm0.07}(E_{\rm jet}^{0.76\pm0.06})$, $P_0\propto E_{\rm jet,iso}^{-0.29\pm0.03}(E_{\rm jet}^{-0.26\pm0.02})$, $B_p\propto E_{\rm jet,iso}^{-0.58\pm0.06}(E_{\rm jet}^{-0.55\pm0.05})$, and $ε\propto E_{\rm jet,iso}^{-0.55\pm0.07}(E_{\rm jet}^{-0.52\pm0.06})$ for our selected EoSs. The universal correlations suggest that a nascent magnetar with the faster $P_0$, lower $B_p$, and lower $ε$ are more inclined to power a more energetic GRB jet, and the $ε$ and $P_0$ of newborn magnetar are likely to originate from the magnetically induced distortion and correspond to the equilibrium spin period as a result of interaction between the magnetar and its accretion disk, respectively. Finally, we found that the GW signals from the remnants of those GW-dominated GRBs with redshift measurements cannot reach aLIGO sensitivity threshold, and only two cases (GRBs 150323A and 170607A) can reach ET sensitivity threshold. Future GW observations could not only offer the first smoking gun that a protomagnetar can serve as the central engine of GRBs but also play a crucial role in precisely constraining the neutron star EoS.

astro-ph.HE↗

$\ell_{1}^{2}-η\ell_{2}^{2}$ sparsity regularization for nonlinear ill-posed problems

In this study, we investigate the $\left\|\cdot\right\|_{\ell_{1}}^{2}-η\left\|\cdot\right\|_{\ell_{2}}^{2}$ sparsity regularization with $0< η\leq 1$, in the context of nonlinear ill-posed inverse problems. We focus on the examination of the well-posedness associated with this regularization approach. Notably, the case where $η=1$ presents weaker theoretical outcomes than $0< η<1$, primarily due to the absence of coercivity and the Radon-Riesz property associated with the regularization term. Under specific conditions pertaining to the nonlinearity of the operator $F$, we establish that every minimizer of the $\left\|\cdot\right\|_{\ell_{1}}^{2}-η\left\|\cdot\right\|_{\ell_{2}}^{2}$ regularization exhibits sparsity. Moreover, for the case where $0<η<1$, we demonstrate convergence rates of $\mathcal{O}\left(δ^{1/2}\right)$ and $\mathcal{O}\left(δ\right)$ for the regularized solution, concerning a sparse exact solution, under differing yet widely accepted conditions related to the nonlinearity of $F$. Additionally, we present the iterative half variation algorithm as an effective method for addressing the $\left\|\cdot\right\|_{\ell_{1}}^{2}-η\left\|\cdot\right\|_{\ell_{2}}^{2}$ regularization in the domain of nonlinear ill-posed equations. Numerical results provided corroborate the effectiveness of the proposed methodology.

math.NA↗

A Nyström Method for Scattering by a Two-layered Medium with a Rough Boundary

This paper is concerned with problems of scattering of time-harmonic acoustic waves by a two-layered medium with a non-locally perturbed boundary (called a rough boundary in this paper) in two dimensions, where a Dirichlet or impedance boundary condition is imposed on the boundary. The two-layered medium is composed of two unbounded media with different physical properties and the interface between the two media is considered to be a planar surface. We formulate the scattering problems considered as boundary value problems and prove the result of the well-posedness of each boundary value problem by utilizing the integral equation method associated with the two-layered Green function. Moreover, we develop a Nyström method for numerically solving the boundary value problems considered, based on the proposed integral equation formulations. We establish the convergence results of the Nyström method with the convergence rates depending on the smoothness of the rough boundary. It is worth noting that in establishing the well-posedness of the boundary value problems as well as the convergence results of the Nyström method, an essential role is played by the investigation of the asymptotic properties of the two-layered Green function for small and large arguments. Finally, numerical experiments are carried out to show the effectiveness of the Nyström method.

math.NA↗

VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning

Reinforcement learning has proven its effectiveness in enhancing the reasoning capabilities of large language models. Recent research efforts have progressively extended this paradigm to multimodal reasoning tasks. Due to the inherent complexity and diversity of multimodal tasks, especially in semantic content and problem formulations, existing models often exhibit unstable performance across various domains and difficulty levels. To address these limitations, we propose VL-Cogito, an advanced multimodal reasoning model trained via a novel multi-stage Progressive Curriculum Reinforcement Learning (PCuRL) framework. PCuRL systematically guides the model through tasks of gradually increasing difficulty, substantially improving its reasoning abilities across diverse multimodal contexts. The framework introduces two key innovations: (1) an online difficulty soft weighting mechanism, dynamically adjusting training difficulty across successive RL training stages; and (2) a dynamic length reward mechanism, which encourages the model to adaptively regulate its reasoning path length according to task complexity, thus balancing reasoning efficiency with correctness. Experimental evaluations demonstrate that VL-Cogito consistently matches or surpasses existing reasoning-oriented models across mainstream multimodal benchmarks spanning mathematics, science, logic, and general understanding, validating the effectiveness of our approach.

cs.CV↗

To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization

Recent advances in mathematical problem-solving with language models (LMs) integrate chain-of-thought (CoT) reasoning and code execution to harness their complementary strengths. However, existing hybrid frameworks exhibit a critical limitation: they depend on externally dictated instructions or rigid code-integration templates, lacking metacognitive awareness -- the capacity to dynamically evaluate intrinsic capabilities and autonomously determine when and how to integrate tools. This rigidity motivates our study of autonomous code integration, enabling models to adapt tool-usage strategies as their reasoning abilities evolve during training. While reinforcement learning (RL) shows promise for boosting LLM reasoning at scale (e.g., DeepSeek-R1), we demonstrate its inefficiency in learning autonomous code integration due to inadequate exploration of the vast combinatorial space of CoT-code interleaving patterns. To address this challenge, we propose a novel Expectation-Maximization (EM) framework that synergizes structured exploration (E-step) with off-policy RL optimization (M-step), creating a self-reinforcing cycle between metacognitive tool-use decisions and evolving capabilities. Experiments reveal our method achieves superior results through improved exploration. Notably, our 7B model improves over 11% on MATH500 and 9.4% on AIME without o1-like CoT.

cs.AI↗

On the residual Monge-Ampère mass of plurisubharmonic functions, III: uniformly directional Lipschitz

The purpose of this article is to study the (residual) Monge-Ampère mass of a plurisubharmonic function with an isolated unbounded locus. A general decomposition formula is obtained under the Sasakian structure of the unit sphere. In complex dimension two, we obtain an $L^{1}$-apriori estimate on the complex Monge-Ampère operator. This induces an upper-bound estimate on the residual mass, provided with the uniform directional Lipschitz continuity. As an application, the zero mass conjecture is confirmed, if the function further separates the circular direction in its alternating part.

math.CV↗

Stability of rotating magnetic levitation

Dynamical magnetic levitation has attracted broad interest in the realm of physics and engineering. The stability analysis of such system is of great significance for practical applications. In this work, we investigate the stable magnetic levitation of a floater magnet above a rotating magnet and copper board system. The conditions for stable levitation are analyzed through both theoretical modeling and experimental observation. This study focuses on the interplay between magnetic forces, damping effects from the copper board, and rotational dynamics. We derive the equilibrium conditions, perform stability analysis, and present phase diagrams in parametric spaces of rotation speed and damping coefficients. The theoretical predictions show qualitative agreement with experimental results, particularly in demonstrating how damping is essential for stable levitation and how the stability region depends on the geometric and magnetic parameters of the system.

physics.class-ph↗

$\ell_{1}^{2}-η\ell_{2}^{2}$ regularization for sparse recovery

This paper presents a regularization technique incorporating a non-convex and non-smooth term, $\ell_{1}^{2}-η\ell_{2}^{2}$, with parameters $0<η\leq 1$ designed to address ill-posed linear problems that yield sparse solutions. We explore the existence, stability, and convergence of the regularized solution, demonstrating that the $\ell_{1}^{2}-η\ell_{2}^{2}$ regularization is well-posed and results in sparse solutions. Under suitable source conditions, we establish a convergence rate of $\mathcal{O}\left(δ\right)$ in the $\ell_{2}$-norm for both a priori and a posteriori parameter choice rules. Additionally, we propose and analyze a numerical algorithm based on a half-variation iterative strategy combined with the proximal gradient method. We prove convergence despite the regularization term being non-smooth and non-convex. The algorithm features a straightforward structure, facilitating implementation. Furthermore, we propose a projected gradient iterative strategy base on surrogate function approach to achieve faster solving. Experimentally, we demonstrate visible improvements of $\ell_{1}^{2}-η\ell_{2}^{2}$ over $\ell_{1}$, $\ell_{1}-η\ell_{2}$, and other nonconvex regularizations for compressive sensing and image deblurring problems. All the numerical results show the efficiency of our proposed approach.

math.OC↗

SVD method for sparse recovery

Sparsity regularization has garnered significant interest across multiple disciplines, including statistics, imaging, and signal processing. Standard techniques for addressing sparsity regularization include iterative soft thresholding algorithms and their accelerated variants. However, these algorithms rely on Landweber iteration, which can be computationally intensive. Therefore, there is a pressing need to develop a more efficient algorithm for sparsity regularization. The Singular Value Decomposition (SVD) method serves as a regularization strategy that does not require Landweber iterations; however, it is confined to classical quadratic regularization. This paper introduces two inversion schemes tailored for situations where the operator $K$ is diagonal within a specific orthogonal basis, focusing on $\ell_{p}$ regularization when $p=1$ and $p=1/2$. Furthermore, we demonstrate that for a general linear compact operator $K$, the SVD method serves as an effective regularization strategy. To assess the efficacy of the proposed methodologies, We conduct several numerical experiments to evaluate the proposed method's effectiveness. The results indicate that our algorithms not only operate faster but also achieve a higher success rate than traditional iterative methods.

math.OC↗