SearcharxivSearch

arXiv subjects

Wen Yang

Publications and source records attributed to Wen Yang.

At least 19 recordsLinked to original sources

EDGE: Experience-Distillation for Guided Exploration in Agentic Reinforcement Learning

Reinforcement learning with outcome-based objectives such as GRPO enables LLM-based agents to solve complex, long-horizon tasks, yet the reusable exploration patterns embedded in interaction trajectories are largely discarded after a single policy update. Existing experience-augmented approaches retrieve historical guidance at inference time, but they apply experiences without accounting for the policy's evolving capability and create persistent dependencies on external retrieval. We propose EDGE (Experience-Distillation for Guided Exploration), a framework that treats retrieved experiences as temporary training-time scaffolds and progressively internalizes their benefits into the parametric policy. Concretely, EDGE partitions each rollout group into experience-conditioned and experience-free trajectories to estimate and admit only positive marginal gains without extra sampling, then distills the induced behavior into the base policy via a reverse-KL objective on its own empirical support. A co-evolutionary experience bank further synthesizes guidance from emerging failure modes and prunes obsolete entries as the policy evolves. Across embodied, web, and search-based QA tasks, EDGE improves over strong RL baselines by up to 12.5 points and remains effective without inference-time scaffolds or a proprietary reflector. The code is available at https://github.com/xvolcano02/EDGE.

cs.CL

On Brezis' open problem 2.2

We prove that the global minimizer of the Ginzburg-Landau energy in the disk of radius $R$ with boundary value $ u(x)=\frac{x}{|x|}$ is the degree-one radial solution of the planar Ginzburg--Landau equation. This gives an affirmative answer to Open Problem~2.2 in Brezis' open-problem list. This is achieved by comparing the radial solution $f$ in the disk with the degree-one radial solution $F$ in the whole plane. Multiplying a disk competitor by $F/f$ enables us to use the known minimality of the whole-plane vortex without changing the boundary trace. The difference of the two energies can be decomposed into Fourier modes. Every nonzero mode is nonnegative, and the zero mode is then handled by a Picone type identity.

math.AP

MEDR: Query-Independent Frame Selection via Multi-Signal Event Modeling and Dynamic Rescoring

Frame selection is a fundamental component of multimodal large language models, enabling long videos to be processed under limited visual-token and computational budgets. Uniform sampling preserves temporal coverage but may miss informative content that appears only briefly. To alleviate this limitation, query-dependent methods can retrieve question-relevant frames. However, because the selected frames depend on the current question, the same visual input cannot be directly shared across different questions, and frame selection must be repeated in multi-turn video dialogue. This motivates us to seek a query-independent frame selection method that preserves the reusability of a fixed visual input while improving the coverage of informative events beyond uniform sampling. We propose Multi-Signal Event Modeling and Dynamic Rescoring (MEDR), a training-free and query-independent frame selection method. Multi-Signal Event Modeling organizes complementary visual, motion, and text signals into signal-specific temporal events. Dynamic Rescoring then iteratively reevaluates each candidate relative to the current selected set, updating its score according to frame-level signal strength, additional event coverage, and temporal proximity. The resulting fixed frame set is constructed without observing the query and can be reused across different questions. On the standard benchmark evaluations, MEDR improves model accuracy by 0.63%-0.89% on Video-MME. On the long-video subset of LongVideoBench, it improves accuracy by up to 1.23% with Qwen3-VL-8B. MEDR further improves overall accuracy by 0.53%, while reusing exactly the same frame set for every question about a video.

cs.CV

Nonradial stable solutions near the Joseph--Lundgren threshold

We study positive stable solutions of the supercritical Lane--Emden equation in the first Joseph--Lundgren interval. For a family of dimensions, we construct nonradial stable entire solutions with exponent close to the upper endpoint of this interval. This disproves a radiality conjecture of Chan and Wei. The construction begins with a smooth positive nonconstant solution on the sphere, which is obtained by matching a polar cap to an inner neck. A sharp expansion of the lowest shifted eigenvalue proves that the resulting singular cone is strictly stable. Finally, a minimal-solution and rescaling argument replaces the cone by a smooth stable entire solution while preserving its sphere variation. To the best of our knowledge, this is the first nontrivial example of nonradial stable solutions for the Emden-Fowler equation.

math.AP

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects

Current end-to-end speech dialogue models are primarily optimized for mainstream languages and remain limited in low-resource dialect scenarios due to the scarcity of dialect speech data. Moreover, during dialect adaptation, the semantic representation space of speech dialogue models continuously evolves, while conventional speech supervision remains unchanged, leading to semantic inconsistency between hidden representations and speech targets and degrading speech stability and naturalness. To address these issues, we propose DialectS2S, an end-to-end speech dialogue model for Chinese dialects. We first develop a scalable dialect speech dialogue synthesis pipeline for efficient data construction. We further introduce a two-stage post-training strategy with self-aligned speech supervision, which aligns the semantic content of speech supervision with the evolved semantic representations of the model to improve dialect speech generation quality. Experimental results show that DialectS2S consistently outperforms existing baselines across multiple Chinese dialects in speech dialogue, achieving substantial improvements in dialect consistency, response quality, and speech intelligibility. Our work provides an efficient and scalable solution for end-to-end speech dialogue modeling in low-resource dialect scenarios. To facilitate future research and practical applications, we fully open-source the DialectS2S framework, including model checkpoints, training datasets, and fine-tuning code.

cs.CL

Nondegeneracy and Morse Index of Ginzburg--Landau Vortices

We prove that the standard degree-two and degree-three vortex solutions of the Ginzburg-Landau equation are nondegenerate. Their Morse indices are also computed. The proof relies on new explicit upper and lower bounds of the modulus of these solutions and a comparison argument. It is expected that our method can be generalized to study higher degree solutions.

math.AP

On Sirakov's equal-frequency uniqueness conjecture

Let $N\in\{2,3\}$, $0<\mu _1\leq\mu _2$, and $0<\beta<\mu _1$. We prove that the equal-frequency two-component cubic Schr\"odinger system \[ -\Delta u+u=\mu _1u^3+\beta uv^2, \qquad -\Delta v+v=\mu _2v^3+\beta u^2v \quad\text{in }\mathbb{R}^N \] has exactly one positive solution in $H^1(\mathbb{R}^N)\times H^1(\mathbb{R}^N)$ modulo simultaneous translations. More precisely, every positive solution is a simultaneous translate of the synchronized state constructed from the unique positive radial solution of $-\Delta w+w=w^3$ in $\mathbb{R}^N$. This settles Sirakov's equal-frequency uniqueness conjecture throughout the weak-coupling range. The main difficulty in the proof is to exclude radial solutions for which the ratio of the normalized components is nonconstant. After normalization, the two components satisfy scalar equations with a common potential. We construct a weighted Pohozaev functional for the system together with a correction term and prove that both the corrected functional and the associated weighted functional are strictly positive. Combining these sign properties with a radial flux identity and an auxiliary quotient associated with the component ratio forces synchronization.

math.AP

Timing and Spectral Analysis of the 2024 Outburst of 2S 1553$-$542 with NuSTAR and NICER

We report a timing and spectral study of the 2024 outburst of the Be/X-ray binary pulsar 2S~1553$-$542 using \textit{NuSTAR} and \textit{NICER} observations. From the \textit{NuSTAR} light curve we measure a pulse period of $9.285022\pm0.000001$~s. The energy-resolved pulse profiles are dominated by a single peak and show a wing-like structure most clearly in the $12$--$22$~keV band. The pulsed fraction remains above 60\% and increases with energy. The phase-averaged \textit{NuSTAR} spectrum is described by an absorbed blackbody plus cutoff power-law continuum, together with an iron emission line and a cyclotron absorption feature. Using the \texttt{cyclabs} model, we obtain a cyclotron energy of $E_{\rm cyc}\simeq24.1$~keV, corresponding to a magnetic field strength of $B\sim3\times10^{12}$~G. Phase-resolved spectroscopy shows that the continuum and cyclotron-line parameters vary with pulse phase, and that the line becomes poorly constrained around the pulse-wing phase. We also searched the short \textit{NICER} GTIs for transient mHz variability using wavelet analysis and a CEEMDAN-based Hilbert--Huang transform. Localized excesses near $\sim10$~mHz and $\sim20$~mHz are found, but the short exposures, COI effects, red-noise fluctuations, and the lack of a well-constrained Fourier peak limit their significance. We therefore treat them as candidate mHz variability rather than firm mHz QPO detections.

astro-ph.HE

Finite Potential Energy for Entire Solutions of the Planar Ginzburg--Landau Equation

We prove that every smooth entire solution $ u\colon\mathbb{R}^2\to\mathbb{R}^2 $ of the Ginzburg--Landau equation $ -\Delta u=u(1-|u|^2) $ with $ |u(x)|\to1 $ as $ |x|\to\infty $ has finite potential energy, i.e., \begin{equation*} \int_{\mathbb{R}^2}(1-|u|^2)^2 \mathrm{d}x<+\infty, \end{equation*} thereby resolving Brezis' Open Problem 2.5 in [4]. The main difficulty stems from the possible presence of a curl-free mode that carries nonzero circulation and decays only like $ |x|^{-1} $; such a mode lies outside $ L^2 $ and does not admit a single-valued potential. By minimizing over $ L^2 $ gradient corrections, we construct a comparison field that solves the homogeneous equation and inherits the same circulation. The Kelvin inversion, combined with the De Giorgi--Nash--Moser theory for quasilinear elliptic equations, then produces the optimal decay $ O(|x|^{-1}) $. For a Ginzburg--Landau solution, the Bernstein estimate and the coercivity of the Jacobi form produce an $ L^2 $ forcing term in the exterior phase equation. The resulting $ L^4 $ bound on the phase field implies $1-|u|^2\in L^2(\mathbb{R}^2)$, and therefore the potential energy is finite.

math.AP

Task Decomposition-Guided Reranking for Adaptive Agent Skill Retrieval

Skill usage can significantly enhance the ability of modern agent systems to complete complex tasks. However, the growing scale of skill libraries makes accurate skill selection increasingly challenging. In real-world scenarios, ambiguous semantic matching often arises between a specific task requirement and multiple generic yet semantically similar candidate skills. Moreover, existing methods tend to overlook the dynamic influence of task difficulty and skill applicability when selecting the optimal target skill set. To address these issues, we propose SkillReranker, an inference-time reranking framework for adaptive skill selection. Specifically, we first perform semantic decomposition on both the task and skill sides, yielding informative subtask and execution-state descriptions as well as transition-state descriptions that characterize each skill's functionality. These descriptions are then used to construct a directed acyclic execution graph, where intermediate task states are modeled as nodes and candidate skills as edges, thereby establishing a structured task-skill correspondence. On this basis, SkillReranker determines whether each state node satisfies the split condition to identify subtask intervals. For each task interval, we employ a cross-encoder to perform comprehensive scoring over candidate skills and select the most suitable ones to form the final target skill set. Experiments on ALFWorld and ScienceWorld with three backbone LLMs show that SkillReranker effectively improves task performance, reduces environment interaction steps, and lowers token consumption compared with existing skill selection baselines.

cs.AI

Antiferromagnetic pseudospintronics without spin splitting

Antiferromagnets (AFMs) are promising for high-density spintronics due to their zero net magnetization, yet conventional AFM spintronics relies on spin splitting-a requirement that excludes many collinear AFMs with compensated spin sublattices. Here we exploit the sublattice degree of freedom in a honeycomb AFM with zero spin splitting. We uncover a coupling between spin and sublattice: the out-of-plane pseudospin polarization is spin-dependent, a mechanism we term partial pseudospin-spin coupling. This allows switching of the pseudospin polarization by reversing the N\'eel vector. Introducing an impurity into a specific sublattice induces Friedel oscillations with a sublattice-resolved amplitude ratio dictated solely by the pseudospin polarization, which is directly measurable by spin-polarized scanning tunneling microscopy. Furthermore, we demonstrate N\'eel-vector-controlled transmission and a large nonvolatile tunneling magnetoresistance in an all-in-one AFM junction, with pronounced resonant enhancement in gate-tunable two-dimensional devices. Our work establishes a new paradigm-AFM pseudospintronics-that utilizes the sublattice pseudospin in zero-spin-splitting AFMs, extending spintronics beyond the conventional spin-splitting paradigm.

cond-mat.mes-hall

Hierarchical Fine-Grained Aerial Object Detection

Fine-grained aerial object detection, driven by the intrinsic granularity of real-world object categories, is crucial for advanced scene understanding in remote sensing. Existing methods largely inherit the paradigm of coarse-grained object detection, relying solely on single-label supervision and thus struggling to distinguish model-level categories with subtle structural differences. However, for each specific model (e.g., Boeing 787), structured prior knowledge such as attributes and hierarchies offers discriminative semantics across multiple granularities. Motivated by this, we present ExpertDet, a scheme that incorporates expert-informed cues to enhance fine-grained aerial object detection. Specifically, we design Vision-aware Masked Attribute Modeling (VMAM), which aligns attribute semantics with visual structures by reconstructing randomly masked attributes from visual cues, enabling the detector to capture subtle structural distinctions. We further propose Hierarchical Visual Instance Promotion (HierVIP), which builds a visual prototype tree based on hierarchical relations and imposes taxonomy-aware constraints to preserve cross-level semantic continuity while enhancing category discrimination. Moreover, we curate a new fine-grained object detection benchmark for Precise recognition of model-specific Ships and Planes from aerial imagery, PSP, covering 106 ship classes and 30 airplane models, respectively, featuring the most extensive collection of model-specific categories among existing aerial object detection datasets to date. We benchmark state-of-the-art object detection algorithms on the PSP benchmark. Extensive evaluation demonstrates that ExpertDet consistently outperforms other fine-grained competitors across hierarchy levels. The dataset, benchmark, and code are available at https://nnnnerd.github.io/PSP-Benchmark/.

cs.CV

Symmetry breaking for the complex sine-Gordon equation

We consider the existence of vortex solutions to the complex sine-Gordon II (CSG2) equation, which can be viewed as an analogy of the Ginzburg-Landau (GL) equation. Using the nontrivial kernels $\eta_{\pm}$ of the linearized CSG2 equation at the standard degree-2 vortex solution $\Psi_2$, we show that it bifurcating to a one-parameter family of symmetry-breaking solutions $\Psi_{2,\alpha}$. Explicit formulas of these solutions are also available, from which we propose a new bilinear system for this equation. Our method can be generalized to higher degree case. Nondegenracy and stability of degree-1 solution are also proved. Finally, we formally discuss the Lyapunov-Schmidt reduction procedure for the multivortex solutions of the CSG2 and GL equation.

math.AP

From Contrast to Consistency: Rethinking Event-based Continuous-Time Optical Flow Estimation

Estimating continuous optical flow is a fundamental yet challenging problem in dynamic visual perception. Event-based cameras, with microsecond latency and high dynamic range, capture brightness changes asynchronously, offering a unique opportunity to model motion with fine temporal precision. However, the scarcity of temporally dense ground-truth annotations limits the effectiveness of supervised learning, while contrast maximization (CM) frameworks, focused on sharpening the Image of Warped Events (IWE), often neglect temporal continuity and structural coherence, leading to distorted trajectories under complex motion. To overcome these challenges, we propose a hybrid-supervised framework for continuous-time optical flow estimation, grounded in the principle of Spatio-temporal Structural Consistency (STSC). This paradigm jointly enforces local structural stability and trajectory continuity, ensuring physically coherent motion across time. To further enhance representation and robustness, we design a bidirectionally complementary multi-scale architecture and employ a curriculum-guided hybrid training strategy, enabling a smooth transition from supervised point constraints to self-supervised manifold regularization. Comprehensive experiments across multiple benchmarks show that our method achieves state-of-the-art performance in both continuous-time and standard optical flow estimation, demonstrating the effectiveness of the proposed learning paradigm.

cs.CV

TokAlign++: Advancing Vocabulary Adaptation via Better Token Alignment

Tokenization is a foundational step in the text process of Large Language Models (LLMs). Texts must be first tokenized into token IDs, which are then input to LLMs. Inefficient tokenization results in long token-ID sequences and will slow down the training and inference of LLMs. The fine-grained knowledge transfer between LLMs, like token-level distillation, is also impeded by the mismatch in vocabulary. To bridge this gap, we introduce a method named TokAlign++ to improve vocabulary adaptation performance by learning better token alignment lexicon. The source and target vocabularies are taken as two different languages, and the bilingual token alignment lexicon is learned from monolingual token representations. Model parameters are rearranged following this bilingual lexicon for new vocabulary, and progressively fine-tuned for adaptation. Experimental results on 15 languages show that our method boosts the multilingual text compression rates and preserves most of the multilingual ability of vanilla models. It costs as few as 1k steps to restore the performance of the vanilla model. After unifying vocabularies between vanilla models, token-level distillation remarkably improves the base model with only 235M tokens.

cs.CL

MOC-3D: Manifold-Order Consistency for Text-to-3D Generation

With the burgeoning development of fields such as the Metaverse, Virtual Reality (VR), and Digital Twins, text-to-3D generation has emerged as a research hotspot in both academia and industry. Currently, optimization methods based on Score Distillation Sampling (SDS) utilizing 2D diffusion priors have become the mainstream technological paradigm in this field. However, due to the view bias of 2D priors and the mode-seeking ambiguity combined with gradient noise induced by high Classifier-Free Guidance (CFG), these methods still suffer from macro-topological inconsistency (e.g., the Janus problem) and micro-geometric discontinuity. To address these challenges, we propose MOC-3D, a text-to-3D generation method based on geometric manifold and semantic view-order consistency. Built upon the ScaleDreamer framework, our method incorporates a Semantic View-Order Constraint Module and a Manifold-based Feature Continuity Module. The former aims to rectify macro-topological inconsistency, while the latter focuses on eliminating micro-geometric discontinuity. Specifically, the Semantic View-Order Constraint Module leverages the prior knowledge of CLIP to impose a Monotonicity Rank Constraint on semantic score representations across different views, thereby providing effective guidance for the global topological structure of 3D objects. Meanwhile, the Manifold-based Feature Continuity Module employs the Riemannian Metric on the Symmetric Positive Definite (SPD) manifold. By measuring the distance of feature statistical distributions in the Riemannian space, it promotes the smooth evolution and continuity of micro-textures across multi-views in a statistical sense. Under the macro-micro synergistic optimization of these two modules, our model can simultaneously improve macro-structural consistency and micro-detail continuity.

cs.CV

UHR-DETR: Efficient End-to-End Small Object Detection for Ultra-High-Resolution Remote Sensing Imagery

Ultra-High-Resolution (UHR) imagery has become essential for modern remote sensing, offering unprecedented spatial coverage. However, detecting small objects in such vast scenes presents a critical dilemma: retaining the original resolution for small objects causes prohibitive memory bottlenecks. Conversely, conventional compromises like image downsampling or patch cropping either erase small objects or destroy context. To break this dilemma, we propose UHR-DETR, an efficient end-to-end transformer-based detector designed for UHR imagery. First, we introduce a Coverage-Maximizing Sparse Encoder that dynamically allocates finite computational resources to informative high-resolution regions, ensuring maximum object coverage with minimal spatial redundancy. Second, we design a Global-Local Decoupled Decoder. By integrating macroscopic scene awareness with microscopic object details, this module resolves semantic ambiguities and prevents scene fragmentation. Extensive experiments on the UHR imagery datasets (e.g., STAR and SODA-A) demonstrate the superiority of UHR-DETR under strict hardware constraints (e.g., a single 24GB RTX 3090). It achieves a 2.8\% mAP improvement while delivering a 10$\times$ inference speedup compared to standard sliding-window baselines on the STAR dataset. Our codes and models will be available at https://github.com/Li-JingFang/UHR-DETR.

cs.CV

Dual-Exposure Imaging with Events

By combining complementary benefits of short- and long-exposure images, Dual-Exposure Imaging (DEI) enhances image quality in low-light scenarios. However, existing DEI approaches inevitably suffer from producing artifacts due to spatial displacement from scene motion and image feature discrepancies from different exposure times. To tackle this problem, we propose a novel Event-based DEI (E-DEI) algorithm, which reconstructs high-quality images from dual-exposure image pairs and events, leveraging high temporal resolution of event cameras to provide accurate inter-/intra-frame dynamic information. Specifically, we decompose this complex task into an integration of two sub-tasks, i.e., event-based motion deblurring and low-light image enhancement tasks, which guides us to design E-DEI network as a dual-path parallel feature propagation architecture. We propose a Dual-path Feature Alignment and Fusion (DFAF) module to effectively align and fuse features extracted from dual-exposure images with assistance of events. Furthermore, we build a real-world Dataset containing Paired low-/normal-light Images and Events (PIED). Experiments on multiple datasets show the superiority of our method. The code and dataset are available at github.

cs.CV