SearcharxivSearch

arXiv subjects

Lin Wu

Publications and source records attributed to Lin Wu.

At least 37 records · Page 2Linked to original sources

A 44-minute periodic radio transient in a supernova remnant

Long-period radio transients (LPTs) are a newly discovered class of radio emitters with periods ranging from minutes to hours. The astrophysical nature remains undetermined, particularly of LPTs with no detectable companions. We report the first evidence for a plausible supernova remnant (SNR) association with an LPT (DART J1832-0911, 2656.23+-0.15 s period), which supports a neutron star origin of such objects. The dispersion measure of this LPT, SNR's CO emission and HI absorption, and low probability of chance of alignment with field pulsars are all consistent with such an association. The source displays either phase-locked circular or nearly 100\% linear polarization, indicating its strong and geometrically stable magnetic field. No detectable optical counterpart was found, even with a 10m-class telescope. The SNR association and the stable polarization suggest that DART J1832-0911 most likely originates from a young neutron star, whose spin could have been braked by supernova's fallback materials. This discovery provides critical insights into the nature of ultra-long period transients and their link to stellar remnants.

astro-ph.HE

Personalized Cell Segmentation: Benchmark and Framework for Reference-Guided Cell Type Segmentation

Accurate cell segmentation is critical for biological and medical imaging studies. Although recent deep learning models have advanced this task, most methods are limited to generic cell segmentation, lacking the ability to differentiate specific cell types. In this work, we introduce the Personalized Cell Segmentation (PerCS) task, which aims to segment all cells of a specific type given a reference cell. To support this task, we establish a benchmark by reorganizing publicly available datasets, yielding 1,372 images and over 110,000 annotated cells. As a pioneering solution, we propose PerCS-DINO, a framework built on the DINOv2 backbone. By integrating image features and reference embeddings via a cross-attention transformer and contrastive learning, PerCS-DINO effectively segments cells matching the reference. Extensive experiments demonstrate the effectiveness of the proposed PerCS-DINO and highlight the challenges of this new task. We expect PerCS to serve as a useful testbed for advancing research in cell-based applications.

cs.CV

Towards a Science of Collective AI: LLM-based Multi-Agent Systems Need a Transition from Blind Trial-and-Error to Rigorous Science

Recent advancements in Large Language Models (LLMs) have greatly extended the capabilities of Multi-Agent Systems (MAS), demonstrating significant effectiveness across a wide range of complex and open-ended domains. However, despite this rapid progress, the field still relies heavily on empirical trial-and-error. It lacks a unified and principled scientific framework necessary for systematic optimization and improvement. This bottleneck stems from the ambiguity of attribution: first, the absence of a structured taxonomy of factors leaves researchers restricted to unguided adjustments; second, the lack of a unified metric fails to distinguish genuine collaboration gain from mere resource accumulation. In this paper, we advocate for a transition to design science through an integrated framework. We advocate to establish the collaboration gain metric ($Γ$) as the scientific standard to isolate intrinsic gains from increased budgets. Leveraging $Γ$, we propose a factor attribution paradigm to systematically identify collaboration-driving factors. To support this, we construct a systematic MAS factor library, structuring the design space into control-level presets and information-level dynamics. Ultimately, this framework facilitates the transition from blind experimentation to rigorous science, paving the way towards a true science of Collective AI.

cs.CL

Enhancing Volumetric Optical Chirality through 2D-3D Structural Design Evolution

Circular dichroism (CD) sensing plays a pivotal role in probing molecular chirality in biomedical sciences. However, engineering superchiral electromagnetic fields that can reliably amplify the faint signatures of chiral analytes remains profoundly challenging. Central to this difficulty is the need to balance two competing demands: maximizing the enhancement of chiral fields while maintaining a sufficiently large interaction volume for effective molecular interrogation. Here, we introduce a figure of merit (FOM) that captures the enhancement and spatial coverage of superchiral fields to benchmark different chiral-field configurations. We examine the effects of helix-geometry evolution on the FOM, including 2D to 3D chirality induction, winding-number escalation, helical-order enhancement, and transverse dilation. By tuning these structural degrees of freedom, the sensing volume can be enlarged without compromising the distribution and enhancement strength of fields. The optimized triple-strand helix markedly enhanced the analyte CD signal, yielding a FOM of 2.43*10^10 nm3, which surpassed prior 2D and 3D configurations by over an order of magnitude. The proposed FOM exhibits a strong linear correlation (R^2 = 0.9256) with the analyte CD signal. Our findings provide a systematic design framework for 3D chiral structures and a robust metric for assessing their chiroptical sensing performance, particularly in scenarios involving clusters of randomly oriented small molecules or a large chiral molecule.

physics.optics

SegMo: Segment-aligned Text to 3D Human Motion Generation

Generating 3D human motions from textual descriptions is an important research problem with broad applications in video games, virtual reality, and augmented reality. Recent methods align the textual description with human motion at the sequence level, neglecting the internal semantic structure of modalities. However, both motion descriptions and motion sequences can be naturally decomposed into smaller and semantically coherent segments, which can serve as atomic alignment units to achieve finer-grained correspondence. Motivated by this, we propose SegMo, a novel Segment-aligned text-conditioned human Motion generation framework to achieve fine-grained text-motion alignment. Our framework consists of three modules: (1) Text Segment Extraction, which decomposes complex textual descriptions into temporally ordered phrases, each representing a simple atomic action; (2) Motion Segment Extraction, which partitions complete motion sequences into corresponding motion segments; and (3) Fine-grained Text-Motion Alignment, which aligns text and motion segments with contrastive learning. Extensive experiments demonstrate that SegMo improves the strong baseline on two widely used datasets, achieving an improved TOP 1 score of 0.553 on the HumanML3D test set. Moreover, thanks to the learned shared embedding space for text and motion segments, SegMo can also be applied to retrieval-style tasks such as motion grounding and motion-to-text retrieval.

cs.CV

Schatten properties of commutators of fractional integrals on spaces of homogeneous type

Extending classical results of Janson and Peetre (1988) on the Schatten class $S^p$ membership of commutators of Riesz potentials on the Euclidean space, we obtain analogous results for commutators $[b,T]$, where $T\in\{T_\varepsilon,\widetilde T_α\}$ belongs to either one of two natural classes of fractional integral operators on a space of homogeneous type. Our approach is based on recent related work of Hytönen and Korte on singular (instead of fractional) integrals; working directly with the kernels, it differs from the Fourier analytic considerations of Janson and Peetre, covering new operators even when specialised to $\mathbb R^d$. The cleanest case of our characterization in spaces of lower dimension $d> 2$ and satisfying a $(1,2)$-Poincaré inequality is as follows. For a parameter $\varepsilon \in (0,\frac{1}{2}-\frac{1}{d})$ describing the order of the fractional integral $T_\varepsilon $, we have a dichotomy: If $\frac{d}{1+d\varepsilon }<p<\frac{1}{\varepsilon}$, then $[b,T_{\varepsilon}]\in S^p$ if and only if $b$ belongs to a suitable Besov (or fractional Sobolev) space. If $0<p\leq \frac{d}{1+d\varepsilon }$, then $[b,T_{\varepsilon}]\in S^p$ if and only if $b$ is constant. This is analogous to the result for singular integrals, where a similar cut-off happens at $p=d$, formally corresponding to fractional order $\varepsilon =0$. We also obtain results for other parameter values, including dimensions $0<d\leq 2$. As an application, these results are used to show Schatten properties of commutators of fractional Bessel operators, complementing recent related results of Fan, Lacey, Li, and Xiong (2025) on commutators of singular integrals in the Bessel setting.

math.FA

PrefPoE: Advantage-Guided Preference Fusion for Learning Where to Explore

Exploration in reinforcement learning remains a critical challenge, as naive entropy maximization often results in high variance and inefficient policy updates. We introduce \textbf{PrefPoE}, a novel \textit{Preference-Product-of-Experts} framework that performs intelligent, advantage-guided exploration via the first principled application of product-of-experts (PoE) fusion for single-task exploration-exploitation balancing. By training a preference network to concentrate probability mass on high-advantage actions and fusing it with the main policy through PoE, PrefPoE creates a \textbf{soft trust region} that stabilizes policy updates while maintaining targeted exploration. Across diverse control tasks spanning both continuous and discrete action spaces, PrefPoE demonstrates consistent improvements: +321\% on HalfCheetah-v4 (1276~$\rightarrow$~5375), +69\% on Ant-v4, +276\% on LunarLander-v2, with consistently enhanced training stability and sample efficiency. Unlike standard PPO, which suffers from entropy collapse, PrefPoE sustains adaptive exploration through its unique dynamics, thereby preventing premature convergence and enabling superior performance. Our results establish that learning \textit{where to explore} through advantage-guided preferences is as crucial as learning how to act, offering a general framework for enhancing policy gradient methods across the full spectrum of reinforcement learning domains. Code and pretrained models are available in supplementary materials.

cs.LG

HOI-Dyn: Learning Interaction Dynamics for Human-Object Motion Diffusion

Generating realistic 3D human-object interactions (HOIs) remains a challenging task due to the difficulty of modeling detailed interaction dynamics. Existing methods treat human and object motions independently, resulting in physically implausible and causally inconsistent behaviors. In this work, we present HOI-Dyn, a novel framework that formulates HOI generation as a driver-responder system, where human actions drive object responses. At the core of our method is a lightweight transformer-based interaction dynamics model that explicitly predicts how objects should react to human motion. To further enforce consistency, we introduce a residual-based dynamics loss that mitigates the impact of dynamics prediction errors and prevents misleading optimization signals. The dynamics model is used only during training, preserving inference efficiency. Through extensive qualitative and quantitative experiments, we demonstrate that our approach not only enhances the quality of HOI generation but also establishes a feasible metric for evaluating the quality of generated interactions.

cs.CV

On the Integration of Spatial-Temporal Knowledge: A Lightweight Approach to Atmospheric Time Series Forecasting

Transformers have gained attention in atmospheric time series forecasting (ATSF) for their ability to capture global spatial-temporal correlations. However, their complex architectures lead to excessive parameter counts and extended training times, limiting their scalability to large-scale forecasting. In this paper, we revisit ATSF from a theoretical perspective of atmospheric dynamics and uncover a key insight: spatial-temporal position embedding (STPE) can inherently model spatial-temporal correlations even without attention mechanisms. Its effectiveness arises from the integration of geographical coordinates and temporal features, which are intrinsically linked to atmospheric dynamics. Based on this, we propose STELLA, a Spatial-Temporal knowledge Embedded Lightweight modeL for ASTF, utilizing only STPE and an MLP architecture in place of Transformer layers. With 10k parameters and one hour of training, STELLA achieves superior performance on five datasets compared to other advanced methods. The paper emphasizes the effectiveness of spatial-temporal knowledge integration over complex architectures, providing novel insights for ATSF. The code is available at https://github.com/GestaltCogTeam/STELLA.

cs.LG

Spectral localization of single-nanoparticle plasmons through photonic substrate engineering

Surface plasmon resonances (SPRs) are crucial for confining light beyond the diffraction limit, yet heavy metal losses often limit their spectral localization. Here, we propose a practical strategy for enabling the spectral localization of single-nanoparticle SPRs through photonic substrate engineering, which creates distinct optical pathways (OPs) to tailor the electromagnetic environments around plasmonic nanoparticles. By analyzing the multiplication factor spectrum of the projected local density of states, we can trace and control these OPs, enabling strong spatial and spectral confinement of single-nanoparticle SPRs. Simulations reveal that a photonic crystal substrate can reduce the mode volume by fivefold and boost the quality factor by over 80 times compared to a metal nanoparticle on a dielectric substrate. Proof-of-concept experiments using two types of leaking Fabry-Perot photonic substrates demonstrate active manipulation of SPRs in both "open" and "closed" OP states. This multidimensional photonic substrate engineering establishes a customizable platform for single-nanoparticle plasmonics, potentially transforming applications that were previously limited by spectral localization.

physics.optics

Light-cone-proximal quasi-BICs for chiral lasing at grazing angles

Chiral quasi-bound states in the continuum (q-BICs) have recently emerged in metaphotonics as resonances that combine ultrahigh quality factors with near-unity circular polarization. However, these states are typically confined to the Gamma-point (normal incidence) due to their symmetry-protected origins. We propose a new mechanism for realizing light-cone-proximal chiral q-BICs at large oblique angles, enabled by the divergence of the radiative local density of states near the light cone. Using dielectric metasurfaces with a monoclinic lattice and broken in-plane mirror symmetry, we demonstrate that tuning the lattice angle allows for robust control of these resonances. The resulting chiral q-BICs exhibit near-unity circular dichroism in transmission and fully circularly polarized emission at angles exceeding 50 degrees from normal. This approach paves the way for directional chiral lasing at grazing angles and for photonic devices operating efficiently in off-normal geometries.

physics.optics

Medical Artificial Intelligence for Early Detection of Lung Cancer: A Survey

Lung cancer remains one of the leading causes of morbidity and mortality worldwide, making early diagnosis critical for improving therapeutic outcomes and patient prognosis. Computer-aided diagnosis systems, which analyze computed tomography images, have proven effective in detecting and classifying pulmonary nodules, significantly enhancing the detection rate of early-stage lung cancer. Although traditional machine learning algorithms have been valuable, they exhibit limitations in handling complex sample data. The recent emergence of deep learning has revolutionized medical image analysis, driving substantial advancements in this field. This review focuses on recent progress in deep learning for pulmonary nodule detection, segmentation, and classification. Traditional machine learning methods, such as support vector machines and k-nearest neighbors, have shown limitations, paving the way for advanced approaches like Convolutional Neural Networks, Recurrent Neural Networks, and Generative Adversarial Networks. The integration of ensemble models and novel techniques is also discussed, emphasizing the latest developments in lung cancer diagnosis. Deep learning algorithms, combined with various analytical techniques, have markedly improved the accuracy and efficiency of pulmonary nodule analysis, surpassing traditional methods, particularly in nodule classification. Although challenges remain, continuous technological advancements are expected to further strengthen the role of deep learning in medical diagnostics, especially for early lung cancer detection and diagnosis. A comprehensive list of lung cancer detection models reviewed in this work is available at https://github.com/CaiGuoHui123/Awesome-Lung-Cancer-Detection.

eess.IV

AnnaAgent: Dynamic Evolution Agent System with Multi-Session Memory for Realistic Seeker Simulation

Constrained by the cost and ethical concerns of involving real seekers in AI-driven mental health, researchers develop LLM-based conversational agents (CAs) with tailored configurations, such as profiles, symptoms, and scenarios, to simulate seekers. While these efforts advance AI in mental health, achieving more realistic seeker simulation remains hindered by two key challenges: dynamic evolution and multi-session memory. Seekers' mental states often fluctuate during counseling, which typically spans multiple sessions. To address this, we propose AnnaAgent, an emotional and cognitive dynamic agent system equipped with tertiary memory. AnnaAgent incorporates an emotion modulator and a complaint elicitor trained on real counseling dialogues, enabling dynamic control of the simulator's configurations. Additionally, its tertiary memory mechanism effectively integrates short-term and long-term memory across sessions. Evaluation results, both automated and manual, demonstrate that AnnaAgent achieves more realistic seeker simulation in psychological counseling compared to existing baselines. The ethically reviewed and screened code can be found on https://github.com/sci-m-wang/AnnaAgent.

cs.CL

Refining CNN-based Heatmap Regression with Gradient-based Corner Points for Electrode Localization

We propose a method for detecting the electrode positions in lithium-ion batteries. The process begins by identifying the region of interest (ROI) in the battery's X-ray image through corner point detection. A convolutional neural network is then used to regress the pole positions within this ROI. Finally, the regressed positions are optimized and corrected using corner point priors, significantly mitigating the loss of localization accuracy caused by operations such as feature map down-sampling and padding during network training. Our findings show that combining traditional pixel gradient analysis with CNN-based heatmap regression for keypoint extraction enhances both accuracy and efficiency, resulting in significant performance improvements.

cs.CV

Referring to Any Person

Humans are undoubtedly the most important participants in computer vision, and the ability to detect any individual given a natural language description, a task we define as referring to any person, holds substantial practical value. However, we find that existing models generally fail to achieve real-world usability, and current benchmarks are limited by their focus on one-to-one referring, that hinder progress in this area. In this work, we revisit this task from three critical perspectives: task definition, dataset design, and model architecture. We first identify five aspects of referable entities and three distinctive characteristics of this task. Next, we introduce HumanRef, a novel dataset designed to tackle these challenges and better reflect real-world applications. From a model design perspective, we integrate a multimodal large language model with an object detection framework, constructing a robust referring model named RexSeek. Experimental results reveal that state-of-the-art models, which perform well on commonly used benchmarks like RefCOCO/+/g, struggle with HumanRef due to their inability to detect multiple individuals. In contrast, RexSeek not only excels in human referring but also generalizes effectively to common object referring, making it broadly applicable across various perception tasks. Code is available at https://github.com/IDEA-Research/RexSeek

cs.CV

Radio dimming associated with filament eruptions in the meter and decimeter wavebands

Filament eruptions are considered to be a common phenomenon on the Sun and other stars, yet they are rarely directly imaged in the meter and decimeter wavebands. Using imaging data from the DAocheng solar Radio Telescope (DART) in the 150-450 MHz frequency range, we present two eruptive filaments that manifest as radio dimmings (i.e., emission depressions). Simultaneously, portions of these eruptive filaments are discernible as dark features in the chromospheric images. The sun-as-a-star flux curves of brightness temperature, derived from the DART images, exhibit obvious radio dimmings. The dimming depths range from 1.5% to 8% of the background level and show a negative correlation with radio frequencies and a positive correlation with filament areas. Our investigation suggests that radio dimming is caused by free-free absorption during filament eruptions obscuring the solar corona. This may provide a new method for detecting stellar filament eruptions.

astro-ph.SR

Exploring the origin of multi-periodic pulsations during a white-light flare

We explored the quasi-periodic pulsations (QPPs) at multiple periods during an X4.0 flare on 2024 May 10 (SOL2024-05-10T06:27), which occurred in the complex active region of NOAA 13664. The flare radiation reveals five prominent periods in multiple wavelengths. A 8-min QPP is simultaneously detected in wavelengths of HXR, radio, UV/EUV, Lya, and white light, which may be associated with nonthermal electrons periodically accelerated by intermittent magnetic reconnection that is modulated by the slow wave. A quasi-period at 14 minutes is observed in the SXR and high-temperature EUV wavebands, and it may be caused by repeatedly heated plasmas in hot flare loops. A quasiperiod at about 18 minutes is only observed by STIX, with reconstructed SXR images suggesting that the 18-min period pulsations should be considered as different flares. Meanwhile, a 3-min QPP is simultaneously detected in wavelengths of HXR, radio, and UV/ EUV, which is directly modulated by the slow magnetoacoustic wave leaking from sunspot umbrae. At last, a 2-min QPP is simultaneously detected in HXR and radio emissions during the pre-flare phase, which is possibly generated by a quasi-periodic regime of magnetic reconnection that is triggered by the kink wave.

astro-ph.SR

A Temporal Modeling Framework for Video Pre-Training on Video Instance Segmentation

Contemporary Video Instance Segmentation (VIS) methods typically adhere to a pre-train then fine-tune regime, where a segmentation model trained on images is fine-tuned on videos. However, the lack of temporal knowledge in the pre-trained model introduces a domain gap which may adversely affect the VIS performance. To effectively bridge this gap, we present a novel video pre-training approach to enhance VIS models, especially for videos with intricate instance relationships. Our crucial innovation focuses on reducing disparities between the pre-training and fine-tuning stages. Specifically, we first introduce consistent pseudo-video augmentations to create diverse pseudo-video samples for pre-training while maintaining the instance consistency across frames. Then, we incorporate a multi-scale temporal module to enhance the model's ability to model temporal relations through self- and cross-attention at short- and long-term temporal spans. Our approach does not set constraints on model architecture and can integrate seamlessly with various VIS methods. Experiment results on commonly adopted VIS benchmarks show that our method consistently outperforms state-of-the-art methods. Our approach achieves a notable 4.0% increase in average precision on the challenging OVIS dataset.

cs.CV