SearcharxivSearch

arXiv subjects

Yajie Zhang

Publications and source records attributed to Yajie Zhang.

At least 19 recordsLinked to original sources

ATP-Bench: Towards Agentic Tool Planning for MLLM Interleaved Generation

Interleaved text-and-image generation represents a significant frontier for Multimodal Large Language Models (MLLMs), offering a more intuitive way to convey complex information. Current paradigms rely on either image generation or retrieval augmentation, yet they typically treat the two as mutually exclusive paths, failing to unify factuality with creativity. We argue that the next milestone in this field is Agentic Tool Planning, where the model serves as a central controller that autonomously determines when, where, and which tools to invoke to produce interleaved responses for visual-critical queries. To systematically evaluate this paradigm, we introduce ATP-Bench, a novel benchmark comprising 7,702 QA pairs (including 1,592 VQA pairs) across eight categories and 25 visual-critical intents, featuring human-verified queries and ground truths. Furthermore, to evaluate agentic planning independent of end-to-end execution and changing tool backends, we propose a Multi-Agent MLLM-as-a-Judge (MAM) system. MAM evaluates tool-call precision, identifies missed opportunities for tool use, and assesses overall response quality without requiring ground-truth references. Our extensive experiments on 10 state-of-the-art MLLMs reveal that models struggle with coherent interleaved planning and exhibit significant variations in tool-use behavior, highlighting substantial room for improvement and providing actionable guidance for advancing interleaved generation. Dataset and code are available at https://github.com/Qwen-Applications/ATP-Bench.

cs.AI

STAIL: Semantic Text-Anchored Incremental Learning for Medical Imaging via Large Language Models

Deep learning models applied to medical image analysis suffer from severe catastrophic forgetting when continually adapting to new clinical tasks in dynamic environments. Mainstream incremental learning methods typically mitigate this by rehearsing raw historical images. However, this pixel-level rehearsal incurs significant storage overhead, raises privacy concerns, and fails to adequately capture the true data distribution with sparse exemplars. Inspired by human cognitive mechanisms, we propose a novel framework termed Semantic Text-Anchored Incremental Learning (STAIL) for sequential clinical tasks. To overcome the rehearsal bottleneck, STAIL introduces an asymmetric semantic consolidation buffer (SCB). By incorporating a minimal set of image anchors and extensive textual descriptions, the SCB enables dense semantic reconstruction of old tasks at a minimal storage cost. Furthermore, we design an LLM-derived Semantic Anchoring Mechanism (LSAM) that leverages the stable semantic space of frozen large language models as developmental priors. This mechanism explicitly anchors evolving visual features to textual representations, guiding and constraining plasticity and stability at both macroscopic and microscopic levels. Extensive experiments across three heterogeneous medical datasets, covering fundus, ultrasound, and X-ray imaging, demonstrate that STAIL acts as a highly effective plug-and-play module. It comprehensively enhances the performance of various existing baselines, achieving average gains of 2.24\% in AAA-AUC for sustained performance and 3.55\% in BWT-AUC for reduced forgetting. Code is available.

cs.CV

A Second-Order Nonlocal Approximation to Manifold Poisson Models with Neumann Boundary

In this paper, we propose a class of nonlocal models to approximate the Poisson model on manifolds with homogeneous Neumann boundary condition, where the manifolds are assumed to be embedded in high dimensional Euclid spaces. In comparison to the existing nonlocal approximation of Poisson models with Neumann boundary, we optimize the truncation error of model by adding an augmented function involving the second order normal derivative along the $2δ$ layer of boundary, with $2δ$ be the nonlocal interaction horizon. The 2nd normal derivative is expressed as the difference between the interior Laplacian and the boundary Laplacian. The concentration of our paper is on the construction of nonlocal model, the well-posedness of model, and its second-order convergence rate to its local counterpart. The localization rate of our nonlocal model is currently optimal among all related works even for the case of high dimensional Euclid spaces.

math.NA

Introducing LongCat-Flash-Thinking: A Technical Report

We present LongCat-Flash-Thinking, an efficient 560-billion-parameter open-source Mixture-of-Experts (MoE) reasoning model. Its advanced capabilities are cultivated through a meticulously crafted training process, beginning with long Chain-of-Thought (CoT) data cold-start and culminating in large-scale Reinforcement Learning (RL). We first employ a well-designed cold-start training strategy, which significantly enhances the reasoning potential and equips the model with specialized skills in both formal and agentic reasoning. Then, a core innovation is our domain-parallel training scheme, which decouples optimization across distinct domains (e.g., STEM, Code, Agentic) and subsequently fuses the resulting expert models into a single, nearly Pareto-optimal model. This entire process is powered by our Dynamic ORchestration for Asynchronous rollout (DORA) system, a large-scale RL framework that delivers a greater than threefold training speedup over synchronous methods on tens of thousands of accelerators. As a result, LongCat-Flash-Thinking achieves state-of-the-art performance among open-source models on a suite of complex reasoning tasks. The model exhibits exceptional efficiency in agentic reasoning, reducing average token consumption by 64.5% (from 19, 653 to 6, 965) on AIME-25, without degrading task accuracy. We release LongCat-Flash-Thinking to promote further advances in reasoning systems and agentic AI research.

cs.AI

Cosmic Ray Detection and Rejection for CSST

As a space telescope, the China Space Station Survey Telescope (CSST) will face significant challenges from cosmic ray (CR) contamination. These CRs will severely degrade image quality and further influence scientific analysis. Due to the CSST's sky survey strategy, traditional multi-frame stacking methods become invalid. The limited revisits prompted us to develop an effective single-image CR processing method for CSST. We retrained the DeepCR model based on CSST simulated images and achieved 97.90+-0.18% recall and 98.67+-0.05% precision on CR detection. Moreover, this paper puts forward an innovative morphology-sensitive inpainting method, which focuses more on areas with higher scientific value. We trained a UNet++ model especially on contaminated stellar/galactic areas, alongside adaptive median filtering for background regions. This method achieves effective for CRs with different intensities and different distances from centers of scientific targets. By this approach, the photometric errors of CR-corrected targets could be restricted to the level comparable to those of uncontaminated sources. Also, it increases the detection rate by 13.6% compared to CR masking. This method will provide a robust CR mitigation for next-generation space telescopes.

astro-ph.IM

Construction of Kondo Chains by Engineering Porphyrin π-Radicals on Au(111)

Quantum manipulation of molecular radical spins provides a crucial platform for exploring emergent phenomena in many-body systems. Here, we combine surface-confined synthesis with scanning tunneling microscopy(STM)tip-induced dehydrogenation to achieve atom-precise engineering of quasi-one-dimensional porphyrin-based Kondo chains (1-7 units) on Au(111). High-resolution STS measurements and low-energy effective modeling collectively demonstrate that π-radicals at each fused-porphyrin unit form Kondo singlets screened by conduction electrons. Adjacent singlets develop direct coherent coupling via quantum-state-overlap-enabled electron tunneling. Crucially, chiral symmetry in the effective model governs zero-mode distribution-present in odd-length chains yet absent in even-length chains-which dictates pronounced odd-even quantum effects in STS spectra of finite chains. Furthermore, the number of parallel porphyrin chains non-monotonically tunes the competition between the Kondo effect and spin exchange, showing opposing trends in strength and demonstrating that both wave-function overlap and the SOMO-LUMO gap collectively govern these interactions. This work simultaneously resolves the dimensional dependence of many-body correlations in confined quantum systems and pioneers approaches for quantum-critical manipulation in molecular spin architectures.

cond-mat.mes-hall

Nonlinear Stability of the Rayleigh-Taylor Problem in Quantum Navier-Stokes Equations

It is well-known that the Rayleigh--Taylor (abbr. RT) instability can be completely inhibited by the quantum effect stabilization in proper circumstances leading to a cutoff wavelength in the \emph{linear} motion equations. Motivated by the linear theory, we further investigate the {stability} for the \emph{nonlinear} RT problem of quantum Navier--Stokes equations in a slab with Navier boundary condition, and rigorously prove the inhibition of RT instability by the quantum effect under a proper setting. More precisely, if the RT density profile $\barρ$ satisfies an additional stabilizing condition, then there is a threshold ${\varepsilon_{c}}$ of the scaled Planck constant, such that if the scaled Planck constant is bigger than ${\varepsilon_{c}}$, the small perturbation solutions around an RT equilibrium state are algebraically stable in time. The mathematical proof is realized by a complicated multi-layer energy method with anisotropic norms of spacial derivatives.

math.AP

SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities by integrating visual and textual inputs, yet modality alignment remains one of the most challenging aspects. Current MLLMs typically rely on simple adapter architectures and pretraining approaches to bridge vision encoders with large language models (LLM), guided by image-level supervision. We identify this paradigm often leads to suboptimal alignment between modalities, significantly constraining the LLM's ability to properly interpret and reason with visual features particularly for smaller language models. This limitation degrades overall performance-particularly for smaller language models where capacity constraints are more pronounced and adaptation capabilities are limited. To address this fundamental limitation, we propose Supervised Embedding Alignment (SEA), a token-level supervision alignment method that enables more precise visual-text alignment during pretraining. SEA introduces minimal computational overhead while preserving language capabilities and substantially improving cross-modal understanding. Our comprehensive analyses reveal critical insights into the adapter's role in multimodal integration, and extensive experiments demonstrate that SEA consistently improves performance across various model sizes, with smaller models benefiting the most (average performance gain of 7.61% for Gemma-2B). This work establishes a foundation for developing more effective alignment strategies for future multimodal systems.

cs.CV

LIGO/Virgo/KAGRA neutron star merger candidate S250206dm: Zwicky Transient Facility observations

We present the searches conducted with the Zwicky Transient Facility (ZTF) in response to S250206dm, a bona fide event with a false alarm rate of one in 25 years, detected by the International Gravitational Wave Network (IGWN). Although the event is significant, the nature of the compact objects involved remains unclear, with at least one likely neutron star. ZTF covered 68% of the localization region, though we did not identify any likely optical counterpart. We describe the ZTF strategy, potential candidates, and the observations that helped rule out candidates, including sources circulated by other collaborations. Similar to Ahumada et al. 2024, we perform a frequentist analysis, using simsurvey, as well as Bayesian analysis, using nimbus, to quantify the efficiency of our searches. We find that, given the nominal distance to this event of 373$\pm$104 Mpc, our efficiencies are above 10% for KNe brighter than $-17.5$ absolute magnitude. Assuming the optical counterpart known as kilonova (KN) lies within the ZTF footprint, our limits constrain the brightest end of the KN parameter space. Through dedicated radiative transfer simulations of KNe from binary neutron star (BNS) and black hole-neutron star (BHNS) mergers, we exclude parts of the BNS KN parameter space. Up to 35% of the models with high wind ejecta mass ($M_{\rm wind} \approx 0.13$ M$_{\odot}$) are ruled out when viewed face-on ($\cosθ_{\rm obs} = 1.0$). Finally, we present a joint analysis using the combined coverage from ZTF and the Gravitational Wave Multimessenger Dark Energy Camera Survey (GW-MMADS). The joint observations cover 73% of the localization region, and the combined efficiency has a stronger impact on rising and slowly fading models, allowing us to rule out 55% of the high-mass KN models viewed face-on.

astro-ph.HE

Kilonova constraints for the LIGO/Virgo/KAGRA neutron star merger candidate S250206dm: GW-MMADS observations

Gravitational wave (GW) neutron star mergers with an associated electromagnetic counterpart constitute powerful probes of binary evolution, the production sites of heavy elements, general relativity, and the expansion of the universe. Only a handful of candidate GW binary mergers during the fourth LIGO/Virgo/KAGRA observing run (O4) so far are believed to include a neutron star. We present optical-near infrared follow-up observations of the candidate neutron-star black hole GW merger S250206dm. This is the first high-significance mass gap neutron star-black hole candidate observed by multiple GW detectors (thus having a significantly smaller sky localization than one-detector events), offering the first opportunity to effectively follow up a GW event of this kind. Our GW MultiMessenger Astronomy DECam Survey (GW-MMADS) campaign consisted of a wide-field search using the Dark Energy Camera (DECam) and T80-South (T80S), as well as galaxy-targeted observations using the Southern Astrophysical Research (SOAR) imager and the Wendelstein 2.1m 3-channel camera. No viable kilonova counterpart was found in our observations. We use our observation depths to place competitive constraints on kilonova models similar to or brighter than the GW170817 kilonova AT 2017gfo within our observed fields, ruling out 100\% of such models with SOAR galaxy-targeted observations and $\sim43$\% (48\%) with DECam (DECam and T80S).

astro-ph.HE

MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence

Spatial perception and reasoning are core components of human cognition, encompassing object recognition, spatial relational understanding, and dynamic reasoning. Despite progress in computer vision, existing benchmarks reveal significant gaps in models' abilities to accurately recognize object attributes and reason about spatial relationships, both essential for dynamic reasoning. To address these limitations, we propose MIRAGE, a multi-modal benchmark designed to evaluate models' capabilities in Counting (object attribute recognition), Relation (spatial relational reasoning), and Counting with Relation. Through diverse and complex scenarios requiring fine-grained recognition and reasoning, MIRAGE highlights critical limitations in state-of-the-art models, underscoring the need for improved representations and reasoning frameworks. By targeting these foundational abilities, MIRAGE provides a pathway toward spatiotemporal reasoning in future research.

cs.CV

Solving Online Resource-Constrained Scheduling for Follow-Up Observation in Astronomy: a Reinforcement Learning Approach

In the astronomical observation field, determining the allocation of observation resources of the telescope array and planning follow-up observations for targets of opportunity (ToOs) are indispensable components of astronomical scientific discovery. This problem is computationally challenging, given the online observation setting and the abundance of time-varying factors that can affect whether an observation can be conducted. This paper presents ROARS, a reinforcement learning approach for online astronomical resource-constrained scheduling. To capture the structure of the astronomical observation scheduling, we depict every schedule using a directed acyclic graph (DAG), illustrating the dependency of timing between different observation tasks within the schedule. Deep reinforcement learning is used to learn a policy that can improve the feasible solution by iteratively local rewriting until convergence. It can solve the challenge of obtaining a complete solution directly from scratch in astronomical observation scenarios, due to the high computational complexity resulting from numerous spatial and temporal constraints. A simulation environment is developed based on real-world scenarios for experiments, to evaluate the effectiveness of our proposed scheduling approach. The experimental results show that ROARS surpasses 5 popular heuristics, adapts to various observation scenarios and learns effective strategies with hindsight.

cs.AI

Atomic-scale observation of $d$-$π$-$d$ spin coupling in coordination structures

Spin coupling between magnetic metal atoms and organic radicals plays a pivotal role in high-performance magnetic materials. The complex interaction involving multi-spin centers in bulk materials makes it challenging to study spin coupling at the atomic scale. Here, we investigate the $d$-$π$-$d$ spin interaction in well-defined metal-organic coordinated structures composed of two iron (Fe) atoms and four all-trans retinoic acid (ReA) molecules, using low-temperature scanning tunneling microscopy and atomic force microscopy. The ReA molecule is turned into a spin-$1/2$ radical state by dehydrogenation, facilitating strong magnetic coupling with the coordinated Fe atoms. Comprehensive theoretical analysis, based on density functional theory and valence bond theory, further elucidates the intrinsic mechanism of ferrimagnetic spin coupling in the coordination structure. Specifically, simultaneous antiferromagnetic coupling of Fe dimer to ReA radicals parallelizes the dimer spin orientation. This work contributes to the fundamental understanding of spin interaction in metal-organic coordination structures and provides microscopic insights for designing advanced magnetic materials.

cond-mat.mtrl-sci

Balancing property optimization and constraint satisfaction for constrained multi-property molecular optimization

Molecular optimization, which aims to discover improved molecules from a vast chemical search space, is a critical step in chemical development. Various artificial intelligence technologies have demonstrated high effectiveness and efficiency on molecular optimization tasks. However, few of these technologies focus on balancing property optimization with constraint satisfaction, making it difficult to obtain high-quality molecules that not only possess desirable properties but also meet various constraints. To address this issue, we propose a constrained multi-property molecular optimization framework (CMOMO), which is a flexible and efficient method to simultaneously optimize multiple molecular properties while satisfying several drug-like constraints. CMOMO improves multiple properties of molecules with constraints based on dynamic cooperative optimization, which dynamically handles the constraints across various scenarios. Besides, CMOMO evaluates multiple properties within discrete chemical spaces cooperatively with the evolution of molecules within an implicit molecular space to guide the evolutionary search. Experimental results show the superior performance of the proposed CMOMO over five state-of-the-art molecular optimization methods on two benchmark tasks of simultaneously optimizing multiple non-biological activity properties while satisfying two structural constraints. Furthermore, the practical applicability of CMOMO is verified on two practical tasks, where it identified a collection of candidate ligands of $β$2-adrenoceptor GPCR and candidate inhibitors of glycogen synthase kinase-3$β$ with high properties and under drug-like constraints.

physics.chem-ph

GRRIS: a real-time intra-site observation scheduling scheme for distributed survey telescope arrays

The distributed telescope array offers promise for conducting large-sky-area, high-frequency time domain surveys. Multiple telescopes can be deployed at each observation site, so intra-site observation task scheduling is crucial for enhancing observation efficiency and quality. Efficient use of observable time and rapid response to special situations are critical to maximize scientific discovery in time domain surveys. Besides, the competing scientific priorities, time-varying observation conditions, and capabilities of observation equipment, lead to a vast search space of the scheduling. So with the increasing number of telescopes and observation fields, balancing computational time with solution quality in observation scheduling poses a significant challenge. Informed by the seminal contributions of earlier studies on a multilevel scheduling model and global scheduler for time domain telescope array, this study is devoted to further exploring the site scheduler. Formulating the observation scheduling of multiple telescopes at the site as a cooperative decision-making problem, this paper proposes GRRIS, a real-time intra-site observation scheduling scheme for telescope array using graph and reinforcement learning. It employs a graph neural network to learn node features that can embed the spatial structure of the observation scheduling. An algorithm based on multi-agent reinforcement learning is designed to efficiently learn the optimum allocation policy of telescope agents to field nodes. Through numerical simulations with real-world scenarios, GRRIS can achieve up to a 22% solution improvement over the most competitive scheme. It offers better scalability and sub-second decision speed, meeting the needs of observation scheduling control for future distributed telescope arrays.

astro-ph.IM

Nonlocal Stokes equation with relaxation on the divergence free equation

In this paper, we consider a new nonlocal approximation to the linear Stokes system with periodic boundary conditions in two and three dimensional spaces . A relaxation term is added to the equation of nonlocal divergence free equation, which is reminiscent to the relaxation of local Stokes equation with small artificial compressibility. Our analysis shows that the well-posedness of the nonlocal system can be established under some mild assumptions on the kernel of nonlocal interactions. Furthermore, the new nonlocal system converges to the conventional, local Stokes system in second order as the horizon parameter of the nonlocal interaction goes to zero. The study provides more theoretical understanding to some numerical methods, such as smoothed particle hydrodynamics, for simulating incompressible viscous flows.

math.AP

On the Inhibition of Rayleigh Taylor Instability by Capillarity in the Navier Stokes Korteweg Model

Bresch--Desjardins--Gisclon--Sart had derived that the capillarity slows down the growth rate of Rayleigh--Taylor (RT) instability in an inhomogeneous incompressible fluid endowed with internal capillarity based on a linearized incompressible Navier--Stokes--Korteweg (NSK) equations in 2008. Later Li--Zhang further obtained another result that the capillarity inhibits RT instability also based on the linearized equations in (SIAM J. Math. Anal. 3287--3315, 2023), if the capillarity coefficient is bigger than some threshold. In this paper, we further rigorously prove such phenomenon of capillarity inhibiting the RT instability in the \emph{nonlinear} incompressible NSK equations in a horizontally periodic slab domain with Navier (slip) boundary conditions. The key idea in the proof is to capture the dissipative estimates of the tangential derivatives of density. Such dissipative estimates result in the decay-in-time of both the velocity and the perturbation density which is very useful to overcome the difficulties arising from the nonlinear terms.

math.AP

A Second-Order Nonlocal Approximation for Manifold Poisson Model with Dirichlet Boundary

Recently, we constructed a class of nonlocal Poisson model on manifold under Dirichlet boundary with global $\mathcal{O}(δ^2)$ truncation error to its local counterpart, where $δ$ denotes the nonlocal horizon parameter. In this paper, the well-posedness of such manifold model is studied. We utilize Poincare inequality to control the lower order terms along the $2δ$-boundary layer in the weak formulation of model. The second order localization rate of model is attained by combining the well-posedness argument and the truncation error analysis. Such rate is currently optimal among all nonlocal models. Besides, we implement the point integral method(PIM) to our nonlocal model through 2 specific numerical examples to illustrate the quadratic rate of convergence on the other side.

math.NA