SearcharxivSearch

arXiv subjects

Jiajun Yang

Publications and source records attributed to Jiajun Yang.

8 recordsLinked to original sources

CloudMamba: An Uncertainty-Guided Dual-Scale Mamba Network for Cloud Detection in Remote Sensing Imagery

Cloud detection in remote sensing imagery is a fundamental, critical, and highly challenging problem. Existing deep learning-based cloud detection methods generally formulate it as a single-stage pixel-wise binary segmentation task with one forward pass. However, such single-stage approaches exhibit ambiguity and uncertainty in thin-cloud regions and struggle to accurately handle fragmented clouds and boundary details. In this paper, we propose a novel deep learning framework termed CloudMamba. To address the ambiguity in thin-cloud regions, we introduce an uncertainty-guided two-stage cloud detection strategy. An embedded uncertainty estimation module is proposed to automatically quantify the confidence of thin-cloud segmentation, and a second-stage refinement segmentation is introduced to improve the accuracy in low-confidence hard regions. To better handle fragmented clouds and fine-grained boundary details, we design a dual-scale Mamba network based on a CNN-Mamba hybrid architecture. Compared with Transformer-based models with quadratic computational complexity, the proposed method maintains linear computational complexity while effectively capturing both large-scale structural characteristics and small-scale boundary details of clouds, enabling accurate delineation of overall cloud morphology and precise boundary segmentation. Extensive experiments conducted on the GF1_WHU and Levir_CS public datasets demonstrate that the proposed method outperforms existing approaches across multiple segmentation accuracy metrics, while offering high efficiency and process transparency. Our code is available at https://github.com/jayoungo/CloudMamba.

cs.CV

MindWatcher: Toward Smarter Multimodal Tool-Integrated Reasoning

Traditional workflow-based agents exhibit limited intelligence when addressing real-world problems requiring tool invocation. Tool-integrated reasoning (TIR) agents capable of autonomous reasoning and tool invocation are rapidly emerging as a powerful approach for complex decision-making tasks involving multi-step interactions with external environments. In this work, we introduce MindWatcher, a TIR agent integrating interleaved thinking and multimodal chain-of-thought (CoT) reasoning. MindWatcher can autonomously decide whether and how to invoke diverse tools and coordinate their use, without relying on human prompts or workflows. The interleaved thinking paradigm enables the model to switch between thinking and tool calling at any intermediate stage, while its multimodal CoT capability allows manipulation of images during reasoning to yield more precise search results. We implement automated data auditing and evaluation pipelines, complemented by manually curated high-quality datasets for training, and we construct a benchmark, called MindWatcher-Evaluate Bench (MWE-Bench), to evaluate its performance. MindWatcher is equipped with a comprehensive suite of auxiliary reasoning tools, enabling it to address broad-domain multimodal problems. A large-scale, high-quality local image retrieval database, covering eight categories including cars, animals, and plants, endows model with robust object recognition despite its small size. Finally, we design a more efficient training infrastructure for MindWatcher, enhancing training speed and hardware utilization. Experiments not only demonstrate that MindWatcher matches or exceeds the performance of larger or more recent models through superior tool invocation, but also uncover critical insights for agent training, such as the genetic inheritance phenomenon in agentic RL.

cs.AI

MCVI-SANet: A lightweight semi-supervised model for LAI and SPAD estimation of winter wheat under vegetation index saturation

Vegetation index (VI) saturation during the dense canopy stage and limited ground-truth annotations of winter wheat constrain accurate estimation of LAI and SPAD. Existing VI-based and texture-driven machine learning methods exhibit limited feature expressiveness. In addition, deep learning baselines suffer from domain gaps and high data demands, which restrict their generalization. Therefore, this study proposes the Multi-Channel Vegetation Indices Saturation Aware Net (MCVI-SANet), a lightweight semi-supervised vision model. The model incorporates a newly designed Vegetation Index Saturation-Aware Block (VI-SABlock) for adaptive channel-spatial feature enhancement. It also integrates a VICReg-based semi-supervised strategy to further improve generalization. Datasets were partitioned using a vegetation height-informed strategy to maintain representativeness across growth stages. Experiments over 10 repeated runs demonstrate that MCVI-SANet achieves state-of-the-art accuracy. The model attains an average R2 of 0.8123 and RMSE of 0.4796 for LAI, and an average R2 of 0.6846 and RMSE of 2.4222 for SPAD. This performance surpasses the best-performing baselines, with improvements of 8.95% in average LAI R2 and 8.17% in average SPAD R2. Moreover, MCVI-SANet maintains high inference speed with only 0.10M parameters. Overall, the integration of semi-supervised learning with agronomic priors provides a promising approach for enhancing remote sensing-based precision agriculture.

cs.CV

Near-field perturbation of laser filament enabling simultaneous far-field THz diagnosis and broadband calculus processing

Terahertz (THz) wave manipulation based on laser filaments-plasma channels formed by femtosecond laser-induced air ionization-has emerged as a promising platform for free-space THz applications. However, in-situ characterization of the spatially confined THz modes within filaments faces significant challenges due to the plasma's ultra-high intensity, which not only hinders direct near-field probing but also limits reliance on indirect far-field reconstruction. Here, we introduce a non-invasive near-field modulation scheme where a metal plate approaches the filament at submillimeter distances (comparable to THz wavelengths), perturbing the dielectric environment to convert the symmetric annular THz mode into an asymmetric state. This controlled transition enables far-field detection of broadband calculus behaviors (first- and second-order differentiation/integration) on time-domain THz waveforms and characteristic spectral transfer functions with 1/f, 1/f^2, f or f^2 dependency (where f is the THz frequency), thereby diagnosing the near-field THz mode confinement. Hence, the proposed approach synergizes near-field modulation efficiency with far-field detection robustness, advancing fundamental understanding of plasma-THz interactions and enabling novel all-optical signal processing for filament-based THz technologies.

physics.optics

LongCat-Flash Technical Report

We introduce LongCat-Flash, a 560-billion-parameter Mixture-of-Experts (MoE) language model designed for both computational efficiency and advanced agentic capabilities. Stemming from the need for scalable efficiency, LongCat-Flash adopts two novel designs: (a) Zero-computation Experts, which enables dynamic computational budget allocation and activates 18.6B-31.3B (27B on average) per token depending on contextual demands, optimizing resource usage. (b) Shortcut-connected MoE, which enlarges the computation-communication overlap window, demonstrating notable gains in inference efficiency and throughput compared to models of a comparable scale. We develop a comprehensive scaling framework for large models that combines hyperparameter transfer, model-growth initialization, a multi-pronged stability suite, and deterministic computation to achieve stable and reproducible training. Notably, leveraging the synergy among scalable architectural design and infrastructure efforts, we complete model training on more than 20 trillion tokens within 30 days, while achieving over 100 tokens per second (TPS) for inference at a cost of \$0.70 per million output tokens. To cultivate LongCat-Flash towards agentic intelligence, we conduct a large-scale pre-training on optimized mixtures, followed by targeted mid- and post-training on reasoning, code, and instructions, with further augmentation from synthetic data and tool use tasks. Comprehensive evaluations demonstrate that, as a non-thinking foundation model, LongCat-Flash delivers highly competitive performance among other leading models, with exceptional strengths in agentic tasks. The model checkpoint of LongCat-Flash is open-sourced to foster community research. LongCat Chat: https://longcat.ai Hugging Face: https://huggingface.co/meituan-longcat GitHub: https://github.com/meituan-longcat

cs.CL

Revealing the terahertz-laser velocity effect during air filamentation via travelling-wave-antenna model

During femtosecond laser filamentation in air, the velocity ratio (K) between the terahertz (THz) phase velocity and the laser group velocity plays a crucial role in THz waves generation. However, K is typically assumed to be unity and its impact has been long overlooked due to the more attention paid to the more easily controlled filament length. Here, we investigate the obscured contribution of K to the THz radiation characteristics by using the improved travelling-wave-antenna (TWA) model. It has been found that, under both single- and two-color laser pumping schemes, K significantly determines the far-field spatial distribution of forward or backward THz radiation, as well as a transition from Bessel- to Cherenkov-type THz emission patterns. These results establish the TWA model as a reliable theoretical tool for studying the mechanisms of THz beam shaping via the designed K. Moreover, for cases of K not being controlled, its value can also be inferred by the proposed TWA model, which could be an effective method to confirm whether the laser ionization front is superluminal or subluminal compared with the generated THz waves.

physics.optics

Progressive Scale-aware Network for Remote sensing Image Change Captioning

Remote sensing (RS) images contain numerous objects of different scales, which poses significant challenges for the RS image change captioning (RSICC) task to identify visual changes of interest in complex scenes and describe them via language. However, current methods still have some weaknesses in sufficiently extracting and utilizing multi-scale information. In this paper, we propose a progressive scale-aware network (PSNet) to address the problem. PSNet is a pure Transformer-based model. To sufficiently extract multi-scale visual features, multiple progressive difference perception (PDP) layers are stacked to progressively exploit the differencing features of bitemporal features. To sufficiently utilize the extracted multi-scale features for captioning, we propose a scale-aware reinforcement (SR) module and combine it with the Transformer decoding layer to progressively utilize the features from different PDP layers. Experiments show that the PDP layer and SR module are effective and our PSNet outperforms previous methods. Our code is public at https://github.com/Chen-Yang-Liu/PSNet

cs.CV

Electronic Structure Topology Associated Domain is Useful to Minimize the Uncertainty of QM/MM Boundary Charge Transfer Effects

The charge transfer effect is an important component in the physical description of realistic proteins. In the hybrid quantum mechanical-molecular mechanical (QM/MM) simulations, the significant charge transfer between the QM/MM boundaries could lead to the slow convergence problem with very large QM regions. In this work, we will discuss how community structure in complex network (electronic structure topology associated domain) can be used to measure QM/MM boundary effects and how it can provide a different perspective on well established concepts in available QM/MM simulations. The graph theory is employed to provide an useful solution to distinguish the significant charge transfer regions between the core active site and the surrounding protein environment in the QM/MM simulations. According to community detection algorithm for complex network, the charge transfer topology associated domain (ctTAD) is suggested to provide an alternative tool for systematically minimizing the charge transfer effects of QM/MM boundary with minor computational costs.

physics.chem-ph