SearcharxivSearch

arXiv subjects

Mengmeng Zhang

Publications and source records attributed to Mengmeng Zhang.

At least 19 recordsLinked to original sources

Multiple Myeloma Lesion Segmentation on Whole-Body Diffusion-Weighted Imaging via Efficient Anatomical Anticipation and Multimodal Confirmation

Whole-body diffusion-weighted imaging (WB-DWI) is widely used for multiple myeloma (MM) assessment, yet automated lesion segmentation remains challenging due to limited anatomical delineation and the low specificity of marrow hyperintensity. Existing studies have introduced bone region-of-interest (ROI) information and apparent diffusion coefficient (ADC) maps to mitigate these ambiguities, but practical limitations remain. Bone ROI construction often relies on costly manual annotation, image registration, or dedicated bone models, while ADC is usually incorporated only through simple channel fusion, limiting its ability to provide complementary structural and lesion-discriminative cues. To address these limitations, we propose a two-stage framework for MM lesion segmentation on WB-DWI. In the first stage, we train a bone ROI generation model from ADC images without dedicated bone labels, providing an efficient and practical anatomical prior for lesion analysis. In the second stage, we propose Anatomy-guided Multimodal U-Net (AMU-Net), which leverages ADC in a manner consistent with clinical lesion assessment rather than treating it as a generic auxiliary modality. Extensive experiments demonstrate the effectiveness and practicality of the proposed method. It achieves the best overall performance among the evaluated methods, with a mean Dice score of 76.2%.

cs.CV

The discrete homotopy hypothesis for directed graphs

We develop a homotopy theory of directed graphs based on cubical homotopy groups, also known as $A$-groups or reduced GLMY homotopy groups. Localizing the category of directed graphs at morphisms that induce isomorphisms on these groups yields an $\infty$-category, denoted by ${\sf DGra}_\infty$. We prove that ${\sf DGra}_\infty$ is equivalent to the $\infty$-category of spaces, establishing a directed version of the discrete homotopy hypothesis of Carranza and Kapulkin.

math.AT

NTIRE 2026 The Second Challenge on Day and Night Raindrop Removal for Dual-Focused Images: Methods and Results

This paper presents an overview of the NTIRE 2026 Second Challenge on Day and Night Raindrop Removal for Dual-Focused Images. Building upon the success of the first edition, this challenge attracted a wide range of impressive solutions, all developed and evaluated on our real-world Raindrop Clarity dataset~\cite{jin2024raindrop}. For this edition, we adjust the dataset with 14,139 images for training, 407 images for validation, and 593 images for testing. The primary goal of this challenge is to establish a strong and practical benchmark for the removal of raindrops under various illumination and focus conditions. In total, 168 teams have registered for the competition, and 17 teams submitted valid final solutions and fact sheets for the testing phase. The submitted methods achieved strong performance on the Raindrop Clarity dataset, demonstrating the growing progress in this challenging task.

cs.CV

A Generalist Model Including Evolved Star Mass and Age

Determining precise stellar ages and masses for evolved giants is crucial for Galactic archaeology but challenged by spectral degeneracies. Gaia's low-resolution XP spectra offer a unique opportunity to infer these parameters on a massive scale using data-driven methods. We extend a transformer-based astronomical foundation model to evolved stars, establishing a unified framework to simultaneously predict atmospheric parameters ($T_{\mathrm{eff}}$, $\log g$, $[\mathrm{M}/\mathrm{H}]$) and evolutionary labels (mass, age) with physical consistency. Treating spectra as token sequences, we integrated mass and age into the model's vocabulary. The model is trained on Gaia XP spectra cross-matched with the APOGEE DR17 DistMass catalog. Our generative approach enables flexible input handling, including spectral inpainting and parameter-to-spectrum generation. On an independent test set, the model achieves a prediction scatter of $\sigma \approx 0.114 \, M_{\odot}$ for mass and $\sigma \approx 1.334$ Gyr for age. Beyond numerical accuracy, it successfully reproduces the giant branch's mass-luminosity relation and autonomously disentangles interstellar extinction from intrinsic temperature variations without explicit physical priors. It also robustly recovers missing spectral data and estimates reliable uncertainties. Validating that foundation models can internalize stellar physics from data, this physically-aware, probabilistic framework offers a powerful tool for unraveling Milky Way history using large-scale spectroscopic surveys.

astro-ph.SR

PhysFormer: A Physics-Embedded Generative Model for Physically Self-Consistent Spectral Synthesis

In scientific and engineering domains, modeling high-dimensional complex systems governed by partial differential equations (PDEs) remains challenging in terms of physical consistency and numerical stability. However, existing approaches, such as physics-informed neural networks (PINNs), typically rely on known physical fields or coefficients and enforce physical constraints via external loss functions, which can lead to training instability and make it difficult to handle high-dimensional or unobservable scenarios. To this end, we propose PhysFormer, a generative modeling framework that is self-consistent at both the data and physical levels. PhysFormer leverages a low-dimensional, physically interpretable latent space to learn key physical quantities directly from data without requiring known high-dimensional physical field parameters, and embeds the physical process of radiative flux generation within the network to ensure the physical consistency of the generated spectra. In high-dimensional, degenerate inversion tasks, PhysFormer constrains generation within physical limits and enhances spectral fidelity and inversion stability under varying signal-to-noise ratios (SNRs). More broadly, this approach shifts the physical processes from external loss functions into the generative mechanism itself, providing a physically consistent generative modeling paradigm for complex systems involving unknown or unobservable physical quantities.

astro-ph.IM

Algebra of Path Integrals on Digraphs

In this paper, we extend the iterated path integrals from smooth manifolds to digraphs and develop the associated algebraic and geometric structures. Iterated path integrals on a digraph naturally give rise to the iterated path algebra and the iterated loop algebra, both defined as quotient algebras of a shuffle algebra, with the latter carrying a canonical Hopf algebra structure. We construct a non-degenerate pairing between elementarily equivalent classes of loops on a digraph and the iterated loop algebra. By restricting to iterated path integrals that are invariant under $C_\partial$-homotopy, a distinguished subalgebra is obtained which, under this pairing, corresponds to the group algebra of the fundamental group. We further show that this subalgebra is a homotopy invariant and forms a Hopf algebra with involutive antipode.

math.AT

MedGround: Bridging the Evidence Gap in Medical Vision-Language Models with Verified Grounding Data

Vision-Language Models (VLMs) can generate convincing clinical narratives, yet frequently struggle to visually ground their statements. We posit that this limitation arises from the scarcity of high-quality, large-scale clinical referring-localization pairs. To address this, we introduce MedGround, an automated pipeline that transforms segmentation resources into high-quality medical referring grounding data. Leveraging expert masks as spatial anchors, MedGround precisely derives localization targets, extracts shape and spatial cues, and guides VLMs to synthesize natural, clinically grounded queries that reflect morphology and location. To ensure data rigor, a multi-stage verification system integrates strict formatting checks, geometry- and medical-prior rules, and image-based visual judging to filter out ambiguous or visually unsupported samples. Finally, we present MedGround-35K, a novel multimodal medical dataset. Extensive experiments demonstrate that VLMs trained with MedGround-35K consistently achieve improved referring grounding performance, enhance multi-object semantic disambiguation, and exhibit strong generalization to unseen grounding settings. This work highlights MedGround as a scalable, data-driven approach to anchor medical language to verifiable visual evidence.

cs.CV

UniTS: Unified Spatio-Temporal Generative Model for Remote Sensing

One of the primary objectives of satellite remote sensing is to capture the complex dynamics of the Earth environment, which encompasses tasks such as reconstructing continuous cloud-free image sequences, detecting land cover changes, and forecasting future surface evolution. However, existing methods typically require specialized models tailored to different tasks, and lack a general framework that can address these multi-level tasks from a unified perspective. In this paper, we propose a Unified Spatio-Temporal Generative Model (UniTS), which integrates several long-separated core tasks, including time series reconstruction, time series cloud removal, time series semantic change detection, and time series forecasting. Based on the flow matching generative paradigm, UniTS constructs a deterministic evolution path from noise to targets under the guidance of task-specific conditions, achieving unified modeling of spatiotemporal representations for multi-level tasks. The UniTS architecture consists of a diffusion transformer with spatiotemporal blocks, where we design an Adaptive Condition Injector (ACor) to enhance the model's conditional perception of multimodal inputs, enabling high-quality controllable generation. Additionally, we design a Spatiotemporal-aware Modulator (STM) to improve the ability of spatiotemporal blocks to capture complex spatiotemporal dependencies. It substantially outperforms existing specialized models, particularly under challenging conditions such as severe cloud contamination, modality absence, and forecasting complex phenological variations.

cs.CV

The FM Agent

Large language models (LLMs) are catalyzing the development of autonomous AI research agents for scientific and engineering discovery. We present FM Agent, a novel and general-purpose multi-agent framework that leverages a synergistic combination of LLM-based reasoning and large-scale evolutionary search to address complex real-world challenges. The core of FM Agent integrates several key innovations: 1) a cold-start initialization phase incorporating expert guidance, 2) a novel evolutionary sampling strategy for iterative optimization, 3) domain-specific evaluators that combine correctness, effectiveness, and LLM-supervised feedback, and 4) a distributed, asynchronous execution infrastructure built on Ray. Demonstrating broad applicability, our system has been evaluated across diverse domains, including operations research, machine learning, GPU kernel optimization, and classical mathematical problems. FM Agent reaches state-of-the-art results autonomously, without human interpretation or tuning -- 1976.3 on ALE-Bench (+5.2\%), 43.56\% on MLE-Bench (+4.0pp), up to 20x speedups on KernelBench, and establishes new state-of-the-art(SOTA) results on several classical mathematical problems. Beyond academic benchmarks, FM Agent shows considerable promise for both large-scale enterprise R\&D workflows and fundamental scientific research, where it can accelerate innovation, automate complex discovery processes, and deliver substantial engineering and scientific advances with broader societal impact.

cs.AI

Entropy Engineering-Regulated Electron-Phonon Coupling for Highly Efficient Photoluminescence in Se-doped WS2

The limited quantum yield of strained monolayer transition metal dichalcogenides grown by vapor-phase methods and during transfer-based stacking poses a fundamental challenge for their optoelectronic applications. Here, we introduce the concept of "entropy engineering" as a transformative strategy to selectively enhance light-matter interactions through controlled electron-phonon coupling. We unveil how tailored entropy introduced via precise selenium doping or interfacial van der Waals proximity can significantly amplify radiative recombination from momentum-dark excitons in WS2 monolayers. Notably, we discover that slight selenium doping drastically enhances the photoluminescence (PL) of WS2 under strain. While both undoped and heavily doped WS2 suffer from strong PL quenching owing to the direct-to-indirect bandgap transition, lightly Se-doped samples exhibit an order-of-magnitude increase in emission intensity. This counterintuitive boost is traced to doping-induced structural disorder, which intensifies electron-phonon interactions and unlocks efficient phonon-assisted emission from otherwise non-radiative indirect excitons. Moreover, we demonstrate that van der Waals coupling to adjacent Se-doped layers can impart interfacial entropy and further augment PL via proximity effects. Our work highlights entropy engineering via controlled doping as a powerful strategy for activating high-efficiency light emission in atomically thin semiconductors.

cond-mat.mtrl-sci

Local-Antisymmetric Flat Band and Coexisting Correlated stripe charge orders in WSe2-Modulated Twisted Bilayer Graphene

Insulating, atomically flat transition metal dichalcogenides (TMDs) like WSe2 are ideal substrates for probing intrinsic graphene properties. Conventionally, their influence on graphene's band structure is assumed negligible, particularly when small moire patterns form. Combining scanning tunneling microscopy/spectroscopy and theoretical analysis, we reveal that the atomic registry in graphene/WSe2 heterostructures profoundly modulates the electronic structure of magic-angle twisted bilayer graphene (MATBG). At special graphene/WSe2 twist angles, an incommensurate moire superlattice hosts three distinct atomic stacking configurations (A, B, X types). These induce position-dependent potentials that asymmetrically shift MATBG's flat bands, transforming them from hole-side to electron-side asymmetric within a single AA-stacked region. This symmetry breaking enables the unprecedented coexistence of orthogonal stripe charge orders in the correlated regime-a phenomenon previously considered mutually exclusive due to Coulomb repulsion. This band modulation arises from the synergistic effects of the graphene/WSe2 interfacial atomic registry and heterostrain within the MATBG, exhibiting multi-field tunability. Our work establishes interfacial atomic registry as a critical, previously overlooked tuning parameter for flat-band physics, opening avenues to engineer correlated quantum states in van der Waals heterostructures.

cond-mat.mes-hall

SpectralX: Parameter-efficient Domain Generalization for Spectral Remote Sensing Foundation Models

Recent advances in Remote Sensing Foundation Models (RSFMs) have led to significant breakthroughs in the field. While many RSFMs have been pretrained with massive optical imagery, more multispectral/hyperspectral data remain lack of the corresponding foundation models. To leverage the advantages of spectral imagery in earth observation, we explore whether existing RSFMs can be effectively adapted to process diverse spectral modalities without requiring extensive spectral pretraining. In response to this challenge, we proposed SpectralX, an innovative parameter-efficient fine-tuning framework that adapt existing RSFMs as backbone while introducing a two-stage training approach to handle various spectral inputs, thereby significantly improving domain generalization performance. In the first stage, we employ a masked-reconstruction task and design a specialized Hyper Tokenizer (HyperT) to extract attribute tokens from both spatial and spectral dimensions. Simultaneously, we develop an Attribute-oriented Mixture of Adapter (AoMoA) that dynamically aggregates multi-attribute expert knowledge while performing layer-wise fine-tuning. With semantic segmentation as downstream task in the second stage, we insert an Attribute-refined Adapter (Are-adapter) into the first stage framework. By iteratively querying low-level semantic features with high-level representations, the model learns to focus on task-beneficial attributes, enabling customized adjustment of RSFMs. Following this two-phase adaptation process, SpectralX is capable of interpreting spectral imagery from new regions or seasons. The codes will be available from the website: https://github.com/YuxiangZhang-BIT.

cs.CV

On the convergence of PINNs for inverse source problem in the complex Ginzburg-Landau equation

This paper addresses the problem of recovering the spatial profile of the source in the complex Ginzburg-Landau equation from regional observation data at fixed times. We establish two types of sufficient measurements for the unique solvability of the inverse problem. The first is to determine the source term by using whole data at one fixed instant. Conditional stability is established by using the eigenfunction expansion argument. Next, using the analytic continuation method, both uniqueness and a stability estimate for recovering the unknown source can be established from local data at two instants. Finally, algorithms based on the physics-informed neural networks (PINNs) are proposed, and several numerical experiments are presented to show the accuracy and efficiency of the algorithm.

math.AP

Cross-domain Hyperspectral Image Classification based on Bi-directional Domain Adaptation

Utilizing hyperspectral remote sensing technology enables the extraction of fine-grained land cover classes. Typically, satellite or airborne images used for training and testing are acquired from different regions or times, where the same class has significant spectral shifts in different scenes. In this paper, we propose a Bi-directional Domain Adaptation (BiDA) framework for cross-domain hyperspectral image (HSI) classification, which focuses on extracting both domain-invariant features and domain-specific information in the independent adaptive space, thereby enhancing the adaptability and separability to the target scene. In the proposed BiDA, a triple-branch transformer architecture (the source branch, target branch, and coupled branch) with semantic tokenizer is designed as the backbone. Specifically, the source branch and target branch independently learn the adaptive space of source and target domains, a Coupled Multi-head Cross-attention (CMCA) mechanism is developed in coupled branch for feature interaction and inter-domain correlation mining. Furthermore, a bi-directional distillation loss is designed to guide adaptive space learning using inter-domain correlation. Finally, we propose an Adaptive Reinforcement Strategy (ARS) to encourage the model to focus on specific generalized feature extraction within both source and target scenes in noise condition. Experimental results on cross-temporal/scene airborne and satellite datasets demonstrate that the proposed BiDA performs significantly better than some state-of-the-art domain adaptation approaches. In the cross-temporal tree species classification task, the proposed BiDA is more than 3\%$\sim$5\% higher than the most advanced method. The codes will be available from the website: https://github.com/YuxiangZhang-BIT/IEEE_TCSVT_BiDA.

cs.CV

Hierarchical Self-Prompting SAM: A Prompt-Free Medical Image Segmentation Framework

Although the Segment Anything Model (SAM) is highly effective in natural image segmentation, it requires dependencies on prompts, which limits its applicability to medical imaging where manual prompts are often unavailable. Existing efforts to fine-tune SAM for medical segmentation typically struggle to remove this dependency. We propose Hierarchical Self-Prompting SAM (HSP-SAM), a novel self-prompting framework that enables SAM to achieve strong performance in prompt-free medical image segmentation. Unlike previous self-prompting methods that remain limited to positional prompts similar to vanilla SAM, we are the first to introduce learning abstract prompts during the self-prompting process. This simple and intuitive self-prompting framework achieves superior performance on classic segmentation tasks such as polyp and skin lesion segmentation, while maintaining robustness across diverse medical imaging modalities. Furthermore, it exhibits strong generalization to unseen datasets, achieving improvements of up to 14.04% over previous state-of-the-art methods on some challenging benchmarks. These results suggest that abstract prompts encapsulate richer and higher-dimensional semantic information compared to positional prompts, thereby enhancing the model's robustness and generalization performance. All models and codes will be released upon acceptance.

cs.CV

Disentangling Complex Systems: IdopNetwork Meets GLMY Homology Theory

The study of complex systems has captured widespread attention in recent years, emphasizing the exploration of interactions and emergent properties among system units. Network analysis based on graph theory has emerged as a powerful approach for analyzing network topology and functions, making them widely adopted in complex systems. IdopNetwork is an advanced statistical physics framework that constructs the interaction within complex systems by integrating large-scale omics data. By combining GLMY theory, the structural characteristics of the network topology can be traced, providing deeper insights into the dynamic evolution of the network. This combination not only offers a novel perspective for dissecting the internal regulation of complex systems from a holistic standpoint but also provides significant support for applied fields such as data science, complex disease, and materials science.

stat.AP

NTIRE 2025 Challenge on Day and Night Raindrop Removal for Dual-Focused Images: Methods and Results

This paper reviews the NTIRE 2025 Challenge on Day and Night Raindrop Removal for Dual-Focused Images. This challenge received a wide range of impressive solutions, which are developed and evaluated using our collected real-world Raindrop Clarity dataset. Unlike existing deraining datasets, our Raindrop Clarity dataset is more diverse and challenging in degradation types and contents, which includes day raindrop-focused, day background-focused, night raindrop-focused, and night background-focused degradations. This dataset is divided into three subsets for competition: 14,139 images for training, 240 images for validation, and 731 images for testing. The primary objective of this challenge is to establish a new and powerful benchmark for the task of removing raindrops under varying lighting and focus conditions. There are a total of 361 participants in the competition, and 32 teams submitting valid solutions and fact sheets for the final testing phase. These submissions achieved state-of-the-art (SOTA) performance on the Raindrop Clarity dataset. The project can be found at https://lixinustc.github.io/CVPR-NTIRE2025-RainDrop-Competition.github.io/.

cs.CV

Calculating Higher Digraph Homotopy Groups

We give the first tractable and systematic examples of nontrivial higher digraph homotopy groups. To do this we define relative digraph homotopy groups and show these satisfy a long exact sequence analogous to the relative homotopy groups of spaces. We then define digraph suspension and Hurewicz homomorphisms and show they commute with each other. The existence of nontrivial digraph homotopy groups then reduces to the existence of corresponding groups in the degree 1 path homology of digraphs.

math.AT