SearcharxivSearch

arXiv subjects

Meng Zhao

Publications and source records attributed to Meng Zhao.

At least 19 recordsLinked to original sources

Domain-Division based Progressive Learning for Source-Free Domain Adaptation

With growing privacy and portability concerns, source-free domain adaptation requires only a source pre-trained model and an unlabeled target domain, allowing for effective adaptation to the target data. Most existing self-training methods focus on selecting and exploiting samples with reliable predictions, often neglecting others. Inspired by the finding that deep models learn clean samples faster than noisy ones, we propose a domain-division based progressive learning method named DPL. Specifically, our approach consists of two alternating stages, each beginning with the division of the target domain into easy-to-adapt and hard-to-adapt subdomains based on adaptation difficulty, followed by neighborhood-based pseudo label assignment. In stage one, we enhance classification accuracy through uncertainty-aware self-training and alignment of corresponding classes between subdomains. Stage two then applies tailored learning strategies to each subdomain, starting with consistency learning on the easy-to-adapt samples and progressing to utilizing local structural information for the more challenging ones, thereby mining the intrinsic properties of the target data. Extensive experiments on several widely used benchmarks validate the effectiveness of our approach, demonstrating superior performance compared to state-of-the-art methods. Our code is available at https://github.com/iamjingli/DPL.

cs.CV

LFM: Leveraging Foundation Models for Source-Free Universal Domain Adaptation

Source-free universal domain adaptation (SF-UniDA) adapts a pre-trained source model to an unlabeled target domain under both covariate and label shifts, without access to source data. However, existing SF-UniDA methods rely on inefficient techniques such as threshold tuning and clustering. Foundation models (FMs), known for their generalization and zero-shot capabilities, remain underexplored in SF-UniDA. In this paper, we propose a framework that leverages foundation models (LFM) for SF-UniDA. We use a vision-language model (VLM) to compute similarities between target samples and text labels, including those for unknown classes generated by prompting a large language model. The label shift type is determined by analyzing the coefficient of variation of a similarity-based sample-level score. Unknown samples are identified using a binary Gaussian mixture model fitted to another similarity-based metric. Under a consensus strategy, the pseudo-labels generated by the VLM are refined by the target model initialized with the pre-trained source model, integrating knowledge from both the source domain and foundation models. Finally, these refined pseudo-labels are used to train the target model. Extensive experiments across all possible label shifts and multiple benchmarks demonstrate the effectiveness and superiority of our proposed LFM framework. Our code is available at https://github.com/iamjingli/LFM.

cs.CV

DaV-Gen: End-to-End Generative Retrieval via Draft-and-Verify

Mainstream industrial information retrieval systems (e.g., search and recommendation) are usually built upon Multi-Stage Cascade Architectures (MCAs), which balance effectiveness and efficiency through a coarse-to-fine ``retrieval-ranking'' pipeline. However, the optimization objectives across different stages are substantially inconsistent, propagating or even amplifying the early-stage errors that ultimately degrade the quality of final results. While emerging end-to-end generative models offer a potential solution by unifying the pipeline, their online serving performance is severely hindered by the auto-regressive process inherited from the standard decoder-only structure. To bridge this gap, we introduce \textbf{DaV-Gen}, a novel unified solution designed to fundamentally refactor the paradigm for both search and recommendation via a ``Draft-and-Verify'' mechanism. Inspired by the process used by speculative decoding, our framework redesigns the generation task into two synergistic operations within a single model. During training, the model is concurrently optimized for both candidate drafting and fine-grained verification. This is achieved by a composite loss function that jointly trains the model on two distinct but related objectives: 1) a contrastive loss that structures the embedding space for efficient drafting, and 2) a fusion loss that combines generative likelihood with vector similarity to produce a superior verification score. This integrated training strategy equips the model with dual capabilities. At inference time, it first performs highly efficient vector-based drafting to generate a candidate set, and then verifies these candidates using the more powerful fused scoring function, thereby achieving both the speed of sparse drafting and the precision of advanced generative models within a unified, end-to-end architecture.

cs.IR

Delayed blow-up by transport noise for the 3D Navier-Stokes equation with Navier-slip boundary conditions

We study the vorticity formulation of the 3D Navier-Stokes equation driven by transport noise in a periodic channel with Navier-slip boundary conditions. We consider both non-degenerate transport noise and degenerate tangential transport noise. For any prescribed $T>0$ and $\epsilon>0$, we prove that, by choosing the noise intensity sufficiently large and concentrating the noise on sufficiently high modes, the solution exists up to $T$ with probability at least $1-\epsilon$. A main contribution of this work is to identify and analyze the interaction between enhanced dissipation induced by transport noise and physical boundary effects. The no-flux condition breaks the isotropy of the noise and changes the scaling limit of the It\^o-Stratonovich corrector. In the non-degenerate case, a boundary feedback term appears in the limiting effective operator; in the degenerate case, the limiting operator is a nonlocal anisotropic tangential dissipation. The proof is based on a combination of a boundary correction operator, a Meyers-type estimate, a scaling-limit analysis of the It\^o-Stratonovich corrector, and resolvent estimates for the deterministic limiting equations.

math.AP

Distribution and Evolution of the Debris Cloud from the Fragmentation of Intelsat 33E

The breakup of Intelsat 33E on 19 October 2024 posed a potential risk to satellites in the Geostationary Earth Orbit (GEO). This study analyzes the evolution and distribution of these fragments using a probabilistic approach. The initial distribution of the fragments, derived from the NASA Standard Breakup Model, indicates the generation of 4,393 fragments larger than 1 centimeter. The spatial propagation of these fragments is modeled analytically in the Earth-Centered Earth-Fixed reference frame, showing the formation of high-density ring structures in the equatorial plane from 24 hours to 28 days after the breakup. The orbits of 36 cataloged fragments are retrieved and compared with the probability density. Furthermore, Monte Carlo simulations validate the probabilistic model and highlight its efficiency in capturing low-probability events. Collision risks to other GEO satellites are assessed, showing that the top 10\% of satellites encounter a collision probability of up to $10^{-8}$ after 28 days. Satellites near the equatorial plane are at higher risk, whereas those with higher inclinations are less affected. These findings underscore the need for enhanced monitoring and mitigation strategies for GEO breakup events, given the challenges in detecting small fragments.

astro-ph.EP

STRIDE-ED: A Strategy-Grounded Stepwise Reasoning Framework for Empathetic Dialogue Systems

Empathetic dialogue requires not only recognizing a user's emotional state but also making strategy-aware, context-sensitive decisions throughout response generation. However, the lack of a comprehensive empathy strategy framework, explicit task-aligned multi-stage reasoning, and high-quality strategy-aware data fundamentally limits existing approaches, preventing them from effectively modeling empathetic dialogue as a complex, multi-stage cognitive and decision-making process. To address these challenges, we propose STRIDE-ED, a STRategy-grounded, Interpretable, and DEep reasoning framework that models Empathetic Dialogue through structured, strategy-conditioned reasoning. To support effective learning, we develop a strategy-aware data refinement pipeline integrating LLM-based annotation, multi-model consistency-weighted evaluation, and dynamic sampling to construct high-quality training data aligned with empathetic strategies. Furthermore, we adopt a two-stage training paradigm that combines supervised fine-tuning with multi-objective reinforcement learning to better align model behaviors with target emotions, empathetic strategies, and response formats. Extensive experiments demonstrate that STRIDE-ED generalizes across diverse open-source LLMs and consistently outperforms existing methods on both automatic metrics and human evaluations. Our data and code are publicly available at https://github.com/jicoder-nwpu/STRIDE-ED.

cs.CL

O-ConNet: Geometry-Aware End-to-End Inference of Over-Constrained Spatial Mechanisms

Deep learning has shown strong potential for scientific discovery, but its ability to model macroscopic rigid-body kinematic constraints remains underexplored. We study this problem on spatial over-constrained mechanisms and propose O-ConNet, an end-to-end framework that infers mechanism structural parameters from only three sparse reachable points while reconstructing the full motion trajectory, without explicitly solving constraint equations during inference. On a self-constructed Bennett 4R dataset of 42,860 valid samples, O-ConNet achieves Param-MAE 0.276 +/- 0.077 and Traj-MAE 0.145 +/- 0.018 (mean +/- std over 10 runs), outperforming the strongest sequence baseline (LSTM-Seq2Seq) by 65.1 percent and 88.2 percent, respectively. These results suggest that end-to-end learning can capture closed-loop geometric structure and provide a practical route for inverse design of spatial over-constrained mechanisms under extremely sparse observations.

cs.RO

Large-scale Integration of Experimental and Computational Data for 2D Materials

The past decade has seen rapid growth in the number of experimentally realized two-dimensional (2D) materials with diverse chemical and physical properties. However, information on their crystal structure, synthesis routes, and measured or predicted properties, remains scattered across thousands of publications. Here we consolidate this fragmented knowledge by establishing X2DB - an open infrastructure that integrates experimental and computational data on 2D materials. Using extensive literature mining and direct community uploads, we identify 370 unique 2D materials that have been realized in monolayer or few-layer form, and link them to their digital counterparts in computational databases, enabling consistent ab initio characterization of their properties across monolayer, bilayer and bulk forms. We describe the structure and content of the database highlighting its support for community uploads, illustrate how it can be used to generate new scientific insight and introduce a hierarchical classification of the known set of 2D materials. Our work provides a foundation for the integration and cross-fertilization of experimental and theoretical knowledge, opening new avenues for data-driven, predictive synthesis of novel 2D materials.

cond-mat.mtrl-sci

Platform and Framework for Time-Resolved Nanoscale Thermal Transport Measurements in STEM

Understanding heat transport at the nanometer scale is critical for semiconductor devices, quantum materials, and thermal management of nanostructures, yet direct local measurements of thermal conductivity and heat capacity remain scarce. We developed a laser-excitation system integrated into a scanning transmission electron microscope (STEM) for nanoscale thermal transport measurements using ultra-high-resolution electron energy-loss spectroscopy (EELS). A fiber-coupled laser is introduced via a modified aperture mechanism, enabling flexible holder geometries and large tilt angles without optical elements in the polepiece gap. Synchronization of pulsed laser excitation with an externally gated direct electron detector provides temporal resolution about 50 ns at <10 meV energy resolution. Local temperatures are determined via the principle of detailed balance, and thermal transport parameters are extracted by fitting a forward-time central-space heat diffusion model including radiative losses. For amorphous carbon films, we obtain a thermal conductivity of 1.24 $\frac{W}{m\cdot K}$ and a heat capacity of 821 $\frac{J}{kg\cdot K}$, consistent with literature. This framework enables time-resolved nanoscale measurements of thermal transport in materials and devices.

cond-mat.mtrl-sci

Toward Temporal Causal Representation Learning with Tensor Decomposition

Temporal causal representation learning is a powerful tool for uncovering complex patterns in observational studies, which are often represented as low-dimensional time series. However, in many real-world applications, data are high-dimensional with varying input lengths and naturally take the form of irregular tensors. To analyze such data, irregular tensor decomposition is critical for extracting meaningful clusters that capture essential information. In this paper, we focus on modeling causal representation learning based on the transformed information. First, we present a novel causal formulation for a set of latent clusters. We then propose CaRTeD, a joint learning framework that integrates temporal causal representation learning with irregular tensor decomposition. Notably, our framework provides a blueprint for downstream tasks using the learned tensor factors, such as modeling latent structures and extracting causal information, and offers a more flexible regularization design to enhance tensor decomposition. Theoretically, we show that our algorithm converges to a stationary point. More importantly, our results fill the gap in theoretical guarantees for the convergence of state-of-the-art irregular tensor decomposition. Experimental results on synthetic and real-world electronic health record (EHR) datasets (MIMIC-III), with extensive benchmarks from both phenotyping and network recovery perspectives, demonstrate that our proposed method outperforms state-of-the-art techniques and enhances the explainability of causal representations.

cs.LG

Level-3 large deviations for the white-forced 2D Navier-Stokes system in a bounded domain

We study the large deviations principle (LDP) of Donsker-Varadhan type for the white-forced Navier-Stokes system in a bounded domain. Under the assumption that the noise is non-degenerate, we establish level-2 and level-3 LDPs with rate functions given by the Donsker-Varadhan formulas. The proof relies on an improved version of Kifer's criterion, a lift argument inspired from [DV83], an improved abstract result on the large-time asymptotics of generalized Markov semigroups, and a delicate approximation scheme utilizing the resolvent operators of the Markov semigroup.

math.AP

Association between nutritional factors, inflammatory biomarkers and cancer types: an analysis of NHANES data using machine learning

Background. Diet and inflammation are critical factors influencing cancer risk. However, the combined impact of nutritional status and inflammatory biomarkers on cancer status and type, using machine learning (ML), remains underexplored. Objectives. This study investigates the association between nutritional factors, inflammatory biomarkers, and cancer status, and whether these relationships differ across cancer types using National Health and Nutrition Examination Survey (NHANES) data. Methods. We analyzed 24 macro- and micronutrients, C-reactive protein (CRP), and the advanced lung cancer inflammation index (ALI) in 26,409 NHANES participants (2,120 with cancer). Multivariable logistic regression assessed associations with cancer prevalence. We also examined whether these features differed across the five most common cancer types. To evaluate predictive value, we applied three ML models - Logistic Regression, Random Forest, and XGBoost - on the full feature set. Results. The cohort's mean age was 49.1 years; 34.7% were obese. Comorbidities such as anemia and liver conditions, along with nutritional factors like protein and several vitamins, were key predictors of cancer status. Among the models, Random Forest performed best, achieving an accuracy of 0.72. Conclusions. Higher-quality nutritional intake and lower levels of inflammation may offer protective effects against cancer. These findings highlight the potential of combining nutritional and inflammatory markers with ML to inform cancer prevention strategies.

q-bio.QM

Selective Complementary Feature Fusion and Modal Feature Compression Interaction for Brain Tumor Segmentation

Efficient modal feature fusion strategy is the key to achieve accurate segmentation of brain glioma. However, due to the specificity of different MRI modes, it is difficult to carry out cross-modal fusion with large differences in modal features, resulting in the model ignoring rich feature information. On the other hand, the problem of multi-modal feature redundancy interaction occurs in parallel networks due to the proliferation of feature dimensions, further increase the difficulty of multi-modal feature fusion at the bottom end. In order to solve the above problems, we propose a noval complementary feature compression interaction network (CFCI-Net), which realizes the complementary fusion and compression interaction of multi-modal feature information with an efficient mode fusion strategy. Firstly, we propose a selective complementary feature fusion (SCFF) module, which adaptively fuses rich cross-modal feature information by complementary soft selection weights. Secondly, a modal feature compression interaction (MFCI) transformer is proposed to deal with the multi-mode fusion redundancy problem when the feature dimension surges. The MFCI transformer is composed of modal feature compression (MFC) and modal feature interaction (MFI) to realize redundancy feature compression and multi-mode feature interactive learning. %In MFI, we propose a hierarchical interactive attention mechanism based on multi-head attention. Evaluations on the BraTS2019 and BraTS2020 datasets demonstrate that CFCI-Net achieves superior results compared to state-of-the-art models. Code: https://github.com/CDmm0/CFCI-Net

eess.IV

Can MLLMs Generalize to Multi-Party dialog? Exploring Multilingual Response Generation in Complex Scenarios

Current multilingual large language models(MLLMs) still focus on simple question-answering formats, often overlooking more complex dialogue scenarios. In other words, their capabilities of multilingual large models have yet to be validated in dialogue tasks with intricate structures. We therefore ask, Q1: How well do LLMs generalize to more complex dialog scenarios? Q2: Can supervised fine-tuning on a high-quality parallel benchmark restore this ability? Q3: Does the "multilingual complementarity" effect survive in the setting? To answer these questions, we introduce XMP, a high-quality parallel Multilingual dataset sourced from Multi-party Podcast dialogues, which is the first parallel dataset focusing on multi-party dialogue scenarios. Most samples in the dataset feature three or more participants, discussing a wide range of topics. Through extensive experiments, we find that, R1: MLLMs fail to generalize to multi-party setting, R2 Fine-tuning on XMP improves only marginally, with the 70B model achieving at most a 1% absolute gain over its 8B counterpart; R3: Mixing languages during SFT is usually detrimental, with any benefits being marginal and limited to isolated cases in the 70B model.

cs.CL

Advancing Multi-Party Dialogue Framework with Speaker-ware Contrastive Learning

Multi-party dialogues, common in collaborative scenarios like brainstorming sessions and negotiations, pose significant challenges due to their complexity and diverse speaker roles. Current methods often use graph neural networks to model dialogue context, capturing structural dynamics but heavily relying on annotated graph structures and overlooking individual speaking styles. To address these challenges, we propose CMR, a Contrastive learning-based Multi-party dialogue Response generation framework. CMR employs a two-stage self-supervised contrastive learning framework. First, it captures global differences in speaking styles across individuals. Then, it focuses on intra-conversation comparisons to identify thematic transitions and contextually relevant facts. To the best of our knowledge, this is the first approach that applies contrastive learning in multi-party dialogue generation. Experimental results demonstrate that CMR not only significantly outperforms state-of-the-art models, but also generalizes well to large pre-trained language models, effectively enhancing their capability in handling multi-party conversations.

cs.CL

Short-Term Evolution and Risks of Debris Cloud Stemming from Collisions in Geostationary Orbit

The increasing population of objects in geostationary orbit has raised concerns about the potential risks posed by debris clouds resulting from fragmentation. The short-term evolution and associated hazards of debris generated by collisions in the geostationary region is investigated in this study. The initial distribution of two debris clouds is modeled using a single probability density function. The combined distribution of the evolved clouds is determined by solving boundary value problems. The risks associated with these debris clouds are evaluated by calculating the instantaneous impact rate and cumulative collision probability. The probability of collisions with millimeter-sized fragments may increase to 1% within 36 hours, while the probability of collisions with fragments 5 cm or larger is approximately $10^{-5}$. These findings underscore the vulnerability of the geostationary region to space traffic accidents.

astro-ph.EP

Slow, Nanometer Light Confinement Observed in Atomically Thin TaS2

Extreme light confinement down to the atomic scale has been theoretically predicted for ultrathin, Ta-based transition metal dichalcogenides (TMDs). In this work, we experimentally demonstrate in 2H-TaS$_2$ monolayers and bilayers a lateral confinement ratio up to 300 at large wave vectors of $q = 0.15 \, \r{A}^{-1}$, and slow light behaviour with a group velocity $\sim 10^{-4}c$. Quantitative momentum-resolved electron energy loss spectroscopy (q-EELS) with a momentum resolution of $0.0056 \, \r{A}^{-1}$ was used as a platform for the nanoscale optical measurements. With it, momentum-dispersed, two-dimensional (2D) plasmon resonances were experimentally observed, showing a transition from 2D to 3D Coulomb interaction in the high-momentum regime, equivalent to light confinement volumes of $1\text{-}2 \, \text{nm}^3$. Remarkably, the resonant modes do not enter the electron-hole continuum, predicting even further enhanced optical field confinements for this material at cryogenic temperatures.

cond-mat.mtrl-sci

Let's Be Self-generated via Step by Step: A Curriculum Learning Approach to Automated Reasoning with Large Language Models

While Chain of Thought (CoT) prompting approaches have significantly consolidated the reasoning capabilities of large language models (LLMs), they still face limitations that require extensive human effort or have performance needs to be improved. Existing endeavors have focused on bridging these gaps; however, these approaches either hinge on external data and cannot completely eliminate manual effort, or they fall short in effectively directing LLMs to generate high-quality exemplary prompts. To address the said pitfalls, we propose a novel prompt approach for automatic reasoning named \textbf{LBS3}, inspired by curriculum learning which better reflects human learning habits. Specifically, LBS3 initially steers LLMs to recall easy-to-hard proxy queries that are pertinent to the target query. Following this, it invokes a progressive strategy that utilizes exemplary prompts stemmed from easy-proxy queries to direct LLMs in solving hard-proxy queries, enabling the high-quality of the proxy solutions. Finally, our extensive experiments in various reasoning-intensive tasks with varying open- and closed-source LLMs show that LBS3 achieves strongly competitive performance compared to the SOTA baselines.

cs.CL