SearcharxivSearch

arXiv subjects

Meng Zhou

Publications and source records attributed to Meng Zhou.

At least 19 recordsLinked to original sources

LOOMSUM:Weaving Quantitative and Narrative Evidence for Faithful Long Text-Table Summarization

Long documents often distribute important information across extensive narrative passages and multiple tables, making faithful summarization particularly challenging. Existing methods may generate individually supported quantitative facts and analytical statements yet associate them incorrectly, producing quantitatively plausible yet analytically unfaithful summaries. In this work, we propose LOOMSUM, a training-free framework that extracts source-grounded atomic evidence, explicitly links table-derived facts with supporting narrative analyses, and plans the discourse structure before generation. We also introduce Table-Grounded Faithfulness (TGF), a claim-level metric that separately evaluates Numeric Grounding, Analysis Support, and Relation Consistency. Experiments on the text--table summarization benchmarks FINDSum and USTT show that LOOMSUM improves analytical faithfulness while maintaining strong summarization quality. Human evaluation finds positive component-level associations with the corresponding human judgments. Our Relation Consistency metric further shows stronger agreement with human relation judgments than generic factuality metrics, indicating that explicit cross-modal linking helps reduce errors in which supported quantities are paired with incorrect narrative interpretations. Together, these findings show that faithful long text--table summarization requires not only grounding individual facts, but also preserving the relations between them.

cs.CL

Beyond Representation Learning: A Systematic Study of Joint-Embedding Predictive Generation for 3D Brain MRI

Joint-embedding predictive architectures (JEPAs) have primarily been developed for self-supervised representation learning. Denoising JEPA (D-JEPA) recently demonstrated strong generative capabilities on natural images, yet the applicability to 3D medical imaging remains unexplored. Building on the D-JEPA framework, we present Med-D-JEPA, a systematic adaptation and evaluation of joint-embedding predictive generation for 3D brain MRI. Med-D-JEPA operates on continuous latent tokens produced by a 3D KL-regularized adversarial variational autoencoder, and combines masked context prediction, representation-level alignment, per-token diffusion, and iterative next-set-of-token sampling. We evaluate unconditional and class-conditional generation quality on BraTS2019 and OASIS-1 datasets; downstream classification utility; and preliminary whole-tumor segmentation on BraTS2020. Across different generation settings, Med-D-JEPA achieves superior or competitive performance compared to several strong baselines on fidelity and diversity metrics. Compared to training with real samples, Med-D-JEPA-based synthetic pretraining improves classification AUC from 0.63 to 0.85 on BraTS2019 and from 0.78 to 0.87 on OASIS-1. In the segmentation study, pretraining on Med-D-JEPA samples improves Dice from 0.74 to 0.80 and reduces HD95 from 13.40 to 9.56 mm. These findings establish joint-embedding predictive generation as a promising direction for 3D medical image synthesis and encourage further research in this direction.

cs.CV

ShiJianBench: From Dialogue to Decision for Long-Horizon Evaluation of Investment Advisors

Conversational investment advisors influence not only what users know, but also how they make subsequent decisions as market conditions evolve. Existing evaluations primarily assess response quality or observed outcomes, leaving the long-horizon pathway from advisor language to investor behavior difficult to audit. We introduce ShiJianBench, an offline framework for evaluating conversational investment advisors through matched investor trajectories under fixed historical market feedback. At its core is a multi-agent investor simulator with explicit evolving state variables, motive-driven deliberation, long-term memory, and dialogue-grounded updates. The simulator is calibrated against aggregate behavioral patterns from 7,199 real users, and advisor policies are evaluated using separate investor-side, service-side, and content-side metrics under a hard compliance gate. Experiments on Chinese fund-market traces from 2021 to 2026 identify a stable leading group of LLM advisors that combines substantially stronger personalized content with competitive investor-side trajectory outcomes. These results reveal a systematic distinction between producing a high-quality response and delivering an effective long-horizon intervention, motivating trajectory-aware evaluation of conversational advisors.

cs.CL

Synthesis imaging with a lunar orbit array: III. Augmented lagrangian Multiplier Imaging using Gradient descent Optimization (AMIGO)

Ground-based radio observations below 30 MHz are severely limited by ionospheric interference and radio frequency interference (RFI) from Earth. A lunar-orbiting radio interferometer mission, the Discovering the Sky at the Longest wavelength (DSL, also known by its Chinese name ``Hongmeng''), has been proposed to overcome these obstacles. However, for such a mission, there are new challenges, such as the nearly all-sky field of view and dynamic 3D baselines, which require a huge computational cost for interferometric image reconstruction. In this work, we present AMIGO (Augmented lagrangian Multiplier Imaging using Gradient descent Optimization), a novel imaging algorithm tailored to lunar-orbiting arrays like DSL, combining the Mini-Batch Gradient Descent (MBGD) method with the Augmented Lagrangian Multiplier (ALM) technique. MBGD reduces the computational complexity and memory cost, enabling efficient handling of large datasets. ALM flexibly incorporates physical priors like non-negative sky temperature and prior angular power spectrum into the imaging algorithm, with adjustable stopping criteria to quantitatively control prior strength. We validate AMIGO using mock visibility data generated under realistic DSL orbit configurations. Reconstructed sky maps at various frequencies and spatial resolutions show that this approach provides a computationally feasible framework for all-sky imaging with a lunar-orbiting array.

astro-ph.IM

HiMatch-AD: DINOv3-driven Hierarchical Matching for Training-free Medical Anomaly Detection

Anomaly detection is essential for medical image analysis, where pathological regions often appear as rare deviations from normal anatomical structures. While training-based methods have achieved promising performance, they require task-specific optimization and extensive normal data, which limits scalability across modalities and institutions. Training-free approaches offer greater flexibility by leveraging pretrained visual representations, yet existing methods typically rely on simple nearest-neighbor retrieval and naive aggregation strategies, which may fail to capture hierarchical semantics and ignore the reliability of multiple anomaly responses. In this work, we propose HiMatch-AD, a DINOv3-driven hierarchical matching framework for training-free medical anomaly detection. Our method first retrieves semantically relevant normal references via dual-branch matching that jointly considers global CLS-token similarity and patch-level representations. Hierarchical anomaly maps are then generated across multiple transformer stages by comparing clustered normal features with query representations. To robustly aggregate anomaly responses, we introduce a unified uncertainty-based fusion mechanism that adaptively weights maps according to their reliability. The entire framework operates without any task-specific training. Extensive experiments on the BMAD benchmark, including brain MRI, liver CT, and retinal OCT datasets, demonstrate that HiMatch-AD consistently outperforms both training-based and DINO-based state-of-the-art methods, which highlights the effectiveness of multi-level matching and uncertainty-aware fusion for scalable medical anomaly detection.

cs.CV

Organizational Control Layer: Governance Infrastructure at the Execution Boundary of LLM Agent Systems

LLM-based agents are increasingly deployed in workflows where generated outputs may trigger state-changing actions, such as price offers, refunds, payments, or tool calls. This creates an execution-boundary problem: a platform must decide whether an agent's proposed action is authorized before the action is executed. We introduce the Organizational Control Layer (OCL), a model-agnostic governance layer that separates proposal generation from environment-facing execution. OCL intercepts generated actions, checks them against role, policy, and economic constraints, and either approves, revises, blocks, or escalates them without modifying the underlying LLM generator. We evaluate OCL on adversarial buyer--seller negotiation environments adapted from AgenticPay. Across multiple frontier LLM backends, OCL reduces observed unsafe executions from 88% to 0% while increasing valid success from 12% to 96%. Ablations show that this gain comes from combining pre-execution enforcement with structured recovery, rather than from prompting or blocking alone. These results suggest that deployment-grade LLM agent systems require explicit governance at the boundary between language generation and executable action.

cs.MA

Synthesis imaging with a lunar orbit array: II. Impacts of instrument-induced phase errors

A lunar orbit interferometer array suffers from a number of systematics. Beyond systematics induced by the imaging algorithm itself and thermal noise considered in Paper I, phase errors due to instrumental inconsistency between receivers, geometric error in baseline determination, and clock synchronization error between satellites will also affect synthesis imaging with the space array. In this paper, we model different sources of phase errors and quantify their impacts on all-sky and patchy-sky map-making, respectively, for the ultra-long wavelength sky ($f\lesssim30$ MHz), using the Discovering the Sky at the Longest wavelength (DSL) mission (also known as the Hongmeng mission) as an example. We find that in the scheme of all-sky imaging, the angular power spectrum can be suppressed uniformly for various sources of phase errors. To ensure a reconstruction of large-scale structures with $\gtrsim 95\%$ of the angular power spectrum, the phase error should be controlled below $\sim 12^\circ$ on the random instrumental component, or below $\sim 12^\circ$ for constant deviation, or below $1.1$ ns on the temporal component. With multiple baseline measurements, the baseline determination errors below $1$ m can also meet the requirement. In the scheme of patchy-sky imaging, the S/N of point source detections does not change significantly, except with instrumental phase errors or at high frequencies. The impact of geometric phase error is relatively stronger in the patchy-sky imaging with higher resolution because longer baselines are used and fewer times of baseline measurements can be averaged over within an integration time. When scaled with wavelength, these results set the basic reference for instrumental requirements for future space interferometers.

astro-ph.IM

FinToolSyn: A forward synthesis Framework for Financial Tool-Use Dialogue Data with Dynamic Tool Retrieval

Tool-use capabilities are vital for Large Language Models (LLMs) in finance, a domain characterized by massive investment targets and data-intensive inquiries. However, existing data synthesis methods typically rely on a reverse synthesis paradigm, generating user queries from pre-sampled tools. This approach inevitably introduces artificial explicitness, yielding queries that fail to capture the implicit, event-driven nature of real-world needs. Moreover, its reliance on static tool sets overlooks the dynamic retrieval process required to navigate massive tool spaces. To address these challenges, we introduce \textit{FinToolSyn}, a forward synthesis framework designed to generate high-quality financial dialogues. Progressing from persona instruction and atomic tool synthesis to dynamic retrieval dialogue generation, our pipeline constructs a repository of 43,066 tools and synthesizes over 148k dialogue instances, incorporating dynamic retrieval to emulate the noisy candidate sets typical of massive tool spaces. We also establish a dedicated benchmark to evaluate tool-calling capabilities in realistic financial scenarios. Extensive experiments demonstrate that models trained on FinToolSyn achieve a 21.06\% improvement, providing a robust foundation for tool learning in financial scenarios.

cs.CL

Optimizing 3D Diffusion Models for Medical Imaging via Multi-Scale Reward Learning

Diffusion models have emerged as powerful tools for 3D medical image generation, yet bridging the gap between standard training objectives and clinical relevance remains a challenge. This paper presents a method to enhance 3D diffusion models using Reinforcement Learning (RL) with multi-scale feedback. We first pretrain a 3D diffusion model on MRI volumes to establish a robust generative prior. Subsequently, we fine-tune the model using Proximal Policy Optimization (PPO), guided by a novel reward system that integrates both 2D slice-wise assessments and 3D volumetric analysis. This combination allows the model to simultaneously optimize for local texture details and global structural coherence. We validate our framework on the BraTS 2019 and OASIS-1 datasets. Our results indicate that incorporating RL feedback effectively steers the generation process toward higher quality distributions. Quantitative analysis reveals significant improvements in Fr\'echet Inception Distance (FID) and, crucially, the synthetic data demonstrates enhanced utility in downstream tumor and disease classification tasks compared to non-optimized baselines.

cs.CV

Introducing LongCat-Flash-Thinking: A Technical Report

We present LongCat-Flash-Thinking, an efficient 560-billion-parameter open-source Mixture-of-Experts (MoE) reasoning model. Its advanced capabilities are cultivated through a meticulously crafted training process, beginning with long Chain-of-Thought (CoT) data cold-start and culminating in large-scale Reinforcement Learning (RL). We first employ a well-designed cold-start training strategy, which significantly enhances the reasoning potential and equips the model with specialized skills in both formal and agentic reasoning. Then, a core innovation is our domain-parallel training scheme, which decouples optimization across distinct domains (e.g., STEM, Code, Agentic) and subsequently fuses the resulting expert models into a single, nearly Pareto-optimal model. This entire process is powered by our Dynamic ORchestration for Asynchronous rollout (DORA) system, a large-scale RL framework that delivers a greater than threefold training speedup over synchronous methods on tens of thousands of accelerators. As a result, LongCat-Flash-Thinking achieves state-of-the-art performance among open-source models on a suite of complex reasoning tasks. The model exhibits exceptional efficiency in agentic reasoning, reducing average token consumption by 64.5% (from 19, 653 to 6, 965) on AIME-25, without degrading task accuracy. We release LongCat-Flash-Thinking to promote further advances in reasoning systems and agentic AI research.

cs.AI

LongCat-Flash Technical Report

We introduce LongCat-Flash, a 560-billion-parameter Mixture-of-Experts (MoE) language model designed for both computational efficiency and advanced agentic capabilities. Stemming from the need for scalable efficiency, LongCat-Flash adopts two novel designs: (a) Zero-computation Experts, which enables dynamic computational budget allocation and activates 18.6B-31.3B (27B on average) per token depending on contextual demands, optimizing resource usage. (b) Shortcut-connected MoE, which enlarges the computation-communication overlap window, demonstrating notable gains in inference efficiency and throughput compared to models of a comparable scale. We develop a comprehensive scaling framework for large models that combines hyperparameter transfer, model-growth initialization, a multi-pronged stability suite, and deterministic computation to achieve stable and reproducible training. Notably, leveraging the synergy among scalable architectural design and infrastructure efforts, we complete model training on more than 20 trillion tokens within 30 days, while achieving over 100 tokens per second (TPS) for inference at a cost of \$0.70 per million output tokens. To cultivate LongCat-Flash towards agentic intelligence, we conduct a large-scale pre-training on optimized mixtures, followed by targeted mid- and post-training on reasoning, code, and instructions, with further augmentation from synthetic data and tool use tasks. Comprehensive evaluations demonstrate that, as a non-thinking foundation model, LongCat-Flash delivers highly competitive performance among other leading models, with exceptional strengths in agentic tasks. The model checkpoint of LongCat-Flash is open-sourced to foster community research. LongCat Chat: https://longcat.ai Hugging Face: https://huggingface.co/meituan-longcat GitHub: https://github.com/meituan-longcat

cs.CL

ClinicalFMamba: Advancing Clinical Assessment using Mamba-based Multimodal Neuroimaging Fusion

Multimodal medical image fusion integrates complementary information from different imaging modalities to enhance diagnostic accuracy and treatment planning. While deep learning methods have advanced performance, existing approaches face critical limitations: Convolutional Neural Networks (CNNs) excel at local feature extraction but struggle to model global context effectively, while Transformers achieve superior long-range modeling at the cost of quadratic computational complexity, limiting clinical deployment. Recent State Space Models (SSMs) offer a promising alternative, enabling efficient long-range dependency modeling in linear time through selective scan mechanisms. Despite these advances, the extension to 3D volumetric data and the clinical validation of fused images remains underexplored. In this work, we propose ClinicalFMamba, a novel end-to-end CNN-Mamba hybrid architecture that synergistically combines local and global feature modeling for 2D and 3D images. We further design a tri-plane scanning strategy for effectively learning volumetric dependencies in 3D images. Comprehensive evaluations on three datasets demonstrate the superior fusion performance across multiple quantitative metrics while achieving real-time fusion. We further validate the clinical utility of our approach on downstream 2D/3D brain tumor classification tasks, achieving superior performance over baseline methods. Our method establishes a new paradigm for efficient multimodal medical image fusion suitable for real-time clinical deployment.

eess.IV

Libra: Assessing and Improving Reward Model by Learning to Think

Reinforcement learning (RL) has significantly improved the reasoning ability of large language models. However, current reward models underperform in challenging reasoning scenarios and predominant RL training paradigms rely on rule-based or reference-based rewards, which impose two critical limitations: 1) the dependence on finely annotated reference answer to attain rewards; and 2) the requirement for constrained output format. These limitations fundamentally hinder further RL data scaling and sustained enhancement of model reasoning performance. To address these limitations, we propose a comprehensive framework for evaluating and improving the performance of reward models in complex reasoning scenarios. We first present a reasoning-oriented benchmark (Libra Bench), systematically constructed from a diverse collection of challenging mathematical problems and advanced reasoning models, to address the limitations of existing reward model benchmarks in reasoning scenarios. We further introduce a novel approach for improving the generative reward model via learning-to-think methodologies. Based on the proposed approach, we develop Libra-RM series, a collection of generative reward models with reasoning capabilities that achieve state-of-the-art results on various benchmarks. Comprehensive downstream experiments are conducted and the experimental results demonstrate the correlation between our Libra Bench and downstream application, and the potential of Libra-RM to further improve reasoning models with unlabeled data.

cs.CL

Fragmented quantum phases in anti-blockade regime of Rydberg atom array

In this Letter, we report a parameter-dependent Hilbert space fragmentation in a one-dimensional Rydberg atom array under anti-blockade conditions. We identify distinct non-equilibrium dynamical phases and show that their quasi-periodic behavior arises from the interplay of multi-path interference and multi-photon cascaded excitations, reflecting fundamental differences in state connectivity and effective dimensionality. By constructing a complete phase diagram, we reveal the transition from thermalization to fragmentation, and further demonstrate a secondary fragmentation process enabled by local constraints. Our results highlight the potential of anti-blockade mechanisms for programmable non-thermal dynamics in many-body systems.

quant-ph

Robust Extraction of Global 21 cm Spectrum from Experiments with a Chromatic Beam Based on Physics-Motivated Error Modeling

The extraction of the sky-averaged 21 cm signal from Cosmic Dawn and the Epoch of Reionization faces significant challenges. The bright and anisotropic Galactic foreground, which is 4 - 5 orders of magnitude brighter than the 21 cm signal, when convolved with the inevitably chromatic beam, introduces additional spectral structures that can easily mimic the real 21 cm signal. In this paper, we investigate the signal extraction for a lunar-orbit experiment, where the antenna moves fast in orbit and data from multiple orbits have to be used. We propose a physics-motivated and correlated modeling of both the foreground and the measurement errors. By dividing the sky into multiple regions according to the spectral index distribution and accounting for the full covariance of modeling errors, we jointly fit both the foreground and the 21 cm signal using simulated data for the Discovering the Sky at the Longest wavelength lunar orbit experiment. This method successfully extracts the 21 cm signals of various amplitudes from the simulated data even for a testing antenna with a relatively high level of chromaticity. This approach, which is robust against moderate beam chromaticity, significantly relaxes the stringent design and manufacturing requirements for the antenna, offering a practical solution for future 21 cm global signal experiments either on the ground or in space.

astro-ph.IM

MulDimIF: A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models

Instruction following refers to the ability of large language models (LLMs) to generate outputs that satisfy all specified constraints. Existing research has primarily focused on constraint categories, offering limited evaluation dimensions and little guidance for improving instruction-following abilities. To address this gap, we introduce MulDimIF, a multi-dimensional constraint framework encompassing three constraint patterns, four constraint categories, and four difficulty levels. Based on this framework, we design a controllable instruction generation pipeline. Through constraint expansion, conflict detection, and instruction rewriting, we construct 9,106 code-verifiable samples. We evaluate 18 LLMs from six model families and find marked performance differences across constraint settings. For instance, average accuracy decreases from 80.82% at Level I to 36.76% at Level IV. Moreover, training with data generated by our framework significantly improves instruction following without compromising general performance. In-depth analysis indicates that these gains stem largely from parameter updates in attention modules, which strengthen constraint recognition and adherence. Code and data are available in https://github.com/Junjie-Ye/MulDimIF.

cs.CL

A flexible Bayesian framework for detecting cross-sample spatial expression variability in heterogeneous tissues

Spatial transcriptomics measures gene expression alongside the spatial coordinates of each capture spot or cell across tissue samples. The detection of spatially variable (SV) genes, whose expression exhibits systematic spatial variation, enables the delineation of functional tissue domains and the identification of region-specific transcriptional changes that underlie disease heterogeneity. However, empirical evidence from spatial transcriptomics data indicates that existing methods are hindered by two practical issues: reliance on predefined spatial patterns that often miss tissue complexity, and a lack of standardized approaches for multi-sample integration, which undermines reproducibility and biological interpretability. To address these issues, we propose a novel integrated Bayesian hierarchical model that combines flexible nonparametric spatial modeling with information sharing across samples. The model uses an adaptive spatial process that can capture a wide range of spatial patterns while remaining interpretable. We also introduce a new prior that borrows strength across samples, enabling robust detection of SV genes from multiple tissue sections. An efficient variational approximation is developed for scalable posterior computation. Analyzing spatial transcriptomics data from human brain and skin cancer tissues, our framework identifies spatially structured SV genes, enabling the delineation of tissue domains and the discovery of functionally coherent gene clusters through pathway enrichment analysis.

stat.AP

Dynamic Attention Mechanism in Spatiotemporal Memory Networks for Object Tracking

Mainstream visual object tracking frameworks predominantly rely on template matching paradigms. Their performance heavily depends on the quality of template features, which becomes increasingly challenging to maintain in complex scenarios involving target deformation, occlusion, and background clutter. While existing spatiotemporal memory-based trackers emphasize memory capacity expansion, they lack effective mechanisms for dynamic feature selection and adaptive fusion. To address this gap, we propose a Dynamic Attention Mechanism in Spatiotemporal Memory Network (DASTM) with two key innovations: 1) A differentiable dynamic attention mechanism that adaptively adjusts channel-spatial attention weights by analyzing spatiotemporal correlations between the templates and memory features; 2) A lightweight gating network that autonomously allocates computational resources based on target motion states, prioritizing high-discriminability features in challenging scenarios. Extensive evaluations on OTB-2015, VOT 2018, LaSOT, and GOT-10K benchmarks demonstrate our DASTM's superiority, achieving state-of-the-art performance in success rate, robustness, and real-time efficiency, thereby offering a novel solution for real-time tracking in complex environments.

cs.CV