SearcharxivSearch

arXiv subjects

Xiaowen Wang

Publications and source records attributed to Xiaowen Wang.

10 recordsLinked to original sources

RAFT: A Stateful Retrieval-Augmented Framework for Troubleshooting Agents

Effective troubleshooting agents in enterprise customer support depend on retrieving actionable guidance from similar historical cases, yet existing retrieval-augmented generation (RAG) systems treat support cases as static documents and overlook their multi-stage, stateful nature. We introduce RAFT (Retrieval-Augmented Framework for Troubleshooting Agents), a stateful RAG framework that abstracts each closed historical case into a directed chain of timeline entries and retrieves at the entry level, surfacing cases whose intermediate states match the active case and returning the parent-case trajectory anchored at the matched state; an optional case-level graph links cases through a configurable similarity representation. We evaluate this retrieval layer directly, which, unlike evaluating a full agent system, requires no production deployment. Because public multi-stage troubleshooting data is extremely rare, we pair a synthetic benchmark built from Microsoft Learn Windows Server documentation with real Apache Jira issues carrying human-created duplicate labels. RAFT improves Case Hit over vanilla RAG and GraphRAG baselines at every stage of case progress, with statistically significant gains over the strongest baseline; the Jira results provide directional evidence that the advantage transfers to real case histories. We release our benchmark, implementation, and the Apache Jira evaluation set.

cs.AI

A high-order Newton multigrid method with a simplified Jacobian for steady-state shallow water equations

A high-order Newton multigrid method is proposed for steady-state shallow water flows in open channels with regular and irregular geometries. The method integrates a finite volume discretization with third-order weighted essentially non-oscillatory (WENO) reconstruction and a Newton multigrid framework with an efficient approximation of the Jacobian matrix for solving the resulting discrete system. In high-order schemes, the computational cost of Jacobian construction becomes dominant due to the wide stencil. Meanwhile, only a small fraction of the non-zero Jacobian entries exhibit large magnitudes. Based on this observation, a simplified Jacobian approximation is introduced using reduced stencils, in which selected off-stencil contributions are neglected, thereby achieving a substantial reduction in computational cost. The proposed approach is verified numerically to show significant efficiency improvement while maintaining comparable convergence behavior to that obtained with the full Jacobian approach. To further enhance performance, a geometric multigrid method incorporating a successive over-relaxation iteration as the smoother is applied to solve the linear systems arising in each Newton step. A variety of numerical experiments, including a one-dimensional smooth subcritical flow, flows over a hump, and a two-dimensional hydraulic jump over a wedge, are carried out to illustrate the third-order accuracy, efficiency, and robustness of the proposed method.

math.NA

GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations

Large Language Model (LLM) agents increasingly serve as personal assistants and workplace collaborators, where their utility depends on memory systems that extract, retrieve, and apply information across long-running conversations. However, both existing memory systems and benchmarks are built around the dyadic, single-user setup, even though real deployments routinely span groups and channels with multiple users interacting with the agent and with each other. This mismatch leaves three properties of group memory unmeasured: (i) group dynamics that go beyond concatenated one-on-one chats, (ii) speaker-grounded belief tracking, where the per-user memory modeling is needed, and (iii) audience-adapted language, where Theory-of-Mind shifts produce role-specific vocabulary. We introduce GroupMemBench, a benchmark that exposes all three. A graph-grounded synthesis pipeline produces multi-party conversations with controllable reply structure and conditions each message on per-user personas and target audiences. An adversarial query pipeline then binds every question to a specific asker across six categories, spanning multi-hop reasoning, knowledge update, term ambiguity, user-implicit reasoning, temporal reasoning, and abstention, and iteratively searches challenging, realistic queries that reflect comprehensive memory capability. Benchmarking leading memory systems exposes a sharp collapse: the strongest one reaches only 46.0% average accuracy, with knowledge update at 27.1% and term ambiguity at 37.7%, while a simple BM25 baseline matches or exceeds most agent memory systems. This indicates current memory ingestion erases the structural and lexical features group memory depends on, leaving multi-user memory far from solved.

cs.CL

CT-Conditioned Diffusion Prior with Physics-Constrained Sampling for PET Super-Resolution

PET super-resolution is highly under-constrained because paired multi-resolution scans from the same subject are rarely available, and effective resolution is determined by scanner-specific physics (e.g., PSF, detector geometry, and acquisition settings). This limits supervised end-to-end training and makes purely image-domain generative restoration prone to hallucinated structures when anatomical and physical constraints are weak. We formulate PET super-resolution as posterior inference under heterogeneous system configurations and propose a CT-conditioned diffusion framework with physics-constrained sampling. During training, a conditional diffusion prior is learned from high-quality PET/CT pairs using cross-attention for anatomical guidance, without requiring paired LR--HR PET data. During inference, measurement consistency is enforced through a scanner-aware forward model with explicit PSF effects and gradient-based data-consistency refinement. Under both standard and OOD settings, the proposed method consistently improves experimental metrics and lesion-level clinical relevance indicators over strong baselines, while reducing hallucination artifacts and improving structural fidelity.

cs.CV

Physics-Informed Cross-Learning for Seismic Acoustic Impedance Inversion and Wavelet Extraction

Seismic acoustic impedance inversion is one of the most challenging tasks in geophysical exploration. Many studies have proposed the use of deep learning for processing; however, most of them are limited by factors such as seismic wavelets and low-frequency initial models. Furthermore, self-supervised frameworks constructed entirely using deep learning models struggle to form direct and effective physical constraints to unlabeled outputs during the multi-model concatenation, which leads to instability in inversion. In this work, we introduced innovations in both the deep learning framework and training strategy. First, we designed a deep learning framework to perform acoustic impedance inversion and seismic wavelet extraction simultaneously. Building on this foundation, considering the scarcity of well data, we proposed a physics-informed cross-learning strategy to impose effective constraints on the framework. We conducted comparative experiments and ablation experiments on both synthetic datasets and field datasets. The results demonstrate that the proposed method achieves a significant improvement compared with semi-supervised learning methods and can extract seismic wavelets with relatively high accuracy. Finally, to ensure the reproducibility of this work, we have made the code open-source.

physics.geo-ph

Encoder-Inverter Framework for Seismic Acoustic Impedance Inversion

Seismic acoustic impedance inversion is a challenging problem in geophysical exploration, primarily due to the scarcity of well-logging data and the inherent nonlinearity of the task. Most existing inversion methods, including semi-supervised learning approaches, still face limitations in accuracy and robustness. In this work, we propose a novel Encoder-Inverter framework that maps continuous seismic traces into high-dimensional linear features, thereby transforming the inversion task into a linear extrapolation or interpolation problem to enhance stability and performance. To achieve this, we introduce two auxiliary models to assist in encoder training and adopt a heterogeneous model structure to prevent shortcut learning, enabling the extraction of more generalizable and effective linear features. We evaluate the proposed method on widely used benchmark datasets, and experimental results demonstrate that our approach achieves superior inversion accuracy and robustness compared to previous methods. To promote reproducibility, we will also open-source the data and code.

physics.geo-ph

Towards Reliable Pediatric Brain Tumor Segmentation: Task-Specific nnU-Net Enhancements

Accurate segmentation of pediatric brain tumors in multi-parametric magnetic resonance imaging (mpMRI) is critical for diagnosis, treatment planning, and monitoring, yet faces unique challenges due to limited data, high anatomical variability, and heterogeneous imaging across institutions. In this work, we present an advanced nnU-Net framework tailored for BraTS 2025 Task-6 (PED), the largest public dataset of pre-treatment pediatric high-grade gliomas. Our contributions include: (1) a widened residual encoder with squeeze-and-excitation (SE) attention; (2) 3D depthwise separable convolutions; (3) a specificity-driven regularization term; and (4) small-scale Gaussian weight initialization. We further refine predictions with two postprocessing steps. Our models achieved first place on the Task-6 validation leaderboard, attaining lesion-wise Dice scores of 0.759 (CC), 0.967 (ED), 0.826 (ET), 0.910 (NET), 0.928 (TC) and 0.928 (WT).

eess.IV

Frequency-Aware Ensemble Learning for BraTS 2025 Pediatric Brain Tumor Segmentation

Pediatric brain tumor segmentation presents unique challenges due to the rarity and heterogeneity of these malignancies, yet remains critical for clinical diagnosis and treatment planning. We propose an ensemble approach integrating nnU-Net, Swin UNETR, and HFF-Net for the BraTS-PED 2025 challenge. Our method incorporates three key extensions: adjustable initialization scales for optimal nnU-Net complexity control, transfer learning from BraTS 2021 pre-trained models to enhance Swin UNETR's generalization on pediatric dataset, and frequency domain decomposition for HFF-Net to separate low-frequency tissue contours from high-frequency texture details. Our final ensemble framework combines nnU-Net ($γ=0.7$), fine-tuned Swin UNETR, and HFF-Net, achieving Dice scores of 62.7% (CC), 83.2% (ED), 72.9% (ET), 85.7% (NET), 91.8% (TC), and 92.6% (WT) on the unseen test dataset, respectively. Our proposed method achieves first place (rank 1st) in the BraTS 2025 Pediatric Brain Tumor Segmentation Challenge.

eess.IV

A Survey of Datasets for Information Diffusion Tasks

Information diffusion across various new media platforms gradually influences perceptions, decisions, and social behaviors of individual users. In communication studies, the famous Five W's of Communication model (5W Model) has displayed the process of information diffusion clearly. At present, although plenty of studies and corresponding datasets about information diffusion have emerged, a systematic categorization of tasks and an integration of datasets are still lacking. To address this gap, we survey a systematic taxonomy of information diffusion tasks and datasets based on the "5W Model" framework. We first categorize the information diffusion tasks into ten subtasks with definitions and datasets analysis, from three main tasks of information diffusion prediction, social bot detection, and misinformation detection. We also collect the publicly available dataset repository of information diffusion tasks with the available links and compare them based on six attributes affiliated to users and content: user information, social network, bot label, propagation content, propagation network, and veracity label. In addition, we discuss the limitations and future directions of current datasets and research topics to advance the future development of information diffusion. The dataset repository can be accessed at our website https://github.com/fuxiaG/Information-Diffusion-Datasets.

cs.SI

MCDAN: a Multi-scale Context-enhanced Dynamic Attention Network for Diffusion Prediction

Information diffusion prediction aims at predicting the target users in the information diffusion path on social networks. Prior works mainly focus on the observed structure or sequence of cascades, trying to predict to whom this cascade will be infected passively. In this study, we argue that user intent understanding is also a key part of information diffusion prediction. We thereby propose a novel Multi-scale Context-enhanced Dynamic Attention Network (MCDAN) to predict which user will most likely join the observed current cascades. Specifically, to consider the global interactive relationship among users, we take full advantage of user friendships and global cascading relationships, which are extracted from the social network and historical cascades, respectively. To refine the model's ability to understand the user's preference for the current cascade, we propose a multi-scale sequential hypergraph attention module to capture the dynamic preference of users at different time scales. Moreover, we design a contextual attention enhancement module to strengthen the interaction of user representations within the current cascade. Finally, to engage the user's own susceptibility, we construct a susceptibility label for each user based on user susceptibility analysis and use the rank of this label for auxiliary prediction. We conduct experiments over four widely used datasets and show that MCDAN significantly overperforms the state-of-the-art models. The average improvements are up to 10.61% in terms of Hits@100 and 9.71% in terms of MAP@100, respectively.

cs.SI