SearcharxivSearch

arXiv subjects

Ao Xu

Publications and source records attributed to Ao Xu.

At least 19 recordsLinked to original sources

ALKEMIE Agent: an autonomous platform for computational materials design

Despite the powerful multi-scale modeling methods and high-throughput infrastructures established in the materials community, real material computation workflows remain fragmented and heavily manual, requiring researchers to constantly bridge software tools, data analysis, and intermediate decisions. This growing gap between methodological capability and practical execution highlights the need for a new kind of autonomous computational framework, one that can coordinate tools, knowledge, and workflows in a more unified and adaptive way. Here, we introduce ALKEMIE Agent, an agentic platform in which retrieval-augmented generation, a materials-computation knowledge base, registered skills, database-supported provenance, AI-assisted structure modeling, bounded task execution, tool-calling iteration, and error-diagnostic assistance are integrated within a traceable control loop. The capabilities of ALKEMIE Agent are demonstrated through applications including materials recommendation, structure modeling, phonon calculations, machine-learned interatomic potential training, LAMMPS simulations, Ab Initio Monte Carlo (AIMC) sampling, and active-learning-based materials screening. Finally, we outline the future directions and challenges for the development of agentic platforms for computational materials design.

cond-mat.mtrl-sci

A Doeblin-Anchored Contrastive Chart for Learning Markov Transition Kernels

Learning a Markov transition model is not merely conditional density estimation: the learned object must be a valid transition kernel before it is iterated in downstream dynamics. This paper introduces a Doeblin-anchored contrastive chart, a statistical-to-dynamical coordinate framework for learning transition kernels from contrastive objectives. Given a restart law and an anchor strength, the chart mixes the target transition with the restart law. The resulting anchored kernel is simultaneously a Doeblin-minorized Markov kernel, the positive conditional law in a binary contrastive experiment, and an explicitly invertible coordinate for the original transition law. We prove that the anchored contrastive risk identifies the anchored transition density and calibrates excess risk to density error. Since inversion of a learned score may produce a signed or unnormalized object, we introduce a measurable Markovization operator that restores kernel validity while preserving integrated $L^1$ accuracy up to a constant factor. Oracle inequalities and H\"older--ReLU approximation bounds yield nonparametric rates for independent transition pairs. For stationary geometrically $\beta$-mixing trajectories, a conservative thinning-and-coupling extension yields the same reconstruction interface with an effective sample size. Occupancy-weighted perturbation bounds transfer one-step kernel error to finite-horizon marginal, path-law, and occupation-measure errors under explicit coverage.

cs.LG

ProRL: Effective Reinforcement Learning for Proactive Recommendation via Rectified Policy Gradient Estimation

Proactive Recommender Systems (PRSs) aim to guide user preference shift toward target items by generating paths of intermediate recommendations. Reinforcement learning (RL) provides a principled framework for optimizing such sequential decision tasks, as path rewards can naturally capture both short-term acceptance and long-term guidance effectiveness. However, naively applying policy gradients to PRS results in deficient gradient estimation. We identify two deficiencies: (1) path-level rewards decompose into step-level rewards with positive mean, creating a length-dependent bias that causes gradients to favor path extension over meaningful exploration; (2) weighting each step by the entire path-level reward ignores the decomposition structure, leading to high gradient variance. To rectify these two deficiencies, we propose an effective RL framework ProRL with two novel mechanisms for proactive recommendation. First, Stepwise Reward Centering subtracts expected rewards to neutralize length-dependent bias, ensuring that path extension yields zero expected gradient signal. Second, Position-Specific Advantage Estimation leverages the reward decomposition structure to compute step-dependent baselines, reducing gradient variance. Together, these mechanisms yield policy gradients that precisely target path quality. Our experiments on three real-world datasets demonstrate that ProRL significantly outperforms state-of-the-art PRSs. Our code is available at https://github.com/hongruhou89/ProRL.

cs.LG

Triggering of extreme events and coherent-structure modulation in wall-turbulence under cyclostationary forces

Atmospheric gusts expose wall-bounded turbulence to severe unsteady forcing, triggering complex non-equilibrium dynamics and extreme aerodynamic loads. In this study, direct numerical simulations are performed to investigate the spatiotemporal modulation of turbulent structures and the triggering mechanisms of near-wall extreme events under Gaussian-type transient forcing. The results reveal that high-amplitude gusts inject energy primarily into the streamwise velocity component, inducing a pronounced non-equilibrium phase lag during turbulent energy redistribution. This process produces hysteresis in wall friction and extends the relaxation time. Spectral and continuous wavelet analyses demonstrate that intense gust forcing suppresses high-frequency random fluctuations and reorganizes turbulent kinetic energy into low-frequency coherent structures. The characteristic frequency of these energetic structures locks onto the gust driving frequency, with a relative deviation of only $2.4\%$. Furthermore, the occurrence probability of extreme near-wall events, including extreme positive (EP) wall-shear-stress events and rare backflow (BF) events, increases by up to an order of magnitude under severe forcing. Using a two-step conditional averaging technique, we demonstrate that BF events are actively driven by intense, localized adverse pressure gradients and energetic ejections, which promote spanwise vortex roll-up in the buffer layer. By contrast, EP events are governed by energetic sweeps of high-speed fluid that compress intense spanwise vorticity into the immediate vicinity of the wall. These findings provide physical insights into non-equilibrium energy transfer and offer theoretical guidance for load alleviation and robust flow control of unmanned aerial vehicles operating in unsteady atmospheric environments.

physics.flu-dyn

Observing Joinings: A Distance-Array Characterization of Furstenberg Disjointness

Joinings are fundamental global objects in ergodic theory, yet in compact metric models one naturally observes only finite orbit-distance patterns. We bridge this gap by introducing multi-particle distance arrays, which sample finite orbit segments and record their joint metric evolution. In the anchored fixed-model setting, this framework yields a purely finite-observable characterization of Furstenberg disjointness: two systems are disjoint if and only if all their anchored multi-orbit distance-array projections are independent. The structural engine behind this criterion is a marked and colored version of the Gromov--Vershik reconstruction principle for exchangeable arrays; unanchored arrays reconstruct the intrinsic twin-free quotient, while anchors recover the actual joining in a fixed model. To quantify this independence, we introduce Wasserstein dependence coefficients, establishing an all-order zero criterion for disjointness, and show that weak neighborhoods of the product joining always admit finite distance-array certificates. Examples from compact rotations, Bernoulli and reversible Markov shifts, common factors, Kronecker factors, and weak mixing demonstrate the strict necessity of the multi-particle level and the broad scope of this approach.

math.DS

Distance-Matrix Wasserstein Statistics for Scalable Gromov--Wasserstein Learning

Gromov--Wasserstein (GW) distances compare graphs, shapes, and point clouds through internal distances, without requiring a common coordinate system. This invariance is powerful, but discrete GW is a nonconvex quadratic optimal transport problem and is difficult to estimate at scale. We propose \emph{Distance-Matrix Wasserstein} (DMW), a hierarchy of Wasserstein statistics comparing laws of random finite distance matrices. Rather than optimizing a global point-level alignment, DMW samples $n$ points from each space, records their pairwise distances, and transports the resulting matrix laws. We prove that DMW is a relaxation and lower bound of GW, and establish a reverse approximation inequality: the GW--DMW gap is controlled by the Wasserstein error of approximating each original measure with $n$ samples. Hence population DMW converges to GW as sampled subspaces become dense. We further give finite-sample bounds, including intrinsic-dimensional rates that depend on the data manifold rather than the ambient matrix dimension $\binom n2$. For scalable computation, we introduce sliced and multi-scale DMW; for $p=1$, the sliced multi-scale dissimilarity yields positive-definite exponential kernels. Experiments on synthetic metric spaces, scalability benchmarks, graph classification, and two-sample testing validate the theory and demonstrate an interpretable GW-style proxy for structural comparison.

cs.LG

Learning to traverse convective flows at moderate to high Rayleigh numbers

We study the navigation of a self-propelled inertial particle in two-dimensional Rayleigh-B\'enard convection at Prandtl number $Pr=0.71$ and cell aspect ratio $\Gamma=4$ for Rayleigh numbers $Ra$ ranging from $10^7$ to $10^{11}$. A reinforcement-learning (RL) controller selects the propulsive acceleration, subject to an upper bound $\mathcal{A}_{\max}$, to achieve a prescribed horizontal displacement. We find that the success rate increases abruptly with $\mathcal{A}_{\max}$ at moderate $Ra$, whereas at higher $Ra$ the transition becomes more gradual and shifts to larger $\mathcal{A}_{\max}$. Moreover, although the completion time increases with $Ra$, the propulsion energy required for successful traversal decreases. Proper orthogonal decomposition indicates that these performance differences are associated with reorganisation of the carrier flow. At moderate $Ra$, the dominant large-scale circulation partitions the domain through persistent transport barriers, requiring a finite thrust surplus to cross them; at higher $Ra$, energy is distributed across many modes, the barriers fragment and transient plume-assisted pathways emerge. Compared with a constant-heading baseline, the learned policy aligns with local currents and consumes significantly less energy. Lagrangian coherent structure analysis further suggests that the RL agent tends to cross repelling barriers and surf along attracting pathways. Finally, by mapping these behaviours onto the local Eulerian flow topology using Voronoi tessellation and the $Q$-criterion, we distil an interpretable, physics-based heuristic strategy that retains robust navigability. These results connect turbulent-flow organisation with autonomous navigation under bounded actuation.

physics.flu-dyn

Revisit eddy viscosity in pressure-driven wall turbulence at high Reynolds number

We investigate eddy-viscosity distributions in pressure-driven wall turbulence for three canonical configurations: plane closed-channel flow, open-channel flow with a free-slip surface, and pipe flow. Using direct numerical simulation (DNS) databases spanning friction Reynolds numbers $Re_{\tau}=$ 2000--12000, we infer the eddy viscosity from one-point statistics through the Boussinesq relation. The DNS-inferred eddy viscosity displays configuration-dependent behavior in the outer region, indicating that a single full-depth expression is not uniformly accurate for all three configurations. Building on the interpretation of eddy viscosity as the product of a velocity scale and a length scale, we extend the log-law scaling into the outer region. Specifically, we adopt a stress-based velocity scale and introduce an outer correction function to capture the remaining dependence on the outer coordinate. We then embed a compact parametric form of this correction into a Cess-type framework with van Driest near-wall damping, yielding a full-depth eddy-viscosity model. We assess the model using eddy-viscosity profiles, the log-law indicator function, and skin friction. The results show that the proposed model yields noticeable improvement for open-channel flow while remaining comparable to the classical Cess model for closed-channel flow and pipe flow. These findings underscore the role of outer boundary conditions in shaping the outer-region eddy viscosity and, consequently, mean-flow predictions.

physics.flu-dyn

Perturb-and-Restore: Simulation-driven Structural Augmentation Framework for Imbalance Chromosomal Anomaly Detection

Detecting structural chromosomal abnormalities is crucial for accurate diagnosis and management of genetic disorders. However, collecting sufficient structural abnormality data is extremely challenging and costly in clinical practice, and not all abnormal types can be readily collected. As a result, deep learning approaches face significant performance degradation due to the severe imbalance and scarcity of abnormal chromosome data. To address this challenge, we propose a Perturb-and-Restore (P&R), a simulation-driven structural augmentation framework that effectively alleviates data imbalance in chromosome anomaly detection. The P&R framework comprises two key components: (1) Structure Perturbation and Restoration Simulation, which generates synthetic abnormal chromosomes by perturbing chromosomal banding patterns of normal chromosomes followed by a restoration diffusion network that reconstructs continuous chromosome content and edges, thus eliminating reliance on rare abnormal samples; and (2) Energy-guided Adaptive Sampling, an energy score-based online selection strategy that dynamically prioritizes high-quality synthetic samples by referencing the energy distribution of real samples. To evaluate our method, we construct a comprehensive structural anomaly dataset consisting of over 260,000 chromosome images, including 4,242 abnormal samples spanning 24 categories. Experimental results demonstrate that the P&R framework achieves state-of-the-art (SOTA) performance, surpassing existing methods with an average improvement of 8.92% in sensitivity, 8.89% in precision, and 13.79% in F1-score across all categories.

cs.CV

Good Reasoning Makes Good Demonstrations: Implicit Reasoning Quality Supervision via In-Context Reinforcement Learning

Reinforcement Learning with Verifiable Rewards (RLVR) improves reasoning in large language models but treats all correct solutions equally, potentially reinforcing flawed traces that arrive at correct answers by chance. We observe that \emph{better reasoning makes better demonstrations}: high-quality solutions serve as more effective in-context examples than low-quality ones. We term this teaching ability \textbf{Demonstration Utility}, and show that the policy model's own in-context learning ability provides an efficient way to measure it, yielding a quality signal termed \textbf{Evidence Gain}. To leverage this signal during training, we introduce \textbf{In-Context RLVR}, which prepends demonstrations before each rollout. Theoretically, we prove that this simple input modification implicitly reweights rewards by a factor approximately proportional to Evidence Gain, assigning higher weights to high-quality traces without requiring costly computation. Experiments on mathematical reasoning benchmarks demonstrate consistent improvements in both accuracy and reasoning quality over standard RLVR baselines. Our codes and datasets are available at https://github.com/Mithas-114/IC-DAPO.

cs.LG

Buoyancy-induced velocity dip in turbulent open Poiseuille--Rayleigh--B\'enard convection

We investigate buoyancy-induced transitions in flow structure and the associated velocity dip in turbulent mixed convection. Numerical simulations are performed for an open Poiseuille--Rayleigh--B\'enard system with a heated no-slip lower wall and a cooled free-slip upper boundary over $10^5 \leq Ra \leq 10^8$, $90 \leq Re_b \leq 5700$, $Pr=0.71$, and $0.013 \leq Ri_b \leq 17$. The flow organisation is governed primarily by the bulk Richardson number $Ri_b$. As buoyancy increases, the flow changes from a shear-dominated state to streamwise-oriented large-scale rolls and then to fragmented rolls. Roll formation coincides with a reorganisation of the velocity statistics about the channel midplane and a displacement of the maximum mean streamwise velocity from the upper boundary into the channel interior, producing a velocity dip. The mean momentum balance shows that spatially uniform streamwise forcing imposes a linear total-stress profile. Wherever the Reynolds shear stress exceeds the local total stress, the viscous stress and mean velocity gradient must be negative. A triple decomposition attributes most of the Reynolds-stress excess in the roll states to slowly varying, roll-associated motions. The roll-associated stress and dip strength vary non-monotonically with $Ri_b$. Quadrant analysis shows that ejections and sweeps sustain the net roll-associated stress, whereas roll fragmentation strengthens the cancellation between positive and negative contributions and weakens the dip. The turbulent kinetic energy (TKE) budget shows a corresponding shift from near-wall shear production to bulk buoyancy production, with shear production becoming negative above the velocity maximum. Finally, a case-specific \textit{a posteriori} simplification of the core-region TKE budget yields an approximate velocity profile that captures the gradient reversal.

physics.flu-dyn

Huayu: Advanced Real-Time Precipitation Estimation from Geostationary Satellite

As climate change drives increased frequency and intensity of extreme precipitation and flooding worldwide, posing escalating threats to public safety and economic assets, accurate and real-time satellite-based precipitation estimation is essential for operational large-scale hydrometeorological analysis and disaster monitoring. NASA's Integrated Multi-satellitE Retrievals for GPM (IMERG Final Run) combines information from "all" satellite microwave observations with gauge correction and climatological adjustment to produce precipitation estimates at 0.1{\deg} spatial and 30-min temporal resolution. However, its latency of approximately 3.5 months restricts its utility for real-time applications, despite outperforming mainstream satellite precipitation datasets in representing rainfall patterns and variability. We present Huayu, a novel machine learning-based real-time satellite precipitation retrieval system that relies solely on infrared observations from the FengYun-4B geostationary satellite to provide a more accurate precipitation estimate at a finer spatiotemporal resolution (15 min, 0.05{\deg}) over a 120{\deg} by 120{\deg} domain. Performance evaluations demonstrate that Huayu achieves strong consistency with rain gauge observations, yielding a critical success index (CSI) of 0.693 - representing a 3.43% improvement over IMERG Final Run (CSI: 0.670). Experimental results confirm that infrared satellite observations can deliver more accurate precipitation estimates than conventional multi-source algorithms.

physics.ao-ph

MAFNet:Multi-frequency Adaptive Fusion Network for Real-time Stereo Matching

Existing stereo matching networks typically rely on either cost-volume construction based on 3D convolutions or deformation methods based on iterative optimization. The former incurs significant computational overhead during cost aggregation, whereas the latter often lacks the ability to model non-local contextual information. These methods exhibit poor compatibility on resource-constrained mobile devices, limiting their deployment in real-time applications. To address this, we propose a Multi-frequency Adaptive Fusion Network (MAFNet), which can produce high-quality disparity maps using only efficient 2D convolutions. Specifically, we design an adaptive frequency-domain filtering attention module that decomposes the full cost volume into high-frequency and low-frequency volumes, performing frequency-aware feature aggregation separately. Subsequently, we introduce a Linformer-based low-rank attention mechanism to adaptively fuse high- and low-frequency information, yielding more robust disparity estimation. Extensive experiments demonstrate that the proposed MAFNet significantly outperforms existing real-time methods on public datasets such as Scene Flow and KITTI 2015, showing a favorable balance between accuracy and real-time performance.

cs.CV

AI Urban Scientist: Multi-Agent Collaborative Automation for Urban Research

Urban research aims to understand how cities operate and evolve as complex adaptive systems. With the rapid growth of urban data and analytical methodologies, the central challenge of the field has shifted from data availability to the integration of heterogeneous data into coherent, verifiable urban knowledge through multidisciplinary approaches. Recent advances in AI, particularly the emergence of large language models (LLMs), have enabled the development of AI scientists capable of autonomous reasoning, hypothesis generation, and data-driven experimentation, demonstrating substantial potential for autonomous urban research. However, most general-purpose AI systems remain misaligned with the domain-specific knowledge, methodological conventions, and inferential standards required in urban studies. Here, we introduce the AI Urban Scientist, a knowledge-driven multi-agent framework designed to support autonomous urban research. Grounded in hypotheses, peer-review feedback, datasets, and research methodologies distilled from large-scale prior studies, the system constructs structured domain knowledge that guides LLM-based agents to automatically generate hypotheses, identify and integrate multi-source urban datasets, conduct empirical analyses and simulations, and iteratively refine analytical methods. Through this process, the framework synthesizes new insights in urban science and accelerates the urban research lifecycle.

cs.CY

Harli: SLO-Aware Co-location of LLM Inference and PEFT-based Finetuning on Model-as-a-Service Platforms

Large language models (LLMs) are increasingly deployed under the Model-as-a-Service (MaaS) paradigm. To meet stringent quality-of-service (QoS) requirements, existing LLM serving systems disaggregate the prefill and decode phases of inference. However, decode instances often experience low GPU utilization due to their memory-bound nature and insufficient batching in dynamic workloads, leaving compute resources underutilized. We introduce Harli, a serving system that improves GPU utilization by co-locating parameter-efficient finetuning (PEFT) tasks with LLM decode instances. PEFT tasks are compute-bound and memory-efficient, making them ideal candidates for safe co-location. Specifically, Harli addresses key challenges--limited memory and unpredictable interference--using three components: a unified memory allocator for runtime memory reuse, a two-stage latency predictor for decode latency modeling, and a QoS-guaranteed throughput-maximizing scheduler for throughput maximization. Experimental results show that Harli improves the finetune throughput by 46.2% on average (up to 92.0%) over state-of-the-art serving systems, while maintaining strict QoS guarantees for inference decode.

cs.DC

Super-resolution reconstruction of turbulent flows from a single Lagrangian trajectory

We studied the reconstruction of turbulent flow fields from trajectory data recorded by actively migrating Lagrangian agents. We propose a deep-learning model, track-to-flow (T2F), which employs a vision transformer as the encoder to capture the spatiotemporal features of a single agent trajectory, and a convolutional neural network as the decoder to reconstruct the flow field. To enhance the physical consistency of the T2F model, we further incorporate a physics-informed loss function inspired by the framework of physics-informed neural network (PINN), yielding a variant model referred to as T2F+PINN. We first evaluate both models in a laminar cylinder wake flow at a Reynolds number of $Re = 800$ as a proof of concept. The results show that the T2F model achieves velocity reconstruction accuracy comparable to that of existing flow reconstruction methods, while the T2F+PINN model reduces the normalised error in vorticity reconstruction relative to the T2F model. We then apply the models in a turbulent Rayleigh-B\'enard convection at a Rayleigh number of $Ra = 10^8$ and a Prandtl number of $Pr = 0.71$. The results show that the T2F model accurately reconstructs both the velocity and temperature fields, whereas the T2F+PINN model further improves the reconstruction accuracy of gradient-related physical quantities, such as temperature gradients, vorticity and the Q value, with a maximum improvement of approximately 60 % compared to the T2F model. Overall, the T2F model is better suited for reconstructing primitive flow variables, while the T2F+PINN model provides advantages in reconstructing gradient-related quantities. Our models open a promising avenue for accurate flow reconstruction from a single Lagrangian trajectory.

physics.flu-dyn

Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution

Embodied AI systems operate in dynamic environments, requiring seamless integration of perception and generation modules to process high-frequency input and output demands. Traditional sequential computation patterns, while effective in ensuring accuracy, face significant limitations in achieving the necessary "thinking" frequency for real-world applications. In this work, we present Auras, an algorithm-system co-designed inference framework to optimize the inference frequency of embodied AI agents. Auras disaggregates the perception and generation and provides controlled pipeline parallelism for them to achieve high and stable throughput. Faced with the data staleness problem that appears when the parallelism is increased, Auras establishes a public context for perception and generation to share, thereby promising the accuracy of embodied agents. Experimental results show that Auras improves throughput by 2.54x on average while achieving 102.7% of the original accuracy, demonstrating its efficacy in overcoming the constraints of sequential computation and providing high throughput.

cs.AI

Interpolation-supplemented lattice Boltzmann simulation of thermal convection on non-uniform meshes

We present a systematic evaluation of an interpolation-supplemented lattice Boltzmann method (ISLBM) for simulating buoyancy-driven thermal convection on non-uniform meshes. The ISLBM extends the standard lattice Boltzmann framework by incorporating quadratic interpolation during the streaming step, enabling flexible mesh refinement near solid boundaries while maintaining algorithmic simplicity and parallel scalability. The method is implemented for a two-dimensional side-heated cavity at high Rayleigh numbers $10^6\leq Ra \leq 10^8$, and for a three-dimensional side-heated cavity at $10^5\leq Ra \leq 10^7$, with the Prandtl number fixed at $Pr=0.71$. Benchmark results show that the ISLBM accurately captures thermal and velocity boundary layers, yielding Nusselt and Reynolds numbers in close agreement with high-fidelity reference data. Grid-convergence studies demonstrate nearly third-order accuracy for global quantities and about second-order for local fields. We further assess the computational performance of the in-house LBM solver against two open-source solvers: Nek5000 based on the spectral element method, and OpenFOAM based on the finite volume method. Performance metrics, including million lattice updates per second (MLUPS) and wall-clock time per dimensionless time unit (WCTpDT), indicate that the ISLBM offers one to three orders of magnitude higher efficiency in large-scale simulations. On GPU architectures, the ISLBM retains high computational performance: throughput on non-uniform meshes reaches 60-70% of that on uniform meshes in terms of MLUPS, while the cost in WCTpDT is about three times higher. These results highlight the potential of interpolation-based LBM approaches for high-fidelity simulations of thermal convection on non-uniform meshes, providing a robust foundation for future extensions to turbulent flows.

physics.flu-dyn