SearcharxivSearch

arXiv subjects

Xiaolei Yang

Publications and source records attributed to Xiaolei Yang.

At least 19 recordsLinked to original sources

DianShi-RxnDB: A Large-Scale, Fine-Grained Organic Reaction Data Platform Built via a Fully Automated Pipeline for Researchers and AI Agents

High-quality structured organic reaction data are essential for developing artificial intelligence for chemistry (AI4Chem), yet much of this knowledge remains dispersed across patent text, images, and reaction schemes. We present DianShi-RxnDB, a large-scale, fine-grained organic reaction data platform built via a fully automated extraction and normalization pipeline integrating patent text, images, and reaction schemes. Its corpus covers organic synthesis patents from the USPTO and EPO published between 1976 and 2025, yielding approximately 24 million reaction instances, of which approximately 14.8 million (61.7%) pass automated qualification checks. Each instance represents a specific single-step experiment recording participants, roles, quantities, temperatures, reaction times, yields, experimental procedures, and provenance links to source patents. In a manual evaluation of 1,300 sampled qualified instances, the micro-averaged field-level accuracy was 92.95%. A matched comparison with Pistachio further indicated advantages in deduplicated record counts, representation granularity, and field-level exact agreement. The platform provides a Web research workbench for searching, filtering, comparing, and source-verifying records, and a Model Context Protocol (MCP) service offering AI agents composable structured retrieval tools. DianShi-RxnDB is available at https://dianshi.opendatalab.org.cn/ .

cs.CL

AgentPanel: Toward a New Paradigm for Human--AI Collaboration in Exploring Scientific Questions

Identifying promising scientific ideas remains an important challenge in research practice. Researchers commonly rely on small-group discussions or one-to-one interactions with a single large language model, yet these approaches often expose them to only a limited range of perspectives and directions. We present AgentPanel, a multi-agent forum for human--AI collaboration in scientific exploration. Heterogeneous agents asynchronously discuss scientific questions in a forum-style environment, while researchers can submit questions, browse and organize candidate ideas, engage agents in follow-up interactions, and optionally generate post-hoc summary reports. We evaluate AgentPanel in terms of idea quality, exploration breadth, interaction effectiveness, candidate-selection efficiency, and practical utility. Offline experiments show that AgentPanel outperforms a centralized multi-agent debate baseline. A human study with 20 participants further shows that users value AgentPanel for perspective diversity and exploration support. In experience-based comparisons with commonly used LLM tools, 65\% of participants favored AgentPanel for both breadth of research directions and overall suitability for early-stage exploration. The platform is publicly available at https://agentpanel.cc/.

cs.AI

Series decomposition of a class of special integrals

In this paper, we propose a new method for calculating integrals for a special class of integrands. As an application, we show how this method can be used to derive optimal pointwise temporal estimates for a class of nonlocal evolution equations. Compared with other methods, our approach can obtain both upper bound and lower bound simultaneously.

math.AP

Law in Silico: Simulating Legal Society with LLM-Based Agents

Since real-world legal experiments are often costly or infeasible, simulating legal societies with Artificial Intelligence (AI) systems provides an effective alternative for verifying and developing legal theory, as well as supporting legal administration. Large Language Models (LLMs), with their world knowledge and role-playing capabilities, are strong candidates to serve as the foundation for legal society simulation. However, the application of LLMs to simulate legal systems remains underexplored. In this work, we introduce Law in Silico, an LLM-based agent framework for simulating legal scenarios with individual decision-making and institutional mechanisms of legislation, adjudication, and enforcement. Our experiments, which compare simulated crime rates with real-world data, demonstrate that LLM-based agents can largely reproduce macro-level crime trends and provide insights that align with real-world observations. At the same time, micro-level simulations reveal that a well-functioning, transparent, and adaptive legal system offers better protection of the rights of vulnerable individuals.

cs.AI

Data-driven modeling of wind farm wake flow based on multi-scale feature recognition

Accurate, efficient prediction of wind flow with wake effects is crucial for wind-farm layout and power forecasting. Existing approaches-physical measurements, numerical simulations, physics-based models, and data-driven models-face trade-offs: the first two are time- and resource-intensive; physics-based models can lack accuracy due to limited physics; data-driven methods leverage abundant, high-quality data and are increasingly popular. We propose a rapid, data-driven wake-flow model inspired by video-frame interpolation and the principle of similarity. Field data are transformed into images; multi-scale feature recognition then identifies, matches, and interpolates wake structures using Scale-Invariant Feature Transform (SIFT) and Dynamic Time Warping (DTW) to generate intermediate flow fields. Six representative mini wind-farm cases validate the approach, spanning variations in turbine spacing, turbine size, combined spacing-size variations, different turbine counts, and wind-direction misalignment. Across cases, the method achieves a mean absolute percentage error (MAPE) of 0.68-2.28%. Because it flexibly computes both 2D and 3D wake fields, the method offers substantial computational-efficiency gains over large-eddy simulation (LES) and Meteodyn WT when 2D accuracy suffices for industrial needs. Accordingly, it provides a practical alternative to measurements, high-fidelity simulations, and simplified physics-based models, enabling efficient expansion of wake-flow databases for wind-farm design and power prediction while balancing speed and accuracy.

physics.flu-dyn

Timeliness-Aware Joint Source and Channel Coding for Adaptive Image Transmission

Accurate and timely image transmission is critical for emerging time-sensitive applications such as remote sensing in satellite-assisted Internet of Things. However, the bandwidth limitation poses a significant challenge in existing wireless systems, making it difficult to fulfill the requirements of both high-fidelity and low-latency image transmission. Semantic communication is expected to break through the performance bottleneck by focusing on the transmission of goal-oriented semantic information rather than raw data. In this paper, we employ a new timeliness metric named the value of information (VoI) and propose an adaptive joint source and channel coding (JSCC) method for image transmission that simultaneously considers both reconstruction quality and timeliness. Specifically, we first design a JSCC framework for image transmission with adaptive code length. Next, we formulate a VoI maximization problem by optimizing the transmission code length of the adaptive JSCC under the reconstruction quality constraint. Then, a deep reinforcement learning-based algorithm is proposed to solve the optimization problem efficiently. Experimental results show that the proposed method significantly outperforms baseline schemes in terms of reconstruction quality and timeliness, particularly in low signal-to-noise ratio conditions, offering a promising solution for efficient and robust image transmission in time-sensitive wireless networks.

eess.SP

Lagrangian-Eulerian learning of flow field and trajectories with TrajectoryFlowNet

Predicting particle transport in complex flows is traditionally achieved by solving the Navier-Stokes equations. While various numerical and experimental methods exist, they typically require deep physical insights and incur high computational costs. Machine learning offers an alternative by learning predictive patterns directly from data, avoiding explicit physical modeling. However, purely data-driven approaches often lack interpretability, physical consistency, and generalizability in sparse data regimes. To this end, we propose TrajectoryFlowNet, a Lagrangian-Eulerian physics-informed neural network architecture, for fluid flow velocimetry and imaging via learning to predict spatiotemporal flow fields and long-range particle trajectories. The salient features of our model include its ability to handle complex flow patterns with irregular boundaries, predict the full-field flows, image the long-range flow trajectory of any arbitrary particle, and ensure physical consistency in predictions based only on very scarce measurement of flow trajectories. We validate TrajectoryFlowNet via both numerical examples (e.g., lid-driven cavity flow and complex cylinder flow) and experimental test cases (e.g., aortic and ventricle blood flows) across diverse flow scenarios. The results demonstrate our model's effectiveness in capturing intricate particle-laden flow dynamics, enabling long-range tracking of particles and accurate construction of flow fields in real-world applications.

physics.flu-dyn

A features-embedded-learning immersed boundary model for large-eddy simulation of turbulent flows with complex boundaries

The hybrid wall-modeled large-eddy simulation (WMLES) and immersed boundary (IB) method offers significant flexibility for simulating high Reynolds number flows involving complex boundaries. However, the approximate boundary conditions (e.g., the wall shear stress boundary condition) developed for body-fitted grids in the literature are not directly applicable to IB methods. In this work, we propose a features-embedded-learning-IB (FEL-IB) wall model to approximate IB boundaries in the hybrid WMLES-IB method, in which the velocity at the IB node (a grid node located in the fluid that has at least one neighbor in the solid) is reconstructed using the power law of the wall, and the momentum flux at the interface between the IB nodes and the fluid nodes is approximated using a neural network model. The neural network model for momentum flux is first pretrained using high-fidelity simulation data of the flow over periodic hills and the logarithmic law, and then learned in the WMLES environment using the ensemble Kalman method. The proposed model is evaluated using two challenging cases: the flow over a body of revolution and the DARPA Suboff submarine model. For the first case, good agreement with the reference data is obtained for the vertical profiles of the streamwise velocity. For the flow over the DARPA Suboff submarine model, WMLES cases with two Reynolds numbers and two grid resolutions are carried out. Overall good a posteriori performance is observed for predicting the mean and root mean square of velocity profiles at various streamwise locations, as well as the skin-friction and pressure coefficients.

physics.flu-dyn

Self-consistent model for active control of wind turbine wakes

Active wake control (AWC) has emerged as a promising strategy for enhancing wind turbine wake recovery, but accurately modelling its underlying fluid mechanisms remains challenging. This study presents a computationally efficient wake model that provides end-to-end prediction capability from rotor actuation to wake recovery enhancement by capturing the coupled dynamics of wake meandering and meanflow modification, requiring only two inputs: a reference wake without control and a user-defined AWC strategy. The model combines physics-based resolvent modelling for large-scale coherent structures and an eddy viscosity modelling for small-scale turbulence. A Reynolds stress model is introduced to account for the influence of both coherent and incoherent wake fluctuations, so that the time-averaged wake recovery enhanced by the AWC can be quantitatively predicted. Validation against large-eddy simulations (LES) across various AWC approaches and actuating frequencies demonstrates the model's predictive capability, accurately capturing AWC-specific and frequency-dependent mean wake recovery with less than 8% error from LES while reducing computational time from thousands of CPU hours to minutes. The efficiency and accuracy of the model makes it a promising tool for practical AWC design and optimization of large-scale wind farms.

physics.flu-dyn

A Survey of mmWave Backscatter: Applications, Platforms, and Technologies

As a key enabling technology of the Internet of Things (IoT) and 5G communication networks, millimeter wave (mmWave) backscatter has undergone noteworthy advancements and brought significant improvement to prevailing sensing and communication systems. Past few years have witnessed growing efforts in innovating mmWave backscatter transmitters (e.g., tags and metasurfaces) and the corresponding techniques, which provide efficient information embedding and fine-grained signal manipulation for mmWave backscatter technologies. These efforts have greatly enabled a variety of appealing applications, such as long-range localization, roadside-to-vehicle communication, coverage optimization and large-scale identification. In this paper, we carry out a comprehensive survey to systematically summarize the works related to the topic of mmWave backscatter. Firstly, we introduce the scope of this survey and provide a taxonomy to distinguish two categories of mmWave backscatter research based on the operating principle of the backscatter transmitter: modulation-based and relay-based. Furthermore, existing works in each category are grouped and introduced in detail, with their common applications, platforms and technologies, respectively. Finally, we elaborate on potential directions and discuss related surveys in this area.

cs.ET

Agents on the Bench: Large Language Model Based Multi Agent Framework for Trustworthy Digital Justice

The justice system has increasingly employed AI techniques to enhance efficiency, yet limitations remain in improving the quality of decision-making, particularly regarding transparency and explainability needed to uphold public trust in legal AI. To address these challenges, we propose a large language model based multi-agent framework named AgentsBench, which aims to simultaneously improve both efficiency and quality in judicial decision-making. Our approach leverages multiple LLM-driven agents that simulate the collaborative deliberation and decision making process of a judicial bench. We conducted experiments on legal judgment prediction task, and the results show that our framework outperforms existing LLM based methods in terms of performance and decision quality. By incorporating these elements, our framework reflects real-world judicial processes more closely, enhancing accuracy, fairness, and society consideration. AgentsBench provides a more nuanced and realistic methods of trustworthy AI decision-making, with strong potential for application across various case types and legal scenarios.

cs.AI

Reinforcement learning-enhanced genetic algorithm for wind farm layout optimization

A reinforcement learning-enhanced genetic algorithm (RLGA) is proposed for wind farm layout optimization (WFLO) problems. While genetic algorithms (GAs) are among the most effective and accessible methods for WFLO, their performance and convergence are highly sensitive to parameter selections. To address the issue, reinforcement learning (RL) is introduced to dynamically select optimal parameters throughout the GA process. To illustrate the accuracy and efficiency of the proposed RLGA, we evaluate the WFLO problem for four layouts (aligned, staggered, sunflower, and unstructured) under unidirectional uniform wind, comparing the results with those from the GA. RLGA achieves similar results to GA for aligned and staggered layouts and outperforms GA for sunflower and unstructured layouts, demonstrating its efficiency. The sunflower and unstructured layouts' complexity highlights RLGA's robustness and efficiency in tackling complex problems. To further validate its capabilities, we investigate larger wind farms with varying turbine placements ($Δx = Δy = 5D$ and 2$D$, where $D$ is the wind turbine diameter) under three wind conditions: unidirectional, omnidirectional, and non-uniform, presenting greater challenges. The proposed RLGA is about three times more efficient than GA, especially for complex problems. This improvement stems from RL's ability to adjust parameters, avoiding local optima and accelerating convergence.

cs.NE

A wall model for separated flows: embedded learning to improve a posteriori performance

The development of a wall model using machine learning methods for the large-eddy simulation (LES) of separated flows is still an unsolved problem. Our approach is to leverage the significance of separated flow data, for which existing theories are not applicable, and the existing knowledge of wall-bounded flows (such as the law of the wall) along with embedded learning to address this issue. The proposed so-called features-embedded-learning (FEL) wall model comprises two submodels: one for predicting the wall shear stress and another for calculating the eddy viscosity at the first off-wall grid nodes. We train the former using the wall-resolved LES data of the periodic hill flow and the law of the wall. For the latter, we propose a modified mixing length model, with the model coefficient trained using the ensemble Kalman method. The proposed FEL model is assessed using the separated flows with different flow configurations, grid resolutions, and Reynolds numbers. Overall good a posteriori performance is observed for predicting the statistics of the recirculation bubble, wall stresses, and turbulence characteristics. The statistics of the modelled subgrid-scale (SGS) stresses at the first off-wall grids are compared with those calculated using the wall-resolved LES data. The comparison shows that the amplitude and distribution of the SGS stresses obtained using the proposed model agree better with the reference data when compared with the conventional wall model.

physics.flu-dyn

Joint SIM Configuration and Power Allocation for Stacked Intelligent Metasurface-assisted MU-MISO Systems with TD3

The stacked intelligent metasurface (SIM) emerges as an innovative technology with the ability to directly manipulate electromagnetic (EM) wave signals, drawing parallels to the operational principles of artificial neural networks (ANN). Leveraging its structure for direct EM signal processing alongside its low-power consumption, SIM holds promise for enhancing system performance within wireless communication systems. In this paper, we focus on SIM-assisted multi-user multi-input and single-output (MU-MISO) system downlink scenarios in the transmitter. We proposed a joint optimization method for SIM phase shift configuration and antenna power allocation based on the twin delayed deep deterministic policy gradient (TD3) algorithm to efficiently improve the sum rate. The results show that the proposed algorithm outperforms both deep deterministic policy gradient (DDPG) and alternating optimization (AO) algorithms. Furthermore, increasing the number of meta-atoms per layer of the SIM is always beneficial. However, continuously increasing the number of layers of SIM does not lead to sustained performance improvement.

eess.SP

Time integration schemes based on neural networks for solving partial differential equations on coarse grids

The accuracy of solving partial differential equations (PDEs) on coarse grids is greatly affected by the choice of discretization schemes. In this work, we propose to learn time integration schemes based on neural networks which satisfy three distinct sets of mathematical constraints, i.e., unconstrained, semi-constrained with the root condition, and fully-constrained with both root and consistency conditions. We focus on the learning of 3-step linear multistep methods, which we subsequently applied to solve three model PDEs, i.e., the one-dimensional heat equation, the one-dimensional wave equation, and the one-dimensional Burgers' equation. The results show that the prediction error of the learned fully-constrained scheme is close to that of the Runge-Kutta method and Adams-Bashforth method. Compared to the traditional methods, the learned unconstrained and semi-constrained schemes significantly reduce the prediction error on coarse grids. On a grid that is 4 times coarser than the reference grid, the mean square error shows a reduction of up to an order of magnitude for some of the heat equation cases, and a substantial improvement in phase prediction for the wave equation. On a 32 times coarser grid, the mean square error for the Burgers' equation can be reduced by up to 35% to 40%.

math.NA

Validation of the dynamic wake meandering model against large eddy simulation for horizontal and vertical steering of wind turbine wakes

This work focuses on the validation of the dynamic wake meandering (DWM) model against large eddy simulation (LES). The wake deficit, mean deflection, and meandering under different wind turbine misalignment angles in yaw and tilt, for the IEA 15MW wind turbine, for two turbulent inflows with different shear and turbulence intensities are compared. Simulation results indicate that the DWM model as implemented in FAST.Farm shows very good agreement with the LES (VFS-Wind) data when predicting the time-averaged horizontal and vertical wake, especially at x > 6D and for cases with positive tilt angles (> 6deg). The wake dynamics captured by the DWM model include the large-eddy-induced wake meandering at low Strouhal number (St < 0.1). Additionally, the wake oscillation induced by the shear layer at St approx. 0.27 is captured only by LES. The mean and standard deviation of the wake deflection, as computed by the DWM, are sensitive to the size of the polar grid used to calculate the spatial-averaged velocity with which the wake planes meander. The power output of a turbine in the wake of a wind turbine in free-wind deflected by a yaw angle gamma = 30deg is almost doubled compared to the fully-waked condition.

physics.flu-dyn

Legal Syllogism Prompting: Teaching Large Language Models for Legal Judgment Prediction

Legal syllogism is a form of deductive reasoning commonly used by legal professionals to analyze cases. In this paper, we propose legal syllogism prompting (LoT), a simple prompting method to teach large language models (LLMs) for legal judgment prediction. LoT teaches only that in the legal syllogism the major premise is law, the minor premise is the fact, and the conclusion is judgment. Then the models can produce a syllogism reasoning of the case and give the judgment without any learning, fine-tuning, or examples. On CAIL2018, a Chinese criminal case dataset, we performed zero-shot judgment prediction experiments with GPT-3 models. Our results show that LLMs with LoT achieve better performance than the baseline and chain of thought prompting, the state-of-art prompting method on diverse reasoning tasks. LoT enables the model to concentrate on the key information relevant to the judgment and to correctly understand the legal meaning of acts, as compared to other methods. Our method enables LLMs to predict judgment along with law articles and justification, which significantly enhances the explainability of models.

cs.CL

On the morphodynamics of a wide class of large-scale meandering rivers: Insights gained by coupling LES with sediment-dynamics

In meandering rivers, interactions between flow, sediment transport, and bed topography affect diverse processes, including bedform development and channel migration. Predicting how these interactions affect the spatial patterns and magnitudes of bed deformation in meandering rivers is essential for various river engineering and geoscience problems. Computational fluid dynamics simulations can predict river morphodynamics at fine temporal and spatial scales but have traditionally been challenged by the large scale of natural rivers. We conducted coupled large-eddy simulation (LES) and bed morphodynamics simulations to create a unique database of hydro-morphodynamic datasets for 42 meandering rivers with a variety of planform shapes and large-scale geometrical features that mimic natural meanders. For each simulated river, the database includes (i) bed morphology, (ii) three-dimensional mean velocity field, and (iii) bed shear stress distribution under bankfull flow conditions. The calculated morphodynamics results at dynamic equilibrium revealed the formation of scour and deposition patterns near the outer and inner banks, respectively, while the location of point bars and scour regions around the apexes of the meander bends is found to vary as a function of the radius of curvature of the bends to the width ratio. A new mechanism is proposed that explains this seemingly paradoxical finding. The high-fidelity simulation results generated in this work provide researchers and scientists with a rich numerical database for morphodynamics and bed shear stress distributions in large-scale meandering rivers to enable systematic investigation of the underlying phenomena and support a range of river engineering applications.

physics.geo-ph