SearcharxivSearch

arXiv subjects

Hongxuan Zhang

Publications and source records attributed to Hongxuan Zhang.

15 recordsLinked to original sources

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents

Reinforcement learning with verifiable rewards (RLVR) offers a verifier-bounded performance ceiling for training multi-turn tool-use agents, yet its trajectory-level credit assignment conflates heterogeneous per-turn outcomes into a single reward signal. On-policy distillation provides dense per-token supervision but is either teacher-bounded or prone to gradient concentration collapse. We introduce $\textbf{CrEST}$, a hierarchical credit assignment framework that retains RL's verifier-bounded ceiling while incorporating dense token-level signals from a privileged self-teacher. $\textbf{CrEST}$ resolves credit at two levels: turn-segmented verified advantages address inter-turn dilution, while entropy-gated self-teacher modulation refines intra-turn token contributions. Experiments on BFCL V3 and WildToolBench show that $\textbf{CrEST}$ consistently outperforms both RL and distillation baselines across two model scales, with the largest gains on long-trajectory and strict session-level metrics. Our work demonstrates that the teacher's role in policy optimization can be reduced from determining update directions to modulating update magnitudes, unlocking dense credit assignment without sacrificing the verifier-bounded ceiling.

cs.AI

Systematic Evaluation of Stencil Configuration, Forcing Scheme, and Resolution Effects in the Stratified Taylor--Green Vortex: A Lattice Boltzmann Study

The rigorous simulation of stratified turbulence remains challenging due to pronounced flow anisotropy, suppressed vertical transport, and high sensitivity to numerical dissipation. This study systematically evaluates the predictive capability of the lattice Boltzmann method (LBM) for a three-dimensional stratified Taylor--Green vortex. Within a double-distribution-function framework under the Boussinesq approximation, we examine the influence of stencil configurations, forcing formulations, and spatial resolutions up to $256^3$, with validation against spectral DNS benchmarks. The results demonstrate that the D3Q27$\times$19 configuration achieves an optimal balance between numerical accuracy and computational efficiency, accurately reproducing the temporal evolution of kinetic and potential energies as well as the characteristic double-peak dissipation structure. Grid-sensitivity analysis further reveals that potential energy and fine-scale turbulent structures are significantly more resolution-dependent than kinetic energy, requiring a minimum resolution of $256^3$ for quantitative convergence. Moreover, under strongly stratified conditions, the velocity-shift forcing schemes outperform discrete source-term approaches, reducing the overall error by approximately 45.54\%. Overall, this work provides practical guidelines for high-fidelity LBM simulations of stratified turbulence and highlights that the coordinated selection of stencil isotropy, spatial resolution, and force discretization is essential for accurately capturing energy cascade and mixing dynamics.

physics.flu-dyn

Don't Just Fine-tune the Agent, Tune the Environment

Large Language Model (LLM) agents show great promise for complex, multi-turn tool-use tasks, but their development is often hampered by the extreme scarcity of high-quality training data. Supervised fine-tuning (SFT) on synthetic data leads to overfitting, whereas standard reinforcement learning (RL) struggles with a critical cold-start problem and training instability. To address these challenges, we introduce $\textbf{Environment Tuning}$, a novel training paradigm that enables agents to learn complex behaviors directly from problem instances without relying on pre-collected expert trajectories. $\textbf{Environment Tuning}$ orchestrates this learning process through a structured curriculum, actionable environment augmentation that provides corrective feedback, and fine-grained progress rewards to ensure stable and efficient exploration. Using only 400 problem instances from Berkeley Function-Calling Leaderboard (BFCL) benchmark, our method not only achieves competitive in-distribution performance against strong baselines but also demonstrates superior out-of-distribution generalization, overcoming the performance collapse common to SFT-based approaches. Our work presents a paradigm shift from supervised fine-tuning on static trajectories to dynamic, environment-based exploration, paving the way for training more robust and data-efficient agents. The code is available at https://github.com/inclusionAI/AWorld-RL/tree/main/EnvTuning.

cs.AI

RAG-R1: Incentivizing the Search and Reasoning Capabilities of LLMs through Multi-query Parallelism

Large Language Models (LLMs), despite their remarkable capabilities, are prone to generating hallucinated or outdated content due to their static internal knowledge. While Retrieval-Augmented Generation (RAG) integrated with Reinforcement Learning (RL) offers a solution, these methods are fundamentally constrained by a single-query mode, leading to prohibitive latency and inherent brittleness. To overcome these limitations, we introduce RAG-R1, a novel two-stage training framework centered around multi-query parallelism. Our framework enables LLMs to adaptively leverage internal and external knowledge during the reasoning process while transitioning from the single-query mode to multi-query parallelism. This architectural shift bolsters reasoning robustness while significantly reducing inference latency. Extensive experiments on seven question-answering benchmarks confirm the superiority of our method, which outperforms the strongest baseline by up to 13.7% and decreases inference time by 11.1%.

cs.CL

Influence of Membrane Characteristics on Efficiency of Vacuum Membrane Distillation: a Lattice Boltzmann Study

With increasing water scarcity, membrane distillation technology has gained widespread attention as an innovative method for seawater desalination.However,existing studies often overlook the influence of membrane characteristics on mass transfer efficiency. This study, based on the lattice Boltzmann method,proposes a model for a novel Poly(tetraethynylpyrene) membrane material to reveal the influence of membrane characteristics on the performance of vacuum membrane distillation. The model considers the factors such as porosity, tortuosity, membrane thickness, pore size, membrane surface wettability and temperature difference on the permeate flux. The results show that the permeate flux increases linearly with the porosity and decreases exponentially with the tortuosity factor. There is an optimal membrane thickness range (2μm) beyond which the permeate flux decreases exponentially. In addition, the permeate flux increases exponentially with increasing temperature difference and pore size. Further analysis of the effect of membrane surface wettability shows that permeate flux increases with increasing hydrophobicity. Finally, the feed temperature and tortuosity factor have the largest effect on permeate flux,followed by membrane thickness and, subsequently, pore size. The model can be further extended to study other configurations of membrane distillation technologies.

physics.flu-dyn

CSR:Achieving 1 Bit Key-Value Cache via Sparse Representation

The emergence of long-context text applications utilizing large language models (LLMs) has presented significant scalability challenges, particularly in memory footprint. The linear growth of the Key-Value (KV) cache responsible for storing attention keys and values to minimize redundant computations can lead to substantial increases in memory consumption, potentially causing models to fail to serve with limited memory resources. To address this issue, we propose a novel approach called Cache Sparse Representation (CSR), which converts the KV cache by transforming the dense Key-Value cache tensor into sparse indexes and weights, offering a more memory-efficient representation during LLM inference. Furthermore, we introduce NeuralDict, a novel neural network-based method for automatically generating the dictionary used in our sparse representation. Our extensive experiments demonstrate that CSR achieves performance comparable to state-of-the-art KV cache quantization algorithms while maintaining robust functionality in memory-constrained environments.

cs.CL

Experimental realization of direct entangling gates between dual-type qubits

Dual-type qubits have become a promising way to suppress the crosstalk error of auxiliary operations in large-scale ion trap quantum computation. Here we demonstrate a direct entangling gate between dual-type qubits encoded in the $S_{1/2}$ and $D_{5/2}$ hyperfine manifolds of $^{137}\mathrm{Ba}^{+}$ ions. Our scheme is economic in the hardware, requiring only a single $532\,$nm laser system to entangle both qubit types by driving their Raman transitions. We achieve a Bell state fidelity of $96.3(4)\%$ for the dual-type Molmer-Sorensen gate between an $S$-$D$ ion pair, comparable to that for the same-type $S$-$S$ or $D$-$D$ gates. This technique can reduce the overhead for back-and-forth conversions between dual-type qubits in the quantum circuit with wide applications in quantum error correction and ion-photon quantum networks.

quant-ph

Electromagnetically-Induced-Transparency Cooling of High-Nuclear-Spin Ions

We report the electromagnetically-induced-transparency (EIT) cooling of $^{137}\mathrm{Ba}^{+}$ ions with a nuclear spin of $I=3/2$, which are a good candidate of qubits for future large-scale trapped ion quantum computing. EIT cooling of atoms or ions with a complex ground-state level structure is challenging due to the lack of an isolated $Λ$ system, as the population can escape from the $Λ$ system to reduce the cooling efficiency. We overcome this issue by leveraging an EIT pumping laser to repopulate the cooling subspace, ensuring continuous and effective EIT cooling. We cool the two radial modes of a single $^{137}\mathrm{Ba}^{+}$ ion to average motional occupations of 0.08(5) and 0.15(7) respectively. Using the same laser parameters, we also cool all the ten radial modes of a five-ion chain to near their ground states. Our approach can be adapted to atomic species possessing similar level structures. It allows engineering of the EIT Fano-like spectrum, which can be useful for simultaneous cooling of modes across a wide frequency range, aiding in large-scale trapped-ion quantum information processing.

quant-ph

Fast Chain-of-Thought: A Glance of Future from Parallel Decoding Leads to Answers Faster

In this work, we propose FastCoT, a model-agnostic framework based on parallel decoding without any further training of an auxiliary model or modification to the LLM itself. FastCoT uses a size-varying context window whose size changes with position to conduct parallel decoding and auto-regressive decoding simultaneously, thus fully utilizing GPU computation resources. In FastCoT, the parallel decoding part provides the LLM with a quick glance of the future composed of approximate tokens, which could lead to faster answers compared to regular autoregressive decoding used by causal transformers. We also provide an implementation of parallel decoding within LLM, which supports KV-cache generation and batch processing. Through extensive experiments, we demonstrate that FastCoT saves inference time by nearly 20% with only a negligible performance drop compared to the regular approach. Additionally, we show that the context window size exhibits considerable robustness for different tasks.

cs.CL

GreenFlow: A Computation Allocation Framework for Building Environmentally Sound Recommendation System

Given the enormous number of users and items, industrial cascade recommendation systems (RS) are continuously expanded in size and complexity to deliver relevant items, such as news, services, and commodities, to the appropriate users. In a real-world scenario with hundreds of thousands requests per second, significant computation is required to infer personalized results for each request, resulting in a massive energy consumption and carbon emission that raises concern. This paper proposes GreenFlow, a practical computation allocation framework for RS, that considers both accuracy and carbon emission during inference. For each stage (e.g., recall, pre-ranking, ranking, etc.) of a cascade RS, when a user triggers a request, we define two actions that determine the computation: (1) the trained instances of models with different computational complexity; and (2) the number of items to be inferred in the stage. We refer to the combinations of actions in all stages as action chains. A reward score is estimated for each action chain, followed by dynamic primal-dual optimization considering both the reward and computation budget. Extensive experiments verify the effectiveness of the framework, reducing computation consumption by 41% in an industrial mobile application while maintaining commercial revenue. Moreover, the proposed framework saves approximately 5000kWh of electricity and reduces 3 tons of carbon emissions per day.

cs.IR

On the Opportunities of Green Computing: A Survey

Artificial Intelligence (AI) has achieved significant advancements in technology and research with the development over several decades, and is widely used in many areas including computing vision, natural language processing, time-series analysis, speech synthesis, etc. During the age of deep learning, especially with the arise of Large Language Models, a large majority of researchers' attention is paid on pursuing new state-of-the-art (SOTA) results, resulting in ever increasing of model size and computational complexity. The needs for high computing power brings higher carbon emission and undermines research fairness by preventing small or medium-sized research institutions and companies with limited funding in participating in research. To tackle the challenges of computing resources and environmental impact of AI, Green Computing has become a hot research topic. In this survey, we give a systematic overview of the technologies used in Green Computing. We propose the framework of Green Computing and devide it into four key components: (1) Measures of Greenness, (2) Energy-Efficient AI, (3) Energy-Efficient Computing Systems and (4) AI Use Cases for Sustainability. For each components, we discuss the research progress made and the commonly used techniques to optimize the AI efficiency. We conclude that this new research direction has the potential to address the conflicts between resource constraints and AI development. We encourage more researchers to put attention on this direction and make AI more environmental friendly.

cs.AI

Spatially Resolved Properties of Supernova Host Galaxies in SDSS-IV MaNGA

We crossmatch galaxies from Mapping Nearby Galaxies at Apache Point Observatory with the Open Supernova Catalog, obtaining a total of 132 SNe within MaNGA bundle. These 132 SNe can be classified into 67 Type Ia and 65 Type CC. We study the global and local properties of supernova host galaxies statistically. Type Ia SNe are distributed in both star-forming galaxies and quiescent galaxies, while Type CC SNe are all distributed along the star-forming main sequence. As the stellar mass increases, the Type Ia/CC number ratio increases. We find: (1) there is no obvious difference in the interaction possibilities and environments between Type Ia SN hosts and a control sample of galaxies with similar stellar mass and SFR distributions, except that Type Ia SNe tend to appear in galaxies which are more bulge-dominated than their controls. For Type CC SNe, there is no difference between their hosts and the control galaxies in galaxy morphology, interaction possibilities as well as environments; (2) the SN locations have smaller velocity dispersion, lower metallicity, and younger stellar population than galaxy centers. This is a natural result of radius gradients of all these parameters. The SN location and the its symmetrical position relative to the galaxy center, as well as regions with similar effective radii have very similar [Mg/Fe], gas-phase metallicity, gas velocity dispersion and stellar population age.

astro-ph.GA

Dynamic DNN Decomposition for Lossless Synergistic Inference

Deep neural networks (DNNs) sustain high performance in today's data processing applications. DNN inference is resource-intensive thus is difficult to fit into a mobile device. An alternative is to offload the DNN inference to a cloud server. However, such an approach requires heavy raw data transmission between the mobile device and the cloud server, which is not suitable for mission-critical and privacy-sensitive applications such as autopilot. To solve this problem, recent advances unleash DNN services using the edge computing paradigm. The existing approaches split a DNN into two parts and deploy the two partitions to computation nodes at two edge computing tiers. Nonetheless, these methods overlook collaborative device-edge-cloud computation resources. Besides, previous algorithms demand the whole DNN re-partitioning to adapt to computation resource changes and network dynamics. Moreover, for resource-demanding convolutional layers, prior works do not give a parallel processing strategy without loss of accuracy at the edge side. To tackle these issues, we propose D3, a dynamic DNN decomposition system for synergistic inference without precision loss. The proposed system introduces a heuristic algorithm named horizontal partition algorithm to split a DNN into three parts. The algorithm can partially adjust the partitions at run time according to processing time and network conditions. At the edge side, a vertical separation module separates feature maps into tiles that can be independently run on different edge nodes in parallel. Extensive quantitative evaluation of five popular DNNs illustrates that D3 outperforms the state-of-the-art counterparts up to 3.4 times in end-to-end DNN inference time and reduces backbone network communication overhead up to 3.68 times.

cs.DC

On the Weyl's law for discretized elliptic operators

In this paper we give an estimate on the asymptotic behavior of eigenvalues of discretized elliptic boundary values problems. We first prove a simple min-max principle for selfadjoint operators on a Hilbert space. Then we show two sided bounds on the $k$-th eigenvalue of the discrete Laplacian by the $k$-th eigenvalue of the continuous Laplacian operator under the assumption that the finite element mesh is quasi-uniform. Combining this result with the well-known Weyl's law, we show that the $k$-th eigenvalue of the discretized isotropic elliptic operators, spectrally equivalent to the discretized Laplacian, is $\mathcal O\left(k^{2/d}\right)$. Finally, we show how these results can be used to obtain an error estimate for finite element approximations of elliptic eigenvalue problems.

math.NA

A unified approach to the design and analysis of AMG

In this work, we present a general framework for the design and analysis of two-level AMG methods. The approach is to find a basis for locally optimal or quasi-optimal coarse space, such as the space of constant vectors for standard discretizations of scalar elliptic partial differential equations. The locally defined basis elements are glued together using carefully designed linear extension maps to form a global coarse space. Such coarse spaces, constructed locally, satisfy global approximation property and by estimating the local Poincar{\' e} constants, we obtain sharp bounds on the convergence rate of the resulting two-level methods. To illustrate the use of the theoretical framework in practice, we prove the uniform convergence of the classical two level AMG method for finite element discretization of a jump coefficient problem on a shape regular mesh.

math.NA