SearcharxivSearch

arXiv subjects

Zihao Wu

Publications and source records attributed to Zihao Wu.

At least 19 recordsLinked to original sources

The Roman eXtreme Deep Field (RXDF)

The Roman eXtreme Deep Field (RXDF) program is one of the five General Astrophysics Survey (GAS) programs approved for observing time with the Nancy Grace Roman Space Telescope in Cycles 1 and 2. It has been allocated 386.41 hours to carry out an imaging survey to AB = 30 mag (5-sigma) over ~140x larger area than the Hubble eXtreme Deep Field (HXDF) full-depth area (ACS+WFC3/IR). The RXDF will cover the full Roman wavelength range with 7 bands, reaching AB = 30 mag in RZYJH, 29 mag in F, and 28 mag in K, over a full-depth area of 678.75 arcmin^2 embedded in a total area of 1,243 arcmin^2, and far exceeding the depths of the Roman Core Community Surveys (CCS). The RXDF is within the Euclid Ultra Deep Field (EUDF) near the North Ecliptic Pole (NEP), a strategic long-term field for generational space facilities, with a wealth of multi-wavelength data including extensive coverage from the James Webb Space Telescope (JWST) NEXUS Treasury program. The observations will cover 3 epochs at a 1-year cadence, each epoch divided into 3 sub-epochs ~10 days apart, enabling time-domain studies on time baselines from ~10 days to over ~2 years. The RXDF is uniquely positioned to address critical questions in reionization, large scale structure (LSS), growth of supermassive black holes (SMBHs), little red dots (LRDs), and high-z supernovae (SNe); the volumes probed by HST+JWST are too small at these extreme depths, and even the deepest CCS tiers are too shallow. In addition to our key objectives, a wealth of additional science will be enabled by engaging the community with our rapidly released datasets, revolutionizing a wide range of science for a lasting legacy. This short document, which is converted from the approved RXDF proposal, aims to provide the community with a summary of the program.

astro-ph.GA

WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation

Massively parallel simulation changes the data regime in which off-policy reinforcement learning (RL) is trained, challenging stabilizers designed for data-limited replay. Through controlled experiments across eight benchmark families, we show that these stabilizers are data-regime-dependent: parameter normalization helps with narrow replay coverage but restricts value fitting when data are abundant, while clipped double-Q can be relaxed in high-throughput manipulation. Age-biased replay weighting improves learning efficiency across regimes, especially with limited network capacity. Based on these findings, we propose WarpSAC, a regime-aware family of off-policy RL algorithms. WarpSAC uses Sample Weight Decay for efficient exploitation and provides two variants: WarpSAC-L (Norm ON, clipped double-Q) for data-limited CPU-scale training, and WarpSAC-A (Norm OFF, single-Q) for data-abundant GPU-parallel training. WarpSAC improves normalized score--step AUC over FlashSAC by 4.5% across nine CPU-scale environments and 23.1% across fourteen GPU-parallel environments. It increases UnitreeG1TransportBox-v1 success rate from 19.8% to 96.4%, improves mean normalized wall-time AUC on MuJoCo Playground by 19.1%, and achieves 36.4% faster sim-to-real deployment on Unitree G1 than FlashSAC. These results show that scalable off-policy RL should adapt its stabilizers to the available data regime.

cs.LG

BioMed-Agent-RL: A Meta Learning, All You Need for Biomedical Applications

The current progress of Clinical Vision Large Language Models (C-VLLMs) has substantially improved digital diagnostics, still these frameworks often endure lesion noises, modality misalignment, hallucination, and missed contextual grounding in complex clinical cases. Moreover, prevailing agent systems usually depend on static and non-adaptable pipelines and lack the versatility necessary for complex medical reasoning. To resolve these difficulties, we present BioMed-Agent-RL, a unified medical agent that incorporates adaptive orchestration, policy, and reward-based reinforcement learning (RL) models for biomedical applications. To ensure reliability, it invokes clinical context-aware preference optimization (CPO), direct preference optimization (DPO), and group relative policy optimization (GRPO) with dynamic entropy regulation. This pipeline utilizes a multimodal meta-learning approach that operates as a field-specific expert and human judgment synthesizer. The agent adaptively utilizes a set of model-level expertise, such as clinical grounding and reasoner, lesion segmenter, and field-specific synthesizer, across various clinical modalities (e.g., X-ray) by utilizing an iterative and adaptive RL approach. The agent learns to seriously synthesize misleading, conflicting vision cues and trust in inherent reasoning, while specialist advice is faulty. An intensive ablation study is conducted across multiple benchmarks, and the agent significantly outperforms existing state of the art models, such as GPT-5, attaining up to ~73% accuracy (gain of ~5%) over contemporary baselines. As a result, the framework suggests a new standard for building factual, reliable, robust, and expert-like intelligent agent systems for independent clinical reasoning.

cs.LG

Low Ly$\alpha$ Visibility in Galaxy Overdensities: Reionization Topology and Neutral-Fraction Ceilings from DIVER over $4.8<z<11$

Ly-alpha emission is widely used to trace cosmic reionization, but its interpretation depends on how Ly-alpha visibility varies with galaxy environment. We use deep JWST/NIRSpec observations from Deep Insights into UV Spectroscopy at the Epoch of Reionization (DIVER) in GOODS-N to measure Ly-alpha visibility for 250 galaxies at 4.8 25 A. We combine these measurements with H-alpha and [O III] emitters from JWST/NIRCam wide-field slitless spectroscopy to map the density field around each DIVER galaxy. Galaxies with high Ly-alpha equivalent widths (W_Lyalpha>25 A) or high effective Ly-alpha escape fractions (f_esc,Lyalpha^eff>0.05) tend to lie farther from nearby H-alpha and [O III] emitters than galaxies with lower Ly-alpha visibility. The clearest signal occurs near the prominent GOODS-N overdensity at z~5.2, where fewer than 15% of galaxies show strong Ly-alpha emission. This trend is opposite to the simplest inside-out reionization expectation that overdensities produce larger ionized regions and enhance Ly-alpha visibility. Possible explanations include circumgalactic and local intergalactic opacity, dense absorbers, and gas kinematics. We also derive an empirical upper envelope for f_esc,Lyalpha^eff and calibrate it with reionization simulations. Interpreting this envelope as a limiting IGM-attenuation signal gives neutral-fraction ceilings of _max=0.36, 0.76, 0.74, 0.84, and 1.0 at z~5.2, 5.8, 6.7, 7.7, and 9.8, respectively. The z~8 ceiling disfavors an almost completely neutral IGM at this epoch. These results support patchy reionization already underway by z~8 and show that galaxy Ly-alpha visibility encodes both large-scale ionization topology and near-source gas structure.

astro-ph.GA

PhotoIFU: NIRCam as a Photometric Integral Field Unit for Mapping Feedback in Galaxies

We present PhotoIFU, a workflow that uses deep multi-band imaging as a low-resolution photometric integral field unit. Applied to PSF-matched JWST/NIRCam imaging, PhotoIFU treats each spatial pixel as a coarse SED element and fits the pixel SEDs with Prospector to map resolved stellar-population and ISM-related properties. We apply this approach to three galaxies at $z=1.3$--3.7 in JADES: two systems with extended ionized line emission and one post-starburst galaxy with an exceptionally strong neutral outflow. Pixel-by-pixel SED fitting gives maps of stellar-mass surface density, specific star formation rate, dust attenuation, gas-phase metallicity, and recent star-formation history. We find that regions selected from the extended-emission or outflow geometry occupy distinct parts of the resolved SED-property distribution compared with the full host. In the systems with extended ionized emission, these regions are generally less dusty, consistent with ionized emission being observed along dust-poor, low-column-density pathways through the host. In the neutral-outflow system, the selected regions show enhanced recent star formation, suggesting that compact rejuvenation may mark the aftermath of an earlier energetic phase. These results show that galactic outflows and extended emission-line structures can be spatially associated with measurable differences in resolved host-galaxy stellar populations and ISM-related properties. PhotoIFU provides an imaging-based method for resolved SED mapping of feedback-related structures in larger galaxy samples where full spectroscopic integral-field mapping is unavailable.

astro-ph.GA

WuYuEval: A Multi-Level Benchmark for Large Language Models in Solid Waste Management

Large language models (LLMs) are increasingly used as technical assistants, but their competence in solid waste management (SWM) remains difficult to assess because existing benchmarks emphasize general knowledge rather than professional decisions under engineering, environmental, and policy constraints. We introduce WuYuEval, a multi-level benchmark for evaluating LLMs in SWM across foundational knowledge, domain reasoning, and expert decision-making. After quality auditing, WuYuEval contains a Foundation Module with 4,590 closed-ended multiple-choice questions across six task types and eight domain categories, together with an Expert Module with 247 scenario-based open-ended questions involving multi-objective optimization, constraint trade-offs, and system design. For expert tasks, we combine anchor-calibrated LLM-as-a-Judge scoring with Elo-based pairwise comparison. Across 33 LLMs, performance varied widely. The leading model reached 94.64\% accuracy on the Foundation Module, but average accuracy still fell from 84.14\% on easy questions to 42.50\% on hard questions, with lower performance concentrated in calculation, experimental design, urban planning, and open-ended expert tasks. Reasoning-oriented Thinking modes improve most matched model pairs after auditing, but the gains depend on baseline capability and are not uniformly positive. These results suggest that visible deliberation helps only when it remains anchored to units, assumptions, and engineering constraints; otherwise, it may drift from decisive answer boundaries. WuYuEval therefore provides both an evaluation resource and an empirical basis for developing SWM-oriented foundation models with professional reasoning chains and explicit constraint control.

cs.CL

JADES: the mass-metallicity relation at $z=1-10$. New calibrations, extremely metal-poor galaxies, and chemical diversity

We present gas-phase metallicities of star-forming galaxies at $z=1$-10 with deep JWST/NIRSpec spectra from the JADES full data release, Dark Horse, and OASIS programmes. We stack $\sim$1500 medium-resolution spectra, yielding detections of the [OIII]$\lambda$4363 auroral line down to $12+\log(\mathrm{O/H})=7.0$ to derive stack-based strong-line calibrations over the metallicity range $12+\log(\mathrm{O/H})=7.0$-8.7. At a fixed metallicity, our stacks exhibit [OIII]$\lambda$5007/H$\beta$ and [OIII]$\lambda$5007/[OII]$\lambda\lambda$3726,3729 values generally lower than calibrations based on high-$z$ individual auroral-line emitters, suggesting an observational bias towards higher excitation introduced when requiring auroral line detections in individual spectra. Based on our new calibrations, we obtain canonical mass-metallicity relations (MZRs) at z$=$1-10, identifying a decrease in metallicities from $z\sim0$ to z$\sim$4-10, without significant change in slope. Moreover, we identify 50 promising candidates of extremely metal-poor galaxies (EMPGs) with $12+\log(\mathrm{O/H})=6.7$-7.3 (1-4\% solar metallicity) at $z=1.2$-9.1. The MZRs of EMPGs are characterised by a large scatter, with those having lower metallicities generally exhibiting lower sSFRs, opposite of what expected from the local Fundamental Metallicity Relation. These results support a stochastic star-formation history involving gas consumption/ejection and metal-poor inflow, strongly affecting metallicities of low-mass galaxies. Furthermore, we identify two Little Red Dots in our EMPG candidates, both exhibiting broad H$\alpha$ and prominent Ly$\alpha$, offering insights into the early black-hole growth in extremely metal-poor environments.

astro-ph.GA

On the spanning cuts consistency problem in the IBP reductions of Feynman integrals

The spanning cuts method is a powerful approach to reduce the cost of IBP reduction while computing Feynman integrals. However, its usage is limited due to the so-called consistency problem. It was unclear why the IBP reduction coefficients can be inconsistent with each other between different cuts. In this paper, we report a mechanism behind this inconsistency. We found that the IBP relations can be violated under the cuts, if we blindly erase the hidden terms that are proportional to the ``vanishing'' Feynman prescription parameters in the relations. In some cases, the cut introduces pinch singularities, which cancel the vanishing Feynman prescription parameters, making the hidden terms finite. In various cases, the error comes from omitting such finite hidden terms. We also claimed that the pinch singularity under the cuts are related to some hidden relations between the propagators. In this paper, we provide an algorithm and its implementation to find the linear hidden relations.

hep-ph

World Models: A Comprehensive Survey of Architectures, Methodologies, Reasoning Paradigms, and Applications

World models, internal simulators that learn the structure and dynamics of an environment, have emerged as a central paradigm in the pursuit of artificial general intelligence, enabling agents to predict, plan, and reason within learned representations. Despite rapid progress across reinforcement learning, robotics, autonomous driving, and video generation, the field lacks a unified framework integrating its diverse architectural choices, training methods, reasoning mechanisms, and application settings. This survey addresses that gap with a multi-axis taxonomy organized along four dimensions: (i) architecture, encompassing representation format, dynamics formulation, input modality, learning paradigm, and downstream application; (ii) methodological family, including state-space and recurrent approaches, transformer-based models, diffusion-based generators, physics-informed networks, and language-augmented multimodal systems; (iii) reasoning strategy, covering imagination-based planning, latent policy learning, counterfactual reasoning, and planning under uncertainty; and (iv) application domain, spanning robotics, autonomous driving, video prediction, multimodal agents, reinforcement learning, scientific modeling, medical imaging, educational measurement, and business and finance. Tracing the field from early cognitive-science foundations to milestone systems such as PlaNet, the Dreamer family, MuZero, Sora, Cosmos, and Genie, we examine how these dimensions interact and highlight the recent convergence of chain-of-thought reasoning with world-model imagination. We review evaluation protocols and benchmarks, identify persistent challenges such as compounding prediction errors, sim-to-real transfer, and fragmented evaluation, and outline future directions toward unified multimodal world models, foundation-scale interactive simulators, and safe deployment in safety-critical domains.

cs.LG

JWST Advanced Deep Extragalactic Survey (JADES) Data Release 5: stellar population catalogue for galaxies in GOODS-N and GOODS-S

We present the galaxy stellar population catalogue from the JWST Advanced Deep Extragalactic Survey (JADES) Data Release 5 (DR5), providing homogeneous Bayesian inference of physical galaxy properties in GOODS-N and GOODS-S. Using deep JWST/NIRCam and MIRI imaging combined with ancillary multi-wavelength data, we model the spectral energy distributions of ~500,000 sources with the Prospector framework. Our modelling incorporates flexible non-parametric star-formation histories (SFHs), nebular emission, dust attenuation, metallicities, and mid-infrared AGN and dust emission. We adopt an evolving star-forming main sequence (SFMS) prior for modelling the SFHs, which provides a physically-motivated long-term shape of SFHs while retaining non-parametric flexibility. The prior links stellar mass growth and SFR through the observed redshift-dependent SFMS, shaping the global behaviour of the inferred SFHs but allowing substantial deviations and scatters wherever supported by the data. We derive posterior distributions for stellar masses, SFRs, SFHs, dust attenuation, metallicities, and AGN contributions. The depth and wavelength coverage of JADES enable robust stellar mass measurements down to low-mass limits, as well as improved constraints on recent star-formation activity for ~350,000 galaxies at z = 1 - 9. The adoption of a physically motivated prior mitigates unphysical solutions and reduces degeneracies between redshift, age, dust, and metallicity, particularly for faint sources. We validate the catalogue through consistency checks and comparison to spectroscopic redshifts where available. The resulting value-added catalogue provides a uniform set of stellar population parameters suitable for statistical studies of galaxy growth, quenching, and the build-up of stellar mass across cosmic time. The full catalogue and posterior summaries are publicly released as part of JADES DR5.

astro-ph.GA

A World Model of Radiologist Reading for Medical Image Representation Learning

Radiologist eye-tracking data provide a rich record of how experts search, compare, and accumulate evidence during image reading; yet, existing methods exploit this signal only partially, either as a static spatial prior or as an auxiliary prediction target decoupled from diagnosis. We propose GazeWorld, a medical imaging world model that treats the image as the world and the radiologist's fixation sequence as a trajectory through it. GazeWorld autoregressively predicts the latent representation of the next fixated patch from all previously visited ones, while a spatial-completion branch covers unvisited regions. At inference, GazeWorld generates a sequence of patch representations from the image alone without requiring real gaze data. Frozen GazeWorld features achieve state-of-the-art diagnostic accuracy across all nine supervised settings on CheXpert, RSNA Pneumonia, and SIIM-ACR Pneumothorax, as well as the highest zero-shot accuracy on all three benchmarks. On the GazeSearch benchmark, a generic decoder trained on the same frozen features outperforms the purpose-built LogitGaze-Med by over 16\% in ScanMatch and 22\% in SED, despite not being explicitly trained to predict gaze. GazeWorld demonstrates that modeling how experts read, not just what they conclude, offers a promising pretraining paradigm for medical imaging AI.

cs.CV

Same Brain, Different Prediction: How Preprocessing Choices Undermine EEG Decoding Reliability

Electroencephalography (EEG) is a cornerstone of brain-computer interfaces and clinical neuroscience, yet deep learning models are typically trained and evaluated under a single, unreported preprocessing pipeline. We formalize preprocessing choices as a counterfactual intervention space and show that EEG predictions are surprisingly unstable under this space: across six datasets spanning four paradigms, up to 42% of trial-level predictions flip when only the preprocessing changes, a variability that standard uncertainty methods do not explicitly quantify because they condition on a fixed preprocessing pipeline. We provide three tools to make this instability measurable, decomposable, and reducible. First, a Walsh-Hadamard decomposition of the 2^7 pipeline space reveals that sensitivity is near-additive in practice under the binary intervention design, enabling efficient step-by-step optimization. Second, we introduce Preprocessing Uncertainty (PU), a per-trial diagnostic that captures a dimension of instability complementary to model-based confidence. Third, we study Normalized Adaptive PGI (NA-PGI), a graph-structured regularizer that exploits the compositional structure of preprocessing interventions as one mitigation strategy with clear scope conditions.

cs.LG

YOTOnet: Zero-Shot Cross-Domain Fault Diagnosis via Domain-Conditioned Mixture of Experts

Mechanical equipment forms the critical backbone of modern industrial production, yet domain shift severely limits the generalization of deep learning based fault diagnosis models across different equipment and operating conditions.Inspired by the success of foundation models in achieving zero-shotgeneralization, we propose YOTOnet (You Only Train Once), a novel architecture specifically designed for cross-domain fault diagnosis in mechanical equipment.YOTOnet comprises three core components: (1) a physics-aware Invariant Feature Distiller that extracts domain-agnostic representations using multi-scale dilated convolutions and FFT-based time-frequency fusion,(2) Domain-Conditioned Sparse Experts (DC-MoE) that adaptively route inputs to specialized processors via learned gating without external meta-data, and (3) a dual-head classification system with auxiliary supervision.Extensive validation on five public bearing datasets (CWRU, MFPT, XJTU,OTTAWA, HUST) through 30 cross-dataset protocols demonstrates the superiority of YOTOnet compared with other state-of-the-art methods. Critically, we observe a clear scaling effect-average test F1 improves from 0.5339(1 training dataset) to 0.705 (4 datasets), with a clear gain when moving from 3 to 4 datasets. These findings provide empirical evidence that foundation model principles can enable robust, train-once deployment for industrial fault diagnosis.

cs.LG

Vibe Medicine: Redefining Biomedical Research Through Human-AI Co-Work

With the emergence of large language models (LLMs) and AI agent frameworks, the human-AI co-work paradigm known as Vibe Coding is changing how people code, making it more accessible and productive. In scientific research, where workflows are more complex and the burden of specialized labor limits independent researchers and those in low-resource areas, the potential impact is even greater, particularly in biomedicine, which involves heterogeneous data modalities and multi-step analytical pipelines. In this paper, we introduce Vibe Medicine, a co-work paradigm in which clinicians and researchers direct skill-augmented AI agents through natural language to execute complex, multi-step biomedical workflows, while retaining the role of research director who specifies objectives, reviews intermediate results, and makes domain-informed decisions. The enabling infrastructure consists of three layers: capable LLMs, agent frameworks such as OpenClaw and Hermes Agent, and the OpenClaw medical skills collection, which includes more than 1,000 curated skills from multiple open-source repositories. We analyze the architecture and skill categories of this collection across ten biomedical domains, and present case studies covering rare disease diagnosis, drug repurposing, and clinical trial design that demonstrate end-to-end workflows in practice. We also identify the principal risks, such as hallucination, data privacy, and over-reliance, and outline directions toward more reliable, trustworthy, and clinically integrated agent-assisted research that advances research and technological equity and reduces health care resource disparities.

cs.AI

Intense and extended CIII] emission suggests a strong outflow in JADES-GS-z14-0

JWST has revealed an overabundance of very bright, blue galaxies at z>10, raising fundamental questions about how star formation and feedback operate at Cosmic Dawn. We present new JWST/NIRSpec MSA PRISM/CLEAR spectroscopy of JADES-GS-z14-0 (z=14.18) obtained with the JADES and OASIS programmes. While the rest-frame UV continuum flux level and shape are consistent between the two datasets, the OASIS spectrum shows a 10$\sigma$ detection of the CIII]$\lambda\lambda1907,1909$ emission line, with a luminosity three times higher than that measured in the JADES data. This difference is naturally explained by the offset in shutter placement between OASIS and JADES, implying that the CIII] emission is spatially displaced by $\sim400$ pc from the stellar continuum. The non-detection of CIII] in NIRCam medium-band imaging indicates that the emitting region is extended on scales $\gtrsim165$ pc, with a surface brightness below the detection threshold. Interpreting this diffuse, carbon-enriched gas as the result of ongoing or past outflows, we infer a mass outflow rate of $\dot{M}_{\rm out}\sim160~{\rm M_\odot\,yr^{-1}}$. We compare it with the star-formation rate (SFR) and derive a mass-loading factor of $\eta = \dot{M}_{\rm out}/{\rm SFR} = 4-15$, suggesting highly efficient feedback at very early times. Finally, we show that, if outflows are one of the mechanisms regulating star formation in JADES-GS-z14-0, the instantaneous star-formation efficiency in massive haloes is constrained to $\epsilon_\star\lesssim0.08$. These results support a scenario in which outflows play a crucial role during the earliest phases of galaxy formation. Comparing our results with the current theoretical galaxy formation model, we conclude that a combination of moderate star-formation efficiency and reduced dust attenuation can account for the emergence of luminous galaxies at the highest redshifts.

astro-ph.GA

The Way We Tally Becomes the Tale: the Impact of Selection Strategies on the Inferred Evolution of Little Red Dots Across Cosmic Time

Little Red Dots (LRDs) have emerged as a key population linked to early black hole growth, yet photometric selections have predominantly targeted only the most extreme red systems, thereby shaping our current understanding of this new population of objects. In this work, we deliberately explore a broad range of optical redness while enforcing stringent compactness and visual inspection to ensure robustness and minimize contamination. Leveraging the depth and multiwavelength coverage of the JWST Advanced Deep Extragalactic Survey (JADES) data in the GOODS-North and GOODS-South fields, we construct the largest photometric census of LRDs to date in these fields, comprising 412 sources over $z\approx2\text{--}11$ across $\approx349.6$ arcmin$^2$. We show that classic extreme color cuts isolate only a minor fraction of this population ($\lesssim25\%$), while the majority of LRDs span a broader, largely unexplored parameter space. We quantify how selection strategies impact UV and optical luminosity functions and number density evolution, finding that current demographic trends of LRDs are strongly driven by selection biases and further limited by incomplete identification at both high and low redshift. Spectroscopically confirmed LRDs reveal a continuous range of spectral shapes, consistent with varying Active Galactic Nucleus (AGN) and host contributions in agreement with recent findings. Our results demonstrate that commonly adopted, purity-driven selections bias current demographic constraints toward the most extreme systems, potentially misrepresenting the diversity and evolution of the LRD population. Accounting for these selection effects is essential for interpreting LRDs and their role in early black hole growth.

astro-ph.GA

Rethinking Token Prediction: Tree-Structured Diffusion Language Model

Discrete diffusion language models have emerged as a competitive alternative to auto-regressive language models, but training them efficiently under limited parameter and memory budgets remains challenging. Modern architectures are predominantly based on a full-vocabulary token prediction layer, which accounts for a substantial fraction of model parameters (e.g., more than 20% in small scale DiT-style designs) and often dominates peak GPU memory usage. This leads to inefficient use of both parameters and memory under constrained training resources. To address this issue, we revisit the necessity of explicit full-vocabulary prediction, and instead exploit the inherent structure among tokens to build a tree-structured diffusion language model. Specifically, we model the diffusion process with intermediate latent states corresponding to a token's ancestor nodes in a pre-constructed vocabulary tree. This tree-structured factorization exponentially reduces the classification dimensionality, makes the prediction head negligible in size, and enables reallocation of parameters to deepen the attention blocks. Empirically, under the same parameter budget, our method reduces peak GPU memory usage by half while matching the perplexity performance of state-of-the-art discrete diffusion language models.

cs.CL

The Rank and Gradient Lost in Non-stationarity: Sample Weight Decay for Mitigating Plasticity Loss in Reinforcement Learning

Deep reinforcement learning (RL) suffers from plasticity loss severely due to the nature of non-stationarity, which impairs the ability to adapt to new data and learn continually. Unfortunately, our understanding of how plasticity loss arises, dissipates, and can be dissolved remains limited to empirical findings, leaving the theoretical end underexplored.To address this gap, we study the plasticity loss problem from the theoretical perspective of network optimization. By formally characterizing the two culprit factors in online RL process: the non-stationarity of data distributions and the non-stationarity of targets induced by bootstrapping, our theory attributes the loss of plasticity to two mechanisms: the rank collapse of the Neural Tangent Kernel (NTK) Gram matrix and the $\Theta(\frac{1}{k})$ decay of gradient magnitude. The first mechanism echoes prior empirical findings from the theoretical perspective and sheds light on the effects of existing methods, e.g., network reset, neuron recycle, and noise injection. Against this backdrop, we focus primarily on the second mechanism and aim to alleviate plasticity loss by addressing the gradient attenuation issue, which is orthogonal to existing methods. We propose Sample Weight Decay -- a lightweight method to restore gradient magnitude, as a general remedy to plasticity loss for deep RL methods based on experience replay. In experiments, we evaluate the efficacy of \methodName upon TD3, \myadded{Double DQN} and SAC with SimBa architecture in MuJoCo, \myadded{ALE} and DeepMind Control Suite tasks. The results demonstrate that \methodName effectively alleviates plasticity loss and consistently improves learning performance across various configurations of deep RL algorithms, UTD, network architectures, and environments, achieving SOTA performance on challenging DMC Humanoid tasks.

cs.LG