SearcharxivSearch

arXiv subjects

Youwei Liu

Publications and source records attributed to Youwei Liu.

4 recordsLinked to original sources

CoMAP: Co-Evolving World Models and Agent Policies for LLM Agents

Equipping language agents with world models enables them to anticipate environment dynamics and evaluate candidate actions before execution. However, existing textual world models are typically fixed after training, preventing them from adapting to the on-policy state-action distributions induced by an evolving agent. Meanwhile, agent-improvement methods often rely on external rewards or verifiers, limiting their applicability in realistic interactive environments. In this paper, we propose COMAP, a novel framework that co-evolves textual world models and agent policies through closed-loop interaction. At each decision step, the world model predicts future state feedback for candidate actions, and the agent performs future-aware reflection by estimating the reliability of this feedback and refining its action accordingly. The resulting on-policy trajectories are then used to update the world model via self-distillation, allowing it to better match the agent's evolving interaction distribution. Across embodied task planning, Web navigation, and tool-use benchmarks, COMAP consistently outperforms competitive baselines, e.g., +16.75% relative improvement with Qwen3-4B. Further analyses show that the co-evolutionary loop improves the world model's prediction accuracy over time and leads to more effective long-horizon decision-making. Our code is available at: https://github.com/loyiv/CoMAP.

cs.AI

Imagine-then-Plan: Agent Learning from Adaptive Lookahead with World Models

Recent advances in world models have shown promise for modeling future dynamics of environmental states, enabling agents to reason and act without accessing real environments. Current methods mainly perform single-step or fixed-horizon rollouts, leaving their potential for complex task planning under-exploited. We propose Imagine-then-Plan (\texttt{ITP}), a unified framework for agent learning via lookahead imagination, where an agent's policy model interacts with the learned world model, yielding multi-step ``imagined'' trajectories. Since the imagination horizon may vary by tasks and stages, we introduce a novel adaptive lookahead mechanism by trading off the ultimate goal and task progress. The resulting imagined trajectories provide rich signals about future consequences, such as achieved progress and potential conflicts, which are fused with current observations, formulating a partially \textit{observable} and \textit{imaginable} Markov decision process to guide policy learning. We instantiate \texttt{ITP} with both training-free and reinforcement-trained variants. Extensive experiments across representative agent benchmarks demonstrate that \texttt{ITP} significantly outperforms competitive baselines. Further analyses validate that our adaptive lookahead largely enhances agents' reasoning capability, providing valuable insights into addressing broader, complex tasks. Our code and data will be publicly available at https://github.com/loyiv/ITP.

cs.CL

Sensitivity of BEACON to Ultra-High Energy Diffuse and Transient Neutrinos

Ultra-high energy neutrinos ($E>10^{17}$ eV) can provide insight into the most powerful accelerators in the universe, however their flux is extremely low. The Beamforming Elevated Array for COsmic Neutrinos (BEACON) is a detector concept which efficiently achieves sensitivity to this flux by employing phased radio arrays on mountains, which search for the radio emission of up-going extensive air showers created by Earth-skimming tau neutrinos. Here, we calculate the point-source effective area of BEACON and characterize its sensitivity to transient neutrino fluences with both short ($<15$ min) and long ($> 1$ day) durations. Additionally, by integrating the effective area, we provide an updated estimate of the diffuse flux sensitivity. With just 100 stations, BEACON achieves sensitivity to short-duration transients such as nearby short gamma-ray bursts. With 1000 stations, BEACON achieves a sensitivity to long-duration transients, as well as the cosmogenic flux, ten times greater than existing experiments at 1 EeV. With an efficient design optimized for ultrahigh energy neutrinos, BEACON is capable of discovering the sources of neutrinos at the highest energies.

astro-ph.HE

Double-edged Role of Interactions in Superconducting Twisted Bilayer Graphene

For the unconventional superconducting phases in moire materials, a critical question is the role played by electronic interactions in the formation of Cooper pairs. In twisted bilayer graphene (tBLG), the strength of electronic interactions can be reduced by increasing the twist angle or screening provided by the dielectric medium. In this work, we place tBLG at 3-4 nm above bulk SrTiO3 substrates, which have a large yet tunable dielectric constant. By raising the dielectric constant in situ in a magic angle device, we observe suppression of both the height and the width of the entire superconducting dome, thus demonstrating that, unlike conventional superconductors, the pairing mechanism in tBLG is strongly dependent on electronic interactions. Interestingly, in contrast to the absence of superconductivity in devices on SiO2 with angle>1.3 deg, we observe a superconducting pocket in a large-angle (angle=1.4 deg) tBLG/STO device while the correlated insulating states are absent. These experimental results are in qualitative agreement with a theoretical model in which the pairing mechanism arises from Coulomb interactions that are screened by plasmons, electron-hole pairs, and longitudinal acoustic phonons. Our results highlight the unconventional nature of the superconductivity in tBLG, the double-edged role played by electronic interactions in its formation, as well as their complex interplay with the correlated insulating states.

cond-mat.mes-hall