SearcharxivSearch

arXiv subjects

Hanyu Liu

Publications and source records attributed to Hanyu Liu.

At least 19 recordsLinked to original sources

Record-Breaking Elemental Superconductivity in Tetralayer Kagome Borophene

Superconductivity above the liquid-nitrogen temperature remains rare in two-dimensional elemental crystals, where strong covalent bonding often yields high phonon frequencies but insufficient electron-phonon coupling. Here, using first-principles calculations and fully anisotropic Migdal-Eliashberg theory, we predict tetralayer kagome borophene (TKB) stabilized by ABAB covalent stacking, as a liquid-nitrogen-temperature elemental superconductor. With a predicted critical temperature of 102 K, TKB sets a record-high value among previously reported elemental superconductors. Unlike known high-Tc boron-based superconductors dominated by in-plane sigma-bonding states and high-frequency in-plane B-B stretching modes, TKB realizes an out-of-plane s-pz-bonding-mediated pairing mechanism, in which interlayer s-pz bonding states at the Fermi level are strongly coupled to low-frequency out-of-plane vibrations of boron atoms. These results reveal a distinct out-of-plane pairing channel in multilayer borophene and establish covalent stacking engineering as a potential route for high-Tc superconductivity in two-dimensional materials.

cond-mat.supr-con

SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models

Long-horizon manipulation is partially observable: the information needed to choose the next action may appear only in observations from minutes earlier. Existing memory mechanisms: retrieval banks, learned compressors, recurrent states must decide what to keep from the past before knowing what a future decision will require. This was motivated by the assumption that minute-scale history is too large to process directly, which modern VLM backbones no longer make true. In this work, we introduce SimpleMemVLA, a VLA without a dedicated memory module. It keeps the sampled history intact and passes it to the backbone in the timestamped video format the backbone was pretrained to process; the hidden states of a generated sub-task then form the only channel from history to a standard flow-matching action head. Since consecutive decisions share most of their history, prefilling the shared prefix during action execution keeps latency close to a single-frame VLA. SimpleMemVLA sets a new state of the art on four memory benchmarks without cost on general-purpose control. Holding the backbone and training setup fixed, it outperforms retrieval, compression and recurrent-state mechanisms by a wide margin, and causal interventions confirm that the policy genuinely reads its history. Code available at https://github.com/wadeKeith/SimpleMemVLA

cs.CV

Understanding and Stabilizing Deep Q-Learning via Controlled Bootstrapping and Regulated Value Dynamics

Deep Q-learning (DQL) has achieved remarkable empirical success in reinforcement learning, yet its training process remains notoriously unstable. Existing studies often attribute instability to isolated factors such as overestimation bias or representation learning issues, lacking a unified understanding of how different sources of instability interact during recursive value estimation. In this work, we provide a systematic analysis of instability in deep Q-learning from three complementary perspectives: operator-level bias in Bellman bootstrapping, estimator-level sensitivity of greedy action selection to regression noise, and parameter-dynamics imbalance under aggressive data reuse. We identify a reward-triggered self-reinforcing trap and characteristic parameter spike dynamics, then derive stabilization principles for controlled bootstrapping, ensemble quantile estimation, and spike-based parameter regulation. Experiments on Atari-100K and Procgen demonstrate competitive performance and improved training stability.

cs.LG

ALKEMIE Agent: an autonomous platform for computational materials design

Despite the powerful multi-scale modeling methods and high-throughput infrastructures established in the materials community, real material computation workflows remain fragmented and heavily manual, requiring researchers to constantly bridge software tools, data analysis, and intermediate decisions. This growing gap between methodological capability and practical execution highlights the need for a new kind of autonomous computational framework, one that can coordinate tools, knowledge, and workflows in a more unified and adaptive way. Here, we introduce ALKEMIE Agent, an agentic platform in which retrieval-augmented generation, a materials-computation knowledge base, registered skills, database-supported provenance, AI-assisted structure modeling, bounded task execution, tool-calling iteration, and error-diagnostic assistance are integrated within a traceable control loop. The capabilities of ALKEMIE Agent are demonstrated through applications including materials recommendation, structure modeling, phonon calculations, machine-learned interatomic potential training, LAMMPS simulations, Ab Initio Monte Carlo (AIMC) sampling, and active-learning-based materials screening. Finally, we outline the future directions and challenges for the development of agentic platforms for computational materials design.

cond-mat.mtrl-sci

AI2Pot: A scalable and unified framework for machine-learning interatomic potential development and large-scale molecular dynamic simulations

Machine-learning interatomic potentials (MLIPs) bridge the accuracy of first-principles calculations and the efficiency required for large-scale molecular dynamics (MD) simulations. However, existing MLIP software remains fragmented across different model architectures, making it difficult to establish unified workflows that support flexible model development, efficient training, and scalable MD deployment. Here, we present AI2Pot, a scalable and unified MLIP framework that seamlessly integrates model training, evaluation, and large-scale MD simulations with PyTorch-compatible ecosystem. Instead of relying on generic automatic differentiation for expensive atomistic operators, AI2Pot re-engineers the core computations of Moment tensor potential (MTP) and Neuroevolution potential (NEP) for both training and inference using hand-crafted C++/CUDA code. These specialized operators constitute a unified computational backend shared by training and inference, improving training-inference consistency and reducing memory usage by avoiding large intermediate caches. As a result, AI2Pot enables fast inference for large-scale atomic systems containing millions of atoms on a single GPU, while retaining the flexibility of PyTorch for model construction, training, and evaluation. Trained models can be deployed in ASE and LAMMPS for MD simulations. Furthermore, AI2Pot provides a companion command-line toolkit (AI2Pot-cli) and Python APIs to facilitate practical MLIP workflows. By unifying high-performance atomistic computing with modern machine-learning ecosystems, AI2Pot offers an user-friendly end-to-end framework for the developing, training, and deploying MLIPs for large scale MD.

cond-mat.mtrl-sci

Hidden ordered compound-layer and its tailoring of the electronic/optical property in Ge2Sb2SexTe5-x alloys

Ge2Sb2SexTe5-x (GSST) alloys represent an emerging class of phase-change materials for integrated photonics. However, the microscopic origins underlying their superior performance compared to the parent compound Ge2Sb2Te5 remain elusive. By using atomic simulations, this work elucidates that the thermal stability and low optical loss of GSST are fundamentally governed by the formation of an in-layer compound-like structure with SeTe2 or Se2Te stoichiometry depending on the Se content, contrasting to the previously believed pure-element-layered model where Se and Te atoms occupy separate layers inside GSST. The newly identified compound-layered structures maintaining stability at temperature above 370 K, yield an enlarged bandgap, weakened antibonding character, and more importantly, a moderate refractive index as well as decreased extinction coefficient which align better with the experiment compared to the previously believed model. The present findings not only help bridge the long-standing theory-experiment gap regarding the optical properties of GSST by redefining its atomic structure, but also establish local chemical ordering as a critical materials design principle for high-performance photonics.

cond-mat.mtrl-sci

HFORD: Hybrid Forward Optimization and Reverse Design Method and Its Applications to On-Chip Millimeter-Wave Inductive Elements

On-chip inductive elements are pivotal in determining both the silicon footprint and performance of millimeter-wave (mmWave) integrated circuits. However, the layout-level synthesis of these passive devices is severely challenged by highly nonlinear geometry-to-performance mappings, computationally expensive full-wave electromagnetic simulations, topology-dependent design spaces, and the inherent non-uniqueness of inverse design. To overcome these bottlenecks, we propose a hybrid forward optimization and reverse design (HFORD) method for the target-to-layout synthesis of mmWave inductive elements. Utilizing a unified core to map device-level requirements to layout-level seeds, HFORD structures direct device targets and translates circuit specifications into a hierarchical synthesis flow. Specifically, sparse-fitting sampling is introduced to improve coverage across critical performance regions, while compact response-fitting coefficients significantly reduce training dimensionality. The HFORD core integrates a random forest for topology selection, a variational autoencoder for spectral feature generation, a mixture density network for probabilistic inverse mapping, and particle swarm optimization for latent space exploration. This integration improves the feasibility of the generated layout seeds under design rule check (DRC) constraints. Two design examples demonstrate that the proposed method accelerates the design cycle from hours to minutes compared to conventional optimization methods.

cs.CE

Beyond Autoregressive RTG: Conditioning via Injection Outside Sequential Modeling in Decision Transformer

Decision Transformer (DT) formulates offline reinforcement learning as autoregressive sequence modeling, achieving promising results by predicting actions from a sequence of Return-to-Go (RTG), state, and action tokens. However, RTG is a scalar that summarizes future rewards, containing far less information than typical state or action vectors, yet it consumes the same computational budget per token. Worse, the self-attention cost of Transformers grows quadratically with sequence length, so including RTG as a separate token adds unnecessary overhead. We propose SlimDT, which removes RTG from the autoregressive sequence. Instead, we inject RTG information into the state representations before the sequential modeling step, allowing the Transformer to process only a compact (state, action) sequence. This reduces the sequence length by one-third, directly improving inference efficiency. On the D4RL benchmark, SlimDT surpasses standard DT across various tasks and achieves performance comparable to existing state-of-the-art methods. Decoupling a sparse conditioning signal from an information-rich sequence thus yields both computational gains and higher task performance.

cs.LG

Liberating LLM Capabilities in Full-Duplex Speech Models

Speech-based large language models are typically constrained to spoken replies, which limits their user-facing outputs to what can be verbalized and suppresses text-native capabilities such as code generation, structured analysis, and multi-step reasoning in realtime interaction, for tasks that require persistent, structured, and inspectable intermediate outputs. Existing work improves spoken reasoning or full-duplex turn-taking, but still treats text as a hidden intermediate state or a subordinate modality rather than a first-class output channel. We propose Listen-Write-Speak (LWS), a text-first tri-channel paradigm in which a single autoregressive LLM continuously listens to user audio, writes visible free-form text as its primary output, and speaks a realtime oral response in parallel under a shared causal attention context. This behavior is implemented entirely through a Token Schema, requiring no architectural modifications, and learned via a two-stage data pipeline that synthesizes per-second cognitive annotations consistent with the revealed input timeline. Empirically, LWS demonstrates strong full-duplex interaction on Full-Duplex-Bench, reaches 4.72 on VoiceBench AlpacaEval, achieves 92.6% writing-speaking consistency, and consistently outperforms its internal ablations on URO-Bench. These results suggest that visible writing can serve as a first-class output channel for speech interaction without sacrificing realtime responsiveness. The code and dataset are available on the project page: https://royalzhang.com/project/lws-page/.

cs.CL

MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

Recent progress in multimodal large language models (MLLMs) has brought AI capabilities from static offline data processing to real-time streaming interaction, yet they still remain far from human-level multimodal interaction. The key bottlenecks are no longer modality coverage or latency alone, but the interaction paradigm itself. First, perception and response are still separated into alternating phases, preventing models from incorporating new inputs for timely adjustment during generation. Second, most current models remain reactive, responding only to explicit user requests instead of acting proactively in the evolving multimodal environment. We present MiniCPM-o 4.5, our latest effort towards human-like multimodal interaction, which mitigates these gaps by real-time full-duplex omni-modal interaction. It can see, listen, and speak simultaneously in real-time, while also exhibiting proactive behaviors such as issuing reminders or comments based on its continuous understanding of the live scene. The key technique behind MiniCPM-o 4.5 is Omni-Flow, a unified streaming framework that aligns omni-modal inputs and outputs along a shared temporal axis. This formulation converts conventional turn-based interaction into a full-duplex, time-aligned process, enabling simultaneous perception and response and allowing proactive behavior to arise within the same framework. With a total of 9B parameters, MiniCPM-o 4.5 approaches Gemini 2.5 Flash in vision-language capabilities, delivering state-of-the-art open-source performance at its scale. It also surpasses Qwen3-Omni-30B-A3B in omni-modal understanding and delivers better speech generation, with significantly higher computation efficiency. Driven by its efficient architecture design and inference optimization, the model can perform real-time full-duplex omni-modal interaction on edge devices with less than 12GB RAM cost.

cs.CL

Crystal structure prediction with nuclear quantum and finite-temperature effects via deep free energy learning

Accurate crystal structure prediction (CSP) requires accounting for finite-temperature and nuclear quantum effects, yet first-principles evaluation of the free energy surface (FES) remains prohibitive for high-throughput searches. We observe that the self-consistent harmonic approximation (SCHA) FES, as a function of nuclear centroid positions, shares the same mathematical structure as a potential-energy surface and can therefore be directly learned by a deep neural network potential. The resulting deep free energy (DF) model, constructed via a two-level concurrent-learning workflow, evaluates free energies, forces, and stresses in a single forward pass. Applied to the La-Sc-H system at 200 GPa and 300 K, DF-based CSP reproduces the stability of the experimentally observed LaH10 and LaSc2H24, and discovers an unreported thermodynamically stable clathrate hydride: P4/mmm LaScH8. Benchmarked on the LaH10 system, the DF model achieves a 1.72*10^6-fold cost reduction relative to DFT-level SSCHA. The DF framework provides a scalable route for incorporating finite-temperature and nuclear quantum effects into high-throughput crystal structure prediction.

cond-mat.mtrl-sci

Interface-dependent Phase Transitions and Ultrafast Hydrogen Superionic Diffusion of H2O Ice

High-pressure experiments using diamond anvils have revealed novel properties and phase behavior of H2O under extreme conditions. When contained in diamond-anvil cells, the H2O samples are usually in direct contact with the diamond anvil. However, the extent to which this interface affects measured pressure-induced properties and behavior, including coexistence lines of ice phases, remains unknown. Combining artificial neural network methods and active learning schemes with large-scale molecular dynamics simulations, we elucidate the interfacial effects on various properties of high-pressure ice phases, including superionic states, solid-solid phase transitions, and melting. The results reveal that the presence of this interface can significantly lower the hydrogen superionic transition temperature. Remarkably, the interface can also induce a spontaneous transition from bcc- to fcc-based ice following the inverse Bain mechanism. Further, we redefined a stability field of bcc and fcc ice below the melting line and predicted the existence of fcc ice at much lower pressures than previously thought. More broadly, the results emphasize the importance of interface effects in understanding a wide range of phenomena reported in experimental studies of ice under pressure, including inconsistencies between theoretical and experimental results of this fundamental system.

cond-mat.mtrl-sci

Decoupling Return-to-Go for Efficient Decision Transformer

The Decision Transformer (DT) has established a powerful sequence modeling approach to offline reinforcement learning. It conditions its action predictions on Return-to-Go (RTG), using it both to distinguish trajectory quality during training and to guide action generation at inference. In this work, we identify a critical redundancy in this design: feeding the entire sequence of RTGs into the Transformer is theoretically unnecessary, as only the most recent RTG affects action prediction. We show that this redundancy can impair DT's performance through experiments. To resolve this, we propose the Decoupled DT (DDT). DDT simplifies the architecture by processing only observation and action sequences through the Transformer, using the latest RTG to guide the action prediction. This streamlined approach not only improves performance but also reduces computational cost. Our experiments show that DDT significantly outperforms DT and establishes competitive performance against state-of-the-art DT variants across multiple offline RL tasks.

cs.AI

Isotropic Superconductivity in Room-temperature Superconductor LaSc$_{2}$H$_{24}$

The discovery of LaSc$_{2}$H$_{24}$ represents a milestone in the quest for room-temperature superconductivity, yet the microscopic mechanism underlying its superior performance remains unclear. Through a comprehensive revisit of theoretical calculations, we uncover a pivotal transition from the anisotropic two-gap superconductivity of LaH$_{10}$ to the isotropic single-gap superconductivity in LaSc$_{2}$H$_{24}$ upon the introduction of scandium, thereby enhancing the superconducting critical temperature ($T_\mathrm{c}$). This enhancement is rooted in a critical dual role of Sc $3d$ electrons: i) the Sc-derived Jahn-Teller effect promotes hydrogen metallization via the elongation of specific interlayer H-H bonds and enhances electron-phonon coupling (EPC) through the softening of associated phonon modes; ii) Sc $3d$ electrons reconstruct the electronic structure into an MgB$_{2}$-like configuration, generating novel Sc-H-Sc $\sigma$- and $\pi$-bonding states with EPC strengths comparable to LaH$_{10}$. Crucially, the pronounced hybridization between Sc and the hydrogen cages effectively unifies these two contributions on the Fermi surface. This Sc-induced gap unification bridges the high-EPC H-H states with widespread Sc-H states, establishing an isotropic single-gap nature with a large overall EPC strength. Our findings identify this Sc-induced gap unification as the fundamental mechanism for achieving room-temperature superconductivity in LaSc$_{2}$H$_{24}$, offering a theoretical blueprint for the future design of superior superconducting hydrides.

cond-mat.supr-con

FOD-Diff: 3D Multi-Channel Patch Diffusion Model for Fiber Orientation Distribution

Diffusion MRI (dMRI) is a critical non-invasive technique to estimate fiber orientation distribution (FOD) for characterizing white matter integrity. Estimating FOD from single-shell low angular resolution dMRI (LAR-FOD) is limited by accuracy, whereas estimating FOD from multi-shell high angular resolution dMRI (HAR-FOD) requires a long scanning time, which limits its applicability. Diffusion models have shown promise in estimating HAR-FOD based on LAR-FOD. However, using diffusion models to efficiently generate HAR-FOD is challenging due to the large number of spherical harmonic (SH) coefficients in FOD. Here, we propose a 3D multi-channel patch diffusion model to predict HAR-FOD from LAR-FOD. We design the FOD-patch adapter by introducing the prior brain anatomy for more efficient patch-based learning. Furthermore, we introduce a voxel-level conditional coordinating module to enhance the global understanding of the model. We design the SH attention module to effectively learn the complex correlations of the SH coefficients. Our experimental results show that our method achieves the best performance in HAR-FOD prediction and outperforms other state-of-the-art methods.

cs.CV

Motus: A Unified Latent Action World Model

While a general embodied agent must function as a unified system, current methods are built on isolated models for understanding, world modeling, and control. This fragmentation prevents unifying multimodal generative capabilities and hinders learning from large-scale, heterogeneous data. In this paper, we propose Motus, a unified latent action world model that leverages existing general pretrained models and rich, sharable motion information. Motus introduces a Mixture-of-Transformer (MoT) architecture to integrate three experts (i.e., understanding, video generation, and action) and adopts a UniDiffuser-style scheduler to enable flexible switching between different modeling modes (i.e., world models, vision-language-action models, inverse dynamics models, video generation models, and video-action joint prediction models). Motus further leverages the optical flow to learn latent actions and adopts a recipe with three-phase training pipeline and six-layer data pyramid, thereby extracting pixel-level "delta action" and enabling large-scale action pretraining. Experiments show that Motus achieves superior performance against state-of-the-art methods in both simulation (a +15% improvement over X-VLA and a +45% improvement over Pi0.5) and real-world scenarios(improved by +11~48%), demonstrating unified modeling of all functionalities and priors significantly benefits downstream robotic tasks.

cs.CV

Revisiting Phase Stability and Superconductivity in Ca-H Superhydrides with Anharmonic Effects

The prediction of superconductivity above 200 K in CaH$_6$ revolutionized research on hydrogen-rich superconductors, and subsequent experiments have verified this prediction, while unidentified peaks in XRD and the decrease in superconducting temperature upon decompression indicate that unresolved issues remain. In this work, we reconstructed the accurate temperature-pressure phase diagram of the Ca-H system and determined the stability ranges of its candidate superconducting phases by considering anharmonic effects. Our results demonstrate that type-I clathrate Ca$_8$H$_{46-delta}$ structures become thermodynamically stable at 0 K when anharmonic effects are considered. Notably, we found that the previously predicted CaH$_6$ phase achieves stability above 500 K, underscoring the significant role of temperature and anharmonic effects in stabilizing this intriguing high-pressure phase. Our findings offer insights into the structure and superconducting mechanisms of hydrides.

cond-mat.supr-con

High-Tc superconductivity above 130 K in cubic MH4 compounds at ambient pressure

Hydrides have long been considered promising candidates for achieving room-temperature superconductivity; however, the extremely high pressures typically required for high critical temperatures remain a major challenge in experiment. Here, we propose a class of high-Tc ambient-pressure superconductors with MH4 stoichiometry. These hydrogen-based compounds adopt the bcc PtHg4 structure type, in which hydrogen atoms occupy the one-quarter body-diagonal sites of metal lattices, with the metal atoms acting as chemical templates for hydrogen assembly. Through comprehensive first-principles calculations, we identify three promising superconductors, PtH4, AuH4 and PdH4, with superconducting critical temperatures of 84 K, 89 K, and 133 K, respectively, all surpassing the liquid-nitrogen temperature threshold of 77 K. The remarkable superconducting properties originate from strong electron-phonon coupling associated with hydrogen vibrations, which in turn arise from phonon softening in the mid-frequency range. Our results provide crucial insights into the design of high-Tc superconductors suitable for future experiments and applications at ambient pressure.

cond-mat.supr-con