SearcharxivSearch

arXiv subjects

Zhuohan Li

Publications and source records attributed to Zhuohan Li.

At least 19 recordsLinked to original sources

Redox-Active Halide Materials for Cathode Applications

Electrochemically redox-active halide (eREAL) materials are an emerging class of materials that combine high Li-ion conductivity with transition-metal redox activity, making them promising candidates for cathode or catholyte applications. As a redox-active catholyte, they could significantly increase the energy density of solid-state batteries. In this work, we perform first-principles calculations on Li-M-Cl (M = 3d transition metals) ternaries to establish such a theoretical foundation for their stability and electrochemical activity. We map the phase stability of eREAL structures with varying metal-to-Cl ratio, transition-metal species, oxidation states, and anion frameworks, and compute cation and anion redox potentials. We find that the high ionicity of metal-Cl bonds elevates cation redox potentials above those of conventional oxide cathodes, but also will promote Cl oxidation and Cl-Cl dimerization at high voltages, which may limit the stability of these materials. Anion substitution effectively tunes both cation and anion redox potentials, with F substitution standing out as a viable route to extend the reversible voltage window. Beyond the anion redox issue, eREAL compounds generally exhibit flat voltage profiles, which potentially poses an electrochemical compatibility challenge when paired with active materials that operate at different voltage values or over wider voltage ranges. Collectively, our study provides a comprehensive analysis for redox behavior of eREAL materials, paving the way for their rational design and optimization in next-generation battery applications.

cond-mat.mtrl-sci

The multiple corrugations in the Galactic disk derived from the LAMOST and Gaia survey data

Large spectroscopic and astrometric surveys have revealed complex wave-like features in the Milky Way disk, suggesting that its kinematic and chemical structures are shaped by time-dependent perturbations. Recent studies have reported oscillatory patterns in the Rg-Vphi-VR space, hinting at a possible structural transition in the outer disk. We aim to characterise the transition between the inner and outer Galactic thin disk and to investigate whether radial corrugations can provide a plausible physical interpretation of the observed features. We analysed two large stellar samples from LAMOST DR8 and Gaia DR3, combining spatial, kinematic, and chemical diagnostics. A simplified corrugation model consisting of two radial waves propagating in opposite directions was constructed and fitted to the observed VR pattern. We further validated the model using N-body simulations. Both LAMOST and Gaia samples reproduce the previously reported wave-like pattern in the Rg-Vphi-VR plane. We identify a clear transition between the inner and outer disks via the variations in rotational velocity and metallicities. The corrugation model naturally reproduces the periodic variation of VR with galactocentric radius, and the superposition of the inward and outward propagating modes gives rise to a comparable oscillatory pattern in both observations and simulations. Our modelling suggests that radial corrugations can provide a plausible interpretation of the observed kinematic signatures. The results highlight the complex, multi-perturber nature of the Galactic disk and motivate further investigation with upcoming surveys.

astro-ph.GA

Galactic Archaeology with the Subaru `\=Onohi`ula Prime Focus Spectrograph Strategic Program

The recently commissioned Subaru `\=Onohi`ula Prime Focus Spectrograph (PFS) will obtain spectra from nearly 2,400 fibers that cover 1.24 square degrees. The 360 night Subaru Strategic Program for PFS is dedicating approximately one-third of its allocation (130 nights) to study the structure and evolution of galaxies in the Local Group. This Galactic Archaeological survey has three pillars. (1) We will determine whether the mass density profiles of dwarf galaxies are consistent with cusps, as expected for cold dark matter, or cores, as expected from alternative dark matter theories or baryonic feedback. We will deduce the density profiles as a function of radius from modeling of the full line-of-sight velocity and abundance distributions for six dwarf galaxies. Our total sample will consist of 18,000 member stars to beyond the nominal tidal radius of each system. (2) From measurements of the [alpha/Fe] abundance ratio, we will learn the difference in assembly history of the two most massive galaxies in the Local Group: M31 and the Milky Way. We will observe 30,000 member stars over 45 square degrees of M31's halo and outer disk. (3) We will uncover how the most fragile (outer) part of the Milky Way responded to accretion events both in the distant past (such as Gaia-Sausage Enceladus) and in more recent history (such as the Sagittarius dwarf spheroidal galaxy). To support this study, PFS will provide velocities and metallicities--from which, in combination with photometry, we will deduce ages--for tens of thousands of main-sequence stars out to a Galactocentric distance of ~30 kpc.

astro-ph.GA

CASCADE: Cumulative Agentic Skill Creation through Autonomous Development and Evolution

Large language model (LLM) agents currently depend on predefined tools or early-stage tool generation, limiting their adaptability and scalability to complex scientific tasks. We introduce CASCADE, a self-evolving agentic framework representing an early instantiation of the transition from "LLM + tool use" to "LLM + skill acquisition". CASCADE enables agents to master complex external tools and codify knowledge through two meta-skills: continuous learning via web search, code extraction, and memory utilization; self-reflection via introspection, knowledge graph exploration, and others. We evaluate CASCADE on SciSkillBench, a benchmark of 116 materials science and chemistry research tasks. CASCADE achieves a 93.3% success rate using GPT-5, compared to 35.4% without evolution mechanisms. We further demonstrate real-world applications in computational analysis, autonomous laboratory experiments, and selective reproduction of published papers. Along with human-agent collaboration and memory consolidation, CASCADE accumulates executable skills that can be shared across agents and scientists, moving toward scalable AI-assisted scientific research.

cs.AI

Hierarchical high-throughput screening of alkaline-stable lithium-ion conductors combining machine learning and first-principles calculations

Solid-state batteries require lithium-ion conductors that combine high ionic conductivity with stability under harsh electrochemical and chemical conditions. Here, we investigate the chemical factors governing the stability of NASICON-type and garnet-type Li-ion conductors in highly alkaline environments. This is particularly relevant to solid-state Li-air cells operated under humidified air where alkaline conditions arise due to the formation of LiOH discharge products. We implement a hierarchical high-throughput screening workflow that consists of a pre-screening step using a universal machine-learning interatomic potential and a more accurate DFT-based screening. This approach enables rapid evaluation of over 320,000 compositions, from which 209 alkaline-stable candidates are identified. We identify specific cation substitutions that improve alkaline stability in NASICON and garnet compounds and reveal the underlying mechanism. More importantly, we highlight design trade-offs that require careful composition optimization to simultaneously enhance synthesizability, operational stability, and Li-ion/electronic conductivities for practical humid Li-air battery applications.

cond-mat.mtrl-sci

Li+/H+ exchange in solid-state oxide Li-ion conductors

Understanding the moisture stability of oxide Li-ion conductors is important for their practical applications in solid-state batteries. Unlike sulfide or halide conductors, oxide conductors generally better resist degradation when in contact with water, but can still undergo topotactic \ch{Li+}/\ch{H+} exchange (LHX). Here, we combine density functional theory (DFT) calculations with a machine-learning interatomic potential model to investigate the thermodynamic driving force of the LHX reaction for two representative oxide Li-ion conductor families: garnets and NASICONs. Li-stuffed garnets exhibit a strong driving force for proton exchange due to their high Li chemical potential. In contrast, NASICONs demonstrate a higher resistance against proton exchange due to the lower Li chemical potential and the lower O-H bond covalency for polyanion-bonded oxygens. Our findings reveal a critical trade-off: Li stuffing enhances conductivity but increases moisture susceptibility. This study underscores the importance of designing Li-ion conductors that possess both high conductivity and high stability in practical environments.

cond-mat.mtrl-sci

Investigating the Degradation of LATP Solid Electrolyte in High Alkaline Li-$O_2$ Batteries

In this study, we address the challenge of electrolyte degradation in all-solid-state humidified Li-$O_2$ batteries, which offer high theoretical energy density and potential cost advantages over conventional lithium-ion batteries. Combining STEM-EELS, and XPS characterizations with DFT calculations, we reveal the leaching of $(PO_4)^{3+}$ and $Al^{3+}$ ions from the $Li_{1.3}Al_{0.3}Ti_{1.7}(PO_4)_3$ (LATP) solid electrolyte upon battery discharge, caused by the highly alkaline environment. Upon charging, the leached ions precipitate as $Li_3PO_4$ and $AlPO_4$, which accumulate on the LATP surface and contribute to battery degradation. A Ti-rich layer is observed at the surface after a few cycles due to depletion of other cations. Our findings suggest that the degradation products are formed through repeated dissolution and precipitation in the discharge-charge cycles. Furthermore, our results indicate that the Ti-rich layer on the LATP surface can potentially reduce parasitic reactions. Our study provides mechanistic understanding of LATP solid electrolyte degradation in humidified Li-$O_2$ cell, paving the way for designing more durable and efficient Li-$O_2$ batteries.

cond-mat.mtrl-sci

gpt-oss-120b & gpt-oss-20b Model Card

We present gpt-oss-120b and gpt-oss-20b, two open-weight reasoning models that push the frontier of accuracy and inference cost. The models use an efficient mixture-of-expert transformer architecture and are trained using large-scale distillation and reinforcement learning. We optimize the models to have strong agentic capabilities (deep research browsing, python tool use, and support for developer-provided functions), all while using a rendered chat format that enables clear instruction following and role delineation. Both models achieve strong results on benchmarks ranging from mathematics, coding, and safety. We release the model weights, inference implementations, tool environments, and tokenizers under an Apache 2.0 license to enable broad use and further research.

cs.CL

The Rotating Bulge and Halo in the Milky Way: Evidence of Angular Momentum Transferred from the Decelerating Bar

Recent observations indicate that both the Milky Way bulge and inner halo exhibit angular momentum, although the origin and evolution of this prograde signature remain ambiguous. One plausible scenario involves secular evolution induced by the central bar and spiral arms. In this study, we identified a component consisting of 1,175,737 stars with net rotation through the application of a neural network (NN) method. To investigate the composition of this rotating sample and the origin of its rotation, we conducted a test particle simulation incorporating an equilibrium axisymmetric background potential together with a central decelerating bar. The test particles were generated using a distribution function (DF) model derived from observational constraints. Our results indicate that the decelerating bar transfers angular momentum to the pseudo-stars, and the rotational profile from our simulation shows strong agreement with observational data. These findings suggest that the rotating sample identified by our NN model predominantly comprises bulge, halo, and thick disk stars, and that the central decelerating bar is pivotal in shaping the inner Galaxy's kinematics through angular momentum transfer.

astro-ph.GA

An Actinide-boost Star Discovered in the Gaia-Sausage-Enceladus

We report the discovery of an actinide-boost, very metal-poor ($\left[\mathrm{Fe/H} \right]=-2.38$), $r$-process-enhanced ($\left[\mathrm{Eu/Fe} \right]=0.80$) star, LAMOST J0804+5740, within the Gaia-Sausage-Enceladus (GSE). Based on the high-resolution ($R\sim36,000\; and \;60,000$) and high signal-to-noise ratio spectra obtained with the High Dispersion Spectrograph on the Subaru Telescope, the abundances of 48 species are determined. Its $\log\epsilon\rm(\mathrm{Th}/\mathrm{Eu}) = -0.22 $ establishes it as the first confirmed actinide-boost star within the GSE. Comparative analysis of its abundance pattern with theoretical $r$-process models reveals that the magnetorotationally driven jet supernova $r$-process model with $\hat{L}v$ = 0.2 provides the best fit and successfully reproduces the actinide-boost signature. Kinematic analysis of actinide-boost stars reveals that approximately two-thirds of them are classified as $\textit{ex-situ}$ stars, suggesting that actinide-boost stars are more likely to originate from accreted dwarf galaxies. As the first actinide-boost star identified within the GSE, J0804+5740 will provide valuable insights into $r$-process nucleosynthesis in accreted dwarf galaxies like the GSE, especially on the production of the heaviest elements.

astro-ph.SR

Jenga: Effective Memory Management for Serving LLM with Heterogeneity

Large language models (LLMs) are widely used but expensive to run, especially as inference workloads grow. To lower costs, maximizing the request batch size by managing GPU memory efficiently is crucial. While PagedAttention has recently been proposed to improve the efficiency of memory management, we find that the growing heterogeneity in the embeddings dimensions, attention, and access patterns of modern LLM architectures introduces new challenges for memory allocation. In this paper, we present Jenga, a novel memory allocation framework for heterogeneous embeddings in LLMs. Jenga tackles two key challenges: (1) minimizing memory fragmentation when managing embeddings of different sizes, and (2) enabling flexible caching and eviction policies tailored to the specific token-dependency patterns of various layers. Jenga employs a two-level memory allocator, leveraging the least common multiple (LCM) of embedding sizes to optimize memory usage and providing APIs to express layer-specific caching logic to enhance memory reuse. We implemente Jenga on vLLM, a state-of-the-art LLM inference engine, and evaluate it with diverse LLMs, datasets, and GPU configurations. Evaluations show that Jenga improves GPU memory utilization by up to 79.6%, and increases serving throughput by up to 4.92x (1.80x on average).

cs.DC

Mapping the Milky Way with Gaia Bp/Rp spectra I: Systematic flux corrections and atmospheric parameters for 68 million stars

Gaia Bp/Rp spectra for over two hundred million stars have great potential for mapping metallicity across the Milky Way. We aim to construct an alternative catalog of atmospheric parameters from Gaia Bp/Rp spectra by fitting them with synthetic spectra based on model atmospheres, and provide corrections to the Bp/Rp fluxes according to stellar colors, magnitudes, and extinction. We use GaiaXPy to obtain calibrated spectra and apply FERRE to match the corrected Bp/Rp spectra with models and infer atmospheric parameters. We train a neural network using stars in APOGEE to predict flux corrections as a function of wavelength for each target. Based on the comparison with APOGEE parameters, we conclude that our estimated parameters have systematic errors and uncertainties in $T_{\mathrm{eff}}$, $\log g$, and [M/H] about $-38 \pm 167$ K, $0.05 \pm 0.40$ dex, and $-0.12 \pm 0.19$ dex, respectively, for stars in the range $4000 \le T_{\mathrm{eff}} \le 7000$ K. The corrected Bp/Rp spectra show better agreement with both models and Hubble Space Telescope CALSPEC data. Our correction increases the precision of the relative spectrophotometry of the Bp/Rp data from $3.2\% - 3.7\%$ to $1.2\% - 2.4\%$. Finally, we have built a catalog of atmospheric parameters for stars within $4000 \le T_{\mathrm{eff}} \le 7000$ K, comprising $68,394,431$ sources, along with a subset of $124,188$ stars with $\mathrm{[M/H]} \le -2.5$. Our results confirm that the Gaia Bp/Rp flux calibrated spectra show systematic patterns as a function of wavelength that are tightly related to colors, magnitudes, and extinction. Our optimization algorithm can give us accurate atmospheric parameters of stars with a clear and direct link to models of stellar atmospheres, and can be used to efficiently search for extremely metal-poor stars.

astro-ph.GA

Hierarchical Absorption in LAMOST low Resolution Normalized Spectra

According to the hierarchical clustering scenario, galaxies like the Milky Way form hierarchically, and many supporting evidences have been found in the Galactic halo. However, most stars in the Milky Way are disk stars. Disk stars have almost lost their spatial distribution and kinematic features at birth, retaining solely chemical signatures. Identifying such substructures using abundances of iron or light elements is difficult due to their high degeneracy. Heavy elements, especially neutron capture elements have limited sources so have lower degeneracy, but spectral line fitting of these elements is tough, requiring mid to high resolution spectra, which are currently limited in sample size. This work utilizes the collective effect of many spectral lines from several elements, especially neutron capture elements to weaken the degeneracy of [Fe/H]. The analysis suggests the presence of at least 12 clusters, due to the hierarchical absorption in LAMOST low resolution normalized Spectra. More detailed work is needed to ascertain whether this hierarchical absorption is natural.

astro-ph.GA

TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput

Large Language Model (LLM) serving systems batch concurrent user requests to achieve efficient serving. However, in real-world deployments, such inter-request parallelism from batching is often limited by external factors such as low request rates or memory constraints. Recent works focus on intra-request parallelism from speculative decoding as a solution to this problem. Unfortunately, benefits from intra-request parallelism are often fragile, as speculative decoding causes overhead, and speculated tokens may miss. We observe that speculative decoding may degrade LLM serving performance if added naively without tuning to the incoming requests and the speculation method. To alleviate the need for expert tuning and make speculative decoding more robust, we present TurboSpec, a speculation control system that automatically profiles the execution environment and utilizes a feedback-based algorithm to dynamically adjust the amount of intra-request parallelism in LLM serving. TurboSpec predicts "goodput" - the amount of successfully generated tokens - to evaluate and adjust intra-request parallelism amount to that with the highest goodput in runtime. We implement TurboSpec on a real-world LLM serving system vLLM and demonstrate its effectiveness across diverse workloads and hardware configurations, providing consistent performance improvements across all test scenarios.

cs.AI

Fairness in Serving Large Language Models

High-demand LLM inference services (e.g., ChatGPT and BARD) support a wide range of requests from short chat conversations to long document reading. To ensure that all client requests are processed fairly, most major LLM inference services have request rate limits, to ensure that no client can dominate the request queue. However, this rudimentary notion of fairness also results in under-utilization of the resources and poor client experience when there is spare capacity. While there is a rich literature on fair scheduling, serving LLMs presents new challenges due to their unpredictable request lengths and their unique batching characteristics on parallel accelerators. This paper introduces the definition of LLM serving fairness based on a cost function that accounts for the number of input and output tokens processed. To achieve fairness in serving, we propose a novel scheduling algorithm, the Virtual Token Counter (VTC), a fair scheduler based on the continuous batching mechanism. We prove a 2x tight upper bound on the service difference between two backlogged clients, adhering to the requirement of work-conserving. Through extensive experiments, we demonstrate the superior performance of VTC in ensuring fairness, especially in contrast to other baseline methods, which exhibit shortcomings under various conditions. The reproducible code is available at https://github.com/Ying1123/VTC-artifact

cs.AI

Overcoming systematic softening in universal machine learning interatomic potentials by fine-tuning

Machine learning interatomic potentials (MLIPs) have introduced a new paradigm for atomic simulations. Recent advancements have seen the emergence of universal MLIPs (uMLIPs) that are pre-trained on diverse materials datasets, providing opportunities for both ready-to-use universal force fields and robust foundations for downstream machine learning refinements. However, their performance in extrapolating to out-of-distribution complex atomic environments remains unclear. In this study, we highlight a consistent potential energy surface (PES) softening effect in three uMLIPs: M3GNet, CHGNet, and MACE-MP-0, which is characterized by energy and force under-prediction in a series of atomic-modeling benchmarks including surfaces, defects, solid-solution energetics, phonon vibration modes, ion migration barriers, and general high-energy states. We find that the PES softening behavior originates from a systematic underprediction error of the PES curvature, which derives from the biased sampling of near-equilibrium atomic arrangements in uMLIP pre-training datasets. We demonstrate that the PES softening issue can be effectively rectified by fine-tuning with a single additional data point. Our findings suggest that a considerable fraction of uMLIP errors are highly systematic, and can therefore be efficiently corrected. This result rationalizes the data-efficient fine-tuning performance boost commonly observed with foundational MLIPs. We argue for the importance of a comprehensive materials dataset with improved PES sampling for next-generation foundational MLIPs.

cond-mat.mtrl-sci

LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset

Studying how people interact with large language models (LLMs) in real-world scenarios is increasingly important due to their widespread use in various applications. In this paper, we introduce LMSYS-Chat-1M, a large-scale dataset containing one million real-world conversations with 25 state-of-the-art LLMs. This dataset is collected from 210K unique IP addresses in the wild on our Vicuna demo and Chatbot Arena website. We offer an overview of the dataset's content, including its curation process, basic statistics, and topic distribution, highlighting its diversity, originality, and scale. We demonstrate its versatility through four use cases: developing content moderation models that perform similarly to GPT-4, building a safety benchmark, training instruction-following models that perform similarly to Vicuna, and creating challenging benchmark questions. We believe that this dataset will serve as a valuable resource for understanding and advancing LLM capabilities. The dataset is publicly available at https://huggingface.co/datasets/lmsys/lmsys-chat-1m.

cs.CL

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Evaluating large language model (LLM) based chat assistants is challenging due to their broad capabilities and the inadequacy of existing benchmarks in measuring human preferences. To address this, we explore using strong LLMs as judges to evaluate these models on more open-ended questions. We examine the usage and limitations of LLM-as-a-judge, including position, verbosity, and self-enhancement biases, as well as limited reasoning ability, and propose solutions to mitigate some of them. We then verify the agreement between LLM judges and human preferences by introducing two benchmarks: MT-bench, a multi-turn question set; and Chatbot Arena, a crowdsourced battle platform. Our results reveal that strong LLM judges like GPT-4 can match both controlled and crowdsourced human preferences well, achieving over 80% agreement, the same level of agreement between humans. Hence, LLM-as-a-judge is a scalable and explainable way to approximate human preferences, which are otherwise very expensive to obtain. Additionally, we show our benchmark and traditional benchmarks complement each other by evaluating several variants of LLaMA and Vicuna. The MT-bench questions, 3K expert votes, and 30K conversations with human preferences are publicly available at https://github.com/lm-sys/FastChat/tree/main/fastchat/llm_judge.

cs.CL