SearcharxivSearch

arXiv subjects

Liu

Publications and source records attributed to Liu.

At least 19 recordsLinked to original sources

FinPersona-Bench: A Benchmark for Longitudinal Psychometric Stability of Autonomous Financial Agents

Large Language Models (LLMs) are increasingly deployed as autonomous financial agents initialized with explicit behavioral mandates such as "preserve capital" or "avoid speculative bets" that are meant to govern every decision throughout deployment. In practice, however, as market context accumulates over long horizons, these mandates gradually lose their behavioral influence, a phenomenon we formalize as Mandate Salience Decay (MSD). To measure MSD objectively, we introduce FinPersona-Bench, a simulation benchmark in which a synthetic market decouples observable price from hidden fundamental value, enabling falsifiable evaluation across three failure modes: trading without signal in calm markets, panic-selling during crashes, and ignoring fundamental value during speculative bubbles. Evaluating 18 leading frontier and open-source LLMs, each assigned one of three behavioral profiles ranging from strict capital preservation to aggressive growth, shows that MSD compounds over time and is model-dependent. In crash scenarios, the behavioral gap between static agents and those receiving periodic mandate re-grounding grows 4.4x from the first to the final quarter of the simulation. The effects of mandate re-grounding are not uniformly positive: it consistently helps conservative agents in low-signal markets but actively worsens behavior for aggressive agents in the same setting. These findings suggest that reliable long-horizon deployment requires selective, mandate-aware re-grounding based on agent profile and market regime.

cs.CL

ECTraj: Enhanced Consistency Training for Multi-Agent Trajectory Prediction

Diffusion models for multi-agent trajectory prediction are limited by iterative denoising, which causes inference latency that hinders their use in time-critical settings like autonomous driving. Fast-sampling variants using DDIM and informed initial noise distributions partially alleviate this issue, but they either fail to achieve true single-step generation or are constrained by the chosen noise distribution. Consistency Models (CMs) offer high-quality one-step generation by mapping noise directly to data, but are difficult to train from scratch. We propose ECTraj, an enhanced CM pipeline with improved training and conditional generation for trajectory prediction. Our framework extends the student-teacher consistency training scheme: the student produces standard outputs, while the teacher explicitly fuses its predictions with parts of the ground truth to give stronger supervision. We also exploit CMs' direct denoising for top-K multi-shot generation during training. Combining conditional generation with this enhanced consistency objective yields faster inference and improved prediction accuracy, establishing competitive new benchmarks on the large-scale Argoverse 2 dataset.

cs.CV

FinReporting: An Agentic Workflow for Localized Reporting of Cross-Jurisdiction Financial Disclosures

Financial reporting systems increasingly leverage Large Language Models (LLMs) to extract and summarize corporate disclosures. However, most existing approaches assume a single-market setting and overlook structural differences across jurisdictions. Variations in accounting taxonomies, tagging infrastructures (e.g., XBRL vs.\ PDF), and aggregation conventions introduce substantial challenges for semantic alignment and reliable verification. Here, we aim to bridge this gap. We present FinReporting, an agentic workflow for localized cross-jurisdiction financial reporting. The system constructs a unified canonical ontology spanning the income statement, balance sheet, and cash flow statement, and decomposes reporting into auditable stages, including filing acquisition, extraction, canonical mapping, and anomaly logging. Rather than treating LLMs as free-form generators, FinReporting employs them as constrained verifiers operating under explicit decision rules with evidence grounding. Evaluated on annual filings from the USA, Japan, and China, FinReporting improves consistency and reliability under heterogeneous reporting regimes. We further release an interactive demo that enables cross-market inspection and supports structured export of localized financial statements. Our demo is available at url{https://huggingface.co/spaces/BoomQ/FinReporting-Demo. A video describing our system is available at https://www.youtube.com/watch?v=f65jdEL31Kk.

cs.CL

CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification

Anthropic proposes the concept of skills for LLM agents to tackle multi-step professional tasks that simple tool invocations cannot address. A tool is a single, self-contained function, whereas a skill is a structured bundle of interdependent multi-file artifacts. Currently, skill generation is not only label-intensive due to manual authoring, but also may suffer from human--machine cognitive misalignment, which can lead to degraded agent performance, as evidenced by evaluations on SkillsBench. Therefore, we aim to enable agents to autonomously generate skills. However, existing self-evolving methods designed for tools cannot be directly applied to skills due to their increased complexity. To address these issues, we propose CoEvoSkills, a self-evolving skills framework that enables agents to autonomously construct complex, multi-file skill packages. Specifically, CoEvoSkills couples a Skill Generator that iteratively refines skills with a Surrogate Verifier that co-evolves to provide informative and actionable feedback without access to ground-truth test content. On SkillsBench, CoEvoSkills outperforms five baselines on both Claude Code and Codex, and generalizes strongly to six additional LLMs. The code is publicly available at https://github.com/Zhang-Henry/CoEvoSkills.

cs.AI

GMPilot: An Expert AI Agent For FDA cGMP Compliance

The pharmaceutical industry is facing challenges with quality management such as high costs of compliance, slow responses and disjointed knowledge. This paper presents GMPilot, a domain-specific AI agent that is designed to support FDA cGMP compliance. GMPilot is based on a curated knowledge base of regulations and historical inspection observations and uses Retrieval-Augmented Generation (RAG) and Reasoning-Acting (ReAct) frameworks to provide real-time and traceable decision support to the quality professionals. In a simulated inspection scenario, GMPilot shows how it can improve the responsiveness and professionalism of quality professionals by providing structured knowledge retrieval and verifiable regulatory and case-based support. Although GMPilot lacks in the aspect of regulatory scope and model interpretability, it is a viable avenue of improving quality management decision-making in the pharmaceutical sector using intelligent approaches and an example of specialized application of AI in highly regulated sectors.

cs.AI

vLLM Semantic Router: Signal Driven Decision Routing for Mixture-of-Modality Models

As large language models (LLMs) diversify across modalities, capabilities, and cost profiles, the problem of intelligent request routing: selecting the right model for each query at inference time, has become a critical systems challenge. We present vLLM Semantic Router, a signal-driven decision routing framework for Mixture-of-Modality (MoM) model deployments. The architecture follows two complementary Shannon-inspired views. In the information-theoretic regime, signal extraction reduces the entropy of "which model?" by distilling routing-relevant information from raw queries. In the Boolean-algebraic regime, the decision engine composes functionally complete routing policies from signal conditions. The central innovation is composable signal orchestration: thirteen heterogeneous signal types, spanning sub-millisecond heuristics and neural classifiers for semantics, safety, and modality, are composed through configurable Boolean decision rules into deployment-specific routing policies, so that fundamentally different scenarios (multi-cloud enterprise, privacy-regulated, cost-optimized) are expressed as different configurations over the same architecture. Matched decisions drive semantic model routing via thirteen selection algorithms, while per-decision plugin chains enforce safety constraints including a three-stage HaluGate hallucination detection pipeline and a lightweight episodic memory system with ReflectionGate for personalized multi-turn context. A typed neural-symbolic DSL specifies these routing policies and compiles them to multiple deployment targets, enabling configuration-first adaptation without code changes. Together, these components show that composable signal orchestration enables a single framework to serve diverse deployment scenarios with differentiated cost, privacy, and safety policies.

cs.NI

Early Architecture Concepts for the Habitable Worlds Observatory -- System Design, Modeling, and Analysis

The Habitable Worlds Observatory (HWO), NASA's next flagship science mission, follows in the tradition of the Nancy Grace Roman Space Telescope and other preceding great observatories. HWO will directly image and characterize Earth-like exoplanet and their atmospheres, with the capability to detect biosignatures and potentially answer the question of whether we are we alone. HWO will also serve as a powerful general astrophysics observatory, enabling breakthroughs in galaxy evolution, stellar astrophysics, and dark matter studies. Currently in pre-formulation, the project has established Exploratory Analytic Cases (EACs), a series of architectural concept designs used to assess the mission's demanding science objectives while exploring challenging engineering parameters. This paper describes the first three EACs, starting with observing strategies and error budget formulation and then progressing to design formulations, trade studies and lessons learned; this paper also discusses the integrated modeling pipeline, a key multidisciplinary system-level analysis capability, and analysis findings as applied to the first EAC. These activities set the stage for the follow on EACs 4 and 5, which will further explore the trade space and prepare for the baseline design that will support the Mission Concept Review (MCR).

astro-ph.IM

Stress-Induced Ferroelectricity in Hafnium Oxide Core-Shell Nanoparticles

In contrast to hafnia (HfO2) thin films, where the appearance of switchable ferroelectric polarization can be induced by strain or defect engineering, reliable methods for controlling ferroelectricity are absent in HfO2 nanoparticles. Direct experimental observations of ferroelectric hysteresis and ferroelectric domains in these nanoparticles are also absent. To the best of our knowledge, stress-induced ferroelectric states in the HfO2 nanoparticles have not been explored. In this work, we study the influence of chemical stress on phase diagrams, dielectric and polar properties of spherical HfO2 core-shell nanoparticles using a Landau-Ginzburg-Devonshire free energy functional that includes trilinear and biquadratic couplings involving polar, antipolar, and nonpolar order parameters. The ferroelectric phase exhibits reentrant behavior as a function of nanoparticle size, such that the spontaneous polarization exists only within a limited range of core radii R_c, namely R_cr^min<R_c<R_cr^max. The minimal critical radius R_cr^min is primarily determined by the size dependence of the depolarization field and correlation effects; the maximal critical radius R_cr^max is primarily determined by the size dependence of chemical stresses induced by the elastic defects in the shell. Thus, this work identifies a stress-driven mechanism for reentrant ferroelectricity stabilization in nanoscale HfO2 systems, arising from the competition between depolarization field-induced suppression of ferroelectricity and its stabilization by shell-induced chemical stress. We revealed that relatively large compressive chemical strains are necessary to induce the ferroelectric phase in the HfO2 nanoparticles. Successful chemical strain engineering opens the way for significant enhancement of nanoscale HfO2 polar properties for applications in advanced memory cells and logic devices.

cond-mat.mtrl-sci

LERA: Reinstating Judgment as a Structural Precondition for Execution in Automated Systems

As automated systems increasingly transition from decision support to direct execution, the problem of accountability shifts from decision quality to execution legitimacy. While optimization, execution, and feedback mechanisms are extensively modeled in contemporary AI and control architectures, the structural role of judgment remains undefined. Judgment is typically introduced as an external intervention rather than a native precondition to execution. This work does not propose a new decision-making algorithm or safety heuristic, but identifies a missing structural role in contemporary AI and control architectures. This paper identifies this absence as a missing Judgment Root Node and proposes LERA (Judgment-Governance Architecture) , a structural framework that enforces judgment as a mandatory, non-bypassable prerequisite for execution. LERA is founded on two axioms: (1) execution is not a matter of system capability, but of structural permission, and (2) execution is not the chronological successor of judgment, but its structural consequence. Together, these axioms decouple execution legitimacy from computational capacity and bind it to judgment completion through a governance gate. LERA does not aim to optimize decisions or automate judgment. Instead, it institutionalizes judgment as a first-class architectural component, ensuring that execution authority remains accountable. By reinstating judgment at the execution boundary, LERA establishes a foundational architecture for judgment-governed automation.

cs.CY

Exploring Protein Language Model Architecture-Induced Biases for Antibody Comprehension

Recent advances in protein language models (PLMs) have demonstrated remarkable capabilities in understanding protein sequences. However, the extent to which different model architectures capture antibody-specific biological properties remains unexplored. In this work, we systematically investigate how architectural choices in PLMs influence their ability to comprehend antibody sequence characteristics and functions. We evaluate three state-of-the-art PLMs-AntiBERTa, BioBERT, and ESM2--against a general-purpose language model (GPT-2) baseline on antibody target specificity prediction tasks. Our results demonstrate that while all PLMs achieve high classification accuracy, they exhibit distinct biases in capturing biological features such as V gene usage, somatic hypermutation patterns, and isotype information. Through attention attribution analysis, we show that antibody-specific models like AntiBERTa naturally learn to focus on complementarity-determining regions (CDRs), while general protein models benefit significantly from explicit CDR-focused training strategies. These findings provide insights into the relationship between model architecture and biological feature extraction, offering valuable guidance for future PLM development in computational antibody design.

cs.LG

gpt-oss-120b & gpt-oss-20b Model Card

We present gpt-oss-120b and gpt-oss-20b, two open-weight reasoning models that push the frontier of accuracy and inference cost. The models use an efficient mixture-of-expert transformer architecture and are trained using large-scale distillation and reinforcement learning. We optimize the models to have strong agentic capabilities (deep research browsing, python tool use, and support for developer-provided functions), all while using a rendered chat format that enables clear instruction following and role delineation. Both models achieve strong results on benchmarks ranging from mathematics, coding, and safety. We release the model weights, inference implementations, tool environments, and tokenizers under an Apache 2.0 license to enable broad use and further research.

cs.CL

Developing an AI-Guided Assistant Device for the Deaf and Hearing Impaired

This study aims to develop a deep learning system for an accessibility device for the deaf or hearing impaired. The device will accurately localize and identify sound sources in real time. This study will fill an important gap in current research by leveraging machine learning techniques to target the underprivileged community. The system includes three main components. 1. JerryNet: A custom designed CNN architecture that determines the direction of arrival (DoA) for nine possible directions. 2. Audio Classification: This model is based on fine-tuning the Contrastive Language-Audio Pretraining (CLAP) model to identify the exact sound classes only based on audio. 3. Multimodal integration model: This is an accurate sound localization model that combines audio, visual, and text data to locate the exact sound sources in the images. The part consists of two modules, one object detection using Yolov9 to generate all the bounding boxes of the objects, and an audio visual localization model to identify the optimal bounding box using complete Intersection over Union (CIoU). The hardware consists of a four-microphone rectangular formation and a camera mounted on glasses with a wristband for displaying necessary information like direction. On a custom collected data set, JerryNet achieved a precision of 91. 1% for the sound direction, outperforming all the baseline models. The CLAP model achieved 98.5% and 95% accuracy on custom and AudioSet datasets, respectively. The audio-visual localization model within component 3 yielded a cIoU of 0.892 and an AUC of 0.658, surpassing other similar models. There are many future potentials to this study, paving the way to creating a new generation of accessibility devices.

cs.LG

Behavior of quantum coherence in the ultrastrong and deep strong coupling regimes of light-matter system

The ultrastrong and deep strong coupling regimes exhibit a variety of intriguing physical phenomena. In this work, we utilize the Hopfield model of a two-mode bosonic system, with each mode interacts with a heat reservoir, to research the behavior of quantum coherence. Our results indicate that a coupled oscillator system can exhibit significant quantum coherence in the ultrastrong and deep strong coupling regimes. In the ground state, the photon-mode and the matter-mode coherences are equal. The larger coherences that encompass the photon mode, the matter mode, and the overall system are achieved at lower optical frequencies and with increased coupling strengths. Notably, the the beam-splitter and phase rotation terms alone does not generate coherences for either total coherence or subsystem coherences; instead, the generation of quantum coherences originates from the one-mode and two-mode squeezing terms. When heat environments are present, the total coherence can be enhanced by the the beam-splitter and phase rotation terms, while it has no effect on subsystem coherences. Moreover, when the one-mode and two-mode squeezing terms and the the beam-splitter and phase rotation terms are considered together, the total coherence increases with stronger coupling. We also observe that lower frequencies maximize total coherence in the deep strong coupling regime. These results demonstrate that the ultrastrong and deep strong coupling regimes give rise to novel characteristics of quantum coherence. This work provides valuable insights into the quantum coherence properties, particularly in the ultrastrong and deep strong coupling regimes between light and matter and may have potential applications in quantum information processing.

quant-ph

The Morphology and Kinematics of a Giant, Symmetric Nebula Around a Radio-Loud Quasar 3C$\,$57: Extended Rotating Gas or Biconical Outflows?

Gas flows between galaxies and the CGM play a crucial role in galaxy evolution. When ionized by a quasar, these gas flows can be directly traced as giant nebulae. We present a study of a giant nebula around a radio-loud quasar, 3C$\,$57 at $z\approx0.672$. Observations from MUSE reveal that the nebula is elongated with a major axis of $70 \, \rm kpc$ and a minor axis of $40 \, \rm kpc$. The nebula displays an approximately symmetric blueshifted-redshifted pattern along the major axis and multi-component emission features in its $\rm[O\,II]$ and $\rm [O\,III]$ profiles. The morphology and kinematics can be explained as rotating gas or biconical outflow, both of which qualitatively reproduce the observed position-velocity diagram. The 3C$\,$57 nebula is significantly more kinematically disturbed, with $\rm W_{80}$ (the line width encompassing 80$\%$ of the flux) of approximately $300{-}400\,\rm km\,s^{-1}$, compared to $\rm H\,I$ gas in local early-type galaxies, which typically shows $\rm W_{80} \approx 50\,\rm km\,s^{-1}$. This velocity dispersion is comparable to the gas in cool-core clusters despite originating in a group 100 times less massive. For biconical outflow models, the inferred $10{-}20^{\circ}$ inclination angle is in tension with the unobscured nature of the quasar, as the dusty torus is expected to be perpendicular to the outflow. Neither a quiescent rotating gas origin nor a biconical outflow fully reproduces the observed kinematics and morphology of the 3C$\,$57 nebula, suggesting a more intricate origin likely involving both rotation and AGN feedback.

astro-ph.GA

Overview of EXL-50 Research Progress and Future Plan

XuanLong-50 (EXL-50) is the first medium-size spherical torus (ST) in China, with the toroidal field at major radius at 50 cm around 0.5T. CS-free and non-inductive current drive via electron cyclotron resonance heating (ECRH) was the main physics research issue for EXL-50. Discharges with plasma currents of 50 kA - 180 kA were routinely obtained in EXL-50, with the current flattop sustained for up to or beyond 2 s. The current drive effectiveness on EXL-50 was as high as 1 A/W for low-density discharges using 28GHz ECRH alone for heating power less than 200 kW. The plasma current reached Ip>80 kA for high-density (5*10e18m-2) discharges with 150 kW 28GHz ECRH. Higher performance discharge (Ip of about 120 kA and core density of about 1*10e19m-3) was achieved with 150 kW 50GHz ECRH. The plasma current in EXL-50 was mainly carried by the energetic electrons.Multi-fluid equilibrium model has been successfully applied to reconstruct the magnetic flux surface and the measured plasma parameters of the EXL-50 equilibrium. The physics mechanisms for the solenoid-free ECRH current drive and the energetic electrons has also been investigated. Preliminary experimental results show that 100 kW of lower hybrid current drive (LHCD) waves can drive 20 kA of plasma current. Several boron injection systems were installed and tested in EXL-50, including B2H6 gas puffing, boron powder injection, boron pellet injection. The research plan of EXL-50U, which is the upgrade machine of EXL-50, is also presented.

physics.plasm-ph

From Defects to Demands: A Unified, Iterative, and Heuristically Guided LLM-Based Framework for Automated Software Repair and Requirement Realization

This manuscript signals a new era in the integration of artificial intelligence with software engineering, placing machines at the pinnacle of coding capability. We present a formalized, iterative methodology proving that AI can fully replace human programmers in all aspects of code creation and refinement. Our approach, combining large language models with formal verification, test-driven development, and incremental architectural guidance, achieves a 38.6% improvement over the current top performer's 48.33% accuracy on the SWE-bench benchmark. This surpasses previously assumed limits, signaling the end of human-exclusive coding and the rise of autonomous AI-driven software innovation. More than a technical advance, our work challenges centuries-old assumptions about human creativity. We provide robust evidence of AI superiority, demonstrating tangible gains in practical engineering contexts and laying the foundation for a future in which computational creativity outpaces human ingenuity.

cs.SE

Understanding GEMM Performance and Energy on NVIDIA Ada Lovelace: A Machine Learning-Based Analytical Approach

Analytical framework for predicting General Matrix Multiplication (GEMM) performance on modern GPUs, focusing on runtime, power consumption, and energy efficiency. Our study employs two approaches: a custom-implemented tiled matrix multiplication kernel for fundamental analysis, and NVIDIA's CUTLASS library for comprehensive performance data collection across advanced configurations. Using the NVIDIA RTX 4070 as our experimental platform, we developed a Random Forest-based prediction model with multi-output regression capability. Through analysis of both naive tiled matrix multiplication with varying tile sizes (1 to 32) and 16,128 CUTLASS GEMM operations across diverse configurations, we identified critical performance patterns related to matrix dimensions, thread block configurations, and memory access patterns. Our framework achieved exceptional accuracy with an R^2 score of 0.98 for runtime prediction (mean error 15.57%) and 0.78 for power prediction (median error 5.42%). The system successfully predicts performance across matrix sizes, demonstrating robust scaling behavior. Our results show that optimal tile size selection can improve performance by up to 3.2x while reducing power consumption by 22% compared to baseline configurations. Analysis of shared memory utilization and SM occupancy reveals that tile sizes of 16x16 achieve the best balance between parallelism and resource usage. The implementation of our framework, including prediction models and analysis tools, is available as an open-source project at GPPerf [https://github.com/pavlyhalim/GPPerf].

cs.DC

Understanding Machine Learning Paradigms through the Lens of Statistical Thermodynamics: A tutorial

This tutorial investigates the convergence of statistical mechanics and learning theory, elucidating the potential enhancements in machine learning methodologies through the integration of foundational principles from physics. The tutorial delves into advanced techniques like entropy, free energy, and variational inference which are utilized in machine learning, illustrating their significant contributions to model efficiency and robustness. By bridging these scientific disciplines, we aspire to inspire newer methodologies in researches, demonstrating how an in-depth comprehension of physical systems' behavior can yield more effective and dependable machine learning models, particularly in contexts characterized by uncertainty.

cs.LG