SearcharxivSearch

arXiv subjects

Ying Liao

Publications and source records attributed to Ying Liao.

13 recordsLinked to original sources

Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?

Large language model (LLM) agents increasingly rely on skills, structured documents that specify when to act, which procedure to follow, and which tools are allowed. Existing evaluations mostly judge the quality of a skill or its contribution to task success, leaving unexamined whether an agent can recognize a relevant skill and apply it on its own. We introduce Skill-Use, a benchmark that evaluates skill use under progressive disclosure, where an agent sees only a skill's name and short description and must retrieve the full procedure before following it. Skill-Use separates three facets of skill use. Trigger measures whether the agent invokes the relevant skill, Compliance measures how faithfully it follows the prescribed procedure, and Boundary measures whether it avoids forbidden operations. A Skill-Use (SU) score combines the three and credits execution only after the skill is triggered. Skill-Use pairs 79 real skills with 177 executable tasks across nine domains, each grounded in real files, run in an isolated Docker sandbox, and scored by a trajectory-based rubric. Evaluating eight LLMs under two agent harnesses, we find that reliable skill use remains out of reach, as the strongest configuration reaches an SU of only 0.613. Triggering and procedural compliance fail as independent bottlenecks, and both scores and model rankings shift with the harness, so skill use behaves as a capability conditioned on the harness rather than a fixed property of the model.

cs.CL

Implied Volatility Expansions for VIX Options in Forward Variance Models

We develop closed-form expansions for the implied volatility of VIX options within the class of forward variance models. Our approach builds on weak-approximation techniques for VIX option prices and yields explicit implied volatility expansions with computable correction terms. The resulting formulas enable fast and accurate calibration without requiring numerical root-finding using option prices. We illustrate the performance of the proposed expansions in both standard and rough Bergomi-type models, as well as in mixed specifications, and demonstrate their accuracy through numerical experiments.

q-fin.CP

GenericAgent: A Token-Efficient Self-Evolving LLM Agent via Contextual Information Density Maximization (V1.0)

Long-horizon large language model (LLM) agents are fundamentally limited by context. As interactions become longer, tool descriptions, retrieved memories, and raw environmental feedback accumulate and push out the information needed for decision-making. At the same time, useful experience gained from tasks is often lost across episodes. We argue that long-horizon performance is determined not by context length, but by how much decision-relevant information is maintained within a finite context budget. We present GenericAgent (GA), a general-purpose, self-evolving LLM agent system built around a single principle: context information density maximization. GA implements this through four closely connected components: a minimal atomic tool set that keeps the interface simple, a hierarchical on-demand memory that only shows a small high-level view by default, a self-evolution mechanism that turns verified past trajectories into reusable SOPs and executable code, and a context truncation and compression layer that maintains information density during long executions. Across task completion, tool use efficiency, memory effectiveness, self-evolution, and web browsing, GA consistently outperforms leading agent systems while using significantly fewer tokens and interactions, and it continues to evolve over time. Project: https://github.com/lsdefine/GenericAgent

cs.CL

Not All Negative Samples Are Equal: LLMs Learn Better from Plausible Reasoning

Learning from negative samples holds great promise for improving Large Language Model (LLM) reasoning capability, yet existing methods treat all incorrect responses as equally informative, overlooking the crucial role of sample quality. To address this, we propose Plausible Negative Samples (PNS), a method that synthesizes high-quality negative samples exhibiting expected format and structural coherence while ultimately yielding incorrect answers. PNS trains a dedicated model via reverse reinforcement learning (RL) guided by a composite reward combining format compliance, accuracy inversion, reward model assessment, and chain-of-thought evaluation, generating responses nearly indistinguishable from correct solutions. We further validate PNS as a plug-and-play data source for preference optimization across three backbone models on seven mathematical reasoning benchmarks. Results demonstrate that PNS consistently outperforms other negative sample synthesis methods, achieving an average improvement of 2.03% over RL-trained models.

cs.LG

Structured Reasoning for Large Language Models

Large language models (LLMs) achieve strong performance by generating long chains of thought, but longer traces always introduce redundant or ineffective reasoning steps. One typical behavior is that they often perform unnecessary verification and revisions even if they have reached the correct answers. This limitation stems from the unstructured nature of reasoning trajectories and the lack of targeted supervision for critical reasoning abilities. To address this, we propose Structured Reasoning (SCR), a framework that decouples reasoning trajectories into explicit, evaluable, and trainable components. We mainly implement SCR using a Generate-Verify-Revise paradigm. Specifically, we construct structured training data and apply Dynamic Termination Supervision to guide the model in deciding when to terminate reasoning. To avoid interference between learning signals for different reasoning abilities, we adopt a progressive two-stage reinforcement learning strategy: the first stage targets initial generation and self-verification, and the second stage focuses on revision. Extensive experiments on three backbone models show that SCR substantially improves reasoning efficiency and self-verification. Besides, compared with existing reasoning paradigms, it reduces output token length by up to 50%.

cs.CL

Your Models Have Thought Enough: Training Large Reasoning Models to Stop Overthinking

Large Reasoning Models (LRMs) have achieved impressive performance on challenging tasks, yet their deep reasoning often incurs substantial computational costs. To achieve efficient reasoning, existing reinforcement learning methods still struggle to construct short reasoning path during the rollout stage, limiting effective learning. Inspired by Evidence Accumulation Models, we find that LRMs have accumulated sufficient information early in reasoning, making further reasoning steps redundant. Based on this insight, we propose Just-Enough Thinking (JET), which trains models to proactively terminate unnecessary reasoning. JET performs trajectory truncation during rollout to expose the model to short, distributionally consistent reasoning paths. Besides, it uses a quality-controlled length reward to better encourage concise reasoning while maintaining correctness. Extensive experiments demonstrate that JET significantly improves reasoning efficiency without sacrificing accuracy. Especially, DeepSeek-Distill-Qwen-1.5B achieves a 4.6% accuracy gain while reducing output length by 46.3% on the Olympiad benchmark. Our code is available in the GitHub.

cs.AI

Efficient calibration of the shifted square-root diffusion model to credit default swap spreads using asymptotic approximations

We derive a closed-form approximation for the credit default swap (CDS) spread in the two-dimensional shifted square-root diffusion (SSRD) model using asymptotic coefficient expansion technique to approximate solutions of nonlinear partial differential equations. Specifically, we identify the Cauchy problems associated with two terms in the CDS spread formula that lack analytical solutions and derive asymptotic approximations for these terms. Our approximation does not require the assumption of uncorrelated interest rate and default intensity processes as typically required for calibration in the SSRD model. Through several calibration studies using market data on CDS spread, we demonstrate the accuracy and efficiency of our proposed formula.

q-fin.MF

Assessment of Precision and Accuracy of Brain White Matter Microstructure using Combined Diffusion MRI and Relaxometry

Joint modeling of diffusion and relaxation has seen growing interest due to its potential to provide complementary information about tissue microstructure. For brain white matter, we designed an optimal diffusion-relaxometry MRI protocol that samples multiple b-values, B-tensor shapes, and echo times (TE). This variable-TE protocol (27 min) has as subsets a fixed-TE protocol (15 min) and a 2-shell dMRI protocol (7 min), both characterizing diffusion only. We assessed the sensitivity, specificity and reproducibility of these protocols with synthetic experiments and in six healthy volunteers. Compared with the fixed-TE protocol, the variable-TE protocol enables estimation of free water fractions while also capturing compartmental $T_2$ relaxation times. Jointly measuring diffusion and relaxation offers increased sensitivity and specificity to microstructure parameters in brain white matter with voxelwise coefficients of variation below 10%.

physics.med-ph

Mapping tissue microstructure of brain white matter in vivo in health and disease using diffusion MRI

Diffusion magnetic resonance imaging offers unique in vivo sensitivity to tissue microstructure in brain white matter, which undergoes significant changes during development and is compromised in virtually every neurological disorder. Yet, the challenge is to develop biomarkers that are specific to micrometer-scale cellular features in a human MRI scan of a few minutes. Here we quantify the sensitivity and specificity of a multicompartment diffusion modeling framework to the density, orientation and integrity of axons. We demonstrate that using a machine learning based estimator, our biophysical model captures the morphological changes of axons in early development, acute ischemia and multiple sclerosis (total N=821). The methodology of microstructure mapping is widely applicable in clinical settings and in large imaging consortium data to study development, aging and pathology.

physics.med-ph

Volume electron microscopy in injured rat brain validates white matter microstructure metrics from diffusion MRI

Biophysical modeling of diffusion MRI (dMRI) offers the exciting potential of bridging the gap between the macroscopic MRI resolution and microscopic cellular features, effectively turning the MRI scanner into a noninvasive in vivo microscope. In brain white matter, the Standard Model (SM) interprets the dMRI signal in terms of axon dispersion, intra- and extra-axonal water fractions and diffusivities. However, for SM to be fully applicable and correctly interpreted, it needs to be carefully evaluated using histology. Here, we perform a comprehensive histological validation of the SM parameters, by characterizing WM microstructure in sham and injured rat brains using volume (3d) electron microscopy (EM) and ex vivo dMRI. Sensitivity is evaluated by how close each SM metric is to its histological counterpart, and specificity by how independent it is from other, non-corresponding histological features. This comparison reveals that SM is sensitive and specific to microscopic properties, clearing the way for the clinical adoption of in vivo dMRI derived SM parameters as biomarkers for neurological disorders.

physics.bio-ph

FC$^2$N: Fully Channel-Concatenated Network for Single Image Super-Resolution

Most current image super-resolution (SR) methods based on convolutional neural networks (CNNs) use residual learning in network structural design, which favors to effective back propagation and hence improves SR performance by increasing model scale. However, residual networks suffer from representational redundancy by introducing identity paths that impede the full exploitation of model capacity. Besides, blindly enlarging network scale can cause more problems in model training, even with residual learning. In this paper, a novel fully channel-concatenated network (FC$^2$N) is presented to make further mining of representational capacity of deep models, in which all interlayer skips are implemented by a simple and straightforward operation, i.e., weighted channel concatenation (WCC), followed by a 1$\times$1 conv layer. Based on the WCC, the model can achieve the joint attention mechanism of linear and nonlinear features in the network, and presents better performance than other state-of-the-art SR models with fewer model parameters. To our best knowledge, FC$^2$N is the first CNN model that does not use residual learning and reaches network depth over 400 layers. Moreover, it shows excellent performance in both largescale and lightweight implementations, which illustrates the full exploitation of the representational capacity of the model.

eess.IV

Limitations in the determination of surface emission distributions on comets through modelling of observational data -- A case study based on Rosetta observations

The European Space Agency's (ESA) Rosetta mission has returned a vast data set of measurements of the inner gas coma of comet 67P/Churyumov-Gerasimenko. These measurements have been used by different groups to determine the distribution of the gas sources at the nucleus surface. The solutions that have been found differ from each other substantially and illustrate the degeneracy of this issue. It is the aim of this work to explore the limitations that current gas models have in linking the coma measurements to the surface. In particular, we discuss the sensitivity of Rosetta's ROSINA/COPS, VIRTIS, and MIRO instruments to differentiate between vastly different spatial distributions of the gas emission from the surface. We have applied a state of the art 3D DSMC gas dynamics code to simulate the inner gas coma of different models that vary in the fraction of the surface that contains ice and in different sizes of active patches. These different distributions result in jet interactions that differ in their dynamical behaviour. We have found that ROSINA/COPS measurements by themselves cannot detect the differences in our models. While ROSINA/COPS measurements are important to constrain the regional inhomogeneities of the gas emission, they can by themselves not determine the surface emission distribution of the gas sources to a spatial accuracy of better than a few hundred metres (400 m). Any solutions fitting the ROSINA/COPS measurements is hence fundamentally degenerate, be it through a forward or inverse model. Only other instruments with complementary measurements can potentially lift this degeneracy as we show here for VIRTIS and MIRO. Finally, as a by-product, we have explored the effect of our activity distributions on lateral flow at the surface that may be responsible for some of the observed aeolian features.

astro-ph.EP

Health Assessment and Prognostics Based on Higher Order Hidden Semi-Markov Models

This paper presents a new and flexible prognostics framework based on a higher order hidden semi-Markov model (HOHSMM) for systems or components with unobservable health states and complex transition dynamics. The HOHSMM extends the basic hidden Markov model (HMM) by allowing the hidden state to depend on its more distant history and assuming generally distributed state duration. An effective Gibbs sampling algorithm is designed for statistical inference of an HOHSMM. The performance of the proposed HOHSMM sampler is evaluated by conducting a simulation experiment. We further design a decoding algorithm to estimate the hidden health states using the learned model. Remaining useful life (RUL) is predicted using a simulation approach given the decoded hidden states. The practical utility of the proposed prognostics framework is demonstrated by a case study on NASA turbofan engines. The results show that the HOHSMM-based prognostics framework provides good hidden health state assessment and RUL estimation for complex systems.

stat.AP