SearcharxivSearch

arXiv subjects

Yifan Nie

Publications and source records attributed to Yifan Nie.

13 recordsLinked to original sources

Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models

Large language models (LLMs) have achieved strong performance on a wide range of natural language tasks, and recent benchmarks suggest that they are increasingly adept at multi-hop reasoning. However, these benchmarks are typically short-horizon, requiring only a small number of retrieval or inference steps, and provide limited evidence of reliability on real-world tasks that involve following manuals spanning hundreds of pages with complex, interdependent guidelines. In this paper, we introduce Tasks over Application Manuals (TAM), a benchmark for evaluating long-horizon procedural reasoning. We construct TAM by curating real-world tasks from two domains: ICD-10-CM clinical coding (mapping medical conditions to diagnostic codes) and U.S. federal sentencing (computing crime sentencing guideline outcomes, specifically offense levels), with human-validated labels. Each task requires following an authoritative manual with tens of thousands of rules and executing a sequence of interdependent steps across different sections to produce an exact answer. We evaluate general-purpose prompting approaches, including retrieval-augmented generation, ReAct-style prompting, and an agent-harness baseline on GPT-5, and find that the best exact-match performance remains extremely low: 1% on ICD-10-CM coding and 15.5% on sentencing tasks. These results show that current benchmarks may overestimate LLM reasoning ability and miss a key challenge: reliably following long, rule-based procedures. The complete TAM data and code are publicly available.

cs.CL

When Synthetic Data Hurts: On Catastrophic Forgetting in Skill Retrieval for LLM Agents

LLM agents increasingly rely on external skills retrieved at runtime, making skill selection from large repositories a critical challenge. We present a production skill router over 34,396 skills and a large-scale study of skill retrieval using limited real supervision and synthetic data. We found that the synthetic-data fine-tuning improves in-distribution retrieval but it causes catastrophic forgetting on real and out-of-distribution (OOD) data. We evaluate several forgetting mitigation fine-tuning approaches inspired by continual learning, including embedding-anchor regularization, Learning without Forgetting (LwF), Elastic Weight Consolidation (EWC), and L2-initialization. The results show that these approaches not only retain the performance on OOD skills retrieval but also improve the retrieval on synthetic in-distribution skills by 13.98\% for 0.6B Qwen retriever and reranker. Our results provide a practical benchmark and a robust fine-tuning recipe for scarce, multi-positive supervision.

cs.IR

Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs

Prompt optimization can improve multi-agent LLM systems, but the prompts being optimized often serve two entangled roles: generating task-relevant content and specifying execution-critical protocols, such as message routing, output formatting, and termination signals, on which the underlying code relies. As a result, a prompt edit intended to improve content generation can inadvertently corrupt the protocol and cause the entire agent pipeline to fail. Our key observation is that these two roles have different representations: execution protocols are typically structured, while task-relevant content is usually expressed in unstructured language. Based on this, we propose control-data flow separation, where execution-critical control is represented as typed, validated program objects, while task-relevant language remains the optimizable data flow for agent communication. This design allows optimizers to improve multi-agent behavior without exposing the routing or formatting interface to prompt drift. Across synthetic reasoning, collaborative review generation, and insurance rating workflows, our framework empirically achieves 100% eventual protocol validity while consistently improving task performance.

cs.AI

Efficient temporal prediction of compressible flows in irregular domains using Fourier neural operators

This paper investigates the temporal evolution of high-speed compressible fluids in irregular flow fields using the Fourier Neural Operator (FNO). We reconstruct the irregular flow field point set into sequential format compatible with FNO input requirements, and then embed temporal bundling technique within a recurrent neural network (RNN) for multi-step prediction. We further employ a composite loss function to balance errors across different physical quantities. Experiments are conducted on three different types of irregular flow fields, including orthogonal and non-orthogonal grid configurations. Then we comprehensively analyze the physical component loss curves, flow field visualizations, and physical profiles. Results demonstrate that our approach significantly surpasses traditional numerical methods in computational efficiency while achieving high accuracy, with maximum relative $L_2$ errors of (0.78, 0.57, 0.35)% for ($p$, $T$, $\mathbf{u}$) respectively. This verifies that the method can efficiently and accurately simulate the temporal evolution of high-speed compressible flows in irregular domains.

physics.flu-dyn

On the Comprehensibility of Multi-structured Financial Documents using LLMs and Pre-processing Tools

The proliferation of complex structured data in hybrid sources, such as PDF documents and web pages, presents unique challenges for current Large Language Models (LLMs) and Multi-modal Large Language Models (MLLMs) in providing accurate answers. Despite the recent advancements of MLLMs, they still often falter when interpreting intricately structured information, such as nested tables and multi-dimensional plots, leading to hallucinations and erroneous outputs. This paper explores the capabilities of LLMs and MLLMs in understanding and answering questions from complex data structures found in PDF documents by leveraging industrial and open-source tools as part of a pre-processing pipeline. Our findings indicate that GPT-4o, a popular MLLM, achieves an accuracy of 56% on multi-structured documents when fed documents directly, and that integrating pre-processing tools raises the accuracy of LLMs to 61.3% for GPT-4o and 76% for GPT-4, and with lower overall cost. The code is publicly available at https://github.com/OGCDS/FinancialQA.

cs.IR

Routine: A Structural Planning Framework for LLM Agent System in Enterprise

The deployment of agent systems in an enterprise environment is often hindered by several challenges: common models lack domain-specific process knowledge, leading to disorganized plans, missing key tools, and poor execution stability. To address this, this paper introduces Routine, a multi-step agent planning framework designed with a clear structure, explicit instructions, and seamless parameter passing to guide the agent's execution module in performing multi-step tool-calling tasks with high stability. In evaluations conducted within a real-world enterprise scenario, Routine significantly increases the execution accuracy in model tool calls, increasing the performance of GPT-4o from 41.1% to 96.3%, and Qwen3-14B from 32.6% to 83.3%. We further constructed a Routine-following training dataset and fine-tuned Qwen3-14B, resulting in an accuracy increase to 88.2% on scenario-specific evaluations, indicating improved adherence to execution plans. In addition, we employed Routine-based distillation to create a scenario-specific, multi-step tool-calling dataset. Fine-tuning on this distilled dataset raised the model's accuracy to 95.5%, approaching GPT-4o's performance. These results highlight Routine's effectiveness in distilling domain-specific tool-usage patterns and enhancing model adaptability to new scenarios. Our experimental results demonstrate that Routine provides a practical and accessible approach to building stable agent workflows, accelerating the deployment and adoption of agent systems in enterprise environments, and advancing the technical vision of AI for Process.

cs.AI

Regex-augmented Domain Transfer Topic Classification based on a Pre-trained Language Model: An application in Financial Domain

A common way to use large pre-trained language models for downstream tasks is to fine tune them using additional layers. This may not work well if downstream domain is a specialized domain whereas the large language model has been pre-trained on a generic corpus. In this paper, we discuss the use of regular expression patterns employed as features for domain knowledge during the process of fine tuning, in addition to domain specific text. Our experiments on real scenario production data show that this method of fine tuning improves the downstream text classification tasks as compared to fine tuning only on domain specific text. We also show that the use of attention network for fine tuning improves results compared to simple linear layers.

cs.CL

Flatbands and Mechanical Deformation Effects in the Moiré Superlattice of MoS$_2$-WSe$_2$ Heterobilayers

It has recently been shown that quantum-confined states can appear in epitaxially grown van der Waals material heterobilayers without a rotational misalignment ($θ=0^\circ$), associated with flat bands in the Brillouin zone of the moiré pattern formed due to the lattice mismatch of the two layers. Peaks in the local density of states and confinement in a MoS$_2$/WSe$_2$ system was qualitatively described only considering local stacking arrangements, which cause band edge energies to vary spatially. In this work, we report the presence of large in-plane strain variation across the moiré unit cell of a $θ=0^\circ$ MoS$_2$/WSe$_2$ heterobilayer, and show that inclusion of strain variation and out-of-plane displacement in density functional theory calculations greatly improves their agreement with the experimental data. We further explore the role of twist-angle by showing experimental data for a twisted MoS$_2$/WSe$_2$ heterobilayer structure with twist angle of $θ=15^\circ$, that exhibits a moiré pattern but no confinement.

cond-mat.mtrl-sci

Higher superconducting transition temperature by breaking the universal pressure relation

By investigating the bulk superconducting state via dc magnetization measurements, we have discovered a common resurgence of the superconductive transition temperatures (Tcs) of the monolayer Bi2Sr2CuO6+δ (Bi2201) and bilayer Bi2Sr2CaCu2O8+δ (Bi2212) to beyond the maximum Tcs (Tc-maxs) predicted by the universal relation between Tc and doping (p) or pressure (P) at higher pressures. The Tc of under-doped Bi2201 initially increases from 9.6 K at ambient to a peak at ~ 23 K at ~ 26 GPa and then drops as expected from the universal Tc-P relation. However, at pressures above ~ 40 GPa, Tc rises rapidly without any sign of saturation up to ~ 30 K at ~ 51 GPa. Similarly, the Tc for the slightly overdoped Bi2212 increases after passing a broad valley between 20-36 GPa and reaches ~ 90 K without any sign of saturation at ~ 56 GPa. We have therefore attributed this Tc-resurgence to a possible pressure-induced electronic transition in the cuprate compounds due to a charge transfer between the Cu 3d_(x^2-y^2 ) and the O 2p bands projected from a hybrid bonding state, leading to an increase of the density of states at the Fermi level, in agreement with our density functional theory calculations. Similar Tc-P behavior has also been reported in the trilayer Br2Sr2Ca2Cu3O10+δ (Bi2223). These observations suggest that higher Tcs than those previously reported for the layered cuprate high temperature superconductors can be achieved by breaking away from the universal Tc-P relation through the application of higher pressures.

cond-mat.supr-con

Quantum Transport and Band Structure Evolution under High Magnetic Field in Few-Layer Tellurene

Quantum Hall effect (QHE) is a macroscopic manifestation of quantized states which only occurs in confined two-dimensional electron gas (2DEG) systems. Experimentally, QHE is hosted in high mobility 2DEG with large external magnetic field at low temperature. Two-dimensional van der Waals materials, such as graphene and black phosphorus, are considered interesting material systems to study quantum transport, because it could unveil unique host material properties due to its easy accessibility of monolayer or few-layer thin films at 2D quantum limit. Here for the first time, we report direct observation of QHE in a novel low-dimensional material system: tellurene.High-quality 2D tellurene thin films were acquired from recently reported hydrothermal method with high hole mobility of nearly 3,000 cm2/Vs at low temperatures, which allows the observation of well-developed Shubnikov-de-Haas (SdH) oscillations and QHE. A four-fold degeneracy of Landau levels in SdH oscillations and QHE was revealed. Quantum oscillations were investigated under different gate biases, tilted magnetic fields and various temperatures, and the results manifest the inherent information of the electronic structure of Te. Anomalies in both temperature-dependent oscillation amplitudes and transport characteristics were observed which are ascribed to the interplay between Zeeman effect and spin-orbit coupling as depicted by the density functional theory (DFT) calculations.

cond-mat.mtrl-sci

Quantum-Confined Electronic States arising from Moiré Pattern of MoS2-WSe2 Hetero-bilayers

A two-dimensional (2D) hetero-bilayer system consisting of MoS2 on WSe2, deposited on epitaxial graphene, is studied by scanning tunneling microscopy and spectroscopy at temperatures of 5 and 80 K. A moiré pattern is observed, arising from lattice mismatch of 3.7% between the MoS2 and WSe2. Significant energy shifts are observed in tunneling spectra observed at the maxima of the moiré corrugation, as compared with spectra obtained at corrugation minima, consistent with prior work. Furthermore, at the minima of the moiré corrugation, sharp peaks in the spectra at energies near the band edges are observed, for spectra acquired at 5 K. The peaks correspond to discrete states that are confined within the moiré unit cells. Conductance mapping is employed to reveal the detailed structure of the wave functions of the states. For measurements at 80 K, the sharp peaks in the spectra are absent, and conductance maps of the band edges reveal little structure.

cond-mat.mes-hall

Magnitude of the Current in Two-Dimensional Interlayer Tunneling Devices

Using the Bardeen tunneling method with first-principles wave functions, computations are made of the tunneling current in graphene / hexagonal-boron-nitride / graphene (G/h-BN/G) vertical structures. Detailed comparison with prior experimental results is made, focusing on the magnitude of the achievable tunnel current. With inclusion of the effects of translational and rotational misalignment of the graphene and the h-BN, predicted currents are found to be about 15x larger than experimental values. A reduction in this discrepancy, to a factor of 2.5x, is achieved by utilizing a realistic size for the band gap of the h-BN, hence affecting the exponential decay constant for the tunneling.

cond-mat.mes-hall

Characteristics of Interlayer Tunneling Field Effect Transistors Computed by a "DFT-Bardeen" Method

Theoretical predictions are made for the current-voltage characteristics of two-dimensional heterojunction interlayer tunneling field-effect transistors (Thin-TFETs), focusing on the magnitude of the current that is achievable in such devices. A theory based on the Bardeen tunneling method is employed, using wavefunctions from first-principles density-functional theory. This method permits convenient incorporation of differing materials into the source and drain electrodes, i.e. with different crystal structures, lattice constants, and/or band structures. Large variations in the tunnel currents are found, depending on the particular two-dimensional materials used for the source and drain electrodes. Tunneling between states derived from the center (Gamma-point) of the Brillouin zone (BZ) is found, in general, to lead to larger current than for zone-edge (e.g. K-point) states. Differences, as large as an order of magnitude, between the present results and various prior predictions are discussed. Predicted values for the tunneling currents, including subthreshold swing, are compared with benchmark values for low-power digital applications. Contact resistance is considered and its effect on the tunneling currents is demonstrated.

cond-mat.mes-hall