SearcharxivSearch

arXiv subjects

Yilin Chen

Publications and source records attributed to Yilin Chen.

At least 19 recordsLinked to original sources

Argus: A General-Purpose Agentic Reasoning Runtime for Long-Horizon Tasks

Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements reveal failure, hidden constraints, or a misspecified objective. We present Argus, a persistent, self-evolving runtime in which Manager, Planner, Engineer, and Reviewer execute bounded missions over durable project state. Argus separates stable user intent from operational objectives, constraints, and verification criteria, and admits memories, skills, procedures, verifiers, routing decisions, and rejected routes only after role-owned review and, when available, task-native verification. Model weights remain fixed; self-evolution occurs through persistent runtime state and control policy, with autonomous execution between operator-owned escalation points. Across seven GPT-5.5 benchmark arenas, Argus achieves about 78% on SWE-Bench Pro versus 59% for Direct Copilot while using 1.41 times the aggregate tokens. After verification-gated self-evolution, mature SWE-Bench waves use 21% fewer solve-input tokens and 15% less active workflow time per task than startup waves, while recording 34 verifier recoveries and 22 strict review-loop rescues. Argus also reaches 76.8% on AARRI-Bench and a 28.0-point gap on mathematical data synthesis, with competitive GPU-kernel and language-model-training results. Beyond benchmarks, an optimized RWKV6 kernel was merged upstream; a multi-day mathematics campaign retained falsified routes and proof-backed frontier updates; and six paper pipelines completed 254 missions with 16 stage rollbacks. These results show that a fixed-weight, self-evolving harness can revise, recover, and accumulate verified approaches while producing structured trajectories for future supervised and reinforcement learning.

cs.AI

MESH: Scaling Up Retrieval with Heterogeneous Content Unification

Optimizing large-scale retrieval hinges on the ability to efficiently surface candidates across diverse content tiers. However, to capture segments such as fresh and long-tail content, modern systems typically resort to a fragmented "zoo" of specialized retrieval models. This operational complexity is attributed to a fundamental challenge in heterogeneous retrieval systems, the Scaling Bias of Heterogeneity, where model capacity gains do not apply equally across diverse content tiers. To bridge this gap, we propose MESH as a unified retrieval scaling framework that mitigates this bias through a modularized architecture integrated with gated bias correction. By partitioning the feature space into independent domains, MESH enforces a structural inductive bias that reduces interference between sparse-item signals and high-frequency engagement features. This protected gradient path leads to improved scaling behavior for sparse content, empirically validated by a 14 times improvement in the power-law scaling exponent for fresh items. In online evaluations on Pinterest's Related Pins platform, a billion scale item-to-item recommendation system, these improvements translate into a +5.5% lift in fresh-item repins, alongside with 55% improvement in funnel efficiency and +0.46% improvement in user retention. Finally, our asynchronous serving strategy ensures production viability by delivering a 2.87 times improvement in system throughput. Our findings suggest MESH as a promising paradigm for consolidating fragmented retrieval infrastructures into more scalable and ecosystem-aware backbones.

cs.IR

Multi-Segment Attention: Enabling Efficient KV-Cache Management for Faster Large Language Model Serving

Large Language Model (LLM) inference relies on key-value (KV) caches to avoid redundant attention computation. While approximate KV cache retention techniques reduce memory usage by sacrificing model accuracy, lossless approaches instead evict KV cache blocks from GPU memory and reconstruct them on demand to preserve exact outputs. Existing lossless KV cache management systems primarily base eviction decisions on access frequency or positional heuristics, without considering how different KV cache blocks affect the execution efficiency of GPU attention kernels. In this paper, we propose AsymCache, a computation-latency-aware KV cache management system for LLM inference that explicitly aligns cache residency decisions with GPU attention kernel performance, including three key components: Multi-Segment Attention (MSA) for efficient non-contiguous KV context processing, a cache eviction policy that jointly optimizes hit rate and position-aware recomputation cost, and an adaptive chunking scheduler for high hardware utilization. Experiments show that AsymCache reduces TTFT by up to 1.90-2.03x and time-per-output-token (TPOT) by 1.62-1.71x over latest baselines, confirming the effectiveness of the method in common workloads and validating its design goal of balancing computational efficiency with cache hit rate. Moreover, the low-level design of AsymCache allows seamless integration into agent serving systems such as Continuum, where it further reduces average job latency by up to 18.1%.

cs.AR

When Can Digital Personas Reliably Approximate Human Survey Findings?

Digital personas powered by Large Language Models (LLMs) are increasingly proposed as substitutes for human survey respondents, yet it remains unclear when they can reliably approximate human survey findings. We answer this question using the LISS panel, constructing personas from respondents' background variables and pre-2023 survey histories, then testing them against the same respondents' held-out post-cutoff answers. Across four persona architectures, three LLMs, and two prediction tasks, we assess performance at the question, respondent, distributional, equity, and clustering levels. Digital personas improve alignment with human response distributions, especially in domains tied to stable attributes and values, but remain limited for individual prediction and fail to recover multivariate respondent structure. Retrieval-augmented architectures provide the clearest gains, but performance depends more on human response structure than on model choice: personas perform best for low-variability questions and common respondent patterns, and worst for subjective, heterogeneous, or rare responses. Our results provide practical guidance on when digital personas could be appropriate for survey research and when human validation remains necessary.

cs.CL

The First Challenge on Mobile Real-World Image Super-Resolution at NTIRE 2026: Benchmark Results and Method Overview

This paper provides a review of the NTIRE 2026 challenge on mobile real-world image super-resolution, highlighting the proposed solutions and the resulting outcomes. The challenge aims to recover high-resolution (HR) images from low-resolution (LR) counterparts generated through unknown degradations with a x4 scaling factor while ensuring the models remain executable on mobile devices. The objective is to develop effective and efficient network designs or solutions that achieve state-of-the-art real-world image super-resolution performance. The track of the challenge evaluates performance using a weighted combination of image quality assessment (IQA) score and speedup ratios. The competition attracted 108 registrants, with 16 teams achieving a valid score in the final ranking. This collaborative effort advances the performance of mobile real-world image super-resolution while offering an in-depth overview of the latest trends in the field.

cs.CV

A defect in diamond with millisecond-scale spin relaxation time at room temperature

Spin defects in diamond are promising platforms for quantum sensing. The longest electron spin relaxation times ($T_1$) at room temperature for solid-state defects are observed in nitrogen vacancy centers in diamond, which can reach 6.67 ms, and substitutional nitrogen ("P1 centers") in diamond, which exhibit a $T_1$ of 2 ms. No other solid-state defect has exhibited millisecond-scale spin relaxation times at room temperature thus far. Here, we characterize the spin properties of the WAR5 defect in diamond with pulsed electron spin resonance. The observed $T_1$ is one of the longest for solid-state spin defects: 0.97(27) ms at room temperature and 14.38(19) min at 4 K. The observed coherence time ($T_2$) is 246(7) $\mu$s, which can be extended to 6.49(34) ms at 4 K with dynamical decoupling. Furthermore, we demonstrate optical spin polarization with a range of wavelengths from 405 nm to 500 nm and propose potential zero-phonon line candidates.

cond-mat.mtrl-sci

Towards Robust Process Reward Modeling via Noise-aware Learning

Process Reward Models (PRMs) have achieved strong results in complex reasoning, but are bottlenecked by costly process-level supervision. A widely used alternative, Monte Carlo Estimation (MCE), defines process rewards as the probability that a policy model reaches the correct final answer from a given reasoning step. However, step correctness is an intrinsic property of the reasoning trajectory, and should be invariant to policy choice. Our empirical findings show that MCE producing policy-dependent rewards that induce label noise, including false positives that reward incorrect steps and false negatives that penalize correct ones. To address above challenges, we propose a two-stage framework to mitigate noisy supervision. In the labeling stage, we introduce a reflection-aware label correction mechanism that uses a large language model (LLM) as a judge to detect reflection and self-correction behaviors related to the current reasoning step, thereby suppressing overestimated rewards. In the training stage, we further propose a \underline{\textbf{N}}oise-\underline{\textbf{A}}ware \underline{\textbf{I}}terative \underline{\textbf{T}}raining framework that enables the PRM to progressively refine noisy labels based on its own confidence. Extensive Experiments show that our method substantially improves step-level correctness discrimination, achieving up to a 27\% absolute gain in average F1 over PRMs trained with noisy supervision.

cs.CL

High-Performance Near-Infrared Quantum Emission from Color Centers in hBN

Color centers hosted in hexagonal boron nitride have emerged as a highly promising platform for single-photon emission and spin-photon technologies relevant to quantum communication and quantum networking. As a wide-bandgap van der Waals material, hBN can host optically active quantum defects across a broad spectral range. Here, we demonstrate a simple and scalable oxygen-plasma process that reproducibly creates single quantum emitters in hBN with blinking-free zero-phonon lines spanning the near-infrared from 700 up to 971 nm. These emitters combine MHz-level brightness, single-photon purity up to 99.9\%, and ultranarrow cryogenic linewidths down to 2.7~GHz under quasi-resonant excitation, placing them in a particularly attractive regime for quantum photonics. Photostability measurements further reveal resistance to photobleaching, sub-nm spectral stability over long timescales, and near-shot-noise-limited intensity fluctuations. Analysis of the phonon sidebands shows weak vibronic coupling and ZPL-dominated emission, with Debye--Waller factors approaching 50\%. Control experiments together with EDS elemental mapping support oxygen incorporation as a necessary ingredient in activating the NIR emitter population, while first-principles calculations identify O$_N$V$_N$ and O$_N$V$_N$H as the leading defect candidates. These results establish a high-performance NIR quantum-emitter platform in hBN for free-space quantum networking and future integrated quantum-photonic architectures.

quant-ph

Efficiently Transforming Neural Networks into Decision Trees: A Path to Ground Truth Explanations with RENTT

Although neural networks are a powerful tool, their widespread use is hindered by the opacity of their decisions and their black-box nature, which result in a lack of trustworthiness. To alleviate this problem, methods in the field of explainable Artificial Intelligence try to unveil how such automated decisions are made. But explainable AI methods are often plagued by missing faithfulness/correctness, meaning that they sometimes provide explanations that do not align with the neural network's decision and logic. Recently, transformations to decision trees have been proposed to overcome such problems. Unfortunately, they typically lack exactness, scalability, or interpretability as the size of the neural network grows. Thus, we generalize these previous results, especially by considering convolutional neural networks, recurrent neural networks, non-ReLU activation functions, and bias terms. Our findings are accompanied by rigorous proofs and we present a novel algorithm RENTT (Runtime Efficient Network to Tree Transformation) designed to compute an exact equivalent decision tree representation of neural networks in a manner that is both runtime and memory efficient. The resulting decision trees are multivariate and thus, possibly too complex to understand. To alleviate this problem, we also provide a method to calculate the ground truth feature importance for neural networks via the equivalent decision trees - for entire models (global), specific input regions (regional), or single decisions (local). All theoretical results are supported by detailed numerical experiments that emphasize two key aspects: the computational efficiency and scalability of our algorithm, and that only RENTT succeeds in uncovering ground truth explanations compared to conventional approximation methods like LIME and SHAP. All code is available at https://github.com/HelenaM23/RENTT .

cs.LG

Human-Agent Collaborative Paper-to-Page Crafting

In the quest for scientific progress, communicating research is as vital as the discovery itself. Yet, researchers are often sidetracked by the manual, repetitive chore of building project webpages to make their dense papers accessible. While automation has tackled static slides and posters, the dynamic, interactive nature of webpages has remained an unaddressed challenge. To bridge this gap, we reframe the problem, arguing that the solution lies not in a single command, but in a collaborative, hierarchical process. We introduce $\textbf{AutoPage}$, a novel multi-agent system that embodies this philosophy. AutoPage deconstructs paper-to-page creation into a coarse-to-fine pipeline from narrative planning to multimodal content generation and interactive rendering. To combat AI hallucination, dedicated "Checker" agents verify each step against the source paper, while optional human checkpoints ensure the final product aligns perfectly with the author's vision, transforming the system from a mere tool into a powerful collaborative assistant. To rigorously validate our approach, we also construct $\textbf{PageBench}$, the first benchmark for this new task. Experiments show AutoPage not only generates high-quality, visually appealing pages but does so with remarkable efficiency in under 15 minutes for less than \$0.1. Code and dataset will be released at $\href{https://mqleet.github.io/AutoPage_ProjectPage/}{Webpage}$.

cs.SE

Harnessing Rule-Based Reinforcement Learning for Enhanced Grammatical Error Correction

Grammatical error correction is a significant task in NLP. Traditional methods based on encoder-decoder models have achieved certain success, but the application of LLMs in this field is still underexplored. Current research predominantly relies on supervised fine-tuning to train LLMs to directly generate the corrected sentence, which limits the model's powerful reasoning ability. To address this limitation, we propose a novel framework based on Rule-Based RL. Through experiments on the Chinese datasets, our Rule-Based RL framework achieves \textbf{state-of-the-art }performance, with a notable increase in \textbf{recall}. This result clearly highlights the advantages of using RL to steer LLMs, offering a more controllable and reliable paradigm for future development in GEC.

cs.CL

Pseudo Empirical Likelihood Inference for Non-Probability Survey Samples

In this paper, the authors first provide an overview of two major developments on complex survey data analysis: the empirical likelihood methods and statistical inference with non-probability survey samples, and highlight the important research contributions to the field of survey sampling in general and the two topics in particular by Canadian survey statisticians. The authors then propose new inferential procedures on analyzing non-probability survey samples through the pseudo empirical likelihood approach. The proposed methods lead to asymptotically equivalent point estimators that have been discussed in the recent literature but possess more desirable features on confidence intervals such as range-respecting and data-driven orientation. Results from a simulation study demonstrate the superiority of the proposed methods in dealing with binary response variables.

stat.ME

Vertical Profile Corrected Satellite NH3 Retrievals Enable Accurate Agricultural Emission Characterization in China

Ammonia (NH3) emissions significantly contribute to atmospheric pollution, yet discrepancies exist between bottom-up inventories and satellite-constrained top-down estimates, with the latter typically one-third higher. This study quantifies how assumptions about NH3 vertical distribution in satellite retrievals contribute to this gap. By implementing spatially and temporally resolved vertical profiles from the Community Multiscale Air Quality model to replace steep gradients in Infrared Atmospheric Sounding Interferometer (IASI) retrievals, we reduced satellite-model column discrepancies from 71% to 18%. We subsequently constrained NH3 emissions across China using a hybrid inversion framework combining iterative mass balance and four-dimensional variational methods. Our posterior emissions showed agreement with the a priori inventory (7.9% lower), suggesting that discrepancies between inventory approaches were amplified by overestimation of near-surface NH3 in baseline satellite retrievals, potentially causing a 43% overestimation of growing season emissions. Evaluation against ground-based measurements confirmed improved model performance, with normalized root-mean-square error reductions of 1-27% across six months. These findings demonstrate that accurate representation of vertical profiles in satellite retrievals is critical for robust NH3 emission estimates and can reconcile the long-standing discrepancy between bottom-up and top-down approaches. Our hybrid inversion methodology, leveraging profile-corrected satellite data, reveals that China's NH3 emissions exhibit greater spatial concentration than previously recognized, reflecting agricultural intensification. This advancement enables timely and accurate characterization of rapidly changing agricultural emission patterns, critical for implementing effective nitrogen pollution control measures.

physics.ao-ph

PinRec: Unified Generative Retrieval for Pinterest Recommender Systems

Generative retrieval methods employ sequential modeling techniques, like transformers, to generate candidate items for recommender systems. These methods have demonstrated promising results in academic benchmarks, surpassing traditional retrieval models such as two-tower architectures. However, a key limitation is that current approaches require a separate model for each product surface, as building a unified model that accommodates the different business needs of various surfaces has proven challenging. Furthermore, existing methods often fail to capture the evolution of user interests over a sequence, focusing instead on only predicting the next item. This paper introduces Pinrec, a novel unified generative retrieval model for all of Pinterest's recommendation surfaces, including home feed, search, and related pins. Pinrec is pretrained on user activity sequences aggregated across surfaces, then fine-tuned for each surface using that surface's impression data. This pretraining-fine-tuning approach enables a single unified model while still adapting to the needs of individual surfaces. To better align recommendations with surface-specific business goals, Pinrec incorporates a novel outcome-conditioned generation mechanism that targets different outcomes for each surface, which further enhances the impact of fine-tuning. Our experiments show that Pinrec balances performance, diversity, and efficiency, delivering significant gains such as +4% increase in search saves. To our knowledge, this paper presents the first rigorous study of a unified generative retrieval model built and deployed at Pinterest scale, marking a significant milestone in the field.

cs.IR

PQCache: Product Quantization-based KVCache for Long Context LLM Inference

As the field of Large Language Models (LLMs) continues to evolve, the context length in inference is steadily growing. Key-Value Cache (KVCache), the intermediate representations of tokens within LLM inference, has now become the primary memory bottleneck due to limited GPU memory. Current methods selectively determine suitable keys and values for self-attention computation in LLMs to address the issue. However, they either fall short in maintaining model quality or result in high serving latency. Drawing inspiration from advanced embedding retrieval techniques prevalent in the data management community, we consider the storage and retrieval of KVCache as a typical embedding retrieval problem. We propose PQCache, which employs Product Quantization (PQ) to manage KVCache, maintaining model quality while ensuring low serving latency. During the prefilling phase, we apply PQ to tokens' keys for each LLM layer and head. During the autoregressive decoding phase, we use PQ codes and centroids to approximately identify important preceding tokens, then fetch the corresponding key-value pairs for self-attention computation. Through meticulous design of overlapping and caching, we minimize any additional computation and communication overhead during both phases. Extensive experiments demonstrate that PQCache achieves both effectiveness and efficiency, with 4.60% score improvement over existing methods on InfiniteBench and low system latency in both prefilling and decoding.

cs.CL

PointSCNet: Point Cloud Structure and Correlation Learning Based on Space Filling Curve-Guided Sampling

Geometrical structures and the internal local region relationship, such as symmetry, regular array, junction, etc., are essential for understanding a 3D shape. This paper proposes a point cloud feature extraction network named PointSCNet, to capture the geometrical structure information and local region correlation information of a point cloud. The PointSCNet consists of three main modules: the space-filling curve-guided sampling module, the information fusion module, and the channel-spatial attention module. The space-filling curve-guided sampling module uses Z-order curve coding to sample points that contain geometrical correlation. The information fusion module uses a correlation tensor and a set of skip connections to fuse the structure and correlation information. The channel-spatial attention module enhances the representation of key points and crucial feature channels to refine the network. The proposed PointSCNet is evaluated on shape classification and part segmentation tasks. The experimental results demonstrate that the PointSCNet outperforms or is on par with state-of-the-art methods by learning the structure and correlation of point clouds effectively.

cs.CV

Multi-Constitutive Neural Network for Large Deformation Poromechanics Problem

In this paper, we study the problem of large-strain consolidation in poromechanics with deep neural networks (DNN). Given different material properties and different loading conditions, the goal is to predict pore pressure and settlement. We propose a novel method "multi-constitutive neural network" (MCNN) such that one model can solve several different constitutive laws. We introduce a one-hot encoding vector as an additional input vector, which is used to label the constitutive law we wish to solve. Then we build a DNN which takes $(\hat{X}, \hat{t})$ as input along with a constitutive law label and outputs the corresponding solution. It is the first time, to our knowledge, that we can evaluate multi-constitutive laws through only one training process while still obtaining good accuracies. We found that MCNN trained to solve multiple PDEs outperforms individual neural network solvers trained with PDE in some cases.

cs.LG

Operator dependent para-controlled calculus, periodic homogenisation and singular PDEs

We develop a variant of the para-controlled distributions framework based on operator dependent (generalised) Besov spaces. These spaces were introduced by Kerkyacharian-Petrushev ([KP15]). In contrast to the Fourier decomposition in the classical setting, they are built from spectral decomposition of general self-adjoint elliptic operators. A major difference is that products of two spectrally localised functions in general have ``frequencies" spread over the whole spectrum. We obtain a decay estimate on ``high frequency" component of the product of two spectrally-localised functions. This enables us to obtain uniform bounds for the naturally associated para-product and commutator operations over a class of operators. This in particular includes the family of periodic homogenisation operators uniform in the oscillation parameter. Next, we show convergence properties of these operations in the periodic homogenisation setting. A key ingredient is the convergence of the generalised Littlewood-Paley block operators with quantitative dependence on the spectrum level. The proof combines a novel clustering argument to group nearby eigenvalues, Weyl's asymptotic formula that implies quantitative bounds on the sizes and locations of the eigenvalue clusters, together with a result by Kenig-Lin-Shen ([KLS13]) on convergence of individual eigenvalues. Finally, we apply this framework to periodic homogenisation problems for dynamical $\Phi^4_3$ and KPZ equations. Assuming the convergence of stochastic objects to their homogenised limit, we show that the solutions and fluxes also converge to the corresponding limits. The convergence of the explicit stochastic terms with suitable renormalisations will be treated in a separate note.

math.AP