SearcharxivSearch

arXiv subjects

Jiachen Yu

Publications and source records attributed to Jiachen Yu.

At least 19 recordsLinked to original sources

FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention

Conventional LLMs keep the full KV cache loaded during decoding, causing a severe GPU memory bottleneck for ultra-long context serving. In this report, we propose \textbf{Lookahead Sparse Attention (LSA)}, a novel inference paradigm powered by a Neural Memory Indexer built upon the DeepSeek-V4 architecture. Rather than passively attending to all historical tokens, LSA proactively predicts future context demands and preserves only the query-critical KV chunks in the GPU memory. Crucially, we instantiate this architecture via a \textbf{backbone-free decoupled training} strategy. By formulating the indexer as a standard dual-encoder architecture, we train it independently using standard retrieval training frameworks without ever loading the massive backbone model into GPU memory. We demonstrate that this ``less is more'' paradigm significantly maximizes serving efficiency while acting as an effective attention denoiser in tasks that rely on long-term global memory. Across primary long-context evaluation suites (e.g., LongBench-v2, LongMemEval, and RULER), \texttt{FM-DS-V4} compresses the average physical KV cache footprint down to merely 13.5\% of the full-context baseline, while consistently preserving or slightly elevating downstream accuracy (+0.6\% absolute margin on average). At 1M context, per-decode-token compute drops to 0.30$\times$ of the baseline and GPU KV cache shrinks by 90\% (3.73$\to$0.37 GB), translating into \textbf{2.8$\times$ aggregate throughput and 2.7$\times$ concurrency gains} in PD-disaggregated serving on 8$\times$H20 GPUs.

cs.LG

Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning

Reinforcement learning from verifiable rewards improves the reasoning ability of large language models, but often suffers from entropy collapse, in which increasingly concentrated policies reduce rollout diversity and useful learning signals. Existing remedies either constrain the RL objective (e.g., entropy regularization) or adjust sampling temperature during rollout collection, but these interventions remain external to the model parameters. We propose Temperature-Scaled On-Policy Self-Distillation (TS-OPSD), a lightweight policy reheating method that internalizes the exploratory effect of temperature into model parameters. Starting from an entropy-collapsed RL checkpoint, TS-OPSD constructs a self-teacher by applying high-temperature scaling to the model's own logits, then distills the resulting smoother distribution back into the student. This policy reheating requires no external teacher, privileged data, or additional inference cost. Experiments on Qwen3-4B-Base and Qwen3-8B-Base show that policy reheating yields a stronger initialization for continued RL than both standard continued RL and rollout-level temperature reheating. Further analyses show that TS-OPSD mainly reduces output sharpness while preserving intermediate representations, top candidate sets, and reasoning capability. These results suggest that entropy restoration can serve as a simple post-collapse intervention for extending reasoning-oriented RL.

cs.CL

Think-with-Rubrics: From External Evaluator to Internal Reasoning Guidance

Rubrics have been extensively utilized for evaluating unverifiable, open-ended tasks, with recent research incorporating them into reward systems for reinforcement learning. However, existing frameworks typically treat rubrics only as external evaluator disjointed from the policy's primary reasoning trace. Such design confines rubrics to post-hoc measurement, leaving them unable to actively guide the model's generation process. In this work, we introduce Think-with-Rubrics, a novel paradigm for instruction following tasks. Think-with-Rubrics integrates rubric generation into the reasoning context, transforming the rubric from an independent artifact into an internal guidance of LLM's generation. During training, LLM sequentially generates a rubric followed by a response, while a trained rubric verifier provides joint supervision by evaluating the consistency between the answer and the self-generated / golden rubrics. Experiments across multiple benchmarks demonstrate that Think-with-Rubrics consistently outperforms the Rubric-as-Reward baseline supervised by golden rubrics by an average of 3.87 points. We have also discussed the mechanism by which Think-with-Rubrics enhances model performance. Experimental results demonstrate that supervision from golden rubrics and self-generated rubrics enhances the performance of Think-with-Rubrics by improving the quality of self-generated rubrics and increasing the internal consistency of responses respectively.

cs.CL

Projection of purification performance for the RELICS experiment

The RELICS (REactor neutrino LIquid xenon Coherent elastic Scattering) experiment employs a dual-phase liquid xenon time projection chamber to search for Coherent Elastic Neutrino-Nucleus Scattering (CE$\nu$NS) induced by reactor neutrinos. To detect these sub-keV nuclear recoils and minimize signal attenuation, it is critical to maintain a sufficiently low impurity concentration in the detector. This work presents a comprehensive purity evolution model developed to describe impurity migration inside the detector. Utilizing measured material outgassing rates as input parameters, the model incorporates non-uniform transport mechanisms of the impurities, including circulation, vaporization, and condensation. The model is validated using data from a dedicated prototype detector. Based on this validated model, projections for the purification performance of the upcoming RELICS-10 and RELICS-50 detectors are provided.

physics.ins-det

CoWork-X: Experience-Optimized Co-Evolution for Multi-Agent Collaboration System

Large language models are enabling language-conditioned agents in interactive environments, but highly cooperative tasks often impose two simultaneous constraints: sub-second real-time coordination and sustained multi-episode adaptation under a strict online token budget. Existing approaches either rely on frequent in-episode reasoning that induces latency and timing jitter, or deliver post-episode improvements through unstructured text that is difficult to compile into reliable low-cost execution. We propose CoWork-X, an active co-evolution framework that casts peer collaboration as a closed-loop optimization problem across episodes, inspired by fast--slow memory separation. CoWork-X instantiates a Skill-Agent that executes via HTN (hierarchical task network)-based skill retrieval from a structured, interpretable, and compositional skill library, and a post-episode Co-Optimizer that performs patch-style skill consolidation with explicit budget constraints and drift regularization. Experiments in challenging Overcooked-AI-like realtime collaboration benchmarks demonstrate that CoWork-X achieves stable, cumulative performance gains while steadily reducing online latency and token usage.

cs.CL

Development of a dual-phase xenon time projection chamber prototype for the RELICS experiment

The RELICS (REactor neutrino LIquid xenon Coherent elastic Scattering) experiment aims to detect coherent elastic neutrino-nucleus scattering from reactor antineutrinos using a dual-phase xenon time projection chamber. To validate the detector concept and ensure technical reliability for the full-scale experiment, a dedicated prototype was designed, constructed, and operated. This work presents an overview of the design, construction, and operational performance of the prototype, with emphasis on its major subsystems, including the TPC, cryogenic and xenon purification systems, slow control, and data acquisition. During operation, the detector demonstrated the capability to achieve a sub-keV energy threshold required for the RELICS physics program, as reflected by a measured single electron gain of 34.30~$\pm$~0.01~(stat.)~PE/e$^-$ and the successful detection of 0.27~keV L-shell decay events from $^{37}$Ar. In addition, essential data analysis techniques and simulation frameworks were developed and validated, establishing the methodological foundation for future RELICS operations. The successful construction and operation of this prototype confirm the feasibility of the core technologies and provide a crucial experimental basis for the final RELICS detector.

physics.ins-det

Visualizing interaction-driven restructuring of quantum Hall edge states

Many topological phases host gapless boundary modes that can be dramatically modified by electronic interactions. Even for the long-studied edge modes of quantum Hall phases, forming at the boundaries of two-dimensional (2D) electron systems, the nature of such interaction-induced changes has been elusive. Despite advances made using local probes, key experimental challenges persist: the lack of direct information about the internal structure of edge states on microscopic scales, and complications from edge disorder. Here, we use scanning tunneling microscopy (STM) to image pristine electrostatically defined quantum Hall edge states in graphene with high spatial resolution and demonstrate how correlations dictate the structures of edge channels on both magnetic and atomic length scales. For integer quantum Hall states in the zeroth Landau level, we show that interactions renormalize the edge velocity, dictate the spatial profile for copropagating modes, and induce unexpected edge valley polarization that differ from those of the bulk. While some of our findings can be understood by mean-field theory, others show breakdown of this picture, highlighting the roles of edge fluctuations and inter-channel couplings. We also extend our measurements to spatially resolve the edge state of fractional quantum Hall phases and detect spectroscopic signatures of interactions in this chiral Luttinger liquid. Our study establishes STM as a promising tool for exploring edge physics of the rapidly expanding 2D topological phases, including newly realized fractional Chern insulators.

cond-mat.mes-hall

Design and characterization of a photosensor system for the RELICS experiment

In this paper, we present the design and characterization of a photosensor system developed for the RELICS experiment. An extended dynamic range base was designed to mitigate photomultiplier tube (PMT) saturation caused by intense cosmic muon backgrounds in the surface-level RELICS detector. The system employs dual readout from the anode and the seventh dynode to extend the linear response range of the PMT. In particular, our characterization and measurements of Hamamatsu R8520-406 PMTs confirm stable operation under positive high-voltage bias, extending the linear response range by more than an order of magnitude. Furthermore, a model of PMT saturation and recovery was developed to evaluate the influence of cosmic muon signals in the RELICS detector. The results demonstrate the system capability to detect coherent elastic neutrino-nucleus scattering signals under surface-level cosmic backgrounds, and suggest the potential to extend the scientific reach of RELICS to MeV-scale interactions.

physics.ins-det

S2J: Bridging the Gap Between Solving and Judging Ability in Generative Reward Models

With the rapid development of large language models (LLMs), generative reward models (GRMs) have been widely adopted for reward modeling and evaluation. Previous studies have primarily focused on training specialized GRMs by optimizing them on preference datasets with the judgment correctness as supervision. While it's widely accepted that GRMs with stronger problem-solving capabilities typically exhibit superior judgment abilities, we first identify a significant solve-to-judge gap when examining individual queries. Specifically, the solve-to-judge gap refers to the phenomenon where GRMs struggle to make correct judgments on some queries (14%-37%), despite being fully capable of solving them. In this paper, we propose the Solve-to-Judge (S2J) approach to address this problem. Specifically, S2J simultaneously leverages both the solving and judging capabilities on a single GRM's output for supervision, explicitly linking the GRM's problem-solving and evaluation abilities during model optimization, thereby narrowing the gap. Our comprehensive experiments demonstrate that S2J effectively reduces the solve-to-judge gap by 16.2%, thereby enhancing the model's judgment performance by 5.8%. Notably, S2J achieves state-of-the-art (SOTA) performance among GRMs built on the same base model while utilizing a significantly smaller training dataset. Moreover, S2J accomplishes this through self-evolution without relying on more powerful external models for distillation.

cs.CL

Quench spectroscopy for Lieb-Liniger bosons in the presence of harmonic trap

Quench spectroscopy has emerged as a novel and powerful technique for probing the energy spectrum of various quantum phases for quantum systems from out-of-equilibrium dynamics. While its efficacy has been demonstrated in the homogeneous systems theoretically, most experimental setups feature a confining potential, such as a harmonic trap, which complicates the practical implementations. In this work, we experimentally probe the quench spectroscopy for one-dimensional bosons in optical lattices with the presence of a harmonic trap, and comparing our results with the density matrix renormalization group simulation. For the Mott insulator phase, although a gap is still observed, the band signal is broadened along the frequency space and cut at the half Brillouin zone, which can be explained by the nearest-neighbor tunneling excitations under harmonic confinement. Comparing with the superfluid spectrum, we can see a clear distinction between the two phases and find the inverse quench with larger amplitude yields the clearest spectrum. Our work offers pivotal insights into conducting quench spectroscopy effectively in practical systems.

cond-mat.quant-gas

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories?

Recent extensive works have demonstrated that by introducing long CoT, the capabilities of MLLMs to solve complex problems can be effectively enhanced. However, the reasons for the effectiveness of such paradigms remain unclear. It is challenging to analysis with quantitative results how much the model's specific extraction of visual cues and its subsequent so-called reasoning during inference process contribute to the performance improvements. Therefore, evaluating the faithfulness of MLLMs' reasoning to visual information is crucial. To address this issue, we first present a cue-driven automatic and controllable editing pipeline with the help of GPT-Image-1. It enables the automatic and precise editing of specific visual cues based on the instruction. Furthermore, we introduce VFaith-Bench, the first benchmark to evaluate MLLMs' visual reasoning capabilities and analyze the source of such capabilities with an emphasis on the visual faithfulness. Using the designed pipeline, we constructed comparative question-answer pairs by altering the visual cues in images that are crucial for solving the original reasoning problem, thereby changing the question's answer. By testing similar questions with images that have different details, the average accuracy reflects the model's visual reasoning ability, while the difference in accuracy before and after editing the test set images effectively reveals the relationship between the model's reasoning ability and visual perception. We further designed specific metrics to expose this relationship. VFaith-Bench includes 755 entries divided into five distinct subsets, along with an additional human-labeled perception task. We conducted in-depth testing and analysis of existing mainstream flagship models and prominent open-source model series/reasoning models on VFaith-Bench, further investigating the underlying factors of their reasoning capabilities.

cs.CV

Improve LLM-as-a-Judge Ability as a General Ability

LLM-as-a-Judge leverages the generative and reasoning capabilities of large language models (LLMs) to evaluate LLM responses across diverse scenarios, providing accurate preference signals. This approach plays a vital role in aligning LLMs with human values, ensuring ethical and reliable AI outputs that align with societal norms. Recent studies have raised many methods to train LLM as generative judges, but most of them are data consuming or lack accuracy, and only focus on LLM's judge ability. In this work, we regard judge ability as a general ability of LLM and implement a two-stage training approach, comprising supervised fine-tuning (SFT) warm-up and direct preference optimization (DPO) enhancement, to achieve judge style adaptation and improve judgment accuracy. Additionally, we introduce an efficient data synthesis method to generate judgmental content. Experimental results demonstrate that our approach, utilizing only about 2% to 40% of the data required by other methods, achieves SOTA performance on RewardBench. Furthermore, our training method enhances the general capabilities of the model by constructing complicated judge task, and the judge signals provided by our model have significantly enhanced the downstream DPO training performance of our internal models in our test to optimize policy model with Judge Model. We also open-source our model weights and training data to facilitate further research.

cs.CL

A Deep Dive Into Large Language Model Code Generation Mistakes: What and Why?

Recent advancements in Large Language Models (LLMs) have led to their widespread application in automated code generation. However, these models can still generate defective code that deviates from the specification. Previous research has mainly focused on the mistakes in LLM-generated standalone functions, overlooking real-world software development situations where the successful generation of the code requires software contexts such as external dependencies. In this paper, we considered both of these code generation situations and identified a range of \textit{non-syntactic mistakes} arising from LLMs' misunderstandings of coding question specifications. Seven categories of non-syntactic mistakes were identified through extensive manual analyses, four of which were missed by previous works. To better understand these mistakes, we proposed six reasons behind these mistakes from various perspectives. Moreover, we explored the effectiveness of LLMs in detecting mistakes and their reasons. Our evaluation demonstrated that GPT-4 with the ReAct prompting technique can achieve an F1 score of up to 0.65 when identifying reasons for LLM's mistakes, such as misleading function signatures. We believe that these findings offer valuable insights into enhancing the quality of LLM-generated code.

cs.SE

Reactor neutrino liquid xenon coherent elastic scattering experiment

Coherent elastic neutrino-nucleus scattering (CEvNS) provides a unique probe for neutrino properties Beyond the Standard Model (BSM) physics. REactor neutrino LIquid xenon Coherent Scattering experiment (RELICS), a proposed reactor neutrino program using liquid xenon time projection chamber (LXeTPC) technology, aims to investigate the CEvNS process of antineutrinos off xenon atomic nuclei. In this work, the design of the experiment is studied and optimized based on Monte Carlo (MC) simulations. To achieve a sufficiently low energy threshold for CEvNS detection, an ionization-only analysis channel is adopted for RELICS. A high emission rate of delayed electrons after a big ionization signal is the major background, leading to an analysis threshold of 120 photo-electrons in the CEvNS search. The second largest background, nuclear recoils induced by cosmic-ray neutrons, is suppressed via a passive water shield. The physics potential of RELICS is explored with a 32 kg*yr exposure at a baseline of 25 m from a reactor core with a 3 GW thermal power. In an energy range of 120 to 300 PE, corresponding to an average nuclear recoil from 0.63 to 1.36 keV considering the liquid xenon response and detector-related effect, we expect 4639.7 CEvNS and 1687.8 background events. The sensitivity of RELICS to the weak mixing angle is investigated at a low momentum transfer. Our study shows that RELICS can further improve the constraints on the non-standard neutrino interaction (NSI) compared to the current best results.

hep-ex

Characterization of two fast-turnaround dry dilution refrigerators for scanning probe microscopy

Low-temperature scanning probe microscopes (SPMs) are critical for the study of quantum materials and quantum information science. Due to the rising costs of helium, cryogen-free cryostats have become increasingly desirable. However, they typically suffer from comparatively worse vibrations than cryogen-based systems, necessitating the understanding and mitigation of vibrations for SPM applications. Here we demonstrate the construction of two cryogen-free dilution refrigerator SPMs with minimal modifications to the factory default and we systematically characterize their vibrational performance. We measure the absolute vibrations at the microscope stage with geophones, and use both microwave impedance microscopy and a scanning single electron transistor to independently measure tip-sample vibrations. Additionally, we implement customized filtering and thermal anchoring schemes, and characterize the cooling power at the scanning stage and the tip electron temperature. This work serves as a reference to researchers interested in cryogen-free SPMs, as such characterization is not standardized in the literature or available from manufacturers.

cond-mat.mes-hall

Harnessing excitons at the nanoscale -- photoelectrical platform for quantitative sensing and imaging

Excitons -- quasiparticles formed by the binding of an electron and a hole through electrostatic attraction -- hold promise in the fields of quantum light confinement and optoelectronic sensing. Atomically thin transition metal dichalcogenides (TMDs) provide a versatile platform for hosting and manipulating excitons, given their robust Coulomb interactions and exceptional sensitivity to dielectric environments. In this study, we introduce a cryogenic scanning probe photoelectrical sensing platform, termed exciton-resonant microwave impedance microscopy (ER-MIM). ER-MIM enables ultra-sensitive probing of exciton polarons and their Rydberg states at the nanoscale. Utilizing this technique, we explore the interplay between excitons and material properties, including carrier density, in-plane electric field, and dielectric screening. Furthermore, we employ deep learning for automated data analysis and quantitative extraction of electrical information, unveiling the potential of exciton-assisted nano-electrometry. Our findings establish an invaluable sensing platform and readout mechanism, advancing our understanding of exciton excitations and their applications in the quantum realm.

cond-mat.mes-hall

Hofstadter states and reentrant charge order in a semiconductor moir\'e lattice

The emergence of moir\'e materials with flat bands provides a platform to systematically investigate and precisely control correlated electronic phases. Here, we report local electronic compressibility measurements of a twisted WSe$_2$/MoSe$_2$ heterobilayer which reveal a rich phase diagram of interpenetrating Hofstadter states and electron solids. We show that this reflects the presence of both flat and dispersive moir\'e bands whose relative energies, and therefore occupations, are tuned by density and magnetic field. At low densities, competition between moir\'e bands leads to a transition from commensurate arrangements of singlets at doubly occupied sites to triplet configurations at high fields. Hofstadter states (i.e., Chern insulators) are generally favored at high densities as dispersive bands are populated, but are suppressed by an intervening region of reentrant charge-ordered states in which holes originating from multiple bands cooperatively crystallize. Our results reveal the key microscopic ingredients that favor distinct correlated ground states in semiconductor moir\'e systems, and they demonstrate an emergent lattice model system in which both interactions and band dispersion can be experimentally controlled.

cond-mat.mes-hall

Spin skyrmion gaps as signatures of strong-coupling insulators in magic-angle twisted bilayer graphene

The flat electronic bands in magic-angle twisted bilayer graphene (MATBG) host a variety of correlated insulating ground states, many of which are predicted to support charged excitations with topologically non-trivial spin and/or valley skyrmion textures. However, it has remained challenging to experimentally address their ground state order and excitations, both because some of the proposed states do not couple directly to experimental probes, and because they are highly sensitive to spatial inhomogeneities in real samples. Here, using a scanning single-electron transistor, we observe thermodynamic gaps at even integer moir\'e filling factors at low magnetic fields. We find evidence of a field-tuned crossover from charged spin skyrmions to bare particle-like excitations, suggesting that the underlying ground state belongs to the manifold of strong-coupling insulators. From the spatial dependence of these states and the chemical potential variation within the flat bands, we infer a link between the stability of the correlated ground states and local twist angle and strain. Our work advances the microscopic understanding of the correlated insulators in MATBG and their unconventional excitations.

cond-mat.mes-hall