SearcharxivSearch

arXiv subjects

Zhaoyi Li

Publications and source records attributed to Zhaoyi Li.

At least 19 recordsLinked to original sources

Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models

On-policy distillation (OPD) transfers teacher capabilities by supervising trajectories sampled from the student's own policy, yet its generalization behavior remains poorly understood, as most studies evaluate OPD on a single domain and on benchmarks close to the training data. We present a controlled study that varies one generalization factor at a time, from in-domain distribution shifts to cross-domain transfer and the multi-teacher setting. We find that OPD transfers a teacher's reasoning behavior rather than its answers to particular problems: training difficulty barely matters, and even problems the teacher never solves are useful. Transfer depends strongly on the origin relationship between teacher and student: same-origin pairs bring the student close to the teacher across languages, reasoning horizons, and even other domains, whereas cross-origin pairs mostly fit the trained distribution. This broad reach is a double-edged sword: since routing prompts to domain experts cannot confine each teacher's influence, combining them yields a mixture-dependent seesaw among their capabilities. These results clarify when OPD generalizes and offer a useful perspective for diagnosing multi-teacher OPD.

cs.CL

Dolomite Mineral-Inspired Equilateral Triangular-Lattice Magnets for Quantum Magnetism

Equilateral triangular lattice magnets provide a versatile materials platform for exploring exotic quantum spin phenomena, while their field-tunable magnetic entropy offers opportunities for low-temperature adiabatic demagnetization refrigeration. Inspired by the natural mineral, we proposed a chemical strategy to achieve equilateral TL magnets, leveraging the high crystal symmetry of a large family of dolomite-type materials. As typical examples, the dolomite-type materials SnM(BO3)2 (M = Co, Mn) were synthesized, and structural analysis reveals that Co2+ and Mn2+ ions form equilateral triangular lattices with an A-B-C stacking fashion. The magnetic susceptibilities and specific heat measurements reveal dominant antiferromagnetic interactions, with Neel temperatures of 0.49K for SnCo(BO3)2 and 0.96K for SnMn(BO3)2, respectively. Our results establish the dolomite-type M'M(X)2 (M'and M sites allow various valence states, e.g., +4/+2 or +3/+3; X = CO32- or BO33-) system as a chemically flexible and structurally perfect material platform for exploring frustrated magnetism and low-temperature magnetocaloric applications.

cond-mat.mtrl-sci

Scaling-optimal purification of noisy qubit unitary channels

We consider the problem of purifying noisy qubit unitary channels. Given the ability to apply an unknown qubit unitary channel followed by depolarizing noise, we aim to construct a superchannel that purifies the noisy unitary back to the original unknown unitary. We first provide numerical evidence that sequential strategies can strictly outperform parallel strategies when the number of channel uses is finite, highlighting the fundamental distinction from state purification. We then provide a concrete $\mathrm{U}(2)$-covariant parallel protocol based on a novel entanglement-assisted quantum error-correcting code that suppresses the first-order noise strength as $O(1/n)$ with $n$ channel uses and show this scaling is asymptotically optimal in the low-noise regime, even when sequential strategies are allowed.

quant-ph

Field-Induced Up-Up-Down State and Frustrated Magnetism in a Non-Kramers Triangular Antiferromagnet

A previously unreported triangular lattice (TL) antiferromagnet, TmZnGaO4, was synthesized as single crystals, and its crystal structure, magnetic susceptibilities, and specific heat were reported. Its crystal structure is isomorphic to that of the transverse-field Ising antiferromagnet TmMgGaO4, with Tm3+ ions located in the TLs, separated by a nonmagnetic bilayer composed mainly of Ga3+ and Zn2+ ions. The magnetic susceptibilities indicate the dominating antiferromagnetic interactions. The magnetization curves (M-H) exhibit strong easy-c-axis anisotropy, with a clear one-third magnetic plateau emerging, consistent with a field-induced up-up-down spin configuration. Instead of forming a conventional long-range magnetic order, the system exhibits two broad anomalies at 0.11 K and 2.81 K in zero-field specific heat measurements, highlighting the persistence of strong spin fluctuations and the potential for exotic quantum spin states. The above results reveal its future interest in exploring exotic quantum spin states in TmZnGaO4.

cond-mat.str-el

LLMSurgeon: Diagnosing Data Mixture of Large Language Models

The pretraining data mixture of Large Language Models (LLMs) constitutes their "digital DNA", shaping model behaviors, capabilities, and failure modes. Yet this composition is rarely disclosed, making post-hoc auditing of data combination or provenance difficult. In this work, we formalize $\textbf{{Data Mixture Surgery (DMS)}}$: given only generated text from a target LLM, estimate the domain-level distribution of its pretraining corpus under a predefined taxonomy. We propose $\textbf{{LLMSurgeon}}$, a strong framework that casts DMS as an inverse problem under the label-shift assumption. Rather than directly aggregating classifier outputs, LLMSurgeon estimates a calibrated $\textit{soft}$ confusion matrix and solves a constrained inverse problem to correct systematic domain confusion and recover the latent mixture prior. To evaluate, we introduce $\textbf{{LLMScan}}$, a recipe-verifiable evaluation suite built from open-source LLMs with transparent pretraining mixtures. Across LLMScan, LLMSurgeon recovers domain mixtures with high fidelity under fixed protocols. Our work presents a practical, post-hoc approach for auditing the digital DNA of foundation models without access to their training data.

cs.CL

An Exponential Sample-Complexity Advantage for Coherent Quantum Inference

Standard quantum inference converts quantum data into classical outputs. We study an alternative inference setting in which the desired output is quantum, preserving coherence. Such settings include quantum purity amplification (QPA), mixed-state approximate purification or cloning, and density matrix exponentiation. We show that such protocols can achieve exponentially lower sample complexity than incoherent, measurement-mediated protocols. For QPA with principal eigenstate targets and $d$-dimensional inputs, coherent processing achieves error $\varepsilon$ using $O(1/\varepsilon)$ copies, versus the $\Omega(d/\varepsilon)$ copies required by any incoherent protocol. Together, these sharp coherent-incoherent separations seed a theory of coherent quantum inference, with an entanglement-breaking limit identifying the optimal incoherent counterpart of each coherent protocol.

quant-ph

Quantum Purity Amplification for Arbitrary Eigenstates and Multiple Outputs

Quantum purity amplification (QPA) is the task of coherently transforming $n$ copies of a mixed state into high-fidelity copies of a chosen eigenstate. We solve QPA in the general setting of $n$ input copies, $m$ output copies, arbitrary target eigenstates, arbitrary local dimension $d$, and generic input spectra. We characterize the optimal channel and derive its all-site and one-site performance laws across output regimes. For the asymptotic analysis, we use a path-graph parametrization to show that, when the target eigenvalue has a constant spectral gap $D_{k,\mathrm{min}}$, achieving all-site error $\varepsilon$ requires a number of input copies independent of $d$ and scaling as $O(m/(\varepsilon D_{k,\mathrm{min}}^2))$. When $m/n$ approaches a constant, the performance exhibits phase-like regimes, which we characterize explicitly. For the nonasymptotic analysis, we develop a theory of generalized Young diagrams that yields tight sample complexity bounds and provides the first dimension-uniform guarantee for optimal QPA. We also provide asymptotically efficient implementations of the optimal protocol. Together, these results establish QPA as a rigorous example of coherent quantum information processing with dimension-uniform sample complexity, supplying the technical foundation for the coherent-incoherent separation developed in the companion work.

quant-ph

On the Role of Reasoning Patterns in the Generalization Discrepancy of Long Chain-of-Thought Supervised Fine-Tuning

Supervised Fine-Tuning (SFT) on long Chain-of-Thought (CoT) trajectories has become a pivotal phase in building large reasoning models. However, how CoT trajectories from different sources influence the generalization performance of models remains an open question. In this paper, we conduct a comparative study using two sources of verified CoT trajectories generated by two competing models, \texttt{DeepSeek-R1-0528} and \texttt{gpt-oss-120b}, with their problem sets controlled to be identical. Despite their comparable performance, we uncover a striking paradox: lower training loss does not translate to better generalization. SFT on \texttt{DeepSeek-R1-0528} data achieves remarkably lower training loss, yet exhibits significantly worse generalization performance on reasoning benchmarks compared to those trained on \texttt{gpt-oss-120b}. To understand this paradox, we perform a multi-faceted analysis probing token-level SFT loss and step-level reasoning behaviors. Our analysis reveals a difference in reasoning patterns. \texttt{gpt-oss-120b} exhibits highly convergent and deductive trajectories, whereas \texttt{DeepSeek-R1-0528} favors a divergent and branch-heavy exploration pattern. Consequently, models trained with \texttt{DeepSeek-R1} data inherit inefficient exploration behaviors, often getting trapped in redundant exploratory branches that hinder them from reaching correct solutions. Building upon this insight, we propose a simple yet effective remedy of filtering out frequently branching trajectories to improve the generalization of SFT. Experiments show that training on selected \texttt{DeepSeek-R1-0528} subsets surprisingly improves reasoning performance by up to 5.1% on AIME25, 5.5% on BeyondAIME, and on average 3.6% on five benchmarks.

cs.CL

BiT-MCTS: A Theme-based Bidirectional MCTS Approach to Chinese Fiction Generation

Generating long-form linear fiction from open-ended themes remains a major challenge for large language models, which frequently fail to guarantee global structure and narrative diversity when using premise-based or linear outlining approaches. We present BiT-MCTS, a theme-driven framework that operationalizes a "climax-first, bidirectional expansion" strategy motivated by Freytag's Pyramid. Given a theme, our method extracts a core dramatic conflict and generates an explicit climax, then employs a bidirectional Monte Carlo Tree Search (MCTS) to expand the plot backward (rising action, exposition) and forward (falling action, resolution) to produce a structured outline. A final generation stage realizes a complete narrative from the refined outline. We construct a Chinese theme corpus for evaluation and conduct extensive experiments across three contemporary LLM backbones. Results show that BiT-MCTS improves narrative coherence, plot structure, and thematic depth relative to strong baselines, while enabling substantially longer, more coherent stories according to automatic metrics and human judgments.

cs.CL

A Kagome-Derived Mosaic Lattice Family A3V9Te13 (A = Cs, Rb) with Tunable Strong Electronic Correlations

The pursuit of geometrically frustrated lattices beyond conventional paradigms remains a central challenge in the design of quantum materials. Herein, we report the discovery of the A3V9Te13 (A = Cs, Rb) family of vanadium-based intermetallic compounds, which host a unique two-dimensional Mosaic lattice derived from the Kagome network, composed of an ordered tessellation of triangles, squares, and pentagons. The Cs compound (CVT) exhibits strong electronic correlations, characterized by non-Fermi liquid behavior at low temperatures, an exceptionally large Sommerfeld coefficient, and a bulk phase transition at T* $\approx$ 47 K with possible charge- or spin-related origin. Inspired by pressure-tuning in related Kagome systems, we demonstrate that the electronic ground state of this lattice is exquisitely tunable via chemical pressure. Systematic substitution of Cs with smaller Rb ions suppresses the T* phase transition and the correlated electronic response, ultimately driving the system into a highly frustrated semiconducting ground state without long-range magnetic order down to 60 mK. This work unveils a new structural platform for exploring the interplay between geometric frustration and strong electron correlations, providing a chemically controllable platform for exploring the phase space between distinct correlated electronic states.

cond-mat.mtrl-sci

HieraMAS: Optimizing Intra-Node LLM Mixtures and Inter-Node Topology for Multi-Agent Systems

Multi-agent systems (MAS) built on large language models (LLMs) have shown strong performance across many tasks. Most existing approaches improve only one aspect at a time, such as the communication topology, role assignment, or LLM routing, while treating each agent as a single, indivisible unit. This misses the opportunity to use mixtures of LLMs within an agent to strengthen role-specific abilities. We propose HieraMAS, a hierarchical collaboration framework that combines intra-node LLM mixtures with an inter-node communication topology. HieraMAS introduces supernodes, where each functional role is implemented by multiple heterogeneous LLMs using a propose-synthesis structure. Optimizing HieraMAS creates unique credit-assignment challenges: final task performance depends heavily on the underlying LLMs' capabilities, which can lead reinforcement methods to incorrectly reward suboptimal configurations. To address this, we use a two-stage algorithm: (1) multi-level reward attribution, which provides fine-grained feedback at both the node level and the overall system level; (2) graph classification for topology selection, which treats choosing the communication structure as a holistic decision rather than optimizing edges one by one. Experiments on reasoning and coding benchmarks show that HieraMAS substantially outperforms existing methods while also delivering better cost-performance trade-offs.

cs.MA

Pushing the Frontier of Black-Box LVLM Attacks via Fine-Grained Detail Targeting

Black-box adversarial attacks on Large Vision-Language Models (LVLMs) are challenging due to missing gradients and complex multimodal boundaries. While prior state-of-the-art transfer-based approaches like M-Attack perform well using local crop-level matching between source and target images, we find this induces high-variance, nearly orthogonal gradients across iterations, violating coherent local alignment and destabilizing optimization. We attribute this to (i) ViT translation sensitivity that yields spike-like gradients and (ii) structural asymmetry between source and target crops. We reformulate local matching as an asymmetric expectation over source transformations and target semantics, and build a gradient-denoising upgrade to M-Attack. On the source side, Multi-Crop Alignment (MCA) averages gradients from multiple independently sampled local views per iteration to reduce variance. On the target side, Auxiliary Target Alignment (ATA) replaces aggressive target augmentation with a small auxiliary set from a semantically correlated distribution, producing a smoother, lower-variance target manifold. We further reinterpret momentum as Patch Momentum, replaying historical crop gradients; combined with a refined patch-size ensemble (PE+), this strengthens transferable directions. Together these modules form M-Attack-V2, a simple, modular enhancement over M-Attack that substantially improves transfer-based black-box attacks on frontier LVLMs: boosting success rates on Claude-4.0 from 8% to 30%, Gemini-2.5-Pro from 83% to 97%, and GPT-5 from 98% to 100%, outperforming prior black-box LVLM attacks. Code and data are publicly available at: https://github.com/vila-lab/M-Attack-V2.

cs.LG

Scaling Reasoning Hop Exposes Weaknesses: Demystifying and Improving Hop Generalization in Large Language Models

Chain-of-thought (CoT) reasoning has become the standard paradigm for enabling Large Language Models (LLMs) to solve complex problems. However, recent studies reveal a sharp performance drop in reasoning hop generalization scenarios, where the required number of reasoning steps exceeds training distributions while the underlying algorithm remains unchanged. The internal mechanisms driving this failure remain poorly understood. In this work, we conduct a systematic study on tasks from multiple domains, and find that errors concentrate at token positions of a few critical error types, rather than being uniformly distributed. Closer inspection reveals that these token-level erroneous predictions stem from internal competition mechanisms: certain attention heads, termed erroneous processing heads (ep heads), tip the balance by amplifying incorrect reasoning trajectories while suppressing correct ones. Notably, removing individual ep heads during inference can often restore the correct predictions. Motivated by these insights, we propose test-time correction of reasoning, a lightweight intervention method that dynamically identifies and deactivates ep heads in the reasoning process. Extensive experiments across different tasks and LLMs show that it consistently improves reasoning hop generalization, highlighting both its effectiveness and potential.

cs.CL

Wide field-of-view and large depth-of-field metalenses

The ability to visualize both macroscopic and microscopic features over an extended field of view is essential for endoscopic imaging and other applications ranging from machine vision to microscopy. However, miniaturizing endoscopes introduces inherent trade-offs between size and optical performance, including field-of-view (FOV), depth-of-field (DOF), and resolution. These constraints limit the use of microendoscopes in clinical settings such as early cancer detection within narrow, hard-to-access anatomical regions, including the lung, ovaries, and pancreas. State-of-the-art microendoscopes typically rely on microlens assemblies that increase both cost and size. Their large f-numbers also hinder the collection of high-resolution information from live tissue. In this work, we present two compact metalens designs that provide wide FOV, extended DOF, and high resolution, enabled by custom-tailored point spread functions (PSFs). The devices achieve a full 172 degrees FOV, an extended DOF from 0.4 mm to beyond 300 mm, and a resolution of 30 line pairs per millimeter, all within a 1 mm x 1 mm x 0.2 mm footprint. A key advantage of our approach is the ability to transition seamlessly between low and high magnification without mechanical refocusing. Final images are reconstructed through backend deconvolution, highlighting the potential of hybrid imaging systems that integrate computational techniques with flat-optics components.

physics.optics

Wafer-scale conformal metasurface optics

Curved and conformal optics offer significant advantages by unlocking additional geometric degrees of freedom for optical design. These capabilities enable enhanced optical performance and are essential for meeting non-optical constraints, such as those imposed by ergonomics, aerodynamics, or wearability. However, existing fabrication techniques such as direct electron or laser beam writing on curved substrates, and soft-stamp-based transfer or nanoimprint lithography suffer from limitations in scalability, yield, geometry control, and alignment accuracy. Here, we present a scalable fabrication strategy for curved and conformal metasurface optics leveraging thermoforming, an industry-standard, high-throughput manufacturing process widely used for shaping thermoplastics. Our approach uniquely enables wafer-scale production of highly curved metasurface optics, achieving sub-millimeter radii of curvature and micron-level alignment precision. To guide the design and fabrication process, we developed a thermorheological model that accurately predicts and compensates for the large strains induced during thermoforming. This allows for precise control of metasurface geometry and preservation of optical function, yielding devices with diffraction-limited performance. As a demonstration, we implemented an artificial compound eye comprising freeform micro-metalens arrays. Compared to traditional micro-optical counterparts, the device exhibits an expanded field of view, reduced aberrations, and improved uniformity, highlighting the potential of thermoformed metasurfaces for next-generation optical systems.

physics.optics

Tunable Multistage Refrigeration via Geometrically Frustrated Triangular Lattice Antiferromagnet for Space Cooling

Low-temperature refrigeration technology constitutes a crucial component in space exploration. The small-scale, low-vibration Stirling-type pulse tube refrigerators hold significant application potential for space cooling. However, the efficient operation of current Stirling-type pulse tube cryocoolers in space cooling applications remains challenging due to the rapid decay of the heat capacity of regenerative materials below 10 K. This study adopts a novel material strategy: using a novel high-spin S = 7/2 magnetic regenerative material, Gd2O2Se, we construct a multistage tunable regenerative material structure to achieve an efficient cooling approach to the liquid helium temperature range. Under substantial geometric frustration from a double-layered triangular lattice, it exhibits two-step specific heat transition peaks at 6.22 K and 2.11 K, respectively. Its ultrahigh specific heat and broad two-step transition temperature range effectively bridge the gap between commercially used high-heat-capacity materials. Experimental verification shows that when Gd2O2Se is combined with Er3Ni and HoCu2 in the Stirling-type pulse tube cryocooler, the cooling efficiency of the pulse tube increases by 66.5 % at 7 K, and the minimum achievable temperature reaches 5.85 K. These results indicate that Gd2O2Se is an ideal magnetic regenerative material for space cooling

cond-mat.mtrl-sci

A Unified Frequency Domain Decomposition Framework for Interpretable and Robust Time Series Forecasting

Current approaches for time series forecasting, whether in the time or frequency domain, predominantly use deep learning models based on linear layers or transformers. They often encode time series data in a black-box manner and rely on trial-and-error optimization solely based on forecasting performance, leading to limited interpretability and theoretical understanding. Furthermore, the dynamics in data distribution over time and frequency domains pose a critical challenge to accurate forecasting. We propose FIRE, a unified frequency domain decomposition framework that provides a mathematical abstraction for diverse types of time series, so as to achieve interpretable and robust time series forecasting. FIRE introduces several key innovations: (i) independent modeling of amplitude and phase components, (ii) adaptive learning of weights of frequency basis components, (iii) a targeted loss function, and (iv) a novel training paradigm for sparse data. Extensive experiments demonstrate that FIRE consistently outperforms state-of-the-art models on long-term forecasting benchmarks, achieving superior predictive performance and significantly enhancing interpretability of time series

cs.LG

Fine-Tuning Flow Matching via Maximum Likelihood Estimation of Reconstructions

Flow Matching (FM) models achieve remarkable results in generative tasks. Building upon diffusion models, FM's simulation-free training paradigm enables simplicity and efficiency but introduces a train-inference gap: model outputs cannot be assessed during training. Moreover, the straight flow assumption suffers from some inherent limitations. To address this, we propose to fine-tune FM via Maximum Likelihood Estimation (MLE) of reconstructions -- enabled by FM's smooth ODE formulation, unlike the stochastic differential equations (SDEs) in diffusion models. We first theoretically analyze the relationship between training loss and inference error in FM under numerical precision constraints. We then propose an easy-to-implement fine-tuning framework based on MLE of reconstructions, with flexibility for sophisticated extensions. Building on this, we incorporate a generalized artificial viscosity term that enhances flow stability and robustness, accompanied by a direct parameterization method and rigorous theoretical guarantees. Experiments demonstrate our method's effectiveness across diverse settings: a toy example provides mechanistic insights into the fine-tuning process, while large-scale evaluations on meteorological forecasting and robotic manipulation policies validate reliable performance improvements.

cs.LG