SearcharxivSearch

arXiv subjects

Yue Yu

Publications and source records attributed to Yue Yu.

At least 55 records · Page 3Linked to original sources

When Rough Data Helps: A Phase Transition in Convergence Rates for Kernel Recovery in Integral Operators

Learning kernels in operators from data is a fundamental task that arises in nonlocal continuum mechanics, operator learning, and interacting particle systems. A central question is how the roughness of input data impacts the accuracy of kernel recovery. We quantify the roughness of the input data via its spectral decay exponent and analyze how it determines the degree of ill-posedness of the inverse problem and, consequently, the convergence rates of the Tikhonov-regularized estimator in the small-noise limit. Within this framework, we identify a phase transition between an under-rough regime, in which rougher data improves recovery, and an over-rough regime, in which further roughening leads to slower rates. These theoretical findings are supported by numerical experiments ranging from idealized settings to more realistic configurations, with quantitative agreement in the former and broad consistency of the main trends in the latter.

math.NA

Scalable Circuit Learning for Interpreting Large Language Models

A prominent research direction in mechanistic interpretability is learning sparse circuits over LLM components to reveal how they jointly produce model behavior. However, raw neurons are polysemantic, making learned circuits hard to interpret. Sparse autoencoder (SAE) features alleviate this, but their high dimensionality makes existing intervention-based circuit learning methods computationally prohibitive. We propose CircuitLasso, a scalable circuit-learning approach based on sparse linear regression. CircuitLasso recovers circuits whose structural accuracy matches that of state-of-the-art intervention-based methods on the benchmark data, at a fraction of the computational cost. For interpretability, CircuitLasso efficiently uncovers relationships among SAE features, showing how human-interpretable semantic features propagate through the model and influence its predictions. Finally, we validate the utility of our learned circuits by leveraging their insights to achieve comparable performance at substantially lower cost on a domain-generalization task.

cs.LG

Adaptive Inference-Time Scaling via Early-Step Latent Verification for Image Editing

Instruction-based image editing has made notable progress with recent advances in generative models. However, the quality of the edited result is still influenced by the randomly sampled initial noise, particularly in complex editing scenarios. An unsuitable initial noise may lead to unsatisfactory editing results. Recent inference-time scaling methods address this issue by sampling multiple initial noises and selecting better candidates. Nevertheless, most of them follow a decode-then-verify scheme which introduces an efficiency-accuracy trade-off. When decoding is performed after limited inference steps, the decoded images often remain too noisy for reliable assessment, whereas sufficiently denoised images require much higher computational cost. To address this issue, we propose VeriLatent, a plug-and-play adaptive inference-time scaling framework with early-step latent verification for image editing. Specifically, we propose a novel verifier that scores each initial noise through a latent-space editing activation map at an early stage. It identifies promising candidates by assessing whether they can induce an effective edit in the correct region. This enables efficient early pruning without decoding latents into images. Building on this, we further develop an adaptive search strategy for inference-time scaling. It allocates inference budgets according to editing difficulty, thereby reducing the number of function evaluations (NFE). Extensive experiments on multiple benchmarks and different base models demonstrate that VeriLatent consistently improves both editing performance and inference-time scaling efficiency.

cs.CV

Bandedge-state-limited single-photon emission from volumetric quantum design of 2D colloidal quantum wells

Present-day solution-processable single-photon sources are dominated by three-dimensionally confined colloidal quantum dot emitters, yet their particle-to-particle variation in single-exciton properties limits reproducibility and scalability. Here, to avoid such heterogeneity, we demonstrate reliable room-temperature single-photon emission from atomically flat two-dimensional (2D) colloidal quantum wells (CQWs) with inherently uniform one-dimensional quantum confinement, despite their long-standing limitations of efficient multiexciton emission and pronounced exciton-surface susceptibility. We resolve these challenges through volumetric quantum design (VQD) of CQWs, yielding a highly localized, single bandedge state. This design laterally confines the bandedge excitonic domain within the exciton coherent area and vertically decouples it from surface states via a thick, strain-relieved quantum-barrier shell that preserves strong confinement, overcoming the daunting thickness-confinement trade-off in 2D CQWs. Statistical single-particle spectroscopy reveals that VQD-CQWs deliver near-blinking-free (on-time >99.5%) and fluence-insensitive antibunching (g(2)(0): 0.041), protected by a bandedge-state-filling bottleneck, together with linear polarization of up to 73% under cavity-free conditions, originating from synergistic transition-dipole and electric-field anisotropies. These advances establish 2D CQWs as a viable, homogenous and scalable platform for quantum technologies.

physics.app-ph

Convergence of parallel overlapping domain decomposition methods with impedance boundary conditions for time-harmonic Maxwell equations in heterogeneous media

This paper analyzes the convergence of parallel overlapping domain-decomposition methods with impedance boundary conditions for the time-harmonic Maxwell equations in heterogeneous media. We prove that the parallel iterative method is well-posed in an appropriate function space, and characterize the error propagation operator through impedance-to-impedance maps that describe interactions between neighboring subdomains. For strip domain decompositions, we derive explicit convergence estimates in terms of the norms of the impedance-to-impedance maps. At the discrete level, we develop the finite-element counterpart of these results based on Nédélec-element discretisations. Under the assumption that the discrete impedance-to-impedance maps approximate their continuous counterparts as the mesh is refined, we show that the discrete method inherits the convergence behavior of the continuous method. We illustrate this theory with numerical experiments for strip domain decompositions, and also present numerical experiments for checkerboard domain decompositions that go beyond our theory.

math.NA

18-dB on-chip vacuum squeezing in an adaptively poled lithium niobate waveguide

Quantum squeezed states of light can enhance measurement sensitivity beyond classical limits and enable quantum information processing, but scalable low-loss sources remain challenging. We demonstrate continuous-wave quantum squeezing on a chip, achieving 18 dB of squeezing and 20 dB of anti-squeezing at 1570 nm in a 1.6-cm traveling-wave adaptively poled thin-film lithium niobate waveguide. A distributed model independently determines facet losses, phase noise, and nonlinear interaction strength without prior assumptions, enabling rigorous inference of on-chip performance. We estimate a 95% confidence interval of [-18.96, -17.25] dB squeezing and [19.96, 21.35] dB anti-squeezing. These values represent the highest squeezing reported for any integrated photonic platform and the first assumption-free statistical validation of integrated squeezing performance. Our results establish thin-film lithium niobate as a high-performance, scalable platform for continuous-variable quantum sensing, communications, and photonic computing.

physics.optics

Self-Creative Text-to-Object Generation using Semantic-Aware Spatial Weighting

Instilling creativity in text-to-image (T2I) generation presents a significant challenge, as it requires synthesized images to exhibit not only visual novelty and surprise, but also artistic value. Current T2I models, however, are largely optimized for literal text-image alignment with their data distribution, and their noise prediction networks constrain the generation to high-probability regions, consequently generating outputs that lack authentic creativity. To address this, we propose a Self-Creative Diffusion (SCDiff) model for meaningful T2I generations featuring two core modules: a learnable spatial weighting (LSW) module and a visual-semantic mixing loss (VSML). The LSW module designs a parametric Kaiser-Bessel window to reinforce central image features, fostering novel and surprising generation. The VSML module introduces a dual loss function: a similarity loss constrains that the new images align with its textual description, while a diversity loss maximizes its distinction from the original image, enhancing both semantic value and visual novelty. Extensive experiments demonstrate that our model substantially improves creativity, semantic alignment, and visual coherence, offering a simple yet powerful framework for generating creative objects.

cs.CV

Skyrmion Phase and Non-Fermi Liquid Behavior in Nonsymmorphic Magnetic Weyl Semimetals

We investigate the interplay between complex magnetic orders and topological electronic states in nonsymmorphic magnetic Weyl semimetals of the ReAlX family (Re is a rare earth element and X is Si or Ge). We show that a Skyrmion lattice can fundamentally alter the behavior of Weyl fermions, driving the system into a non-Fermi liquid state and producing large, sign-tunable Hall responses. To this end, we construct a lattice model incorporating conduction Weyl fermions coupled to localized magnetic moments via Kondo interaction. Considering a multi-${\bf Q}$ cycloid magnetic configuration that evolves into a Skyrmion lattice under an in-plane Zeeman field, we analyze its profound impact on the band structure through magnetic Brillouin zone and band-folding. Using the Kubo formula, we calculate the conductivity tensor and examine the transport properties in the clean limit. Our results reveal that the Skyrmion lattice induces significant changes in both longitudinal and Hall conductivities. Remarkably, the temperature-dependent resistivity deviates from standard Fermi-liquid behavior ($ρ_{xx}\sim T^2$), exhibiting a non-Fermi liquid power-law scaling ($ρ_{xx}\sim T^α$ with $α$ between 3 and 5). This work provides a unified theoretical framework connecting multi-${\bf Q}$ magnetic textures, Skyrmion physics, and anomalous transport in topological semimetals, bridging the fields of topological magnetism and topological fermions.

cond-mat.mes-hall

Intelligence Delivery Network: Toward an Internet Architecture for the AI Age

The rapid emergence of AI-powered applications is reshaping the role of the Internet. Users increasingly rely on the network to obtain intelligence services derived from large foundation models, rather than merely to reach remote endpoints or retrieve specific content. Today's dominant deployment paradigm for AI services remains cloud-centric, where user requests are transmitted to remote data centers for centralized inference. Although operationally convenient, this paradigm suffers from latency and jitter, heavy wide-area traffic, limited utilization of distributed heterogeneous compute resources, and growing privacy and governance concerns. In this paper, we propose the Intelligence Delivery Network (IDN), an Internet architecture that treats AI capabilities as deliverable network services. The key idea is to position, select, reuse, and verify intelligence across cloud, regional, edge, and local environments according to demand locality, resource availability, and policy constraints. We present the system assumptions of IDN, define its core architectural mechanisms, and discuss how capability abstraction, compute resource integration, demand-driven deployment, service routing, state-aware caching, and trust management can jointly support distributed AI services. We believe that IDN provides a practical path toward an Internet architecture for the AI age, making AI capabilities more accessible, efficient, trustworthy, and responsive to diverse application needs.

cs.NI

CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives

Autoregressive video generation aims at real-time, open-ended synthesis. Yet, cinematic storytelling is not merely the endless extension of a single scene; it requires progressing through evolving events, viewpoint shifts, and discrete shot boundaries. Existing autoregressive models often struggle in this setting. Trained primarily for short-horizon continuation, they treat long sequences as extended single shots, inevitably suffering from motion stagnation and semantic drift during long rollouts. To bridge this gap, we introduce CausalCine, an interactive autoregressive framework that transforms multi-shot video generation into an online directing process. CausalCine generates causally across shot changes, accepts dynamic prompts on the fly, and reuses context without regenerating previous shots. To achieve this, we first train a causal base model on native multi-shot sequences to learn complex shot transitions prior to acceleration. We then propose Content-Aware Memory Routing (CAMR), which dynamically retrieves historical KV entries according to attention-based relevance scores rather than temporal proximity, preserving cross-shot coherence under bounded active memory. Finally, we distill the causal base model into a few-step generator for real-time interactive generation. Extensive experiments demonstrate that CausalCine significantly outperforms autoregressive baselines and approaches the capability of bidirectional models while unlocking the streaming interactivity of causal generation. Demo available at https://yihao-meng.github.io/CausalCine/

cs.CV

EGL-SCA: Structural Credit Assignment for Co-Evolving Instructions and Tools in Graph Reasoning Agents

Graph reasoning agents operating from natural-language inputs must solve a coupled problem: they must reconstruct a structured graph instance from text, decide whether existing computational assets are sufficient, interact with tools under a strict execution protocol, and satisfy an external verifier that checks structured correctness rather than textual plausibility. Existing approaches usually improve either the instruction side or the tool side in isolation, which leaves unclear what should be updated after failure. We propose EGL-SCA, a verifier-centric dual-space framework that models a graph reasoning agent using two collaborative components: an instruction-side policy space for reasoning strategies, and a tool-side program space for executable algorithmic tools. Our central mechanism is structural credit assignment, which maps trajectory evidence to conditional updates, precisely routing failures to either prompt optimization or tool synthesis and repair. To provide sufficient learning signals for dual-space adaptation, we introduce a training distribution stratified by task family, coupled with a Pareto-style retention strategy to balance success, generality, and parsimony. Experiments on four graph reasoning benchmarks show that EGL-SCA achieves a state-of-the-art 92.0\% average success rate. By effectively co-evolving instructions and tools, our framework significantly outperforms both pure-prompting and fixed-toolbox baselines.

cs.AI

Revisiting Graph-Tokenizing Large Language Models: A Systematic Evaluation of Graph Token Understanding

The remarkable success of large language models (LLMs) has motivated researchers to adapt them as universal predictors for various graph tasks. As a widely recognized paradigm, Graph-Tokenizing LLMs (GTokenLLMs) compress complex graph data into graph tokens and treat them as prefix tokens for querying LLMs, leading many to believe that LLMs can understand graphs more effectively and efficiently. In this paper, we challenge this belief: \textit{Do GTokenLLMs fully understand graph tokens in the natural-language embedding space?} Motivated by this question, we formalize a unified framework for GTokenLLMs and propose an evaluation pipeline, \textbf{GTEval}, to assess graph-token understanding via instruction transformations at the format and content levels. We conduct extensive experiments on 6 representative GTokenLLMs with GTEval. The primary findings are as follows: (1) Existing GTokenLLMs do not fully understand graph tokens. They exhibit over-sensitivity or over-insensitivity to instruction changes, and rely heavily on text for reasoning; (2) Although graph tokens preserve task-relevant graph information and receive attention across LLM layers, their utilization varies across models and instruction variants; (3) Additional instruction tuning can improve performance on the original and seen instructions, but it does not fully address the challenge of graph-token understanding, calling for further improvement.

cs.CL

2D Optical Beam Scanning using Integrated Acousto-Optics and a Frequency Comb

Optical beam steering is an essential technology for free-space optical communication, reconfigurable optical networks and quantum information systems. Yet conventional steering methods either require bulky mechanical mechanisms, or rely on complex arrays of individually controlled light emitting elements. Integrated acousto-optic beam steering (AOBS) offers non-mechanical, continuous one-dimensional steering on-chip by using traveling acoustic waves with variable frequency to deflect light. In this work, we combine AOBS with an optical frequency comb and optical gratings to enable two-dimensional beam steering from a single aperture. Azimuthal scanning is controlled via acoustic frequency while polar coverage is realized by dispersing frequency comb lines with the gratings. We demonstrate this architecture by sequentially selecting and steering 11 comb lines spanning 1540-1570 nm, achieving a field of view of 18.2 by 4.3 degrees. Validation with a tunable laser extends polar coverage to 11.4 degrees. Both components are realized on the same thin-film lithium niobate platform, providing a pathway toward monolithic integration.

physics.optics

Bridging Discrete Planning and Continuous Execution for Redundant Robot

Voxel-grid reinforcement learning is widely adopted for path planning in redundant manipulators due to its simplicity and reproducibility. However, direct execution through point-wise numerical inverse kinematics on 7-DoF arms often yields step-size jitter, abrupt joint transitions, and instability near singular configurations. This work proposes a bridging framework between discrete planning and continuous execution without modifying the discrete planner itself. On the planning side, step-normalized 26-neighbor Cartesian actions and a geometric tie-breaking mechanism are introduced to suppress unnecessary turns and eliminate step-size oscillations. On the execution side, a task-priority damped least-squares (TP-DLS) inverse kinematics layer is implemented. This layer treats end-effector position as a primary task, while posture and joint centering are handled as subordinate tasks projected into the null space, combined with trust-region clipping and joint velocity constraints. On a 7-DoF manipulator in random sparse, medium, and dense environments, this bridge raises planning success in dense scenes from about 0.58 to 1.00, shortens representative path length from roughly 1.53 m to 1.10 m, and while keeping end-effector error below 1 mm, reduces peak joint accelerations by over an order of magnitude, substantially improving the continuous execution quality of voxel-based RL paths on redundant manipulators.

cs.RO

MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling

We present MiroThinker v1.0, an open-source research agent designed to advance tool-augmented reasoning and information-seeking capabilities. Unlike previous agents that only scale up model size or context length, MiroThinker explores interaction scaling at the model level, systematically training the model to handle deeper and more frequent agent-environment interactions as a third dimension of performance improvement. Unlike LLM test-time scaling, which operates in isolation and risks degradation with longer reasoning chains, interactive scaling leverages environment feedback and external information acquisition to correct errors and refine trajectories. Through reinforcement learning, the model achieves efficient interaction scaling: with a 256K context window, it can perform up to 600 tool calls per task, enabling sustained multi-turn reasoning and complex real-world research workflows. Across four representative benchmarks-GAIA, HLE, BrowseComp, and BrowseComp-ZH-the 72B variant achieves up to 81.9%, 37.7%, 47.1%, and 55.6% accuracy respectively, surpassing previous open-source agents and approaching commercial counterparts such as GPT-5-high. Our analysis reveals that MiroThinker benefits from interactive scaling consistently: research performance improves predictably as the model engages in deeper and more frequent agent-environment interactions, demonstrating that interaction depth exhibits scaling behaviors analogous to model size and context length. These findings establish interaction scaling as a third critical dimension for building next-generation open research agents, complementing model capacity and context windows.

cs.CL

DPC: Training-Free Text-to-SQL Candidate Selection via Dual-Paradigm Consistency

While Large Language Models (LLMs) demonstrate impressive proficiency in generating SQL queries, they fundamentally lack the capability to self-evaluate correctness without an execution oracle. This limitation creates a stark Generation-Selection Gap, where high potential accuracy (Pass@K) fails to translate into execution accuracy (Pass@1). Although supervised verifiers offer mitigation, they incur prohibitive annotation costs and suffer from domain fragility. Consequently, recent research has pivoted to the training-free setting. However, existing methods--such as Self-Consistency or LLM-as-a-Judge--remain hampered by systematic bias (consensus on hallucinations) and symbolic blindness (inability to simulate execution states). We introduce DPC (Dual-Paradigm Consistency), a multi-agent framework that reformulates SQL selection from a probabilistic guessing task on hidden data into a deterministic verification task on visible data. Specifically, DPC employs a SLICER and a TESTER agent to collaboratively construct a Minimal Distinguishing Database (MDD)--an adversarial, fully observable micro-environment engineered to expose logical discrepancies between candidates. To break the self-correction bias, a SOLVER agent then verifies the SQL candidates by cross-referencing their execution against a parallel Python/Pandas solution. By validating execution consistency between declarative (SQL) and imperative (Python) paradigms, DPC robustly discriminates correct logic from systematic hallucinations. Experiments on BIRD and Spider across multiple LLMs demonstrate that our method consistently outperforms existing selection baselines, achieving absolute accuracy improvements of up to 2.2% over strong competitors like Self-Consistency.

cs.DB

Topologically non-trivial gap function and topology-induced time-reversal symmetry breaking in a superconductor with singular dynamical interaction

In many strongly correlated electron systems, non-Fermi liquid behavior and unconventional superconductivity can be viewed as emerging from an effective 4-fermion interaction with a singular frequency dependence. A pairing instability in such a system is qualitatively different from that in a Fermi liquid and generally gives rise to multiple pairing states with topologically distinct gap functions. However, in the systems studied so far, a topologically trivial solution has the lowest energy. Here we show that a repulsive Hubbard-type interaction with a finite cutoff added to a model with a singular dynamical interaction selects, in some parameter range, the theretofore subleading, topologically nontrivial solution. We consider a minimal model that displays this behavior and show that the transformation between the topologically trivial and nontrivial gap functions necessarily occurs via an intermediate phase with topology-induced breaking of time-reversal symmetry.

cond-mat.str-el

Judge Like Human Examiners: A Weighted Importance Multi-Point Evaluation Framework for Generative Tasks with Long-form Answers

Evaluating the quality of model responses remains challenging in generative tasks with long-form answers, as the expected answers usually contain multiple semantically distinct yet complementary factors that should be factorized for fine-grained assessment. Recent evaluation methods resort to relying on either task-level rubrics or question-aware checklists. However, they still 1) struggle to assess whether a response is genuinely grounded in provided contexts; 2) fail to capture the heterogeneous importance of different aspects of reference answers. Inspired by human examiners, we propose a Weighted Importance Multi-Point Evaluation (WIMPE) framework, which factorizes each reference answer into weighted context-bound scoring points. Two complementary metrics, namely Weighted Point-wise Alignment (WPA) and Point-wise Conflict Penalty (PCP), are designed to measure the alignment and contradiction between model responses and reference answers. Extensive experiments on 10 generative tasks demonstrate that WIMPE achieves higher correlations with human annotations.

cs.CL