SearcharxivSearch

arXiv subjects

Yantao Li

Publications and source records attributed to Yantao Li.

At least 19 recordsLinked to original sources

Is Multimodal Speculative Decoding Ready for Diffusion-Based Parallel Drafting? A Survey and Empirical Diagnosis

Speculative decoding accelerates autoregressive generation by allowing a lightweight drafter to propose future tokens while a target model verifies them in parallel. Its lossless guarantee has motivated a line of work that pushes the drafter itself toward parallel generation. The most recent paradigm is block-parallel generative drafting, including diffusion-based methods such as DFlash and DSpark, achieving up to 3.6x speedup on common daily chatting tasks. While this transition is well studied in text-only LLMs, its applicability to multimodal models remains an open question. Existing multimodal speculative decoding efforts focus on input compression, adapter alignment, candidate coverage, or modality-specific verification; however, block-parallel generative drafting remains largely unexplored. To bridge this gap, this paper combines a modality-centered survey with a cross-architecture empirical study to ask: Is multimodal speculative decoding ready for diffusion-based parallel drafting? In this survey, we systematically analyze a wide spectrum of multimodal models, spanning Vision-Language, Video-Language, Audio, and Vision-Language-Action (VLA) architectures, from the dual perspectives of drafting parallelism and cross-modal information interaction. We introduce a unified taxonomy that isolates drafter-side parallelism from orthogonal design choices such as tree construction and verification strategies. Furthermore, we provide a comprehensive empirical comparison of existing methods under varying degrees of parallelism across standardized multimodal benchmarks, including OCR, VQA, visual reasoning, and image captioning. Finally, we summarize the limitations of current approaches, discuss open challenges, and outline promising future directions for this rapidly evolving field.

cs.AI

Dynamical Polarization from Hidden Spin and Orbital Textures in p-Wave Magnets

Period-averaged descriptions often miss essential features of driven quantum matter. We show that the micromotion of an optically driven $p$-wave magnet unveils a hidden net spin polarization, absent from both the equilibrium and period-averaged spin textures, which remain odd in momentum. This spin polarization oscillates at the drive frequency and is resonantly enhanced at the interband gap set by nonrelativistic exchange splitting. The drive further activates an orbital angular momentum governed by interband quantum geometry. While its linear response remains momentum-odd, nonlinear rectification yields a static, momentum-even orbital polarization for suitably oriented driving fields. These results establish $p$-wave magnets as a source of resonant ac spin and rectified dc orbital polarization: effects invisible to any period-averaged treatment.

cond-mat.mes-hall

OTCache: Optimal Transport for Geometry-Aware Caching in Diffusion Models

We propose OTCache, a training-free framework for accelerating diffusion sampling via caching schedule prediction. Existing graph-based caching methods reduce redundant computation by optimizing shortest-path objectives, but rely on an additive independence assumption, which often breaks down in the low NFE regime. To address this issue, OTCache models caching schedules across inference budgets as a smooth evolution in policy space, inspired by Optimal Transport (OT). The framework consists of three stages: (1) obtaining a high-fidelity \textbf{reference schedule} using a graph-based caching method under a conservative budget; (2) performing a lightweight anchor search under an extreme low-budget setting via Optuna optimization with an end-to-end perceptual objective; and (3) predicting schedules for target budgets via quantile interpolation between the reference and anchor policies using continuous warping representations. Experiments on FLUX.1 [dev], Qwen-Image, and HunyuanVideo show that OTCache achieves 4.5x, 4.7x, and 3.66x acceleration, respectively, while consistently improving generation fidelity over state-of-the-art caching baselines. This work provides a new perspective on accelerating diffusion models through Optimal-Transport-inspired schedule modeling. Code:https://github.com/UnicomAI/OTCache

cs.LG

Token-Operations-Oriented Inference Optimization Techniques for Large Models

Large model inference optimization serves as a key foundation for supporting the scalable, low-cost, and highly stable operation of large model services. Centered on token-oriented inference optimization technology, this paper proposes for the first time a four-layer technical architecture consisting of Multi-model Fusion, Model Optimization, Compute-Model Fusion, and Compute-Network-Model Fusion. It systematically reviews the key technologies and current industry status across these four levels and analyzes the application value of related technologies in real-world business scenarios. This paper provides a practical technical path for reducing token production costs, improving token service efficiency, ensuring the stability of token supply, and driving the transition of large model services from being merely callable to being operable.

cs.SE

VeriScale: Adversarial Test-Suite Scaling for Verifiable Code Generation

As large language models (LLMs) are increasingly deployed for software engineering, constructing high-quality benchmarks is crucial for evaluating not just the functional correctness, but also the formal verifiability of generated code. However, existing benchmarks are limited by the quantity and quality of positive and negative test cases, leading to an overestimation of model capabilities in generating specifications and implementations. To address this, we propose VeriScale, a novel framework driven by the adversarial implementations. It consists of two stages: test-suite expansion to construct diverse and challenging test cases, and test-suite reduction to distill them into compact yet discriminative suites. While VeriScale is general, we instantiate it on Verina to construct VerinaPlus, which expands the original test suites by over 83$\times$, and VerinaLite, a lightweight 14$\times$ variant. Our experiments across eight state-of-the-art LLMs demonstrate that VerinaPlus exposes substantial model weaknesses hidden by the original benchmark, evidenced by sharp score drops on both SpecGen and CodeGen tasks, whereas VerinaLite maintains this discriminative power at a fraction of the evaluation cost. The enhanced benchmarks and source code are publicly available at https://github.com/XiaoyangLiu-sjtu/VeriScale.

cs.LG

MediaClaw: Multimodal Intelligent-Agent Platform Technical Report

MediaClaw is a multimodal agent platform built on the OpenClaw ecosystem. Its core design follows a three-layer architecture of unified abstraction, pluginized extension, and workflow orchestration. The system is intended to address practical deployment pain points in AIGC adoption, including fragmented capabilities, heterogeneous interfaces, disconnected production processes, and limited reuse of high-quality production workflows. \system{} abstracts full-category AIGC capabilities into a unified invocation model, uses plugins to support hot-pluggable capability expansion, and uses task-oriented Skills to turn complex production processes into reusable workflow assets. This report focuses on the architectural design philosophy of MediaClaw, the design logic of its core capability model, and the key engineering trade-offs in implementation. It aims to provide reusable practical reference for building multimodal capability platforms.

cs.AI

Decompose, Structure, and Repair: A Neuro-Symbolic Framework for Autoformalization via Operator Trees

Statement autoformalization acts as a critical bridge between human mathematics and formal mathematics by translating natural language problems into formal language. While prior works have focused on data synthesis and diverse training paradigms to optimize end-to-end Large Language Models (LLMs), they typically treat formal code as flat sequences, neglecting the hierarchical logic inherent in mathematical statements. In this work, we introduce Decompose, Structure, and Repair (DSR), a neuro-symbolic framework that restructures autoformalization into a modular pipeline. DSR decomposes statements into logical components and maps them to structured operator trees, leveraging this topological blueprint to precisely localize and repair errors via sub-tree refinement. Furthermore, we introduce PRIME, a benchmark of 156 undergraduate and graduate-level theorems selected from canonical textbooks and expertly annotated in Lean 4. Experimental results demonstrate that DSR establishes a new state-of-the-art, consistently outperforming baselines under equivalent computational budgets. The datasets, model, and code are available at https://github.com/XiaoyangLiu-sjtu/DSR.

cs.LG

$P$-wave Orbital Magnetism

Realization of unconventional odd-parity magnets usually requires noncollinear spin textures of the underlying lattice. We propose a different concept of $p$-wave magnetism that originates from an orbital texture induced by loop currents. The resulting $p$-wave orbital magnetism is protected by the combined translation and time-reversal symmetry, with even-parity components arising when the symmetry is broken. Our proposal is exemplified by a two-dimensional (2D) lattice model whose energy spectrum contains Dirac points and which is characterized by a nontrivial topology controlled by the magnitude of the loop currents. Since the odd-parity magnetism precludes macroscopic magnetization, we suggest measuring it via orbital Hall conductivity. Our work establishes orbital degrees of freedom as an additional platform for unconventional $p$-wave magnetism beyond noncollinear spin textures, as well as makes a step forward to bridging odd-parity magnetism and topology.

cond-mat.mes-hall

PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment

Reinforcement learning has recently improved the reasoning ability of Large Language Models and Multimodal LLMs, yet prevailing reward designs emphasise final-answer correctness and consequently tolerate process hallucinations--cases where models reach the right answer while misperceiving visual evidence. We address this process-level misalignment with PaLMR, a framework that aligns not only outcomes but also the reasoning process itself. PaLMR comprises two complementary components: a perception-aligned data layer that constructs process-aware reasoning data with structured pseudo-ground-truths and verifiable visual facts, and a process-aligned optimisation layer that constructs a hierarchical reward fusion scheme with a process-aware scoring function to encourage visually faithful chains-of-thought and improve training stability. Experiments on Qwen2.5-VL-7B show that our approach substantially reduces reasoning hallucinations and improves visual reasoning fidelity, achieving state-of-the-art results on HallusionBench while maintaining strong performance on MMMU, MathVista, and MathVerse. These findings indicate that PaLMR offers a principled and practical route to process-aligned multimodal reasoning, advancing the reliability and interpretability of MLLMs.

cs.CV

Formalization of Two Fixed-Point Algorithms in Hilbert Spaces

Iterative algorithms are fundamental tools for approximating fixed-points of nonexpansive operators in real Hilbert spaces. Among them, Krasnosel'ski\u{\i}--Mann iteration and Halpern iteration are two widely used schemes. In this work, we formalize the convergence of these two fixed-point algorithms in the interactive theorem prover Lean4 based on type dependent theory. To this end, weak convergence and topological properties in the infinite-dimensional real Hilbert space are formalized. Definition and properties of nonexpansive operators are also provided. As a useful tool in convex analysis, we then formalize the Fej\'{e}r monotone sequence. Building on these foundations, we verify the convergence of both the iteration schemes. Our formalization provides reusable components for machine-checked convergence analysis of fixed-point iterations and theories of convex analysis in real Hilbert spaces. Our code is available at https://github.com/TTony2019/fixed-point-iterations-in-lean.

math.OC

Tunable Non-Equilibrium Magic and Minimum Twist Angles in AA-Stacked Twisted Multilayer Graphene

We report the discovery of a series of non-equilibrium magic angles at which isolated topological flat quasienergy bands form in AA-stacked twisted multilayer graphene under circularly polarized light. These non-equilibrium magic angles can be traced back to specific static twist angles where the bandwidth reaches a minimum \textit{without} the formation of isolated flat bands. We refer to these as minimum twist angles, in contrast to the magic angles observed in twisted bilayer graphene. We show that an applied displacement field can further flatten the optically induced topological flat bands accompanied by larger non-equilibrium magic angles. The discovery of these electrically tunable topological flat quasienergy bands is expected to open up a new avenue of exploring exotic Floquet-driven phenomena in AA-stacked twisted multilayer graphene.

cond-mat.str-el

Vision-Language Models Can Self-Improve Reasoning via Reflection

Chain-of-thought (CoT) has proven to improve the reasoning capability of large language models (LLMs). However, due to the complexity of multimodal scenarios and the difficulty in collecting high-quality CoT data, CoT reasoning in multimodal LLMs has been largely overlooked. To this end, we propose a simple yet effective self-training framework, R3V, which iteratively enhances the model's Vision-language Reasoning by Reflecting on CoT Rationales. Our framework consists of two interleaved parts: (1) iteratively bootstrapping positive and negative solutions for reasoning datasets, and (2) reflection on rationale for learning from mistakes. Specifically, we introduce the self-refine and self-select losses, enabling the model to refine flawed rationale and derive the correct answer by comparing rationale candidates. Experiments on a wide range of vision-language tasks show that R3V consistently improves multimodal LLM reasoning, achieving a relative improvement of 23 to 60 percent over GPT-distilled baselines. Additionally, our approach supports self-reflection on generated solutions, further boosting performance through test-time computation.

cs.LG

Collective modes in terahertz field response of superconductors with paramagnetic impurities

We consider a problem of nonlinear response to an external electromagnetic radiation of conventional disordered superconductors which contain a small amount of weak magnetic impurities. We focus on the diffusive limit and use Usadel equation to analyze the collective excitations and obtain the dispersion relations for the collective modes. We determine the resonant frequency and dispersion of both amplitude and phase (Carlson-Goldman) modes for moderate strength of magnetic scattering. We find that the Carlson-Goldman and superconducting plasmon modes can only be excited at some finite value of the threshold momentum which increases with an increase in spin-flip scattering rate while the amplitude mode is diffusive and becomes strongly suppressed with the increase in spin-flip scattering. The value of the threshold momentum is determined by the distance between the two consecutive spin-flip scattering events. Furthermore, we also find that the superconducting plasmon mode becomes gapless in the presence of the pair breaking processes. Possible ways towards experimental verification of our results are also discussed.

cond-mat.supr-con

SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Graphical User Interface (GUI) agents are designed to automate complex tasks on digital devices, such as smartphones and desktops. Most existing GUI agents interact with the environment through extracted structured data, which can be notably lengthy (e.g., HTML) and occasionally inaccessible (e.g., on desktops). To alleviate this issue, we propose a novel visual GUI agent -- SeeClick, which only relies on screenshots for task automation. In our preliminary study, we have discovered a key challenge in developing visual GUI agents: GUI grounding -- the capacity to accurately locate screen elements based on instructions. To tackle this challenge, we propose to enhance SeeClick with GUI grounding pre-training and devise a method to automate the curation of GUI grounding data. Along with the efforts above, we have also created ScreenSpot, the first realistic GUI grounding benchmark that encompasses mobile, desktop, and web environments. After pre-training, SeeClick demonstrates significant improvement in ScreenSpot over various baselines. Moreover, comprehensive evaluations on three widely used benchmarks consistently support our finding that advancements in GUI grounding directly correlate with enhanced performance in downstream GUI agent tasks. The model, data and code are available at https://github.com/njucckevin/SeeClick.

cs.HC

Amplitude Higgs mode in superconductors with magnetic impurities

We study nonlinear response of conventional superconducting alloys with weak magnetic impurities to an external alternating electromagnetic field. In particular, we calculate a correction to superconducting order parameter up to the second order in external vector potential. We show that frequency dependence of the order parameter amplitude has characteristic resonant shape with a maximum at frequency which is smaller than twice the magnitude of the pairing amplitude in equilibrium and at the same time exceeds the single-particle threshold energy. Our results suggest that in the presence of magnetic impurities the dynamics of the pairing amplitude in the collisionless regime will remain robust with respect to dissipative processes. We also evaluate the third harmonic contribution to the current as a function of the probe frequency and for various concentrations of magnetic impurities.

cond-mat.supr-con

Topological Mixed Valence Model in Magic-Angle Twisted Bilayer Graphene

We develop a model to describe the mixed valence regime in magic-angle twisted bilayer graphene (MATBG) using the recently developed heavy-fermion framework. By employing the large-$N$ slave-boson approach, we derive the self-consistent mean field equations and solve them numerically. We find that the SU(8) symmetry constraint moir\'e system exhibits novel mixed-valence properties which are different from conventional heavy-fermions systems. We find the solutions describing the physics at the filling near the Mott insulator regime in the limit of strong Coulomb interactions between the flat-band fermions. Our model can provide additional insight into the possible microscopic origin of unconventional superconductivity in MATBG.

cond-mat.supr-con

Enhancement of High Harmonic Generation in Bulk Floquet Systems

We formulate a theory of bulk optical current for a periodically driven system, which accounts for the mixing of external drive and laser field frequencies and, therefore, the broadening of the harmonic spectrum compared to the undriven system. We express the current in terms of Floquet-Bloch bands and their non-adiabatic Berry connection and curvature. Using this expression, we relate spatio-temporal symmetries of the driven model to selection rules for current harmonics. We illustrate the application of this theory by studying high harmonic generation in the periodically driven Su-Schrieffer-Heeger model. In the high frequency and low field amplitude limit, we find analytical expressions for current harmonics. We also calculate the current numerically beyond the high frequency limit and verify that when the drive breaks a temporal symmetry, harmonics forbidden in the undriven model become available. Moreover, we find significant enhancement in higher harmonics when the system is driven, even for low field amplitudes. Our work offers a unified Floquet approach to nonlinear optical properties of solids, which is useful for realistic calculations of high harmonic spectra of electronic systems subject to multiple periodic drives.

cond-mat.mes-hall

Renormalized Magic Angles in Asymmetric Twisted Graphene Multilayers

Stacked graphene multilayers with a small relative twist angle between each of the layers have been found to host flat bands at a series of magic angles. We consider the effect that Dirac point asymmetry between the layers, and in particular different Fermi velocities in each layer, may have on this phenomenon. Such asymmetry may be introduced by unequal Fermi velocity renormalizations through Coulomb interactions with a dielectric substrate. It also arises in an approximate way in tetralayer systems, in which the outer twist angles are large enough that there is a dominant moire periodicity from the stacking of the inner two layers. We find in such models that the flat band phenomenon persists in spite of this asymmetry, and that the magic angles acquire a degree of tunability through either controlling the screening in the bilayer system or the twist angles of the outer layers in the tetralayer system. Notably, we find in our models that the quantitative values of the magic angles are increased.

cond-mat.mes-hall