SearcharxivSearch

arXiv subjects

Tao Yu

Publications and source records attributed to Tao Yu.

At least 109 records · Page 6Linked to original sources

MuonBP: Faster Muon via Block-Periodic Orthogonalization

Gradient orthogonalization is a simple strategy that shows great utility in speeding up gradient descent. The Muon optimizer (Jordan, Jin, et al., 2024) combines gradient orthogonalization with first-order momentum and achieves significant improvement in data efficiency over Adam/AdamW (Loshchilov and Hutter, 2019) for language model training. However, when using model parallelism, gradient orthogonalization introduces additional overhead compared to coordinate-wise optimizers (such as AdamW) due to additional gather and scatter operations on gradient matrix shards from different devices. This additional communication can amount to a throughput hit of 5%-10% compared to Adam/AdamW. To remedy this, we propose Muon with Block-Periodic Orthogonalization (MuonBP), which applies orthogonalization independently to matrix shards on each device and periodically performs full orthogonalization to maintain training stability at scale. We show how to adjust the learning rate from the baseline to MuonBP and give convergence guarantees for this algorithm. Crucially, our theory dictates that we use two stepsizes: one for the blockwise orthogonalization steps, and one for the full orthogonalization steps. Our method is simple, requires minimal hyperparameter adjustments, and achieves competitive iteration complexity compared with baseline Muon while providing per-iteration throughput comparable to coordinate-wise methods such as AdamW. When training an 8B model with eight-way tensor parallelism and ZeRO optimizer state sharding, MuonBP achieves 8% throughput increase compared to Muon with no degradation in performance.

cs.LG

Magnon Correlation Enables Spin Injection, Dephasing, and Transport in Canted Antiferromagnets

Thermal and electrical injection and transport of magnon spins in magnetic insulators is conventionally understood by the non-equilibrium population of magnons. However, this view is challenged by several recent experiments in noncollinear antiferromagnets, which urge a thorough theoretical investigation at the fundamental level. We find that the magnon spin in antiferromagnets is described by a matrix, so even when the diagonal terms -- spins carried by population -- vanish, the off-diagonal correlations transmit magnon spins. Our quantum theory shows that a net spin-flip of electrons in adjacent conductors creates quantum coherence between magnon states, which transports magnon spins in canted antiferromagnets, even without a definite phase difference between magnon modes in the incoherent process. It reveals that the pumped magnon correlation is not conserved due to an intrinsic spin torque, which causes dephasing and strong spatial spin oscillations during transport; both are enhanced by magnetic fields. Spin transfer to proximity conductors can cause extrinsic dephasing, which suppresses spin oscillations and thereby gates spin transport.

cond-mat.mes-hall

Evanescent Orbital Pumping by Magnetization Dynamics Free of Spin-Orbit Coupling

Converting magnetization spin to orbital current often relies on strong spin-orbit interaction that may cause additional angular momentum dissipation. We report that coherent magnetization dynamics in magnetic nanostructures can evanescently pump an orbital current into adjacent semiconductors due to the coupling between its stray electromagnetic field and electron orbitals without relying on spin-orbit coupling. The underlying photonic spin of the electromagnetic field governs the orbital polarization that flows along the gradient of the driven field. Due to the joint effect of the electric and magnetic fields, the orbital Hall current that flows perpendicularly to the gradient of the time-varying field is also generated and does not suffer from the orbital torque. These findings extend the paradigm of orbital pumping to include photonic angular momentum and pave the way for developing low-dissipation orbitronic devices.

cond-mat.mes-hall

Chiral Locking of Magnon Flow and Electron Spin Accumulation in Their Near-Field Radiative Spin Transfer

We report a non-contact mechanism for directional injection of magnons in magnetic films when driven by a spin accumulation $\pmbμ_s$ of electrons of a nearby metallic layer, governed by the long-range dipolar coupling between magnons and electron spins, which spontaneously generates a magnon current ${\bf J}_m$ flowing in the film plane. Crucially, in such near-field radiative spin transfer, the magnon flow ${\bf J}_m$ is always perpendicular to the spin accumulation $\pmbμ_s$, showing a universal chiral locking relation. The spin injection is efficient even when $\pmbμ_s$ is parallel to the magnetization, a feature breaking the limitation of the spin transfer by contact exchange interaction. Our findings reveal the critical role of dipolar chirality in driving the magnon thermal current and paving the way for the functional design of magnonic devices based on near-field radiative spin transfer.

cond-mat.mes-hall

BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions

Efficiently solving real-world problems with LLMs increasingly hinges on their ability to interact with dynamic web environments and autonomously acquire external information. While recent research like Search-R1 and WebDancer demonstrates strong performance in solving web tasks, they heavily rely on additional tools to convert the interactive web environment into static text content. This is in contrast to human browsing behaviors, which involve diverse interactions with the browser, such as scrolling, clicking, and typing. In this paper, we propose BrowserAgent, a more interactive agent that solves complex tasks through human-inspired browser actions. BrowserAgent operates directly on raw web pages via Playwright through a set of predefined browser actions. We adopt a two-stage training (Supervised Fine-Tuning (SFT) and Rejection Fine-Tuning (RFT)) to improve the model's generalization abilities. Despite using significantly less training data than Search-R1, BrowserAgent achieves more competitive results across different Open-QA tasks. Additionally, we introduce an explicit memory mechanism to store key conclusions across steps, further enhancing the model's reasoning capabilities for long-horizon tasks. Notably, BrowserAgent-7B can achieve around 20\% improvement over Search-R1 on multi-hop QA tasks like HotpotQA, 2Wiki, and Bamboogle. These results indicate that BrowserAgent can serve as a more advanced framework for more interactive and scalable web agents.

cs.CL

OpenCUA: Open Foundations for Computer-Use Agents

Vision-language models have demonstrated impressive capabilities as computer-use agents (CUAs) capable of automating diverse computer tasks. As their commercial potential grows, critical details of the most capable CUA systems remain closed. As these agents will increasingly mediate digital interactions and execute consequential decisions on our behalf, the research community needs access to open CUA frameworks to study their capabilities, limitations, and risks. To bridge this gap, we propose OpenCUA, a comprehensive open-source framework for scaling CUA data and foundation models. Our framework consists of: (1) an annotation infrastructure that seamlessly captures human computer-use demonstrations; (2) AgentNet, the first large-scale computer-use task dataset spanning 3 operating systems and 200+ applications and websites; (3) a scalable pipeline that transforms demonstrations into state-action pairs with reflective long Chain-of-Thought reasoning that sustain robust performance gains as data scales. Our end-to-end agent models demonstrate strong performance across CUA benchmarks. In particular, OpenCUA-72B achieves an average success rate of 45.0% on OSWorld-Verified, establishing a new state-of-the-art (SOTA) among open-source models. Further analysis confirms that our approach generalizes well across domains and benefits significantly from increased test-time computation. We release our annotation tool, datasets, code, and models to build open foundations for further CUA research.

cs.AI

On the Bonahon--Wong--Yang invariants of pseudo-Anosov maps

We conjecture (and prove for once-punctured torus bundles) that the Bonahon--Wong--Yang invariants of pseudo-Anosov homeomorphisms of a punctured surface at roots of unity coincide with the 1-loop invariant of their mapping torus at roots of unity. This explains the topological invariance of the BWY invariants and how their volume conjecture, to all orders, and with exponentially small terms included, follows from the quantum modularity conjecture. Using the numerical methods of Zagier and the first author, we illustrate how to efficiently compute the invariants and their asymptotics to arbitrary order in perturbation theory, using as examples the $LR$ and the $LLR$ pseudo-Anosov monodromies of the once-punctured torus. Finally, we introduce descendant versions of the 1-loop and BWY invariants and conjecture (and numerically check for pseudo-Anosov monodromies of $L/R$-length at most 5) that they are related by a Fourier transform. This edition includes statements and proofs for roots of unity of all order, even and odd.

math.GT

Digital Twin-based Cooperative Autonomous Driving in Smart Intersections: A Multi-Agent Reinforcement Learning Approach

Unsignalized intersections pose safety and efficiency challenges due to complex traffic flows and blind spots. In this paper, a digital twin (DT)-based cooperative driving system with roadside unit (RSU)-centric architecture is proposed for enhancing safety and efficiency at unsignalized intersections. The system leverages comprehensive bird-eye-view (BEV) perception to eliminate blind spots and employs a hybrid reinforcement learning (RL) framework combining offline pre-training with online fine-tuning. Specifically, driving policies are initially trained using conservative Q-learning (CQL) with behavior cloning (BC) on real datasets, then fine-tuned using multi-agent proximal policy optimization (MAPPO) with self-attention mechanisms to handle dynamic multi-agent coordination. The RSU implements real-time commands via vehicle-to-infrastructure (V2I) communications. Experimental results show that the proposed method yields failure rates below 0.03\% coordinating up to three connected autonomous vehicles (CAVs), significantly outperforming traditional methods. In addition, the system exhibits sub-linear computational scaling with inference times under 40 ms. Furthermore, it demonstrates robust generalization across diverse unsignalized intersection scenarios, indicating its practicality and readiness for real-world deployment.

eess.SY

TREE:Token-Responsive Energy Efficiency Framework For Green AI-Integrated 6G Networks

As wireless networks evolve toward AI-integrated intelligence, conventional energy-efficiency metrics fail to capture the value of AI tasks. In this paper, we propose a novel EE metric called Token-Responsive Energy Efficiency (TREE), which incorporates the token throughput of large models as network utility carriers into the system utility. Based on this metric, we analyze the design principles of AI-integrated 6G networks from the perspective of three critical AI elements, namely computing power, model and data. Case studies validate TREE's unique capability to expose energy-service asymmetries in hybrid traffic scenarios where conventional metrics prove inadequate. Although it is impossible to determine every design detail of AI-integrated 6G network at current time, we believe that the proposed TREE based framework will help the network operators to quantify the operating energy cost of AI services and continue to evolve towards sustainable 6G networks.

eess.SY

Can Uncertainty Quantification Improve Learned Index Benefit Estimation?

Index tuning is crucial for optimizing database performance by selecting optimal indexes based on workload. The key to this process lies in an accurate and efficient benefit estimator. Traditional methods relying on what-if tools often suffer from inefficiency and inaccuracy. In contrast, learning-based models provide a promising alternative but face challenges such as instability, lack of interpretability, and complex management. To overcome these limitations, we adopt a novel approach: quantifying the uncertainty in learning-based models' results, thereby combining the strengths of both traditional and learning-based methods for reliable index tuning. We propose Beauty, the first uncertainty-aware framework that enhances learning-based models with uncertainty quantification and uses what-if tools as a complementary mechanism to improve reliability and reduce management complexity. Specifically, we introduce a novel method that combines AutoEncoder and Monte Carlo Dropout to jointly quantify uncertainty, tailored to the characteristics of benefit estimation tasks. In experiments involving sixteen models, our approach outperformed existing uncertainty quantification methods in the majority of cases. We also conducted index tuning tests on six datasets. By applying the Beauty framework, we eliminated worst-case scenarios and more than tripled the occurrence of best-case scenarios.

cs.DB

Training LLMs with MXFP4

Low precision (LP) datatypes such as MXFP4 can accelerate matrix multiplications (GEMMs) and reduce training costs. However, directly using MXFP4 instead of BF16 during training significantly degrades model quality. In this work, we present the first near-lossless training recipe that uses MXFP4 GEMMs, which are $2\times$ faster than FP8 on supported hardware. Our key insight is to compute unbiased gradient estimates with stochastic rounding (SR), resulting in more accurate model updates. However, directly applying SR to MXFP4 can result in high variance from block-level outliers, harming convergence. To overcome this, we use the random Hadamard tranform to theoretically bound the variance of SR. We train GPT models up to 6.7B parameters and find that our method induces minimal degradation over mixed-precision BF16 training. Our recipe computes $>1/2$ the training FLOPs in MXFP4, enabling an estimated speedup of $>1.3\times$ over FP8 and $>1.7\times$ over BF16 during backpropagation.

cs.LG

IKOD: Mitigating Visual Attention Degradation in Large Vision-Language Models

Recent advancements in Large Vision-Language Models (LVLMs) have demonstrated significant progress across multiple domains. However, these models still face the inherent challenge of integrating vision and language for collaborative inference, which often leads to "hallucinations", outputs that are not grounded in the corresponding images. Many efforts have been made to address these issues, but each comes with its own limitations, such as high computational cost or expensive dataset annotation. Recent research shows that LVLMs exhibit a long-term bias where hallucinations increase as the sequence length grows, yet the underlying cause remains poorly understood. Building on extensive research into attention mechanisms in LVLMs, we analyze the relationship between this long-term bias and visual attention. In our research, we identify a consistent phenomenon in current LVLMs: the model's attention to visual input diminishes as the generated sequence grows, which we hypothesize to be a key factor contributing to observed increasing hallucinations. Based on these insights, we propose Image attention-guided Key-value merging cOllaborative Decoding (IKOD), a collaborative decoding strategy generating more image-focused sequences. This method derives logits from shorter sequences with higher image attention through key-value merging and combines them with those from the original decoding, effectively mitigating attention degradation and suppressing hallucinations while not incurring too much inference cost. Extensive experiments on both hallucination and comprehensive benchmarks demonstrate IKOD's superior effectiveness in mitigating hallucinations and improving comprehensive capacities for LVLMs. Importantly, IKOD requires no additional training or external tools, making it a lightweight and efficient framework applicable to various models.

cs.CV

Directional entanglement of spin-orbit locked nitrogen-vacancy centers by magnons

We address that the stray magnetic field emitted by the excited quantum states of the nitrogen-vacancy (NV) centers is spin-momentum locked, such that the spin transfer to nearby ferromagnetic nanostructures is unidirectional. This may allow the controlled excitation of propagating magnons by NV centers in diamond. A pair of NV spin qubits exchange virtual magnons in a magnetic nanowire in a chiral manner that leads to directional quantum entanglement. A magnon-based ``quantum-entanglement isolator" should be a useful device in future quantum information technology.

cond-mat.mes-hall

Image of the time-dependent black hole

The Event Horizon Telescope's 2024 observations report a shift in the position angle of the brightness asymmetry in M87*, revealing time variability in the black hole's image. In this analysis, we investigate the time-dependent of a Vaidya black hole. By introducing a mass function that increases linearly with time, along with a conformal transformation, we derive the conformal Vaidya metric and define a new time coordinate $t_c$. Using the semi-analytical approach, we analyze the ray trajectories and radiation flux of the Vaidya black hole in the background of a thin accretion disk. We discuss how the observed flux in the Vaidya spacetime evolves as a function of the new time coordinate $t_c$. The results show that the facula on the observable plane undergoes radial displacement as $t_c$ increases, revealing the time-dependent evolution of black hole images.

gr-qc

Kwai Keye-VL Technical Report

While Multimodal Large Language Models (MLLMs) demonstrate remarkable capabilities on static images, they often fall short in comprehending dynamic, information-dense short-form videos, a dominant medium in today's digital landscape. To bridge this gap, we introduce \textbf{Kwai Keye-VL}, an 8-billion-parameter multimodal foundation model engineered for leading-edge performance in short-video understanding while maintaining robust general-purpose vision-language abilities. The development of Keye-VL rests on two core pillars: a massive, high-quality dataset exceeding 600 billion tokens with a strong emphasis on video, and an innovative training recipe. This recipe features a four-stage pre-training process for solid vision-language alignment, followed by a meticulous two-phase post-training process. The first post-training stage enhances foundational capabilities like instruction following, while the second phase focuses on stimulating advanced reasoning. In this second phase, a key innovation is our five-mode ``cold-start'' data mixture, which includes ``thinking'', ``non-thinking'', ``auto-think'', ``think with image'', and high-quality video data. This mixture teaches the model to decide when and how to reason. Subsequent reinforcement learning (RL) and alignment steps further enhance these reasoning capabilities and correct abnormal model behaviors, such as repetitive outputs. To validate our approach, we conduct extensive evaluations, showing that Keye-VL achieves state-of-the-art results on public video benchmarks and remains highly competitive on general image-based tasks (Figure 1). Furthermore, we develop and release the \textbf{KC-MMBench}, a new benchmark tailored for real-world short-video scenarios, where Keye-VL shows a significant advantage.

cs.CV

ELGAR: Expressive Cello Performance Motion Generation for Audio Rendition

The art of instrument performance stands as a vivid manifestation of human creativity and emotion. Nonetheless, generating instrument performance motions is a highly challenging task, as it requires not only capturing intricate movements but also reconstructing the complex dynamics of the performer-instrument interaction. While existing works primarily focus on modeling partial body motions, we propose Expressive ceLlo performance motion Generation for Audio Rendition (ELGAR), a state-of-the-art diffusion-based framework for whole-body fine-grained instrument performance motion generation solely from audio. To emphasize the interactive nature of the instrument performance, we introduce Hand Interactive Contact Loss (HICL) and Bow Interactive Contact Loss (BICL), which effectively guarantee the authenticity of the interplay. Moreover, to better evaluate whether the generated motions align with the semantic context of the music audio, we design novel metrics specifically for string instrument performance motion generation, including finger-contact distance, bow-string distance, and bowing score. Extensive evaluations and ablation studies are conducted to validate the efficacy of the proposed methods. In addition, we put forward a motion generation dataset SPD-GEN, collated and normalized from the MoCap dataset SPD. As demonstrated, ELGAR has shown great potential in generating instrument performance motions with complicated and fast interactions, which will promote further development in areas such as animation, music education, interactive art creation, etc.

cs.GR

Kimi-VL Technical Report

We present Kimi-VL, an efficient open-source Mixture-of-Experts (MoE) vision-language model (VLM) that offers advanced multimodal reasoning, long-context understanding, and strong agent capabilities - all while activating only 2.8B parameters in its language decoder (Kimi-VL-A3B). Kimi-VL demonstrates strong performance across challenging domains: as a general-purpose VLM, Kimi-VL excels in multi-turn agent tasks (e.g., OSWorld), matching flagship models. Furthermore, it exhibits remarkable capabilities across diverse challenging vision language tasks, including college-level image and video comprehension, OCR, mathematical reasoning, and multi-image understanding. In comparative evaluations, it effectively competes with cutting-edge efficient VLMs such as GPT-4o-mini, Qwen2.5-VL-7B, and Gemma-3-12B-IT, while surpassing GPT-4o in several key domains. Kimi-VL also advances in processing long contexts and perceiving clearly. With a 128K extended context window, Kimi-VL can process diverse long inputs, achieving impressive scores of 64.5 on LongVideoBench and 35.1 on MMLongBench-Doc. Its native-resolution vision encoder, MoonViT, further allows it to see and understand ultra-high-resolution visual inputs, achieving 83.2 on InfoVQA and 34.5 on ScreenSpot-Pro, while maintaining lower computational cost for common tasks. Building upon Kimi-VL, we introduce an advanced long-thinking variant: Kimi-VL-Thinking-2506. Developed through long chain-of-thought (CoT) supervised fine-tuning (SFT) and reinforcement learning (RL), the latest model exhibits strong long-horizon reasoning capabilities (64.0 on MMMU, 46.3 on MMMU-Pro, 56.9 on MathVision, 80.1 on MathVista, 65.2 on VideoMMMU) while obtaining robust general abilities. Code and models are publicly accessible at https://github.com/MoonshotAI/Kimi-VL.

cs.CV

Electromagnetic Proximity Effect: Superconducting Magnonics and Beyond

The exchange interaction at interfaces between superconductors (SCs) and ferromagnets (FMs) has been a central topic in condensed matter physics for many decades, starting with the prediction of exotic phases such as the Fulde-Ferrell-Larkin-Ovchinnikov states and leading to the discovery of triplet superconductivity. This review focuses on new phenomena in SC$|$FM heterostructures caused by the \textit{non-contact dipolar interaction} between magnons, i.e., the quanta of spin wave excitations in the ferromagnet, and the superconducting order. A universal non-relativistic spin-orbit coupling locks the polarization and momentum of their evanescent stray magnetic fields and leads to chiral screening by proximate superconductors. The interaction-induced hybrid quasiparticles are magnon-Meissner collective modes, magnon-cooparon, Josephson plasmonic modes, and nodal magnon-photon polaritons. Superconducting and normal metallic gates modulate and control the magnetodipolar interaction and thereby magnetization and energy transport at interfaces and in thin films.

cond-mat.supr-con