SearcharxivSearch

arXiv subjects

Tao Yu

Publications and source records attributed to Tao Yu.

At least 73 records · Page 4Linked to original sources

Predicting Networks Before They Happen: Experimentation on a Real-Time V2X Digital Twin

Emerging safety-critical Vehicle-to-Everything (V2X) applications require networks to proactively adapt to rapid environmental changes rather than merely reacting to them. While Network Digital Twins (NDTs) offer a pathway to such predictive capabilities, existing solutions typically struggle to reconcile high-fidelity physical modeling with strict real-time constraints. This paper presents a novel, end-to-end real-time V2X Digital Twin framework that integrates live mobility tracking with deterministic channel simulation. By coupling the Tokyo Mobility Digital Twin-which provides live sensing and trajectory forecasting-with VaN3Twin-a full-stack simulator with ray tracing-we enable the prediction of network performance before physical events occur. We validate this approach through an experimental proof-of-concept deployed in Tokyo, Japan, featuring connected vehicles operating on 60 GHz links. Our results demonstrate the system's ability to predict Received Signal Strength (RSSI) with a maximum average error of 1.01 dB and reliably forecast Line-of-Sight (LoS) transitions within a maximum average end-to-end system latency of 250 ms, depending on the ray tracing level of detail. Furthermore, we quantify the fundamental trade-offs between digital model fidelity, computational latency, and trajectory prediction horizons, proving that high-fidelity and predictive digital twins are feasible in real-world urban environments.

cs.NI

Ultrastrong magnon-photon coupling in superconductor/antiferromagnet/superconductor heterostructures at terahertz frequencies

We predict the realization of ultrastrong coupling between magnons of antiferromagnets and photons in superconductor/antiferromagnet/superconductor heterostructures at terahertz frequencies, from both quantum and classical perspectives. The hybridization of the two magnon modes with photons strongly depends on the applied magnetic field: at zero magnetic field, only a single antiferromagnetic mode with a lower frequency couples to the photon, forming a magnon-polariton, while using a magnetic field activates coupling for both antiferromagnetic modes. The coupling between magnon and photon is ultrastrong with the coupling constant $\sim$ 100 GHz exceeding 10% of the antiferromagnetic resonant frequency. The superconductor modulates the spin of the resulting magnon-polaritons and the group velocity, achieving values amounting to several tenths of the speed of light, which promises strong tunability of magnon transport in antiferromagnets by superconductors.

cond-mat.supr-con

Realizing Immersive Volumetric Video: A Multimodal Framework for 6-DoF VR Engagement

Fully immersive experiences that tightly integrate 6-DoF visual and auditory interaction are essential for virtual and augmented reality. While such experiences can be achieved through computer-generated content, constructing them directly from real-world captured videos remains largely unexplored. We introduce Immersive Volumetric Videos, a new volumetric media format designed to provide large 6-DoF interaction spaces, audiovisual feedback, and high-resolution, high-frame-rate dynamic content. To support IVV construction, we present ImViD, a multi-view, multi-modal dataset built upon a space-oriented capture philosophy. Our custom capture rig enables synchronized multi-view video-audio acquisition during motion, facilitating efficient capture of complex indoor and outdoor scenes with rich foreground--background interactions and challenging dynamics. The dataset provides 5K-resolution videos at 60 FPS with durations of 1-5 minutes, offering richer spatial, temporal, and multimodal coverage than existing benchmarks. Leveraging this dataset, we develop a dynamic light field reconstruction framework built upon a Gaussian-based spatio-temporal representation, incorporating flow-guided sparse initialization, joint camera temporal calibration, and multi-term spatio-temporal supervision for robust and accurate modeling of complex motion. We further propose, to our knowledge, the first method for sound field reconstruction from such multi-view audiovisual data. Together, these components form a unified pipeline for immersive volumetric video production. Extensive benchmarks and immersive VR experiments demonstrate that our pipeline generates high-quality, temporally stable audiovisual volumetric content with large 6-DoF interaction spaces. This work provides both a foundational definition and a practical construction methodology for immersive volumetric videos.

cs.CV

A Unified Framework for Analysis of Randomized Greedy Matching Algorithms

Randomized greedy algorithms form one of the simplest yet most effective approaches for computing approximate matchings in graphs. In this paper, we focus on the class of vertex-iterative (VI) randomized greedy matching algorithms, which process the vertices of a graph $G=(V,E)$ in some order $π$ and, for each vertex $v$, greedily match it to the first available neighbor according to a preference order $σ(v)$. Various VI algorithms have been studied, each corresponding to a different distribution over $π$ and $σ(v)$. We develop a unified framework for analyzing this family of algorithms and use it to obtain improved approximation ratios for Ranking and FRanking, the state-of-the-art randomized greedy algorithms for the random-order and adversarial-order settings, respectively. In Ranking, the decision order is drawn uniformly at random and used as the common preference order, whereas FRanking uses an adversarial decision order and a uniformly random preference order shared by all vertices. We obtain an approximation ratio of $0.560$ for Ranking, improving on the $0.5469$ bound of Derakhshan et al. [SODA 2026]. For FRanking, we obtain a ratio of $0.539$, improving on the $0.521$ bound of Huang et al. [JACM 2020]. These results also imply state-of-the-art approximation ratios for oblivious matching and fully online matching problems on general graphs. Our analysis framework also enables us to prove improved approximation ratios for graphs with no short odd cycles. Such graphs form an intermediate class between general graphs and bipartite graphs. In particular, we show that Ranking is at least $0.570$-competitive for graphs that are both triangle-free and pentagon-free. For graphs whose shortest odd cycle has length at least $129$, we prove that Ranking is at least $0.615$-competitive.

cs.DS

Pumping of spin supercurrent in unitary triplet superconductors

One efficient mechanism for generating a charge supercurrent is Andreev reflection, in which the electric current injected from a normal metal into a conventional superconductor is converted into a supercurrent, thereby preserving charge conservation. We here propose a general principle for generating spin supercurrents in triplet superconductors by analogy with such charge transport, i.e., assuming spin conservation. We find a spin torque that is proportional to the triplet superconducting order parameter and, in the spin-conservation scenario, converts the particle spin to that of Cooper pairs. Based on this general principle, we propose an implementation to efficiently generate a spin supercurrent in unitary triplet superconductors, even though Cooper pairs carry no spin polarization at equilibrium, by the magnetization dynamics ${\bf M}(t)$ of a proximity magnetic nanostructure. The efficiency of this spin pumping is not solely limited to the $d{\bf M}/dt\times {\bf M}$ due to the emergent particle-hole symmetry, thereby going beyond the conventional spin pumping of electrons. This general principle provides an efficient approach to generating and manipulating dissipationless spin currents in many unconventional superconductors.

cond-mat.supr-con

Frequency Comb of Electric-Polarization Waves

Frequency combs are a spectrum of equally spaced frequency components with very high time-frequency accuracy, which have been widely used in the optical and microwave frequency ranges. We propose the realization of a frequency comb operating at the terahertz regime in terms of the nonlinear dynamics of electric-polarization waves, or ferrons as their quanta, in the ferroelectric materials. The efficiency of the frequency comb of the electric-polarization waves is exactly proportional to the static electric polarization carried by the ferron modes, which thereby offers new opportunities for the direct observation and application of the intrinsic properties of ferrons.

cond-mat.mes-hall

The equivalence of precompactness, zero maximal pattern entropy and bounded mean complexity for finite partitions

In this paper, we investigate several types of low complexity of finite partitions, including precompactness, zero maximal pattern entropy, bounded mean complexity and mean equicontinuity. We first show that a collection of finite partitions in a standard probability space is precompact in the Rokhlin metric if and only if it has zero maximal pattern entropy if and only if the collection of the characteristic functions of atoms in those partitions is precompact in $L^2$ if and only if it has bounded mean complexity with respect the Hamming distance. Next, we show that for a countably infinite discrete amenable group acting on a standard probability space, a finite partition has zero maximal pattern entropy if and only if each characteristic function of atom in the partition is almost periodic if and only if it has bounded mean complexity with respect to some (and hence any) Følner sequence if and only if it is mean equicontinuous with respect to some (and hence any) tempered Følner sequence.

math.DS

CUBE: A Standard for Unifying Agent Benchmarks

The proliferation of agent benchmarks has created critical fragmentation that threatens research productivity. Each new benchmark requires substantial custom integration, creating an "integration tax" that limits comprehensive evaluation. We propose CUBE (Common Unified Benchmark Environments), a universal protocol standard built on MCP and Gym that allows benchmarks to be wrapped once and used everywhere. By separating task, benchmark, package, and registry concerns into distinct API layers, CUBE enables any compliant platform to access any compliant benchmark for evaluation, RL training, or data generation without custom integration. We call on the community to contribute to the development of this standard before platform-specific implementations deepen fragmentation as benchmark production accelerates through 2026.

cs.AI

Observation of Long-Lifetime Magnon Pairs by Fano Resonance of Photons

Mode fluctuations with a long lifetime are essential for quantum information and logic operations in magnonic devices. We probe the broadband nonlinear magnetization dynamics of a high-quality ferromagnet under a strong microwave drive using microwave spectroscopy. We observe an \textit{unexpected} Fano resonance in the microwave transmission when the driven amplitude of the magnetization is large and the drive frequency $ω_d$ is close to but not at the ferromagnetic resonance. We interpret this Fano resonance by a scattering theory of photons considering the three-magnon interaction between the Kittel magnon and magnon pairs with opposite wave vectors of frequency $ω_d/2$. The theoretical model suggests that the microwave spectroscopy measures the dynamics of the fluctuation $δ\hatα$ of the Kittel magnon and $δ\hatβ_{\pm k}$ of the magnon pairs over the driven steady states, which are coupled coherently by the steady-state amplitudes. With the damping of $δ\hatβ_{\pm k}$ much smaller than that of $δ\hatα$, the theoretical calculation well reproduces the observed Fano resonance, indicating the magnon pairs hold a recorded long lifetime.

cond-mat.mes-hall

PvP: Data-Efficient Humanoid Robot Learning with Proprioceptive-Privileged Contrastive Representations

Achieving efficient and robust whole-body control (WBC) is essential for enabling humanoid robots to perform complex tasks in dynamic environments. Despite the success of reinforcement learning (RL) in this domain, its sample inefficiency remains a significant challenge due to the intricate dynamics and partial observability of humanoid robots. To address this limitation, we propose PvP, a Proprioceptive-Privileged contrastive learning framework that leverages the intrinsic complementarity between proprioceptive and privileged states. PvP learns compact and task-relevant latent representations without requiring hand-crafted data augmentations, enabling faster and more stable policy learning. To support systematic evaluation, we develop SRL4Humanoid, the first unified and modular framework that provides high-quality implementations of representative state representation learning (SRL) methods for humanoid robot learning. Extensive experiments on the LimX Oli robot across velocity tracking and motion imitation tasks demonstrate that PvP significantly improves sample efficiency and final performance compared to baseline SRL methods. Our study further provides practical insights into integrating SRL with RL for humanoid WBC, offering valuable guidance for data-efficient humanoid robot learning.

cs.RO

Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents

Public research results on large-scale supervised finetuning of AI agents remain relatively rare, since the collection of agent training data presents unique challenges. In this work, we argue that the bottleneck is not a lack of underlying data sources, but that a large variety of data is fragmented across heterogeneous formats, tools, and interfaces. To this end, we introduce the agent data protocol (ADP), a light-weight representation language that serves as an "interlingua" between agent datasets in diverse formats and unified agent training pipelines downstream. The design of ADP is expressive enough to capture a large variety of tasks, including API/tool use, browsing, coding, software engineering, and general agentic workflows, while remaining simple to parse and train on without engineering at a per-dataset level. In experiments, we unified a broad collection of 13 existing agent training datasets into ADP format, and converted the standardized ADP data into training-ready formats for multiple agent frameworks. We performed SFT on these data, and demonstrated an average performance gain of ~20% over corresponding base models, and delivers state-of-the-art or near-SOTA performance on standard coding, browsing, tool use, and research benchmarks, without domain-specific tuning. All code and data are released publicly, in the hope that ADP could help lower the barrier to standardized, scalable, and reproducible agent training.

cs.CL

Monocular Mesh Recovery and Body Measurement of Female Saanen Goats

The lactation performance of Saanen dairy goats, renowned for their high milk yield, is intrinsically linked to their body size, making accurate 3D body measurement essential for assessing milk production potential, yet existing reconstruction methods lack goat-specific authentic 3D data. To address this limitation, we establish the FemaleSaanenGoat dataset containing synchronized eight-view RGBD videos of 55 female Saanen goats (6-18 months). Using multi-view DynamicFusion, we fuse noisy, non-rigid point cloud sequences into high-fidelity 3D scans, overcoming challenges from irregular surfaces and rapid movement. Based on these scans, we develop SaanenGoat, a parametric 3D shape model specifically designed for female Saanen goats. This model features a refined template with 41 skeletal joints and enhanced udder representation, registered with our scan data. A comprehensive shape space constructed from 48 goats enables precise representation of diverse individual variations. With the help of SaanenGoat model, we get high-precision 3D reconstruction from single-view RGBD input, and achieve automated measurement of six critical body dimensions: body length, height, chest width, chest girth, hip width, and hip height. Experimental results demonstrate the superior accuracy of our method in both 3D reconstruction and body measurement, presenting a novel paradigm for large-scale 3D vision applications in precision livestock farming.

cs.CV

Symmetry-Aware Fusion of Vision and Tactile Sensing via Bilateral Force Priors for Robotic Manipulation

Insertion tasks in robotic manipulation demand precise, contact-rich interactions that vision alone cannot resolve. While tactile feedback is intuitively valuable, existing studies have shown that naïve visuo-tactile fusion often fails to deliver consistent improvements. In this work, we propose a Cross-Modal Transformer (CMT) for visuo-tactile fusion that integrates wrist-camera observations with tactile signals through structured self- and cross-attention. To stabilize tactile embeddings, we further introduce a physics-informed regularization that encourages bilateral force balance, reflecting principles of human motor control. Experiments on the TacSL benchmark show that CMT with symmetry regularization achieves a 96.59% insertion success rate, surpassing naïve and gated fusion baselines and closely matching the privileged "wrist + contact force" configuration (96.09%). These results highlight two central insights: (i) tactile sensing is indispensable for precise alignment, and (ii) principled multimodal fusion, further strengthened by physics-informed regularization, unlocks complementary strengths of vision and touch, approaching privileged performance under realistic sensing.

cs.RO

PaperX: A Unified Framework for Multimodal Academic Presentation Generation with Scholar DAG

Transforming scientific papers into multimodal presentation content is essential for research dissemination but remains labor intensive. Existing automated solutions typically treat each format as an isolated downstream task, leading to redundant processing and semantic inconsistency. We introduce PaperX, a unified framework that models academic presentation generation as a structural transformation and rendering process. Central to our approach is the Scholar DAG, an intermediate representation that decouples the paper's logical structure from its final presentation syntax. By applying adaptive graph traversal strategies, PaperX generates diverse, high quality outputs from a single source. Comprehensive evaluations demonstrate that our framework achieves the state of the art performance in content fidelity and aesthetic quality while significantly improving cost efficiency compared to specialized single task agents.

cs.DL

Beyond Closed-Pool Video Retrieval: A Benchmark and Agent Framework for Real-World Video Search and Moment Localization

Traditional video retrieval benchmarks focus on matching precise descriptions to closed video pools, failing to reflect real-world searches characterized by fuzzy, multi-dimensional memories on the open web. We present \textbf{RVMS-Bench}, a comprehensive system for evaluating real-world video memory search. It consists of \textbf{1,440 samples} spanning \textbf{20 diverse categories} and \textbf{four duration groups}, sourced from \textbf{real-world open-web videos}. RVMS-Bench utilizes a hierarchical description framework encompassing \textbf{Global Impression, Key Moment, Temporal Context, and Auditory Memory} to mimic realistic multi-dimensional search cues, with all samples strictly verified via a human-in-the-loop protocol. We further propose \textbf{RACLO}, an agentic framework that employs abductive reasoning to simulate the human ``Recall-Search-Verify'' cognitive process, effectively addressing the challenge of searching for videos via fuzzy memories in the real world. Experiments reveal that existing MLLMs still demonstrate insufficient capabilities in real-world Video Retrieval and Moment Localization based on fuzzy memories. We believe this work will facilitate the advancement of video retrieval robustness in real-world unstructured scenarios.

cs.CV

A Survey of Behavior Foundation Model: Next-Generation Whole-Body Control System of Humanoid Robots

Humanoid robots are drawing significant attention as versatile platforms for complex motor control, human-robot interaction, and general-purpose physical intelligence. However, achieving efficient whole-body control (WBC) in humanoids remains a fundamental challenge due to sophisticated dynamics, underactuation, and diverse task requirements. While learning-based controllers have shown promise for complex tasks, their reliance on labor-intensive and costly retraining for new scenarios limits real-world applicability. To address these limitations, behavior(al) foundation models (BFMs) have emerged as a new paradigm that leverages large-scale pre-training to learn reusable primitive skills and broad behavioral priors, enabling zero-shot or rapid adaptation to a wide range of downstream tasks. In this paper, we present a comprehensive overview of BFMs for humanoid WBC, tracing their development across diverse pre-training pipelines. Furthermore, we discuss real-world applications, current limitations, urgent challenges, and future opportunities, positioning BFMs as a key approach toward scalable and general-purpose humanoid intelligence. Finally, we provide a curated and regularly updated collection of BFM papers and projects to facilitate more subsequent research, which is available at https://github.com/yuanmingqi/awesome-bfm-papers.

cs.RO

Learning Nonlinear Systems In-Context: From Synthetic Data to Real-World Motor Control

LLMs have shown strong in-context learning (ICL) abilities, but have not yet been extended to signal processing systems. Inspired by their design, we have proposed for the first time ICL using transformer models applicable to motor feedforward control, a critical task where classical PI and physics-based methods struggle with nonlinearities and complex load conditions. We propose a transformer based model architecture that separates signal representation from system behavior, enabling both few-shot finetuning and one-shot ICL. Pretrained on a large corpus of synthetic linear and nonlinear systems, the model learns to generalize to unseen system dynamics of real-world motors only with a handful of examples. In experiments, our approach generalizes across multiple motor load configurations, transforms untuned examples into accurate feedforward predictions, and outperforms PI controllers and physics-based feedforward baselines. These results demonstrate that ICL can bridge synthetic pretraining and real-world adaptability, opening new directions for data efficient control of physical systems.

cs.LG

Ferron-Polaritons in Superconductor/Ferroelectric/Superconductor Heterostructures

We predict the formation of ferron-polariton - a hybrid light-matter quasiparticle arising from the coupling between collective ferroelectric excitations (ferrons) and Swihart photons in a superconductor/ferroelectric/superconductor heterostructure. The coupling provides direct evidence for ferrons and reaches the ultrastrong-coupling regime, with a spectral gap in the terahertz range, orders of magnitude larger than those in magnetic analogues, reflecting the superior strength of electric dipole interactions. Our work establishes superconductor-ferroelectric heterostructures as a novel platform for exploring extreme light-matter coupling and for developing high-speed, ferroelectric-based quantum technologies at terahertz frequencies.

cond-mat.supr-con