SearcharxivSearch

arXiv subjects

Boyuan Zhao

Publications and source records attributed to Boyuan Zhao.

8 recordsLinked to original sources

TacForcing: Streaming Action Generation with Execution-Time Tactile Feedback

Contact-rich manipulation requires adapting to contact states that can evolve substantially within an action horizon. However, chunk-based vision-language-action models predict complete action chunks from observations collected before execution, leaving tactile conditioning stale during execution. Existing tactile-reactive approaches typically rely on separate high-frequency controllers, which increase both architectural and training complexity. In this paper, we introduce TacForcing, a streaming action-generation framework that effectively incorporates execution-time tactile feedback. Instead of employing a separate reactive controller, TacForcing replaces the standard action expert with a streaming action expert to generate actions conditioned on the evolving tactile observations acquired during execution. TacForcing also introduces Execution-Aware Tactile Attention (EATA), which restricts tactile conditioning to actions nearing execution, thereby reducing the temporal mismatch between tactile acquisition and action execution. Across six simulated UniVTAC tasks and three real-world contact-rich manipulation tasks, TacForcing achieves average success rates of 65% and 69%, respectively, outperforming strong baselines in both settings.

cs.RO

World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis

We propose world-language-action (WLA) models as a new class of embodied foundation models. WLA takes textual instructions, images, and robot states as inputs to jointly predict textual subtasks, subgoal images, and robot actions, conjoining the \emph{world modeling interface} to learn from extensive egocentric videos as in the world-action model (WAM) and the \emph{language reasoning} capacities to solve complex long-horizon tasks as in vision-language-action (VLA) models. At the core of WLA lies an \emph{autoregressive (AR)} Transformer backbone, instead of a bidirectional diffusion Transformer as in WAMs, to predict the \emph{next state}, comprising the \emph{semantic-level} textual intention and complementary \emph{fine-grained} physical dynamics. The physical dynamics are supervised by the world modeling objective based on a dedicated World Expert, and are leveraged to ease the characterization of the state-action correlation for the Action Expert. WLA leverages meta-queries to make the world prediction \emph{implicitly} impact the action generation so that the former can be disabled during inference. The world prediction can also be activated to enable test-time scaling for improved robot control. Our WLA-0 prototype, with 2B active parameters, achieves 40 ms per inference on an NVIDIA RTX 5090. Evaluations across simulated and real-world environments demonstrate that WLA-0 achieves state-of-the-art multi-task and long-horizon learning abilities, e.g., 92.94\% success rate on RoboTwin2.0 Clean and 56.5\% success rate on RMBench. WLA-0 also holds the promise to learn novel tasks directly from \emph{cross-embodiment robot videos} without action annotations.

cs.RO

Convergence of the Birkhoff spectrum for nonintegrable observables

We consider interval maps with countably many full branches and observables with polynomial tails. We show that the Birkhoff spectrum is real analytic and that its convergence to the Hausdorff dimension of the repeller is governed by the polynomial tail exponent. This result extends previous work by Arima on more regular observables and demonstrates how the tail behaviour influences the structure of the Birkhoff spectrum. Our proof relies on techniques from thermodynamic formalism and tail estimates for the observable and our applications are to natural observations on Gauss maps, L\"uroth transformations as well as to a the first return time for a class of induced Manneville-Pomeau maps.

math.DS

Almost sure convergence of cover times for $\psi$-mixing systems

Given a topologically transitive system on the unit interval, one can investigate the cover time, i.e. time for an orbit to reach certain level of resolution in the repeller. We introduce a new notion of dimension, namely the stretched Minkowski dimension, and show that under mixing conditions, the asymptotics of typical cover times are determined by Minkowski dimensions when they are finite, or by stretched Minkowski dimensions otherwise. For application, we show that for countably full-branched affine maps, results using the usual Minkowski dimensions fail to produce a finite log limit of cover times whilst the stretched version gives an finite limit. In addition, cover times of irrational rotations are explicitly calculated as counterexamples, due to the absence of mixing.

math.DS

Closest Distance between Iterates of Typical Points

The shortest distance between the first $n$ iterates of a typical point can be quantified with a log rule for some dynamical systems admitting Gibbs measures. We show this in two settings. For topologically mixing Markov shifts with at most countably infinite alphabet admitting a Gibbs measure with respect to a locally Hölder potential, we prove the asymptotic length of the longest common substring for a typical point converges and the limit depends on the Rényi entropy. For interval maps with the Gibbs-Markov structure, we prove a similar rule relating the correlation dimension of Gibbs measures with the shortest distance between two iterates in the orbit generated by a typical point.

math.DS

Countable Markov shifts with exponential mixing

We give a set of equivalent conditions for a potential on a Countable Markov Shift to have strong positive recurrence, which is also equivalent to having exponential decay of correlations. A key ingredient of our proofs is quantifying how the shift behaves at its boundary.

math.DS

Recordism: A social-scientific prospect of blockchain from social, legal, financial, and technological perspectives

Blockchain technology has the potential to revolutionize the architecture of cyberspace by transforming the way information is stored, circulated, and exchanged in cyberspace through decentralization, transparency, and de-identification. This means that ordinary participants can simultaneously become traders, miners, retailers, and customers, thus breaking down barriers, reducing the information gap between participants in the community, and contributing to the futuristic metaverse with an open, progressive, and equal ideology. The impact of this information transformation empowered by blockchain extends to our understanding of methodology, legal governance in cyberspace, and financial and technological development. This study asks: what are the implications of the blockchain-driven information revolution for society and social sciences? In order to answer this main question, the paper focuses on four key perspectives: methodological, legal, financial, and technical. Through the analysis of these four perspectives, the paper provides a comprehensive understanding of the impact of blockchain on society, the social sciences, and technology, making a contribution to current scholarship. It finds that blockchain is not only an innovative cognition method, but also a community representative, serving as a source of trust, a governance watchdog, an enforcer of cyber laws, and an incubator for future technologies. Despite some challenges in integrating blockchain with existing social structures, this paper concludes that blockchain has the potential to play a significant role in shaping the future.

cs.CY

Psychologically-Inspired Music Recommendation System

In the last few years, automated recommendation systems have been a major focus in the music field, where companies such as Spotify, Amazon, and Apple are competing in the ability to generate the most personalized music suggestions for their users. One of the challenges developers still fail to tackle is taking into account the psychological and emotional aspects of the music. Our goal is to find a way to integrate users' personal traits and their current emotional state into a single music recommendation system with both collaborative and content-based filtering. We seek to relate the personality and the current emotional state of the listener to the audio features in order to build an emotion-aware MRS. We compare the results both quantitatively and qualitatively to the output of the traditional MRS based on the Spotify API data to understand if our advancements make a significant impact on the quality of music recommendations.

cs.IR