SearcharxivSearch

arXiv subjects

Yusuke Kato

Publications and source records attributed to Yusuke Kato.

At least 19 recordsLinked to original sources

Early warning signals for synchronization transitions from partial observations

Anticipating the onset of collective synchronization is important in many networked systems, yet observing every oscillator is often impractical. We investigate whether synchronization transitions can be detected from a small set of monitored, or sentinel, nodes. Using a stochastic Kuramoto model on networks, we numerically compare three early warning signals: the local order parameter, its temporal variance, and the variance of individual oscillator phases after removing their mean rotational trends. We also compare sentinel-selection strategies based on node dynamics, degree, and random sampling. We show that, under partial observation, the local order parameter and the variance of detrended phases provide substantially stronger warning signals than the variance of the local order parameter. Selecting nodes according to their dynamical behavior near the transition consistently improves performance over other sentinel-selection methods. With only $\lfloor\ln N\rfloor$ dynamically selected sentinels, two success indicators approach the performance obtained by observing all $N$ nodes. These results demonstrate that synchronization transitions can be anticipated from sparse observations when the warning signal and monitored nodes are chosen appropriately.

physics.soc-ph

Theory of Magnetic Excitations in the Heavy-Fermion Spin-Triplet Superconductor UTe$_2$

We study the dynamical spin response of UTe$_2$ by using a mixed-dimensional periodic Anderson model. Within the BCS-RPA formalism, we examine how the $f$-orbital character of the quasiparticles affects magnetic excitations in both the normal and superconducting (SC) states. In the normal state, finite mixing between localized $f$ electrons and conduction electrons produces a hybridization gap and enhances the spin response at $\mathbf{Q}_{\mathrm{Y}} = (0,\pi,0)$, indicating that the magnetic excitation originates from particle--hole scattering across the hybridization gap. In the SC state, we compare four odd-parity irreducible representations, $A_u$, $B_{1u}$, $B_{2u}$, and $B_{3u}$, for the spin-triplet order parameter. We find that, for the component of the spin susceptibility parallel to the $\mathbf{d}$ vector, a pronounced superconductivity-induced spin resonance appears at $\mathbf{Q}_{\mathrm{Y}}$ only in the $B_{2u}$ state. This behavior arises because the $B_{2u}$ order parameter remains finite and changes its sign between the relevant $f$-electron-dominated Fermi-surface regions connected by $\mathbf{Q}_\mathrm{Y}$ near $k_z=\pi$. The sign-change criterion is applicable to multiband superconductors in three-dimensional heavy-fermion systems, in the presence of (i) low-dimensional portions of the Fermi surface connected by a nesting vector, (ii) the dominance of the $f$-electron character, and (iii) the finite amplitude of SC gap on the portions.

cond-mat.supr-con

Bayesian comparison of Langevin dynamics for cell motility from positional observation

We develop a Bayesian framework for model comparison of second-order Langevin dynamics from position-only trajectories. While approximate increment likelihoods for nonlinear position-only inference have been formulated previously, a unified evidence-based framework for comparing multiple second-order models under positional observation has remained lacking. Here we address this problem by combining exact increment likelihoods for linear Gaussian models with a previously proposed approximate likelihood for nonlinear dynamics. Synthetic-data benchmarks show reliable recovery of the generating model at fine sampling intervals and progressive loss of identifiability under coarse temporal sampling. Application to Dictyostelium discoideum trajectories demonstrates that the statistically supported model depends strongly on temporal resolution. Moreover, the selected models reproduce key statistical properties of the experimental trajectories, providing additional support for the model-comparison results. Our framework therefore offers a practical approach to evidence-based comparison of partially observed stochastic dynamics.

q-bio.CB

Contrastive Action-Image Pre-training for Visuomotor Control

Existing vision encoders for robotics face a fundamental bottleneck: robotic datasets lack the scale necessary for large-scale pre-training. Prior work circumvents this data scarcity by turning to internet-scale image and language data or egocentric human video. While these models show promise, neither paradigm learns from paired vision and action data, which downstream visuomotor control policies require. However, robot trajectories, the most direct source of this paired signal, are not available at pre-training scale, motivating us to extract action signals from abundant human video instead. To this end, we introduce CAIP (Contrastive Action-Image Pre-training), a vision encoder that treats human hand poses from large-scale egocentric video as a proxy for end-effector actions. By extracting 3D hand keypoints, a representation that aligns naturally with downstream robot action spaces, CAIP learns a unified action-image representation through a contrastive objective. Leveraging 32,041 hours of egocentric human video and only 88 hours of robotic manipulation data, CAIP outperforms state-of-the-art vision encoders including DINOv2, SigLIP, MVP, and R3M. Evaluated on a challenging real-world dexterous manipulation setup using Dexmate Vega and Sharpa Wave hands, CAIP yields performance gains of more than 30% on tasks involving folding, pouring, and fine-grained manipulation. Our results show that our method of contrastive action-centric pre-training yields a scalable path to achieving robust visual representations better suited for physical interaction.

cs.RO

Quantum-geometric origin of superfluid weight in quasicrystals with critical states

A distinctive feature of many quasiperiodic systems is the presence of critical states that are neither extended nor exponentially localized. We investigate the geometric effect on the superfluid weight in quasiperiodic systems with critical states at zero temperature. We employ both real-space and momentum-space approaches to superfluid weight in quasicrystals, which allows us to separate the conventional and quantum geometric contributions. We find that the superfluid weight is dominated by the geometric contribution in quasiperiodic systems with critical states. This finding reveals a fundamental interplay between superconductivity and critical states in quasicrystals.

cond-mat.supr-con

Uncertainty-Aware Sparse Identification of Dynamical Systems via Bayesian Model Averaging

In many problems of data-driven modeling for dynamical systems, the governing equations are not known a priori and must be selected phenomenologically from a large set of candidate interactions and basis functions. In such situations, point estimates alone can be misleading, because multiple model components may explain the observed data comparably well, especially when the data are limited or the dynamics exhibit poor identifiability. Quantifying the uncertainty associated with model selection is therefore essential for constructing reliable dynamical models from data. In this work, we develop a Bayesian sparse identification framework for dynamical systems with coupled components, aimed at inferring both interaction structure and functional form together with principled uncertainty quantification. The proposed method combines sparse modeling with Bayesian model averaging, yielding posterior inclusion probabilities that quantify the credibility of each candidate interaction and basis component. Through numerical experiments on oscillator networks, we show that the framework accurately recovers sparse interaction structures with quantified uncertainty, including higher-order harmonic components, phase-lag effects, and multi-body interactions. We also demonstrate that, even in a phenomenological setting where the true governing equations are not contained in the assumed model class, the method can identify effective functional components with quantified uncertainty. These results highlight the importance of Bayesian uncertainty quantification in data-driven discovery of dynamical models.

stat.AP

Boundary condition for phonon distribution functions at a smooth crystal interface and interfacial angular momentum transfer

We theoretically elucidate the boundary conditions for phonon distribution functions of long-wavelength acoustic phonons at smooth crystal interfaces. We first derive boundary conditions that fully incorporate reflection, transmission, and mode conversion. We obtain these conditions for phonons from those for classical lattice vibrations, using the correspondence between the quantum and classical descriptions. This formulation provides a theoretical foundation for the acoustic mismatch model, widely used to analyze Kapitza resistance. We then refine the boundary conditions to include spatial dependence parallel to the interface. The refined form captures transverse shifts of elastic wave packets, analogous to the optical Imbert--Fedorov shift, and ensures conservation of total angular momentum. Consequently, circularly polarized phonons carrying spin angular momentum (SAM) generate phonon orbital angular momentum (OAM) at the interface. We analytically determine the spatial profile of this OAM and demonstrate that SAM and OAM are both involved in the interfacial diffusion of chiral phonons. Our theory provides concise boundary conditions for phonons, with applications ranging from heat transport to phonon angular momentum transport.

cond-mat.mes-hall

MobileWorldBench: Towards Semantic World Modeling For Mobile Agents

World models have shown great utility in improving the task performance of embodied agents. While prior work largely focuses on pixel-space world models, these approaches face practical limitations in GUI settings, where predicting complex visual elements in future states is often difficult. In this work, we explore an alternative formulation of world modeling for GUI agents, where state transitions are described in natural language rather than predicting raw pixels. First, we introduce MobileWorldBench, a benchmark that evaluates the ability of vision-language models (VLMs) to function as world models for mobile GUI agents. Second, we release MobileWorld, a large-scale dataset consisting of 1.4M samples, that significantly improves the world modeling capabilities of VLMs. Finally, we propose a novel framework that integrates VLM world models into the planning framework of mobile agents, demonstrating that semantic world models can directly benefit mobile agents by improving task success rates. The code and dataset is available at https://github.com/jacklishufan/MobileWorld

cs.AI

LaViDa: A Large Diffusion Language Model for Multimodal Understanding

Modern Vision-Language Models (VLMs) can solve a wide range of tasks requiring visual reasoning. In real-world scenarios, desirable properties for VLMs include fast inference and controllable generation (e.g., constraining outputs to adhere to a desired format). However, existing autoregressive (AR) VLMs like LLaVA struggle in these aspects. Discrete diffusion models (DMs) offer a promising alternative, enabling parallel decoding for faster inference and bidirectional context for controllable generation through text-infilling. While effective in language-only settings, DMs' potential for multimodal tasks is underexplored. We introduce LaViDa, a family of VLMs built on DMs. We build LaViDa by equipping DMs with a vision encoder and jointly fine-tune the combined parts for multimodal instruction following. To address challenges encountered, LaViDa incorporates novel techniques such as complementary masking for effective training, prefix KV cache for efficient inference, and timestep shifting for high-quality sampling. Experiments show that LaViDa achieves competitive or superior performance to AR VLMs on multi-modal benchmarks such as MMMU, while offering unique advantages of DMs, including flexible speed-quality tradeoff, controllability, and bidirectional reasoning. On COCO captioning, LaViDa surpasses Open-LLaVa-Next-8B by +4.1 CIDEr with 1.92x speedup. On bidirectional tasks, it achieves +59% improvement on Constrained Poem Completion. These results demonstrate LaViDa as a strong alternative to AR VLMs. Code and models will be released in the camera-ready version.

cs.CV

Systematic construction of asymptotic quantum many-body scar states and their relation to supersymmetric quantum mechanics

We develop a systematic method for constructing asymptotic quantum many-body scar (AQMBS) states. While AQMBS states are closely related to quantum many-body scar (QMBS) states, they exhibit key differences. Unlike QMBS states, AQMBS states are not energy eigenstates of the Hamiltonian, making their construction more challenging. We demonstrate that, under appropriate conditions, AQMBS states can be obtained as low-lying gapless excited states of a parent Hamiltonian, which has a QMBS state as its ground state. Furthermore, our formalism reveals a connection between QMBS and supersymmetric (SUSY) quantum mechanics. The QMBS state can be interpreted as a SUSY-unbroken ground state.

cond-mat.stat-mech

Uncovering influence of football players' behaviour on team performance in ball possession through dynamical modelling

A quest for uncovering influence of behaviour on team performance involves understanding individual behaviour, interactions with others and environment, variations across groups, and effects of interventions. Although insights into each of these areas have accumulated in sports science literature on football, it remains unclear how one can enhance team performance. We analyse influence of football players' behaviour on team performance in three-versus-one ball possession game by constructing and analysing a dynamical model. We developed a model for the motion of the players and the ball, which mathematically represented our hypotheses on players' behaviour and interactions. The model's plausibility was examined by comparing simulated outcomes with our experimental result. Possible influences of interventions were analysed through sensitivity analysis, where causal effects of several aspects of behaviour such as pass speed and accuracy were found. Our research highlights the potential of dynamical modelling for uncovering influence of behaviour on team effectiveness.

physics.soc-ph

Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context Reflection

The predominant approach to advancing text-to-image generation has been training-time scaling, where larger models are trained on more data using greater computational resources. While effective, this approach is computationally expensive, leading to growing interest in inference-time scaling to improve performance. Currently, inference-time scaling for text-to-image diffusion models is largely limited to best-of-N sampling, where multiple images are generated per prompt and a selection model chooses the best output. Inspired by the recent success of reasoning models like DeepSeek-R1 in the language domain, we introduce an alternative to naive best-of-N sampling by equipping text-to-image Diffusion Transformers with in-context reflection capabilities. We propose Reflect-DiT, a method that enables Diffusion Transformers to refine their generations using in-context examples of previously generated images alongside textual feedback describing necessary improvements. Instead of passively relying on random sampling and hoping for a better result in a future generation, Reflect-DiT explicitly tailors its generations to address specific aspects requiring enhancement. Experimental results demonstrate that Reflect-DiT improves performance on the GenEval benchmark (+0.19) using SANA-1.0-1.6B as a base model. Additionally, it achieves a new state-of-the-art score of 0.81 on GenEval while generating only 20 samples per prompt, surpassing the previous best score of 0.80, which was obtained using a significantly larger model (SANA-1.5-4.8B) with 2048 samples under the best-of-N approach.

cs.CV

Nonequilibrium electron distribution function in a voltage-biased nanowire: A nonequilibrium Green's function approach

We develop a theoretical framework to determine distribution functions in nonequilibrium systems coupled to equilibrium reservoirs, by using the nonequilibrium Green's function technique. As a paradigmatic example, we consider the nonequilibrium distribution function in a nanowire under a bias voltage. We model the system as a tight-binding chain connected to reservoirs with different electrochemical potentials at both ends. For electron scattering processes in the wire, we consider both elastic scattering from impurities and inelastic scattering from phonons within the self-consistent Born approximation. We demonstrate that the nonequilibrium distribution functions, as well as the electrostatic potential profiles, in various scattering regimes are well described within our framework. This scheme will contribute to advancing our understanding of quantum many-body phenomena driven by nonequilibrium distribution functions that have different functional forms from the equilibrium ones.

cond-mat.mes-hall

Bayesian estimation of coupling strength and heterogeneity in a coupled oscillator model from macroscopic quantities

Various macroscopic oscillations, such as the heartbeat and the flashing of fireflies, are created by synchronizing oscillatory units (oscillators). To elucidate the mechanism of synchronization, several coupled oscillator models have been devised and extensively analyzed. Although parameter estimation of these models has also been actively investigated, most of the proposed methods are based on the data from individual oscillators, not from macroscopic quantities. In the present study, we propose a Bayesian framework to estimate the model parameters of coupled oscillator models, using the time series data of the Kuramoto order parameter as the only given data. We adopt the exchange Monte Carlo method for the efficient estimation of the posterior distribution and marginal likelihood. Numerical experiments are performed to confirm the validity of our method and examine the dependence of the estimation error on the observational noise and system size.

physics.data-an

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows

We introduce OmniFlow, a novel generative model designed for any-to-any generation tasks such as text-to-image, text-to-audio, and audio-to-image synthesis. OmniFlow advances the rectified flow (RF) framework used in text-to-image models to handle the joint distribution of multiple modalities. It outperforms previous any-to-any models on a wide range of tasks, such as text-to-image and text-to-audio synthesis. Our work offers three key contributions: First, we extend RF to a multi-modal setting and introduce a novel guidance mechanism, enabling users to flexibly control the alignment between different modalities in the generated outputs. Second, we propose a novel architecture that extends the text-to-image MMDiT architecture of Stable Diffusion 3 and enables audio and text generation. The extended modules can be efficiently pretrained individually and merged with the vanilla text-to-image MMDiT for fine-tuning. Lastly, we conduct a comprehensive study on the design choices of rectified flow transformers for large-scale audio and text generation, providing valuable insights into optimizing performance across diverse modalities. The Code will be available at https://github.com/jacklishufan/OmniFlows.

cs.MM

SegLLM: Multi-round Reasoning Segmentation

We present SegLLM, a novel multi-round interactive reasoning segmentation model that enhances LLM-based segmentation by exploiting conversational memory of both visual and textual outputs. By leveraging a mask-aware multimodal LLM, SegLLM re-integrates previous segmentation results into its input stream, enabling it to reason about complex user intentions and segment objects in relation to previously identified entities, including positional, interactional, and hierarchical relationships, across multiple interactions. This capability allows SegLLM to respond to visual and text queries in a chat-like manner. Evaluated on the newly curated MRSeg benchmark, SegLLM outperforms existing methods in multi-round interactive reasoning segmentation by over 20%. Additionally, we observed that training on multi-round reasoning segmentation data enhances performance on standard single-round referring segmentation and localization tasks, resulting in a 5.5% increase in cIoU for referring expression segmentation and a 4.5% improvement in Acc@0.5 for referring expression localization.

cs.CV

Theory of phonon angular momentum transport across a smooth crystal interface

We theoretically elucidate the transfer of phonon angular momentum by acoustic modes across a smooth interface between crystals. We analyze this process, which is difficult to describe with the conventional acoustic mismatch model, using a reformulated boundary condition and the Boltzmann theory. For an interface between a chiral and an achiral crystal, our analysis reveals that thermal gradients in the chiral crystal induce angular momentum, which diffuses into the achiral crystal even without heat flow. Notably, the density of angular momentum can be enhanced near the interface. These findings advance our understanding of phonon transport and its interplay with electron spins.

cond-mat.mes-hall

Chirality-dependent spin polarization in metals: linear and quadratic responses

We study spin polarization induced by locally injected electric currents in a metal whose spin--orbit coupling reflects its structural chirality. We reveal both spin polarization in the bulk in the linear response and antiparallel spin polarization near the interface in the quadratic response to external electric currents, and reproduce the experimentally observed correlation between the chirality of the metal and the direction of spin polarization. In particular, we elucidate that the sign of the spin polarization in the quadratic response is opposite to that expected from the bulk spin current. This sign discrepancy originates from spin polarization induced by dipole-like charge distribution appearing in the quadratic response.

cond-mat.mes-hall