Searcharxiv⌕ Search

arXiv subjects

Yusuke Kato

Publications and source records attributed to Yusuke Kato.

At least 37 records · Page 2Linked to original sources

Uncovering influence of football players' behaviour on team performance in ball possession through dynamical modelling

A quest for uncovering influence of behaviour on team performance involves understanding individual behaviour, interactions with others and environment, variations across groups, and effects of interventions. Although insights into each of these areas have accumulated in sports science literature on football, it remains unclear how one can enhance team performance. We analyse influence of football players' behaviour on team performance in three-versus-one ball possession game by constructing and analysing a dynamical model. We developed a model for the motion of the players and the ball, which mathematically represented our hypotheses on players' behaviour and interactions. The model's plausibility was examined by comparing simulated outcomes with our experimental result. Possible influences of interventions were analysed through sensitivity analysis, where causal effects of several aspects of behaviour such as pass speed and accuracy were found. Our research highlights the potential of dynamical modelling for uncovering influence of behaviour on team effectiveness.

physics.soc-ph↗

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows

We introduce OmniFlow, a novel generative model designed for any-to-any generation tasks such as text-to-image, text-to-audio, and audio-to-image synthesis. OmniFlow advances the rectified flow (RF) framework used in text-to-image models to handle the joint distribution of multiple modalities. It outperforms previous any-to-any models on a wide range of tasks, such as text-to-image and text-to-audio synthesis. Our work offers three key contributions: First, we extend RF to a multi-modal setting and introduce a novel guidance mechanism, enabling users to flexibly control the alignment between different modalities in the generated outputs. Second, we propose a novel architecture that extends the text-to-image MMDiT architecture of Stable Diffusion 3 and enables audio and text generation. The extended modules can be efficiently pretrained individually and merged with the vanilla text-to-image MMDiT for fine-tuning. Lastly, we conduct a comprehensive study on the design choices of rectified flow transformers for large-scale audio and text generation, providing valuable insights into optimizing performance across diverse modalities. The Code will be available at https://github.com/jacklishufan/OmniFlows.

cs.MM↗

Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context Reflection

The predominant approach to advancing text-to-image generation has been training-time scaling, where larger models are trained on more data using greater computational resources. While effective, this approach is computationally expensive, leading to growing interest in inference-time scaling to improve performance. Currently, inference-time scaling for text-to-image diffusion models is largely limited to best-of-N sampling, where multiple images are generated per prompt and a selection model chooses the best output. Inspired by the recent success of reasoning models like DeepSeek-R1 in the language domain, we introduce an alternative to naive best-of-N sampling by equipping text-to-image Diffusion Transformers with in-context reflection capabilities. We propose Reflect-DiT, a method that enables Diffusion Transformers to refine their generations using in-context examples of previously generated images alongside textual feedback describing necessary improvements. Instead of passively relying on random sampling and hoping for a better result in a future generation, Reflect-DiT explicitly tailors its generations to address specific aspects requiring enhancement. Experimental results demonstrate that Reflect-DiT improves performance on the GenEval benchmark (+0.19) using SANA-1.0-1.6B as a base model. Additionally, it achieves a new state-of-the-art score of 0.81 on GenEval while generating only 20 samples per prompt, surpassing the previous best score of 0.80, which was obtained using a significantly larger model (SANA-1.5-4.8B) with 2048 samples under the best-of-N approach.

cs.CV↗

Proposal for simulating quantum spin models with the Dzyaloshinskii-Moriya interaction using Rydberg atoms and the construction of asymptotic quantum many-body scar states

We have developed a method to simulate quantum spin models with the Dzyaloshinskii-Moriya interaction (DMI) using Rydberg atom quantum simulators. Our approach involves a two-photon Raman transition and a transformation to the spin-rotating frame, both of which are feasible with current experimental techniques. As a model that can be simulated in our setup but not in solid-state systems, we consider an $S=\frac{1}{2}$ spin chain with a Hamiltonian consisting of the DMI and Zeeman energy. We study the magnetization curve in the ground state of this model and quench dynamics. Further, we show the existence of quantum many-body scar states and asymptotic quantum many-body scar states. The observed nonergodicity in this model demonstrates the importance of the highly tunable DMI that can be realized by the proposed quantum simulator.

cond-mat.quant-gas↗

SegLLM: Multi-round Reasoning Segmentation

We present SegLLM, a novel multi-round interactive reasoning segmentation model that enhances LLM-based segmentation by exploiting conversational memory of both visual and textual outputs. By leveraging a mask-aware multimodal LLM, SegLLM re-integrates previous segmentation results into its input stream, enabling it to reason about complex user intentions and segment objects in relation to previously identified entities, including positional, interactional, and hierarchical relationships, across multiple interactions. This capability allows SegLLM to respond to visual and text queries in a chat-like manner. Evaluated on the newly curated MRSeg benchmark, SegLLM outperforms existing methods in multi-round interactive reasoning segmentation by over 20%. Additionally, we observed that training on multi-round reasoning segmentation data enhances performance on standard single-round referring segmentation and localization tasks, resulting in a 5.5% increase in cIoU for referring expression segmentation and a 4.5% improvement in Acc@0.5 for referring expression localization.

cs.CV↗

Aligning Diffusion Models by Optimizing Human Utility

We present Diffusion-KTO, a novel approach for aligning text-to-image diffusion models by formulating the alignment objective as the maximization of expected human utility. Since this objective applies to each generation independently, Diffusion-KTO does not require collecting costly pairwise preference data nor training a complex reward model. Instead, our objective requires simple per-image binary feedback signals, e.g. likes or dislikes, which are abundantly available. After fine-tuning using Diffusion-KTO, text-to-image diffusion models exhibit superior performance compared to existing techniques, including supervised fine-tuning and Diffusion-DPO, both in terms of human judgment and automatic evaluation metrics such as PickScore and ImageReward. Overall, Diffusion-KTO unlocks the potential of leveraging readily available per-image binary signals and broadens the applicability of aligning text-to-image diffusion models with human preferences.

cs.CV↗

Anisotropy-induced spin parity effects

Spin parity effects refer to those special situations where a dichotomy in the physical behavior of a system arises, solely depending on whether the relevant spin quantum number is integral or half-odd integral. As is the case with the Haldane conjecture in antiferromagnetic spin chains, their pursuit often derives deep insights and invokes new developments in quantum condensed matter physics. Here, we put forth a simple and general scheme for generating such effects in any spatial dimension through the use of anisotropic interactions, and a setup within reasonable reach of state-of-the-art cold-atom implementations. We demonstrate its utility through a detailed analysis of the magnetization behavior of a specific one-dimensional spin chain model, an anisotropic antiferromagnet in a transverse magnetic field, unraveling along the way the quantum origin of finite-size effects observed in the magnetization curve that had previously been noted but not clearly understood.

cond-mat.stat-mech↗

Periodic forces combined with feedback induce quenching in a bistable oscillator

The coexistence of an abnormal rhythm and a normal steady state is often observed in nature (e.g., epilepsy). Such a system is modeled as a bistable oscillator that possesses both a limit cycle and a fixed point. Although bistable oscillators under several perturbations have been addressed in the literature, the mechanism of oscillation quenching, a transition from a limit cycle to a fixed point, has not been fully understood. In this study, we analyze quenching using the extended Stuart-Landau oscillator driven by periodic forces. Numerical simulations suggest that the entrainment to the periodic force induces the amplitude change of a limit cycle. By reducing the system with an averaging method, we investigate the bifurcation structures of the periodically-driven oscillator. We find that oscillation quenching occurs by the homoclinic bifurcation when we use a periodic force combined with quadratic feedback. In conclusion, we develop a state-transition method in a bistable oscillator using periodic forces, which would have the potential for practical applications in controlling and annihilating abnormal oscillations. Moreover, we clarify the rich and diverse bifurcation structures behind periodically-driven bistable oscillators, which we believe would contribute to further understanding the complex behaviors in non-autonomous systems.

nlin.AO↗

Time-dependent Ginzburg-Landau theory of the vortex spin Hall effect

We develop a time-dependent Ginzburg-Landau theory of the vortex spin Hall effect, i.e., a spin Hall effect that is driven by the motion of superconducting vortices. For the direct vortex spin Hall effect in which an input charge current drives the transverse spin current accompanying the vortex motion, we start from the well-known Schmid-Caroli-Maki solution for the time-dependent Ginzburg-Landau equation under the applied electric field, and find out the expression of the induced spin current. For the inverse vortex spin Hall effect in which an input spin current drives the longitudinal vortex motion and produces the transverse charge current, we microscopically construct the time-dependent Ginzburg-Landau equation under the applied spin accumulation gradient, and calculate the induced transverse charge current as well as the open circuit voltage. The time-dependent Ginzburg-Landau equation and its analytical solution developed here can be a basis for more quantitative numerical simulations of the vortex spin Hall effect.

cond-mat.supr-con↗

World-Model-Based Control for Industrial box-packing of Multiple Objects using NewtonianVAE

The process of industrial box-packing, which involves the accurate placement of multiple objects, requires high-accuracy positioning and sequential actions. When a robot is tasked with placing an object at a specific location with high accuracy, it is important not only to have information about the location of the object to be placed, but also the posture of the object grasped by the robotic hand. Often, industrial box-packing requires the sequential placement of identically shaped objects into a single box. The robot's action should be determined by the same learned model. In factories, new kinds of products often appear and there is a need for a model that can easily adapt to them. Therefore, it should be easy to collect data to train the model. In this study, we designed a robotic system to automate real-world industrial tasks, employing a vision-based learning control model. We propose in-hand-view-sensitive Newtonian variational autoencoder (ihVS-NVAE), which employs an RGB camera to obtain in-hand postures of objects. We demonstrate that our model, trained for a single object-placement task, can handle sequential tasks without additional training. To evaluate efficacy of the proposed model, we employed a real robot to perform sequential industrial box-packing of multiple objects. Results showed that the proposed model achieved a 100% success rate in industrial box-packing tasks, thereby outperforming the state-of-the-art and conventional approaches, underscoring its superior effectiveness and potential in industrial tasks.

cs.RO↗

Hierarchical Open-vocabulary Universal Image Segmentation

Open-vocabulary image segmentation aims to partition an image into semantic regions according to arbitrary text descriptions. However, complex visual scenes can be naturally decomposed into simpler parts and abstracted at multiple levels of granularity, introducing inherent segmentation ambiguity. Unlike existing methods that typically sidestep this ambiguity and treat it as an external factor, our approach actively incorporates a hierarchical representation encompassing different semantic-levels into the learning process. We propose a decoupled text-image fusion mechanism and representation learning modules for both "things" and "stuff". Additionally, we systematically examine the differences that exist in the textual and visual features between these types of categories. Our resulting model, named HIPIE, tackles HIerarchical, oPen-vocabulary, and unIvErsal segmentation tasks within a unified framework. Benchmarked on over 40 datasets, e.g., ADE20K, COCO, Pascal-VOC Part, RefCOCO/RefCOCOg, ODinW and SeginW, HIPIE achieves the state-of-the-art results at various levels of image comprehension, including semantic-level (e.g., semantic segmentation), instance-level (e.g., panoptic/referring segmentation and object detection), as well as part-level (e.g., part/subpart segmentation) tasks. Our code is released at https://github.com/berkeley-hipie/HIPIE.

cs.CV↗

Synchronization and oscillation quenching in interacting metronomes on a movable platform: simple model and bifurcation analysis

Various oscillatory phenomena occur in the world. Because some oscillations are related to abnormal states (e.g., particular diseases), establishing state-transition methods from an oscillatory to a resting state is important. In this study, we construct a simple metronome model and analyze the oscillation-quenching phenomenon of metronomes on a platform as an example of such state transitions. Although numerous studies were conducted on the metronome dynamics, most of them focused on the synchronization, and few studies treated the oscillation quenching because of the difficulty in analysis. To facilitate the analysis, we model a metronome as a linear spring pendulum with an impulsive force (escapement mechanism) described by a fifth-order polynomial. By performing an averaging approximation, we obtain a phase diagram for the in-phase synchronization, anti-phase synchronization, and oscillation quenching. We also numerically integrate the equation of motion and confirm the agreement between the analytical and numerical results. Despite the simplicity, our model successfully reproduces essential phenomena in interacting mechanical clocks, such as the bistability of in-phase and anti-phase synchrony and oscillation quenching occurring for a large mass ratio between the oscillator and the platform. We believe that our simple model will contribute to future analyses of other dynamics observed in metronomes.

nlin.AO↗

Synchronization and stability analysis of an exponentially diverging solution in a mathematical model of asymmetrically interacting agents

This study deals with an existing mathematical model of asymmetrically interacting agents. We analyze the following two previously unfocused features of the model: (i) synchronization of growth rates and (ii) initial value dependence of damped oscillation. By applying the techniques of variable transformation and time-scale separation, we perform the stability analysis of a diverging solution. We find that (i) all growth rates synchronize to the same value that is as small as the smallest growth rate and (ii) oscillatory dynamics appear if the initial value of the slowest-growing agent is sufficiently small. Furthermore, our analytical method proposes a way to apply stability analysis to an exponentially diverging solution, which we believe is also a contribution of this study. Although the employed model is originally proposed as a model of infectious disease, we do not discuss its biological relevance but merely focus on the technical aspects.

nlin.AO↗

Spin Relaxation, Diffusion and Edelstein Effect in Chiral Metal Surface

We study electron spin transport at spin-splitting surface of chiral-crystalline-structured metals and Edelstein effect at the interface, by using the Boltzmann transport equation beyond the relaxation time approximation. We first define spin relaxation time and spin diffusion length for two-dimensional systems with anisotropic spin--orbit coupling through the spectrum of the integral kernel in the collision integral. We then explicitly take account of the interface between the chiral metal and a nonmagnetic metal with finite thickness. For this composite system, we derive analytical expressions for efficiency of the charge current--spin current interconversion as well as other coefficients found in the Edelstein effect. We also develop the Onsager's reciprocity in the Edelstein effect along with experiments so that it relates local input and output, which are respectively defined in the regions separated by the interface. We finally provide a transfer matrix corresponding to the Edelstein effect through the interface, with which we can easily represent the Onsager's reciprocity as well as the charge--spin conversion efficiencies we have obtained. We confirm the validity of the Boltzmann transport equation in the present system starting from the Keldysh formalism in the supplemental material. Our formulation also applies to the Rashba model and other spin-splitting systems.

cond-mat.mes-hall↗

Refine and Represent: Region-to-Object Representation Learning

Recent works in self-supervised learning have demonstrated strong performance on scene-level dense prediction tasks by pretraining with object-centric or region-based correspondence objectives. In this paper, we present Region-to-Object Representation Learning (R2O) which unifies region-based and object-centric pretraining. R2O operates by training an encoder to dynamically refine region-based segments into object-centric masks and then jointly learns representations of the contents within the mask. R2O uses a "region refinement module" to group small image regions, generated using a region-level prior, into larger regions which tend to correspond to objects by clustering region-level features. As pretraining progresses, R2O follows a region-to-object curriculum which encourages learning region-level features early on and gradually progresses to train object-centric representations. Representations learned using R2O lead to state-of-the art performance in semantic segmentation for PASCAL VOC (+0.7 mIOU) and Cityscapes (+0.4 mIOU) and instance segmentation on MS COCO (+0.3 mask AP). Further, after pretraining on ImageNet, R2O pretrained models are able to surpass existing state-of-the-art in unsupervised object segmentation on the Caltech-UCSD Birds 200-2011 dataset (+2.9 mIoU) without any further training. We provide the code/models from this work at https://github.com/KKallidromitis/r2o.

cs.CV↗

Spin parity effects in monoaxial chiral ferromagnetic chain

We present a fully quantum mechanical account of a novel {\it spin parity effect} -- physics which depend sharply on whether the spin quantum number $S$ is half-odd integral or integral, which we find to be present in models of monoaxial chiral ferromagnetic spin chains.

cond-mat.stat-mech↗

Vortex dynamics in the two-dimensional BCS-BEC crossover

The Bardeen-Cooper-Schrieffer (BCS) condensation and Bose-Einstein condensation (BEC) are the two limiting ground states of paired Fermion systems, and the crossover between these two limits has been a source of excitement for both fields of high temperature superconductivity and cold atom superfluidity. For superconductors, ultra-low doping systems like graphene and LixZrNCl successfully approached the crossover starting from the BCS-side. These superconductors offer new opportunities to clarify the nature of charged-particles transport towards the BEC regime. Here we report the study of vortex dynamics within the crossover using their Hall effect as a probe in LixZrNCl. We observed a systematic enhancement of the Hall angle towards the BCS-BEC crossover, which was qualitatively reproduced by the phenomenological time-dependent Ginzburg-Landau (TDGL) theory. LixZrNCl exhibits a band structure free from various electronic instabilities, allowing us to achieve a comprehensive understanding of the vortex Hall effect and thereby propose a global picture of vortex dynamics within the crossover. These results demonstrate that gate-controlled superconductors are ideal platforms towards investigations of unexplored properties in BEC superconductors.

cond-mat.supr-con↗

Fluctuation contribution to Spin Hall Effect in Superconductors

We theoretically study the contribution of superconducting fluctuation to extrinsic spin Hall effects in two- and three-dimensional electron gas and intrinsic spin Hall effects in two-dimensional electron gas with Rashba-type spin-orbit interaction. The Aslamazov-Larkin, Density-of-States, Maki-Thompson terms have logarithmic divergence $\lnε$ in the limit $ε=(T- T_{\mathrm{c}})/T_{\mathrm{c}} \rightarrow +0$ in two-dimensional systems for both extrinsic and intrinsic spin Hall effects except the Maki-Thompson terms in extrinsic effect, which are proportional to $(ε-γ_φ)^{-1}\lnε$ with a cutoff $γ_φ$ in two-dimensional systems. We found that the fluctuation effects on the extrinsic spin Hall effect have an opposite sign to that in the normal state, while those on the intrinsic spin Hall effect have the same sign.

cond-mat.supr-con↗