SearcharxivSearch

arXiv subjects

Sota Shimizu

Publications and source records attributed to Sota Shimizu.

3 recordsLinked to original sources

Bridging integrated information theory and the free-energy principle in living neuronal networks

Integrated Information Theory (IIT) links consciousness to integrated causal structure, whereas the Free-Energy Principle (FEP) explains self-organization through variational free-energy minimization. Their relationship in living neural systems remains unclear. We analyzed dissociated neuronal cultures learning to infer hidden signal sources. Across repeated stimulation, variational free energy decreased, while inference accuracy and Bayesian surprise, defined as the divergence between prior and posterior beliefs, increased. An IIT-inspired integrated-information proxy and main-complex size followed a non-monotonic, hill-shaped trajectory. The proxy correlated most strongly with Bayesian surprise and more weakly with accuracy and variational free energy. An Ising-model analysis indicated that Bayesian surprise and integrated information can be jointly amplified near shared positive critical modes and suggested how early connectivity development followed by response stabilization could generate the observed trajectory. These results link belief updating to integrated information in living neuronal networks and provide an empirical point of contact between IIT and the FEP.

q-bio.NC

Gaze-Based Intention Recognition for Human-Robot Collaboration

This work aims to tackle the intent recognition problem in Human-Robot Collaborative assembly scenarios. Precisely, we consider an interactive assembly of a wooden stool where the robot fetches the pieces in the correct order and the human builds the parts following the instruction manual. The intent recognition is limited to the idle state estimation and it is needed to ensure a better synchronization between the two agents. We carried out a comparison between two distinct solutions involving wearable sensors and eye tracking integrated into the perception pipeline of a flexible planning architecture based on Hierarchical Task Networks. At runtime, the wearable sensing module exploits the raw measurements from four 9-axis Inertial Measurement Units positioned on the wrists and hands of the user as an input for a Long Short-Term Memory Network. On the other hand, the eye tracking relies on a Head Mounted Display and Unreal Engine. We tested the effectiveness of the two approaches with 10 participants, each of whom explored both options in alternate order. We collected explicit metrics about the attractiveness and efficiency of the two techniques through User Experience Questionnaires as well as implicit criteria regarding the classification time and the overall assembly time. The results of our work show that the two methods can reach comparable performances both in terms of effectiveness and user preference. Future development could aim at joining the two approaches two allow the recognition of more complex activities and to anticipate the user actions.

cs.RO

An Effective Transformer-based Contextual Model and Temporal Gate Pooling for Speaker Identification

Wav2vec2 has achieved success in applying Transformer architecture and self-supervised learning to speech recognition. Recently, these have come to be used not only for speech recognition but also for the entire speech processing. This paper introduces an effective end-to-end speaker identification model applied Transformer-based contextual model. We explored the relationship between the hyper-parameters and the performance in order to discern the structure of an effective model. Furthermore, we propose a pooling method, Temporal Gate Pooling, with powerful learning ability for speaker identification. We applied Conformer as encoder and BEST-RQ for pre-training and conducted an evaluation utilizing the speaker identification of VoxCeleb1. The proposed method has achieved an accuracy of 87.1% with 28.5M parameters, demonstrating comparable precision to wav2vec2 with 317.7M parameters. Code is available at https://github.com/HarunoriKawano/speaker-identification-with-tgp.

cs.SD