SearcharxivSearch

arXiv subjects

Kazuki Sakai

Publications and source records attributed to Kazuki Sakai.

10 recordsLinked to original sources

Human-robot conversation with multiple participants in noisy public spaces

For noisy real-world environments such as those in open public spaces, spoken dialogue systems for both autonomous robots and avatars should be carefully designed to provide enhanced speech signals. These signals can be used either for speech recognition or, in the case of an avatar system, transmitted as clean speech to a remote operator. This work proposes an audio system that can be used for both these scenarios and was demonstrated as a proof-of-concept at the 2025 World Expo in Osaka. The first scenario is an attentive listening system with the android ERICA, and the second is a conversation support system with mobile Teleco robots, with one of them acting as an avatar for a remote operator. Both systems feature multi-party conversation and use a single multi-channel microphone array. We describe how our audio system not only enhances the speech of multiple speakers in a noisy environment, but provides a form of spatial audio which allows for more immersiveness in avatar-based conversational interactions.

cs.RO

Closed-Form Steepest Descent Direction toward Flat Minima: Reducing Upper Bounds on the Loss Hessian Eigenspectrum in Neural Networks

The flatness hypothesis suggests that flatness of the loss landscape, as measured by the eigenvalues of the loss Hessian, correlates with better neural network generalization. While various algorithms reduce these eigenvalues, most focus on procedural design, leaving it unclear how data distributions and NN parameters structurally determine directions toward flat minima. Characterizing these directions analytically is generally intractable. To overcome this mathematical difficulty, recent studies derived the Wolkowicz-Styan (WS) upper bound on the maximum eigenvalue of the cross-entropy loss Hessian in three-layer NNs. Although this upper bound is differentiable, its gradient was not derived. Therefore, we analytically derive the gradient of the WS upper bound to characterize directions leading to flat minima. Based on this, we propose Hessian Spectral Range (HSR) Regularization, which updates parameters along the steepest descent direction of the WS bound. Experiments demonstrate that HSR Regularization narrows the Hessian eigenvalue spectrum, avoids sharp minima and saddle points, and promotes convergence to flat minima. Although the applicability of this method is currently limited to cross-entropy loss and three-layer architectures, to the best of the authors' knowledge, this is the first study to report a closed-form gradient that promotes convergence to flat minima without numerical approximations. Therefore, the theoretical analysis of this gradient is expected to contribute to the further development of NNs.

cs.LG

Wolkowicz-Styan Upper Bound on the Hessian Eigenspectrum for Cross-Entropy Loss in Nonlinear Smooth Neural Networks

Neural networks (NNs) are central to modern machine learning and achieve state-of-the-art results in many applications. However, the relationship between loss geometry and generalization is still not well understood. The local geometry of the loss function near a critical point is well-approximated by its quadratic form, obtained through a second-order Taylor expansion. The coefficients of the quadratic term correspond to the Hessian matrix, whose eigenspectrum allows us to evaluate the sharpness of the loss at the critical point. Extensive research suggests flat critical points generalize better, while sharp ones lead to higher generalization error. However, sharpness requires the Hessian eigenspectrum, but general matrix characteristic equations have no closed-form solution. Therefore, most existing studies on evaluating loss sharpness rely on numerical approximation methods. Existing closed-form analyses of the eigenspectrum are primarily limited to simplified architectures, such as linear or ReLU-activated networks; consequently, theoretical analysis of smooth nonlinear multilayer neural networks remains limited. Against this background, this study focuses on nonlinear, smooth multilayer neural networks and derives a closed-form upper bound for the maximum eigenvalue of the Hessian with respect to the cross-entropy loss by leveraging the Wolkowicz-Styan bound. Specifically, the derived upper bound is expressed as a function of the affine transformation parameters, hidden layer dimensions, and the degree of orthogonality among the training samples. The primary contribution of this paper is an analytical characterization of loss sharpness in smooth nonlinear multilayer neural networks via a closed-form expression, avoiding explicit numerical eigenspectrum computation. We hope that this work provides a small yet meaningful step toward unraveling the mysteries of deep learning.

cs.LG

MolLIBRA: Genetic Molecular Optimization with Multi-Fingerprint Surrogates and Text-Molecule Aligned Critic

We study sample-efficient molecular optimization under a limited budget of oracle evaluations. We propose MolLIBRA (MultimOdaLity and Language Integrated Bayesian and evolutionaRy optimizAtion), a genetic algorithm based framework that pre-ranks candidate molecules using multiple critics before oracle calls: (i) an ensemble of Gaussian process (GP) surrogates defined over multiple molecular fingerprints and (ii) a pretrained text-molecule aligned encoder CLAMP. The GP ensemble enables adaptive selection of task-appropriate fingerprints, while CLAMP provides a zero-shot scoring signal from task descriptions by measuring the similarity between molecular and text embeddings. On the Practical Molecular Optimization (PMO) benchmark with a budget of 1,000 evaluations (PMO-1K), MolLIBRA-L, our variant with a language-model-based candidate generator, attains the best Top-10 AUC on 14/22 tasks and the highest overall sum of Top-10 AUC across tasks among prior methods.

cs.NE

Computing the Wave: Where the Gravitational Wave Community benefits from High-Energy Physics, and where it differs ?

High-Energy Physics (HEP) and Gravitational Wave (GW) communities serve different scientific purposes. However, their methodologies might potentially offer mutual enrichment through common software developments. A suite of libraries is currently being prototyped and made available at https://git.ligo.org/kagra/libraries-addons/root, extending at no cost the CERN ROOT data analysis framework toward advanced signal processing. We will also present a performance benchmark comparing the FFTW and KFR library performances.

astro-ph.IM

Parameter estimation of protoneutron stars from gravitational wave signals using the Hilbert-Huang transform

Core-collapse supernovae (CCSNe) are potential multimessenger events detectable by current and future gravitational wave (GW) detectors. The GW signals emitted during these events are expected to provide insights into the explosion mechanism and the internal structures of neutron stars. In recent years, several studies have empirically derived the relationship between the frequencies of the GW signals originating from the oscillations of protoneutron stars (PNSs) and the physical parameters of these stars. This study applies the Hilbert-Huang transform (HHT) [Proc. R. Soc. A 454, 903 (1998)] to extract the frequencies of these modes to infer the physical properties of the PNSs. The results exhibit comparable accuracy to a short-time Fourier transform-based estimation, highlighting the potential of this approach as a complementary method for extracting physical information from GW signals of CCSNe.

gr-qc

Black hole spectroscopy for KAGRA future prospect in O5

Ringdown gravitational waves of compact binary mergers are an important target to test general relativity. The main components of the ringdown waveform after merger are black hole quasinormal modes. In general relativity, all multipolar quasinormal modes of a black hole should give the same values of black hole parameters. Although the observed binary black hole events so far are not significant enough to perform the test with ringdown gravitational waves, it is expected that the test will be achieved in third generation detectors. The Japanese gravitational wave detector KAGRA, called bKAGRA for the current configuration, has started observation, and discussions for the future upgrade plans have also started. In this study, we consider which KAGRA upgrade plan is the best to detect the subdominant quasinormal modes of black holes in the aim of testing general relativity. We use a numerical relativity waveform as injected signals that contains two multipolar modes and analyze each mode by matched filtering. Our results suggest that the plan FDSQZ, which improves the sensitivity of KAGRA for broad frequency range, is the most suitable configuration for black hole spectroscopy.

gr-qc

Comparison of various methods to extract ringdown frequency from gravitational wave data

The ringdown part of gravitational waves in the final stage of merger of compact objects tells us the nature of strong gravity which can be used for testing the theories of gravity. The ringdown waveform, however, fades out in a very short time with a few cycles, and hence it is challenging for gravitational wave data analysis to extract the ringdown frequency and its damping time scale. We here propose to build up a suite of mock data of gravitational waves to compare the performance of various approaches developed to detect quasi-normal modes from a black hole. In this paper we present our initial results of comparisons of the following five methods; (1) plain matched filtering with ringdown part (MF-R) method, (2) matched filtering with both merger and ringdown parts (MF-MR) method, (3) Hilbert-Huang transformation (HHT) method, (4) autoregressive modeling (AR) method, and (5) neural network (NN) method. After comparing their performance, we discuss our future projects.

gr-qc

Estimation of starting times of quasinormal modes in ringdown gravitational waves with the Hilbert-Huang transform

It is known that a quasinormal mode (QNM) of a remnant black hole dominates a ringdown gravitational wave (GW) in a binary black hole (BBH) merger. To study properties of the QNMs, it is important to determine the time when the QNMs appear in a GW signal as well as to calculate its frequency and amplitude. In this paper, we propose a new method of estimating the starting time of the QNM and calculating the QNM frequency and amplitude of BBH GWs. We apply it to simulated merger waveforms by numerical relativity and the observed data of GW150914. The results show that the obtained QNM frequencies and time evolutions of amplitudes are consistent with the theoretical values within 1% accuracy for pure waveforms free from detector noise. In addition, it is revealed that there is a correlation between the starting time of the QNM and the spin of the remnant black hole. In the analysis of GW150914, we show that the parameters of the remnant black hole estimated through our method are consistent with those given by LIGO and a reasonable starting time of the QNM is determined.

gr-qc

Amplitude-based detection method for gravitational wave bursts with the Hilbert-Huang Transform

We propose a new detection method for gravitational wave bursts. It analyzes observed data with the Hilbert-Huang transform, which is an approach of time-frequency analysis constructed with the aim of manipulating non-linear and non-stationary data. Using the simulated time-series noise data and waveforms from rotating core-collapse supernovae at 30 kpc, we performed simulation to evaluate the performance of our method and it revealed the total detection probability to be 0.94 without false alerms, which corresponds to the false alarm rate < 0.001 Hz. The detection probability depends on the characteristics of the waveform, but it was found that the parameter determining the degree of differential rotation of the collapsing star is the most important for the performance of our method.

astro-ph.IM