Searcharxiv⌕ Search

arXiv subjects

Keita Goto

Publications and source records attributed to Keita Goto.

13 recordsLinked to original sources

Electronic state of vortices at twin boundaries in a nematic superconductor

Local electronic states of vortices in an s $\pm$ d wave nematic superconductor are studied both in the absence and presence of twin boundaries. The Bogoliubov-de Gennes theory for a tight-binding model is used with its nematicity represented by the anisotropy in the transfer integrals and attractive interactions between the nearest-neighbor sites. We evaluate s and d wave components of the pair potentials and the local density of states, and analyze the effects of nematicity on the vortex core structures with/without twin boundaries. We find that a single vortex trapped at the twin boundary is composed of a bound pair of fractional vortices accompanied by weakly-induced s $\pm$ id wave components, despite that such s $\pm$ id wave components do not appear in a zero magnetic field. The calculated spatial structures of the local electronic states are compared with the vortex image measured by STM experiments in an iron-based superconductor FeSe.

cond-mat.supr-con↗

Bagpiper: Solving Open-Ended Audio Tasks via Rich Captions

Current audio foundation models typically rely on rigid, task-specific supervision (e.g., speech recognition), addressing isolated factors of audio rather than the whole. In contrast, human processes audio holistically, seamlessly bridging raw audio waveform with abstract cognitive concepts (e.g., all perception details of audio events) to execute complex tasks. Grounded in this philosophy, we introduce Bagpiper, an 8B audio foundation model that interprets physical audio via rich captions, i.e., comprehensive natural language descriptions that encapsulate the critical cognitive concepts inherent in the audio. By pre-training on a massive corpus of 600B tokens, the model establishes a robust bidirectional mapping between raw audio and this high-level conceptual space. During fine-tuning, Bagpiper adopts a caption-then-process workflow, simulating an intermediate cognitive reasoning step to solve diverse tasks without knowing prior task-specific practice. Experimentally, Bagpiper achieves universal generation that can uniformly generate speech, sound effects, music, and their arbitrary combinations. It also maintains comparable performance with the 7B Qwen-2.5-Omni for audio understanding. To the best of our knowledge, Bagpiper is among the first works that achieve open-ended audio understanding and generation on speech, sound, and music. Model, data, and code will be released at Bagpiper Home Page.

cs.CL↗

Bagpiper-TTS: Natural Language Guided Universal Speech Synthesis

Classical TTS systems typically rely on rigid input formats and predefined metadata slots, limiting their ability to fulfill flexible user requirements. This paper introduces Bagpiper-TTS, a universal speech synthesis system that deals with diverse natural language user requests. Given a natural language prompt, Bagpiper-TTS first reasons over the users' intent to derive a rich caption, i.e., a comprehensive textual blueprint encompassing both transcription and nuanced metadata. Subsequently, this caption guides the synthesis of the target speech. Our model inherently supports a broad spectrum of tasks besides classical TTS applications, including multi-talker, intent-to-speech, role-play synthesis, singing voice synthesis, and more. Experimental results demonstrate that Bagpiper-TTS achieves an 1.7% Word Error Rate (WER) on the Seed-TTS-Eval benchmark and match the performance of dedicated models in both LLM-as-a-judge and human subjective evaluations across multiple applications.

cs.CL↗

Online Predictive Coding for Dual-Mode Self-Supervised Speech Model

Dual-mode self-supervised speech models are pre-trained to handle streaming and non-streaming conditions simultaneously. However, their attention is computed over different context ranges, which often makes optimization difficult. In previous work, we proposed online registers, additional tokens intended to compensate for missing future context in streaming mode, but the gains remained limited. To address these issues, we introduce two improvements for robust dual-mode pre-training: (1) Online Predictive Coding (OPC), which regularizes the registers through multi-step future prediction, and (2) Dual-mode Layer Normalization, which stabilizes optimization. We fine-tune the proposed dual-mode self-supervised speech models for speech recognition on LibriSpeech and WSJ. Results show that OPC consistently reduces the online-offline performance gap; at 160 ms latency on LibriSpeech, word error rates improve from 3.65% to 3.40% on test-clean and from 10.15% to 9.65% on test-other.

cs.SD↗

Non-Archimedean balanced metrics and their application to totally degenerate abelian varieties

For a polarized complex manifold with discrete automorphism group, it is known that if the first Chern class admits a cscK metric, then the balanced metrics, which are characterized in terms of the algebro-geometric notion of Chow stability, approximate this cscK metric. In this paper, we study a non-Archimedean analogue of this phenomenon. In particular, we prove that such an analogue holds for polarized totally degenerate abelian varieties. As an application, we also show that, for a totally degenerating family of polarized abelian varieties, the validity of this non-Archimedean analogue yields a uniform estimate for the Calabi--Yau metrics on fibers sufficiently close to the degenerate fiber.

math.AG↗

Online Register for Dual-Mode Self-Supervised Speech Models: Mitigating The Lack of Future Context

Dual-mode self-supervised speech models (S3Ms), which jointly pre-trained in the offline and online mode, suffer from attention mismatch in streaming scenarios due to missing future context. To address this challenge, we proposed online registers, learnable tokens appended to each chunk in online mode. These tokens act as virtual placeholders for unseen future frames, enabling the model to compensate for missing context without introducing additional latency. Furthermore, we introduce a future prediction loss that explicitly guides the registers to capture predictive cues, thereby enriching their ability to retain future information. Experiments on LibriSpeech, and out-of-domain benchmarks demonstrate that online registers consistently reduce the performance gap between offline and online modes, achieving a 3.4% relative improvement on LibriSpeech with 160 ms chunks, especially in low-latency settings.

cs.SD↗

Evaluating Self-Supervised Speech Models via Text-Based LLMS

Self-Supervised Learning (SSL) has gained traction for its ability to learn rich representations with low labeling costs, applicable across diverse downstream tasks. However, assessing the downstream-task performance remains challenging due to the cost of extra training and evaluation. Existing methods for task-agnostic evaluation also require extra training or hyperparameter tuning. We propose a novel evaluation metric using large language models (LLMs). By inputting discrete token sequences and minimal domain cues derived from SSL models into LLMs, we obtain the mean log-likelihood; these cues guide in-context learning, rendering the score more reliable without extra training or hyperparameter tuning. Experimental results show a correlation between LLM-based scores and automatic speech recognition task. Additionally, our findings reveal that LLMs not only functions as an SSL evaluation tools but also provides inference-time embeddings that are useful for speaker verification task.

cs.SD↗

OpusLM: A Family of Open Unified Speech Language Models

This paper presents Open Unified Speech Language Models (OpusLMs), a family of open foundational speech language models (SpeechLMs) up to 7B. Initialized from decoder-only text language models, the OpusLMs are continuously pre-trained on 213K hours of speech-text pairs and 292B text-only tokens. We demonstrate our OpusLMs achieve comparable (or even superior) performance with existing SpeechLMs in speech recognition, speech synthesis, and text-only capabilities. Technically, this paper articulates our SpeechLM designs on tokenization, multi-stream language models, and multi-stage training strategies. We experimentally demonstrate the importance of model size scaling and the effect of annealing data selection. The OpusLMs are all built from publicly available materials and are fully transparent models. We release our code, data, checkpoints, and training logs to facilitate open SpeechLM research

cs.CL↗

On The Two Types Of Affine Structures For Degenerating Kummer Surfaces -Non-Archimedean VS Gromov-Hausdorff Limits-

Kontsevich and Soibelman constructed integral affine manifolds with singularities (IAMS, for short) for maximal degenerations of polarized Calabi-Yau manifolds in a non-Archimedean way. On the other hand, for each maximally degenerating family of polarized Calabi-Yau manifolds, we can consider the Gromov-Hausdorff limit of the fibers. It is expected that this Gromov-Hausdorff limit carries an IAMS-structure. Kontsevich and Soibelman conjectured that these two types of IAMS are the same. This conjecture is believed in the mirror symmetry context. In this paper, we prove the above conjecture for maximal degenerations of polarized Kummer surfaces.

math.AG↗

On the Berkovich double residue fields and birational models

Just as a residue field can be considered for a point of an algebraic variety, we can also consider a residue field for a point of a Berkovich analytic space. This residue field is a valuation field in the algebraic sense. Then we can consider its residue field as a valuation field. We call it the Berkovich double residue field at the point. In this paper, we consider a point $x$ of the Berkovich analytification of an algebraic variety and identify the Berkovich double residue field at $x$ with the union of the residue fields at the center of $x$ in birational models. Besides, we concretely compute the Berkovich double residue field for any quasi monomial valuation.

math.AG↗

Toric degenerations of Calabi--Yau complete intersections and metric SYZ conjecture

We consider a toric degeneration $\mathcal{X}$ of Calabi--Yau complete intersections of Batyrev--Borisov in the Gross--Siebert program. For the toric degeneration $\mathcal{X}$, we study the real Monge--Ampère equation corresponding to the non-archimedean Monge--Ampère equation that yields the non-archimedean Calabi--Yau metric. Our main theorem describes the real Monge--Ampère equation in terms of tropical geometry and proves the metric SYZ conjecture for the toric degeneration $\mathcal{X}$ supposing the existence of its solution.

math.AG↗

Special Lagrangian fibrations, Berkovich retraction, and crystallographic groups

We explicitly construct special Lagrangian fibrations on finite quotients of maximally degenerating abelian varieties, glue with Berkovich retraction in non-Archimedean geometry by using "hybrid" technique. We also study their symmetries explicitly which can be regarded as crystallographic groups. In particular, a conjecture of Kontsevich-Soibelman is solved at an enhanced level for finite quotients of abelian varieties in any dimension.

math.AG↗

Semi-Supervised Contrastive Learning with Generalized Contrastive Loss and Its Application to Speaker Recognition

This paper introduces a semi-supervised contrastive learning framework and its application to text-independent speaker verification. The proposed framework employs generalized contrastive loss (GCL). GCL unifies losses from two different learning frameworks, supervised metric learning and unsupervised contrastive learning, and thus it naturally determines the loss for semi-supervised learning. In experiments, we applied the proposed framework to text-independent speaker verification on the VoxCeleb dataset. We demonstrate that GCL enables the learning of speaker embeddings in three manners, supervised learning, semi-supervised learning, and unsupervised learning, without any changes in the definition of the loss function.

eess.AS↗