Searcharxiv⌕ Search

arXiv subjects

Akira Takahashi

Publications and source records attributed to Akira Takahashi.

At least 19 recordsLinked to original sources

Passing: An Endless Journey through Reconstructed Spacetime with AI-Generated Sound

This paper introduces Passing, an interactive audiovisual installation that generates an endless journey from a single continuous monorail-window recording by reconstructing it as a spatiotemporal volume. Rather than replaying the footage linearly, the work resamples its spatial and temporal structure along nonlinear trajectories, producing a continuously passing landscape whose depth, speed, and temporal order become unstable. A camera-based viewer-presence detection system estimates whether a viewer is present in the viewing zone and uses this presence state to influence transitions among rendered video sequences. The resulting video stream is fed into SpecMaskFoley, a real-time video-to-audio synthesis model that generates a synchronized soundscape for the reconfigured image. The model is not used to reconstruct an objectively correct soundtrack, but functions as a speculative listener, proposing a possible auditory interpretation of a world whose conventional spatial and temporal premises have been disrupted. Passing distributes creative agency across the artist, who defines the rules of spacetime reconstruction; the AI model, which interprets the emergent visual flow as sound; and the audience, whose embodied presence influences the audiovisual trajectory. Through this structure, the work investigates how authorship and listening may be negotiated among human intention, machine inference, and audience interpretation. Artwork page: https://ryufurusawa.com/passing

cs.SD↗

Certified in Theory, Broken in Practice: Assumption Gaps in Cryptographic Model Certification

Privacy-preserving machine learning auditing protocols allow auditors to assess models for properties such as accuracy or fairness, without revealing their internals or training data. This makes them especially attractive for auditing models deployed in sensitive domains such as healthcare or finance. For these protocols to be meaningful in real-world audit settings, though, their guarantees must reflect how the model will behave once deployed, rather than merely certifying its behavior during an audit. Existing security definitions often miss this mark: most certify model behavior only on a fixed audit dataset, without ensuring that the same guarantees generalize to other datasets drawn from the same distribution. As we show, this gap allows a model provider to attack many cryptographic model certification (CMC) schemes built on secure zero knowledge proofs (ZKP) by carefully engineering training data, resulting in models that exhibit benign behavior during an audit, but pathological behavior in practice. For example, we empirically demonstrate that an attacker can certify that a model achieves over 99% accuracy on an audit dataset, but less than 30% accuracy on fresh samples from the same distribution. To address this gap, we formalize rigorous cryptographic security notions tailored to CMC frameworks, introduce a generic protocol template, and prove that it satisfies these requirements. Our results thus offer both cautionary evidence about existing approaches and constructive guidance for designing secure, privacy-preserving ML auditing protocols.

cs.CR↗

ZKBoost: Zero-Knowledge Verifiable Training for XGBoost

Gradient boosted decision trees, particularly XGBoost, are among the most effective methods for tabular data. As deployment in sensitive settings increases, cryptographic guarantees of model integrity become essential. We present ZKBoost, the first zero-knowledge proof of training (zkPoT) protocol for XGBoost, enabling model owners to prove correct training on a committed dataset without revealing data or model parameters. Naively re-executing XGBoost training in ZK would incur prohibitive costs, primarily due to the oblivious partitioning of training samples and unknown tree splits. Moreover, previous work on ZKP of training and inference had subtle security issues, such as leakage of tree topology and soundness gaps allowing cheating model providers to deviate from the correct execution of training and inference. We make two key contributions to address these challenges: (1) a generic zkPoT template for XGBoost that can be instantiated with any general-purpose ZKP backend, significantly improving prover costs compared to naive re-execution of the training process; and (2) a VOLE-based instantiation that overcomes the security issues of previous ZK proofs of training at minimal costs. To maximize efficiency, we develop a fixed-point version of XGBoost, which is particularly well suited for efficient instantiation of ZKP, and show it matches standard XGBoost accuracy to within 1\% on real-world datasets.

cs.CR↗

MMAudioReverbs: Video-Guided Acoustic Modeling for Dereverberation and Room Impulse Response Estimation

Although recent video-to-audio (V2A) models excelled at synthesizing semantically plausible sounds from visual inputs, they do not explicitly model room-acoustic effects such as reverberation or room impulse responses (RIRs), and thus offer limited controllability over these effects. However, we hypothesize that such V2A models implicitly have semantic knowledge of the relationship between spatial audio and the corresponding vision cues. In this paper, we revisit a V2A model for the sake of the above, and propose the way to utilize the pretrained model as prior for physically grounded room-acoustic processing. Based on one of the state-of-the-art V2A models, MMAudio, we propose MMAudioReverbs that is a unified framework dealing with i) dereverberation and ii) room impulse response (RIR) estimation without network architectural modification, and fine-tuned on a small dataset. Experimental results showed that audio and visual cues respectively have advantage depending on the type of physical room acoustics. It implies that foundation V2A models can be used for physically grounded room-acoustic analysis.

cs.SD↗

MMAudio-LABEL: Audio Event Labeling via Audio Generation for Silent Video

Recent advances in multimodal generation have enabled high-quality audio generation from silent videos. Practical applications, such as sound production, demand not only the generated audio but also explicit sound event labels detailing the type and timing of sounds. One straightforward approach involves applying a standard sound event detection to the generated audio. However, this post-hoc pipeline is inherently limited, as it is prone to error accumulation. To address this limitation, we propose MMAudio-LABEL (LAtent-Based Event Labeling), an event-aware audio generation framework built on a foundational audio generation model as its backbone that jointly generates audio and frame-aligned sound event predictions from silent videos. We evaluate our method on the Greatest Hits dataset for onset detection and 17-class material classification. Our approach improves onset-detection accuracy from 46.7% to 75.0% and material-classification accuracy from 40.6% to 61.0% over baselines. These results suggest that jointly learning audio generation and event prediction enables a more interpretable and practical video-to-audio synthesis.

cs.SD↗

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation

We introduce MMAudioSep, a generative model for video/text-queried sound separation that is founded on a pretrained video-to-audio model. By leveraging knowledge about the relationship between video/text and audio learned through a pretrained audio generative model, we can train the model more efficiently, i.e., the model does not need to be trained from scratch. We evaluate the performance of MMAudioSep by comparing it to existing separation models, including models based on both deterministic and generative approaches, and find it is superior to the baseline models. Furthermore, we demonstrate that even after acquiring functionality for sound separation via fine-tuning, the model retains the ability for original video-to-audio generation. This highlights the potential of foundational sound generation models to be adopted for sound-related downstream tasks. Our code is available at https://github.com/sony/mmaudiosep.

cs.SD↗

ZEBRA-Prop: A Zero-Shot Embedding-Based Rapid and Accessible Regression Model for Materials Properties

Large language models (LLMs) exhibit substantial potential across diverse scientific disciplines, including materials science. A property prediction framework, ZEBRA-Prop (Zero-Shot Embedding-Based Rapid and Accessible Regression Model for Materials Properties), is presented here as an extension of LLM-Prop. In contrast to LLM-Prop, which requires task-specific fine-tuning of the LLM, ZEBRA-Prop eliminates fine-tuning, thereby reducing computational cost and enabling rapid model training. The framework employs MatTPUSciBERT, an LLM specialized for materials science, to enhance predictive capability. Multiple textual embeddings are incorporated through a learnable weighting mechanism, which alleviates the context-length constraints inherent in LLM-Prop and facilitates effective integration of diverse textual representations. Evaluation is conducted using two datasets: the TextEdge dataset (approximately 140,000 entries) and an in-house dataset (approximately 2,000 entries) derived from the Materials Project database, with physical properties obtained from first-principles calculations. The predictive performance of ZEBRA-Prop is close to that of LLM-Prop, while the training time is reduced by approximately 95%. The performance improvements are attributable to three principal factors: domain-specific LLM utilization, diversified textual descriptions, and systematic text preprocessing. ZEBRA-Prop constitutes a scalable and computationally efficient framework for materials property prediction and supports accelerated materials discovery, particularly under limited computational resources.

cond-mat.mtrl-sci↗

Do Foundational Audio Encoders Understand Music Structure?

In music information retrieval (MIR) research, the use of pretrained foundational audio encoders (FAEs) has recently become a trend. FAEs pretrained on large amounts of music and audio data have been shown to improve performance on MIR tasks such as music tagging and automatic music transcription. However, their use for music structure analysis (MSA) remains underexplored: only a small subset of FAEs has been examined for MSA, and the impact of factors such as learning methods, training data, and model context length on MSA performance remains unclear. In this study, we conduct comprehensive experiments on 11 types of FAEs to investigate how these factors affect MSA performance. Our results demonstrate that FAEs using self-supervised learning with masked language modeling on music data are particularly effective for MSA. These findings pave the way for future research in FAE and MSA.

cs.SD↗

'Studies for': A Human-AI Co-Creative Sound Artwork Using a Real-time Multi-channel Sound Generation Model

This paper explores the integration of AI technologies into the artistic workflow through the creation of Studies for, a generative sound installation developed in collaboration with sound artist Evala (https://www.ntticc.or.jp/en/archive/works/studies-for/). The installation employs SpecMaskGIT, a lightweight yet high-quality sound generation AI model, to generate and playback eight-channel sound in real-time, creating an immersive auditory experience over the course of a three-month exhibition. The work is grounded in the concept of a "new form of archive," which aims to preserve the artistic style of an artist while expanding beyond artists' past artworks by continued generation of new sound elements. This speculative approach to archival preservation is facilitated by training the AI model on a dataset consisting of over 200 hours of Evala's past sound artworks. By addressing key requirements in the co-creation of art using AI, this study highlights the value of the following aspects: (1) the necessity of integrating artist feedback, (2) datasets derived from an artist's past works, and (3) ensuring the inclusion of unexpected, novel outputs. In Studies for, the model was designed to reflect the artist's artistic identity while generating new, previously unheard sounds, making it a fitting realization of the concept of "a new form of archive." We propose a Human-AI co-creation framework for effectively incorporating sound generation AI models into the sound art creation process and suggest new possibilities for creating and archiving sound art that extend an artist's work beyond their physical existence. Demo page: https://sony.github.io/studies-for/

cs.SD↗

Deep Learning-Based Extraction of Promising Material Groups and Common Features from High-Dimensional Data: A Case of Optical Spectra of Inorganic Crystals

We report an interpretation method for deep learning models that allows us to handle high-dimensional spectral data in materials science. The proposed method uses feature extraction and clustering analysis to categorize materials into classes based on similarities in both spectral data and chemical characteristics such as elemental composition and atomic arrangement. As a demonstration, we apply this method to an atomistic line graph neural network (ALIGNN) model trained on first-principles calculation data of 2,681 metal oxides, chalcogenides, and related compounds for optical absorption spectrum prediction. Our analysis reveals key elemental species and their coordination environments that influence optical absorption onset characteristics. The method proposed herein is broadly applicable to the classification and interpretation of diverse spectral data, extending beyond the optical absorption spectra of inorganic crystals.

cond-mat.mtrl-sci↗

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet

Foley synthesis aims to synthesize high-quality audio that is both semantically and temporally aligned with video frames. Given its broad application in creative industries, the task has gained increasing attention in the research community. To avoid the non-trivial task of training audio generative models from scratch, adapting pretrained audio generative models for video-synchronized foley synthesis presents an attractive direction. ControlNet, a method for adding fine-grained controls to pretrained generative models, has been applied to foley synthesis, but its use has been limited to handcrafted human-readable temporal conditions. In contrast, from-scratch models achieved success by leveraging high-dimensional deep features extracted using pretrained video encoders. We have observed a performance gap between ControlNet-based and from-scratch foley models. To narrow this gap, we propose SpecMaskFoley, a method that steers the pretrained SpecMaskGIT model toward video-synchronized foley synthesis via ControlNet. To unlock the potential of a single ControlNet branch, we resolve the discrepancy between the temporal video features and the time-frequency nature of the pretrained SpecMaskGIT via a frequency-aware temporal feature aligner, eliminating the need for complicated conditioning mechanisms widely used in prior arts. Evaluations on a common foley synthesis benchmark demonstrate that SpecMaskFoley could even outperform strong from-scratch baselines, substantially advancing the development of ControlNet-based foley synthesis models. Demo page: https://zzaudio.github.io/SpecMaskFoley_Demo/

cs.SD↗

SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond

Recent advances in generative models that iteratively synthesize audio clips sparked great success to text-to-audio synthesis (TTA), but with the cost of slow synthesis speed and heavy computation. Although there have been attempts to accelerate the iterative procedure, high-quality TTA systems remain inefficient due to hundreds of iterations required in the inference phase and large amount of model parameters. To address the challenges, we propose SpecMaskGIT, a light-weighted, efficient yet effective TTA model based on the masked generative modeling of spectrograms. First, SpecMaskGIT synthesizes a realistic 10s audio clip by less than 16 iterations, an order-of-magnitude less than previous iterative TTA methods. As a discrete model, SpecMaskGIT outperforms larger VQ-Diffusion and auto-regressive models in the TTA benchmark, while being real-time with only 4 CPU cores or even 30x faster with a GPU. Next, built upon a latent space of Mel-spectrogram, SpecMaskGIT has a wider range of applications (e.g., the zero-shot bandwidth extension) than similar methods built on the latent wave domain. Moreover, we interpret SpecMaskGIT as a generative extension to previous discriminative audio masked Transformers, and shed light on its audio representation learning potential. We hope our work inspires the exploration of masked audio modeling toward further diverse scenarios.

cs.SD↗

A charge model as an effective model of one-dimensional Hubbard and extended Hubbard systems: its application to linear optical spectrum calculations in large systems based upon many-body Wannier functions

We propose an effective model called the "charge model", for the half-filled one-dimensional Hubbard and extended Hubbard models. In this model, spin-charge separation, which has been justified from an infinite on-site repulsion ($U$) in the strict sense, is compatible with charge fluctuations. Our analyses based on the many-body Wannier functions succeeded in determining the optical conductivity spectra in large systems. The obtained spectra reproduce the spectra for the original models well even in the intermediate $U$ region of $U=5-10T$, with $T$ being the nearest-neighbor electron hopping energy. These results indicate that the spin-charge separation works fairly well in this intermediate $U$ region against the usual expectation and that the charge model is an effective model that applies to actual quasi-one-dimensional materials classified as strongly correlated electron systems.

cond-mat.str-el↗

Electrically Benign Defect Behavior in Zinc Tin Nitride Revealed from First Principles

Zinc tin nitride (ZnSnN2) is attracting growing interest as a non-toxic and earth-abundant photoabsorber for thin-film photovoltaics. Carrier transport in ZnSnN2 and consequently cell performance are strongly affected by point defects with deep levels acting as carrier recombination centers. In this study, the point defects in ZnSnN2 are revisited by careful first-principles modeling based on recent experimental and theoretical findings. It is shown that ZnSnN2 does not have low-energy defects with deep levels, in contrast to previously reported results. Therefore, ZnSnN2 is more promising as a photoabsorber material than formerly considered.

cond-mat.mtrl-sci↗

Linearized machine-learning interatomic potentials for non-magnetic elemental metals: Limitation of pairwise descriptors and trend of predictive power

Machine-learning interatomic potential (MLIP) has been of growing interest as a useful method to describe the energetics of systems of interest. In the present study, we examine the accuracy of linearized pairwise MLIPs and angular-dependent MLIPs for 31 elemental metals. Using all of the optimal MLIPs for 31 elemental metals, we show the robustness of the linearized frameworks, the general trend of the predictive power of MLIPs and the limitation of pairwise MLIPs. As a result, we obtain accurate MLIPs for all 31 elements using the same linearized framework. This indicates that the use of numerous descriptors is the most important practical feature for constructing MLIPs with high accuracy. An accurate MLIP can be constructed using only pairwise descriptors for most non-transition metals, whereas it is very important to consider angular-dependent descriptors when expressing interatomic interactions of transition metals.

cond-mat.mtrl-sci↗

Conceptual and practical bases for the high accuracy of machine learning interatomic potential

Machine learning interatomic potentials (MLIPs) based on a large dataset obtained by density functional theory (DFT) calculation have been developed recently. This study gives both conceptual and practical bases for the high accuracy of MLIPs, although MLIPs have been considered to be simply an accurate black-box description of atomic energy. We also construct the most accurate MLIP of the elemental Ti ever reported using a linearized MLIP framework and many angular-dependent descriptors, which also corresponds to a generalization of the modified embedded atom method (MEAM) potential.

cond-mat.mtrl-sci↗

Representation of compounds for machine-learning prediction of physical properties

The representations of a compound, called "descriptors" or "features", play an essential role in constructing a machine-learning model of its physical properties. In this study, we adopt a procedure for generating a systematic set of descriptors from simple elemental and structural representations. First it is applied to a large dataset composed of the cohesive energy for about 18000 compounds computed by density functional theory (DFT) calculation. As a result, we obtain a kernel ridge prediction model with a prediction error of 0.041 eV/atom, which is close to the "chemical accuracy" of 1 kcal/mol (0.043 eV/atom). The procedure is also applied to two smaller datasets, i.e., a dataset of the lattice thermal conductivity (LTC) for 110 compounds computed by DFT calculation and a dataset of the experimental melting temperature for 248 compounds. We examine the performance of the descriptor sets on the efficiency of Bayesian optimization in addition to the accuracy of the kernel ridge regression models. They exhibit good predictive performances.

cond-mat.mtrl-sci↗

First-principles interatomic potentials for ten elemental metals via compressed sensing

Interatomic potentials have been widely used in atomistic simulations such as molecular dynamics. Recently, frameworks to construct accurate interatomic potentials that combine a systematic set of density functional theory (DFT) calculations with machine learning techniques have been proposed. One of these methods is to use compressed sensing to derive a sparse representation for the interatomic potential. This facilitates the control of the accuracy of interatomic potentials. In this study, we demonstrate the applicability of compressed sensing to deriving the interatomic potential of ten elemental metals, namely, Ag, Al, Au, Ca, Cu, Ga, In, K, Li and Zn. For each elemental metal, the interatomic potential is obtained from DFT calculations using elastic net regression. The interatomic potentials are found to have prediction errors of less than 3.5 meV/atom, 0.03 eV/Å and 0.15 GPa for the energy, force and the stress tensor, respectively, which enable the accurate prediction of physical properties such as lattice constants and the phonon dispersion relationship.

cond-mat.mtrl-sci↗