Searcharxiv⌕ Search

arXiv · 2609.39044

Game Sound-Effect Completion with Event-Level Transformation Hints

Abstract

Creating sound effects for a new game-character skin requires a distinct acoustic identity while preserving gameplay-event roles. The challenge is to complete a coherent set of related sounds whose required degrees of redesign differ. We formulate this task as completion conditioned on base-skin audio, completed target assets, and a textual design description. We develop a pipeline to collect, process, and align corresponding events across League of Legends skins. Building on Stable Audio 3's pretrained audio prior, we fine-tune a latent inpainting model to jointly complete missing events. A signed soft retention mask encodes available audio and an adjustable transformation hint for each missing event, specifying the requested balance between retention and redesign. Experiments on held-out skins show improved reconstruction over the evaluated general-purpose audio editors. Target-derived hints further improve paired similarity, with three-level hints retaining most of the benefit of continuous guidance.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Xinrui Jiang, Heng Yu. 2026-09-30. Game Sound-Effect Completion with Event-Level Transformation Hints. https://arxiv.org/abs/2609.39044

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Trigger Sound Suppression for Misophonia

Misophonia, a disorder of decreased tolerance to specific sounds, affects 5-20% of the population, yet sufferers have no good options: therapy helps a minority, and earplugs or noise cancellation silence everything. We present a study for neural trigger sound suppression for misophonia, selectively removing trigger sounds. We curate a dataset covering the 10 most common trigger classes. Using streaming dual-path networks operating on 6 ms audio chunks, we explore both one-hot and multi-hot-conditioned models that suppress 1-3 triggers from the acoustic scene. We validate our model outputs in a listening study with 30 adults with clinically elevated misophonia impairment. Participants reported significantly lower distress and arousal, and improved valence, for suppressed audio.

cs.SD↗

RAST: Resolution-Aware Privileged Structure Transfer for Low-Resolution Audio Activity Recognition

Audio is increasingly used for human activity recognition (HAR) because it captures object interactions, environmental events, and contextual cues in everyday environments. High-resolution (HR) audio provides rich acoustic information for model development but incurs substantial energy and storage costs and may expose sensitive speech content. Low-resolution (LR) audio offers a more privacy-preserving and resource-efficient alternative for deployment, but reduced sampling rates can remove acoustic cues essential for activity recognition, leading to significant performance degradation. We formulate this training-deployment mismatch as sensor-resolution privileged learning, in which HR audio is available during training, while inference relies exclusively on LR audio. We propose RAST, a resolution-aware transfer framework that compresses HR teacher representations by preserving token-level information and neighborhood structure before performing localized HR-LR alignment. Experiments on the SAMoSA and AudioIMU datasets show that RAST consistently outperforms LR-only training and direct teacher-transfer baselines, improving LR-only recognition by up to approximately 7.8% while requiring only LR audio at inference.

cs.SD↗

Audio Token Attention Is Predictable Before the Language Model Runs

A large audio language model (LALM) turns a minute of speech into 750-1,500 tokens and prefills every one. Image-token pruning often cuts after the language model's first layers, where image tokens draw little attention. Audio tokens draw much more attention there, and their ranking is still far from final, so audio needs a ranking before the language model runs. Surprisingly, the attention an audio token will receive across the language model is already linearly predictable from its encoder output, before the language model runs. A linear map, fitted in closed form without labels, predicts this all-layer attention ranking at $ρ\geq .69$ on eleven of thirteen LALMs. Our method, Triage, cuts audio tokens by this prediction and, on multiple choice, cuts again at layer 2, correcting the prediction with the attention observed there. Triage sets its compression without labels, under two budgets that limit how far its output may differ from the model's own full-audio output. At the conservative budget, its word error rate and accuracy stay within .04 of full audio. At the aggressive budget, Triage beats every baseline in all twelve transcription cases. On multiple choice, at 2.2-5x compression, it outperforms DART, the strongest baseline on average, by .043 in mean accuracy. Because it cuts before the language model, it raises the audio that fits in Qwen2.5-Omni-3B's context window from 21.8 to about 62 minutes. At its most compressive point, Triage lets one GPU serve 4x as many concurrent 5-minute streams of that model. Project page: https://audio-triage.github.io

cs.SD↗