Searcharxiv⌕ Search

arXiv · 2609.34907

SincDPNet: Interpretable Raw-Waveform Bathroom Activity Recognition for Assistive Living

Abstract

Bathroom acoustic-event recognition can support ambient assisted living in settings where continuous video monitoring is undesirable. However, practical deployment requires models that are compact, interpretable, and robust to changes in the recording environment. This work introduces \dataset{}, a seven-class bathroom acoustic-event dataset containing 21{,}387 annotated clips recorded across five environments, and proposes SincDPNet, a compact raw-waveform classifier with a learnable sinc filter bank followed by a depthwise-separable convolutional body. Each sinc filter is controlled by two frequency parameters, allowing the learned passbands to be inspected directly in hertz while keeping the front end small. To reduce room-specific leakage, recording sessions and environments are separated before overlapping windows are assigned to the training, validation, and test partitions. We further use multi-objective Bayesian optimization as a design tool to examine the validation performance--model-size trade-off across 24 configurations. The selected designs span different operating points: the best-performing model achieves 80.2\% accuracy and 0.760 macro-F1 with 14{,}040 parameters, while the compact $N_f=25$ configuration uses only 2{,}848 parameters and achieves 75.7\% accuracy, 0.661 macro-F1, and 0.716 MCC on the held-out environment. Analysis of the learned filters and confusion patterns shows that spectral overlap contributes to confusion among water-related events, while the \textit{Door}/\textit{Walker/Crutch} errors also reflect similarities in their transient temporal structure.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Debolina Chowdhury, Suman Samui, Sujoy Saha. 2026-09-28. SincDPNet: Interpretable Raw-Waveform Bathroom Activity Recognition for Assistive Living. https://arxiv.org/abs/2609.34907

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Localized time-frequency representation learning for bioacoustic classification in complex soundscapes

Prevailing bioacoustic classifiers assign species labels to fixed time-frequency windows rather than to individual vocalizations. When multiple vocalizations occur within the same window, predictions cannot be unambiguously linked to specific calls, which limits analyses at the level of individual vocalizations. This work introduces a framework for time-frequency localized bird classification. A Local-Context Classifier (LCC) identifies species from localized time-frequency events (TFEs), while a Dual-Context Classifier (DCC) combines local and global acoustic context through a fine-tuned bioacoustic foundation model. On an in-distribution dataset from Singapore comprising 306 vocalization classes from 103 bird species, the LCC achieves an F1-micro score of 79.3%, while combining local and global context through the DCC yields the highest overall performance (94.6%). To reduce labeled data requirements, the LCC is pre-trained via self-supervised contrastive learning, achieving an 18.8% relative gain on an out-of-distribution dataset. A focused evaluation on continuous soundscape recordings further demonstrates the potential of the framework for long-term monitoring applications. By preserving the time-frequency localization of individual vocalizations, the proposed framework supports both ecological monitoring and vocalization-level studies of animal acoustic behavior.

cs.SD↗

When Demonstrations Fail: Diagnosing the Limits of In-Context Learning in Large Audio-Language Models with Progressive Cue Removal

While Large Audio-Language Models (LALMs) have been shown to exhibit degraded instruction-following capabilities, their ability to infer task patterns from in-context examples with audio remains understudied. To address this gap, we design a three-stage evaluation pipeline that progressively reduces textual guidance to systematically evaluate LALMs' in-context learning ability in the audio modality. Evaluating six LALMs across four audio understanding tasks under two output constraint categories, we uncover a consistent asymmetry across LALMs: in-context demonstrations reliably improve format compliance but fail to improve the core task performance. This suggests that LALMs can glean surface-level formatting patterns from demonstrations but may struggle to leverage cross-modal semantic grounding to reliably infer task objectives from examples with audio, highlighting potential limitations in current cross-modal integration. We further probe how demonstrations are used through two complementary analyses, demonstration label shuffling and attention knockout on demonstration spans, both showing that LALMs leverage in-context examples primarily to establish the output label space and format rather than to learn a meaningful input-output correspondence.

cs.SD↗

Time-frequency localization of bird calls in dense soundscapes

Passive acoustic monitoring enables large-scale wildlife observation. Most bioacoustic classifiers predict species presence in a time window without localizing vocalizations precisely in time or frequency, limiting downstream analyses. We formulate time-frequency localization of bird calls as object detection on spectrograms and compare three computer vision model families (YOLO11, SAM 3, RF-DETR) against a non-learnable baseline. We introduce Intersection over Minimum (IoMin), an evaluation metric that better handles ambiguous acoustic boundaries than IoU. We also open-source a browser-based tool for efficient bounding-box labeling. The best RF-DETR model nearly doubles baseline performance on in-distribution, dense soundscapes from Singapore (83.4% vs. 42.1% IoMin@50 F1-score) and generalizes better to out-of-distribution recordings from Hawaii (63.2% vs. 48.6%). These results indicate that fine-tuned computer-vision models are well suited for time-frequency localization of bird vocalizations in complex soundscapes.

cs.SD↗