ECHOv2: A Frequency-Structured Pre-trained Acoustic Representation Model with Cross-Band Modeling for Machine Anomalous Sound Detection
Machine anomalous sound detection (ASD) is an important technology for industrial acoustic monitoring, where robust acoustic representation learning remains challenging due to limited anomalous samples and complex machine sound characteristics. Existing pre-trained acoustic representation models do not fully capture frequency-specific characteristics of machine sounds. To address this, we propose ECHOv2, a frequency-structured pre-trained acoustic representation model. The model learns localized intra-band representations to capture fine-grained spectral patterns while also incorporating a two-level self-distillation strategy with explicit inter-band supervision to model cross-frequency dependencies. The inter-band branch performs global context alignment and masked sub-band reconstruction, and multiple summary tokens are introduced for structured aggregation with controllable frequency granularity, enabling region-aware interaction across sub-bands during training. This design allows ECHOv2 to robustly handle diverse machine types and noisy operating conditions while maintaining stable representation quality. To enable fair and consistent evaluation of pre-trained audio backbones, we establish a unified ASD benchmark over DCASE 2020--2025 with two complementary protocols: embedding-based evaluation for frozen representation discriminability and adaptation-based evaluation for downstream transferability. Ablation studies confirm the effectiveness of intra-band learning, inter-band supervision, and structured aggregation granularity for robust ASD representation learning. These findings demonstrate that structured cross-band modeling improves acoustic representation learning for machine monitoring and provides an effective pre-trained representation model for industrial anomalous sound detection. The model and benchmark are publicly available to promote reproducible research.