Searcharxiv⌕ Search

arXiv subjects

Ertuğrul Mutlu

Publications and source records attributed to Ertuğrul Mutlu.

2 recordsLinked to original sources

Behavioral Convergence Without Representational Convergence: Persistent Training-History Dependence in Neural Networks

Neural networks trained toward the same final objective can reach similar predictive performance while retaining internal representations shaped by earlier training history. We study this effect using controlled sequential-training experiments in which paired convolutional networks start from identical weights, experience reversed task orders, and then receive the same deterministic common-relaxation distribution. Across 20 paired MNIST runs, 16 satisfy a predeclared behavioral-matching criterion, yet their matched representations retain a mean history score of 0.139 (95% bootstrap CI: 0.127-0.153) and approximately 3.1% prediction disagreement. Extending common relaxation to 50,000 optimizer updates does not erase the measured difference: across five paired seeds, the representation-history score remains 0.190 (95% bootstrap CI: 0.161-0.219) at the end of the measured horizon while the mean accuracy gap is only 0.18 percentage points. Fresh linear probes show that, with sufficient labeled data, the two histories retain practically equivalent linearly accessible class information. A same-label rotated-MNIST control reproduces the effect: all five paired seeds reach behavioral matching while retaining a mean representation-history score of 0.162. Finally, a matched-learning-rate ReLU-LeakyReLU control reduces the 50,000-update representation residue by 0.040 on average in all five paired seeds, providing directional evidence that activation-mediated plasticity contributes to the persistence of training-history effects. These results provide protocol-scoped evidence that behavioral convergence need not imply representational convergence and that optimization history can leave measurable internal traces after prolonged common training.

cs.LG↗

Wavelet-Based Parity Detection Revisited: Representation Dependence, Generalization, and Mechanistic Analysis

Parity is exactly determined by the least significant bit (LSB), so machine learning is unnecessary for solving the task. This work instead uses parity as a controlled diagnostic for studying how a classical signal-processing pipeline preserves, amplifies, or suppresses access to symbolic information embedded in a numerical representation. Integers from 0 to 10,000 are encoded as fixed-width 32-bit binary signals, decomposed with a level-3 Daubechies-2 (db2) discrete wavelet transform, summarized by mean absolute coefficient magnitude, and clustered independently per subband with k-means. Clustering is unsupervised, while cluster-to-parity calibration uses training labels only under a stratified 60/20/20 train/validation/test protocol. The frozen configuration reaches 84.26% accuracy on the held-out test set (95% Wilson CI: 82.60-85.79%), and 84.20% +/- 0.57% across 20 stratified resplits. Raw binary summary statistics reach 61.30%, while masking the LSB reduces validation accuracy to 48.15%, consistent with chance. The A3 approximation band alone retains 83.20%, whereas detail bands remain near chance. Performance is strongly representation-dependent: moving the parity bit changes validation accuracy up to 98.60%, and changing the wavelet boundary mode ranges from 54.45% to 83.20%. Cross-magnitude extrapolation degrades from 79.98% on 10,001-20,000 to 59.69% on 100,001-1,000,000. These results do not show that wavelets discover the arithmetic rule of parity. Rather, they show that information already present in the binary encoding becomes more or less recoverable depending on multiscale filtering, spatial alignment, and boundary handling.

cs.LG↗