SearcharxivSearch

arXiv subjects

Cooper Doyle

Publications and source records attributed to Cooper Doyle.

7 recordsLinked to original sources

BOND: License to Train with Black-Box Functions

We introduce Bounded Numerical Differentiation (BOND), a perturbative method for estimating the gradients of black-box functions. BOND is distinguished by its formulation, which adaptively bounds perturbations to ensure accurate sign estimation, and by its implementation, which operates at black-box interfaces. This enables BOND to be more accurate and scalable compared to existing methods, facilitating end-to-end training of architectures that incorporate non-autodifferentiable modules. We observe that these modules, implemented in our experiments as frozen networks, can enhance model performance without increasing the number of trainable parameters. Our findings highlight the potential of leveraging fixed transformations to expand model capacity, pointing to hybrid analogue - digital devices as a path to scaling networks, and provides insights into the dynamics of adaptive optimizers.

cs.LG

FeatureCuts: Feature Selection for Large Data by Optimizing the Cutoff

In machine learning, the process of feature selection involves finding a reduced subset of features that captures most of the information required to train an accurate and efficient model. This work presents FeatureCuts, a novel feature selection algorithm that adaptively selects the optimal feature cutoff after performing filter ranking. Evaluated on 14 publicly available datasets and one industry dataset, FeatureCuts achieved, on average, 15 percentage points more feature reduction and up to 99.6% less computation time while maintaining model performance, compared to existing state-of-the-art methods. When the selected features are used in a wrapper method such as Particle Swarm Optimization (PSO), it enables 25 percentage points more feature reduction, requires 66% less computation time, and maintains model performance when compared to PSO alone. The minimal overhead of FeatureCuts makes it scalable for large datasets typically seen in enterprise applications.

cs.LG

Your Absorbing Discrete Diffusion Secretly Models the Bayesian Posterior

Discrete diffusion language models learn to reconstruct text from randomly masked inputs, yet under mild assumptions their denoiser already implements the exact Bayesian posterior over the original tokens. We prove that the expected denoiser output under the forward corruption distribution recovers the true posterior, and that a simple Monte Carlo estimator converges to this posterior at rate O(1/sqrt(K)) with finite-sample concentration bounds. Building on this insight, we introduce an inference-time ensemble that runs K independent denoising passes and aggregates both posterior means and variances without any extra training. On WikiText-2, our MC-marginal sampler recovers the analytic lambda-DCE zero-shot perplexity (approximately 39) to within a few points at K=128, and its per-token variance shows a strong rank correlation with reconstruction error (Spearman rho = 0.996). This cost-proportional procedure yields calibrated uncertainty estimates and a direct trade-off between compute and posterior fidelity in discrete diffusion LMs.

cs.CL

Localising Dropout Variance in Twin Networks

Accurate individual treatment-effect estimation demands not only reliable point predictions but also uncertainty measures that help practitioners \emph{locate} the source of model failure. We introduce a layer-wise variance decomposition for deep twin-network models: by toggling Monte Carlo Dropout independently in the shared encoder and the outcome heads, we split total predictive variance into an \emph{encoder component} ($\sigma_{\mathrm{enc}}^2$) and a \emph{head component} ($\sigma_{\mathrm{head}}^2$), with $\sigma_{\mathrm{enc}}^2 + \sigma_{\mathrm{head}}^2 \approx \sigma_{\mathrm{tot}}^2$ by the law of total variance. Across three synthetic covariate-shift regimes, the encoder component dominates under distributional shift ($\rho_{\mathrm{enc}}=0.53$) while the head component becomes informative only once encoder uncertainty is controlled. On a real-world twins cohort with induced multivariate shift, only $\sigma_{\mathrm{enc}}^2$ spikes on out-of-distribution samples and becomes the primary error predictor ($\rho_{\mathrm{enc}}\!\approx\!0.89$), while $\sigma_{\mathrm{head}}^2$ remains flat. The decomposition adds negligible cost over standard MC Dropout and provides a practical diagnostic for deciding whether to collect more diverse covariates or more outcome data.

cs.LG

Learning Adapter Rank via Symmetry Breaking

Low-rank adaptation is effective partly because downstream updates lie in a low-dimensional subspace, but the latent rank coordinates of LoRA are not identifiable: any invertible reparameterization of the adapter factors leaves the weight update unchanged. We show that variational inference with a diagonal rank-wise posterior turns this non-identifiability into a useful inductive bias. By breaking LoRA's rotational gauge symmetry, the variational objective selects a preferred basis in rank space, enabling automatic relevance determination over rank directions. This yields Low-Rank Variational Dropout (LRVD), a Bayesian framework that performs inference directly in the low-rank adaptation space rather than the ambient weight space. As an instantiation, BayesLoRA jointly learns effective adapter rank and predictive uncertainty with only $\mathcal{O}(r)$ additional parameters. Empirically, BayesLoRA induces stable rank structure aligned with the dominant singular directions of learned updates, yields compact predictive calibration and matches or exceeds strong low-rank sparsification baselines at comparable training cost.

cs.LG

Biphoton entanglement of topologically-distinct modes

The robust generation and manipulation of entangled multiphoton states on-chip has an essential role in quantum computation and communication. Lattice topology has emerged as a means of protecting photonic states from disorder but entanglement across different topologies remained unexplored. We report biphoton entanglement between topologically distinct spatial modes in a bipartite array of silicon waveguides. The results highlight topology as an additional degree of freedom for entanglement and open avenues for investigating information teleportation between trivial and topological modes.

quant-ph

Topologically protected entangled photonic states

Entangled multiphoton states lie at the heart of quantum information, computing, and communications. In recent years, topology has risen as a new avenue to robustly transport quantum states in the presence of fabrication defects, disorder and other noise sources. Whereas topological protection of single photons and correlated photons has been recently demonstrated experimentally, the observation of topologically protected entangled states has thus far remained elusive. Here, we experimentally demonstrate the topological protection of spatially-entangled biphoton states. We observe robustness in crucial features of the topological biphoton correlation map in the presence of deliberately introduced disorder in the silicon nanophotonic structure, in contrast with the lack of robustness in nontopological structures. The topological protection is shown to ensure the coherent propagation of the entangled topological modes, which may lead to robust propagation of quantum information in disordered systems.

physics.optics