SearcharxivSearch

arXiv subjects

Wilson Wu

Publications and source records attributed to Wilson Wu.

11 recordsLinked to original sources

Quantum-Limited Subdiffraction Telescopy Requires Genuine Multi-Telescope Interference

Conventional stellar interferometry reconstructs incoherent sources from pairwise mutual coherences between telescopes. Are such pairwise measurements sufficient for quantum-limited subdiffraction imaging with a telescope array? We show that for generic image-moment estimation, they are not. We consider weak incoherent light from a generic extended source observed by an array of telescopes, each supporting a single optical mode. For an N-telescope array, we derive the quantum Fisher information (QFI) scaling of image moments up to the cutoff 2N-2 and prove that arbitrary measurements restricted to telescope pairs attain the full-array QFI scaling only up to second order. Thus, estimating higher-order moments at the quantum limit requires genuinely multi-telescope interference. Inspired by spatial-mode demultiplexing (SPADE) from single-aperture subdiffraction imaging, we construct array-SPADE measurements that attain the optimal QFI scaling up to the finite-array cutoff. Finally, we show that these measurements can, in principle, be embedded in ancilla- and memory-assisted quantum-network architectures for long-baseline telescopy.

quant-ph

Estimating the expected output of wide random MLPs more efficiently than sampling

By far the most common way to estimate an expected loss in machine learning is to draw samples, compute the loss on each one, and take the empirical average. However, sampling is not necessarily optimal. Given an MLP at initialization, we show how to estimate its expected output over Gaussian inputs without running samples through the network at all. Instead, we produce approximate representations of the distributions of activations at each layer, leveraging tools such as cumulants and Hermite expansions. We show both theoretically and empirically that for sufficiently wide networks, our estimator achieves a target mean squared error using substantially fewer FLOPs than Monte Carlo sampling. We find moreover that our methods perform particularly well at estimating the probabilities of rare events, and additionally demonstrate how they can be used for model training. Together, these findings suggest a path to producing models with a greatly reduced probability of catastrophic tail risks.

cs.LG

Bayesian Influence Functions for Hessian-Free Data Attribution

Classical influence functions face significant challenges when applied to deep neural networks, primarily due to non-invertible Hessians and high-dimensional parameter spaces. We propose the local Bayesian influence function (BIF), an extension of classical influence functions that replaces Hessian inversion with loss landscape statistics that can be estimated via stochastic-gradient MCMC sampling. This Hessian-free approach captures higher-order interactions among parameters and scales efficiently to neural networks with billions of parameters. We demonstrate state-of-the-art results on predicting retraining experiments.

cs.LG

C*-like modules and matrix $p$-operator norms

We present a generalization of H\"older duality to algebra-valued pairings via $L^p$-modules. H\"older duality states that if $p \in (1, \infty)$ and $p^{\prime}$ are conjugate exponents, then the dual space of $L^p(\mu)$ is isometrically isomorphic to $L^{p^{\prime}}(\mu)$. In this work we study certain pairs $(\mathsf{Y},\mathsf{X})$, as generalizations of the pair $(L^{p^{\prime}}(\mu), L^p(\mu))$, that have an $L^p$-operator algebra valued pairing $\mathsf{Y} \times \mathsf{X} \to A$. When the $A$-valued version of H\"older duality still holds, we say that $(\mathsf{Y},\mathsf{X})$ is C*-like. We show that finite and countable direct sums of the C*-like module $(A,A)$ are still C*-like when $A$ is any block diagonal subalgebra of $d \times d$ matrices. We provide counterexamples when $A \subset M_d^p(\mathbb{C})$ is not block diagonal.

math.FA

Feasibility study of frequency-encoded photonic qubits over a free-space channel

Frequency-bin quantum encoding shows great promise for quantum communication given its high-dimensional scaling, compatibility with photonic integrated circuits and synergy with classical optical communication technology. However, to date all demonstrations have been performed over single-mode and static channels, while the transmission over fluctuating and turbulent channels has not been addressed. We propose and demonstrate a novel approach that leverages field-widened interferometers to decode frequency-bins transmitted over free-space channels without any adaptive optics or modal filtering. Moreover, we investigate the phase stability requirements so that frequency-bin encoding could be feasible for satellite to ground quantum links. Our passive approach expands the versatility of frequency-bin encoding, paving the way towards long-range and fluctuating channels.

quant-ph

Towards a unified and verified understanding of group-operation networks

A recent line of work in mechanistic interpretability has focused on reverse-engineering the computation performed by neural networks trained on the binary operation of finite groups. We investigate the internals of one-hidden-layer neural networks trained on this task, revealing previously unidentified structure and producing a more complete description of such models in a step towards unifying the explanations of previous works (Chughtai et al., 2023; Stander et al., 2024). Notably, these models approximate equivariance in each input argument. We verify that our explanation applies to a large fraction of networks trained on this task by translating it into a compact proof of model performance, a quantitative evaluation of the extent to which we faithfully and concisely explain model internals. In the main text, we focus on the symmetric group S5. For models trained on this group, our explanation yields a guarantee of model accuracy that runs 3x faster than brute force and gives a >=95% accuracy bound for 45% of the models we trained. We were unable to obtain nontrivial non-vacuous accuracy bounds using only explanations from previous works.

cs.LG

Do language models plan ahead for future tokens?

Do transformers "think ahead" during inference at a given position? It is known transformers prepare information in the hidden states of the forward pass at time step $t$ that is then used in future forward passes $t+\tau$. We posit two explanations for this phenomenon: pre-caching, in which off-diagonal gradient terms present during training result in the model computing features at $t$ irrelevant to the present inference task but useful for the future, and breadcrumbs, in which features most relevant to time step $t$ are already the same as those that would most benefit inference at time $t+\tau$. We test these hypotheses by training language models without propagating gradients to past timesteps, a scheme we formalize as myopic training. In a constructed synthetic data setting, we find clear evidence for pre-caching. In the autoregressive language modeling setting, our experiments are more suggestive of the breadcrumbs hypothesis, though pre-caching increases with model scale.

cs.LG

Learning Deterministic Finite Automata from Confidence Oracles

We discuss the problem of learning a deterministic finite automaton (DFA) from a confidence oracle. That is, we are given access to an oracle $Q$ with incomplete knowledge of some target language $L$ over an alphabet $\Sigma$; the oracle maps a string $x\in\Sigma^*$ to a score in the interval $[-1,1]$ indicating its confidence that the string is in the language. The interpretation is that the sign of the score signifies whether $x\in L$, while the magnitude $|Q(x)|$ represents the oracle's confidence. Our goal is to learn a DFA representation of the oracle that preserves the information that it is confident in. The learned DFA should closely match the oracle wherever it is highly confident, but it need not do this when the oracle is less sure of itself.

cs.FL

Towards Fully Passive Time-Bin Quantum Key Distribution over Moving Free-Space Channels

Encoding quantum information in photonic time-bin states is typically considered impractical for moving free-space quantum communication due to the difficulties with phase stabilization of distant quantum time-bin interferometers and turbulence of free-space channels. We demonstrate a novel approach using reference frame independent time-bin quantum key distribution that completely avoids the need for active relative phase stabilization while simultaneously overcoming a highly multi-mode channel without any active mode filtering. This scheme enables passive, self-compensating time-bin quantum communication without any mode filtering, mode sorting, adaptive optics, active basis selection, or active phase alignment. We realize a proof-of-concept demonstration using hybrid polarization and time-bin entangled photons that demonstrates a sustained asymptotic secure key rate greater than 0.07 bits/coincidence over a 15m multi-mode fiber optical channel and showing entanglement correlations over a moving 38.5dB loss free-space channel, including system losses. The scheme simplifies the use of time-bin encoding and can be readily applied over various spatially multi-mode and fluctuating channels involving rapidly moving platforms, including airborne and satellite systems.

quant-ph

Distributed Verifiers in PCP

Traditional proof systems involve a resource-bounded verifier communicating with a powerful (but untrusted) prover. Distributed verifier proof systems are a new family of proof models that involve a network of verifier nodes communicating with a single independent prover that has access to the complete network structure of the verifiers. The prover is tasked with convincing all verifiers of some global property of the network graph. In addition, each individual verifier may be given some input string they will be required to verify during the course of computation. Verifier nodes are allowed to exchange messaged with nodes a constant distance away, and accept / reject the input after some computation. Because individual nodes are limited to a local view, communication with the prover is potentially necessary to prove global properties about the network graph of nodes, which only the prover has access to. In this system of models, the entire model accepts the input if and only if every individual node has accepted. There are three models in the distributed verifier proof system family: $\mathsf{LCP}$, $\mathsf{dIP}$, and our proposed $\mathsf{dPCP}$, with the fundamental difference between these coming from the type of communication established between the verifiers and the prover. In this paper, we will first go over the past work in the $\mathsf{LCP}$ and $\mathsf{dIP}$ space before showing properties and proofs in our $\mathsf{dPCP}$ system.

cs.CC

Analyzing and Improving Neural Networks by Generating Semantic Counterexamples through Differentiable Rendering

Even as deep neural networks (DNNs) have achieved remarkable success on vision-related tasks, their performance is brittle to transformations in the input. Of particular interest are semantic transformations that model changes that have a basis in the physical world, such as rotations, translations, changes in lighting or camera pose. In this paper, we show how differentiable rendering can be utilized to generate images that are informative, yet realistic, and which can be used to analyze DNN performance and improve its robustness through data augmentation. Given a differentiable renderer and a DNN, we show how to use off-the-shelf attacks from adversarial machine learning to generate semantic counterexamples -- images where semantic features are changed as to produce misclassifications or misdetections. We validate our approach on DNNs for image classification and object detection. For classification, we show that semantic counterexamples, when used to augment the dataset, (i) improve generalization performance (ii) enhance robustness to semantic transformations, and (iii) transfer between models. Additionally, in comparison to sampling-based semantic augmentation, our technique generates more informative data in a sample efficient manner.

cs.LG