SearcharxivSearch

arXiv subjects

Kilian Müller

Publications and source records attributed to Kilian Müller.

12 recordsLinked to original sources

Streamlined optical training of large-scale modern deep learning architectures with direct feedback alignment

Modern deep learning relies nearly exclusively on dedicated electronic hardware accelerators. Photonic approaches, with low consumption and high operation speed, are increasingly considered for inference but, to date, remain mostly limited to relatively basic tasks. Simultaneously, the problem of training deep and complex neural networks, overwhelmingly performed through backpropagation, remains a significant limitation to the size and, consequently, the performance of current architectures and a major compute and energy bottleneck. Here, we experimentally implement a versatile and scalable training algorithm, called direct feedback alignment, on a hybrid electronic-photonic platform. An optical processing unit performs large-scale random matrix multiplications, which is the central operation of this algorithm, at speeds up to 1500 TeraOPS under 30 Watts of power. We perform optical training of modern deep learning architectures, including Transformers, with more than 1B parameters, and obtain good performances on language, vision, and diffusion-based generative tasks. We study the scaling of the training time, and demonstrate a potential advantage of our hybrid opto-electronic approach for ultra-deep and wide neural networks, thus opening a promising route to sustain the exponential growth of modern artificial intelligence beyond traditional von Neumann approaches.

cs.ET

Linear Optical Random Projections Without Holography

We introduce a novel method to perform linear optical random projections without the need for holography. Our method consists of a computationally trivial combination of multiple intensity measurements to mitigate the information loss usually associated with the absolute-square non-linearity imposed by optical intensity measurements. Both experimental and numerical findings demonstrate that the resulting matrix consists of real-valued, independent, and identically distributed (i.i.d.) Gaussian random entries. Our optical setup is simple and robust, as it does not require interference between two beams. We demonstrate the practical applicability of our method by performing dimensionality reduction on high-dimensional data, a common task in randomized numerical linear algebra with relevant applications in machine learning.

physics.optics

No Time Like the Present: Effects of Language Change on Automated Comment Moderation

The spread of online hate has become a significant problem for newspapers that host comment sections. As a result, there is growing interest in using machine learning and natural language processing for (semi-) automated abusive language detection to avoid manual comment moderation costs or having to shut down comment sections altogether. However, much of the past work on abusive language detection assumes that classifiers operate in a static language environment, despite language and news being in a state of constant flux. In this paper, we show using a new German newspaper comments dataset that the classifiers trained with naive ML techniques like a random-test train split will underperform on future data, and that a time stratified evaluation split is more appropriate. We also show that classifier performance rapidly degrades when evaluated on data from a different period than the training data. Our findings suggest that it is necessary to consider the temporal dynamics of language when developing an abusive language detection system or risk deploying a model that will quickly become defunct.

cs.CL

LightOn Optical Processing Unit: Scaling-up AI and HPC with a Non von Neumann co-processor

We introduce LightOn's Optical Processing Unit (OPU), the first photonic AI accelerator chip available on the market for at-scale Non von Neumann computations, reaching 1500 TeraOPS. It relies on a combination of free-space optics with off-the-shelf components, together with a software API allowing a seamless integration within Python-based processing pipelines. We discuss a variety of use cases and hybrid network architectures, with the OPU used in combination of CPU/GPU, and draw a pathway towards "optical advantage".

cs.AR

Photonic co-processors in HPC: using LightOn OPUs for Randomized Numerical Linear Algebra

Randomized Numerical Linear Algebra (RandNLA) is a powerful class of methods, widely used in High Performance Computing (HPC). RandNLA provides approximate solutions to linear algebra functions applied to large signals, at reduced computational costs. However, the randomization step for dimensionality reduction may itself become the computational bottleneck on traditional hardware. Leveraging near constant-time linear random projections delivered by LightOn Optical Processing Units we show that randomization can be significantly accelerated, at negligible precision loss, in a wide range of important RandNLA algorithms, such as RandSVD or trace estimators.

stat.ML

High-Resolution Imaging of Cold Atoms through a Multimode Fiber

We developed an ultra-compact high-resolution imaging system for cold atoms. Its only in-vacuum element is a multimode optical fiber with a diameter of $230\,μ$m, which simultaneously collects light and guides it out of the vacuum chamber. External adaptive optics allow us to image cold Rb atoms with a $\sim 1\,μ$m resolution over a $100 \times 100\,μ$m$^2$ field of view. These optics can be easily rearranged to switch between fast absorption imaging and high-sensitivity fluorescence imaging. This system is particularly suited for hybrid quantum engineering platforms where cold atoms are combined with optical cavities, superconducting circuits or optomechanical devices restricting the optical access.

physics.atom-ph

Hardware Beyond Backpropagation: a Photonic Co-Processor for Direct Feedback Alignment

The scaling hypothesis motivates the expansion of models past trillions of parameters as a path towards better performance. Recent significant developments, such as GPT-3, have been driven by this conjecture. However, as models scale-up, training them efficiently with backpropagation becomes difficult. Because model, pipeline, and data parallelism distribute parameters and gradients over compute nodes, communication is challenging to orchestrate: this is a bottleneck to further scaling. In this work, we argue that alternative training methods can mitigate these issues, and can inform the design of extreme-scale training hardware. Indeed, using a synaptically asymmetric method with a parallelizable backward pass, such as Direct Feedback Alignement, communication needs are drastically reduced. We present a photonic accelerator for Direct Feedback Alignment, able to compute random projections with trillions of parameters. We demonstrate our system on benchmark tasks, using both fully-connected and graph convolutional networks. Our hardware is the first architecture-agnostic photonic co-processor for training neural networks. This is a significant step towards building scalable hardware, able to go beyond backpropagation, and opening new avenues for deep learning.

cs.LG

Light-in-the-loop: using a photonics co-processor for scalable training of neural networks

As neural networks grow larger and more complex and data-hungry, training costs are skyrocketing. Especially when lifelong learning is necessary, such as in recommender systems or self-driving cars, this might soon become unsustainable. In this study, we present the first optical co-processor able to accelerate the training phase of digitally-implemented neural networks. We rely on direct feedback alignment as an alternative to backpropagation, and perform the error projection step optically. Leveraging the optical random projections delivered by our co-processor, we demonstrate its use to train a neural network for handwritten digits recognition.

cs.LG

FakeYou! -- A Gamified Approach for Building and Evaluating Resilience Against Fake News

Nowadays fake news are heavily discussed in public and political debates. Even though the phenomenon of intended false information is rather old, misinformation reaches a new level with the rise of the internet and participatory platforms. Due to Facebook and Co., purposeful false information - often called fake news - can be easily spread by everyone. Because of a high data volatility and variety in content types (text, images,...) debunking of fake news is a complex challenge. This is especially true for automated approaches, which are prone to fail validating the veracity of the information. This work focuses on an a gamified approach to strengthen the resilience of consumers towards fake news. The game FakeYou motivates its players to critically analyze headlines regarding their trustworthiness. Further, the game follows a "learning by doing strategy": by generating own fake headlines, users should experience the concepts of convincing fake headline formulations. We introduce the game itself, as well as the underlying technical infrastructure. A first evaluation study shows, that users tend to use specific stylistic devices to generate fake news. Further, the results indicate, that creating good fakes and identifying correct headlines are challenging and hard to learn.

cs.CY

Manipulating cold atoms through a high-resolution compact system based on a multimode fiber

To manipulate cold atoms in spatially constrained quantum engineering platforms, we developed a lensless optical system with a $\sim$1 $μ$m resolution and a transverse size of only 225 $μ$m. We use a multimode optical fiber with a high numerical aperture, which directly guides light inside the ultra-high-vacuum system. Spatial light modulators allow us to generate control beams at the in-vacuum fiber end by digital optical phase conjugation. As a demonstration, we use this system to optically transport cold atoms towards the in-vacuum fiber end, to load them in optical microtraps and to re-cool them in optical molasses. This work opens new perspectives for setups combining cold atoms with other optical, electronic or opto-mechanical systems with limited optical access.

physics.atom-ph

Suppression and Revival of Weak Localization through Control of Time-Reversal Symmetry

We report on the observation of suppression and revival of coherent backscattering of ultra-cold atoms launched in an optical disorder and submitted to a short dephasing pulse, as proposed in a recent paper of T. Micklitz \textit{et al.} [arXiv:1406.6915]. This observation, in a quasi-2D geometry, demonstrates a novel and general method to study weak localization by manipulating time reversal symmetry in disordered systems. In future experiments, this scheme could be extended to investigate higher order localization processes at the heart of Anderson (strong) localization.

cond-mat.quant-gas

Coherent Backscattering of Ultracold Atoms

We report on the direct observation of coherent backscattering (CBS) of ultracold atoms, in a quasi-two-dimensional configuration. Launching atoms with a well-defined momentum in a laser speckle disordered potential, we follow the progressive build up of the momentum scattering pattern, consisting of a ring associated with multiple elastic scattering, and the CBS peak in the backward direction. Monitoring the depletion of the initial momentum component and the formation of the angular ring profile allows us to determine microscopic transport quantities. The time resolved evolution of the CBS peak is studied and is found a fair agreement with predictions, at long times as well as at short times. The observation of CBS can be considered a direct signature of coherence in quantum transport of particles in disordered media. It is responsible for the so called weak localization phenomenon, which is the precursor of Anderson localization.

physics.atom-ph