SearcharxivSearch

arXiv subjects

Daniel Brunner

Publications and source records attributed to Daniel Brunner.

At least 19 recordsLinked to original sources

Power law scaling for classification accuracy in physical neural networks

Physical neural networks (PNNs) harness the intrinsic complexity of physical systems to perform neural computation, potentially at speeds and energy efficiencies inaccessible to conventional digital hardware. Yet, a principled framework for quantifying and predicting their computing accuracy across diverse substrates has remained elusive. Here we introduce the Hotelling Trace Criterion (HTC), a task-conditioned measure of PNN- state separability that can be evaluated without training. We demonstrate that it predicts PNN classification performance with high fidelity across highly nonlinear optical fibres, vertical-cavity surface-emitting lasers, and coupled nonlinear oscillator networks, for benchmark tasks of different difficulty. Classification loss follows a power law in HTC, with Pearson correlation coefficients exceeding 0.99 for MNIST and $\approx$0.97 for Fashion-MNIST, noteworthy experimental and simulated data from physically distinct systems collapse onto a single scaling curve determined by the task rather than the substrate. Applying HTC layer-by-layer during training further reveals that gradient-based optimisation distributes representational capacity unevenly across PNN layers, providing a quantitative diagnostic of training and architecture efficiency invisible to standard loss monitoring. Crucially, once the scaling exponent is established from a small number of trained calibration systems, all further performance predictions require no training since performance can be derived from the much more efficient HTC measurement. These results establish HTC as a substrate-agnostic figure of merit for comparing and scaling PNNs, advancing the field further towards a complete theory connecting fundamental hardware parameters to task performance through universal scaling laws.

cs.ET

Measuring & Mitigating Over-Alignment for LLMs in Multilingual Criminal Law Courts

While the wider applicability of LLMs in the legal field is currently debated due to their reliability and the gravity of any errors, narrow uses with well-understood and mitigated risks have emerged. Notably the Swiss Federal Supreme Court uses small on-premises models for tentative translations and short-passage summarization across the four official languages. However, such usage is challenging in the context of Criminal Law. Since rulings and cases employees work on routinely can contain detailed descriptions of violent and sexual offenses, their legitimate work is compromised by refusals and disclaimers due to the activation of model guardrails (over-alignment). To measure this phenomenon, we introduce TF-RefusalBench, a multilingual benchmark for criminal-law translation and summarization derived from public Swiss Supreme Court rulings. TF-RefusalBench contains 5,200 total prompts across French, German, Italian, and English, corresponding to common task prompts and passages likely to trigger refusal. We then use TF-RefusalBench to show that over-alignment is a multifaceted phenomenon, influenced by the model and the prompt and text languages being processed, and that its impact cannot be evaluated solely from an over-refusal perspective, given the disclaimer's impact on task faithfulness. Finally, we evaluate approaches to enable on-premises LLMs for Criminal Law Tasks, demonstrating that while prompting can be effective, abliteration (refusal directions ablation) eliminates refusal with minimal impact on task performance.

cs.CL

3D Photonic integration leveraging hybrid-confinement circuits

Three-dimensional (3D) photonic integration offers a pathway to overcome the fundamental scaling limitations of planar platforms by enabling enhanced routing flexibility for compact, low-loss, and highly interconnected photonic circuits. In this work, we fabricate 3D photonic circuits combining high-confinement air-clad waveguides for compact routing with low-confinement polymer-clad waveguides for robust single-mode operation within a monolithic platform. Efficient mode transition between polymer-clad and air-clad waveguides is demonstrated with a loss of 0.25 dB per interface. We also realize compact, Euler S- and U-shaped bends with minimal bending radii of 10 $\mu$m and losses as low as 0.5 dB and 0.4 dB, respectively, along with compact adiabatic air-clad splitters exhibiting a splitting loss of 0.6~dB over a length of 52 $\mu$m. Finally, full fabrication of a compact hybrid circuit is demonstrated, highlighting the feasibility and scalability of the approach. Our work represents a significant step in 3D photonic integration for applications including optical neural networks, photonic wire bonding and their potential for novel integrated photonic applications.

physics.optics

Dynamic Consumer Demand at Large Scale

We study consumer demand in large-scale retail settings with many products, multiple categories and repeated purchase behavior. While inertia and brand loyalty are well documented, existing discrete choice models typically focus on single categories or become computationally infeasible in high-dimensional environments. We propose a dynamic product-level factor model that captures heterogeneity in baseline preferences, price sensitivity and inertia through a shared latent factor structure. By factorizing individual-product coefficients, the model pools information across individuals and categories and allows for correlated heterogeneity. We estimate the model using Bayesian variational inference, enabling scalable estimation with tens of thousands of parameters. In a simulation study calibrated to realistic retail data, we show that the dynamic factor model substantially improves predictive performance relative to static factor models and mixed logit benchmarks, particularly when individual purchase histories are sparse. Accounting for inertia also leads to more elastic demand estimates, underscoring the importance of dynamics for measuring consumer responsiveness. Our results highlight dynamic factor models as a scalable and flexible approach for demand estimation in modern, high-dimensional retail markets.

econ.EM

The thin line for optical neural networks towards broad practical relevance

Optical neural networks promise unmatched efficiency, bandwidth, and latency, critical benefits as demand for neural network hardware surges. However, their practical value for general-purpose acceleration or specialized applications must be proven under application-realistic conditions. We discuss recent insights and outline key research priorities.

physics.optics

LEXam: Benchmarking Legal Reasoning on 340 Law Exams

Long-form legal reasoning remains a key challenge for large language models (LLMs) in spite of recent advances in test-time scaling. To address this, we introduce LEXam, a novel benchmark derived from 340 law exams spanning 116 law school courses across a range of subjects and degree levels. The dataset comprises 7,537 law exam questions in English and German. It includes both long-form, open-ended questions and multiple-choice questions with varying numbers of options. Besides reference answers, the open questions are also accompanied by explicit guidance outlining the expected legal reasoning approach such as issue spotting, rule recall, or rule application. Our evaluation on both open-ended and multiple-choice questions present significant challenges for current LLMs; in particular, they notably struggle with open questions that require structured, multi-step legal reasoning. Moreover, our results underscore the effectiveness of the dataset in differentiating between models with varying capabilities. Deploying an ensemble LLM-as-a-Judge paradigm with rigorous human expert validation, we demonstrate how model-generated reasoning steps can be evaluated consistently and accurately, closely aligning with human expert assessments. Our evaluation setup provides a scalable method to assess legal reasoning quality beyond simple accuracy metrics. Project page: https://lexam-benchmark.github.io/.

cs.CL

Model-free front-to-end training of a large high performance laser neural network

Artificial neural networks (ANNs), have become ubiquitous and revolutionized many applications ranging from computer vision to medical diagnoses. However, they offer a fundamentally connectionist and distributed approach to computing, in stark contrast to classical computers that use the von Neumann architecture. This distinction has sparked renewed interest in developing unconventional hardware to support more efficient implementations of ANNs, rather than merely emulating them on traditional systems. Photonics stands out as a particularly promising platform, providing scalability, high speed, energy efficiency, and the ability for parallel information processing. However, fully realized autonomous optical neural networks (ONNs) with in-situ learning capabilities are still rare. In this work, we demonstrate a fully autonomous and parallel ONN using a multimode vertical cavity surface emitting laser (VCSEL) using off-the-shelf components. Our ONN is highly efficient and is scalable both in network size and inference bandwidth towards the GHz range. High performance hardware-compatible optimization algorithms are necessary in order to minimize reliance on external von Neumann computers to fully exploit the potential of ONNs. As such we present and extensively study several algorithms which are broadly compatible with a wide range of systems. We then apply these algorithms to optimize our ONN, and benchmark them using the MNIST dataset. We show that our ONN can achieve high accuracy and convergence efficiency, even under limited hardware resources. Crucially, we compare these different algorithms in terms of scaling and optimization efficiency in term of convergence time which is crucial when working with limited external resources. Our work provides some guidance for the design of future ONNs as well as a simple and flexible way to train them.

cs.LG

Limits of nonlinear and dispersive fiber propagation for an optical fiber-based extreme learning machine

We report a generalized nonlinear Schr\"odinger equation simulation model of an extreme learning machine (ELM) based on optical fiber propagation. Using the MNIST handwritten digit dataset as a benchmark, we study how accuracy depends on propagation dynamics, as well as parameters governing spectral encoding, readout, and noise. For this dataset and with quantum noise limited input, test accuracies of : over 91% and 93% are found for propagation in the anomalous and normal dispersion regimes respectively. Our results also suggest that quantum noise on the input pulses introduces an intrinsic penalty to ELM performance.

physics.optics

SwiLTra-Bench: The Swiss Legal Translation Benchmark

In Switzerland legal translation is uniquely important due to the country's four official languages and requirements for multilingual legal documentation. However, this process traditionally relies on professionals who must be both legal experts and skilled translators -- creating bottlenecks and impacting effective access to justice. To address this challenge, we introduce SwiLTra-Bench, a comprehensive multilingual benchmark of over 180K aligned Swiss legal translation pairs comprising laws, headnotes, and press releases across all Swiss languages along with English, designed to evaluate LLM-based translation systems. Our systematic evaluation reveals that frontier models achieve superior translation performance across all document types, while specialized translation systems excel specifically in laws but under-perform in headnotes. Through rigorous testing and human expert validation, we demonstrate that while fine-tuning open SLMs significantly improves their translation quality, they still lag behind the best zero-shot prompted frontier models such as Claude-3.5-Sonnet. Additionally, we present SwiLTra-Judge, a specialized LLM evaluation system that aligns best with human expert assessments.

cs.CL

Expressivity of Quantum Reservoir Computers

Using Hamiltonian encoding to inject an input into parameterized quantum circuits (PQCs), the output of the PQC can be written as truncated Fourier series. In recent years, the expressivity of PQCs was established as the number of frequencies contained in this Fourier series. While this concept has also been applied to other quantum machine learning (QML) paradigms, a clear notion of expressivity for temporal information processing with quantum systems is still lacking. Here, we introduce such a notion to the field of quantum reservoir computing (QRC). We analytically derive an expression for the readouts showing that the output of a QRC can be interpreted as a multi-dimensional Fourier series. We give a formula for the growth of expressivity induced by the sequential information injection, which we corroborate with numerical simulations, calculating explicitly the number of multi-dimensional output functions which can be generated from the readouts. Our results show that the specific interplay between system size, input encoding, and memory time gives rise to a boundary on the system size beyond which it is obstructive to further increase the reservoir size in extreme scrambling systems. We propose a recipe for determining this maximal system size for a given QRC setup.

quant-ph

Roadmap on Neuromorphic Photonics

This roadmap consolidates recent advances while exploring emerging applications, reflecting the remarkable diversity of hardware platforms, neuromorphic concepts, and implementation philosophies reported in the field. It emphasizes the critical role of cross-disciplinary collaboration in this rapidly evolving field.

cs.ET

Principles and Metrics of Extreme Learning Machines Using a Highly Nonlinear Fiber

Optical computing offers potential for ultra high-speed and low latency computation by leveraging the intrinsic properties of light. Here, we explore the use of highly nonlinear optical fibers (HNLFs) as platforms for optical computing based on the concept of Extreme Learning Machines. Task-independent evaluations are introduced to the field for the first time and focus on the fundamental metrics of effective dimensionality and consistency, which we experimentally characterize for different nonlinear and dispersive conditions. We show that input power and fiber characteristics significantly influence the dimensionality of the computational system, with longer fibers and higher dispersion producing up to 100 principal components (PCs) at input power levels of 30 mW, where the PC correspond to the linearly independent dimensions of the system. The spectral distribution of the PC's eigenvectors reveals that the high-dimensional dynamics facilitating computing through dimensionality expansion are located within 40~nm of the pump wavelength at 1560~nm, providing general insight for computing with nonlinear Schr\"odinger equation systems. Task-dependent results demonstrate the effectiveness of HNLFs in classifying MNIST dataset images. Using input data compression through PC analysis, we inject MNIST images of various input dimensionality into the system and study the impact of input power upon classification accuracy. At optimized power levels we achieve a classification test accuracy of 88\%, significantly surpassing the baseline of 83.7\% from linear systems. Noteworthy, we find that best performance is not obtained at maximal input power, i.e. maximal system dimensionality, but at more than one order of magnitude lower. The same is confirmed regarding the MNIST image's compression, where accuracy is substantially improved when strongly compressing the image to less than 50 PCs.

physics.optics

Reviving holographic photonic integration

Photonic integration of thick holograms in waveguiding structures could be considered the chimera of photonics; multi-faceted and hard to tame. It is the fundamental, and hence indispensable, concept behind compact and monolithically integrated linear optical transformation1. The true relevance of this becomes apparent in the high-dimensional context of unconventional optical computing, that is, in optical neural networks. Yet, integrating such holographic connections is very challenging. It demands high fabrication accuracy, and numerical design of the circuit is often non-tractable for large architectures. Both challenges are intrinsically linked to the usually large refractive index differences between sections of such holographic optical waveguides when using standard techniques of silicon photonics.

physics.optics

Experimental reservoir computing with diffractively coupled VCSELs

We present experiments on reservoir computing (RC) using a network of vertical-cavity surface-emitting lasers (VCSELs) that we diffractively couple via an external cavity. Our optical reservoir computer consists of 24 physical VCSEL nodes. We evaluate the system's memory and solve the 2-bit XOR task and the 3-bit header recognition (HR) task with bit error ratios (BERs) below 1\,\% and the 2-bit digital-to-analog conversion (DAC) task with a root-mean-square error (RMSE) of 0.067.

cs.ET

A spiking photonic neural network of 40.000 neurons, trained with rank-order coding for leveraging sparsity

Spiking neural networks are neuromorphic systems that emulate certain aspects of biological neurons, offering potential advantages in energy efficiency and speed by for example leveraging sparsity. While CMOS-based electronic SNN hardware has shown promise, scalability and parallelism challenges remain. Photonics provides a promising platform for SNNs due to the speed of excitable photonic devices standing in as neurons and the parallelism and low-latency of optical signal conduction. Here, we present a photonic SNN comprising 40,000 neurons using off-the-shelf components, including a spatial light modulator and a CMOS camera, enabling scalable and cost-effective implementations for photonic SNN proof of concept studies. The system is governed by a modified Ikeda map, were adding additional inhibitory feedback forcing introduces excitability akin to biological dynamics. Using latency encoding and sparsity, the network achieves 83.5% accuracy on MNIST using 22% of neurons, and 77.5% with 8.5% neuron utilization. Training is performed via liquid state machine concepts combined with the hardware-compatible SPSA algorithm, marking its first use in photonic neural networks. This demonstration integrates photonic nonlinearity, excitability, and sparse computation, paving the way for efficient large-scale photonic neuromorphic systems.

cs.ET

Impact of white noise in artificial neural networks trained for classification: performance and noise mitigation strategies

In recent years, the hardware implementation of neural networks, leveraging physical coupling and analog neurons has substantially increased in relevance. Such nonlinear and complex physical networks provide significant advantages in speed and energy efficiency, but are potentially susceptible to internal noise when compared to digital emulations of such networks. In this work, we consider how additive and multiplicative Gaussian white noise on the neuronal level can affect the accuracy of the network when applied for specific tasks and including a softmax function in the readout layer. We adapt several noise reduction techniques to the essential setting of classification tasks, which represent a large fraction of neural network computing. We find that these adjusted concepts are highly effective in mitigating the detrimental impact of noise.

cs.LG

Annealing-inspired training of an optical neural network with ternary weights

Artificial neural networks (ANNs) represent a fundamentally connectionnist and distributed approach to computing, and as such they differ from classical computers that utilize the von Neumann architecture. This has revived research interest in new unconventional hardware to enable more efficient implementations of ANNs rather than emulating them on traditional machines. In order to fully leverage the capabilities of this new generation of ANNs, optimization algorithms that take into account hardware limitations and imperfections are necessary. Photonics represents a particularly promising platform, offering scalability, high speed, energy efficiency, and the capability for parallel information processing. Yet, fully fledged implementations of autonomous optical neural networks (ONNs) with in-situ learning remain scarce. In this work, we propose a ternary weight architecture high-dimensional semiconductor laser-based ONN. We introduce a simple method for achieving ternary weights with Boolean hardware, significantly increasing the ONN's information processing capabilities. Furthermore, we design a novel in-situ optimization algorithm that is compatible with, both, Boolean and ternary weights, and provide a detailed hyperparameter study of said algorithm for two different tasks. Our novel algorithm results in benefits, both in terms of convergence speed and performance. Finally, we experimentally characterize the long-term inference stability of our ONN and find that it is extremely stable with a consistency above 99\% over a period of more than 10 hours, addressing one of the main concerns in the field. Our work is of particular relevance in the context of in-situ learning under restricted hardware resources, especially since minimizing the power consumption of auxiliary hardware is crucial to preserving efficiency gains achieved by non-von Neumann ANN implementations.

cs.ET

Training of Physical Neural Networks

Physical neural networks (PNNs) are a class of neural-like networks that leverage the properties of physical systems to perform computation. While PNNs are so far a niche research area with small-scale laboratory demonstrations, they are arguably one of the most underappreciated important opportunities in modern AI. Could we train AI models 1000x larger than current ones? Could we do this and also have them perform inference locally and privately on edge devices, such as smartphones or sensors? Research over the past few years has shown that the answer to all these questions is likely "yes, with enough research": PNNs could one day radically change what is possible and practical for AI systems. To do this will however require rethinking both how AI models work, and how they are trained - primarily by considering the problems through the constraints of the underlying hardware physics. To train PNNs at large scale, many methods including backpropagation-based and backpropagation-free approaches are now being explored. These methods have various trade-offs, and so far no method has been shown to scale to the same scale and performance as the backpropagation algorithm widely used in deep learning today. However, this is rapidly changing, and a diverse ecosystem of training techniques provides clues for how PNNs may one day be utilized to create both more efficient realizations of current-scale AI models, and to enable unprecedented-scale models.

physics.app-ph