SearcharxivSearch

arXiv subjects

Michael Brand

Publications and source records attributed to Michael Brand.

17 recordsLinked to original sources

An Axiomatisation of Error Intolerant Estimation

Point estimation is a fundamental statistical task. Given the wide selection of available point estimators, it is unclear, however, what, if any, would be universally-agreed theoretical reasons to generally prefer one such estimator over another. In this paper, we define a class of estimation scenarios which includes commonly-encountered problem situations such as both ``high stakes'' estimation and scientific inference, and introduce a new class of estimators, Error Intolerance Candidates (EIC) estimators, which we prove is optimal for it. EIC estimators are parameterised by an externally-given loss function. We prove, however, that even without such a loss function if one accepts a small number of incontrovertible-seeming assumptions regarding what constitutes a reasonable loss function, the optimal EIC estimator can be characterised uniquely. The optimal estimator derived in this second case is a previously-studied combination of maximum a posteriori (MAP) estimation and Wallace-Freeman (WF) estimation which has long been advocated among Minimum Message Length (MML) researchers, where it is derived as an approximation to the information-theoretic Strict MML estimator. Our results provide a novel justification for it that is purely Bayesian and requires neither approximations nor coding, placing both MAP and WF as special cases in the larger class of EIC estimators.

math.ST

Applying Trust for Operational States of ICT-Enabled Power Grid Services

Digitalization enables the automation required to operate modern cyber-physical energy systems (CPESs), leading to a shift from hierarchical to organic systems. However, digitalization increases the number of factors affecting the state of a CPES (e.g., software bugs and cyber threats). In addition to established factors like functional correctness, others like security become relevant but are yet to be integrated into an operational viewpoint, i.e. a holistic perspective on the system state. Trust in organic computing is an approach to gain a holistic view of the state of systems. It consists of several facets (e.g., functional correctness, security, and reliability), which can be used to assess the state of CPES. Therefore, a trust assessment on all levels can contribute to a coherent state assessment. This paper focuses on the trust in ICT-enabled grid services in a CPES. These are essential for operating the CPES, and their performance relies on various data aspects like availability, timeliness, and correctness. This paper proposes to assess the trust in involved components and data to estimate data correctness, which is crucial for grid services. The assessment is presented considering two exemplary grid services, namely state estimation and coordinated voltage control. Furthermore, the interpretation of different trust facets is also discussed.

cs.CR

Semantically enriched spatial modelling of industrial indoor environments enabling location-based services

This paper presents a concept for a software system called RAIL representing industrial indoor environments in a dynamic spatial model, aimed at easing development and provision of location-based services. RAIL integrates data from different sensor modalities and additional contextual information through a unified interface. Approaches to environmental modelling from other domains are reviewed and analyzed for their suitability regarding the requirements for our target domains; intralogistics and production. Subsequently a novel way of modelling data representing indoor space, and an architecture for the software system are proposed.

cs.RO

Verniered Optical Phased Arrays for Grating Lobe Suppression and Extended FOV

Optical phased arrays (OPAs) which beam-steer in 2D have so far been unable to pack emitting elements at $λ/2$ spacing, leading to grating lobes which limit the field-of-view, introduce signal ambiguity, and reduce optical efficiency. Vernier schemes, which use paired transmitter and receiver phased arrays with different periodicity, deliberately misalign the transmission and receive patterns so that only a single pairing of transmit/receive lobes permit a signal to be detected. A pair of OPAs designed to exploit this effect thereby effectively suppress the effects of grating lobes and recover the system's field-of-view, avoid potential ambiguities, and reduce excess noise. Here we analytically evaluate Vernier schemes with arbitrary phase control to find optimal configurations, as well as elucidate the manner in which a Vernier scheme can recover the full field-of-view. We present the first experimental implementation of a Vernier scheme and demonstrate grating lobe suppression using a pair of 2D wavelength-steered OPAs. These results present a route forward for addressing the pervasive issue of grating lobes, significantly alleviating the need for dense emitter pitches.

physics.app-ph

Serpentine optical phased arrays for scalable integrated photonic LIDAR beam steering

Optical phased arrays (OPAs) implemented in integrated photonic circuits could enable a variety of 3D sensing, imaging, illumination, and ranging applications, and their convergence in new LIDAR technology. However, current integrated OPA approaches do not scale - in control complexity, power consumption, and optical efficiency - to the large aperture sizes needed to support medium to long range LIDAR. We present the serpentine optical phased array (SOPA), a new OPA concept that addresses these fundamental challenges and enables architectures that scale up to large apertures. The SOPA is based on a serially interconnected array of low-loss grating waveguides and supports fully passive, two-dimensional (2D) wavelength-controlled beam steering. A fundamentally space-efficient design that folds the feed network into the aperture also enables scalable tiling of SOPAs into large apertures with a high fill-factor. We experimentally demonstrate the first SOPA, using a 1450 - 1650 nm wavelength sweep to produce 16,500 addressable spots in a 27x610 array. We also demonstrate, for the first time, far-field interference of beams from two separate OPAs on a single silicon photonic chip, as an initial step towards long-range computational imaging LIDAR based on novel active aperture synthesis schemes.

physics.optics

MML is not consistent for Neyman-Scott

Strict Minimum Message Length (SMML) is an information-theoretic statistical inference method widely cited (but only with informal arguments) as providing estimations that are consistent for general estimation problems. It is, however, almost invariably intractable to compute, for which reason only approximations of it (known as MML algorithms) are ever used in practice. Using novel techniques that allow for the first time direct, non-approximated analysis of SMML solutions, we investigate the Neyman-Scott estimation problem, an oft-cited showcase for the consistency of MML, and show that even with a natural choice of prior neither SMML nor its popular approximations are consistent for it, thereby providing a counterexample to the general claim. This is the first known explicit construction of an SMML solution for a natural, high-dimensional problem.

stat.ML

In-database connected component analysis

We describe a Big Data-practical, SQL-implementable algorithm for efficiently determining connected components for graph data stored in a Massively Parallel Processing (MPP) relational database. The algorithm described is a linear-space, randomised algorithm, always terminating with the correct answer but subject to a stochastic running time, such that for any $ε>0$ and any input graph $G=\langle V, E \rangle$ the algorithm terminates after $\mathop{\text{O}}(\log |V|)$ SQL queries with probability of at least $1-ε$, which we show empirically to translate to a quasi-linear runtime in practice.

cs.DS

A taxonomy of estimator consistency on discrete estimation problems

We describe a four-level hierarchy mapping both all discrete estimation problems and all estimators on these problems, such that the hierarchy describes each estimator's consistency guarantees on each problem class. We show that no estimator is consistent for all estimation problems, but that some estimators, such as Maximum A Posteriori, are consistent for the widest possible class of discrete estimation problems. For Maximum Likelihood and Approximate Maximum Likelihood estimators we show that they do not provide consistency on as wide a class, but define a sub-class of problems characterised by their consistency. Lastly, we show that some popular estimators, specifically Strict Minimum Message Length, do not provide consistency guarantees even within the sub-class.

math.ST

Risk-averse estimation, an axiomatic approach to inference, and Wallace-Freeman without MML

We define a new class of Bayesian point estimators, which we refer to as risk averse. Using this definition, we formulate axioms that provide natural requirements for inference, e.g. in a scientific setting, and show that for well-behaved estimation problems the axioms uniquely characterise an estimator. Namely, for estimation problems in which some parameter values have a positive posterior probability (such as, e.g., problems with a discrete hypothesis space), the axioms characterise Maximum A Posteriori (MAP) estimation, whereas elsewhere (such as in continuous estimation) they characterise the Wallace-Freeman estimator. Our results provide a novel justification for the Wallace-Freeman estimator, which previously was derived only as an approximation to the information-theoretic Strict Minimum Message Length estimator. By contrast, our derivation requires neither approximations nor coding.

stat.ML

RKL: a general, invariant Bayes solution for Neyman-Scott

Neyman-Scott is a classic example of an estimation problem with a partially-consistent posterior, for which standard estimation methods tend to produce inconsistent results. Past attempts to create consistent estimators for Neyman-Scott have led to ad-hoc solutions, to estimators that do not satisfy representation invariance, to restrictions over the choice of prior and more. We present a simple construction for a general-purpose Bayes estimator, invariant to representation, which satisfies consistency on Neyman-Scott over any non-degenerate prior. We argue that the good attributes of the estimator are due to its intrinsic properties, and generalise beyond Neyman-Scott as well.

stat.ML

The IMP game: Learnability, approximability and adversarial learning beyond $Σ^0_1$

We introduce a problem set-up we call the Iterated Matching Pennies (IMP) game and show that it is a powerful framework for the study of three problems: adversarial learnability, conventional (i.e., non-adversarial) learnability and approximability. Using it, we are able to derive the following theorems. (1) It is possible to learn by example all of $Σ^0_1 \cup Π^0_1$ as well as some supersets; (2) in adversarial learning (which we describe as a pursuit-evasion game), the pursuer has a winning strategy (in other words, $Σ^0_1$ can be learned adversarially, but $Π^0_1$ not); (3) some languages in $Π^0_1$ cannot be approximated by any language in $Σ^0_1$. We show corresponding results also for $Σ^0_i$ and $Π^0_i$ for arbitrary $i$.

cs.LO

Arbitrary Sequence RAMs

It is known that in some cases a Random Access Machine (RAM) benefits from having an additional input that is an arbitrary number, satisfying only the criterion of being sufficiently large. This is known as the ARAM model. We introduce a new type of RAM, which we refer to as the Arbitrary Sequence RAM (ASRAM), that generalises the ARAM by allowing the generation of additional arbitrary large numbers at will during execution time. We characterise the power contribution of this ability under several RAM variants. In particular, we demonstrate that an arithmetic ASRAM is more powerful than an arithmetic ARAM, that a sufficiently equipped ASRAM can recognise any language in the arithmetic hierarchy in constant time (and more, if it is given more time), and that, on the other hand, in some cases the ASRAM is no more powerful than its underlying RAM.

cs.CC

On the density of nice Friedmans

A Friedman number is a positive integer which is the result of an expression combining all of its own digits by use of the four basic operations, exponentiation and digit concatenation. A "nice" Friedman number is a Friedman number for which the expression constructing the number from its own digits can be represented with the original order of the digits unchanged. One of the fundamental questions regarding Friedman numbers, and particularly regarding nice Friedman numbers, is how common they are among the integers. In this paper, we prove that nice Friedman numbers have density 1, when considered in binary, ternary or base four.

math.NT

The RAM equivalent of P vs. RP

One of the fundamental open questions in computational complexity is whether the class of problems solvable by use of stochasticity under the Random Polynomial time (RP) model is larger than the class of those solvable in deterministic polynomial time (P). However, this question is only open for Turing Machines, not for Random Access Machines (RAMs). Simon (1981) was able to show that for a sufficiently equipped Random Access Machine, the ability to switch states nondeterministically does not entail any computational advantage. However, in the same paper, Simon describes a different (and arguably more natural) scenario for stochasticity under the RAM model. According to Simon's proposal, instead of receiving a new random bit at each execution step, the RAM program is able to execute the pseudofunction $\textit{RAND}(y)$, which returns a uniformly distributed random integer in the range $[0,y)$. Whether the ability to allot a random integer in this fashion is more powerful than the ability to allot a random bit remained an open question for the last 30 years. In this paper, we close Simon's open problem, by fully characterising the class of languages recognisable in polynomial time by each of the RAMs regarding which the question was posed. We show that for some of these, stochasticity entails no advantage, but, more interestingly, we show that for others it does.

cs.CC

Computing with and without arbitrary large numbers

In the study of random access machines (RAMs) it has been shown that the availability of an extra input integer, having no special properties other than being sufficiently large, is enough to reduce the computational complexity of some problems. However, this has only been shown so far for specific problems. We provide a characterization of the power of such extra inputs for general problems. To do so, we first correct a classical result by Simon and Szegedy (1992) as well as one by Simon (1981). In the former we show mistakes in the proof and correct these by an entirely new construction, with no great change to the results. In the latter, the original proof direction stands with only minor modifications, but the new results are far stronger than those of Simon (1981). In both cases, the new constructions provide the theoretical tools required to characterize the power of arbitrary large numbers.

cs.CC

Lower bounds on the Münchhausen problem

"The Baron's omni-sequence", B(n), first defined by Khovanova and Lewis (2011), is a sequence that gives for each n the minimum number of weighings on balance scales that can verify the correct labeling of n identically-looking coins with distinct integer weights between 1 gram and n grams. A trivial lower bound on B(n) is log_3(n), and it has been shown that B(n) is log_3(n) + O(log log n). In this paper we give a first nontrivial lower bound to the Münchhausen problem, showing that there is an infinite number of n values for which B(n) does not equal ceil(log_3 n). Furthermore, we show that if N(k) is the number of n values for which k = ceil(log_3 n) and B(n) does not equal k, then N(k) is an unbounded function of k.

cs.IT

Compressed Genotyping

Significant volumes of knowledge have been accumulated in recent years linking subtle genetic variations to a wide variety of medical disorders from Cystic Fibrosis to mental retardation. Nevertheless, there are still great challenges in applying this knowledge routinely in the clinic, largely due to the relatively tedious and expensive process of DNA sequencing. Since the genetic polymorphisms that underlie these disorders are relatively rare in the human population, the presence or absence of a disease-linked polymorphism can be thought of as a sparse signal. Using methods and ideas from compressed sensing and group testing, we have developed a cost-effective genotyping protocol. In particular, we have adapted our scheme to a recently developed class of high throughput DNA sequencing technologies, and assembled a mathematical framework that has some important distinctions from 'traditional' compressed sensing ideas in order to address different biological and technical constraints.

q-bio.GN