SearcharxivSearch

arXiv subjects

Jorge F. Silva

Publications and source records attributed to Jorge F. Silva.

At least 19 recordsLinked to original sources

A review on fundamental bounds and estimators for photometry and astrometry of celestial point sources using array detectors, from first principles

Precise astrometric and photometric measurements of celestial point sources are fundamental to modern astronomy. These measurements, used to determine object positions, motions, and fluxes, are based on observational models that have evolved from empirical centroiding rules to rigorous probabilistic formulations at the pixel level. This review summarizes key contributions that formalized this transition and analyzes seminal works addressing both the theoretical limits and the empirical performance of estimators. Central to these developments is the derivation of fundamental bounds, such as the Cramér-Rao Lower Bound (CRLB), and the assessment of widely used estimators, including Maximum Likelihood (ML), Least Squares (LS), and Weighted Least Squares (WLS). These studies show that, while the CRLB sets a theoretical benchmark, practical estimators achieve it only under specific signal-to-noise ratio (SNR) regimes, with notable discrepancies in high-SNR conditions. Moreover, recent results demonstrate that jointly estimating source flux and background significantly improves photometric precision compared to sequential approaches. Looking ahead, the increasing complexity of astronomical surveys, driven by massive data volumes, dynamic observational conditions, and the integration of machine learning, poses new challenges to reliable inference. In this context, tools from statistical theory, including performance bounds and theoretically grounded estimators, remain critical to guide algorithm design and ensure robust astrometric and photometric pipelines.

astro-ph.IM

A novel Information-Driven Strategy for Optimal Regression Assessment

In Machine Learning (ML), a regression algorithm aims to minimize a loss function based on data. An assessment method in this context seeks to quantify the discrepancy between the optimal response for an input-output system and the estimate produced by a learned predictive model (the student). Evaluating the quality of a learned regressor remains challenging without access to the true data-generating mechanism, as no data-driven assessment method can ensure the achievability of global optimality. This work introduces the Information Teacher, a novel data-driven framework for evaluating regression algorithms with formal performance guarantees to assess global optimality. Our novel approach builds on estimating the Shannon mutual information (MI) between the input variables and the residuals and applies to a broad class of additive noise models. Through numerical experiments, we confirm that the Information Teacher is capable of detecting global optimality, which is aligned with the condition of zero estimation error with respect to the -- inaccessible, in practice -- true model, working as a surrogate measure of the ground truth assessment loss and offering a principled alternative to conventional empirical performance metrics.

stat.ML

The impact of the point spread function fitting radius on photometric uncertainty based on the Fisher information matrix

In point spread function (PSF) photometry, the selection of the fitting aperture radius plays a critical role in determining the precision of flux and background estimations. Traditional methods often rely on maximizing the signal-to-noise ratio (S/N) as a criterion for aperture selection. However, S/N-based approaches do not necessarily provide the optimal precision for joint estimation problems as they do not account for the statistical limits imposed by the Fisher information in the context of the Cramér-Rao lower bound (CRLB). This study aims to establish an alternative criterion for selecting the optimal fitting radius based on Fisher information rather than S/N. Fisher information serves as a fundamental measure of estimation precision, providing theoretical guarantees on the achievable accuracy for parameter estimation. By leveraging Fisher information, we seek to define an aperture selection strategy that minimizes the loss of precision. We conducted a series of numerical experiments that analyze the behavior of Fisher information and estimator performance as a function of the PSF aperture radius. Specifically, we revisited fundamental photometric models and explored the relationship between aperture size and information content. We compared the empirical variance of classical estimators, such as maximum likelihood and stochastic weighted least squares, against the theoretical CRLB derived from the Fisher information matrix. Our results indicate that aperture selection based on the Fisher information provides a more robust framework for achieving optimal estimation precision.

astro-ph.IM

Fault Detection and Monitoring using a Data-Driven Information-Based Strategy: Method, Theory, and Application

The ability to detect when a system undergoes an incipient fault is of paramount importance in preventing a critical failure. Classic methods for fault detection (including model-based and data-driven approaches) rely on thresholding error statistics or simple input-residual dependencies but face difficulties with non-linear or non-Gaussian systems. Behavioral methods (e.g., those relying on digital twins) address these difficulties but still face challenges when faulty data is scarce, decision guarantees are required, or working with already-deployed models is required. In this work, we propose an information-driven fault detection method based on a novel concept drift detector, addressing these challenges. The method is tailored to identifying drifts in input-output relationships of additive noise models (i.e., model drifts) and is based on a distribution-free mutual information (MI) estimator. Our scheme does not require prior faulty examples and can be applied distribution-free over a large class of system models. Our core contributions are twofold. First, we demonstrate the connection between fault detection, model drift detection, and testing independence between two random variables. Second, we prove several theoretical properties of the proposed MI-based fault detection scheme: (i) strong consistency, (ii) exponentially fast detection of the non-faulty case, and (iii) control of both significance levels and power of the test. To conclude, we validate our theory with synthetic data and the benchmark dataset N-CMAPSS of aircraft turbofan engines. These empirical results support the usefulness of our methodology in many practical and realistic settings, and the theoretical results show performance guarantees that other methods cannot offer.

eess.SP

Understanding Encoder-Decoder Structures in Machine Learning Using Information Measures

We present new results to model and understand the role of encoder-decoder design in machine learning (ML) from an information-theoretic angle. We use two main information concepts, information sufficiency (IS) and mutual information loss (MIL), to represent predictive structures in machine learning. Our first main result provides a functional expression that characterizes the class of probabilistic models consistent with an IS encoder-decoder latent predictive structure. This result formally justifies the encoder-decoder forward stages many modern ML architectures adopt to learn latent (compressed) representations for classification. To illustrate IS as a realistic and relevant model assumption, we revisit some known ML concepts and present some interesting new examples: invariant, robust, sparse, and digital models. Furthermore, our IS characterization allows us to tackle the fundamental question of how much performance (predictive expressiveness) could be lost, using the cross entropy risk, when a given encoder-decoder architecture is adopted in a learning setting. Here, our second main result shows that a mutual information loss quantifies the lack of expressiveness attributed to the choice of a (biased) encoder-decoder ML design. Finally, we address the problem of universal cross-entropy learning with an encoder-decoder design where necessary and sufficiency conditions are established to meet this requirement. In all these results, Shannon's information measures offer new interpretations and explanations for representation learning.

cs.LG

Optimal photometry of point sources: Joint source flux and background determination on array detectors -- from theory to practical implementation

In this paper we study the joint determination of source and background flux for point sources as observed by digital array detectors. We explicitly compute the two-dimensional Cramér-Rao absolute lower bound (CRLB) as well as the performance bounds for high-dimensional implicit estimators from a generalized Taylor expansion. This later approach allows us to obtain computable prescriptions for the bias and variance of the joint estimators. We compare these prescriptions with empirical results from numerical simulations in the case of the weighted least squares estimator (introducing an improved version, denoted stochastic weighted least-squares) as well as with the maximum likelihood estimator, finding excellent agreement. We demonstrate that these estimators provide quasi-unbiased joint estimations of the flux and background, with a variance that approaches the CRLB very tightly and are, hence, optimal, unlike the case of sequential estimation used commonly in astronomical photometry which is sub-optimal. We compare our predictions with numerical simulations of realistic observations, as well as with observations of a bona-fide non-variable stellar source observed with TESS, and compare it to the results from the sequential estimation of background and flux, confirming our theoretical expectations. Our practical estimators can be used as benchmarks for general photometric pipelines, or for applications that require maximum precision and accuracy in absolute photometry.

astro-ph.IM

Gaussian process deconvolution

Let us consider the deconvolution problem, that is, to recover a latent source $x(\cdot)$ from the observations $\mathbf{y} = [y_1,\ldots,y_N]$ of a convolution process $y = x\star h + η$, where $η$ is an additive noise, the observations in $\mathbf{y}$ might have missing parts with respect to $y$, and the filter $h$ could be unknown. We propose a novel strategy to address this task when $x$ is a continuous-time signal: we adopt a Gaussian process (GP) prior on the source $x$, which allows for closed-form Bayesian nonparametric deconvolution. We first analyse the direct model to establish the conditions under which the model is well defined. Then, we turn to the inverse problem, where we study i) some necessary conditions under which Bayesian deconvolution is feasible, and ii) to which extent the filter $h$ can be learnt from data or approximated for the blind deconvolution case. The proposed approach, termed Gaussian process deconvolution (GPDC) is compared to other deconvolution methods conceptually, via illustrative examples, and using real-world datasets.

stat.ML

Optimal observational scheduling framework for binary and multiple stellar systems

The optimal instant of observation of astrophysical phenomena for objects that vary on human time-sales is an important problem, as it bears on the cost-effective use of usually scarce observational facilities. In this paper we address this problem for the case of tight visual binary systems through a Bayesian framework based on the maximum entropy sampling principle. Our proposed information-driven methodology exploits the periodic structure of binary systems to provide a computationally efficient estimation of the probability distribution of the optimal observation time. We show the optimality of the proposed sampling methodology in the Bayes sense and its effectiveness through direct numerical experiments. We successfully apply our scheme to the study of two visual-spectroscopic binaries, and one purely astrometric triple hierarchical system. We note that our methodology can be applied to any time-evolving phenomena, a particularly interesting application in the era of dedicated surveys, where a definition of the cadence of observations can have a crucial impact on achieving the science goals.

astro-ph.SR

Lossy Compression for Robust Unsupervised Time-Series Anomaly Detection

A new Lossy Causal Temporal Convolutional Neural Network Autoencoder for anomaly detection is proposed in this work. Our framework uses a rate-distortion loss and an entropy bottleneck to learn a compressed latent representation for the task. The main idea of using a rate-distortion loss is to introduce representation flexibility that ignores or becomes robust to unlikely events with distinctive patterns, such as anomalies. These anomalies manifest as unique distortion features that can be accurately detected in testing conditions. This new architecture allows us to train a fully unsupervised model that has high accuracy in detecting anomalies from a distortion score despite being trained with some portion of unlabelled anomalous data. This setting is in stark contrast to many of the state-of-the-art unsupervised methodologies that require the model to be only trained on "normal data". We argue that this partially violates the concept of unsupervised training for anomaly detection as the model uses an informed decision that selects what is normal from abnormal for training. Additionally, there is evidence to suggest it also effects the models ability at generalisation. We demonstrate that models that succeed in the paradigm where they are only trained on normal data fail to be robust when anomalous data is injected into the training. In contrast, our compression-based approach converges to a robust representation that tolerates some anomalous distortion. The robust representation achieved by a model using a rate-distortion loss can be used in a more realistic unsupervised anomaly detection scheme.

cs.LG

Bayesian inference in single-line spectroscopic binaries with a visual orbit

We present a Bayesian inference methodology for the estimation of orbital parameters on single-line spectroscopic binaries with astrometric data, based on the No-U-Turn sampler Markov chain Monte Carlo algorithm. Our approach is designed to provide a precise and efficient estimation of the joint posterior distribution of the orbital parameters in the presence of partial and heterogeneous observations. This scheme allows us to directly incorporate prior information about the system - in the form of a trigonometric parallax, and an estimation of the mass of the primary component from its spectral type - to constrain the range of solutions, and to estimate orbital parameters that cannot be usually determined (e.g. the individual component masses), due to the lack of observations or imprecise measurements. Our methodology is tested by analyzing the posterior distributions of well-studied double-line spectroscopic binaries treated as single-line binaries by omitting the radial velocity data of the secondary object. Our results show that the system's mass ratio can be estimated with an uncertainty smaller than 10% using our approach. As a proof of concept, the proposed methodology is applied to twelve single-line spectroscopic binaries with astrometric data that lacked a joint astrometric-spectroscopic solution, for which we provide full orbital elements. Our sample-based methodology allows us also to study the impact of different posterior distributions in the corresponding observations space. This novel analysis provides a better understanding of the effect of the different sources of information on the shape and uncertainty of the orbit and radial velocity curve.

astro-ph.SR

Studying the Interplay between Information Loss and Operation Loss in Representations for Classification

Information-theoretic measures have been widely adopted in the design of features for learning and decision problems. Inspired by this, we look at the relationship between i) a weak form of information loss in the Shannon sense and ii) the operation loss in the minimum probability of error (MPE) sense when considering a family of lossy continuous representations (features) of a continuous observation. We present several results that shed light on this interplay. Our first result offers a lower bound on a weak form of information loss as a function of its respective operation loss when adopting a discrete lossy representation (quantization) instead of the original raw observation. From this, our main result shows that a specific form of vanishing information loss (a weak notion of asymptotic informational sufficiency) implies a vanishing MPE loss (or asymptotic operational sufficiency) when considering a general family of lossy continuous representations. Our theoretical findings support the observation that the selection of feature representations that attempt to capture informational sufficiency is appropriate for learning, but this selection is a rather conservative design principle if the intended goal is achieving MPE in classification. Supporting this last point, and under some structural conditions, we show that it is possible to adopt an alternative notion of informational sufficiency (strictly weaker than pure sufficiency in the mutual information sense) to achieve operational sufficiency in learning.

cs.LG

Universal Weak Variable-Length Source Coding on Countable Infinite Alphabets

Motivated from the fact that universal source coding on countably infinite alphabets is not feasible, this work introduces the notion of almost lossless source coding. Analog to the weak variable-length source coding problem studied by Han (IEEE TIT, 2000, 46, 1217-1226), almost lossless source coding aims at relaxing the lossless block-wise assumption to allow an average per-letter distortion that vanishes asymptotically as the block-length tends to infinity. In this setup, we show on one hand that Shannon entropy characterizes the minimum achievable rate (similarly to the case of finite alphabet sources) while on the other that almost lossless universal source coding becomes feasible for the family of finite-entropy stationary memoryless sources with infinite alphabets. Furthermore, we study a stronger notion of almost lossless universality that demands uniform convergence of the average per-letter distortion to zero, where we establish a necessary and sufficient condition for the so-called family of envelope distributions to achieve it. Remarkably, this condition is the same necessary and sufficient condition needed for the existence of a strongly minimax (lossless) universal source code for the family of envelope distributions. Finally, we show that an almost lossless coding scheme offers faster rate of convergence for the (minimax) redundancy compared to the well-known information radius developed for the lossless case at the expense of tolerating a non-zero distortion that vanishes to zero as the block-length grows. This shows that even when lossless universality is feasible, an almost lossless scheme can offer different regimes on the rates of convergence of the (worst case) redundancy versus the (worst case) distortion.

cs.IT

On the Exponential Approximation of Type II Error Probability of Distributed Test of Independence

This paper studies distributed binary test of statistical independence under communication (information bits) constraints. While testing independence is very relevant in various applications, distributed independence test is particularly useful for event detection in sensor networks where data correlation often occurs among observations of devices in the presence of a signal of interest. By focusing on the case of two devices because of their tractability, we begin by investigating conditions on Type I error probability restrictions under which the minimum Type II error admits an exponential behavior with the sample size. Then, we study the finite sample-size regime of this problem. We derive new upper and lower bounds for the gap between the minimum Type II error and its exponential approximation under different setups, including restrictions imposed on the vanishing Type I error probability. Our theoretical results shed light on the sample-size regimes at which approximations of the Type II error probability via error exponents became informative enough in the sense of predicting well the actual error probability. We finally discuss an application of our results where the gap is evaluated numerically, and we show that exponential approximations are not only tractable but also a valuable proxy for the Type II probability of error in the finite-length regime.

math.ST

Finite-Length Bounds on Hypothesis Testing Subject to Vanishing Type I Error Restrictions

A central problem in Binary Hypothesis Testing (BHT) is to determine the optimal tradeoff between the Type I error (referred to as false alarm) and Type II (referred to as miss) error. In this context, the exponential rate of convergence of the optimal miss error probability -- as the sample size tends to infinity -- given some (positive) restrictions on the false alarm probabilities is a fundamental question to address in theory. Considering the more realistic context of a BHT with a finite number of observations, this paper presents a new non-asymptotic result for the scenario with monotonic (sub-exponential decreasing) restriction on the Type I error probability, which extends the result presented by Strassen in 2009. Building on the use of concentration inequalities, we offer new upper and lower bounds to the optimal Type II error probability for the case of finite observations. Finally, the derived bounds are evaluated and interpreted numerically (as a function of the number samples) for some vanishing Type I error restrictions.

cs.IT

Data-Driven Representations for Testing Independence: Modeling, Analysis and Connection with Mutual Information Estimation

This work addresses testing the independence of two continuous and finite-dimensional random variables from the design of a data-driven partition. The empirical log-likelihood statistic is adopted to approximate the sufficient statistics of an oracle test against independence (that knows the two hypotheses). It is shown that approximating the sufficient statistics of the oracle test offers a learning criterion for designing a data-driven partition that connects with the problem of mutual information estimation. Applying these ideas in the context of a data-dependent tree-structured partition (TSP), we derive conditions on the TSP's parameters to achieve a strongly consistent distribution-free test of independence over the family of probabilities equipped with a density. Complementing this result, we present finite-length results that show our TSP scheme's capacity to detect the scenario of independence structurally with the data-driven partition as well as new sampling complexity bounds for this detection. Finally, some experimental analyses provide evidence regarding our scheme's advantage for testing independence compared with some strategies that do not use data-driven representations.

stat.ML

On Universal D-Semifaithful Coding for Memoryless Sources with Infinite Alphabets

The problem of variable length and fixed-distortion universal source coding (or D-semifaithful source coding) for stationary and memoryless sources on countably infinite alphabets ($\infty$-alphabets) is addressed in this paper. The main results of this work offer a set of sufficient conditions (from weaker to stronger) to obtain weak minimax universality, strong minimax universality, and corresponding achievable rates of convergences for the worse-case redundancy for the family of stationary memoryless sources whose densities are dominated by an envelope function (or the envelope family) on $\infty$-alphabets. An important implication of these results is that universal D-semifaithful source coding is not feasible for the complete family of stationary and memoryless sources on $\infty$-alphabets. To demonstrate this infeasibility, a sufficient condition for the impossibility is presented for the envelope family. Interestingly, it matches the well-known impossibility condition in the context of lossless (variable-length) universal source coding. More generally, this work offers a simple description of what is needed to achieve universal D-semifaithful coding for a family of distributions $Λ$. This reduces to finding a collection of quantizations of the product space at different block-lengths -- reflecting the fixed distortion restriction -- that satisfy two asymptotic requirements: the first is a universal quantization condition with respect to $Λ$, and the second is a vanishing information radius (I-radius) condition for $Λ$ reminiscent of the condition known for lossless universal source coding.

cs.IT

Compressibility Analysis of Asymptotically Mean Stationary Processes

This work provides new results for the analysis of random sequences in terms of $\ell_p$-compressibility. The results characterize the degree in which a random sequence can be approximated by its best $k$-sparse version under different rates of significant coefficients (compressibility analysis). In particular, the notion of strong $\ell_p$-characterization is introduced to denote a random sequence that has a well-defined asymptotic limit (sample-wise) of its best $k$-term approximation error when a fixed rate of significant coefficients is considered (fixed-rate analysis). The main theorem of this work shows that the rich family of asymptotically mean stationary (AMS) processes has a strong $\ell_p$-characterization. Furthermore, we present results that characterize and analyze the $\ell_p$-approximation error function for this family of processes. Adding ergodicity in the analysis of AMS processes, we introduce a theorem demonstrating that the approximation error function is constant and determined in closed-form by the stationary mean of the process. Our results and analyses contribute to the theory and understanding of discrete-time sparse processes and, on the technical side, confirm how instrumental the point-wise ergodic theorem is to determine the compressibility expression of discrete-time processes even when stationarity and ergodicity assumptions are relaxed.

stat.ME

Bayes-based orbital elements estimation in triple hierarchical stellar systems

Under certain rather prevalent conditions (driven by dynamical orbital evolution), a hierarchical triple stellar system can be well approximated, from the standpoint of orbital parameter estimation, as two binary star systems combined. Even under this simplifying approximation, the inference of orbital elements is a challenging technical problem because of the high dimensionality of the parameter space, and the complex relationships between those parameters and the observations (astrometry and radial velocity). In this work we propose a new methodology for the study of triple hierarchical systems using a Bayesian Markov-Chain Monte Carlo-based framework. In particular, graphical models are introduced to describe the probabilistic relationship between parameters and observations in a dynamically self-consistent way. As information sources we consider the cases of isolated astrometry, isolated radial velocity, as well as the joint case with both types of measurements. Graphical models provide a novel way of performing a factorization of the joint distribution (of parameter and observations) in terms of conditional independent components (factors), so that the estimation can be performed in a two-stage process that combines different observations sequentially. Our framework is tested against three well-studied benchmark cases of triple systems, where we determine the inner and outer orbital elements, coupled with the mutual inclination of the orbits, and the individual stellar masses, along with posterior probability (density) distributions for all these parameters. Our results are found to be consistent with previous studies. We also provide a mathematical formalism to reduce the dimensionality in the parameter space for triple hierarchical stellar systems in general.

astro-ph.SR