SearcharxivSearch

arXiv subjects

Itai Dattner

Publications and source records attributed to Itai Dattner.

15 recordsLinked to original sources

A nonparametric approach to understand multivariate quantile dynamics in financial time series

Over the last decade, nonparametric methods have gained increasing attention for modeling complex data structures due to their flexibility and minimal structural assumptions. In this paper, we study a general multivariate nonparametric regression framework that encompasses a broad class of parametric models commonly used in financial econometrics. Both the response and the covariate processes are allowed to be multivariate with fixed finite dimensions, and the framework accommodates temporal dependence, thereby introducing additional modeling and theoretical hurdles. To address these challenges, we adopt a functional dependence structure which permits flexible dynamic behavior while maintaining tractable asymptotic analysis. Within this setting, we establish strong and weak convergence results for the estimators of the conditional mean and volatility functions. In addition, we investigate conditional geometric quantiles in the multivariate time series context and prove their consistency under mild regularity conditions. The finite sample performance is examined through comprehensive simulation studies, and the methodology is illustrated by modeling the stock returns of Maersk and Lockheed Martin as a nonparametric function of a geopolitical risk index.

stat.ME

Adaptive Physics-Guided Neural Network

This paper introduces an adaptive physics-guided neural network (APGNN) framework for predicting quality attributes from image data by integrating physical laws into deep learning models. The APGNN adaptively balances data-driven and physics-informed predictions, enhancing model accuracy and robustness across different environments. Our approach is evaluated on both synthetic and real-world datasets, with comparisons to conventional data-driven models such as ResNet. For the synthetic data, 2D domains were generated using three distinct governing equations: the diffusion equation, the advection-diffusion equation, and the Poisson equation. Non-linear transformations were applied to these domains to emulate complex physical processes in image form. In real-world experiments, the APGNN consistently demonstrated superior performance in the diverse thermal image dataset. On the cucumber dataset, characterized by low material diversity and controlled conditions, APGNN and PGNN showed similar performance, both outperforming the data-driven ResNet. However, in the more complex thermal dataset, particularly for outdoor materials with higher environmental variability, APGNN outperformed both PGNN and ResNet by dynamically adjusting its reliance on physics-based versus data-driven insights. This adaptability allowed APGNN to maintain robust performance across structured, low-variability settings and more heterogeneous scenarios. These findings underscore the potential of adaptive physics-guided learning to integrate physical constraints effectively, even in challenging real-world contexts with diverse environmental conditions.

stat.ME

Physics-Guided Inverse Regression for Crop Quality Assessment

We present an innovative approach leveraging Physics-Guided Neural Networks (PGNNs) for enhancing agricultural quality assessments. Central to our methodology is the application of physics-guided inverse regression, a technique that significantly improves the model's ability to precisely predict quality metrics of crops. This approach directly addresses the challenges of scalability, speed, and practicality that traditional assessment methods face. By integrating physical principles, notably Fick`s second law of diffusion, into neural network architectures, our developed PGNN model achieves a notable advancement in enhancing both the interpretability and accuracy of assessments. Empirical validation conducted on cucumbers and mushrooms demonstrates the superior capability of our model in outperforming conventional computer vision techniques in postharvest quality evaluation. This underscores our contribution as a scalable and efficient solution to the pressing demands of global food supply challenges.

stat.ME

Model Selection for Ordinary Differential Equations: a Statistical Testing Approach

Ordinary differential equations (ODEs) are foundational in modeling intricate dynamics across a gamut of scientific disciplines. Yet, a possibility to represent a single phenomenon through multiple ODE models, driven by different understandings of nuances in internal mechanisms or abstraction levels, presents a model selection challenge. This study introduces a testing-based approach for ODE model selection amidst statistical noise. Rooted in the model misspecification framework, we adapt foundational insights from classical statistical paradigms (Vuong and Hotelling) to the ODE context, allowing for the comparison and ranking of diverse causal explanations without the constraints of nested models. Our simulation studies validate the theoretical robustness of our proposed test, revealing its consistent size and power. Real-world data examples further underscore the algorithm's applicability in practice. To foster accessibility and encourage real-world applications, we provide a user-friendly Python implementation of our model selection algorithm, bridging theoretical advancements with hands-on tools for the scientific community.

stat.ME

Modeling Motion Dynamics in Psychotherapy: a Dynamical Systems Approach

This study introduces a novel mechanistic modeling and statistical framework for analyzing motion energy dynamics within psychotherapy sessions. We transform raw motion energy data into an interpretable narrative of therapist-patient interactions, thereby revealing unique insights into the nature of these dynamics. Our methodology is established through three detailed case studies, each shedding light on the complexities of dyadic interactions. A key component of our approach is an analysis spanning four years of one therapist's sessions, allowing us to distinguish between trait-like and state-like dynamics. This research represents a significant advancement in the quantitative understanding of motion dynamics in psychotherapy, with the potential to substantially influence both future research and therapeutic practice.

q-bio.NC

A Robust Server-Effort Policy for Fluid Processing Networks

Multi-Class Processing Networks describe a set of servers that perform multiple classes of jobs on different items. A useful and tractable way to find an optimal control for such a network is to approximate it by a fluid model, resulting in a Separated Continuous Linear Programming (SCLP) problem. Clearly, arrival and service rates in such systems suffer from inherent uncertainty. A recent study addressed this issue by formulating a Robust Counterpart for SCLP models with budgeted uncertainty which provides a solution in terms of processing rates. This solution is transformed into a sequencing policy. However, in cases where servers can process several jobs simultaneously, a sequencing policy cannot be implemented. In this paper, we propose to use in these cases a a resource allocation policy, namely, the proportion of server effort per class. We formulate Robust Counterparts of both processing rates and server-effort uncertain models for four types of uncertainty sets: box, budgeted, one-sided budgeted, and polyhedral. We prove that server-effort model provides a better robust solution than any algebraic transformation of the robust solution of the processing rates model. Finally, to get a grasp of how much our new model improves over the processing rates robust model, we provide results of some numerical experiments.

math.OC

A statistical methodology for data-driven partitioning of infectious disease incidence into age-groups

Understanding age-group dynamics of infectious diseases is a fundamental issue for both scientific study and policymaking. Age-structure epidemic models were developed in order to study and improve our understanding of these dynamics. By fitting the models to incidence data of real outbreaks one can infer estimates of key epidemiological parameters. However, estimation of the transmission in an age-structured populations requires first to define the age-groups of interest. Misspecification in representing the heterogeneity in the age-dependent transmission rates can potentially lead to biased estimation of parameters. We develop the first statistical, data-driven methodology for deciding on the best partition of incidence data into age-groups. The method employs a top-down hierarchical partitioning algorithm, with a metric distance built for maximizing mathematical identifiability of the transmission matrix, and a stopping criteria based on significance testing. The methodology is tested using simulations showing good statistical properties. The methodology is then applied to influenza incidence data of 14 seasons in order to extract the significant age-group clusters in each season.

q-bio.PE

Separable nonlinear least-squares parameter estimation for complex dynamic systems

Nonlinear dynamic models are widely used for characterizing functional forms of processes that govern complex biological pathway systems. Over the past decade, validation and further development of these models became possible due to data collected via high-throughput experiments using methods from molecular biology. While these data are very beneficial, they are typically incomplete and noisy, so that inferring parameter values for complex dynamic models is associated with serious computational challenges. Fortunately, many biological systems have embedded linear mathematical features, which may be exploited, thereby improving fits and leading to better convergence of optimization algorithms. In this paper, we explore options of inference for dynamic models using a novel method of {\it separable nonlinear least-squares optimization}, and compare its performance to the traditional nonlinear least-squares method. The numerical results from extensive simulations suggest that the proposed approach is at least as accurate as the traditional nonlinear least-squares, but usually superior, while also enjoying a substantial reduction in computational time.

stat.ME

simode: R Package for statistical inference of ordinary differential equations using separable integral-matching

In this paper we describe simode: Separable Integral Matching for Ordinary Differential Equations. The statistical methodologies applied in the package focus on several minimization procedures of an integral-matching criterion function, taking advantage of the mathematical structure of the differential equations like separability of parameters from equations. Application of integral based methods to parameter estimation of ordinary differential equations was shown to yield more accurate and stable results comparing to derivative based ones. Linear features such as separability were shown to ease optimization and inference. We demonstrate the functionalities of the package using various systems of ordinary differential equations.

stat.CO

A Guided FP-growth algorithm for multitude-targeted mining of big data

In this paper we present the GFP-growth (Guided FP-growth) algorithm, a novel method for multitude-targeted mining: finding the count of a given large list of itemsets in large data. The GFP-growth algorithm is designed to focus on the specific multitude itemsets of interest and optimizes the time and memory costs. We prove that the GFP-growth algorithm yields the exact frequency-counts for the required itemsets. We show that for a number of different problems, a solution can be devised which takes advantage of the efficient implementation of multitude-targeted mining for boosting the performance. In particular, we study in detail the problem of generating the minority-class rules from imbalanced data, a scenario that appears in many real-life domains such as medical applications, failure prediction, network and cyber security, and maintenance. We develop the Minority-Report Algorithm that uses the GFP-growth for boosting performance. We prove some theoretical properties of the Minority-Report Algorithm and demonstrate its performance gain using simulations and real data.

cs.DB

Application of one-step method to parameter estimation in ODE models

In this paper we study application of Le Cam's one-step method to parameter estimation in ordinary differential equations models. This computationally simple technique can serve as an alternative to numerical evaluation of the popular nonlinear least squares estimator, which typically requires the use of a multi-step iterative algorithm and repetitive numerical integration of the ODE system. The one-step method starts from a preliminary $\sqrt{n}$-consistent estimator of the parameter of interest and next turns it into an asymptotic (as the sample size $n\rightarrow\infty$) equivalent of the least squares estimator through a numerically straightforward procedure. We demonstrate performance of the one-step estimator via extensive simulations and real data examples. The method enables the researcher to obtain both point and interval estimates. The preliminary $\sqrt{n}$-consistent estimator that we use depends on nonparametric smoothing, and we provide a data driven methodology for choosing its tuning parameter and support it by theory. An easy implementation scheme of the one-step method for practical use is pointed out.

stat.ME

A two-stage approach for estimating the parameters of an age-group epidemic model from incidence data

Age-dependent dynamics is an important characteristic of many infectious diseases. Age-group epidemic models describe the infection dynamics in different age-groups by allowing to set distinct parameter values for each. However, such models are highly nonlinear and may have a large number of unknown parameters. Thus, parameter estimation of age-group models, while becoming a fundamental issue for both the scientific study and policy making in infectious diseases, is not a trivial task in practice. In this paper, we examine the estimation of the so called next-generation matrix using incidence data of a single entire outbreak, and extend the approach to deal with recurring outbreaks. Unlike previous studies, we do not assume any constraints regarding the structure of the matrix. A novel two-stage approach is developed, which allows for efficient parameter estimation from both statistical and computational perspectives. Simulation studies corroborate the ability to estimate accurately the parameters of the model for several realistic scenarios. The model and estimation method are applied to real data of influenza-like-illness in Israel. The parameter estimates of the key relevant epidemiological parameters and the recovered structure of the estimated next-generation matrix are in line with results obtained in previous studies.

stat.ME

Consistency of direct integral estimator for partially observed systems of ordinary differential equations linear in the parameters

Dynamic systems are ubiquitous in nature and are used to model many processes in biology, chemistry, physics, medicine, and engineering. In particular, systems of ordinary differential equations are commonly used for the mathematical modelling of the rate of change of dynamic processes. In many practical applications, the process can only be partially measured, a fact that renders estimation of parameters of the system extremely challenging. Recently, a 'direct integral estimator' for partially observed systems of ordinary differential equations was introduced. The practical performance of the integral estimator was demonstrated, but its theoretical properties were not derived. In this paper we use the sieve framework to prove that the estimator is consistent.

math.ST

Adaptive quantile estimation in deconvolution with unknown error distribution

Quantile estimation in deconvolution problems is studied comprehensively. In particular, the more realistic setup of unknown error distributions is covered. Our plug-in method is based on a deconvolution density estimator and is minimax optimal under minimal and natural conditions. This closes an important gap in the literature. Optimal adaptive estimation is obtained by a data-driven bandwidth choice. As a side result, we obtain optimal rates for the plug-in estimation of distribution functions with unknown error distributions. The method is applied to a real data example.

math.ST

Optimal Rate of Direct Estimators in Systems of Ordinary Differential Equations Linear in Functions of the Parameters

Many processes in biology, chemistry, physics, medicine, and engineering are modeled by a system of differential equations. Such a system is usually characterized via unknown parameters and estimating their 'true' value is thus required. In this paper we focus on the quite common systems for which the derivatives of the states may be written as sums of products of a function of the states and a function of the parameters. For such a system linear in functions of the unknown parameters we present a necessary and sufficient condition for identifiability of the parameters. We develop an estimation approach that bypasses the heavy computational burden of numerical integration and avoids the estimation of system states derivatives, drawbacks from which many classic estimation methods suffer. We also suggest an experimental design for which smoothing can be circumvented. The optimal rate of the proposed estimators, i.e., their $\sqrt n$-consistency, is proved and simulation results illustrate their excellent finite sample performance and compare it to other estimation approaches.

math.ST