SearcharxivSearch

arXiv subjects

Marcin Pitera

Publications and source records attributed to Marcin Pitera.

At least 19 recordsLinked to original sources

WANDR: A Benchmark for Wide and Deep Research

WANDR (Wide ANd Deep Research) is a benchmark of 500 realistic, challenging data-collection tasks for research agents. Each task requires a system to discover a large set of entities that satisfy specified criteria (breadth), investigate each entity through multiple coordinated web searches (depth), and return independently verifiable records with supporting sources and excerpts. Tasks are represented as qualification key hierarchies that specify the entities, relationships, evidence, and required count at each level; a hierarchy with n companies, m employees per company, and k sources per employee requires n x m x k records. This structure supports diverse workflows such as market mapping, due diligence, literature review, product comparison, and talent sourcing, with targets ranging from dozens to thousands of records. WANDR replaces static gold answer sets with task-specific judges that refetch cited pages and verify each record against its evidence, allowing evaluation of current and changing facts. Record verdicts are aggregated into soft and hard precision, recall, and F1 scores that distinguish factual quality, coverage, and hierarchical completeness. The tasks are derived from de-identified product-usage logs and produced through a semi-automated pipeline with automated checks, empirical audits, and human review where needed. We evaluate six production research systems and find that the benchmark is far from saturated: at high effort, the strongest system reaches only 0.363 soft F1 and 0.133 hard F1. Performance degrades as target volume and hierarchy depth increase, with incomplete discovery, missing enrichment, and incomplete evidence construction remaining major bottlenecks. The benchmark and evaluation harness are available at https://github.com/perplexityai/wandr.

cs.LG

Blackwell optimality in risk-sensitive stochastic control

In this paper, we consider a discrete-time Markov Decision Process (MDP) on a finite state-action space with a long-run risk-sensitive criterion used as the objective function. We discuss the concept of Blackwell optimality and comment on intricacies which arise when the risk-neutral expectation is replaced by the risk-sensitive entropy. Also, we show the relation between the Blackwell optimality and ultimate stationarity and provide an illustrative example that helps to better understand the structural difference between these two concepts.

math.OC

Policy stability and ultimate stationarity in discounted risk-sensitive stochastic control

We study discrete-time Markov Decision Processes (MDPs) on finite state-action spaces and analyze the stability of optimal policies and value functions in the long-run discounted risk-sensitive objective setting. Our analysis addresses robustness with respect to perturbations of the risk-aversion parameter and the discount factor, the emergence of ultimate stationarity, and the interaction between discounted and averaged formulations under suitable mixing assumptions. We further investigate limiting regimes associated with vanishing discount and vanishing risk sensitivity, and discuss the role of Blackwell-type stability properties in the discounted setting. Finally, we provide numerical illustrations that highlight the intrinsic non-stationarity of optimal discounted risk-sensitive policies.

math.OC

Coherent estimation of risk measures

We develop a statistical framework for risk estimation, inspired by the axiomatic theory of risk measures. Coherent risk estimators -- functionals of P\&L samples inheriting the economic properties of risk measures -- are defined and characterized through robust representations linked to $L$-estimators. The framework provides a canonical methodology for constructing estimators with sound financial and statistical properties, unifying risk measure theory, principles for capital adequacy, and practical statistical challenges in market risk. Numerical illustrations based on simulated and market data demonstrate that coherence of a risk measure does not necessarily carry over to its estimators and show that alternative admissible weight structures within the CRE representation can lead to substantially different capital adequacy outcomes.

q-fin.RM

Statistical applications of the 20/60/20 rule in risk management and portfolio optimization

This paper explores the applications of the 20/60/20 rule-a heuristic method that segments data into top-performing, average-performing, and underperforming groups-in mathematical finance. We review the statistical foundations of this rule and demonstrate its usefulness in risk management and portfolio optimization. Our study highlights three key applications. First, we apply the rule to stock market data, showing that it enables effective population clustering. Second, we introduce a novel, easy-to-implement method for extracting heavy-tail characteristics in risk management. Third, we integrate spatial reasoning based on the 20/60/20 rule into portfolio optimization, enhancing robustness and improving performance. To support our findings, we develop a new measure for quantifying tail heaviness and employ conditional statistics to reconstruct the unconditional distribution from the core data segment. This reconstructed distribution is tested on real financial data to evaluate whether the 20/60/20 segmentation effectively balances capturing extreme risks with maintaining the stability of central returns. Our results offer insights into financial data behavior under heavy-tailed conditions and demonstrate the potential of the 20/60/20 rule as a complementary tool for decision-making in finance.

q-fin.PM

Blackwell optimality and policy stability for long-run risk sensitive stochastic control

This paper analyzes the stability of optimal policies in the long-run stochastic control framework with an averaged risk-sensitive criterion for discrete-time MDPs on finite state-action space. In particular, we study the robustness of optimal controls when perturbations to the risk-aversion parameter are applied, and investigate the Blackwell property, together with its link to the risk-sensitive vanishing discount approximation framework. Finally, we present examples that help to better understand the intricacies of the risk-sensitive control framework.

math.OC

Conditional correlation estimation and serial dependence identification

It has been recently shown in Jaworski, P., Jelito, D. and Pitera, M. (2024), 'A note on the equivalence between the conditional uncorrelation and the independence of random variables', Electronic Journal of Statistics 18(1), that one can characterise the independence of random variables via the family of conditional correlations on quantile-induced sets. This effectively shows that the localized linear measure of dependence is able to detect any form of nonlinear dependence for appropriately chosen conditioning sets. In this paper, we expand this concept, focusing on the statistical properties of conditional correlation estimators and their potential usage in serial dependence identification. In particular, we show how to estimate conditional correlations in generic and serial dependence setups, discuss key properties of the related estimators, define the conditional equivalent of the autocorrelation function, and provide a series of examples which prove that the proposed framework could be efficiently used in many practical econometric applications.

stat.ME

Gaussian dependence structure pairwise goodness-of-fit testing based on conditional covariance and the 20/60/20 rule

We present a novel data-oriented statistical framework that assesses the presumed Gaussian dependence structure in a pairwise setting. This refers to both multivariate normality and normal copula goodness-of-fit testing. The proposed test clusters the data according to the 20/60/20 rule and confronts conditional covariance (or correlation) estimates on the obtained subsets. The corresponding test statistic has a natural practical interpretation, desirable statistical properties, and asymptotic pivotal distribution under the multivariate normality assumption. We illustrate the usefulness of the introduced framework using extensive power simulation studies and show that our approach outperforms popular benchmark alternatives. Also, we apply the proposed methodology to commodities market data.

stat.ME

A novel scaling approach for unbiased adjustment of risk estimators

The assessment of risk based on historical data faces many challenges, in particular due to the limited amount of available data, lack of stationarity, and heavy tails. While estimation on a short-term horizon for less extreme percentiles tends to be reasonably accurate, extending it to longer time horizons or extreme percentiles poses significant difficulties. The application of theoretical risk scaling laws to address this issue has been extensively explored in the literature. This paper presents a novel approach to scaling a given risk estimator, ensuring that the estimated capital reserve is robust and conservatively estimates the risk. We develop a simple statistical framework that allows efficient risk scaling and has a direct link to backtesting performance. Our method allows time scaling beyond the conventional square-root-of-time rule, enables risk transfers, such as those involved in economic capital allocation, and could be used for unbiased risk estimation in small sample settings. To demonstrate the effectiveness of our approach, we provide various examples related to the estimation of value-at-risk and expected shortfall together with a short empirical study analysing the impact of our method.

q-fin.RM

Goodness-of-fit tests for the one-sided L\'evy distribution based on quantile conditional moments

In this paper we introduce a novel statistical framework based on the first two quantile conditional moments that facilitates effective goodness-of-fit testing for one-sided L\'evy distributions. The scale-ratio framework introduced in this paper extends our previous results in which we have shown how to extract unique distribution features using conditional variance ratio for the generic class of {\alpha}-stable distributions. We show that the conditional moment-based goodness-of-fit statistics are a good alternative to other methods introduced in the literature tailored to the one-sided L\'evy distributions. The usefulness of our approach is verified using an empirical test power study. For completeness, we also derive the asymptotic distributions of the test statistics and show how to apply our framework to real data.

stat.ME

Utility-based acceptability indices

In this short paper we introduce a new class of performance measures based on certainty equivalents defined via scaled utility functions. We analyse their properties, show that the corresponding portfolio optimization problem is well-posed under generic conditions, and analyse the link between portfolio dynamics, benchmark process, and utility function choice in the long-run setting.

q-fin.RM

Existence of bounded solutions to multiplicative Poisson equations under mixing property

In this paper we study the problem of Multiplicative Poisson Equation (MPE) bounded solution existence in the generic discrete-time setting. Assuming mixing and boundedness of the risk-reward function, we investigate what conditions should be imposed on the underlying non-controlled probability kernel or the reward function in order for the MPE bounded solution to always exists. In particular, we consolidate span-norm framework based results and derive an explicit sharp bound that needs to be imposed on the cost function to guarantee the bounded solution existence under mixing. Also, we study the properties which the probability kernel must satisfy to ensure existence of bounded MPE for any generic risk-reward function and characterise process behaviour in the complement of the invariant measure support. Finally, we present numerous examples and stochastic-dominance based arguments that help to better understand the intricacies that emerge when the ergodic risk-neutral mean operator is replaced with ergodic risk-sensitive entropy.

math.PR

Estimation of stability index for symmetric {\alpha}-stable distribution using quantile conditional variance ratios

The class of $\alpha$-stable distributions is widely used in various applications, especially for modelling heavy-tailed data. Although the $\alpha$-stable distributions have been used in practice for many years, new methods for identification, testing, and estimation are still being refined and new approaches are being proposed. The constant development of new statistical methods is related to the low efficiency of existing algorithms, especially when the underlying sample is small or the underlying distribution is close to Gaussian. In this paper we propose a new estimation algorithm for stability index, for samples from the symmetric $\alpha$-stable distribution. The proposed approach is based on quantile conditional variance ratio. We study the statistical properties of the proposed estimation procedure and show empirically that our methodology often outperforms other commonly used estimation algorithms. Moreover, we show that our statistic extracts unique sample characteristics that can be combined with other methods to refine existing methodologies via ensamble methods. Although our focus is set on the symmetric $\alpha$-stable case, we demonstrate that the considered statistic is insensitive to the skewness parameter change, so that our method could be also used in a more generic framework. For completeness, we also show how to apply our method on real data linked to plasma physics.

stat.ME

A note on the equivalence between the conditional uncorrelation and the independence of random variables

It is well known that while the independence of random variables implies zero correlation, the opposite is not true. Namely, uncorrelated random variables are not necessarily independent. In this note we show that the implication could be reversed if we consider the localised version of the correlation coefficient. More specifically, we show that if random variables are conditionally (locally) uncorrelated for any quantile conditioning sets, then they are independent. For simplicity, we focus on the absolutely continuous case. Also, we illustrate potential usefulness of the stated result using two simple examples.

math.ST

Estimating value at risk: LSTM vs. GARCH

Estimating value-at-risk on time series data with possibly heteroscedastic dynamics is a highly challenging task. Typically, we face a small data problem in combination with a high degree of non-linearity, causing difficulties for both classical and machine-learning estimation algorithms. In this paper, we propose a novel value-at-risk estimator using a long short-term memory (LSTM) neural network and compare its performance to benchmark GARCH estimators. Our results indicate that even for a relatively short time series, the LSTM could be used to refine or monitor risk estimation processes and correctly identify the underlying risk dynamics in a non-parametric fashion. We evaluate the estimator on both simulated and market data with a focus on heteroscedasticity, finding that LSTM exhibits a similar performance to GARCH estimators on simulated data, whereas on real market data it is more sensitive towards increasing or decreasing volatility and outperforms all existing estimators of value-at-risk in terms of exception rate and mean quantile score.

q-fin.RM

Estimating and backtesting risk under heavy tails

While the {estimation} of risk is an important question in the daily business of banking and insurance, many existing plug-in estimation procedures suffer from an unnecessary bias. This often leads to the underestimation of risk and negatively impacts backtesting results, especially in small sample cases. In this article we show that the link between estimation bias and backtesting can be traced back to the dual relationship between risk measures and the corresponding performance measures, and discuss this in reference to value-at-risk, expected shortfall and expectile value-at-risk. Motivated by the consistent underestimation of risk by plug-in procedures, we propose a new algorithm for bias correction and show how to apply it for generalized Pareto distributions to the i.i.d. setting and to a GARCH(1,1) time series. In particular, we show that the application of our algorithm leads to gain in efficiency when heavy tails or heteroscedasticity exists in the data.

q-fin.RM

Estimating and backtesting risk under heavy tails

While the estimation of risk is an important question in the daily business of banking and insurance, many existing plug-in estimation procedures suffer from an unnecessary bias. This often leads to the underestimation of risk and negatively impacts backtesting results, especially in small sample cases. In this article we show that the link between estimation bias and backtesting can be traced back to the dual relationship between risk measures and the corresponding performance measures, and discuss this in reference to value-at-risk, expected shortfall and expectile value-at-risk. Motivated by the consistent underestimation of risk by plug-in procedures, we propose a new algorithm for bias correction and show how to apply it for generalized Pareto distributions to the i.i.d.\ setting and to a GARCH(1,1) time series. In particular, we show that the application of our algorithm leads to gain in efficiency when heavy tails or heteroscedasticity exists in the data.

q-fin.RM

Discrete-time risk sensitive portfolio optimization with proportional transaction costs

In this paper we consider a discrete-time risk sensitive portfolio optimization over a long time horizon with proportional transaction costs. We show that within the log-return i.i.d. framework the solution to a suitable Bellman equation exists under minimal assumptions and can be used to characterize the optimal strategies for both risk-averse and risk-seeking cases. Moreover, using numerical examples, we show how a Bellman equation analysis can be used to construct or refine optimal trading strategies in the presence of transaction costs.

q-fin.PM