SearcharxivSearch

arXiv subjects

Mehmet Caner

Publications and source records attributed to Mehmet Caner.

14 recordsLinked to original sources

Designing Agentic AI-Based Screening for Portfolio Investment

We introduce a new agentic artificial intelligence (AI) platform for portfolio management. Our architecture consists of three layers. First, two large language model (LLM) agents are assigned specialized tasks: one agent screens for firms with desirable fundamentals, while a sentiment analysis agent screens for firms with desirable news. Second, these agents deliberate to generate and agree upon buy and sell signals from a large portfolio, substantially narrowing the pool of candidate assets. Finally, we apply a high-dimensional precision matrix estimation procedure to determine optimal portfolio weights. We show, through information acquisition theory, that screening with agentic AI can bring utility gains in screening compared with humans. We introduce the concept of \emph{sensible screening} and establish that, under mild screening errors, the squared Sharpe ratio of the screened portfolio consistently estimates its target. Empirically, our method achieves superior Sharpe ratios relative to an unscreened baseline portfolio and to conventional screening approaches, evaluated on S\&P~500 data over both short and medium terms.

q-fin.PM

Portfolio Analysis in High Dimensions with TE and Weight Constraints

This paper explores the statistical properties of forming constrained optimal portfolios within a high-dimensional set of assets. We examine portfolios with tracking error constraints, those with simultaneous tracking error and weight restrictions, and portfolios constrained solely by weight. Tracking error measures portfolio performance against a benchmark (typically an index), while weight constraints determine asset allocation based on regulatory requirements or fund prospectuses. Our approach employs a novel statistical learning technique that integrates factor models with nodewise regression, named the Constrained Residual Nodewise Optimal Weight Regression (CROWN) method. We demonstrate its estimation consistency in large dimensions, even when assets outnumber the portfolio's time span. Convergence rate results for constrained portfolio weights, risk, and Sharpe Ratio are provided, and simulation and empirical evidence highlight the method's outstanding performance.

q-fin.PM

Deep Learning Based Residuals in Non-linear Factor Models: Precision Matrix Estimation of Returns with Low Signal-to-Noise Ratio

This paper introduces a consistent estimator and rate of convergence for the precision matrix of asset returns in large portfolios using a non-linear factor model within the deep learning framework. Our estimator remains valid even in low signal-to-noise ratio environments typical for financial markets and is compatible with weak factors. Our theoretical analysis establishes uniform bounds on expected estimation risk based on deep neural networks for an expanding number of assets. Additionally, we provide a new consistent data-dependent estimator of error covariance in deep neural networks. Our models demonstrate superior accuracy in extensive simulations and the empirics.

stat.ML

Sharpe Ratio Analysis in High Dimensions: Residual-Based Nodewise Regression in Factor Models

We provide a new theory for nodewise regression when the residuals from a fitted factor model are used. We apply our results to the analysis of the consistency of Sharpe ratio estimators when there are many assets in a portfolio. We allow for an increasing number of assets as well as time observations of the portfolio. Since the nodewise regression is not feasible due to the unknown nature of idiosyncratic errors, we provide a feasible-residual-based nodewise regression to estimate the precision matrix of errors which is consistent even when number of assets, p, exceeds the time span of the portfolio, n. In another new development, we also show that the precision matrix of returns can be estimated consistently, even with an increasing number of factors and p>n. We show that: (1) with p>n, the Sharpe ratio estimators are consistent in global minimum-variance and mean-variance portfolios; and (2) with p>n, the maximum Sharpe ratio estimator is consistent when the portfolio weights sum to one; and (3) with p<<n, the maximum-out-of-sample Sharpe ratio estimator is consistent.

q-fin.PM

Shoiuld Humans Lie to Machines: The Incentive Compatibility of Lasso and General Weighted Lasso

We consider situations where a user feeds her attributes to a machine learning method that tries to predict her best option based on a random sample of other users. The predictor is incentive-compatible if the user has no incentive to misreport her covariates. Focusing on the popular Lasso estimation technique, we borrow tools from high-dimensional statistics to characterize sufficient conditions that ensure that Lasso is incentive compatible in large samples. We extend our results to the Conservative Lasso estimator and provide new moment bounds for this generalized weighted version of Lasso. Our results show that incentive compatibility is achieved if the tuning parameter is kept above some threshold. We present simulations that illustrate how this can be done in practice.

econ.EM

Generalized Linear Models with Structured Sparsity Estimators

In this paper, we introduce structured sparsity estimators in Generalized Linear Models. Structured sparsity estimators in the least squares loss are introduced by Stucky and van de Geer (2018) recently for fixed design and normal errors. We extend their results to debiased structured sparsity estimators with Generalized Linear Model based loss. Structured sparsity estimation means penalized loss functions with a possible sparsity structure used in the chosen norm. These include weighted group lasso, lasso and norms generated from convex cones. The significant difficulty is that it is not clear how to prove two oracle inequalities. The first one is for the initial penalized Generalized Linear Model estimator. Since it is not clear how a particular feasible-weighted nodewise regression may fit in an oracle inequality for penalized Generalized Linear Model, we need a second oracle inequality to get oracle bounds for the approximate inverse for the sample estimate of second-order partial derivative of Generalized Linear Model. Our contributions are fivefold: 1. We generalize the existing oracle inequality results in penalized Generalized Linear Models by proving the underlying conditions rather than assuming them. One of the key issues is the proof of a sample one-point margin condition and its use in an oracle inequality. 2. Our results cover even non sub-Gaussian errors and regressors. 3. We provide a feasible weighted nodewise regression proof which generalizes the results in the literature from a simple l_1 norm usage to norms generated from convex cones. 4. We realize that norms used in feasible nodewise regression proofs should be weaker or equal to the norms in penalized Generalized Linear Model loss. 5. We can debias the first step estimator via getting an approximate inverse of the singular-sample second order partial derivative of Generalized Linear Model loss.

stat.ML

An Upper Bound for Functions of Estimators in High Dimensions

We provide an upper bound as a random variable for the functions of estimators in high dimensions. This upper bound may help establish the rate of convergence of functions in high dimensions. The upper bound random variable may converge faster, slower, or at the same rate as estimators depending on the behavior of the partial derivative of the function. We illustrate this via three examples. The first two examples use the upper bound for testing in high dimensions, and third example derives the estimated out-of-sample variance of large portfolios. All our results allow for a larger number of parameters, p, than the sample size, n.

econ.EM

A Nodewise Regression Approach to Estimating Large Portfolios

This paper investigates the large sample properties of the variance, weights, and risk of high-dimensional portfolios where the inverse of the covariance matrix of excess asset returns is estimated using a technique called nodewise regression. Nodewise regression provides a direct estimator for the inverse covariance matrix using the Least Absolute Shrinkage and Selection Operator (Lasso) of Tibshirani (1994) to estimate the entries of a sparse precision matrix. We show that the variance, weights, and risk of the global minimum variance portfolios and the Markowitz mean-variance portfolios are consistently estimated with more assets than observations. We show, empirically, that the nodewise regression-based approach performs well in comparison to factor models and shrinkage methods.

math.ST

High Dimensional Linear GMM

This paper proposes a desparsified GMM estimator for estimating high-dimensional regression models allowing for, but not requiring, many more endogenous regressors than observations. We provide finite sample upper bounds on the estimation error of our estimator and show how asymptotically uniformly valid inference can be conducted in the presence of conditionally heteroskedastic error terms. We do not require the projection of the endogenous variables onto the linear span of the instruments to be sparse; that is we do not impose the instruments to be sparse for our inferential procedure to be asymptotically valid. Furthermore, the variables of the model are not required to be sub-gaussian and we also explain how our results carry over to the classic linear dynamic panel data model. Simulations show that our estimator has a low mean square error and does well in terms of size and power of the tests constructed based on the estimator.

math.ST

Inference in partially identified models with many moment inequalities using Lasso

This paper considers inference in a partially identified moment (in)equality model with many moment inequalities. We propose a novel two-step inference procedure that combines the methods proposed by Chernozhukov, Chetverikov and Kato (2018a) (CCK18, hereafter) with a first step moment inequality selection based on the Lasso. Our method controls asymptotic size uniformly, both in underlying parameter and data distribution. Also, the power of our method compares favorably with that of the corresponding two-step method in CCK18 for large parts of the parameter space, both in theory and in simulations. Finally, we show that our Lasso-based first step can be implemented by thresholding standardized sample averages, and so it is straightforward to implement.

math.ST

Asymptotically Honest Confidence Regions for High Dimensional Parameters by the Desparsified Conservative Lasso

In this paper we consider the conservative Lasso which we argue penalizes more correctly than the Lasso and show how it may be desparsified in the sense of van de Geer et al. (2014) in order to construct asymptotically honest (uniform) confidence bands. In particular, we develop an oracle inequality for the conservative Lasso only assuming the existence of a certain number of moments. This is done by means of the Marcinkiewicz-Zygmund inequality. We allow for heteroskedastic non-subgaussian error terms and covariates. Next, we desparsify the conservative Lasso estimator and derive the asymptotic distribution of tests involving an increasing number of parameters. Our simulations reveal that the desparsified conservative Lasso estimates the parameters more precisely than the desparsified Lasso, has better size properties and produces confidence bands with superior coverage rates.

math.ST

Delta Theorem in the Age of High Dimensions

We provide a new version of delta theorem, that takes into account of high dimensional parameter estimation. We show that depending on the structure of the function, the limits of functions of estimators have faster or slower rate of convergence than the limits of estimators. We illustrate this via two examples. First, we use it for testing in high dimensions, and second in estimating large portfolio risk. Our theorem works in the case of larger number of parameters, $p$, than the sample size, $n$: $p>n$.

math.ST

Sharp Threshold Detection Based on Sup-norm Error rates in High-dimensional Models

We propose a new estimator, the thresholded scaled Lasso, in high dimensional threshold regressions. First, we establish an upper bound on the $\ell_\infty$ estimation error of the scaled Lasso estimator of Lee et al. (2012). This is a non-trivial task as the literature on high-dimensional models has focused almost exclusively on $\ell_1$ and $\ell_2$ estimation errors. We show that this sup-norm bound can be used to distinguish between zero and non-zero coefficients at a much finer scale than would have been possible using classical oracle inequalities. Thus, our sup-norm bound is tailored to consistent variable selection via thresholding. Our simulations show that thresholding the scaled Lasso yields substantial improvements in terms of variable selection. Finally, we use our estimator to shed further empirical light on the long running debate on the relationship between the level of debt (public and private) and GDP growth.

stat.ME

Oracle Inequalities for Convex Loss Functions with Non-Linear Targets

This paper consider penalized empirical loss minimization of convex loss functions with unknown non-linear target functions. Using the elastic net penalty we establish a finite sample oracle inequality which bounds the loss of our estimator from above with high probability. If the unknown target is linear this inequality also provides an upper bound of the estimation error of the estimated parameter vector. These are new results and they generalize the econometrics and statistics literature. Next, we use the non-asymptotic results to show that the excess loss of our estimator is asymptotically of the same order as that of the oracle. If the target is linear we give sufficient conditions for consistency of the estimated parameter vector. Next, we briefly discuss how a thresholded version of our estimator can be used to perform consistent variable selection. We give two examples of loss functions covered by our framework and show how penalized nonparametric series estimation is contained as a special case and provide a finite sample upper bound on the mean square error of the elastic net series estimator.

math.ST