SearcharxivSearch

arXiv · 2107.05201

Deep Risk Model: A Deep Learning Solution for Mining Latent Risk Factors to Improve Covariance Matrix Estimation

Abstract

Modeling and managing portfolio risk is perhaps the most important step to achieve growing and preserving investment performance. Within the modern portfolio construction framework that built on Markowitz's theory, the covariance matrix of stock returns is a required input to calculate portfolio risk. Traditional approaches to estimate the covariance matrix are based on human-designed risk factors, which often require tremendous time and effort to design better risk factors to improve the covariance estimation. In this work, we formulate the quest of mining risk factors as a learning problem and propose a deep learning solution to effectively ``design'' risk factors with neural networks. The learning objective is also carefully set to ensure the learned risk factors are effective in explaining the variance of stock returns as well as having desired orthogonality and stability. Our experiments on the stock market data demonstrate the effectiveness of the proposed solution: our method can obtain $1.9\%$ higher explained variance measured by $R^2$ and also reduce the risk of a global minimum variance portfolio. The incremental analysis further supports our design of both the architecture and the learning objective.

Explore related subjects

Keep this discovery

BibTeXRIS

Hengxu Lin, Dong Zhou, Weiqing Liu, Jiang Bian. 2021-07-12. Deep Risk Model: A Deep Learning Solution for Mining Latent Risk Factors to Improve Covariance Matrix Estimation. https://arxiv.org/abs/2107.05201

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Simplifying Cyber Cat(astrophe)s with Cyber Kittens: Power Law Plausibility for Cyber Insurance Risks

Cyber insurance requires accurate modeling of worst-case catastrophic (cat) events, but the field lacks robust quantitative approaches for estimating upper-bound losses. Building on a recent dataset of 24 cyber cat events over 30 years, this work tests whether cyber economic losses follow a power law distribution. We analyze "cyber kittens" - sub-1B USD events distinguished from cat events (1B+ USD) only by magnitude - extracted via LLM from cyber insurance claims data (2020-2024). Using victim count (weighted by claim year) as a proxy for economic loss, we link kitten-sized events to known cat events to estimate losses. The kitten distribution proved consistent with the cat dataset, and power laws were statistically plausible: each order-of-magnitude increase in event size corresponds to a 5-7x drop in probability. Extrapolating, an event 100x the largest 2020-2024 cat event is expected roughly every 206 years, translating to 100-250B USD in losses - catastrophic, but not extraordinary relative to other insurance lines.

q-fin.RM

Pricing the DeFi Tail: Do Protocols or Depositors Price Operational Risk?

Similar to banks, DeFi protocols expose depositors to operational risk (USD 9.45 billion across 1,075 events since 2020). Unlike banks, they are not required to hold capital against it. A protocol may maintain a buffer voluntarily. Absent one, the risk falls on the depositor, who should then demand a risk premium in the supply yield. I quantify the underlying tail on one benchmark, a per-sector Basel loss-distribution approach fitted to a new operational risk event dataset, and test both margins against it. Tails in the four core sectors are no heavier than the Moscadelli banking band $[0.85, 1.39]$. Bridge, Derivatives, and the residual Other sector exhibit cyber-loss-level tails ($\hat\xi \approx 1.6$), with point estimates past the infinite-mean boundary. The Lending tail implies a $\mathrm{VaR}_{99.9}$ capital buffer of 18% of TVL and of the ten largest Lending venues, the four holding a buffer cover on average 5% of it. Under market discipline, depositors should demand a higher yield in compensation where a venue does not maintain a buffer. I find that venues without a buffer pay a higher premium than those with (a 125-bps gap in medians): evidence the market discriminates in the right direction. However, the premium falls far short of an adequately priced tail. This unpriced tail falls disproportionately on the retail depositor, who sees only the posted rate but lacks the information and skills to price it. Because these products are not bank-regulated, I recommend disclosure over capital mandates: protocols, and any service providers that front access to it, should publish standardized losses, existing capital buffers and tail coverage.

q-fin.RM

Illiquidity at Risk

Market efficiency relies fundamentally on stable liquidity. Consequently, forecasting liquidity dynamics is a priority for both investors and regulators. We introduce a new tail-risk metric, Illiquidity-at-Risk (IlliQaR), designed to quantify the magnitude of extreme liquidity dry-ups. Relying upon the realized Amihud (a precise illiquidity measurement derived from high-frequency data as the ratio of realized volatility to trading volume) we assess the predictive power of various linear and non-linear econometric models, with a specific focus on the impact of discontinuous jump components. Accounting for these jumps is essential for achieving accurate probability coverage and better IlliQaR predictions during periods of systemic stress, where standard continuous models systematically underestimate the severity of liquidity evaporation. Our empirical analysis, encompassing the S&P 500 index and a cross-section of 25 large U.S. equities, demonstrates that incorporating jumps significantly improves forecasts of illiquidity. Our results suggest that individual stock IlliQaR violations often cluster during periods of S&P 500 liquidity stress. This indicates that Illiquidity at Risk is not just a localized concern but a systemic one, where the main index acts as a leading indicator for extreme dry-ups in individual stock liquidity.

q-fin.RM