SearcharxivSearch

arXiv subjects

Willie Yu

Publications and source records attributed to Willie Yu.

At least 19 recordsLinked to original sources

ETF Risk Models

We discuss how to build ETF risk models. Our approach anchors on i) first building a multilevel (non-)binary classification/taxonomy for ETFs, which is utilized in order to define the risk factors, and ii) then building the risk models based on these risk factors by utilizing the heterotic risk model construction of https://ssrn.com/abstract=2600798 (for binary classifications) or general risk model construction of https://ssrn.com/abstract=2722093 (for non-binary classifications). We discuss how to build an ETF taxonomy using ETF constituent data. A multilevel ETF taxonomy can also be constructed by appropriately augmenting and expanding well-built and granular third-party single-level ETF groupings.

q-fin.RM

Machine Learning Treasury Yields

We give explicit algorithms and source code for extracting factors underlying Treasury yields using (unsupervised) machine learning (ML) techniques, such as nonnegative matrix factorization (NMF) and (statistically deterministic) clustering. NMF is a popular ML algorithm (used in computer vision, bioinformatics/computational biology, document classification, etc.), but is often misconstrued and misused. We discuss how to properly apply NMF to Treasury yields. We analyze the factors based on NMF and clustering and their interpretation. We discuss their implications for forecasting Treasury yields in the context of out-of-sample ML stability issues.

stat.ME

iCurrency?

We discuss the idea of a purely algorithmic universal world iCurrency set forth in [Kakushadze and Liew, 2014] (https://ssrn.com/abstract=2542541) and expanded in [Kakushadze and Liew, 2017] (https://ssrn.com/abstract=3059330) in light of recent developments, including Libra. Is Libra a contender to become iCurrency? Among other things, we analyze the Libra proposal, including the stability and volatility aspects, and discuss various issues that must be addressed. For instance, one cannot expect a cryptocurrency such as Libra to trade in a narrow band without a robust monetary policy. The presentation in the main text of the paper is intentionally nontechnical. It is followed by an extensive appendix with a mathematical description of the dynamics of (crypto)currency exchange rates in target zones, mechanisms for keeping the exchange rate from breaching the band, the role of volatility, etc.

q-fin.GN

Machine Learning Risk Models

We give an explicit algorithm and source code for constructing risk models based on machine learning techniques. The resultant covariance matrices are not factor models. Based on empirical backtests, we compare the performance of these machine learning risk models to other constructions, including statistical risk models, risk models based on fundamental industry classifications, and also those utilizing multilevel clustering based industry classifications.

q-fin.PM

Altcoin-Bitcoin Arbitrage

We give an algorithm and source code for a cryptoasset statistical arbitrage alpha based on a mean-reversion effect driven by the leading momentum factor in cryptoasset returns discussed in https://ssrn.com/abstract=3245641. Using empirical data, we identify the cross-section of cryptoassets for which this altcoin-Bitcoin arbitrage alpha is significant and discuss it in the context of liquidity considerations as well as its implications for cryptoasset trading.

q-fin.PM

Betas, Benchmarks and Beating the Market

We give an explicit formulaic algorithm and source code for building long-only benchmark portfolios and then using these benchmarks in long-only market outperformance strategies. The benchmarks (or the corresponding betas) do not involve any principal components, nor do they require iterations. Instead, we use a multifactor risk model (which utilizes multilevel industry classification or clustering) specifically tailored to long-only benchmark portfolios to compute their weights, which are explicitly positive in our construction.

q-fin.PM

Stock Market Visualization

We provide complete source code for a front-end GUI and its back-end counterpart for a stock market visualization tool. It is built based on the "functional visualization" concept we discuss, whereby functionality is not sacrificed for fancy graphics. The GUI, among other things, displays a color-coded signal (computed by the back-end code) based on how "out-of-whack" each stock is trading compared with its peers ("mean-reversion"), and the most sizable changes in the signal ("momentum"). The GUI also allows to efficiently filter/tier stocks by various parameters (e.g., sector, exchange, signal, liquidity, market cap) and functionally display them. The tool can be run as a web-based or local application.

q-fin.PM

Notes on Fano Ratio and Portfolio Optimization

We discuss - in what is intended to be a pedagogical fashion - generalized "mean-to-risk" ratios for portfolio optimization. The Sharpe ratio is only one example of such generalized "mean-to-risk" ratios. Another example is what we term the Fano ratio (which, unlike the Sharpe ratio, is independent of the time horizon). Thus, for long-only portfolios optimizing the Fano ratio generally results in a more diversified and less skewed portfolio (compared with optimizing the Sharpe ratio). We give an explicit algorithm for such optimization. We also discuss (Fano-ratio-inspired) long-short strategies that outperform those based on optimizing the Sharpe ratio in our backtests.

q-fin.PM

Dead Alphas as Risk Factors

We give an explicit algorithm and source code for extracting equity risk factors from dead (a.k.a. "flatlined" or "hockey-stick") alphas and using them to improve performance characteristics of good (tradable) alphas. In a nutshell, we use dead alphas to extract directions in the space of stock returns along which there is no money to be made (and/or those bets are too volatile). In practice the number of dead alphas can be large compared with the number of underlying stocks and care is required in identifying the aforesaid directions.

q-fin.PM

Estimating Cost Savings from Early Cancer Diagnosis

We estimate treatment cost-savings from early cancer diagnosis. For breast, lung, prostate and colorectal cancers and melanoma, which account for more than 50% of new incidences projected in 2017, we combine published cancer treatment cost estimates by stage with incidence rates by stage at diagnosis. We extrapolate to other cancer sites by using estimated national expenditures and incidence rates. A rough estimate for the U.S. national annual treatment cost-savings from early cancer diagnosis is in 11 digits. Using this estimate and cost-neutrality, we also estimate a rough upper bound on the cost of a routine early cancer screening test.

q-bio.TO

Decoding Stock Market with Quant Alphas

We give an explicit algorithm and source code for extracting expected returns for stocks from expected returns for alphas. Our algorithm altogether bypasses combining alphas with weights into "alpha combos". Simply put, we have developed a new method for trading alphas which does not involve combining them. This yields substantial cost savings as alpha combos cost hedge funds around 3% of the P&L, while alphas themselves cost around 10%. Also, the extra layer of alpha combos, which our new method avoids, adds noise and suboptimality. We also arrive at our algorithm independently by explicitly constructing alpha risk models based on position data.

q-fin.PM

Mutation Clusters from Cancer Exome

We apply our statistically deterministic machine learning/clustering algorithm *K-means (recently developed in https://ssrn.com/abstract=2908286) to 10,656 published exome samples for 32 cancer types. A majority of cancer types exhibit mutation clustering structure. Our results are in-sample stable. They are also out-of-sample stable when applied to 1,389 published genome samples across 14 cancer types. In contrast, we find in- and out-of-sample instabilities in cancer signatures extracted from exome samples via nonnegative matrix factorization (NMF), a computationally costly and non-deterministic method. Extracting stable mutation structures from exome data could have important implications for speed and cost, which are critical for early-stage cancer diagnostics such as novel blood-test methods currently in development.

q-bio.GN

Open Source Fundamental Industry Classification

We provide complete source code for building a fundamental industry classification based on publically available and freely downloadable data. We compare various fundamental industry classifications by running a horserace of short-horizon trading signals (alphas) utilizing open source heterotic risk models (https://ssrn.com/abstract=2600798) built using such industry classifications. Our source code includes various stand-alone and portable modules, e.g., for downloading/parsing web data, etc.

q-fin.GN

*K-means and Cluster Models for Cancer Signatures

We present *K-means clustering algorithm and source code by expanding statistical clustering methods applied in https://ssrn.com/abstract=2802753 to quantitative finance. *K-means is statistically deterministic without specifying initial centers, etc. We apply *K-means to extracting cancer signatures from genome data without using nonnegative matrix factorization (NMF). *K-means' computational cost is a fraction of NMF's. Using 1,389 published samples for 14 cancer types, we find that 3 cancers (liver cancer, lung cancer and renal cell carcinoma) stand out and do not have cluster-like structures. Two clusters have especially high within-cluster correlations with 11 other cancers indicating common underlying structures. Our approach opens a novel avenue for studying such structures. *K-means is universal and can be applied in other fields. We discuss some potential applications in quantitative finance.

q-bio.GN

Statistical Industry Classification

We give complete algorithms and source code for constructing (multilevel) statistical industry classifications, including methods for fixing the number of clusters at each level (and the number of levels). Under the hood there are clustering algorithms (e.g., k-means). However, what should we cluster? Correlations? Returns? The answer turns out to be neither and our backtests suggest that these details make a sizable difference. We also give an algorithm and source code for building "hybrid" industry classifications by improving off-the-shelf "fundamental" industry classifications by applying our statistical industry classification methods to them. The presentation is intended to be pedagogical and geared toward practical applications in quantitative trading.

q-fin.PM

Factor Models for Cancer Signatures

We present a novel method for extracting cancer signatures by applying statistical risk models (http://ssrn.com/abstract=2732453) from quantitative finance to cancer genome data. Using 1389 whole genome sequenced samples from 14 cancers, we identify an "overall" mode of somatic mutational noise. We give a prescription for factoring out this noise and source code for fixing the number of signatures. We apply nonnegative matrix factorization (NMF) to genome data aggregated by cancer subtype and filtered using our method. The resultant signatures have substantially lower variability than those from unfiltered data. Also, the computational cost of signature extraction is cut by about a factor of 10. We find 3 novel cancer signatures, including a liver cancer dominant signature (96% contribution) and a renal cell carcinoma signature (70% contribution). Our method accelerates finding new cancer signatures and improves their overall stability. Reciprocally, the methods for extracting cancer signatures could have interesting applications in quantitative finance.

q-bio.GN

How to Combine a Billion Alphas

We give an explicit algorithm and source code for computing optimal weights for combining a large number N of alphas. This algorithm does not cost O(N^3) or even O(N^2) operations but is much cheaper, in fact, the number of required operations scales linearly with N. We discuss how in the absence of binary or quasi-binary clustering of alphas, which is not observed in practice, the optimization problem simplifies when N is large. Our algorithm does not require computing principal components or inverting large matrices, nor does it require iterations. The number of risk factors it employs, which typically is limited by the number of historical observations, can be sizably enlarged via using position data for the underlying tradables.

q-fin.PM

Statistical Risk Models

We give complete algorithms and source code for constructing statistical risk models, including methods for fixing the number of risk factors. One such method is based on eRank (effective rank) and yields results similar to (and further validates) the method set forth in an earlier paper by one of us. We also give a complete algorithm and source code for computing eigenvectors and eigenvalues of a sample covariance matrix which requires i) no costly iterations and ii) the number of operations linear in the number of returns. The presentation is intended to be pedagogical and oriented toward practical applications.

q-fin.PM