Searcharxiv⌕ Search

arXiv subjects

Colin M. Van Oort

Publications and source records attributed to Colin M. Van Oort.

8 recordsLinked to original sources

Testing replication for an agent-based model of market fragmentation and latency arbitrage

This study strengthens the foundations of multi-venue market modeling by attempting an independent replication of Wah and Wellman's 2016 model of latency arbitrage in a fragmented market. We find that faithful replication is hindered by missing implementation details in the original paper and limited quantitative reporting. We demonstrate that increasing the number of simulation runs beyond the original design allows for the creation of bootstrap confidence intervals to support rigorous tests of quantitative alignment, compensating for lacking distributional information (e.g. variance). We also demonstrate that increased complexity across the modeled scenarios corresponds with increased difficulty aligning to the original results. We draw on a codebase released by the original authors in connection with a later paper to recover additional implementation details; however, we reject quantitative alignment between that codebase and the published results. Combining information from the paper and the released code, we achieve relational equivalence for most metrics but reject quantitative alignment for model settings where latency is non-zero. We show that many of the qualitative takeaways from the original paper on the effects of market fragmentation and latency arbitrage are sensitive to the specifics of a `greedy strategy' extension given to the zero-intelligence (ZI) trader agents. Under an alternative interpretation of this strategy, we find that market fragmentation decreases execution times in all experiments and increases trader welfare in most experiments. Finally, to facilitate future replication, critique, and extension, we provide an ODD (Overview, Design concepts, Details) protocol for our implementations of the model.

q-fin.TR↗

Revisiting Cont's Stylized Facts for Modern Stock Markets

In 2001, Rama Cont introduced a now-widely used set of 'stylized facts' to synthesize empirical studies of financial price changes (returns), resulting in 11 statistical properties common to a large set of assets and markets. These properties are viewed as constraints a model should be able to reproduce in order to accurately represent returns in a market. It has not been established whether the characteristics Cont noted in 2001 still hold for modern markets following significant regulatory shifts and technological advances. It is also not clear whether a given time series of financial returns for an asset will express all 11 stylized facts. We test both of these propositions by attempting to replicate each of Cont's 11 stylized facts for intraday returns of the individual stocks in the Dow 30, using the same authoritative data as that used by the U.S. regulator from October 2018 - March 2019. We find conclusive evidence for eight of Cont's original facts and no support for the remaining three. Our study represents the first test of Cont's 11 stylized facts against a consistent set of stocks, therefore providing insight into how these stylized facts should be viewed in the context of modern stock markets.

q-fin.ST↗

Adaptive Agents and Data Quality in Agent-Based Financial Markets

We present our Agent-Based Market Microstructure Simulation (ABMMS), an Agent-Based Financial Market (ABFM) that captures much of the complexity present in the US National Market System for equities (NMS). Agent-Based models are a natural choice for understanding financial markets. Financial markets feature a constrained action space that should simplify model creation, produce a wealth of data that should aid model validation, and a successful ABFM could strongly impact system design and policy development processes. Despite these advantages, ABFMs have largely remained an academic novelty. We hypothesize that two factors limit the usefulness of ABFMs. First, many ABFMs fail to capture relevant microstructure mechanisms, leading to differences in the mechanics of trading. Second, the simple agents that commonly populate ABFMs do not display the breadth of behaviors observed in human traders or the trading systems that they create. We investigate these issues through the development of ABMMS, which features a fragmented market structure, communication infrastructure with propagation delays, realistic auction mechanisms, and more. As a baseline, we populate ABMMS with simple trading agents and investigate properties of the generated data. We then compare the baseline with experimental conditions that explore the impacts of market topology or meta-reinforcement learning agents. The combination of detailed market mechanisms and adaptive agents leads to models whose generated data more accurately reproduce stylized facts observed in actual markets. These improvements increase the utility of ABFMs as tools to inform design and policy decisions.

q-fin.TR↗

Augmenting semantic lexicons using word embeddings and transfer learning

Sentiment-aware intelligent systems are essential to a wide array of applications. These systems are driven by language models which broadly fall into two paradigms: Lexicon-based and contextual. Although recent contextual models are increasingly dominant, we still see demand for lexicon-based models because of their interpretability and ease of use. For example, lexicon-based models allow researchers to readily determine which words and phrases contribute most to a change in measured sentiment. A challenge for any lexicon-based approach is that the lexicon needs to be routinely expanded with new words and expressions. Here, we propose two models for automatic lexicon expansion. Our first model establishes a baseline employing a simple and shallow neural network initialized with pre-trained word embeddings using a non-contextual approach. Our second model improves upon our baseline, featuring a deep Transformer-based network that brings to bear word definitions to estimate their lexical polarity. Our evaluation shows that both models are able to score new words with a similar accuracy to reviewers from Amazon Mechanical Turk, but at a fraction of the cost.

cs.CL↗

Scaling of inefficiencies in the U.S. equity markets: Evidence from three market indices and more than 2900 securities

Using the most comprehensive, commercially-available dataset of trading activity in U.S. equity markets, we catalog and analyze quote dislocations between the SIP National Best Bid and Offer (NBBO) and a synthetic BBO constructed from direct feeds. We observe a total of over 3.1 billion dislocation segments in the Russell 3000 during trading in 2016, roughly 525 per second of trading. However, these dislocations do not occur uniformly throughout the trading day. We identify a characteristic structure that features more dislocations near the open and close. Additionally, around 23% of observed trades executed during dislocations. These trades may have been impacted by stale information, leading to estimated opportunity costs on the order of $ 2 billion USD. A subset of the constituents of the S&P 500 index experience the greatest amount of opportunity cost and appear to drive inefficiencies in other stocks. These results quantify impacts of the physical structure of the U.S. National Market System.

q-fin.TR↗

Fragmentation and inefficiencies in US equity markets: Evidence from the Dow 30

Using the most comprehensive source of commercially available data on the US National Market System, we analyze all quotes and trades associated with Dow 30 stocks in 2016 from the vantage point of a single and fixed frame of reference. We find that inefficiencies created in part by the fragmentation of the equity marketplace are relatively common and persist for longer than what physical constraints may suggest. Information feeds reported different prices for the same equity more than 120 million times, with almost 64 million dislocation segments featuring meaningfully longer duration and higher magnitude. During this period, roughly 22% of all trades occurred while the SIP and aggregated direct feeds were dislocated. The current market configuration resulted in a realized opportunity cost totaling over $160 million when compared with a single feed, single exchange alternative---a conservative estimate that does not take into account intra-day offsetting events.

q-fin.TR↗

CASI: A Convolutional Neural Network Approach for Shell Identification

We utilize techniques from deep learning to identify signatures of stellar feedback in simulated molecular clouds. Specifically, we implement a deep neural network with an architecture similar to U-Net and apply it to the problem of identifying wind-driven shells and bubbles using data from magneto-hydrodynamic simulations of turbulent molecular clouds with embedded stellar sources. The network is applied to two tasks, dense regression and segmentation, on two varieties of data, simulated density and synthetic 12 CO observations. Our Convolutional Approach for Shell Identification (CASI) is able to obtain a true positive rate greater than 90\%, while maintaining a false positive rate of 1\%, on two segmentation tasks and also performs well on related regression tasks. The source code for CASI is available on GitLab.

astro-ph.IM↗

Simon's fundamental rich-get-richer model entails a dominant first-mover advantage

Herbert Simon's classic rich-get-richer model is one of the simplest empirically supported mechanisms capable of generating heavy-tail size distributions for complex systems. Simon argued analytically that a population of flavored elements growing by either adding a novel element or randomly replicating an existing one would afford a distribution of group sizes with a power-law tail. Here, we show that, in fact, Simon's model does not produce a simple power law size distribution as the initial element has a dominant first-mover advantage, and will be overrepresented by a factor proportional to the inverse of the innovation probability. The first group's size discrepancy cannot be explained away as a transient of the model, and may therefore be many orders of magnitude greater than expected. We demonstrate how Simon's analysis was correct but incomplete, and expand our alternate analysis to quantify the variability of long term rankings for all groups. We find that the expected time for a first replication is infinite, and show how an incipient group must break the mechanism to improve their odds of success. We present an example of citation counts for a specific field that demonstrates a first-mover advantage consistent with our revised view of the rich-get-richer mechanism. Our findings call for a reexamination of preceding work invoking Simon's model and provide an expanded understanding going forward.

physics.soc-ph↗