SearcharxivSearch

arXiv subjects

Hai Qian

Publications and source records attributed to Hai Qian.

11 recordsLinked to original sources

Emergent invariance and scaling properties in the collective return dynamics of a stock market

Several works have observed heavy-tailed behavior in the distributions of returns in different markets, which are observable indicators of underlying complex dynamics. Such prior works study return distributions that are marginalized across the individual stocks in the market, and do not track statistics about the joint distributions of returns conditioned on different stocks, which would be useful for optimizing inter-stock asset allocation strategies. As a step towards this goal, we study emergent phenomena in the distributions of returns as captured by their pairwise correlations. In particular, we consider the pairwise (between stocks $i,j$) partial correlations of returns with respect to the market mode, $c_{i,j}(τ)$, (thus, correcting for the baseline return behavior of the market), over different time horizons ($τ$), and discover two novel emergent phenomena: (i) the standardized distributions of the $c_{i,j}(τ)$'s are observed to be invariant of $τ$ ranging from from $1000 \textrm{min}$ (2.5 days) to $30000 \textrm{min}$ (2.5 months); (ii) the scaling of the standard deviation of $c_{i,j}(τ)$'s with $τ$ admits \iffalse within this regime is empirically observed to \fi good fits to simple model classes such as a power-law $τ^{-λ}$ or stretched exponential function $e^{-τ^β}$ ($λ,β> 0$). Moreover, the parameters governing these fits provide a summary view of market health: for instance, in years marked by unprecedented financial crises -- for example $2008$ and $2020$ -- values of $λ$ (scaling exponent) are substantially lower. Finally, we demonstrate that the observed emergent behavior cannot be adequately supported by existing generative frameworks such as single- and multi-factor models. We introduce a promising agent-based Vicsek model that closes this gap.

cond-mat.stat-mech

Interpretable Learning-to-Rank with Generalized Additive Models

Interpretability of learning-to-rank models is a crucial yet relatively under-examined research area. Recent progress on interpretable ranking models largely focuses on generating post-hoc explanations for existing black-box ranking models, whereas the alternative option of building an intrinsically interpretable ranking model with transparent and self-explainable structure remains unexplored. Developing fully-understandable ranking models is necessary in some scenarios (e.g., due to legal or policy constraints) where post-hoc methods cannot provide sufficiently accurate explanations. In this paper, we lay the groundwork for intrinsically interpretable learning-to-rank by introducing generalized additive models (GAMs) into ranking tasks. Generalized additive models (GAMs) are intrinsically interpretable machine learning models and have been extensively studied on regression and classification tasks. We study how to extend GAMs into ranking models which can handle both item-level and list-level features and propose a novel formulation of ranking GAMs. To instantiate ranking GAMs, we employ neural networks instead of traditional splines or regression trees. We also show that our neural ranking GAMs can be distilled into a set of simple and compact piece-wise linear functions that are much more efficient to evaluate with little accuracy loss. We conduct experiments on three data sets and show that our proposed neural ranking GAMs can achieve significantly better performance than other traditional GAM baselines while maintaining similar interpretability.

cs.IR

Transfer of Machine Learning Fairness across Domains

If our models are used in new or unexpected cases, do we know if they will make fair predictions? Previously, researchers developed ways to debias a model for a single problem domain. However, this is often not how models are trained and used in practice. For example, labels and demographics (sensitive attributes) are often hard to observe, resulting in auxiliary or synthetic data to be used for training, and proxies of the sensitive attribute to be used for evaluation of fairness. A model trained for one setting may be picked up and used in many others, particularly as is common with pre-training and cloud APIs. Despite the pervasiveness of these complexities, remarkably little work in the fairness literature has theoretically examined these issues. We frame all of these settings as domain adaptation problems: how can we use what we have learned in a source domain to debias in a new target domain, without directly debiasing on the target domain as if it is a completely new problem? We offer new theoretical guarantees of improving fairness across domains, and offer a modeling approach to transfer to data-sparse target domains. We give empirical results validating the theory and showing that these modeling approaches can improve fairness metrics with less data.

cs.LG

Toward a better trade-off between performance and fairness with kernel-based distribution matching

As recent literature has demonstrated how classifiers often carry unintended biases toward some subgroups, deploying machine learned models to users demands careful consideration of the social consequences. How should we address this problem in a real-world system? How should we balance core performance and fairness metrics? In this paper, we introduce a MinDiff framework for regularizing classifiers toward different fairness metrics and analyze a technique with kernel-based statistical dependency tests. We run a thorough study on an academic dataset to compare the Pareto frontier achieved by different regularization approaches, and apply our kernel-based method to two large-scale industrial systems demonstrating real-world improvements.

cs.LG

Fairness in Recommendation Ranking through Pairwise Comparisons

Recommender systems are one of the most pervasive applications of machine learning in industry, with many services using them to match users to products or information. As such it is important to ask: what are the possible fairness risks, how can we quantify them, and how should we address them? In this paper we offer a set of novel metrics for evaluating algorithmic fairness concerns in recommender systems. In particular we show how measuring fairness based on pairwise comparisons from randomized experiments provides a tractable means to reason about fairness in rankings from recommender systems. Building on this metric, we offer a new regularizer to encourage improving this metric during model training and thus improve fairness in the resulting rankings. We apply this pairwise regularization to a large-scale, production recommender system and show that we are able to significantly improve the system's pairwise fairness.

cs.CY

Putting Fairness Principles into Practice: Challenges, Metrics, and Improvements

As more researchers have become aware of and passionate about algorithmic fairness, there has been an explosion in papers laying out new metrics, suggesting algorithms to address issues, and calling attention to issues in existing applications of machine learning. This research has greatly expanded our understanding of the concerns and challenges in deploying machine learning, but there has been much less work in seeing how the rubber meets the road. In this paper we provide a case-study on the application of fairness in machine learning research to a production classification system, and offer new insights in how to measure and address algorithmic fairness issues. We discuss open questions in implementing equality of opportunity and describe our fairness metric, conditional equality, that takes into account distributional differences. Further, we provide a new approach to improve on the fairness metric during model training and demonstrate its efficacy in improving performance for a real-world product

cs.LG

Growth of Order in An Anisotropic Swift-Hohenberg Model

We have studied the ordering kinetics of a two-dimensional anisotropic Swift-Hohenberg (SH) model numerically. The defect structure for this model is simpler than for the isotropic SH model. One finds only dislocations in the aligned ordering striped system. The motion of these point defects is strongly influenced by the anisotropic nature of the system. We developed accurate numerical methods for following the trajectories of dislocations. This allows us to carry out a detailed statistical analysis of the dynamics of the dislocations. The average speeds for the motion of the dislocations in the two orthogonal directions obey power laws in time with different amplitudes but the same exponents. The position and velocity distribution functions are only weakly anisotropic.

cond-mat.soft

The Vortex Kinetics of Conserved and Non-conserved O(n) Models

We study the motion of vortices in the conserved and non-conserved phase-ordering models. We give an analytical method for computing the speed and position distribution functions for pairs of annihilating point vortices based on heuristic scaling arguments. In the non-conserved case this method produces a speed distribution function consistent with previous analytic results. As two special examples, we simulate the conserved and non-conserved O(2) model in two dimensional space numerically. The numerical results for the non-conserved case are consistent with the theoretical predictions. The speed distribution of the vortices in the conserved case is measured for the first time. Our theory produces a distribution function with the correct large speed tail but does not accurately describe the numerical data at small speeds. The position distribution functions for both models are measured for the first time and we find good agreement with our analytic results. We are also able to extend this method to models with a scalar order parameter.

cond-mat.soft

A Model for Striped Growth

We introduce a model for describing the defected growth of striped patterns. This model, while roughly related to the Swift-Hohenberg model, generates a quite different mixture of defects during phase ordering. We find two characteristic lengths in the system: the scaling length L(t), and the average width of the domain walls. The growth law exponent is larger than the value of 1/2 found in typical point defect systems.

cond-mat.soft

Vortex Dynamics in a Coarsening Two Dimensional XY Model

The vortex velocity distribution function for a 2-dimensional coarsening non-conserved O(2) time-dependent Ginzburg-Landau model is determined numerically and compared to theoretical predictions. In agreement with these predictions the distribution function scales with the average vortex speed which is inversely proportional to t^x, where t is the time after the quench and x is near to 1/2. We find the entire curve, including a large speed algebraic tail, in good agreement with the theory.

cond-mat.soft

Defect Structures in the Growth Kinetics of the Swift-Hohenberg Model

The growth of striped order resulting from a quench of the two-dimensional Swift-Hohenberg model is studied in the regime of a small control parameter and quenches to zero temperature. We introduce an algorithm for finding and identifying the disordering defects (dislocations, disclinations and grain boundaries) at a given time. We can track their trajectories separately. We find that the coarsening of the defects and lowering of the effective free energy in the system are governed by a growth law $L(t)\approx t^{x}$ with an exponent x near 1/3. We obtain scaling for the correlations of the nematic order parameter with the same growth law. The scaling for the order parameter structure factor is governed, as found by others, by a growth law with an exponent smaller than x and near to 1/4. By comparing two systems with different sizes, we clarify the finite size effect. We find that the system has a very low density of disclinations compared to that for dislocations and fraction of points in grain boundaries. We also measure the speed distributions of the defects at different times and find that they all have power-law tails and the average speed decreases as a power law.

cond-mat.soft