Searcharxiv⌕ Search

arXiv subjects

Chao Yu

Publications and source records attributed to Chao Yu.

At least 163 records · Page 9Linked to original sources

Robust optimal investment and risk control for an insurer with general insider information

In this paper, we study the robust optimal investment and risk control problem for an insurer who owns the insider information about the financial market and the insurance market under model uncertainty. Both financial risky asset process and insurance risk process are assumed to be very general jump diffusion processes. The insider information is of the most general form rather than the initial enlargement type. We use the theory of forward integrals to give the first half characterization of the robust optimal strategy and transform the anticipating stochastic differential game problem into the nonanticipative stochastic differential game problem. Then we adopt the stochastic maximum principle to obtain the total characterization of the robust strategy. We discuss the two typical situations when the insurer is `small' and `large' by Malliavin calculus. For the `small' insurer, we obtain the closed-form solution in the continuous case and the half closed-form solution in the case with jumps. For the `large' insurer, we reduce the problem to the quadratic backward stochastic differential equation (BSDE) and obtain the closed-form solution in the continuous case without model uncertainty. We discuss some impacts of the model uncertainty, insider information and the `large' insurer on the optimal strategy.

math.NA↗

Malliavin calculus and its application to robust optimal portfolio for an insider

Insider information and model uncertainty are two unavoidable problems for the portfolio selection theory in reality. This paper studies the robust optimal portfolio strategy for an investor who owns general insider information under model uncertainty. On the aspect of the mathematical theory, we improve some properties of the forward integral and use Malliavin calculus to derive the anticipating Itô formula . Then we use forward integrals to formulate the insider-trading problem with model uncertainty. We give the half characterization of the robust optimal portfolio and obtain the semimartingale decomposition of the driving noise $W$ with respect to the insider information filtration, which turns the problem turns to the nonanticipative stochastic differential game problem. We give the total characterization by the stochastic maximum principle. When considering two typical situations where the insider is `small' and `large', we give the corresponding BSDEs to characterize the robust optimal portfolio strategy, and derive the closed form of the portfolio and the value function in the case of the small insider by the Donsker $δ$ functional. We present the simulation result and give the economic analysis of optimal strategies under different situations.

math.NA↗

ESCM$^2$: Entire Space Counterfactual Multi-Task Model for Post-Click Conversion Rate Estimation

Accurate estimation of post-click conversion rate is critical for building recommender systems, which has long been confronted with sample selection bias and data sparsity issues. Methods in the Entire Space Multi-task Model (ESMM) family leverage the sequential pattern of user actions, i.e. $impression\rightarrow click \rightarrow conversion$ to address data sparsity issue. However, they still fail to ensure the unbiasedness of CVR estimates. In this paper, we theoretically demonstrate that ESMM suffers from the following two problems: (1) Inherent Estimation Bias (IEB), where the estimated CVR of ESMM is inherently higher than the ground truth; (2) Potential Independence Priority (PIP) for CTCVR estimation, where there is a risk that the ESMM overlooks the causality from click to conversion. To this end, we devise a principled approach named Entire Space Counterfactual Multi-task Modelling (ESCM$^2$), which employs a counterfactual risk miminizer as a regularizer in ESMM to address both IEB and PIP issues simultaneously. Extensive experiments on offline datasets and online environments demonstrate that our proposed ESCM$^2$ can largely mitigate the inherent IEB and PIP issues and achieve better performance than baseline models.

cs.AI↗

Constrained Sequence-to-Tree Generation for Hierarchical Text Classification

Hierarchical Text Classification (HTC) is a challenging task where a document can be assigned to multiple hierarchically structured categories within a taxonomy. The majority of prior studies consider HTC as a flat multi-label classification problem, which inevitably leads to "label inconsistency" problem. In this paper, we formulate HTC as a sequence generation task and introduce a sequence-to-tree framework (Seq2Tree) for modeling the hierarchical label structure. Moreover, we design a constrained decoding strategy with dynamic vocabulary to secure the label consistency of the results. Compared with previous works, the proposed approach achieves significant and consistent improvements on three benchmark datasets.

cs.CL↗

Passive Motion Detection via mmWave Communication System

In this paper, an integrated passive sensing and communication system working in 60 GHz band is elaborated, and the sensing performance is investigated in an application of hand gesture recognition. Specifically, in this integrated system, there are two radio frequency (RF) chains at the receiver and one at the transmitter. Each RF chain is connected with one phased array for analog beamforming. To facilitate simultaneous sensing and communication, the transmitter delivers one stream of information-bearing signals via two beam lobes, one is aligned with the main signal propagation path and the other is directed to the sensing target. Signals from the two lobes are received by the two RF chains at the receiver, respectively. By cross ambiguity coherent processing, the time-Doppler spectrograms of hand gestures can be obtained. Relying on the passive sensing system, a dataset of received signals, where three types of hand gestures are sensed, is collected by using Line-of-Sight (LoS) and Non-Line-of-Sight (NLoS) paths as the reference channel respectively. Then a neural network is trained by the dataset for motion detection. It is shown that the classification accuracy rate is high as long as sufficient sensing time is assured. Finally, an empirical model characterizing the relation between the classification accuracy and sensing duration is derived analytically.

eess.SP↗

Quasi-Monte Carlo-Based Conditional Malliavin Method for Continuous-Time Asian Option Greeks

Although many methods for computing the Greeks of discrete-time Asian options are proposed, few methods to calculate the Greeks of continuous-time Asian options are known. In this paper, we develop an integration by parts formula in the multi-dimensional Malliavin calculus, and apply it to obtain the Greeks formulae for continuous-time Asian options in the multi-asset situation. We combine the Malliavin method with the quasi-Monte Carlo method to calculate the Greeks in simulation. We discuss the asymptotic convergence of simulation estimates for the continuous-time Asian option Greeks obtained by Malliavin derivatives. We propose to use the conditional quasi-Monte Carlo method to smooth Malliavin Greeks, and show that the calculation of conditional expectations analytically is viable for many types of Asian options. We prove that the new estimates for Greeks have good smoothness. For binary Asian options, Asian call options and up-and-out Asian call options, for instance, our estimates are infinitely times differentiable. We take the gradient principal component analysis method as a dimension reduction technique in simulation. Numerical experiments demonstrate the large efficiency improvement of the proposed method, especially for Asian options with discontinuous payoff functions.

math.NA↗

High-harmonic generation approaching the quantum critical point of strongly correlated systems

By employing the exact diagonalization method, we investigate the high-harmonic generation (HHG) of the correlated systems under the strong laser irradiation. For the extended Hubbard model on a periodic chain, HHG close to the quantum critical point (QCP) is more significant compared to two neighboring gapped phases (i.e., charge-density-wave and spin-density wave states), especially in low-frequencies. We confirm that the systems in the vicinity of the QCP are supersensitive to the external field and more optical-transition channels via excited states are responsible for HHG. This feature holds the potential of obtaining high-efficiency harmonics by making use of materials approaching to QCP. Based on two-dimensional Haldane model, we further propose that the even- or odd-order components of generated harmonics can be promisingly regarded as spectral signals to distinguish the topologically ordered phases from locally ordered ones. Our findings in this work pave the way to achieve ultrafast light source from HHG in strongly correlated materials and to study quantum phase transition by nonlinear optics in strong laser fields.

cond-mat.str-el↗

Note on surface growth approach for bulk reconstruction

In a recent paper, a novel surface growth approach for reconstructing bulk geometry and matter fields was proposed, it was shown that this picture can be explicitly realized by the one-shot entanglement distillation tensor network and the surface state correspondence. In the present paper, we give direct analysis for the growth of the bulk minimal surfaces in asymptotically AdS 3 spacetime and show that bulk geometry can be efficiently reproduced in this way, which provides further support for the surface growth approach in entanglement wedge reconstruction.

hep-th↗

Solid-like high harmonic generation from rotationally periodic systems

High harmonic generation (HHG) from crystals in strong laser fields has been understood by the band theory of solid, which is based on the periodic boundary condition (PBC) of translational invariant. For systems having PBC of rotational invariant, in principles an analogous Bloch theorem can be developed and applied. Taking a ring-type cluster of cyclo[18]carbon as a representative, we theoretically suggest a quasi-band model and study its HHG by solving time-dependent Liouville-von Neumann equation. Under the irradiation of circularly polarized laser, explicit selection rules for left-handed and right-handed harmonics are observed, while in linearly polarized laser field, cyclo[18]carbon exhibits solid-like HHG originated from intra-band oscillations and inter-band transitions, which in turn is promising to optically detect the symmetry and geometry of controversial structures. In a sense, this work presents a connection linking the high harmonics of gases and solids.

physics.optics↗

The Role of Shift Vector in High-Harmonic Generation from Non-Centrosymmetric Topological Insulators under Strong Laser Fields

As a promising avenue to obtain new extreme ultraviolet light source and detect electronic properties, high-harmonic generation (HHG) has been actively developed in both theory and experiment. In solids lacking inversion symmetry, when electrons undergo a nonadiabatic transition, a directional charge shift occurs and is characterized by shift vector, which measures the real-space shift of the photoexcited electron and hole. For the first time, we have revealed that shift vector plays prominent roles in the real-space tunneling mechanism of three-step model for electrons under strong laser fields. Since shift vector is determined by the topological properties of related wave functions, we expect HHG with its contribution can provide direct knowledge on the band topology in noncentrosymmetric topological insulators (TIs). In both Kane-Mele model and realistic material BiTeI, we have found that the shift vector reverses when band inversion happens during the topological phase transition between normal and topological insulators. Under oscillating strong laser fields, the reversal of shift vector leads to completely opposite radiation time of high-order harmonics. This makes HHG a feasible all-optical strong-field method to directly identify the band inversion in non-centrosymmetric TIs.

physics.optics↗

Multi-Agent Vulnerability Discovery for Autonomous Driving with Hazard Arbitration Reward

Discovering hazardous scenarios is crucial in testing and further improving driving policies. However, conducting efficient driving policy testing faces two key challenges. On the one hand, the probability of naturally encountering hazardous scenarios is low when testing a well-trained autonomous driving strategy. Thus, discovering these scenarios by purely real-world road testing is extremely costly. On the other hand, a proper determination of accident responsibility is necessary for this task. Collecting scenarios with wrong-attributed responsibilities will lead to an overly conservative autonomous driving strategy. To be more specific, we aim to discover hazardous scenarios that are autonomous-vehicle responsible (AV-responsible), i.e., the vulnerabilities of the under-test driving policy. To this end, this work proposes a Safety Test framework by finding Av-Responsible Scenarios (STARS) based on multi-agent reinforcement learning. STARS guides other traffic participants to produce Av-Responsible Scenarios and make the under-test driving policy misbehave via introducing Hazard Arbitration Reward (HAR). HAR enables our framework to discover diverse, complex, and AV-responsible hazardous scenarios. Experimental results against four different driving policies in three environments demonstrate that STARS can effectively discover AV-responsible hazardous scenarios. These scenarios indeed correspond to the vulnerabilities of the under-test driving policies, thus are meaningful for their further improvements.

cs.RO↗

Lifelong Reinforcement Learning with Temporal Logic Formulas and Reward Machines

Continuously learning new tasks using high-level ideas or knowledge is a key capability of humans. In this paper, we propose Lifelong reinforcement learning with Sequential linear temporal logic formulas and Reward Machines (LSRM), which enables an agent to leverage previously learned knowledge to fasten learning of logically specified tasks. For the sake of more flexible specification of tasks, we first introduce Sequential Linear Temporal Logic (SLTL), which is a supplement to the existing Linear Temporal Logic (LTL) formal language. We then utilize Reward Machines (RM) to exploit structural reward functions for tasks encoded with high-level events, and propose automatic extension of RM and efficient knowledge transfer over tasks for continuous learning in lifetime. Experimental results show that LSRM outperforms the methods that learn the target tasks from scratch by taking advantage of the task decomposition using SLTL and knowledge transfer over RM during the lifelong learning process.

cs.AI↗

Coordinated Proximal Policy Optimization

We present Coordinated Proximal Policy Optimization (CoPPO), an algorithm that extends the original Proximal Policy Optimization (PPO) to the multi-agent setting. The key idea lies in the coordinated adaptation of step size during the policy update process among multiple agents. We prove the monotonicity of policy improvement when optimizing a theoretically-grounded joint objective, and derive a simplified optimization objective based on a set of approximations. We then interpret that such an objective in CoPPO can achieve dynamic credit assignment among agents, thereby alleviating the high variance issue during the concurrent update of agent policies. Finally, we demonstrate that CoPPO outperforms several strong baselines and is competitive with the latest multi-agent PPO method (i.e. MAPPO) under typical multi-agent settings, including cooperative matrix games and the StarCraft II micromanagement tasks.

cs.AI↗

Enabling Large-Reach TLBs for High-Throughput Processors by Exploiting Memory Subregion Contiguity

Accelerators, like GPUs, have become a trend to deliver future performance desire, and sharing the same virtual memory space between CPUs and GPUs is increasingly adopted to simplify programming. However, address translation, which is the key factor of virtual memory, is becoming the bottleneck of performance for GPUs. In GPUs, a single TLB miss can stall hundreds of threads due to the SIMT execute model, degrading performance dramatically. Through real system analysis, we observe that the OS shows an advanced contiguity (e.g., hundreds of contiguous pages), and more large memory regions with advanced contiguity tend to be allocated with the increase of working sets. Leveraging the observation, we propose MESC to improve the translation efficiency for GPUs. The key idea of MESC is to divide each large page frame (2MB size) in virtual memory space into memory subregions with fixed size (i.e., 64 4KB pages), and store the contiguity information of subregions and large page frames in L2PTEs. With MESC, address translations of up to 512 pages can be coalesced into single TLB entry, without the needs of changing memory allocation policy (i.e., demand paging) and the support of large pages. In the experimental results, MESC achieves 77.2% performance improvement and 76.4% reduction in dynamic translation energy for translation-sensitive workloads.

cs.AR↗

Reinforcement Learning with Expert Trajectory For Quantitative Trading

In recent years, quantitative investment methods combined with artificial intelligence have attracted more and more attention from investors and researchers. Existing related methods based on the supervised learning are not very suitable for learning problems with long-term goals and delayed rewards in real futures trading. In this paper, therefore, we model the price prediction problem as a Markov decision process (MDP), and optimize it by reinforcement learning with expert trajectory. In the proposed method, we employ more than 100 short-term alpha factors instead of price, volume and several technical factors in used existing methods to describe the states of MDP. Furthermore, unlike DQN (deep Q-learning) and BC (behavior cloning) in related methods, we introduce expert experience in training stage, and consider both the expert-environment interaction and the agent-environment interaction to design the temporal difference error so that the agents are more adaptable for inevitable noise in financial data. Experimental results evaluated on share price index futures in China, including IF (CSI 300) and IC (CSI 500), show that the advantages of the proposed method compared with three typical technical analysis and two deep leaning based methods.

cs.LG↗

A Joint Training Dual-MRC Framework for Aspect Based Sentiment Analysis

Aspect based sentiment analysis (ABSA) involves three fundamental subtasks: aspect term extraction, opinion term extraction, and aspect-level sentiment classification. Early works only focused on solving one of these subtasks individually. Some recent work focused on solving a combination of two subtasks, e.g., extracting aspect terms along with sentiment polarities or extracting the aspect and opinion terms pair-wisely. More recently, the triple extraction task has been proposed, i.e., extracting the (aspect term, opinion term, sentiment polarity) triples from a sentence. However, previous approaches fail to solve all subtasks in a unified end-to-end framework. In this paper, we propose a complete solution for ABSA. We construct two machine reading comprehension (MRC) problems and solve all subtasks by joint training two BERT-MRC models with parameters sharing. We conduct experiments on these subtasks, and results on several benchmark datasets demonstrate the effectiveness of our proposed framework, which significantly outperforms existing state-of-the-art methods.

cs.CL↗

Discovering Diverse Multi-Agent Strategic Behavior via Reward Randomization

We propose a simple, general and effective technique, Reward Randomization for discovering diverse strategic policies in complex multi-agent games. Combining reward randomization and policy gradient, we derive a new algorithm, Reward-Randomized Policy Gradient (RPG). RPG is able to discover multiple distinctive human-interpretable strategies in challenging temporal trust dilemmas, including grid-world games and a real-world game Agar.io, where multiple equilibria exist but standard multi-agent policy gradient algorithms always converge to a fixed one with a sub-optimal payoff for every player even using state-of-the-art exploration techniques. Furthermore, with the set of diverse strategies from RPG, we can (1) achieve higher payoffs by fine-tuning the best policy from the set; and (2) obtain an adaptive agent by using this set of strategies as its training opponents. The source code and example videos can be found in our website: https://sites.google.com/view/staghuntrpg.

cs.AI↗

Single-photon imaging over 200 km

Long-range active imaging has widespread applications in remote sensing and target recognition. Single-photon light detection and ranging (lidar) has been shown to have high sensitivity and temporal resolution. On the application front, however, the operating range of practical single-photon lidar systems is limited to about tens of kilometers over the Earth's atmosphere, mainly due to the weak echo signal mixed with high background noise. Here, we present a compact coaxial single-photon lidar system capable of realizing 3D imaging at up to 201.5 km. It is achieved by using high-efficiency optical devices for collection and detection, and what we believe is a new noise-suppression technique that is efficient for long-range applications. We show that photon-efficient computational algorithms enable accurate 3D imaging over hundreds of kilometers with as few as 0.44 signal photons per pixel. The results represent a significant step toward practical, low-power lidar over extra-long ranges.

eess.IV↗