SearcharxivSearch

arXiv subjects

Neal Batra

Publications and source records attributed to Neal Batra.

2 recordsLinked to original sources

From Optimal Actions to World Models: Identifiability of Transition Kernels in Discounted MDPs

We study what can be recovered about the transition probabilities of a Markov decision process from optimal actions alone. This is closely related to the inverse problem considered by Letcher et al., who ask when the dynamics can be recovered from numerical \(Q\)-values. Here the numerical values themselves are not observed; only the optimal actions are known, for every reward in a given class. For state-action rewards \(r(s,a)\), knowing the optimal actions for every reward also tells us how much better one action is than another when each is followed by the same fixed policy. This is still not enough to determine the transition probabilities uniquely. We prove that two kernels give the same optimal actions for every reward exactly when \[ Q_{s,a} = \Bigl(P_{s,a}+\tfrac1\gamma e_s^{\mathsf T}(L-I)\Bigr)L^{-1} \] for one invertible matrix \(L\) satisfying \(L\mathbf 1=\mathbf 1\). Near a kernel with strictly positive entries, there is an \(n(n-1)\)-dimensional family of different kernels with this property. The result is unchanged if we consider only rewards having a unique optimal action at every state. We then compare this with rewards of the forms \(r(s)\) and \(r(s,a,s')\). Rewards that depend on the next state can usually recover the transition kernel itself: every row at a state with at least two actions is determined, and we describe exactly when a row at a state with one action can remain hidden. State rewards reveal less: two kernels give the same optimal actions exactly when every deterministic policy is optimal for the same set of rewards. The results show how the form of the reward affects what can be learned about the dynamics from optimal actions alone.

cs.LG

Modelling High-Frequency Data with Bivariate Hawkes Processes: Power-Law vs. Exponential Kernels

This study explores the application of Hawkes processes to model high-frequency data in the context of limit order books. Two distinct Hawkes-based models are proposed and analyzed: one utilizing exponential kernels and the other employing power-law kernels. These models are implemented within a bivariate framework. The performance of each model is evaluated using high-frequency trading data, with a focus on their ability to reproduce key statistical properties of limit order books. Through a comprehensive comparison, we identify the strengths and limitations of each kernel type, providing insights into their suitability for modeling high-frequency financial data. Simulations are conducted to validate the models, and the results are interpreted. Based on these insights, a trading strategy is formulated.

q-fin.MF