SearcharxivSearch

arXiv subjects

Sanchayan Dutta

Publications and source records attributed to Sanchayan Dutta.

6 recordsLinked to original sources

Active Inference as Context Acquisition for AI Agents

Interactive AI agents must acquire the right context as efficiently as possible. When a user omits a constraint, preference, file, or task variable, an agent can proceed with a default assumption or spend tokens on a clarifying question, retrieval call, tool call, or prompt trial. We formulate this tradeoff as active inference for context acquisition. An inner inference step updates beliefs over a latent task state, and an outer decision selects the next context action, task action, or stop action to minimize expected free energy under cost. In deterministic settings, the epistemic term reduces to expected information gain, optionally normalized by token cost. We instantiate the framework in Optimal Question Asking (OQA), with exact posteriors and a dynamic programming oracle, and benchmark frontier language models on binary and multiway categorical tasks from 25 to 300 candidates. We also study clarification before generation and automated prompt optimization under token budgets. The formulation is model-agnostic and views active inference as a design principle for the context-acquisition layer of AI agents.

cs.AI

A Note on Clifford Stabilizer Codes for Ising Anyons

We provide a streamlined elaboration on existing ideas that link Ising anyon (or equivalently, Majorana) stabilizer codes to certain classes of binary classical codes. The groundwork for such Majorana-based quantum codes can be found in earlier works (including, for example, Bravyi (arXiv:1004.3791) and Vijay et al. (arXiv:1703.00459)), where it was observed that commuting families of fermionic (Clifford) operators can often be systematically lifted from weakly self-dual or self-orthogonal binary codes. Here, we recast and unify these ideas into a classification theorem that explicitly shows how q-isotropic subspaces in $\mathbb{F}_2^{2n}$ yield commuting Clifford operators relevant to Ising anyons, and how these subspaces naturally correspond to punctured self-orthogonal codes in $\mathbb{F}_2^{2n+1}$.

quant-ph

Toward generalizable learning of all (linear) first-order methods via memory augmented Transformers

We show that memory-augmented Transformers can implement the entire class of linear first-order methods (LFOMs), a class that contains gradient descent (GD) and more advanced methods such as conjugate gradient descent (CGD), momentum methods and all other variants that linearly combine past gradients. Building on prior work that studies how Transformers simulate GD, we provide theoretical and empirical evidence that memory-augmented Transformers can learn more advanced algorithms. We then take a first step toward turning the learned algorithms into actually usable methods by developing a mixture-of-experts (MoE) approach for test-time adaptation to out-of-distribution (OOD) samples. Lastly, we show that LFOMs can themselves be treated as learnable algorithms, whose parameters can be learned from data to attain strong performance.

cs.LG

Riemannian Bilevel Optimization

We develop new algorithms for Riemannian bilevel optimization. We focus in particular on batch and stochastic gradient-based methods, with the explicit goal of avoiding second-order information such as Riemannian hyper-gradients. We propose and analyze $\mathrm{RF^2SA}$, a method that leverages first-order gradient information to navigate the complex geometry of Riemannian manifolds efficiently. Notably, $\mathrm{RF^2SA}$ is a single-loop algorithm, and thus easier to implement and use. Under various setups, including stochastic optimization, we provide explicit convergence rates for reaching $ε$-stationary points. We also address the challenge of optimizing over Riemannian manifolds with constraints by adjusting the multiplier in the Lagrangian, ensuring convergence to the desired solution without requiring access to second-order derivatives.

math.OC

Quantum Circuit Design Methodology for Multiple Linear Regression

Multiple linear regression assumes an imperative role in supervised machine learning. In 2009, Harrow et al. [Phys. Rev. Lett. 103, 150502 (2009)] showed that their HHL algorithm can be used to sample the solution of a linear system $\mathbf{Ax=b}$ exponentially faster than any existing classical algorithm, with some manageable caveats. The entire field of quantum machine learning gained considerable traction after the discovery of this celebrated algorithm. However, effective practical applications and experimental implementations of HHL are still sparse in the literature. Here, we demonstrate a potential practical utility of HHL, in the context of regression analysis, using the remarkable fact that there exists a natural reduction of any multiple linear regression problem to an equivalent linear systems problem. We put forward a $7$-qubit quantum circuit design, motivated from an earlier work by Cao et al. [Mol. Phys. 110, 1675 (2012)], to solve a $3$-variable regression problem, using only elementary quantum gates. We also implement the Group Leaders Optimization Algorithm (GLOA) [Mol. Phys. 109 (5), 761 (2011)] and elaborate on the advantages of using such stochastic algorithms in creating low-cost circuit approximations for the Hamiltonian simulation. We believe that this application of GLOA and similar stochastic algorithms in circuit approximation will boost time- and cost-efficient circuit designing for various quantum machine learning protocols. Further, we discuss our Qiskit simulation and explore certain generalizations to the circuit design.

quant-ph

Euler Number and Percolation Threshold on a Square Lattice with Diagonal Connection Probability and Revisiting the Island-Mainland Transition

We report some novel properties of a square lattice filled with white sites, randomly occupied by black sites (with probability $p$). We consider connections up to second nearest neighbours, according to the following rule. Edge-sharing sites, i.e. nearest neighbours of similar type are always considered to belong to the same cluster. A pair of black corner-sharing sites, i.e. second nearest neighbours may form a 'cross-connection' with a pair of white corner-sharing sites. In this case assigning connected status to both pairs simultaneously, makes the system quasi-three dimensional, with intertwined black and white clusters. The two-dimensional character of the system is preserved by considering the black diagonal pair to be connected with a probability $q$, in which case the crossing white pair of sites are deemed disjoint. If the black pair is disjoint, the white pair is considered connected. In this scenario we investigate (i) the variation of the Euler number $χ(p) \ [=N_B(p)-N_W(p)]$ versus $p$ graph for varying $q$, (ii) variation of the site percolation threshold with $q$ and (iii) size distribution of the black clusters for varying $p$, when $q=0.5$. Here $N_B$ is the number of black clusters and $N_W$ is the number of white clusters, at a certain probability $p$. We also discuss the earlier proposed 'Island-Mainland' transition (Khatun, T., Dutta, T. & Tarafdar, S. Eur. Phys. J. B (2017) 90: 213) and show mathematically that the proposed transition is not, in fact, a critical phase transition and does not survive finite size scaling. It is also explained mathematically why clusters of size 1 are always the most numerous.

cond-mat.stat-mech