Searcharxiv⌕ Search

arXiv subjects

Abhishek Bhattacharjee

Publications and source records attributed to Abhishek Bhattacharjee.

At least 19 recordsLinked to original sources

Exact Bayes Regret and Asymptotic Optimality in High-Dimensional Gaussian Bandits

We study Bayesian linear bandits with an isotropic Gaussian parameter, independent Gaussian candidate arms, and Gaussian reward noise when the horizon is proportional to the dimension. The normalized posterior uncertainty has an explicit limit that is uniform over all causal policies. Gaussian posterior identities then determine the limiting parameter overlaps without an assumed closure of the adaptive recursion. These results yield exact regret curves for Thompson sampling, posterior-mean greedy selection, and a family of policies that scale the posterior sampling covariance. The normalized realized cumulative regret converges in L1, uniformly on compact proportional-time intervals. A policy-uniform lower bound identifies the limiting optimal Bayes regret and proves that posterior-mean greedy selection attains it. Thompson sampling incurs a strictly larger leading regret; its instantaneous regret ratio relative to greedy selection lies between one and two and approaches two at long proportional horizons. Closed-form cumulative curves also identify a different comparison in the vanishing-noise limit. Finally, the instantaneous regret converges to a nondegenerate Gaussian decision-loss distribution, rather than to its mean. The analysis separates the amount of information acquired by a bandit policy from the quality of the decisions made using that information.

stat.ML↗

Kinetic energy functional constructed from exact gradient expansion of second order in uniform gas limit

Orbital-Free Density Functional Theory (OFDFT) has re-emerged as a viable alternative to Kohn-Sham DFT, driven by recent advances in kinetic energy density functionals (KEDFs). Nonlocal (NL) KEDFs have significantly extended OFDFT's applicability, particularly for bulk solids, but their high computational cost and dependence of system-specific parameters limit their universality. In this work, we propose a semilocal KEDF at the Generalized Gradient Approximation (GGA) level that achieves accuracy comparable to state-of-the-art NL and meta-GGA functionals, while remaining entirely parameter-free. Our construction revives the Thomas-Fermi-von Weizsacker (TFvW) framework by modulating the relative contributions of TF and vW terms through physically motivated constraints and preserving the exact second-order gradient expansion. Despite its simple form, the proposed functional (KGE2) performs remarkably well across both extended systems (metals and semiconductors) and finite systems (clusters), without any need for parameter tuning. These results mark a step toward a transferable, computationally efficient, and general-purpose KEDF suitable for large-scale OFDFT simulations.

cond-mat.mtrl-sci↗

Distribution-Free Changepoint Inference for Rainfall Patterns

Long records of rainfall contain critical information about the temporal organization of precipitation within a year, including wet and dry spells. Detecting when this temporal pattern changes is central to climate-impact assessment and environmental monitoring.A common data structure features observations from m locations over T years. For location i and year t, the record is a vector of daily measurements $X_{i,t} = (X_{i,t,1}, ..., X_{i,t,D})^\top$, with D = 365. The goal is to identify an unknown year $t_0$ when the temporal rainfall pattern changes across the region. Because changes in the intra-annual pattern affect the frequency-domain behavior of the daily sequence, we represent each yearly curve through a spectral feature capturing the strength of temporal oscillation at a fixed frequency.Classical changepoint methods often rely on parametric likelihoods or Gaussian approximations, which are hard to justify for skewed and heavy-tailed rainfall data. Furthermore, many procedures only return a point estimate. A confidence set for the change year is crucial in environmental applications to quantify temporal uncertainty and distinguish sharp transitions from weak evidence.This paper develops a distribution-free framework for detecting a synchronized structural change in yearly rainfall patterns across independently monitored locations. For each location-year pair, we compute a fixed-frequency nonparametric spectral density estimate. These features form a $T \times m$ data matrix. By exploiting the exchangeability of yearly spectral features before and after the true changepoint, alongside spatial independence, we construct finite-sample valid confidence sets for the unknown change year.

stat.ME↗

Cluster-Based Dimensionality Reduction by Nonparametric Distributional Screening

We consider dimensionality reduction for high-dimensional observations accompanied by a supplied partition into two or more clusters. The objective is not to construct a low-rank projection, but to retain an interpretable subset of the original coordinates that preserves the distributional information distinguishing the clusters. For each coordinate, the proposed procedure compares the cluster-specific empirical distribution functions through a several-sample Kolmogorov-Smirnov separation statistic. We formalize the resulting marginal cluster support and establish simultaneous finite-sample concentration over all coordinates, explicit bounds for false inclusions and omissions, and exact support recovery when the minimum distributional separation dominates the high-dimensional stochastic error. We also quantify the dimension inflation induced by using an unadjusted testing level and give a familywise-error-controlled version. Under a conditional sufficiency condition, sure screening preserves the full-data posterior cluster probabilities, mutual information, and Bayes risk; an additional result characterizes robustness to imperfectly estimated cluster labels. The procedure is invariant to strictly increasing coordinate transformations and can retain low-variance cluster signals that principal components may discard. We further develop average dual information, a criterion combining partition agreement after transformation with structural coverage of cluster-relevant coordinates, and derive its basic properties and consistency. Simulations illustrate the theory, the interpretability of the selected coordinates, and the distinction between cluster-directed screening and variance-directed projection.

stat.ME↗

Artificial Rosetta Stone: Constrained Maximum A Posteriori (MAP) Reconstruction of Symbolic Raga Sequences via Order-k Markov Models

Reconstructing a damaged musical fragment is an inverse problem: the observed sequence contains partial information, while a raga encodes constraints limiting allowable completions. This paper formalizes a mathematical framework for this, proposing the Artificial Rosetta Stone (ARS). We separate three claims often conflated: a symbolic sequence can be reconstructed probabilistically; a sequence can be consistent with an explicit grammar; and a historical performance can be authenticated. We only support the first two. We model a raga via a finite alphabet and constraint system, using an order-k Markov model for melodic probabilities. A symmetric Dirichlet prior yields a tractable posterior. We pose missing-note reconstruction as a constrained MAP problem. For fixed-length sequences and finite-order constraints, optimization admits an exact dynamic-programming solution with worst-case time complexity $O(TN^{k+1})$. We derive the parameter count $N^k(N - 1)$, prove a concentration bound under explicit mixing assumptions, and analyze estimation error propagation. A reproducible synthetic experiment uses six raga-inspired alphabets, orders $k \in \{1, 2, 3\}$, and masking rates up to 50%. This is a proof of concept, not historical reconstruction. A real-audio feasibility pilot evaluates 30 usable sequences from 42 Yaman clips via automated pitch extraction, segmentation, and quantization. Lacking documented provenance and relying on automated transcription, this is not expert-validated archival reconstruction. Claims are tied to stated conditions, not universal properties of Hindustani music. Code: https://github.com/mathacker23/ArtificialRosettaStone.

cs.SD↗

Design-Based Inference under Deep Domain Stratification: Language of Instruction and Private-Institution Choice in India's NSS 71st Round

Large household surveys support precise national estimates but can become statistically fragile after repeated disaggregation by geography, sector, sex, age, and outcome category. This paper develops an auditable design-based framework for deciding how far such disaggregation can be taken in a stratified multistage survey. The framework is built around a nested contribution ledger that reconstructs each domain total through the first stage probability proportional to size expansion, the certainty-plus-random hamlet-group selection, and the second-stage household expansion. Nonlinear domain parameters are expressed as ratios of these totals and analyzed by first-order linearization. The two independent National Sample Survey subsamples then provide a natural replication variance estimator. A granularity-stability profile combines the resulting relative standard error with replicate support and concentration diagnostics, so that a detailed estimate is accompanied by evidence about whether the design can sustain it. Finite population unbiasedness of the total estimator, asymptotic validity of the ratio linearization, and unbiasedness of the two-subsample variance estimator for linearized totals are established. The method is illustrated with the 71st-round Social Consumption: Education survey, focusing on home language versus medium of instruction and reported reasons for preferring private educational institutions in India and Himachal Pradesh. The application preserves the substantive analysis in the original project while replacing ad hoc calculation with a reproducible inferential workflow.

stat.ME↗

Bayesian Distilled Clustering for High-Dimensional Mixture Models

Latent subgroup analysis is central to fields such as genomics, precision medicine, and social science, where the goal is to identify heterogeneous populations with distinct covariate structures or response behaviors. Mixture models provide a natural probabilistic framework for this task, representing the data-generating distribution as a weighted combination of subgroup-specific laws with unobserved labels. In high-dimensional regimes, these analyses face significant challenges. Often, only a small subset of covariates drives meaningful subgroup separation; the remaining variables may introduce noise or redundancy. Standard clustering methods typically treat all dimensions as equal, but in high-dimensional spaces, irrelevant coordinates can distort distances and obscure the low-dimensional structures defining latent classes. This paper introduces a Bayesian distilled clustering framework for high-dimensional mixture models. We propose that clustering should occur within a statistically justified subspace rather than the full ambient space. Our method utilizes a Bayesian variable selection model to estimate posterior inclusion probabilities, quantifying the evidence that each covariate contributes to subgroup separation or response behavior. A "distilled" covariate set is then identified by controlling the expected false-discovery proportion. Clustering is performed on this reduced subspace, followed by conditional independence diagnostics to examine subgroup-specific dependencies among selected variables. Critically, this framework is model-based: the distillation step is tied directly to the mixture structure and response model rather than a generic dimension-reduction criterion. This ensures the resulting subspace remains aligned with the scientific objective: identifying latent subgroups that differ in both distributional structure and behavior.

stat.ME↗

Physically motivated iso-orbital indicator for meta-GGA exchange functionals

The iso-orbital indicator $α= (τ- τ^\mathrm{vW})/τ^\mathrm{UEG}$ is a key ingredient of meta-generalized gradient approximation (meta-GGA) functionals, but diverges in low-density tails , causing unphysical exchange potentials and systematic band gap errors as noted in [J. Chem. Phys. 150, 161101 (2019)]. We replace the denominator of $α$ with a physically motivated Pauli KED drawn from the orbital-free DFT literature, eliminating the divergence in the low density atomic tail without any empirical regularization parameter. Testing two such enhancement factors: LKT and PGS, within the r$^2$SCAN and MS2 exchange functionals, we find that the modified indicators suppress spurious oscillations in the semilocal exchange potential and restore correct electron localization in atomic tails. For a ten-member cubic semiconductor benchmark, the band gap mean absolute error is reduced by 41.1 % for r$^2$SCAN@PGS and 48.8 % for MS2@PGS, while cohesive energy accuracy is largely preserved. The consistent improvement across two functionals with distinct constructions confirms a physical rather than functional specific origin, and motivates further development of meta-GGA functionals with constraint satisfying iso-orbital indicators.

cond-mat.mtrl-sci↗

Nonlocal Orbital-Free Kinetic Energy Functional from the Jellium-with-Gap Model for Finite Systems

The quasi-linear scaling of orbital-free density functional theory (OF-DFT) with system size makes it a computationally efficient alternative to conventional Kohn--Sham density functional theory for many condensed-matter applications. However, its applicability remains limited, particularly for finite systems such as molecular clusters, due to the lack of accurate kinetic energy density functionals. In this context, the development of nonlocal kinetic energy density functionals (NL-KEDFs) has significantly advanced the practical utility of OF-DFT. Here, following an alternative formulation based on the linear-response kernel derived from the jellium-with-gap model (JGM), we develop an NL-KEDF capable of accurately describing the diverse density regimes characteristic of finite systems, including molecular clusters. Benchmark calculations, together with an analysis of the corresponding Pauli potentials, demonstrate that the proposed functional achieves higher accuracy than state-of-the-art orbital-free approaches for finite systems. Furthermore, the optical properties computed using the present method show good agreement with reference results, highlighting its reliability. These results indicate that the proposed NL-KEDF provides a robust and efficient framework for extending OF-DFT to finite systems, with potential implications for nanomaterial design and a deeper understanding of nanoscale phenomena.

cond-mat.mtrl-sci↗

CounterPoint: Using Hardware Event Counters to Refute and Refine Microarchitectural Assumptions (Extended Version)

Hardware event counters offer the potential to reveal not only performance bottlenecks but also detailed microarchitectural behavior. In practice, this promise is undermined by their vague specifications, opaque designs, and multiplexing noise, making event counter data hard to interpret. We introduce CounterPoint, a framework that tests user-specified microarchitectural models - expressed as $μ$path Decision Diagrams - for consistency with performance counter data. When mismatches occur, CounterPoint pinpoints plausible microarchitectural features that could explain them, using multi-dimensional counter confidence regions to mitigate multiplexing noise. We apply CounterPoint to the Haswell Memory Management Unit as a case study, shedding light on multiple undocumented and underdocumented microarchitectural behaviors. These include a load-store queue-side TLB prefetcher, merging page table walkers, abortable page table walks, and more. Overall, CounterPoint helps experts reconcile noisy hardware performance counter measurements with their mental model of the microarchitecture - uncovering subtle, previously hidden hardware features along the way.

cs.AR↗

Regulating Next-Generation Implantable Brain-Computer Interfaces: Recommendations for Ethical Development and Implementation

Brain-computer interfaces offer significant therapeutic opportunities for a variety of neurophysiological and neuropsychiatric disorders and may perhaps one day lead to augmenting the cognition and decision-making of the healthy brain. However, existing regulatory frameworks designed for implantable medical devices are inadequate to address the unique ethical, legal, and social risks associated with next-generation networked brain-computer interfaces. In this article, we make nine recommendations to support developers in the design of BCIs and nine recommendations to support policymakers in the application of BCIs, drawing insights from the regulatory history of IMDs and principles from AI ethics. We begin by outlining the historical development of IMDs and the regulatory milestones that have shaped their oversight. Next, we summarize similarities between IMDs and emerging implantable BCIs, identifying existing provisions for their regulation. We then use two case studies of emerging cutting-edge BCIs, the HALO and SCALO computer systems, to highlight distinctive features in the design and application of next-generation BCIs arising from contemporary chip architectures, which necessitate reevaluating regulatory approaches. We identify critical ethical considerations for these BCIs, including unique conceptions of autonomy, identity, and mental privacy. Based on these insights, we suggest potential avenues for the ethical regulation of BCIs, emphasizing the importance of interdisciplinary collaboration and proactive mitigation of potential harms. The goal is to support the responsible design and application of new BCIs, ensuring their safe and ethical integration into medical practice.

cs.HC↗

Fiduciary AI for the Future of Brain-Technology Interactions

Brain foundation models represent a new frontier in AI: instead of processing text or images, these models interpret real-time neural signals from EEG, fMRI, and other neurotechnologies. When integrated with brain-computer interfaces (BCIs), they may enable transformative applications-from thought controlled devices to neuroprosthetics-by interpreting and acting on brain activity in milliseconds. However, these same systems pose unprecedented risks, including the exploitation of subconscious neural signals and the erosion of cognitive liberty. Users cannot easily observe or control how their brain signals are interpreted, creating power asymmetries that are vulnerable to manipulation. This paper proposes embedding fiduciary duties-loyalty, care, and confidentiality-directly into BCI-integrated brain foundation models through technical design. Drawing on legal traditions and recent advancements in AI alignment techniques, we outline implementable architectural and governance mechanisms to ensure these systems act in users' best interests. Placing brain foundation models on a fiduciary footing is essential to realizing their potential without compromising self-determination.

cs.CY↗

Meta-GGA dielectric-dependent and range-separated screened hybrid functional for reliable prediction of material properties

We propose a range-separated hybrid exchange-correlation functional to calculate solid-state material properties. The functional mixes Hartree-Fock exchange with the semilocal exchange of the meta-generalized gradient approximation (meta-GGA) and the fraction of Hartree-Fock exchange is determined from the dielectric function. First-principles calculations and comparison with other meta-GGA approximations show that the functional leads to reasonably good performance for the band gap and optical properties. We also show that the present functional also successfully resolves the well-known ``band gap problem'' of narrow gap Cu-based semiconductors, such as Cu3SbSe4 and Cu3AsSe4, where, in general, a considerably large band inversion energy leads to a ``false'' negative or metallic band gap for all other methods. Furthermore, reasonable accuracy for the occupied d-bands and transition energies is also obtained for bulk solids. Thus, overall, our results demonstrate the predictive power of range-separated meta-GGA hybrid functionals for quantum materials simulations.

cond-mat.mtrl-sci↗

PULSE: Accelerating Distributed Pointer-Traversals on Disaggregated Memory (Extended Version)

Caches at CPU nodes in disaggregated memory architectures amortize the high data access latency over the network. However, such caches are fundamentally unable to improve performance for workloads requiring pointer traversals across linked data structures. We argue for accelerating these pointer traversals closer to disaggregated memory in a manner that preserves expressiveness for supporting various linked structures, ensures energy efficiency and performance, and supports distributed execution. We design PULSE, a distributed pointer-traversal framework for rack-scale disaggregated memory to meet all the above requirements. Our evaluation of PULSE shows that it enables low-latency, high-throughput, and energy-efficient execution for a wide range of pointer traversal workloads on disaggregated memory that fare poorly with caching alone.

cs.DC↗

The Interplay of Computing, Ethics, and Policy in Brain-Computer Interface Design

Brain-computer interfaces (BCIs) connect biological neurons in the brain with external systems like prosthetics and computers. They are increasingly incorporating processing capabilities to analyze and stimulate neural activity, and consequently, pose unique design challenges related to ethics, law, and policy. For the first time, this paper articulates how ethical, legal, and policy considerations can shape BCI architecture design, and how the decisions that architects make constrain or expand the ethical, legal, and policy frameworks that can be applied to them.

cs.AR↗

Towards Forever Access for Implanted Brain-Computer Interfaces

Designs for implanted brain-computer interfaces (BCIs) have increased significantly in recent years. Each device promises better clinical outcomes and quality-of-life improvements, yet due to severe and inflexible safety constraints, progress requires tight co-design from materials to circuits and all the way up the stack to applications and algorithms. This trend has become more aggressive over time, forcing clinicians and patients to rely on vendor-specific hardware and software for deployment, maintenance, upgrades, and replacement. This over-reliance is ethically problematic, especially if companies go out-of-business or business objectives diverge from clinical promises. Device heterogeneity additionally burdens clinicians and healthcare facilities, adding complexity and costs for in-clinic visits, monitoring, and continuous access. Reliability, interoperability, portability, and future-proofed design is needed, but this unfortunately comes at a cost. These system features sap resources that would have otherwise been allocated to reduce power/energy and improve performance. Navigating this trade-off in a systematic way is critical to providing patients with forever access to their implants and reducing burdens placed on healthcare providers and caretakers. We study the integration of on-device storage to highlight the sensitivity of this trade-off and establish other points of interest within BCI design that require careful investigation. In the process, we revisit relevant problems in computer architecture and medical devices from the current era of hardware specialization and modern neurotechnology.

cs.AR↗

Swapping-Centric Neural Recording Systems

Neural interfaces read the activity of biological neurons to help advance the neurosciences and offer treatment options for severe neurological diseases. The total number of neurons that are now being recorded using multi-electrode interfaces is doubling roughly every 4-6 years \cite{Stevenson2011}. However, processing this exponentially-growing data in real-time under strict power-constraints puts an exorbitant amount of pressure on both compute and storage within traditional neural recording systems. Existing systems deploy various accelerators for better performance-per-watt while also integrating NVMs for data querying and better treatment decisions. These accelerators have direct access to a limited amount of fast SRAM-based memory that is unable to manage the growing data rates. Swapping to the NVM becomes inevitable; however, naive approaches are unable to complete during the refractory period of a neuron -- i.e., a few milliseconds -- which disrupts timely disease treatment. We propose co-designing accelerators and storage, with swapping as a primary design goal, using theoretical and practical models of compute and storage respectively to overcome these limitations.

cs.AR↗

The QUATRO Application Suite: Quantum Computing for Models of Human Cognition

Research progress in quantum computing has, thus far, focused on a narrow set of application domains. Expanding the suite of quantum application domains is vital for the discovery of new software toolchains and architectural abstractions. In this work, we unlock a new class of applications ripe for quantum computing research -- computational cognitive modeling. Cognitive models are critical to understanding and replicating human intelligence. Our work connects computational cognitive models to quantum computer architectures for the first time. We release QUATRO, a collection of quantum computing applications from cognitive models. The development and execution of QUATRO shed light on gaps in the quantum computing stack that need to be closed to ease programming and drive performance. Among several contributions, we propose and study ideas pertaining to quantum cloud scheduling (using data from gate- and annealing-based quantum computers), parallelization, and more. In the long run, we expect our research to lay the groundwork for more versatile quantum computer systems in the future.

cs.CE↗