SearcharxivSearch

arXiv subjects

Anish Acharya

Publications and source records attributed to Anish Acharya.

At least 19 recordsLinked to original sources

From Canonical to Tunable Phase Diagrams in Open Quantum Long-Range Systems

We investigate the dissipative dynamics of a generalized Lipkin-Meshkov-Glick (LMG) model coupled to a thermal environment. In this generalized model, in addition to the conventional quadratic interaction, one considers quartic interactions between spin-$1/2$'s coupled all-to-all and evolving in presence of a transverse field. Employing the usual linear Lindblad master equation with thermally-balanced jump processes, we derive magnetisation evolution equations, and demonstrate that the corresponding stationary solution reproduces the canonical equilibrium phase diagram of the model. We then extend our analysis to a nonlinear Lindblad equation that incorporates imperfect quantum-jump processes in terms of jump-retention parameters. Here, remarkably, the system relaxes to a genuine nonequilibrium stationary state whose properties differ qualitatively from those obtained in the linear case. The jump-retention parameters provide tunable knobs that shift the phase boundaries and even modify the nature of the phase transitions with respect to the linear case. Our exact results establish a direct connection between dissipative relaxation dynamics and stationary-state behavior, while identifying controlled quantum-jump retention as a mechanism for engineering nonequilibrium phases in long-range interacting open quantum systems.

quant-ph

RoPoLL: Robust Panel of LLM Judges

The LLM Jury, a Panel of LLM Evaluators (PoLL) reporting consensus scores, has become a practical alternative to single-judge LLM evaluation, yet its statistical behavior remains poorly understood. We formalize the LLM Jury under the Huber contamination model and show that PoLL incurs unbounded bias under any positive contamination, regardless of jury size, whenever a single judge fails in a biased, LLM-typical way (mode collapse, sycophancy, safety refusal). Framing jury consensus as classical robust mean estimation, we propose RoPoLL (Robust Panel of LLM-as-Judge), which preserves the PoLL panel but replaces the aggregation function with a robust mean estimator, instantiated with the geometric median (GM): tuning-free, with the optimal finite-sample breakdown point 1/2. A finite-sample error bound and a matching information-theoretic minimax lower bound agree on the parametric rate sigma*sqrt(d/N) and differ on the breakdown floor by a factor of sqrt(d), a statistical-computational gap that polynomial-time RoPoLL pays relative to the intractable Tukey halfspace median. Across 13 open-weight judges (4B-675B), three reward-model benchmarks, and four corruption regimes at rates up to 50%, RoPoLL dominates PoLL on every biased corruption type: by about 19% on cross-dimensional attacks at matched compute, and by orders of magnitude on heavy-tailed Byzantine adversaries. A 3-judge RoPoLL committee at 38B beats Mistral-Large-3 (675B) by 1.31x on HelpSteer-2 under 30% bimodal-random corruption, an 18x parameter advantage at better accuracy; a Noisy-GT control confirms the premium is paid against biased contamination, not benign imprecision.

cs.AI

Analytical approach to subsystem resetting in generalized Kuramoto models

Stochastic resetting has emerged as a powerful mechanism for driving systems into nonequilibrium stationary states with tunable properties. While most existing studies focus on global resetting, where all degrees of freedom are simultaneously reset, recent work has shown that resetting only a subset of degrees of freedom (subsystem resetting) can qualitatively alter collective behavior in interacting many-body systems. In this work, we develop a general theoretical framework for analysing subsystem resetting in Kuramoto-type coupled-oscillator systems. Building on a continued-fraction approach, we derive self-consistent equations for the stationary-state order parameter of the non-reset subsystem, applicable to both noisy and noiseless dynamics and to models with arbitrary interaction harmonics. Using this framework, we systematically investigate how the stationary state and phase transitions depend on the resetting rate, the size of the reset subsystem, and the reset configuration. We show that subsystem resetting can shift or even suppress synchronization transitions, and can give rise to nontrivial features such as re-entrant behavior and restructuring of phase boundaries. In specific cases, including the noiseless Kuramoto model with a Lorentzian frequency distribution, our results recover known analytical predictions and extend them to more general settings. These results establish subsystem resetting as a versatile control protocol for engineering collective dynamics in nonequilibrium interacting systems.

cond-mat.stat-mech

Dynamics and steady states of tight-binding chains in presence of isolated defects

Reduced transport and localization in isolated quantum systems are typically attributed to spatially-extended disorder, but may also emerge from the influence of a few controllable defects. We show here how a single defect profoundly reshapes wave-function spreading on a finite and periodic tight-binding lattice. Adapting the defect technique from classical random-walk studies, we obtain exact time-resolved site-occupation probabilities and several observables of interest. Even a single defect induces remarkable nonlinear effects, including non-monotonic suppression of transport, enhanced localization at distant sites, and strong sensitivity to the initial particle position at long times. These results demonstrate that minimal perturbations can generate nontrivial long-time transport signatures, giving rise to a microscopic defect-driven mechanism of quantum localization. Although the main results presented pertain to a single isolated defect, we show that the developed formalism may naturally extend to multiple as well as to a wider class of defects.

quant-ph

Stationary-state dynamics of interacting phase oscillators in presence of noise and stochastic resetting

We explore the impact of global resetting on Kuramoto-type models of coupled limit-cycle oscillators with distributed frequencies both in absence and presence of noise. The dynamics comprises repeated interruption of the bare dynamics at random times with simultaneous resetting of phases of all the oscillators to a predefined state. To characterize the stationary-state behavior, we develop an analytical framework that spans across different generalizations of the Kuramoto model involving either quenched or annealed disorder or both, and for any choice of the natural frequency distribution. The framework applies to the dynamics both in absence and presence of resetting, and is employed to obtain in particular the stationary-state synchronization order parameter of the system, which is a measure of spontaneous ordering among the oscillator phases. A key finding is the pivotal role of correlations in shaping the ordering dynamics under resettling.

cond-mat.stat-mech

Manipulating phases in many-body interacting systems with subsystem resetting

Stabilizing thermodynamically unstable phases in many-body systems, such as suppressing pathological neuronal synchronization in Parkinson's disease or maintaining magnetic order across broad temperature ranges, remains a persistent challenge. In traditional approaches, such phases are stabilized through intervening in the dynamics of all system constituents or introducing additional interactions. Here, we offer a hitherto-unexplored alternative, namely, subsystem resetting, whereby intervention in the dynamics of only a part of the system, and that too only occasionally in time, is implemented through resetting its state to a reset configuration. Just playing with a few parameters, e.g., the nature of the reset configuration and the size of the reset subsystem, one achieves a remarkable and robust control over the phase diagram of the bare dynamics. We demonstrate that these universal effects span a wide variety of scenarios, including equilibrium and non-equilibrium, mean-field and non-mean-field dynamics, with and without quenched disorder. Despite the challenges posed by memory effects, we obtain explicit analytical predictions, validated by simulations.

cond-mat.stat-mech

Geometric Median Matching for Robust k-Subset Selection from Noisy Data

Data pruning -- the combinatorial task of selecting a small and representative subset from a large dataset, is crucial for mitigating the enormous computational costs associated with training data-hungry modern deep learning models at scale. Since large scale data collections are invariably noisy, developing data pruning strategies that remain robust even in the presence of corruption is critical in practice. However, existing data pruning methods often fail under high corruption rates due to their reliance on empirical mean estimation, which is highly sensitive to outliers. In response, we propose Geometric Median (GM) Matching, a novel k-subset selection strategy that leverages Geometric Median -- a robust estimator with an optimal breakdown point of 1/2; to enhance resilience against noisy data. Our method iteratively selects a k-subset such that the mean of the subset approximates the GM of the (potentially) noisy dataset, ensuring robustness even under arbitrary corruption. We provide theoretical guarantees, showing that GM Matching enjoys an improved O(1/k) convergence rate -- a quadratic improvement over random sampling, even under arbitrary corruption. Extensive experiments across image classification and image generation tasks demonstrate that GM Matching consistently outperforms existing pruning approaches, particularly in high-corruption settings and at high pruning rates; making it a strong baseline for robust data pruning.

cs.LG

Neural Distributed Source Coding

Distributed source coding (DSC) is the task of encoding an input in the absence of correlated side information that is only available to the decoder. Remarkably, Slepian and Wolf showed in 1973 that an encoder without access to the side information can asymptotically achieve the same compression rate as when the side information is available to it. While there is vast prior work on this topic, practical DSC has been limited to synthetic datasets and specific correlation structures. Here we present a framework for lossy DSC that is agnostic to the correlation structure and can scale to high dimensions. Rather than relying on hand-crafted source modeling, our method utilizes a conditional Vector-Quantized Variational Autoencoder (VQ-VAE) to learn the distributed encoder and decoder. We evaluate our method on multiple datasets and show that our method can handle complex correlations and achieves state-of-the-art PSNR. Our code is made available at https://github.com/acnagle/neural-dsc.

cs.IT

Geometric Median (GM) Matching for Robust Data Pruning

Large-scale data collections in the wild, are invariably noisy. Thus developing data pruning strategies that remain robust even in the presence of corruption is critical in practice. In this work, we propose Geometric Median ($\gm$) Matching -- a herding style greedy algorithm that yields a $k$-subset such that the mean of the subset approximates the geometric median of the (potentially) noisy dataset. Theoretically, we show that $\gm$ Matching enjoys an improved $\gO(1/k)$ scaling over $\gO(1/\sqrt{k})$ scaling of uniform sampling; while achieving {\bf optimal breakdown point} of {\bf 1/2} even under {\bf arbitrary} corruption. Extensive experiments across several popular deep learning benchmarks indicate that $\gm$ Matching consistently improves over prior state-of-the-art; the gains become more profound at high rates of corruption and aggressive pruning rates; making $\gm$ Matching a strong baseline for future research in robust data pruning.

cs.LG

Positive Unlabeled Contrastive Learning

Self-supervised pretraining on unlabeled data followed by supervised fine-tuning on labeled data is a popular paradigm for learning from limited labeled examples. We extend this paradigm to the classical positive unlabeled (PU) setting, where the task is to learn a binary classifier given only a few labeled positive samples, and (often) a large amount of unlabeled samples (which could be positive or negative). We first propose a simple extension of standard infoNCE family of contrastive losses, to the PU setting; and show that this learns superior representations, as compared to existing unsupervised and supervised approaches. We then develop a simple methodology to pseudo-label the unlabeled samples using a new PU-specific clustering scheme; these pseudo-labels can then be used to train the final (positive vs. negative) classifier. Our method handily outperforms state-of-the-art PU methods over several standard PU benchmark datasets, while not requiring a-priori knowledge of any class prior (which is a common assumption in other PU methods). We also provide a simple theoretical analysis that motivates our methods.

cs.LG

Understanding Contrastive Representation Learning from Positive Unlabeled (PU) Data

Pretext Invariant Representation Learning (PIRL) followed by Supervised Fine-Tuning (SFT) has become a standard paradigm for learning with limited labels. We extend this approach to the Positive Unlabeled (PU) setting, where only a small set of labeled positives and a large unlabeled pool -- containing both positives and negatives are available. We study this problem under two regimes: (i) without access to the class prior, and (ii) when the prior is known or can be estimated. We introduce Positive Unlabeled Contrastive Learning (puCL), an unbiased and variance reducing contrastive objective that integrates weak supervision from labeled positives judiciously into the contrastive loss. When the class prior is known, we propose Positive Unlabeled InfoNCE (puNCE), a prior-aware extension that re-weights unlabeled samples as soft positive negative mixtures. For downstream classification, we develop a pseudo-labeling algorithm that leverages the structure of the learned embedding space via PU aware clustering. Our framework is supported by theory; offering bias-variance analysis, convergence insights, and generalization guarantees via augmentation concentration; and validated empirically across standard PU benchmarks, where it consistently outperforms existing methods, particularly in low-supervision regimes.

cs.LG

Tight-binding model subject to conditional resets at random times

We investigate the dynamics of a quantum system subjected to a time-dependent and conditional resetting protocol. Namely, we ask: what happens when the unitary evolution of the system is repeatedly interrupted at random time instants with an instantaneous reset to a specified set of reset configurations taking place with a probability that depends on the current configuration of the system at the instant of reset? Analyzing the protocol in the framework of the so-called tight-binding model describing the hopping of a quantum particle to nearest-neighbour sites in a one-dimensional open lattice, we obtain analytical results for the probability of finding the particle on the different sites of the lattice. We explore a variety of dynamical scenarios, including the one in which the resetting time intervals are sampled from an exponential as well as from a power-law distribution, and a set-up that includes a Floquet-type Hamiltonian involving an external periodic forcing. Under exponential resetting, and in both presence and absence of the external forcing, the system relaxes to a stationary state characterized by localization of the particle around the reset sites. The choice of the reset sites plays a defining role in dictating the relative probability of finding the particle at the reset sites as well as in determining the overall spatial profile of the site-occupation probability. Indeed, a simple choice can be engineered that makes the spatial profile highly asymmetric even when the bare dynamics does not involve the effect of any bias. Furthermore, analyzing the case of power-law resetting serves to demonstrate that the attainment of the stationary state in this quantum problem is not always evident and depends crucially on whether the distribution of reset time intervals has a finite or an infinite mean.

cond-mat.stat-mech

LDKP: A Dataset for Identifying Keyphrases from Long Scientific Documents

Identifying keyphrases (KPs) from text documents is a fundamental task in natural language processing and information retrieval. Vast majority of the benchmark datasets for this task are from the scientific domain containing only the document title and abstract information. This limits keyphrase extraction (KPE) and keyphrase generation (KPG) algorithms to identify keyphrases from human-written summaries that are often very short (approx 8 sentences). This presents three challenges for real-world applications: human-written summaries are unavailable for most documents, the documents are almost always long, and a high percentage of KPs are directly found beyond the limited context of title and abstract. Therefore, we release two extensive corpora mapping KPs of ~1.3M and ~100K scientific articles with their fully extracted text and additional metadata including publication venue, year, author, field of study, and citations for facilitating research on this real-world problem.

cs.CL

Faster Non-Convex Federated Learning via Global and Local Momentum

We propose \texttt{FedGLOMO}, a novel federated learning (FL) algorithm with an iteration complexity of $\mathcal{O}(ε^{-1.5})$ to converge to an $ε$-stationary point (i.e., $\mathbb{E}[\|\nabla f(\bm{x})\|^2] \leq ε$) for smooth non-convex functions -- under arbitrary client heterogeneity and compressed communication -- compared to the $\mathcal{O}(ε^{-2})$ complexity of most prior works. Our key algorithmic idea that enables achieving this improved complexity is based on the observation that the convergence in FL is hampered by two sources of high variance: (i) the global server aggregation step with multiple local updates, exacerbated by client heterogeneity, and (ii) the noise of the local client-level stochastic gradients. By modeling the server aggregation step as a generalized gradient-type update, we propose a variance-reducing momentum-based global update at the server, which when applied in conjunction with variance-reduced local updates at the clients, enables \texttt{FedGLOMO} to enjoy an improved convergence rate. Moreover, we derive our results under a novel and more realistic client-heterogeneity assumption which we verify empirically -- unlike prior assumptions that are hard to verify. Our experiments illustrate the intrinsic variance reduction effect of \texttt{FedGLOMO}, which implicitly suppresses client-drift in heterogeneous data distribution settings and promotes communication efficiency.

stat.ML

DISCO : efficient unsupervised decoding for discrete natural language problems via convex relaxation

In this paper we study test time decoding; an ubiquitous step in almost all sequential text generation task spanning across a wide array of natural language processing (NLP) problems. Our main contribution is to develop a continuous relaxation framework for the combinatorial NP-hard decoding problem and propose Disco - an efficient algorithm based on standard first order gradient based. We provide tight analysis and show that our proposed algorithm linearly converges to within $ε$ neighborhood of the optima. Finally, we perform preliminary experiments on the task of adversarial text generation and show superior performance of Disco over several popular decoding approaches.

cs.CL

Robust Training in High Dimensions via Block Coordinate Geometric Median Descent

Geometric median (\textsc{Gm}) is a classical method in statistics for achieving a robust estimation of the uncorrupted data; under gross corruption, it achieves the optimal breakdown point of 0.5. However, its computational complexity makes it infeasible for robustifying stochastic gradient descent (SGD) for high-dimensional optimization problems. In this paper, we show that by applying \textsc{Gm} to only a judiciously chosen block of coordinates at a time and using a memory mechanism, one can retain the breakdown point of 0.5 for smooth non-convex problems, with non-asymptotic convergence rates comparable to the SGD with \textsc{Gm}.

cs.LG

Alexa Conversations: An Extensible Data-driven Approach for Building Task-oriented Dialogue Systems

Traditional goal-oriented dialogue systems rely on various components such as natural language understanding, dialogue state tracking, policy learning and response generation. Training each component requires annotations which are hard to obtain for every new domain, limiting scalability of such systems. Similarly, rule-based dialogue systems require extensive writing and maintenance of rules and do not scale either. End-to-End dialogue systems, on the other hand, do not require module-specific annotations but need a large amount of data for training. To overcome these problems, in this demo, we present Alexa Conversations, a new approach for building goal-oriented dialogue systems that is scalable, extensible as well as data efficient. The components of this system are trained in a data-driven manner, but instead of collecting annotated conversations for training, we generate them using a novel dialogue simulator based on a few seed dialogues and specifications of APIs and entities provided by the developer. Our approach provides out-of-the-box support for natural conversational phenomena like entity sharing across turns or users changing their mind during conversation without requiring developers to provide any such dialogue flows. We exemplify our approach using a simple pizza ordering task and showcase its value in reducing the developer burden for creating a robust experience. Finally, we evaluate our system using a typical movie ticket booking task and show that the dialogue simulator is an essential component of the system that leads to over $50\%$ improvement in turn-level action signature prediction accuracy.

cs.CL

GupShup: An Annotated Corpus for Abstractive Summarization of Open-Domain Code-Switched Conversations

Code-switching is the communication phenomenon where speakers switch between different languages during a conversation. With the widespread adoption of conversational agents and chat platforms, code-switching has become an integral part of written conversations in many multi-lingual communities worldwide. This makes it essential to develop techniques for summarizing and understanding these conversations. Towards this objective, we introduce abstractive summarization of Hindi-English code-switched conversations and develop the first code-switched conversation summarization dataset - GupShup, which contains over 6,831 conversations in Hindi-English and their corresponding human-annotated summaries in English and Hindi-English. We present a detailed account of the entire data collection and annotation processes. We analyze the dataset using various code-switching statistics. We train state-of-the-art abstractive summarization models and report their performances using both automated metrics and human evaluation. Our results show that multi-lingual mBART and multi-view seq2seq models obtain the best performances on the new dataset

cs.CL