SearcharxivSearch

arXiv subjects

Sheryl Mathew

Publications and source records attributed to Sheryl Mathew.

8 recordsLinked to original sources

Battery Locality Is Necessary in Noncommuting Quantum Charging Bounds

Bounds on the charging power of quantum batteries in direct charging protocols are often interpreted as charging-side constraints. Although the locality of the charging Hamiltonian is known to restrict attainable power in certain protocols, whether battery locality is itself operationally necessary has remained unresolved. Here we show that it is necessary for spin batteries whose interactions do not pairwise commute. We construct commuting and noncommuting high-locality batteries with identical interaction supports and local energy scale, driven by the same charging. The commuting battery saturates the battery-locality-independent bound, whereas the noncommuting battery exceeds it by a parametrically growing factor, demonstrating that this bound cannot apply generally. The enhanced norm of the commutator is dynamically attained from a suitable joint state supported entirely in the battery ground-energy sector, although this state may be correlated with the auxiliary fermionic degrees of freedom. Our results establish battery locality as an operational resource enabled by noncommuting battery interactions.

quant-ph

Super-Extensive Charging Power in the Absence of Global Operations

Quantum batteries have emerged as a platform for investigating whether quantum effects can accelerate energy storage beyond classical limits. Although a variety of charging schemes have reported signatures of quantum advantage, the fundamental physical requirements for achieving superextensive charging power remain insufficiently understood. Here, we show that, in addition to Hamiltonian locality, a key structural property, g-extensiveness, quantifying the distribution of interaction energy across lattice sites places a fundamental bound on charging performance in spin-lattice models. We prove that superextensive power scaling is possible only when the interaction-energy distribution becomes increasingly nonuniform, with the maximal local weight growing with system size. This criterion explains why many previously studied protocols fail to exhibit superextensive power, even when the Hamiltonians involve large participation numbers. We further demonstrate that this condition is realized in an experimentally relevant interacting model, where, despite fixed interaction order, the charging power scales superextensively. Our results establish g-extensiveness as a necessary resource for quantum advantage in direct-charging protocols and provide a systematic framework for identifying and engineering physically feasible quantum batteries capable of outperforming classical counterparts in charging power.

quant-ph

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning

In reinforcement learning with human feedback (RLHF), reward models can efficiently learn and amplify latent biases within multimodal datasets, which can lead to imperfect policy optimization through flawed reward signals and decreased fairness. Bias mitigation studies have often applied passive constraints, which can fail under causal confounding. Here, we present a counterfactual reward model that introduces causal inference with multimodal representation learning to provide an unsupervised, bias-resilient reward signal. The heart of our contribution is the Counterfactual Trust Score, an aggregated score consisting of four components: (1) counterfactual shifts that decompose political framing bias from topical bias; (2) reconstruction uncertainty during counterfactual perturbations; (3) demonstrable violations of fairness rules for each protected attribute; and (4) temporal reward shifts aligned with dynamic trust measures. We evaluated the framework on a multimodal fake versus true news dataset, which exhibits framing bias, class imbalance, and distributional drift. Following methodologies similar to unsupervised drift detection from representation-based distances [1] and temporal robustness benchmarking in language models [2], we also inject synthetic bias across sequential batches to test robustness. The resulting system achieved an accuracy of 89.12% in fake news detection, outperforming the baseline reward models. More importantly, it reduced spurious correlations and unfair reinforcement signals. This pipeline outlines a robust and interpretable approach to fairness-aware RLHF, offering tunable bias reduction thresholds and increasing reliability in dynamic real-time policy making.

cs.LG

VinaBench: Benchmark for Faithful and Consistent Visual Narratives

Visual narrative generation transforms textual narratives into sequences of images illustrating the content of the text. However, generating visual narratives that are faithful to the input text and self-consistent across generated images remains an open challenge, due to the lack of knowledge constraints used for planning the stories. In this work, we propose a new benchmark, VinaBench, to address this challenge. Our benchmark annotates the underlying commonsense and discourse constraints in visual narrative samples, offering systematic scaffolds for learning the implicit strategies of visual storytelling. Based on the incorporated narrative constraints, we further propose novel metrics to closely evaluate the consistency of generated narrative images and the alignment of generations with the input textual narrative. Our results across three generative vision models demonstrate that learning with VinaBench's knowledge constraints effectively improves the faithfulness and cohesion of generated visual narratives.

cs.CV

Multimodal Fusion with LLMs for Engagement Prediction in Natural Conversation

Over the past decade, wearable computing devices (``smart glasses'') have undergone remarkable advancements in sensor technology, design, and processing power, ushering in a new era of opportunity for high-density human behavior data. Equipped with wearable cameras, these glasses offer a unique opportunity to analyze non-verbal behavior in natural settings as individuals interact. Our focus lies in predicting engagement in dyadic interactions by scrutinizing verbal and non-verbal cues, aiming to detect signs of disinterest or confusion. Leveraging such analyses may revolutionize our understanding of human communication, foster more effective collaboration in professional environments, provide better mental health support through empathetic virtual interactions, and enhance accessibility for those with communication barriers. In this work, we collect a dataset featuring 34 participants engaged in casual dyadic conversations, each providing self-reported engagement ratings at the end of each conversation. We introduce a novel fusion strategy using Large Language Models (LLMs) to integrate multiple behavior modalities into a ``multimodal transcript'' that can be processed by an LLM for behavioral reasoning tasks. Remarkably, this method achieves performance comparable to established fusion techniques even in its preliminary implementation, indicating strong potential for further research and optimization. This fusion method is one of the first to approach ``reasoning'' about real-world human behavior through a language model. Smart glasses provide us the ability to unobtrusively gather high-density multimodal data on human behavior, paving the way for new approaches to understanding and improving human communication with the potential for important societal benefits. The features and data collected during the studies will be made publicly available to promote further research.

cs.AI

Fault-tolerant hyperbolic Floquet quantum error correcting codes

A central goal in quantum error correction is to reduce the overhead of fault-tolerant quantum computing by increasing noise thresholds and reducing the number of physical qubits required to sustain a logical qubit. We introduce a potential path towards this goal based on a family of dynamically generated quantum error correcting codes that we call "hyperbolic Floquet codes.'' These codes are defined by a specific sequence of non-commuting two-body measurements arranged periodically in time that stabilize a topological code on a hyperbolic manifold with negative curvature. We focus on a family of lattices for $n$ qubits that, according to our prescription that defines the code, provably achieve a finite encoding rate $(1/8+2/n)$ while still requiring only two-body measurements. Similar to hyperbolic surface codes, the distance of the code at each time-step scales at most logarithmically in $n$. The family of lattices we choose indicates that this scaling is achievable in practice. We develop and benchmark an efficient matching-based decoder that provides evidence of a threshold near 0.1% in a phenomenological noise model and 0.25% in an entangling measurements noise model. Utilizing weight-two check operators and a qubit connectivity of 3, one of our hyperbolic Floquet codes uses 400 physical qubits to encode 52 logical qubits with a code distance of 8, i.e., it is a $[[400,52,8]]$ code. At small error rates, comparable logical error suppression to this code requires 5x as many physical qubits (1924) when using the honeycomb Floquet code with the same noise model and decoder.

quant-ph

Difference-Masking: Choosing What to Mask in Continued Pretraining

The self-supervised objective of masking-and-predicting has led to promising performance gains on a variety of downstream tasks. However, while most approaches randomly mask tokens, there is strong intuition that deciding what to mask can substantially improve learning outcomes. We investigate this in continued pretraining setting in which pretrained models continue to pretrain on domain-specific data before performing some downstream task. We introduce Difference-Masking, a masking strategy that automatically chooses what to mask during continued pretraining by considering what makes a task domain different from the pretraining domain. Empirically, we find that Difference-Masking outperforms baselines on continued pretraining settings across four diverse language-only and multimodal video tasks.

cs.LG

Human activity recognition using deep learning approaches and single frame cnn and convolutional lstm

Human activity recognition is one of the most important tasks in computer vision and has proved useful in different fields such as healthcare, sports training and security. There are a number of approaches that have been explored to solve this task, some of them involving sensor data, and some involving video data. In this paper, we aim to explore two deep learning-based approaches, namely single frame Convolutional Neural Networks (CNNs) and convolutional Long Short-Term Memory to recognise human actions from videos. Using a convolutional neural networks-based method is advantageous as CNNs can extract features automatically and Long Short-Term Memory networks are great when it comes to working on sequence data such as video. The two models were trained and evaluated on a benchmark action recognition dataset, UCF50, and another dataset that was created for the experimentation. Though both models exhibit good accuracies, the single frame CNN model outperforms the Convolutional LSTM model by having an accuracy of 99.8% with the UCF50 dataset.

cs.CV