SearcharxivSearch

arXiv subjects

Ayushi Mishra

Publications and source records attributed to Ayushi Mishra.

7 recordsLinked to original sources

Mechanistic Interpretability of Brain-to-Speech Models Across Speech Modes

Brain-to-speech decoding models demonstrate robust performance in vocalized, mimed, and imagined speech; yet, the fundamental mechanisms via which these models capture and transmit information across different speech modalities are less explored. In this work, we use mechanistic interpretability to causally investigate the internal representations of a neural speech decoder. We perform cross-mode activation patching of internal activations across speech modes, and use tri-modal interpolation to examine whether speech representations vary discretely or continuously. We use coarse-to-fine causal tracing and causal scrubbing to find localized causal structure, allowing us to find internal subspaces that are sufficient for cross-mode transfer. In order to determine how finely distributed these effects are within layers, we perform neuron-level activation patching. We discover that small but not distributed subsets of neurons, rather than isolated units, affect the cross-mode transfer. Our results show that speech modes lie on a shared continuous causal manifold, and cross-mode transfer is mediated by compact, layer-specific subspaces rather than diffuse activity. Together, our findings give a causal explanation for how speech modality information is organized and used in brain-to-speech decoding models, revealing hierarchical and direction-dependent representational structure across speech modes.

cs.LG

Spatial Audio Processing with Large Language Model on Wearable Devices

Integrating spatial context into large language models (LLMs) has the potential to revolutionize human-computer interaction, particularly in wearable devices. In this work, we present a novel system architecture that incorporates spatial speech understanding into LLMs, enabling contextually aware and adaptive applications for wearable technologies. Our approach leverages microstructure-based spatial sensing to extract precise Direction of Arrival (DoA) information using a monaural microphone. To address the lack of existing dataset for microstructure-assisted speech recordings, we synthetically create a dataset called OmniTalk by using the LibriSpeech dataset. This spatial information is fused with linguistic embeddings from OpenAI's Whisper model, allowing each modality to learn complementary contextual representations. The fused embeddings are aligned with the input space of LLaMA-3.2 3B model and fine-tuned with lightweight adaptation technique LoRA to optimize for on-device processing. SING supports spatially-aware automatic speech recognition (ASR), achieving a mean error of $25.72^\circ$-a substantial improvement compared to the 88.52$^\circ$ median error in existing work-with a word error rate (WER) of 5.3. SING also supports soundscaping, for example, inference how many people were talking and their directions, with up to 5 people and a median DoA error of 16$^\circ$. Our system demonstrates superior performance in spatial speech understanding while addressing the challenges of power efficiency, privacy, and hardware constraints, paving the way for advanced applications in augmented reality, accessibility, and immersive experiences.

cs.SD

MalDicom: A Memory Forensic Framework for Detecting Malicious Payload in DICOM Files

Digital Imaging and Communication System (DICOM) is widely used throughout the public health sector for portability in medical imaging. However, these DICOM files have vulnerabilities present in the preamble section. Successful exploitation of these vulnerabilities can allow attackers to embed executable codes in the 128-Byte preamble of DICOM files. Embedding the malicious executable will not interfere with the readability or functionality of DICOM imagery. However, it will affect the underline system silently upon viewing these files. This paper shows the infiltration of Windows malware executables into DICOM files. On viewing the files, the malicious DICOM will get executed and eventually infect the entire hospital network through the radiologist's workstation. The code injection process of executing malware in DICOM files affects the hospital networks and workstations' memory. Memory forensics for the infected radiologist's workstation is crucial as it can detect which malware disrupts the hospital environment, and future detection methods can be deployed. In this paper, we consider the machine learning (ML) algorithms to conduct memory forensics on three memory dump categories: Trojan, Spyware, and Ransomware, taken from the CIC-MalMem-2022 dataset. We obtain the highest accuracy of 75% with the Random Forest model. For estimating the feature importance for ML model prediction, we leveraged the concept of Shapley values.

cs.CR

MediHunt: A Network Forensics Framework for Medical IoT Devices

The Medical Internet of Things (MIoT) has enabled small, ubiquitous medical devices to communicate with each other to facilitate interconnected healthcare delivery. These devices interact using communication protocols like MQTT, Bluetooth, and Wi-Fi. However, as MIoT devices proliferate, these networked devices are vulnerable to cyber-attacks. This paper focuses on the vulnerabilities present in the Message Queuing Telemetry and Transport (MQTT) protocol. The MQTT protocol is prone to cyber-attacks that can harm the system's functionality. The memory-constrained MIoT devices enforce a limitation on storing all data logs that are required for comprehensive network forensics. This paper solves the data log availability challenge by detecting the attack in real-time and storing the corresponding logs for further analysis with the proposed network forensics framework: MediHunt. Machine learning (ML) techniques are the most real safeguard against cyber-attacks. However, these models require a specific dataset that covers diverse attacks on the MQTT-based IoT system for training. The currently available datasets do not encompass a variety of applications and TCP layer attacks. To address this issue, we leveraged the usage of a flow-based dataset containing flow data for TCP/IP layer and application layer attacks. Six different ML models are trained with the generated dataset to evaluate the effectiveness of the MediHunt framework in detecting real-time attacks. F1 scores and detection accuracy exceeded 0.99 for the proposed MediHunt framework with our custom dataset.

cs.CR

Gravitational Wave Formation from the Collapse of Dark Energy Field Configurations

Dark Energy is the dominant component of the energy density of the universe. In a previous paper, we have shown that the collapse of dark energy fields leads to the formation of Super Massive Black Holes with masses comparable to the masses of Black Holes at the centers of galaxies. Thus it becomes a pressing issue to investigate the other physical consequences of the collapse of Dark Energy fields. Given that the primary interactions of Dark Energy fields with the rest of the Universe are gravitational, it is particularly interesting to investigate the gravitational wave signals emitted during the process of the collapse of Dark Energy fields. This is the focus of the current work described in this paper. We describe and use the 3+1 BSSN formalism to follow the evolution of the dark energy fields coupled with gravity and to extract the gravitational wave signals. Finally, we describe the results of our numerical computations and the gravitational wave signals produced as a result of the collapse of the dark energy fields.

physics.gen-ph

Active modulation of surfactant driven flow instabilities by swarming bacteria

Models based on surfactant driven instabilities have been employed to describe pattern formation by swarming bacteria. However, by definition, such models cannot account for the effect of bacterial sensing and decision making. Here we present a more complete model for bacterial pattern formation which accounts for these effects by coupling active bacterial motility to the passive fluid dynamics. We experimentally identify behaviours which cannot be captured by previous models based on passive population dispersal and show that a more accurate description is provided by our model. It is seen that the coupling of bacterial motility to the fluid dynamics significantly alters the phase space of surfactant driven pattern formation. We also show that our formalism is applicable across bacterial species.

physics.bio-ph

Spatial Awareness of a Bacterial Swarm

Bacteria are perhaps the simplest living systems capable of complex behaviour involving sensing and coherent, collective behaviour an example of which is the phenomena of swarming on agar surfaces. Two fundamental questions in bacterial swarming is how the information gathered by individual members of the swarm is shared across the swarm leading to coordinated swarm behaviour and what specific advantages does membership of the swarm provide its members in learning about their environment. In this article, we show a remarkable example of the collective advantage of a bacterial swarm which enables it to sense inert obstacles along its path. Agent based computational model of swarming revealed that independent individual behaviour in response to a two-component signalling mechanism could produce such behaviour. This is striking because independent individual behaviour without any explicit communication between agents was found to be sufficient for the swarm to effectively compute the gradient of signalling molecule concentration across the swarm and respond to it.

physics.bio-ph