SearcharxivSearch

arXiv subjects

Ynes Ineza

Publications and source records attributed to Ynes Ineza.

4 recordsLinked to original sources

Intermittent File Encryption in Ransomware: Measurement, Modeling, and Detection

File-encrypting ransomware increasingly employs intermittent encryption techniques, encrypting only parts of files to evade classical detection methods.This paper provides a systematic empirical characterization of byte-level statistics under intermittent encryption across common file types, establishing a baseline for how partial encryption reshapes data structure. Guided by these measurements, we model intermittent encryption as a convex mixture of ciphertext and cleartext and, via a classical KL-divergence bound, derive file-type-specific detectability limits for histogram-based detectors. Leveraging these insights, we evaluate convolutional neural network (CNN) detectors trained on realistic intermittent-encryption configurations from leading ransomware families. Our findings show that localized, chunk-level CNNs consistently outperform whole-file analysis, highlighting a practical, robust baseline for future detection systems.

cs.CR

Beyond the Voice: Inertial Sensing of Mouth Motion for High Security Speech Verification

Voice interfaces are increasingly used in high-stakes domains such as mobile banking, smart-home security, and hands-free healthcare. Meanwhile, modern generative models have made high-quality voice forgeries inexpensive and easy to create, eroding confidence in voice authentication alone. To strengthen protection against such attacks, we present a second authentication factor that combines acoustic evidence with the unique motion patterns of a speaker's lower face. By placing lightweight inertial sensors around the mouth to capture mouth opening and evolving lower-facial geometry, our system records a distinct motion signature with strong discriminative power across individuals. We built a prototype and recruited 43 participants to evaluate the system under four conditions: seated, walking on level ground, walking on stairs, and speaking with different language backgrounds (native vs. non-native English). Across all scenarios, our approach consistently achieved a median equal-error rate (EER) of 0.01 or lower, indicating that mouth-movement data remain robust under variations in gait, posture, and spoken language. We discuss specific use cases where this second line of defense could provide tangible security benefits to voice authentication systems.

cs.CR

Guarding Your Conversations: Privacy Gatekeepers for Secure Interactions with Cloud-Based AI Models

The interactive nature of Large Language Models (LLMs), which closely track user data and context, has prompted users to share personal and private information in unprecedented ways. Even when users opt out of allowing their data to be used for training, these privacy settings offer limited protection when LLM providers operate in jurisdictions with weak privacy laws, invasive government surveillance, or poor data security practices. In such cases, the risk of sensitive information, including Personally Identifiable Information (PII), being mishandled or exposed remains high. To address this, we propose the concept of an "LLM gatekeeper", a lightweight, locally run model that filters out sensitive information from user queries before they are sent to the potentially untrustworthy, though highly capable, cloud-based LLM. Through experiments with human subjects, we demonstrate that this dual-model approach introduces minimal overhead while significantly enhancing user privacy, without compromising the quality of LLM responses.

cs.CR

Unmasking the giant: A comprehensive evaluation of ChatGPT's proficiency in coding algorithms and data structures

The transformative influence of Large Language Models (LLMs) is profoundly reshaping the Artificial Intelligence (AI) technology domain. Notably, ChatGPT distinguishes itself within these models, demonstrating remarkable performance in multi-turn conversations and exhibiting code proficiency across an array of languages. In this paper, we carry out a comprehensive evaluation of ChatGPT's coding capabilities based on what is to date the largest catalog of coding challenges. Our focus is on the python programming language and problems centered on data structures and algorithms, two topics at the very foundations of Computer Science. We evaluate ChatGPT for its ability to generate correct solutions to the problems fed to it, its code quality, and nature of run-time errors thrown by its code. Where ChatGPT code successfully executes, but fails to solve the problem at hand, we look into patterns in the test cases passed in order to gain some insights into how wrong ChatGPT code is in these kinds of situations. To infer whether ChatGPT might have directly memorized some of the data that was used to train it, we methodically design an experiment to investigate this phenomena. Making comparisons with human performance whenever feasible, we investigate all the above questions from the context of both its underlying learning models (GPT-3.5 and GPT-4), on a vast array sub-topics within the main topics, and on problems having varying degrees of difficulty.

cs.SE