Searcharxiv⌕ Search

arXiv subjects

Gabrielle De Micheli

Publications and source records attributed to Gabrielle De Micheli.

5 recordsLinked to original sources

Lightweight, Practical Encrypted Face Recognition with GPU Support

Face recognition typically operates in a client-server setting, where the client extracts a compact face embedding and the server performs similarity search over a template database. Since facial data is highly sensitive, this raises significant privacy concerns. Fully homomorphic encryption (FHE) addresses these concerns by enabling end-to-end encrypted similarity search. However, existing FHE-based protocols are computationally costly and, especially, impose high memory overhead due to large rotation-key sets and bandwidth-bound homomorphic operations. Building on prior work, HyDia (PoPETS 2025), we introduce algorithmic and system-level improvements targeting real-world deployment with resource-constrained (edge) clients. First, we propose BSGS-Diagonal, a fast and memory-efficient similarity computation algorithm that applies a Baby-Step/Giant-Step strategy with precomputed rotations reused across consecutive matrix--vector products. This yields a 91% reduction in rotation keys (~14GB less client memory) and cuts peak server-side CPU RAM usage from over 33GB to 11GB for databases up to 1M entries, with runtime improvements of up to 1.57x for membership verification and 1.43x for identification. Second, we introduce GPU-optimized similarity computation kernels, including an efficient homomorphic Chebyshev evaluator built upon FIDESlib (ISPASS 2025), a CKKS-level GPU library based on OpenFHE. Rather than offloading individual CKKS primitives, our integrated kernels fuse operations to avoid repeated CPU--GPU ciphertext movement and costly FIDESlib/OpenFHE data-structure conversions. Our HyDia and BSGS GPU results achieve up to 9x and 21x speedups on single GPU (and up to 287x and 211x using multi-GPUs), respectively, enabling sub-second encrypted face recognition for databases up to 2^16 entries (or 2^19 entries in a multi-GPU setting), while further reducing host memory usage.

cs.CR↗

Understanding Calibration and Truncation Error Propagation in Training-Free Low-Rank Compression for LLMs

Training-free low-rank compression frameworks have been gaining prominence for LLM compression given their effectiveness in reducing model parameter count while maintaining task-level accuracy. However, existing SOTA frameworks share two key limitations: (1) residual errors in calibration data activations accumulate across layers during compression, causing misalignment between representations simulated at compression time and those experienced at inference; (2) the assumption that layer importance distribution is preserved post-compression does not hold. Together, these two effects introduce misalignment in the compression process in relation to the deployed model. We study these effects and propose a simple, training-free methodology compatible with existing frameworks to mitigate them, comprising: (1) Layer-by-Layer Compression with Calibration Correction; (2) Iterative Compression with Rank Allocation Correction. Implemented atop an existing SOTA decomposition framework, and evaluated on Llama and Qwen3 models across various benchmarks and compression rates, our approach demonstrates up to ~1-2.5 accuracy point improvements over per-weight and joint decomposition baselines on zero-shot tasks.

cs.AI↗

HE-LRM: Encrypted Deep Learning Recommendation Models using Fully Homomorphic Encryption

Fully Homomorphic Encryption (FHE) enables computation directly on encrypted data and privacy-preserving neural inference in the cloud. Existing solutions focus on models with dense inputs (e.g., CNNs and MLPs). Recommendation models (e.g., DLRM) pose a different challenge: sparse categorical inputs require private lookups into large embedding tables, which must be implemented using FHE's restrictive operators. Naive lookups incur significant communication and memory costs; prior work proposes compressing embedding tables at the expense of introducing large server-side compute costs (i.e., indicator function) and revealing embedding-table structure. We present HE-LRM, a performance optimized solution for executing recommendation with FHE. First, we develop an embedding compression technique using client-side digit decomposition that achieves 56$\times$ speedup over the state-of-the-art. Next, we propose a multi-embedding packing strategy that enables ciphertext SIMD-parallel lookups across multiple tables. We integrate HE-LRM into the open-source Orion FHE framework to demonstrate end-to-end encrypted DLRM inference. We evaluate HE-LRM on UCI (health prediction) and Criteo (click prediction), achieving inference latencies of 24 seconds on UCI and 228 to 489 seconds, respectively, on a single-threaded CPU. Finally, we show how GPU and ASIC FHE acceleration can reduce end-to-end latencies to seconds and even sub-seconds. Our code can be found at https://github.com/baahl-nyu/orion/tree/criteo-helrm.

cs.CR↗

PAC to the Future: Zero-Knowledge Proofs of PAC Private Systems

Privacy concerns in machine learning systems have grown significantly with the increasing reliance on sensitive user data for training large-scale models. This paper introduces a novel framework combining Probably Approximately Correct (PAC) Privacy with zero-knowledge proofs (ZKPs) to provide verifiable privacy guarantees in trustless computing environments. Our approach addresses the limitations of traditional privacy-preserving techniques by enabling users to verify both the correctness of computations and the proper application of privacy-preserving noise, particularly in cloud-based systems. We leverage non-interactive ZKP schemes to generate proofs that attest to the correct implementation of PAC privacy mechanisms while maintaining the confidentiality of proprietary systems. Our results demonstrate the feasibility of achieving verifiable PAC privacy in outsourced computation, offering a practical solution for maintaining trust in privacy-preserving machine learning and database systems while ensuring computational integrity.

cs.CR↗

CARVQ: Corrective Adaptor with Group Residual Vector Quantization for LLM Embedding Compression

Large Language Models (LLMs) typically rely on a large number of parameters for token embedding, leading to substantial storage requirements and memory footprints. In particular, LLMs deployed on edge devices are memory-bound, and reducing the memory footprint by compressing the embedding layer not only frees up the memory bandwidth but also speeds up inference. To address this, we introduce CARVQ, a post-training novel Corrective Adaptor combined with group Residual Vector Quantization. CARVQ relies on the composition of both linear and non-linear maps and mimics the original model embedding to compress to approximately 1.6 bits without requiring specialized hardware to support lower-bit storage. We test our method on pre-trained LLMs such as LLaMA-3.2-1B, LLaMA-3.2-3B, LLaMA-3.2-3B-Instruct, LLaMA-3.1-8B, Qwen2.5-7B, Qwen2.5-Math-7B and Phi-4, evaluating on common generative, discriminative, math and reasoning tasks. We show that in most cases, CARVQ can achieve lower average bitwidth-per-parameter while maintaining reasonable perplexity and accuracy compared to scalar quantization. Our contributions include a novel compression technique that is compatible with state-of-the-art transformer quantization methods and can be seamlessly integrated into any hardware supporting 4-bit memory to reduce the model's memory footprint in memory-constrained devices. This work demonstrates a crucial step toward the efficient deployment of LLMs on edge devices.

cs.LG↗