Searcharxiv⌕ Search

arXiv subjects

Shashank Gupta

Publications and source records attributed to Shashank Gupta.

At least 55 records · Page 3Linked to original sources

Safe Deployment for Counterfactual Learning to Rank with Exposure-Based Risk Minimization

Counterfactual learning to rank (CLTR) relies on exposure-based inverse propensity scoring (IPS), a LTR-specific adaptation of IPS to correct for position bias. While IPS can provide unbiased and consistent estimates, it often suffers from high variance. Especially when little click data is available, this variance can cause CLTR to learn sub-optimal ranking behavior. Consequently, existing CLTR methods bring significant risks with them, as naively deploying their models can result in very negative user experiences. We introduce a novel risk-aware CLTR method with theoretical guarantees for safe deployment. We apply a novel exposure-based concept of risk regularization to IPS estimation for LTR. Our risk regularization penalizes the mismatch between the ranking behavior of a learned model and a given safe model. Thereby, it ensures that learned ranking models stay close to a trusted model, when there is high uncertainty in IPS estimation, which greatly reduces the risks during deployment. Our experimental results demonstrate the efficacy of our proposed method, which is effective at avoiding initial periods of bad performance when little data is available, while also maintaining high performance at convergence. For the CLTR field, our novel exposure-based risk minimization method enables practitioners to adopt CLTR methods in a safer manner that mitigates many of the risks attached to previous methods.

cs.IR↗

Cryptanalysis of quantum permutation pad

Cryptanalysis increases the level of confidence in cryptographic algorithms. We analyze the security of a symmetric cryptographic algorithm - quantum permutation pad (QPP) [8]. We found the instances of ciphertext the same as plaintext even after the action of QPP with the probability 1/N when the entire set of permutation matrices of dimension N is used and with the probability 1/N^m when an incomplete set of m permutation matrices of dimension N are used. We visually show such instances in a cipher image created by QPP of 256 permutation matrices of different dimensions. For any practical usage of QPP, we recommend a set of 256 permutation matrices of a dimension more or equal to 2048.

cs.CR↗

Quantum contextuality provides communication complexity advantage

Despite the conceptual importance of contextuality in quantum mechanics, there is a hitherto limited number of applications requiring contextuality but not entanglement. Here, we show that for any quantum state and observables of sufficiently small dimensions producing contextuality, there exists a communication task with quantum advantage. Conversely, any quantum advantage in this task admits a proof of contextuality whenever an additional condition holds. We further show that given any set of observables allowing for quantum state-independent contextuality, there exists a class of communication tasks wherein the difference between classical and quantum communication complexities increases as the number of inputs grows. Finally, we show how to convert each of these communication tasks into a semi-device-independent protocol for quantum key distribution.

quant-ph↗

Device-Independent Quantum Key Distribution Using Random Quantum States

We Haar uniformly generate random states of various ranks and study their performance in an entanglement-based quantum key distribution (QKD) task. In particular, we analyze the efficacy of random two-qubit states in realizing device-independent (DI) QKD. We first find the normalized distribution of entanglement and Bell-nonlocality which are the key resource for DI-QKD for random states ranging from rank-1 to rank-4. The number of entangled as well as Bell-nonlocal states decreases as rank increases. We observe that decrease of the secure key rate is more pronounced in comparison to that of the quantum resource with increase in rank. We find that the pure state and Werner state provide the upper and lower bound, respectively, on the minimum secure key rate of all mixed two-qubit states possessing the same magnitude of entanglement under general as well as optimal collective attack strategies.

quant-ph↗

Genuine three qubit Einstein-Podolsky-Rosen steering under decoherence: Revealing hidden genuine steerability via pre-processing

The behaviour of genuine EPR steering of three qubit states under various environmental noises is investigated. In particular, we consider the two possible steering scenarios in the tripartite setting: (1 -> 2), where Alice demonstrates genuine steering to Bob-Charlie, and (2 -> 1), where Alice-Bob together demonstrates genuine steering to Charlie. In both these scenarios, we analyze the genuine steerability of the generalized Greenberger-Horne-Zeilinger (gGHZ) states or the W-class states under the action of noise modeled by amplitude damping (AD), phase flip (PF), bit flip (BF), and phase damping (PD) channels. In each case, we consider three different interactions with the noise depending upon the number of parties undergoing decoherence. We observed that the tendency to demonstrate genuine steering decreases as the number of parties undergoing decoherence increases from one to three. We have observed several instances where the genuine steerability of the state revives after collapsing if one keeps on increasing the damping. However, the hidden genuine steerability of a state cannot be revealed solely from the action of noise. So, the parties having a characterized subsystem, perform local pre-processing operations depending upon the steering scenario and the state shared with the dual intent of revealing hidden genuine steerability or enhancing it.

quant-ph↗

A universal whitening algorithm for commercial random number generators

Random number generators are imperfect due to manufacturing bias and technological imperfections. These imperfections are removed using post-processing algorithms that in general compress the data and do not work in every scenario. In this work, we present a universal whitening algorithm using n-qubit permutation matrices to remove the imperfections in commercial random number generators without compression. Specifically, we demonstrate the efficacy of our algorithm in several categories of random number generators and its comparison with cryptographic hash functions and block ciphers. We have achieved improvement in almost every randomness parameter evaluated using ENT randomness test suite. The modified random number files obtained after the application of our algorithm in the raw random data file pass the NIST SP 800-22 tests in both the cases: 1. The raw file does not pass all the tests. 2. The raw file also passes all the tests.

quant-ph↗

Quantum entropy expansion using n-qubit permutation matrices in Galois field

Random numbers are critical for any cryptographic application. However, the data that is flowing through the internet is not secure because of entropy deprived pseudo random number generators and unencrypted IoTs. In this work, we address the issue of lesser entropy of several data formats. Specifically, we use the large information space associated with the n-qubit permutation matrices to expand the entropy of any data without increasing the size of the data. We take English text with the entropy in the range 4 - 5 bits per byte. We manipulate the data using a set of n-qubit (n $\leq$ 10) permutation matrices and observe the expansion of the entropy in the manipulated data (to more than 7.9 bits per byte). We also observe similar behaviour with other data formats like image, audio etc. (n $\leq$ 15).

quant-ph↗

Sparsely Activated Mixture-of-Experts are Robust Multi-Task Learners

Traditional multi-task learning (MTL) methods use dense networks that use the same set of shared weights across several different tasks. This often creates interference where two or more tasks compete to pull model parameters in different directions. In this work, we study whether sparsely activated Mixture-of-Experts (MoE) improve multi-task learning by specializing some weights for learning shared representations and using the others for learning task-specific information. To this end, we devise task-aware gating functions to route examples from different tasks to specialized experts which share subsets of network weights conditioned on the task. This results in a sparsely activated multi-task model with a large number of parameters, but with the same computational cost as that of a dense model. We demonstrate such sparse networks to improve multi-task learning along three key dimensions: (i) transfer to low-resource tasks from related tasks in the training mixture; (ii) sample-efficient generalization to tasks not seen during training by making use of task-aware routing from seen related tasks; (iii) robustness to the addition of unrelated tasks by avoiding catastrophic forgetting of existing tasks.

cs.LG↗

Knowledge Infused Decoding

Pre-trained language models (LMs) have been shown to memorize a substantial amount of knowledge from the pre-training corpora; however, they are still limited in recalling factually correct knowledge given a certain context. Hence, they tend to suffer from counterfactual or hallucinatory generation when used in knowledge-intensive natural language generation (NLG) tasks. Recent remedies to this problem focus on modifying either the pre-training or task fine-tuning objectives to incorporate knowledge, which normally require additional costly training or architecture modification of LMs for practical applications. We present Knowledge Infused Decoding (KID) -- a novel decoding algorithm for generative LMs, which dynamically infuses external knowledge into each step of the LM decoding. Specifically, we maintain a local knowledge memory based on the current context, interacting with a dynamically created external knowledge trie, and continuously update the local memory as a knowledge-aware constraint to guide decoding via reinforcement learning. On six diverse knowledge-intensive NLG tasks, task-agnostic LMs (e.g., GPT-2 and BART) armed with KID outperform many task-optimized state-of-the-art models, and show particularly strong performance in few-shot scenarios over seven related knowledge-infusion techniques. Human evaluation confirms KID's ability to generate more relevant and factual language for the input context when compared with multiple baselines. Finally, KID also alleviates exposure bias and provides stable generation quality when generating longer sequences. Code for KID is available at https://github.com/microsoft/KID.

cs.CL↗

Constructive Feedback of Non-Markovianity on Resources in Random Quantum States

We explore the impact of non-Markovian channels on the quantum correlations (QCs) of Haar uniformly generated random two-qubit input states with different ranks -- either one of the qubits (single-sided) or both the qubits independently (double-sided) are passed through noisy channels. Under dephasing and depolarizing channels with varying non-Markovian strength, entanglement and quantum discord of the output states collapse and revive with the increase of noise. We find that in case of the depolarizing double-sided channel, both the QCs of random states show a higher number of revivals on average than that of the single-sided ones with a fixed non-Markovianity strength, irrespective of the rank of the states -- we call such a counter-intuitive event as a constructive feedback of non-Markovianity. Consequently, the average noise at which QCs of random states show first revival decreases with the increase of the strength of non-Markovian noise, thereby indicating the role of non-Markovian channels on the regenerations of QCs even in presence of a high amount of noise. However, we observe that non-Markovianity does not play any role to increase the robustness in random quantum states which can be measured by the mean value of critical noise at which quantum correlations first collapse. Moreover, we observe that the tendency of a state to show regeneration increases with the increase of average QCs of the random input states along with non-Markovianity.

quant-ph↗

"All-versus-nothing" proof of genuine tripartite steering and entanglement certification in the two-sided device-independent scenario

We consider the task of certification of genuine entanglement of tripartite states. For this purpose, we first present an "all-versus-nothing" proof of genuine tripartite Einstein-Podolsky-Rosen (EPR) steering by demonstrating the non-existence of a hybrid local hidden state (LHS) model in the tripartite network as a motivation to our main result. A full logical contradiction of the predictions of the hybrid LHS model with quantum mechanical outcome statistics for any three-qubit generalized Greenberger-Horne-Zeilinger (GGHZ) states and pure W-class states is shown. Using logical contradiction, we can distinguish between the GGHZ and W-class state in a two-sided device-independent (2SDI) steering scenario. We next formulate a 2SDI steering inequality which is a generalization of the fine-grained steering inequality (FGI) derived in \cite{PKM14} for the tripartite scenario. We show that the maximum quantum violation of this tripartite FGI can be used to certify genuine entanglement of three-qubit pure states.

quant-ph↗

Lung-Originated Tumor Segmentation from Computed Tomography Scan (LOTUS) Benchmark

Lung cancer is one of the deadliest cancers, and in part its effective diagnosis and treatment depend on the accurate delineation of the tumor. Human-centered segmentation, which is currently the most common approach, is subject to inter-observer variability, and is also time-consuming, considering the fact that only experts are capable of providing annotations. Automatic and semi-automatic tumor segmentation methods have recently shown promising results. However, as different researchers have validated their algorithms using various datasets and performance metrics, reliably evaluating these methods is still an open challenge. The goal of the Lung-Originated Tumor Segmentation from Computed Tomography Scan (LOTUS) Benchmark created through 2018 IEEE Video and Image Processing (VIP) Cup competition, is to provide a unique dataset and pre-defined metrics, so that different researchers can develop and evaluate their methods in a unified fashion. The 2018 VIP Cup started with a global engagement from 42 countries to access the competition data. At the registration stage, there were 129 members clustered into 28 teams from 10 countries, out of which 9 teams made it to the final stage and 6 teams successfully completed all the required tasks. In a nutshell, all the algorithms proposed during the competition, are based on deep learning models combined with a false positive reduction technique. Methods developed by the three finalists show promising results in tumor segmentation, however, more effort should be put into reducing the false positive rate. This competition manuscript presents an overview of the VIP-Cup challenge, along with the proposed algorithms and results.

eess.IV↗

Exploring Low-Cost Transformer Model Compression for Large-Scale Commercial Reply Suggestions

Fine-tuning pre-trained language models improves the quality of commercial reply suggestion systems, but at the cost of unsustainable training times. Popular training time reduction approaches are resource intensive, thus we explore low-cost model compression techniques like Layer Dropping and Layer Freezing. We demonstrate the efficacy of these techniques in large-data scenarios, enabling the training time reduction for a commercial email reply suggestion system by 42%, without affecting the model relevance or user engagement. We further study the robustness of these techniques to pre-trained model and dataset size ablation, and share several insights and recommendations for commercial applications.

cs.CL↗

Eavesdropping a Quantum Key Distribution network using sequential quantum unsharp measurement attacks

We investigate the possibility of eavesdropping on a quantum key distribution network by local sequential quantum unsharp measurement attacks by the eavesdropper. In particular, we consider a pure two-qubit state shared between two parties Alice and Bob, sharing quantum steerable correlations that form the one-sided device-independent quantum key distribution network. One qubit of the shared state is with Alice and the other one while going to the Bob's place is intercepted by multiple sequential eavesdroppers who perform quantum unsharp measurement attacks thus gaining some positive key rate while preserving the quantum steerable correlations for the Bob. In this way, Bob will also have a positive secret key rate although reduced. However, this reduction is not that sharp and can be perceived due to decoherence and imperfection of the measurement devices. At the end, we show that an unbounded number of eavesdroppers can also get secret information in some specific scenario.

quant-ph↗

Distillation of genuine tripartite Einstein-Podolsky-Rosen steering

We show that a perfectly genuine tripartite EPR steerable assemblage can be distilled from partially genuine tripartite EPR steerable assemblages. In particular, we consider two types of hybrid scenarios: one-sided device-independent (1SDI) scenario (where one observer is untrusted, and other two observers are trusted) and two-sided device-independent (2SDI) scenario (where two observers are untrusted, and one observer is trusted). In both the scenarios, we show distillation of perfectly genuine steerable assemblage of three-qubit Greenberger-Horne-Zeilinger (GHZ) states or three-qubit W states from many copies of initial partially genuine steerable assemblages of the corresponding states. In each of these cases, we demonstrate that at least one copy of a perfectly genuine steerable assemblage can be distilled with certainty from infinitely many copies of initial assemblages. In case of practical scenarios employing finite copies, we show that the efficiency of our distillation protocols reaches near perfect levels using only a few number of initial assemblages.

quant-ph↗

Low Resistance III-V Hetero-contacts to N-Ge

We experimentally study III-V/Ge heterostructure and demonstrate InGaAs hetero-contacts to n-Ge with a wide range of In % and achieve low contact resistivity ($ρ_C$) of $5\times10^{-8} Ω\cdot cm^2$ for Ge doping of $3 \times 10^{19} cm^{-3}$. This results from re-directing the charge neutrality level (CNL) near the conduction band and benefiting from low effective mass for high electron transmission. For the first time, we observe that the heterointerface presents no temperature dependence despite the two different conduction minimum valley locations of III-V ($Γ$-valley) and Ge (L-valley), which potentially stems from elastic trap-assisted tunneling through defect states at the interface generated by dislocations. The hetero-interface plays a dominant role in the overall $ρ_C$ below $\approx 1 \times 10^{-7} Ω\cdot cm^2$, which can be further improved with large active dopant concentration in Ge by co-doping.

cond-mat.mtrl-sci↗

A Survey on Serverless Computing

The Internet is responsible for accelerating growth in several fields such as digital media, healthcare, the military. Furthermore, the Internet was founded on the principle of allowing clients to communicating with servers. However, serverless computing is one such field that tries to break free from this paradigm. Event-driven compute services allow users to build more agile applications using capacity provisioning and a pay-for-value billing model. This paper provides a formal account of the research contributions in the field of Serverless computing.

cs.DC↗

Feature Extraction Functions for Neural Logic Rule Learning

Combining symbolic human knowledge with neural networks provides a rule-based ante-hoc explanation of the output. In this paper, we propose feature extracting functions for integrating human knowledge abstracted as logic rules into the predictive behavior of a neural network. These functions are embodied as programming functions, which represent the applicable domain knowledge as a set of logical instructions and provide a modified distribution of independent features on input data. Unlike other existing neural logic approaches, the programmatic nature of these functions implies that they do not require any kind of special mathematical encoding, which makes our method very general and flexible in nature. We illustrate the performance of our approach for sentiment classification and compare our results to those obtained using two baselines.

cs.AI↗