Searcharxiv⌕ Search

arXiv subjects

Nishant Kumar

Publications and source records attributed to Nishant Kumar.

At least 37 records · Page 2Linked to original sources

Normalizing Flow based Feature Synthesis for Outlier-Aware Object Detection

Real-world deployment of reliable object detectors is crucial for applications such as autonomous driving. However, general-purpose object detectors like Faster R-CNN are prone to providing overconfident predictions for outlier objects. Recent outlier-aware object detection approaches estimate the density of instance-wide features with class-conditional Gaussians and train on synthesized outlier features from their low-likelihood regions. However, this strategy does not guarantee that the synthesized outlier features will have a low likelihood according to the other class-conditional Gaussians. We propose a novel outlier-aware object detection framework that distinguishes outliers from inlier objects by learning the joint data distribution of all inlier classes with an invertible normalizing flow. The appropriate sampling of the flow model ensures that the synthesized outliers have a lower likelihood than inliers of all object classes, thereby modeling a better decision boundary between inlier and outlier objects. Our approach significantly outperforms the state-of-the-art for outlier-aware object detection on both image and video datasets. Code available at https://github.com/nish03/FFS

cs.CV↗

Hematite $α-Fe_{2}O_{3}(0001)$ in top and side view: resolving long-standing controversies about its surface structure

Hematite $α-Fe_{2}O_{3}(0001)$ is the most-investigated iron oxide model system in photo and electrocatalytic research. The rich chemistry of Fe and O allows for many bulk and surface transformations, but their control is challenging. This has led to controversies regarding the structure of the topmost layers. This comprehensive study combines surface methods (nc-AFM, STM, LEED, and XPS) complemented by structural and chemical analysis of the near-surface bulk (HRTEM and EELS). The results show that a compact 2D layer constitutes the topmost surface of $α-Fe_{2}O_{3}(0001)$; it is locally corrugated due to the mismatch with the bulk. Assessing the influence of naturally-occurring impurities shows that these can force the formation of surface phases that are not stable on pure samples. Impurities can also cause the formation of ill-defined inclusions in the subsurface and modify the oxidation phase diagram of hematite. The results provide a significant step forward in determining the hematite surface structure that is crucial for accurately modeling catalytic reactions. Combining surface and cross-sectional imaging provided the full view that is essential for understanding the evolution of the near-surface region of oxide surfaces under oxidative conditions.

cond-mat.mtrl-sci↗

A Direct and New Construction of Near-Optimal Multiple ZCZ Sequence Sets

In this paper, for the first time, we present a direct and new construction of multiple zero-correlation zone (ZCZ) sequence sets with inter-set zero-cross correlation zone (ZCCZ) from generalised Boolean function. Tang \emph{et al.} in their 2010 paper, proposed an open problem to construct $N$ binary ZCZ sequence sets such that each of these ZCZ sequence sets is optimal and if the union of these $N$ sets is taken then that union is again an optimal ZCZ sequence set. The proposed construction partially settles this open problem by presenting a construction of optimal ZCZ sequence sets such that their union is a near-optimal ZCZ sequence set. Further, the performance parameter of each binary ZCZ sequence set in the proposed construction is $1$ and tends to $1$ for their union. The proposed construction is presented by a two-layer graphical representation and compared with the existing state-of-the-art. Finally, novel multi-cluster quasi synchronous-code division multiple access (QS-CDMA) system model is provided by using the proposed multiple ZCZ sequence sets.

cs.IT↗

HAMMER: Multi-Level Coordination of Reinforcement Learning Agents via Learned Messaging

Cooperative multi-agent reinforcement learning (MARL) has achieved significant results, most notably by leveraging the representation-learning abilities of deep neural networks. However, large centralized approaches quickly become infeasible as the number of agents scale, and fully decentralized approaches can miss important opportunities for information sharing and coordination. Furthermore, not all agents are equal -- in some cases, individual agents may not even have the ability to send communication to other agents or explicitly model other agents. This paper considers the case where there is a single, powerful, \emph{central agent} that can observe the entire observation space, and there are multiple, low-powered \emph{local agents} that can only receive local observations and are not able to communicate with each other. The central agent's job is to learn what message needs to be sent to different local agents based on the global observations, not by centrally solving the entire problem and sending action commands, but by determining what additional information an individual agent should receive so that it can make a better decision. In this work we present our MARL algorithm \algo, describe where it would be most applicable, and implement it in the cooperative navigation and multi-agent walker domains. Empirical results show that 1) learned communication does indeed improve system performance, 2) results generalize to heterogeneous local agents, and 3) results generalize to different reward structures.

cs.MA↗

Bulk Modulus along Jamming Transition Lines of Bidisperse Granular Packings

We present 3D DEM simulations of bidisperse granular packings to investigate their jamming densities, $ϕ_J$, and dimensionless bulk moduli, $K$, as a function of the size ratio, $δ$, and the concentration of small particles, $X_{\mathrm S}$. We determine the partial and total bulk moduli for each packing and report the jamming transition diagram, i.e., the density or volume fraction marking both the first and second transitions of the system. At a large enough size difference, e.g., $δ\le 0.22$, $X^{*}_{\mathrm S}$ divides the diagram with most small particles either non-jammed or jammed jointly with large ones. We find that the bulk modulus $K$ jumps at $X^{*}_{\mathrm S}(δ= 0.15) \approx 0.21$, at the maximum jamming density, where both particle species mix most efficiently, while for $X_{\mathrm S} < X^{*}_{\mathrm S}$ $K$ is decoupled in two scenarios as a result of the first and second jamming transition. Along the second transition, $K$ rises relative to the values found at the first transition, however, is still small compared to $K$ at $X^{*}_{\mathrm S}$. While the first transition is sharp, the second is smooth, carried by small-large interactions, while the small-small contacts display a transition. This demonstrates that for low enough $δ$ and $X_{\mathrm S}$, the jamming of small particles indeed impacts the internal resistance of the system. Our new results will allow tuning the bulk modulus $K$ or other properties, such as the wave speed, by choosing specific sizes and concentrations based on a better understanding of whether small particles contribute to the jammed structure or not, and how the micromechanical structure behaves at either transition.

cond-mat.soft↗

Enhancing Fairness of Visual Attribute Predictors

The performance of deep neural networks for image recognition tasks such as predicting a smiling face is known to degrade with under-represented classes of sensitive attributes. We address this problem by introducing fairness-aware regularization losses based on batch estimates of Demographic Parity, Equalized Odds, and a novel Intersection-over-Union measure. The experiments performed on facial and medical images from CelebA, UTKFace, and the SIIM-ISIC melanoma classification challenge show the effectiveness of our proposed fairness losses for bias mitigation as they improve model fairness while maintaining high classification performance. To the best of our knowledge, our work is the first attempt to incorporate these types of losses in an end-to-end training scheme for mitigating biases of visual attribute predictors. Our code is available at https://github.com/nish03/FVAP.

cs.CV↗

A New Framework for Quantum Oblivious Transfer

We present a new template for building oblivious transfer from quantum information that we call the "fixed basis" framework. Our framework departs from prior work (eg., Crepeau and Kilian, FOCS '88) by fixing the correct choice of measurement basis used by each player, except for some hidden trap qubits that are intentionally measured in a conjugate basis. We instantiate this template in the quantum random oracle model (QROM) to obtain simple protocols that implement, with security against malicious adversaries: 1. Non-interactive random-input bit OT in a model where parties share EPR pairs a priori. 2. Two-round random-input bit OT without setup, obtained by showing that the protocol above remains secure even if the (potentially malicious) OT receiver sets up the EPR pairs. 3. Three-round chosen-input string OT from BB84 states without entanglement or setup. This improves upon natural variations of the CK88 template that require at least five rounds. Along the way, we develop technical tools that may be of independent interest. We prove that natural functions like XOR enable seedless randomness extraction from certain quantum sources of entropy. We also use idealized (i.e. extractable and equivocal) bit commitments, which we obtain by proving security of simple and efficient constructions in the QROM.

quant-ph↗

A Direct Construction of Complete Complementary Code with Zero Correlation Zone property for Prime-Power Length

In this paper, we propose a direct construction of a novel type of code set, which has combined properties of complete complementary code (CCC) and zero-correlation zone (ZCZ) sequences and called it complete complementary-ZCZ (CC-ZCZ) code set. The code set is constructed by using multivariable functions. The proposed construction also provides Golay-ZCZ codes with new lengths, i.e., prime-power lengths. The proposed Golay-ZCZ codes are optimal and asymptotically optimal for binary and non-binary cases, respectively, by \emph{Tang-Fan-Matsufuzi} bound. Furthermore, the proposed direct construction provides novel ZCZ sequences of length $p^k$, where $k$ is an integer $\geq 2$. We establish a relationship between the proposed CC-ZCZ code set and the first-order generalized Reed-Muller (GRM) code, and proved that both have the same Hamming distance. We also counted the number of CC-ZCZ code set in first-order GRM codes. The column sequence peak-to-mean envelope power ratio (PMEPR) of the proposed CC-ZCZ construction is derived and compared with existing works. The proposed construction is also deduced to Golay-ZCZ and ZCZ sequences which are compared to the existing work. The proposed construction generalizes many of the existing work.

cs.IT↗

TransDrift: Modeling Word-Embedding Drift using Transformer

In modern NLP applications, word embeddings are a crucial backbone that can be readily shared across a number of tasks. However as the text distributions change and word semantics evolve over time, the downstream applications using the embeddings can suffer if the word representations do not conform to the data drift. Thus, maintaining word embeddings to be consistent with the underlying data distribution is a key problem. In this work, we tackle this problem and propose TransDrift, a transformer-based prediction model for word embeddings. Leveraging the flexibility of transformer, our model accurately learns the dynamics of the embedding drift and predicts the future embedding. In experiments, we compare with existing methods and show that our model makes significantly more accurate predictions of the word embedding than the baselines. Crucially, by applying the predicted embeddings as a backbone for downstream classification tasks, we show that our embeddings lead to superior performance compared to the previous methods.

cs.CL↗

InFlow: Robust outlier detection utilizing Normalizing Flows

Normalizing flows are prominent deep generative models that provide tractable probability distributions and efficient density estimation. However, they are well known to fail while detecting Out-of-Distribution (OOD) inputs as they directly encode the local features of the input representations in their latent space. In this paper, we solve this overconfidence issue of normalizing flows by demonstrating that flows, if extended by an attention mechanism, can reliably detect outliers including adversarial attacks. Our approach does not require outlier data for training and we showcase the efficiency of our method for OOD detection by reporting state-of-the-art performance in diverse experimental settings. Code available at https://github.com/ComputationalRadiationPhysics/InFlow .

cs.LG↗

Applications of deep learning in traffic congestion detection, prediction and alleviation: A survey

Detecting, predicting, and alleviating traffic congestion are targeted at improving the level of service of the transportation network. With increasing access to larger datasets of higher resolution, the relevance of deep learning for such tasks is increasing. Several comprehensive survey papers in recent years have summarised the deep learning applications in the transportation domain. However, the system dynamics of the transportation network vary greatly between the non-congested state and the congested state -- thereby necessitating the need for a clear understanding of the challenges specific to congestion prediction. In this survey, we present the current state of deep learning applications in the tasks related to detection, prediction, and alleviation of congestion. Recurring and non-recurring congestion are discussed separately. Our survey leads us to uncover inherent challenges and gaps in the current state of research. Finally, we present some suggestions for future research directions as answers to the identified challenges.

cs.LG↗

FuseVis: Interpreting neural networks for image fusion using per-pixel saliency visualization

Image fusion helps in merging two or more images to construct a more informative single fused image. Recently, unsupervised learning based convolutional neural networks (CNN) have been utilized for different types of image fusion tasks such as medical image fusion, infrared-visible image fusion for autonomous driving as well as multi-focus and multi-exposure image fusion for satellite imagery. However, it is challenging to analyze the reliability of these CNNs for the image fusion tasks since no groundtruth is available. This led to the use of a wide variety of model architectures and optimization functions yielding quite different fusion results. Additionally, due to the highly opaque nature of such neural networks, it is difficult to explain the internal mechanics behind its fusion results. To overcome these challenges, we present a novel real-time visualization tool, named FuseVis, with which the end-user can compute per-pixel saliency maps that examine the influence of the input image pixels on each pixel of the fused image. We trained several image fusion based CNNs on medical image pairs and then using our FuseVis tool, we performed case studies on a specific clinical application by interpreting the saliency maps from each of the fusion methods. We specifically visualized the relative influence of each input image on the predictions of the fused image and showed that some of the evaluated image fusion methods are better suited for the specific clinical application. To the best of our knowledge, currently, there is no approach for visual analysis of neural networks for image fusion. Therefore, this work opens up a new research direction to improve the interpretability of deep fusion networks. The FuseVis tool can also be adapted in other deep neural network based image processing applications to make them interpretable.

cs.CV↗

CrypTFlow2: Practical 2-Party Secure Inference

We present CrypTFlow2, a cryptographic framework for secure inference over realistic Deep Neural Networks (DNNs) using secure 2-party computation. CrypTFlow2 protocols are both correct -- i.e., their outputs are bitwise equivalent to the cleartext execution -- and efficient -- they outperform the state-of-the-art protocols in both latency and scale. At the core of CrypTFlow2, we have new 2PC protocols for secure comparison and division, designed carefully to balance round and communication complexity for secure inference tasks. Using CrypTFlow2, we present the first secure inference over ImageNet-scale DNNs like ResNet50 and DenseNet121. These DNNs are at least an order of magnitude larger than those considered in the prior work of 2-party DNN inference. Even on the benchmarks considered by prior work, CrypTFlow2 requires an order of magnitude less communication and 20x-30x less time than the state-of-the-art.

cs.CR↗

Activity-based contact network scaling and epidemic propagation in metropolitan areas

Given the growth of urbanization and emerging pandemic threats, more sophisticated models are required to understand disease propagation and investigate the impacts of intervention strategies across various city types. We introduce a fully mechanistic, activity-based and highly spatio-temporally resolved epidemiological model which leverages on person-trajectories obtained from integrated mobility demand and supply models in full-scale cities. Simulating COVID-19 evolution in two full-scale cities with representative synthetic populations and mobility patterns, we analyze activity-based contact networks. We observe that transit contacts are scale-free in both cities, work contacts are Weibull distributed, and shopping or leisure contacts are exponentially distributed. We also investigate the impact of the transit network, finding that its removal dampens disease propagation, while work is also critical to post-peak disease spreading. Our framework, validated against existing case and mortality data, demonstrates the potential for tracking and tracing, along with detailed socio-demographic and mobility analyses of epidemic control strategies.

physics.soc-ph↗

CrypTFlow: Secure TensorFlow Inference

We present CrypTFlow, a first of its kind system that converts TensorFlow inference code into Secure Multi-party Computation (MPC) protocols at the push of a button. To do this, we build three components. Our first component, Athos, is an end-to-end compiler from TensorFlow to a variety of semi-honest MPC protocols. The second component, Porthos, is an improved semi-honest 3-party protocol that provides significant speedups for TensorFlow like applications. Finally, to provide malicious secure MPC protocols, our third component, Aramis, is a novel technique that uses hardware with integrity guarantees to convert any semi-honest MPC protocol into an MPC protocol that provides malicious security. The malicious security of the protocols output by Aramis relies on integrity of the hardware and semi-honest security of MPC. Moreover, our system matches the inference accuracy of plaintext TensorFlow. We experimentally demonstrate the power of our system by showing the secure inference of real-world neural networks such as ResNet50 and DenseNet121 over the ImageNet dataset with running times of about 30 seconds for semi-honest security and under two minutes for malicious security. Prior work in the area of secure inference has been limited to semi-honest security of small networks over tiny datasets such as MNIST or CIFAR. Even on MNIST/CIFAR, CrypTFlow outperforms prior work.

cs.CR↗

Additional transition line in jammed asymmetric bidisperse granular packings

We present numerical evidence for an additional discontinuous transition inside the jammed regime for an asymmetric bidisperse granular packing upon compression. This additional transition line separates jammed states with networks of predominantly large particles from jammed networks formed by both large and small particles, and the transition is indicated by a discontinuity in the number of particles contributing to the jammed network. The additional transition line emerges from the curves of jamming transitions and terminates in an end-point where the discontinuity vanishes. The additional line is starting at a size ratio around $δ= 0.22$ and grows longer for smaller $δ$. For $δ\to 0$, the additional transition line approaches a limit that can be derived analytically. The observed jamming scenarios are reminiscent of glass-glass transitions found in colloidal glasses.

cond-mat.soft↗

Visualisation of Medical Image Fusion and Translation for Accurate Diagnosis of High Grade Gliomas

The medical image fusion combines two or more modalities into a single view while medical image translation synthesizes new images and assists in data augmentation. Together, these methods help in faster diagnosis of high grade malignant gliomas. However, they might be untrustworthy due to which neurosurgeons demand a robust visualisation tool to verify the reliability of the fusion and translation results before they make pre-operative surgical decisions. In this paper, we propose a novel approach to compute a confidence heat map between the source-target image pair by estimating the information transfer from the source to the target image using the joint probability distribution of the two images. We evaluate several fusion and translation methods using our visualisation procedure and showcase its robustness in enabling neurosurgeons to make finer clinical decisions.

cs.CV↗

Structural Similarity based Anatomical and Functional Brain Imaging Fusion

Multimodal medical image fusion helps in combining contrasting features from two or more input imaging modalities to represent fused information in a single image. One of the pivotal clinical applications of medical image fusion is the merging of anatomical and functional modalities for fast diagnosis of malignant tissues. In this paper, we present a novel end-to-end unsupervised learning-based Convolutional Neural Network (CNN) for fusing the high and low frequency components of MRI-PET grayscale image pairs, publicly available at ADNI, by exploiting Structural Similarity Index (SSIM) as the loss function during training. We then apply color coding for the visualization of the fused image by quantifying the contribution of each input image in terms of the partial derivatives of the fused image. We find that our fusion and visualization approach results in better visual perception of the fused image, while also comparing favorably to previous methods when applying various quantitative assessment metrics.

eess.IV↗