Searcharxiv⌕ Search

arXiv subjects

Arun Kumar

Publications and source records attributed to Arun Kumar.

At least 73 records · Page 4Linked to original sources

Morphological Segmentation Inside-Out

Morphological segmentation has traditionally been modeled with non-hierarchical models, which yield flat segmentations as output. In many cases, however, proper morphological analysis requires hierarchical structure -- especially in the case of derivational morphology. In this work, we introduce a discriminative, joint model of morphological segmentation along with the orthographic changes that occur during word formation. To the best of our knowledge, this is the first attempt to approach discriminative segmentation with a context-free model. Additionally, we release an annotated treebank of 7454 English words with constituency parses, encouraging future research in this area.

cs.CL↗

A novel method of fuzzy time series forecasting based on interval index number and membership value using support vector machine

Fuzzy time series forecasting methods are very popular among researchers for predicting future values as they are not based on the strict assumptions of traditional time series forecasting methods. Non-stochastic methods of fuzzy time series forecasting are preferred by the researchers as they provide more significant forecasting results. There are generally, four factors that determine the performance of the forecasting method (1) number of intervals (NOIs) and length of intervals to partition universe of discourse (UOD) (2) fuzzification rules or feature representation of crisp time series (3) method of establishing fuzzy logic rule (FLRs) between input and target values (4) defuzzification rule to get crisp forecasted value. Considering the first two factors to improve the forecasting accuracy, we proposed a novel non-stochastic method fuzzy time series forecasting in which interval index number and membership value are used as input features to predict future value. We suggested a simple rounding-off range and suitable step size method to find the optimal number of intervals (NOIs) and used fuzzy c-means clustering process to divide UOD into intervals of unequal length. We implement support vector machine (SVM) to establish FLRs. To test our proposed method we conduct a simulated study on five widely used real time series and compare the performance with some recently developed models. We also examine the performance of the proposed model by using multi-layer perceptron (MLP) instead of SVM. Two performance measures RSME and SMAPE are used for performance analysis and observed better forecasting accuracy by the proposed model.

cs.LG↗

Skellam Type Processes of Order K and Beyond

In this article, we introduce Skellam process of order k and its running average. We also discuss the time-changed Skellam process of order k. In particular we discuss space-fractional Skellam process and tempered space-fractional Skellam process via time changes in Poisson process by independent stable subordinator and tempered stable subordinator, respectively. We derive the marginal probabilities, Levy measures, governing difference-differential equations of the introduced processes. Our results generalize Skellam process and running average of Poisson process in several directions.

math.PR↗

A Methodology to Assess the Human Factors Associated with Lunar Teleoperated Assembly Tasks

Low-latency telerobotics can enable more intricate surface tasks on extraterrestrial planetary bodies than has ever been attempted. For humanity to create a sustainable lunar presence, well-developed collaboration between humans and robots is necessary to perform complex tasks. This paper presents a methodology to assess the human factors, situational awareness (SA) and cognitive load (CL), associated with teleoperated assembly tasks. Currently, telerobotic assembly on an extraterrestrial body has never been attempted, and a valid methodology to assess the associated human factors has not been developed. The Telerobotics Laboratory at the University of Colorado-Boulder created the Telerobotic Simulation System (TSS) which enables remote operation of a rover and a robotic arm. The TSS was used in a laboratory experiment designed as an analog to a lunar mission. The operator's task was to assemble a radio interferometer. Each participant completed this task under two conditions, remote teleoperation (limited SA) and local operation (optimal SA). The goal of the experiment was to establish a methodology to accurately measure the operator's SA and CL while performing teleoperated assembly tasks. A successful methodology would yield results showing greater SA and lower CL while operating locally. Performance metrics showed greater SA and lower CL in the local environment, supported by a 27% increase in the mean time to completion of the assembly task when operating remotely. Subjective measurements of SA and CL did not align with the performance metrics. Results from this experiment will guide future work attempting to accurately quantify the human factors associated with telerobotic assembly. Once an accurate methodology has been developed, we will be able to measure how new variables affect an operator's SA and CL to optimize the efficiency and effectiveness of telerobotic assembly tasks.

cs.RO↗

Potential Theory of Normal Tempered Stable Process

In this article, we study the potential theory of normal tempered stable process which is obtained by time-changing the Brownian motion with a tempered stable subordinator. Precisely, we study the asymptotic behavior of potential density and Levy density associated with tempered stable subordinator and the Green function and the Levy density associated with the normal tempered stable process. We also provide the corresponding results for normal inverse Gaussian process which is a well studied process in literature.

math.PR↗

Hayward black holes in Einstein-Gauss-Bonnet gravity

The Hayward metric is a spherically symmetric charged regular black holes, a modification of the Reisnner-Nordstr$\ddot{o}$m black holes of Einstein's equations coupled to nonlinear electrodynamics. We consider Einstein-Gauss-Bonnet gravity (EGB) coupled to nonlinear electrodynamics to present an exact five dimension ($5D$) Hayward black holes with a regular center, having inner (Cauchy) and outer (event) horizons which go over to Boulware-Desser black holes when the charge is switched off ($e=0$). The presence of charge $e$ leads the modification in thermodynamical quantities, and it has also been shown that the Hawking-Page like phase transition can be achieved. The specific heat shows divergence at the horizon radius $r=r_C$ (critical radius), where the temperature has a maximum. Our result in the limit, $e\to0$, reduces vis-a-vis to the $5D$ Boulware-Desser solutions.

gr-qc↗

Understanding and Benchmarking the Impact of GDPR on Database Systems

The General Data Protection Regulation (GDPR) provides new rights and protections to European people concerning their personal data. We analyze GDPR from a systems perspective, translating its legal articles into a set of capabilities and characteristics that compliant systems must support. Our analysis reveals the phenomenon of metadata explosion, wherein large quantities of metadata needs to be stored along with the personal data to satisfy the GDPR requirements. Our analysis also helps us identify new workloads that must be supported under GDPR. We design and implement an open-source benchmark called GDPRbench that consists of workloads and metrics needed to understand and assess personal-data processing database systems. To gauge the readiness of modern database systems for GDPR, we follow best practices and developer recommendations to modify Redis, PostgreSQL, and a commercial database system to be GDPR compliant. Our experiments demonstrate that the resulting GDPR compliant systems achieve poor performance on GPDR workloads, and that performance scales poorly as the volume of personal data increases. We discuss the real-world implications of these findings, and identify research challenges towards making GDPR compliance efficient in production environments. We release all of our software artifacts and datasets at http://www.gdprbench.org

cs.DB↗

MLSys: The New Frontier of Machine Learning Systems

Machine learning (ML) techniques are enjoying rapidly increasing adoption. However, designing and implementing the systems that support ML models in real-world deployments remains a significant obstacle, in large part due to the radically different development and deployment profile of modern ML methods, and the range of practical concerns that come with broader adoption. We propose to foster a new systems machine learning research community at the intersection of the traditional systems and ML communities, focused on topics such as hardware systems for ML, software systems for ML, and ML optimized for metrics beyond predictive accuracy. To do this, we describe a new conference, MLSys, that explicitly targets research at the intersection of systems and machine learning with a program committee split evenly between experts in systems and ML, and an explicit focus on topics at the intersection of the two.

cs.LG↗

Mixtures of Tempered Stable Subordinators

In this article, we introduce mixtures of tempered stable subordinators (TSS). These mixtures define a class of subordinators which generalize tempered stable subordinators. The main properties like probability density function (pdf), Levy density, moments, governing Fokker-Planck-Kolmogorov (FPK) type equations, asymptotic form of potential density and asymptotic form of the renewal function for the corresponding inverse subordinator are discussed. We also generalize these results to n-th order mixtures of TSS.

math.PR↗

Predicting Eating Events in Free Living Individuals -- A Technical Report

This technical report records the experiments of applying multiple machine learning algorithms for predicting eating and food purchasing behaviors of free-living individuals. Data was collected with accelerometer, global positioning system (GPS), and body-worn cameras called SenseCam over a one week period in 81 individuals from a variety of ages and demographic backgrounds. These data were turned into minute-level features from sensors as well as engineered features that included time (e.g., time since last eating) and environmental context (e.g., distance to nearest grocery store). Algorithms include Logistic Regression, RBF-SVM, Random Forest, and Gradient Boosting. Our results show that the Gradient Boosting model has the highest mean accuracy score (0.7289) for predicting eating events before 0 to 4 minutes. For predicting food purchasing events, the RBF-SVM model (0.7395) outperforms others. For both prediction models, temporal and spatial features were important contributors to predicting eating and food purchasing events.

cs.LG↗

A novel Algorithm for Optimal Placement of Multiple Inertial Sensors to Improve the Sensing Accuracy

This paper proposes a novel algorithm to determine the optimal placement of redundant inertial sensors such as accelerometers and gyroscopes (gyros) for increasing the sensing accuracy. In this paper, we have proposed a novel iterative algorithm to find the optimal sensor configuration. The proposed algorithm utilizes the majorization-minimization (MM) algorithm and the duality principle to find the optimal configuration. Unlike the state-of-the-art which are mainly geometrical in nature and restricted to certain noise statistics, the proposed algorithm gives the exact positions of the sensors, and moreover, the proposed algorithm is independent of the nature of the noise at different sensors. The proposed alogrithm has been implemented and tested via numerical simulation in the MATLAB. The simulation results show that the algorithm converges to the optimal configurations and show the effectiveness of the proposed algorithm.

eess.SP↗

Fractional Risk Process in Insurance

Important models in insurance, for example the Carm{é}r--Lundberg theory and the Sparre Andersen model, essentially rely on the Poisson process. The process is used to model arrival times of insurance claims. This paper extends the classical framework for ruin probabilities by proposing and involving the fractional Poisson process as a counting process and addresses fields of applications in insurance. The interdependence of the fractional Poisson process is an important feature of the process, which leads to initial stress of the surplus process. On the other hand we demonstrate that the average capital required to recover a company after ruin does not change when switching to the fractional Poisson regime. We finally address particular risk measures, which allow simple evaluations in an environment governed by the fractional Poisson process.

math.ST↗

On the infinite divisibility of distributions of some inverse subordinators

We consider the infinite divisibility of distributions of some well-known inverse subordinators. Using a tail probability bound, we establish that distributions of many of the inverse subordinators used in the literature are not infinitely divisible. We further show that the distribution of a renewal process time-changed by an inverse stable subordinator is not infinitely divisible, which in particular implies that the distribution of the fractional Poisson process is not infinitely divisible.

math.PR↗

Belief dynamics extraction

Animal behavior is not driven simply by its current observations, but is strongly influenced by internal states. Estimating the structure of these internal states is crucial for understanding the neural basis of behavior. In principle, internal states can be estimated by inverting behavior models, as in inverse model-based Reinforcement Learning. However, this requires careful parameterization and risks model-mismatch to the animal. Here we take a data-driven approach to infer latent states directly from observations of behavior, using a partially observable switching semi-Markov process. This process has two elements critical for capturing animal behavior: it captures non-exponential distribution of times between observations, and transitions between latent states depend on the animal's actions, features that require more complex non-markovian models to represent. To demonstrate the utility of our approach, we apply it to the observations of a simulated optimal agent performing a foraging task, and find that latent dynamics extracted by the model has correspondences with the belief dynamics of the agent. Finally, we apply our model to identify latent states in the behaviors of monkey performing a foraging task, and find clusters of latent states that identify periods of time consistent with expectant waiting. This data-driven behavioral model will be valuable for inferring latent cognitive states, and thereby for measuring neural representations of those states.

cs.AI↗

Tuple-oriented Compression for Large-scale Mini-batch Stochastic Gradient Descent

Data compression is a popular technique for improving the efficiency of data processing workloads such as SQL queries and more recently, machine learning (ML) with classical batch gradient methods. But the efficacy of such ideas for mini-batch stochastic gradient descent (MGD), arguably the workhorse algorithm of modern ML, is an open question. MGD's unique data access pattern renders prior art, including those designed for batch gradient methods, less effective. We fill this crucial research gap by proposing a new lossless compression scheme we call tuple-oriented compression (TOC) that is inspired by an unlikely source, the string/text compression scheme Lempel-Ziv-Welch, but tailored to MGD in a way that preserves tuple boundaries within mini-batches. We then present a suite of novel compressed matrix operation execution techniques tailored to the TOC compression scheme that operate directly over the compressed data representation and avoid decompression overheads. An extensive empirical evaluation with real-world datasets shows that TOC consistently achieves substantial compression ratios by up to 51x and reduces runtimes for MGD workloads by up to 10.2x in popular ML systems.

cs.LG↗

Document Structure Measure for Hypernym discovery

Hypernym discovery is the problem of finding terms that have is-a relationship with a given term. We introduce a new context type, and a relatedness measure to differentiate hypernyms from other types of semantic relationships. Our Document Structure measure is based on hierarchical position of terms in a document, and their presence or otherwise in definition text. This measure quantifies the document structure using multiple attributes, and classes of weighted distance functions.

cs.CL↗

Fine Grained Classification of Personal Data Entities

Entity Type Classification can be defined as the task of assigning category labels to entity mentions in documents. While neural networks have recently improved the classification of general entity mentions, pattern matching and other systems continue to be used for classifying personal data entities (e.g. classifying an organization as a media company or a government institution for GDPR, and HIPAA compliance). We propose a neural model to expand the class of personal data entities that can be classified at a fine grained level, using the output of existing pattern matching systems as additional contextual features. We introduce new resources, a personal data entities hierarchy with 134 types, and two datasets from the Wikipedia pages of elected representatives and Enron emails. We hope these resource will aid research in the area of personal data discovery, and to that effect, we provide baseline results on these datasets, and compare our method with state of the art models on OntoNotes dataset.

cs.CL↗

Temporal Proximity induces Attributes Similarity

Users consume their favorite content in temporal proximity of consumption bundles according to their preferences and tastes. Thus, the underlying attributes of items implicitly match user preferences, however, current recommender systems largely ignore this fundamental driver in identifying matching items. In this work, we introduce a novel temporal proximity filtering method to enable items-matching. First, we demonstrate that proximity preferences exist. Second, we present an induced similarity metric in temporal proximity driven by user tastes and third, we show that this induced similarity can be used to learn items pairwise similarity in attribute space. The proposed model does not rely on any knowledge outside users' consumption bundles and provide a novel way to devise user preferences and tastes driven novel items recommender.

cs.IR↗