Searcharxiv⌕ Search

arXiv subjects

Rajesh Sharma

Publications and source records attributed to Rajesh Sharma.

At least 73 records · Page 4Linked to original sources

Exploring Collaborative Game Play with Robots to Encourage Good Hand Hygiene Practises among Children

This paper presents the design, implementation, and evaluation of a novel collaborative educational game titled "Land of Hands", involving children and a customized social robot that we designed (HakshE). Through this gaming platform, we aim to teach proper hand hygiene practises to children and explore the extent of interactions that take place between a pro-social robot and children in such a setting. We blended gamification with Computers as Social Actors (CASA) paradigm to model the robot as a social actor or a fellow player in the game. The game was developed using Godot's 2D engine and Alice 3. In this study, 32 participants played the game online through a video teleconferencing platform Zoom. To understand the influence a pro-social robot's nudges has on children's interactions, we split our study into two conditions: With-Nudges and Without-Nudges. Detailed analysis of rubrics and video analyses of children's interactions show that our platform helped children learn good hand hygiene practises. We also found that using a pro-social robot creates enjoyable interactions and greater social engagement between the children and the robot although learning itself wasn't influenced by the pro-sociality of the robot.

cs.RO↗

Pre-Trained Language Transformers are Universal Image Classifiers

Facial images disclose many hidden personal traits such as age, gender, race, health, emotion, and psychology. Understanding these traits will help to classify the people in different attributes. In this paper, we have presented a novel method for classifying images using a pretrained transformer model. We apply the pretrained transformer for the binary classification of facial images in criminal and non-criminal classes. The pretrained transformer of GPT-2 is trained to generate text and then fine-tuned to classify facial images. During the finetuning process with images, most of the layers of GT-2 are frozen during backpropagation and the model is frozen pretrained transformer (FPT). The FPT acts as a universal image classifier, and this paper shows the application of FPT on facial images. We also use our FPT on encrypted images for classification. Our FPT shows high accuracy on both raw facial images and encrypted images. We hypothesize the meta-learning capacity FPT gained because of its large size and trained on a large size with theory and experiments. The GPT-2 trained to generate a single word token at a time, through the autoregressive process, forced to heavy-tail distribution. Then the FPT uses the heavy-tail property as its meta-learning capacity for classifying images. Our work shows one way to avoid bias during the machine classification of images.The FPT encodes worldly knowledge because of the pretraining of one text, which it uses during the classification. The statistical error of classification is reduced because of the added context gained from the text.Our paper shows the ethical dimension of using encrypted data for classification.Criminal images are sensitive to share across the boundary but encrypted largely evades ethical concern.FPT showing good classification accuracy on encrypted images shows promise for further research on privacy-preserving machine learning.

cs.CV↗

What goes on inside rumour and non-rumour tweets and their reactions: A Psycholinguistic Analyses

In recent years, the problem of rumours on online social media (OSM) has attracted lots of attention. Researchers have started investigating from two main directions. First is the descriptive analysis of rumours and secondly, proposing techniques to detect (or classify) rumours. In the descriptive line of works, where researchers have tried to analyse rumours using NLP approaches, there isnt much emphasis on psycho-linguistics analyses of social media text. These kinds of analyses on rumour case studies are vital for drawing meaningful conclusions to mitigate misinformation. For our analysis, we explored the PHEME9 rumour dataset (consisting of 9 events), including source tweets (both rumour and non-rumour categories) and response tweets. We compared the rumour and nonrumour source tweets and then their corresponding reply (response) tweets to understand how they differ linguistically for every incident. Furthermore, we also evaluated if these features can be used for classifying rumour vs. non-rumour tweets through machine learning models. To this end, we employed various classical and ensemble-based approaches. To filter out the highly discriminative psycholinguistic features, we explored the SHAP AI Explainability tool. To summarise, this research contributes by performing an in-depth psycholinguistic analysis of rumours related to various kinds of events.

cs.CL↗

Do Facial Trait Correlates with Roll Call Voting in Parliament? Using fWHR to Study Performance in Politics

Research has shown that people recognize and select leaders based on their facial appearance. However, considering the correlation between the performance of leaders and their facial traits, empirical findings are mixed. This paper adds to the debate by focusing on two previously understudied aspects of facial traits among political leaders: (i) previous studies have focused on electoral success and achievement drive of politicians omitting their actual daily performance after elections; (ii) previous research has analyzed individual politicians omitting the context of social circumstances which potentially influence their performance. We address these issues by analyzing Ukrainian members of parliament (MPs) who voted for bills in six consecutive Verkhovna Rada starting from Rada 4 (2002-06) to Rada 9 (2019-Present) to study politicians' performance, which is defined as co-voting or cooperation between MPs in voting on the same bill. In simple words, we analyze whether politicians tend to follow leaders when voting. This ability to summon the votes of others is interpreted as better performance. To measure performance, we proposed a generic methodology named Feature Importance For Measuring Performance (FIMP) that can be used in various scenarios. Using FIMP, our data suggest that MPs vote has no impact from their colleagues with higher or lower facial width-to-height ratio (fWHR), a popular measure of the facial trait.

cs.SI↗

Longitudinal change in language behaviour during protests: A case study of Euromaidan in Ukraine

In the last decade, online social media has become the primary platform for protesters to organize and express their agenda in various parts of the world. Nevertheless, scholars still debate whether online tools induce offline protests or facilitate them. Unfortunately, studies of protests often lack panel data and cannot address how particular users change their behaviour over time in line with the protest agenda. To this end, we analyze a new dataset of the Facebook page EuroMaydan that was explicitly created to facilitate the protest in Ukraine from November 2013 to February 2014. Moreover, our analysis follows this page even after the end of the protest till June 2014. In total, the dataset includes 26,631 posts and 1,470,593 comments that were generated by 124,790 users during and after the protests. We use this panel data to test how particular users switch between the two most popular languages of this page: Ukrainian and Russian. Previous studies have puzzled with language behaviour during the Ukrainian protest. Although researchers expected to see more Ukrainian language due to national mobilization, studies discovered that Ukrainian protesters frequently used Russian, especially after the end of the protest. A hypothesis was suggested that protesters use language strategically depending on circumstances (e.g., to maximize outreach). However, previous studies relied on aggregated data and did not explore the within-subject variation of language behaviour. Our study adds to this scholarship by adding a longitudinal analysis of language behaviour. Considering the broader contribution of this paper, we add to the literature on protests by validating previous findings derived from surveys that activists change their behaviour to reflect prior preferences and situational goals rather than modify their preferences.

cs.SI↗

Understanding Cycling Mobility: Bologna Case Study

Understanding human mobility in urban environments is of the utmost importance to manage traffic and for deploying new resources and services. In recent years, the problem is exacerbated due to rapid urbanization and climate changes. In an urban context, human mobility has many facets, and cycling represents one of the most eco-friendly and efficient/effective ways to move in touristic and historical cities. The main objective of this work is to study the cycling mobility within the city of Bologna, Italy. We used six months dataset that consists of 320,118 self-reported bike trips. In particular, we performed several descriptive analysis to understand spatial and temporal patterns of bike users for understanding popular roads, and most favorite points within the city. This analysis involved several other public datasets in order to explore variables that can possibly affect the cycling activity, such as weather, pollution, and events. The main results of this study indicate that bike usage is more correlated to temperature, and precipitation and has no correlation to wind speed and pollution. In addition, we also exploited various machine learning and deep learning approaches for predicting short-term trips in the near future (that is for the following 30, and 60 minutes), that could help local governmental agencies for urban planning. Our best model achieved an R square of 0.91, a Mean Absolute Error of 5.38 and a Root Mean Squared Error of 8.12 for the 30-minutes time interval.

cs.CY↗

Identifying Possible Rumor Spreaders on Twitter: A Weak Supervised Learning Approach

Online Social Media (OSM) platforms such as Twitter, Facebook are extensively exploited by the users of these platforms for spreading the (mis)information to a large audience effortlessly at a rapid pace. It has been observed that the misinformation can cause panic, fear, and financial loss to society. Thus, it is important to detect and control the misinformation in such platforms before it spreads to the masses. In this work, we focus on rumors, which is one type of misinformation (other types are fake news, hoaxes, etc). One way to control the spread of the rumors is by identifying users who are possibly the rumor spreaders, that is, users who are often involved in spreading the rumors. Due to the lack of availability of rumor spreaders labeled dataset (which is an expensive task), we use publicly available PHEME dataset, which contains rumor and non-rumor tweets information, and then apply a weak supervised learning approach to transform the PHEME dataset into rumor spreaders dataset. We utilize three types of features, that is, user, text, and ego-network features, before applying various supervised learning approaches. In particular, to exploit the inherent network property in this dataset (user-user reply graph), we explore Graph Convolutional Network (GCN), a type of Graph Neural Network (GNN) technique. We compare GCN results with the other approaches: SVM, RF, and LSTM. Extensive experiments performed on the rumor spreaders dataset, where we achieve up to 0.864 value for F1-Score and 0.720 value for AUC-ROC, shows the effectiveness of our methodology for identifying possible rumor spreaders using the GCN technique.

cs.AI↗

Misinformation Detection on YouTube Using Video Captions

Millions of people use platforms such as YouTube, Facebook, Twitter, and other mass media. Due to the accessibility of these platforms, they are often used to establish a narrative, conduct propaganda, and disseminate misinformation. This work proposes an approach that uses state-of-the-art NLP techniques to extract features from video captions (subtitles). To evaluate our approach, we utilize a publicly accessible and labeled dataset for classifying videos as misinformation or not. The motivation behind exploring video captions stems from our analysis of videos metadata. Attributes such as the number of views, likes, dislikes, and comments are ineffective as videos are hard to differentiate using this information. Using caption dataset, the proposed models can classify videos among three classes (Misinformation, Debunking Misinformation, and Neutral) with 0.85 to 0.90 F1-score. To emphasize the relevance of the misinformation class, we re-formulate our classification problem as a two-class classification - Misinformation vs. others (Debunking Misinformation and Neutral). In our experiments, the proposed models can classify videos with 0.92 to 0.95 F1-score and 0.78 to 0.90 AUC ROC.

cs.LG↗

Overall Behavioural Index (OBI) For Measuring Segregation

Segregation, defined as the degree of separation between two or more population groups, helps to understand a complex social environment and subsequently provides a basis for public policy intervention. To measure segregation, past works often propose indexes that are criticized for being over-simplified and over-reduced. In other words, these indexes use the highly aggregated information to measure segregation. In this paper, we propose three novel indexes to measure segregation, namely: (i) Individual Segregation Index (ISI), (ii) Individual Inclination Index (III), and (iii) Overall Behavioural Index (OBI). The ISI index measures individuals' segregation, and the III index reports the individuals' inclination towards other population groups. The OBI index, calculated using both III and ISI index, is non-simplified and not only recognizes individuals' connectivity behaviour but group's connectivity behavioural distribution as well. By considering commonly used Freeman's segregation and homophily index as baseline indexes, we compare the OBI index on real call data records (CDR) dataset of Estonia to show the effectiveness of the proposed indexes.

cs.SI↗

Studying Leaders & Their Concerns Using Online Social Media During The Times Of Crisis -- A COVID Case Study

Online social media (OSM) has emerged as a prominent platform for debate on a wide range of issues. Even celebrities and public figures often share their opinions on a variety of topics through OSM platforms. One such subject that has gained a lot of coverage on Twitter is the Novel Coronavirus, officially known as COVID-19, which has become a pandemic and has sparked a crisis in human history. In this study, we examine 29 million tweets over three months to study highly influential users, whom we refer to as leaders. We recognize these leaders through social network techniques and analyze their tweets using text analysis. Using a community detection algorithm, we categorize these leaders into four clusters: research, news, health, and politics, with each cluster containing Twitter handles (accounts) of individual users or organizations. E.g., the health cluster includes the World Health Organization (@WHO), the Director-General of WHO (@DrTedros), and so on. The emotion analysis reveals that (i) all clusters show an equal amount of fear in their tweets, (ii) research and news clusters display more sadness than others, and (iii) health and politics clusters are attempting to win public trust. According to the text analysis, the (i) research cluster is more concerned with recognizing symptoms and the development of vaccination; (ii) news and politics clusters are mostly concerned with travel. We then show that we can use our findings to classify tweets into clusters with a score of 96% AUC ROC.

cs.SI↗

A Graph Convolutional Neural Network based Framework for Estimating Future Citations Count of Research Articles

Scientific publications play a vital role in the career of a researcher. However, some articles become more popular than others among the research community and subsequently drive future research directions. One of the indicative signs of popular articles is the number of citations an article receives. The citation count, which is also the basis with various other metrics, such as the journal impact factor score, the $h$-index, is an essential measure for assessing a scientific paper's quality. In this work, we proposed a Graph Convolutional Network (GCN) based framework for estimating future research publication citations for both the short-term (1-year) and long-term (for 5-years and 10-years) duration. We have tested our proposed approach over the AMiner dataset, specifically on research articles from the computer science domain, consisting of more than 0.8 million articles.

cs.DL↗

Edge Computing Enabled by Unmanned Autonomous Vehicles

Pervasive applications are revolutionizing the perception that users have towards the environment. Indeed, pervasive applications perform resource intensive computations over large amounts of stream sensor data collected from multiple sources. This allows applications to provide richer and deep insights into the natural characteristics that govern everything that surrounds us. A key limitation of these applications is that they have high energy footprints, which in turn hampers the quality of experience of users. While cloud and edge computing solutions can be applied to alleviate the problem, these solutions are hard to adopt in existing architecture and far from become ubiquitous. Fortunately, cloudlets are becoming portable enough, such that they can be transported and integrated into any environment easily and dynamically. In this article, we investigate how cloudlets can be transported by unmanned autonomous vehicles (UAV)s to provide computation support on the edge. Based on our study, we develop GEESE, a novel UAVbased system that enables the dynamic deployment of an edge computing infrastructure through the cooperation of multiple UAVs carrying cloudlets. By using GEESE, we conduct rigorous experiments to analyze the effort to deliver cloudlets using aerial, ground, and underwater UAVs. Our results indicate that UAVs can work in a cooperative manner to enable edge computing in the wild.

cs.DC↗

COVID-19 and the stock market: evidence from Twitter

COVID-19 has had a much larger impact on the financial markets compared to previous epidemics because the news information is transferred over the social networks at a speed of light. Using Twitter's API, we compiled a unique dataset with more than 26 million COVID-19 related Tweets collected from February 2nd until May 1st, 2020. We find that more frequent use of the word "stock" in daily Tweets is associated with a substantial decline in log returns of three key US indices - Dow Jones Industrial Average, S&P500, and NASDAQ. The results remain virtually unchanged in multiple robustness checks.

econ.GN↗

Are Social Networks Watermarking Us or Are We (Unawarely) Watermarking Ourself?

In the last decade, Social Networks (SNs) have deeply changed many aspects of society, and one of the most widespread behaviours is the sharing of pictures. However, malicious users often exploit shared pictures to create fake profiles leading to the growth of cybercrime. Thus, keeping in mind this scenario, authorship attribution and verification through image watermarking techniques are becoming more and more important. In this paper, firstly, we investigate how 13 most popular SNs treat the uploaded pictures, in order to identify a possible implementation of image watermarking techniques by respective SNs. Secondly, on these 13 SNs, we test the robustness of several image watermarking algorithms. Finally, we verify whether a method based on the Photo-Response Non-Uniformity (PRNU) technique can be successfully used as a watermarking approach for authorship attribution and verification of pictures on SNs. The proposed method is robust enough in spite of the fact that the pictures get downgraded during the uploading process by SNs. The results of our analysis on a real dataset of 8,400 pictures show that the proposed method is more effective than other watermarking techniques and can help to address serious questions about privacy and security on SNs.

cs.MM↗

Which bills are lobbied? Predicting and interpreting lobbying activity in the US

Using lobbying data from OpenSecrets.org, we offer several experiments applying machine learning techniques to predict if a piece of legislation (US bill) has been subjected to lobbying activities or not. We also investigate the influence of the intensity of the lobbying activity on how discernible a lobbied bill is from one that was not subject to lobbying. We compare the performance of a number of different models (logistic regression, random forest, CNN and LSTM) and text embedding representations (BOW, TF-IDF, GloVe, Law2Vec). We report results of above 0.85% ROC AUC scores, and 78% accuracy. Model performance significantly improves (95% ROC AUC, and 88% accuracy) when bills with higher lobbying intensity are looked at. We also propose a method that could be used for unlabelled data. Through this we show that there is a considerably large number of previously unlabelled US bills where our predictions suggest that some lobbying activity took place. We believe our method could potentially contribute to the enforcement of the US Lobbying Disclosure Act (LDA) by indicating the bills that were likely to have been affected by lobbying but were not filed as such.

econ.GN↗

Mobility Based SIR Model For Pandemics -- With Case Study Of COVID-19

In the last decade, humanity has faced many different pandemics such as SARS, H1N1, and presently novel coronavirus (COVID-19). On one side, scientists are focusing on vaccinations, and on the other side, there is a need to propose models that can help us in understanding the spread of these pandemics as it can help governmental and other concerned agencies to be well prepared, especially from pandemics, which spreads faster like COVID-19. The main reason for some epidemic turning into pandemics is the connectivity among different regions of the world, which makes it easier to affect a wider geographical area, often worldwide. In addition, the population distribution and social coherence in the different regions of the world is non-uniform. Thus, once the epidemic enters a region, then the local population distribution plays an important role. Inspired by these ideas, we proposed a mobility-based SIR model for epidemics, which especially takes into account pandemic situations. To the best of our knowledge, this model is first of its kind, which takes into account the population distribution and connectivity of different geographic locations across the globe. In addition to presenting the mathematical proof of our model, we have performed extensive simulations using synthetic data to demonstrate our model's generalizability. To demonstrate the wider scope of our model, we used our model to forecast the COVID-19 cases for Estonia.

cs.SI↗