SearcharxivSearch

arXiv subjects

Michael D. Porter

Publications and source records attributed to Michael D. Porter.

13 recordsLinked to original sources

Quantifying the Influence of User Behaviors on the Dissemination of Fake News on Twitter with Multivariate Hawkes Processes

Fake news has emerged as a pervasive problem within Online Social Networks, leading to a surge of research interest in this area. Understanding the dissemination mechanisms of fake news is crucial in comprehending the propagation of disinformation/misinformation and its impact on users in Online Social Networks. This knowledge can facilitate the development of interventions to curtail the spread of false information and inform affected users to remain vigilant against fraudulent/malicious content. In this paper, we specifically target the Twitter platform and propose a Multivariate Hawkes Point Processes model that incorporates essential factors such as user networks, response tweet types, and user stances as model parameters. Our objective is to investigate and quantify their influence on the dissemination process of fake news. We derive parameter estimation expressions using an Expectation Maximization algorithm and validate them on a simulated dataset. Furthermore, we conduct a case study using a real dataset of fake news collected from Twitter to explore the impact of user stances and tweet types on dissemination patterns. This analysis provides valuable insights into how users are influenced by or influence the dissemination process of disinformation/misinformation, and demonstrates how our model can aid in intervening in this process.

cs.SI

Reevaluating Data Partitioning for Emotion Detection in EmoWOZ

This paper focuses on the EmoWoz dataset, an extension of MultiWOZ that provides emotion labels for the dialogues. MultiWOZ was partitioned initially for another purpose, resulting in a distributional shift when considering the new purpose of emotion recognition. The emotion tags in EmoWoz are highly imbalanced and unevenly distributed across the partitions, which causes sub-optimal performance and poor comparison of models. We propose a stratified sampling scheme based on emotion tags to address this issue, improve the dataset's distribution, and reduce dataset shift. We also introduce a special technique to handle conversation (sequential) data with many emotional tags. Using our proposed sampling method, models built upon EmoWoz can perform better, making it a more reliable resource for training conversational agents with emotional intelligence. We recommend that future researchers use this new partitioning to ensure consistent and accurate performance evaluations.

cs.CL

Learning affective meanings that derives the social behavior using Bidirectional Encoder Representations from Transformers

Predicting the outcome of a process requires modeling the system dynamic and observing the states. In the context of social behaviors, sentiments characterize the states of the system. Affect Control Theory (ACT) uses sentiments to manifest potential interaction. ACT is a generative theory of culture and behavior based on a three-dimensional sentiment lexicon. Traditionally, the sentiments are quantified using survey data which is fed into a regression model to explain social behavior. The lexicons used in the survey are limited due to prohibitive cost. This paper uses a fine-tuned Bidirectional Encoder Representations from Transformers (BERT) model to develop a replacement for these surveys. This model achieves state-of-the-art accuracy in estimating affective meanings, expanding the affective lexicon, and allowing more behaviors to be explained.

cs.CL

A tale of two metrics: Polling and financial contributions as a measure of performance

Campaign analysis is an integral part of American democracy and has many complexities in its dynamics. Experts have long sought to understand these dynamics and evaluate campaign performance using a variety of techniques. We explore campaign financing and standing in the polls as two components of campaign performance in the context of the 2020 Democratic primaries. We show where these measures exhibit represent similar dynamics and where they differ. We focus on identifying change points in the trend for all candidates using joinpoint regression models. We find how these change points identify major events such as failure or success in a debate. Joinpoint regression reveals who the voters support when they stop supporting a specific candidate. This study demonstrates the value of joinpoint regression in political campaign analysis and it represents a crossover of this technique into the political domain building a foundation for continued exploration and use of this method.

stat.AP

How emoji and word embedding helps to unveil emotional transitions during online messaging

During online chats, body-language and vocal characteristics are not part of the communication mechanism making it challenging to facilitate an accurate interpretation of feelings, emotions, and attitudes. The use of emojis to express emotional feeling is an alternative approach in these types of communication. In this project, we focus on modeling a customer's emotion in an online messaging session with a chatbot. We use Affect Control Theory (ACT) to predict emotional change during the interaction. To let the customer use emojis, we also extend the affective dictionaries used by ACT. For this purpose, we mapped Emoji2vec embedding to the affective space. Our framework can find emotional change during messaging and how a customer's reaction is changed accordingly.

cs.HC

Detecting, identifying, and localizing radiological material in urban environments using scan statistics

A method is proposed, based on scan statistics, to detect, identify, and localize illicit radiological material using mobile sensors in an urban environment. Our method handles varying levels of background radiation that change according to an (unknown) environment. Our method can accurately determine if a source is present along a street segment as well as identify which of six possible sources generated the radiation. Our method can also localize the source, when detected, to within a few seconds. We have presented our results across a range of decision thresholds allowing stakeholders to evaluate the performance at different false alarm rates. Due to the simplicity of our approach, our models can be trained in a few minutes with very little training data and holds the potential to score a run in real-time. Our method was one of the top performing submissions in the 'Detecting Radiological Threats in Urban Areas' competition.

eess.SP

Enabling the next generation of scientific discoveries by embracing photonic technologies

The fields of Astronomy and Astrophysics are technology limited, where the advent and application of new technologies to astronomy usher in a flood of discoveries altering our understanding of the Universe (e.g., recent cases include LIGO and the GRAVITY instrument at the VLTI). Currently, the field of astronomical spectroscopy is rapidly approaching an impasse: the size and cost of instruments, especially multi-object and integral field spectrographs for extremely large telescopes (ELTs), are pushing the limits of what is feasible, requiring optical components at the very edge of achievable size and performance. For these reasons, astronomers are increasingly looking for innovative solutions like photonic technologies that promote instrument miniaturization and simplification, while providing superior performance. Astronomers have long been aware of the potential of photonic technologies. The goal of this white paper is to draw attention to key photonic technologies and developments over the past two decades and demonstrate there is new momentum in this arena. We outline where the most critical efforts should be focused over the coming decade in order to move towards realizing a fully photonic instrument. A relatively small investment in this technology will advance astronomical photonics to a level where it can reliably be used to solve challenging instrument design limitations. For the benefit of both ground and space borne instruments alike, an endorsement from the National Academy of Sciences decadal survey will ensure that such solutions are set on a path to their full scientific exploitation, which may one day address a broad range of science cases outlined in the KSPs.

astro-ph.IM

Modelling the Proliferation of Terrorism via Diffusion and Contagion

The proliferation of terrorism is a serious concern in national and international security, as its spread is seen as an existential threat to Western liberal democracies. Understanding and effectively modelling the spread of terrorism provides useful insight into formulating effective responses. A mathematical model capturing the theoretical constructs of contagion and diffusion is constructed for explaining the spread of terrorist activity and used to analyse data from the Global Terrorism Database from 2000--2016 for Afghanistan, Iraq, and Israel.

stat.AP

Optimal Bayesian clustering using non-negative matrix factorization

Bayesian model-based clustering is a widely applied procedure for discovering groups of related observations in a dataset. These approaches use Bayesian mixture models, estimated with MCMC, which provide posterior samples of the model parameters and clustering partition. While inference on model parameters is well established, inference on the clustering partition is less developed. A new method is developed for estimating the optimal partition from the pairwise posterior similarity matrix generated by a Bayesian cluster model. This approach uses non-negative matrix factorization (NMF) to provide a low-rank approximation to the similarity matrix. The factorization permits hard or soft partitions and is shown to perform better than several popular alternatives under a variety of penalty functions.

stat.ME

A Statistical Approach to Crime Linkage

The object of this paper is to develop a statistical approach to criminal linkage analysis that discovers and groups crime events that share a common offender and prioritizes suspects for further investigation. Bayes factors are used to describe the strength of evidence that two crimes are linked. Using concepts from agglomerative hierarchical clustering, the Bayes factors for crime pairs are combined to provide similarity measures for comparing two crime series. This facilitates crime series clustering, crime series identification, and suspect prioritization. The ability of our models to make correct linkages and predictions is demonstrated under a variety of real-world scenarios with a large number of solved and unsolved breaking and entering crimes. For example, a naïve Bayes model for pairwise case linkage can identify 82\% of actual linkages with a 5\% false positive rate. For crime series identification, 77\%-89\% of the additional crimes in a crime series can be identified from a ranked list of 50 incidents.

stat.AP

Self-exciting hurdle models for terrorist activity

A predictive model of terrorist activity is developed by examining the daily number of terrorist attacks in Indonesia from 1994 through 2007. The dynamic model employs a shot noise process to explain the self-exciting nature of the terrorist activities. This estimates the probability of future attacks as a function of the times since the past attacks. In addition, the excess of nonattack days coupled with the presence of multiple coordinated attacks on the same day compelled the use of hurdle models to jointly model the probability of an attack day and corresponding number of attacks. A power law distribution with a shot noise driven parameter best modeled the number of attacks on an attack day. Interpretation of the model parameters is discussed and predictive performance of the models is evaluated.

stat.AP

Mixture Likelihood Ratio Scan Statistic for Disease Outbreak Detection

Early detection of disease outbreaks is of paramount importance to implementing intervention strategies to mitigate the severity and duration of the outbreak. We build methodology that utilizes the characteristic profile of disease outbreaks to reduce the time to detection and false positive rate. We model daily counts through a Poisson distribution with additive background plus outbreak components. The outbreak component has a parametric form with unknown underlying parameters. A mixture likelihood ratio scan statistic is developed to maximize parameters over a window in time. This provides an alert statistic with early time to detection and low false positive rate. The methodology is demonstrated on three simulated data sets meant to represent E. coli, Cryptosporidium, and Influenza outbreaks.

stat.ME