SearcharxivSearch

arXiv subjects

Jongwook Woo

Publications and source records attributed to Jongwook Woo.

16 recordsLinked to original sources

Predictive Analysis of CFPB Consumer Complaints Using Machine Learning

This paper introduces the Consumer Feedback Insight & Prediction Platform, a system leveraging machine learning to analyze the extensive Consumer Financial Protection Bureau (CFPB) Complaint Database, a publicly available resource exceeding 4.9 GB in size. This rich dataset offers valuable insights into consumer experiences with financial products and services. The platform itself utilizes machine learning models to predict two key aspects of complaint resolution: the timeliness of company responses and the nature of those responses (e.g., closed, closed with relief etc.). Furthermore, the platform employs Latent Dirichlet Allocation (LDA) to delve deeper, uncovering common themes within complaints and revealing underlying trends and consumer issues. This comprehensive approach empowers both consumers and regulators. Consumers gain valuable insights into potential response wait times, while regulators can utilize the platform's findings to identify areas where companies may require further scrutiny regarding their complaint resolution practices.

cs.DC

Cyberattack Data Analysis in IoT Environments using Big Data

In the landscape of the Internet of Things (IoT), transforming various industries, our research addresses the growing connectivity and security challenges, including interoperability and standardized protocols. Despite the anticipated exponential growth in IoT connections, network security remains a major concern due to inadequate datasets that fail to fully encompass potential cyberattacks in realistic IoT environments. Using Apache Hadoop and Hive, our in-depth analysis of security vulnerabilities identified intricate patterns and threats, such as attack behavior, network traffic anomalies, TCP flag usage, and targeted attacks, underscoring the critical need for robust data platforms to enhance IoT security.

cs.CR

US College Net Price Prediction Comparing ML Regression Models

This paper will illustrate the usage of Machine Learning algorithms on US College Scorecard datasets. For this paper, we will use our knowledge, research, and development of a predictive model to compare the results of all the models and predict the public and private net prices. This paper focuses on analyzing US College Scorecard data from data published on government websites. Our goal is to use four machine learning regression models to develop a predictive model to forecast the equitable net cost for every college, encompassing both public institutions and private, whether for-profit or nonprofit.

cs.CY

Insuring Smiles: Predicting routine dental coverage using Spark ML

Finding suitable health insurance coverage can be challenging for individuals and small enterprises in the USA. The Health Insurance Exchange Public Use Files (Exchange PUFs) dataset provided by CMS offers valuable information on health and dental policies [1]. In this paper, we leverage machine learning algorithms to predict if a health insurance plan covers routine dental services for adults. By analyzing plan type, region, deductibles, out-of-pocket maximums, and copayments, we employ Logistic Regression, Decision Tree, Random Forest, Gradient Boost, Factorization Model and Support Vector Machine algorithms. Our goal is to provide a clinical strategy for individuals and families to select the most suitable insurance plan based on income and expenses.

cs.LG

Using Spark Machine Learning Models to Perform Predictive Analysis on Flight Ticket Pricing Data

This paper discusses predictive performance and processes undertaken on flight pricing data utilizing r2(r-square) and RMSE that leverages a large dataset, originally from Expedia.com, consisting of approximately 20 million records or 4.68 gigabytes. The project aims to determine the best models usable in the real world to predict airline ticket fares for non-stop flights across the US. Therefore, good generalization capability and optimized processing times are important measures for the model. We will discover key business insights utilizing feature importance and discuss the process and tools used for our analysis. Four regression machine learning algorithms were utilized: Random Forest, Gradient Boost Tree, Decision Tree, and Factorization Machines utilizing Cross Validator and Training Validator functions for assessing performance and generalization capability.

cs.LG

CFPB Consumer Complaints Analysis Using Hadoop

Consumer complaints are a crucial source of information for companies, policymakers, and consumers alike. They provide insight into the problems faced by consumers and help identify areas for improvement in products, services, and regulatory frameworks. This paper aims to analyze Consumer Complaints Dataset provided by Consumer Financial Protection Bureau (CFPB) and provide insights into the nature and patterns of consumer complaints in the USA. We begin by describing the dataset and its features, including the types of complaints, companies involved, and geographic distribution. We then conduct exploratory data analysis to identify trends and patterns in the data, such as the most common types of complaints, the companies with the highest number of complaints, and the states with the most complaints. We have also performed descriptive and inferential statistics to test hypotheses and draw conclusions about the data. We have investigated whether there are significant differences in the types of complaints or companies involved based on geographic location. Overall, our analysis provides valuable insights into the nature of consumer complaints in the USA and helps stakeholders make informed decisions to improve the consumer experience.

cs.DC

Amazon Books Rating prediction & Recommendation Model

This paper uses the dataset of Amazon to predict the books ratings listed on Amazon website. As part of this project, we predicted the ratings of the books, and also built a recommendation cluster. This recommendation cluster provides the recommended books based on the column's values from dataset, for instance, category, description, author, price, reviews etc. This paper provides a flow of handling big data files, data engineering, building models and providing predictions. The models predict book ratings column using various PySpark Machine Learning APIs. Additionally, we used hyper-parameters and parameters tuning. Also, Cross Validation and TrainValidationSplit were used for generalization. Finally, we performed a comparison between Binary Classification and Multiclass Classification in their accuracies. We converted our label from multiclass to binary to see if we could find any difference between the two classifications. As a result, we found out that we get higher accuracy in binary classification than in multiclass classification.

cs.IR

Consumer's Behavior Analysis of Electric Vehicle using Cloud Computing in the State of New York

Sales of Electric Vehicles (EVs) in the United States have grown fast in the past decade. We analyze the Electric Vehicle Drive Clean Rebate data from the New York State Energy Research and Development Authority (NYSERDA) to understand consumer behavior in EV purchasing and their potential environmental impact. Based on completed rebate applications since 2017, this dataset features the make and model of the EV that consumers purchased, the geographic location of EV consumers, transaction type to obtain the EV, projected environmental impact, and tax incentive issued. This analysis consists of a mapped and calculated statistical data analysis over an established period. Using the SAP Analytics Cloud (SAC), we first import and clean the data to generate statistical snapshots for some primary attributes. Next, different EV options were evaluated based on environmental carbon footprints and rebate amounts. Finally, visualization, geo, and time-series analysis presented further insights and recommendations. This analysis helps the reader to understand consumers' EV buying behavior, such as the change of most popular maker and model over time, acceptance of EVs in different regions in New York State, and funds required to support clean air initiatives. Conclusions from the current study will facilitate the use of renewable energy, reduce reliance on fossil fuels, and accelerate economic growth sustainably, in addition to analyzing the trend of rebate funding size over the years and predicting future funding.

cs.CY

LA City Bike Lane Infrastructure and its Effects on Business Closures From 2016-2021

This paper analyzes the presence of bike lanes within the city of Los Angeles and how their presence could offer positive economic benefits to businesses. City wide shutdowns initiated in March of 2020 acted as a demand shock to local businesses. Over the next two years, businesses struggled to keep the doors open. We explored the possibility of bicycle infrastructure playing a role in supporting/insulating local businesses from closure and positive economic impacts. Our analysis shows that the businesses along bike lane in Los Angeles runs positively, and that COVID-19 is not related to, significantly.

physics.soc-ph

Spread of COVID-19: Adult Detention Facilities in LA County

We analyze the spread of COVID-19 cases within adult detention facilities in Los Angeles (LA) county. Throughout the analysis we review the data to explore the range of positive cases in each center and see the percentage of people who were positive for COVID-19 against the amount of people who were tested. Additionally, we see if there is any correlation between the surrounding community of each detention center and the number of positive cases in each center and explore the protocols in place at each detention center. We use the cloud visualization tool SAP Analytics Cloud (SAC) with the data from the California government website through adult detention facilities in LA County. We found that (1) the number of confirmed cases at the facilities and the surrounding communities are not related, (2) the data does not represent all positive cases at the facility, and (3) there are not enough tests at the facilities.

cs.DC

Crime Patterns in Los Angeles County Before and After Covid19 (2018-2021)

The objective of our research is to present the change in crime rates in Los Angeles post-Covid19. Using data analysis with Geo-Mapping, bubbles, Marimekko, and a time series charts, we can illustrate which areas have the largest crime rate, and how it has changed. Through regression modeling, we can interpret which locations may also have a correlation to crime versus income, race, type of crime, and gender. The story will help to uncover whether the areas associated with crime are due to demographic or income variance. In showing the details of crimes in Los Angeles along with the factors at play we hope to see a compelling relationship between crime rates and recent events from 2020 to the present, along with changes in crime type trends during these periods. We use Excel to clean the data for SAP SAC to model effectively, as well as resources from other studies a comparison.

cs.DC

Scalable Analysis for Covid-19 and Vaccine Data

This paper explains the scalable methods used for extracting and analyzing the Covid-19 vaccine data. Using Big Data such as Hadoop and Hive, we collect and analyze the massive data set of the confirmed, the fatality, and the vaccination data set of Covid-19. The data size is about 3.2 Giga-Byte. We show that it is possible to store and process massive data with Big Data. The paper proceeds tempo-spatial analysis, and visual maps, charts, and pie charts visualize the result of the investigation. We illustrate that the more vaccinated, the fewer the confirmed cases.

cs.DC

Scalable Traffic Predictive Analysis using GPU in Big Data

The paper adopts parallel computing systems for predictive analysis in both CPU and GPU leveraging Spark Big Data platform. The traffic dataset is adopted to predict the traffic jams in Los Angeles County. It is collected from a popular platform in the USA for tracking information on the road using the device information and reports shared by the users. Large-scale traffic data set can be stored and processed using both GPU and CPU in this Scalable Big Data systems. The major contribution of this paper is to improve the performance of machine learning in distributed parallel computing systems with GPU to predict the traffic congestion. We show that the parallel computing can be achieve using both GPU and CPU with the existing Apache Spark platform. Our method can be applicable to other large scale datasets in different domains. The process modeling, as well as results, are interpreted using computing time and metrics: AUC, Precision and Recall. It should help the traffic management in Smart City.

cs.DC

Scalable Predictive Time-Series Analysis of COVID-19: Cases and Fatalities

COVID 19 is an acute disease that started spreading throughout the world, beginning in December 2019. It has spread worldwide and has affected more than 7 million people, and 200 thousand people have died due to this infection as of Oct 2020. In this paper, we have forecasted the number of deaths and the confirmed cases in Los Angeles and New York of the United States using the traditional and Big Data platforms based on the Times Series: ARIMA and ETS. We also implemented a more sophisticated time-series forecast model using Facebook Prophet API. Furthermore, we developed the classification models: Logistic Regression and Random Forest regression to show that the Weather does not affect the number of the confirmed cases. The models are built and run in legacy systems (Azure ML Studio) and Big Data systems (Oracle Cloud and Databricks). Besides, we present the accuracy of the models.

cs.DC

Yelp Dataset Analysis using Scalable Big Data

Yelp has served and will continue to serve as a data-driven application. Yelp has published a dataset containing business information, reviews, user information, and check-in information. This paper will examine this dataset to provide descriptive analytics to understand business performance, geo-spatial distribution of businesses, reviewers' rating and other characteristics, and temporal distribution of check-ins in business premises. With these analysis we are able to establish that yelp reviews, tips, elite users and check ins have started to plummet over the years. Coincidentally, the paper also establishes that Canadians have a more stable star ratings as well as sentiment ratings when compared to Americans.

cs.DC

Avocado Buying Trends in the United States Using SAC

The purpose of our paper is to analyze the dataset from Hass Avocado Board (HAB). The data features historical data on avocado prices and sales volume in multiple cities, states, and regions of the United States, ranging from 2015 to 2020. The paper consists of a mapped and calculated statistical analysis of the data over an established period. Using the cloud visualization tool SAP Analytics Cloud (SAC) to import and clean the data, we will complete comprehensive investigations and employ class lecture information as needed. Then, we will present insights using visualization, and time-series analysis. Our research intends to reveal insights into consumer buying statistics, such as the most popular type of avocado, preferences of organic to conventional and seasonality trends. The research is relevant as health-conscious trends have become increasingly popular, and avocado purchases indicate this.

cs.DC