Twitter Classification of Complaints

Page created by Jean Floyd
 
CONTINUE READING
Twitter Classification of Complaints
© April 2023| IJIRT | Volume 9 Issue 11 | ISSN: 2349-6002

                     Twitter Classification of Complaints

             Aniket Solanki1, Harsh Tripathi2, Harsheet Soni3, Kunal Singh4, Naina Kaushik5
   1234
       Student Computer Engineering, MCT’s Rajiv Gandhi Institute of Technology, Mumbai University
               5
                 Professor, MCT’s Rajiv Gandhi Institute of Technology, Mumbai University

Abstract—This project addresses the problem of               negative response (since twitter allows us to download
sentiment analysis in twitter; that is classifying tweets    stream of geo-tagged tweets for particular locations. If
according to the sentiment expressed in them: positive,      firms can get this information they can analyze the
negative or neutral. Twitter is an online micro-blogging     reasons behind geographically differentiated response,
and social-networking platform which allows users to         and so they can market their product in a more
write short status updates of maximum length 140             optimized manner by looking for appropriate solutions
characters. It is a rapidly expanding service with over      like creating suitable market segments. Predicting the
200 million registered users out of which 100 million are    results of popular political elections and polls is also
active users and half of them log on twitter on a daily      an emerging application to sentiment analysis. One
basis - generating nearly 250 million tweets per day. Due
                                                             such study was conducted by Tumasjan et al. in
to this large amount of usage we hope to achieve a
                                                             Germany for predicting the outcome of federal
reflection of public sentiment by analysing the
                                                             elections in which concluded that twitter is a good
sentiments expressed in the tweets. Analysing the public
                                                             reflection         of         offline        sentiment.
sentiment is important for many applications such as
firms trying to find out the response of their products in
the market, predicting political elections and predicting
socioeconomic phenomena like stock exchange. The aim
of this project is to develop a functional classifier for
accurate and automatic sentiment classification of an
unknown tweet stream.
Index Terms— Sentiments, Predicting.

                   I. INTRODUCTION
We have chosen to work with twitter since we feel it is
a better approximation of public sentiment as opposed
to conventional internet articles and web blogs. The
reason is that the amount of relevant data is much
larger for twitter, as compared to traditional blogging
sites. Moreover the response on twitter is more prompt
                                                                         Fig 1: System Architecture
and also more general (since the number of users who
tweet is substantially more than those who write web                      II. LITERATURE REVIEW
blogs on a daily basis).
                                                             A. Literature Review
Sentiment analysis of public is highly critical in           Due to the enormous amount of data generated on
macro-scale     socioeconomic      phenomena       like      social media networks, Twitter sentiment analysis is a
predicting the stock market rate of a particular firm.       burgeoning field of study. An overview of the most
This could be done by analysing overall public               recent studies on Twitter sentiment analysis is
sentiment towards that firm with respect to time and         provided in this review of the literature.
using economics tools for finding the correlation
between public sentiment and the firm’s stock market         1)"Deep Learning Techniques for Sentiment Analysis
value. Firms can also estimate how well their product        on Twitter Data: A Review" by R. Selvi et al. (2021).
is responding in the market, which areas of the market       Deep learning techniques for sentiment analysis on
is it having a favourable response and in which a            Twitter are examined in detail in this work. The

IJIRT 159154               INTERNATIONAL JOURNAL OF INNOVATIVE RESEARCH IN TECHNOLOGY                           495
Twitter Classification of Complaints
© April 2023| IJIRT | Volume 9 Issue 11 | ISSN: 2349-6002

authors explore several pre-processing and feature         2) Lexalytics: A text analytics software provider called
extraction strategies in addition to comparing various     Lexalytics provides sentiment analysis solutions,
deep learning architectures. They also emphasize the       including a Twitter sentiment analysis solution. It
difficulties with Twitter sentiment analysis and offer     analyses text and identifies sentiment using machine
suggestions for further study.                             learning and natural language processing techniques.
2) "Sentiment Analysis of Twitter Data: A Complete         Several languages of text can be analysed using
Study," by P. Pakhare et al. (2021), offers a thorough     Lexalytics, which can also be tailored to certain
analysis of numerous methods such as lexicon-based,        fields.Text is given a subjectivity score and a
machine learning-based, and deep learning-based            sentiment polarity score by Lexalytics. The emotion of
techniques for sentiment analysis on Twitter data. The     the text is indicated by the polarity score, which is a
authors examine the difficulties sarcasm and irony         float value between -1 and 1. (negative to positive).
provide and make recommendations for future                3) CoreNLP: A sentiment analysis module is part of
research topics.                                           the Java-based natural language processing toolbox
3) First most recent techniques and uses of sentiment      known as Stanford CoreNLP. It analyses text using
analysis on Twitter are covered by M. I. Ali et al. in     deep learning techniques and assigns a sentiment
their paper "Twitter Sentiment Analysis: A Survey of       score. Stanford CoreNLP can be tailored to particular
Techniques and Applications" from 2020. They               fields and can analyse text in a variety of
contrast various strategies, including lexicon-based,      languages.Text is given a sentiment polarity score and
machine learning-based, and deep learning-based            a subjectivity score using Stanford CoreNLP. The
ones, and offer suggestions for further research.          emotion of the text is indicated by the polarity score,
4) In "Sentiment Analysis on Twitter Data: A               which is a float value.
Review," R. Kumar et al. (2020), they concentrate on       4) Vader Sentiment Analysis: The rule-based
the various methods for pre-processing, feature            sentiment analysis program Vader (Valence Aware
extraction, and sentiment classification. The authors      Dictionary and Sentiment Reasoner) employs a
talk about the difficulties in performing sentiment        dictionary of words and phrases to assign a sentiment
analysis on Twitter data and make recommendations          score to text. The lexicon includes positive and
for future study areas.                                    negative sentiment scores for words and phrases and
                                                           is based on human-labeled data. Moreover, Vader
The informality of tweets and the usage of irony and       considers negations, intensifiers, and other textual
sarcasm, according to the literature, make Twitter         elements that may have an impact on emotion.
sentiment analysis a difficult endeavour. Whereas          Vader has demonstrated strong performance on
recent developments in deep learning algorithms have       Twitter and other social media metrics. It features a
yielded encouraging results, future research should        straightforward API for text analysis and is available
concentrate on enhancing the precision of sentiment        as a Python package.
analysis using Twitter data.                               5) TextBlob: It is a Python package that offers a
                                                           straightforward API for carrying out typical natural
B. Survey of Existing Systems                              language processing operations, such as sentiment
1) IBM Watson: A sentiment analysis service is part        analysis. It categorises text as either positive, negative,
of IBM Watson, a collection of AI-powered products.        or neutral using a Naive Bayes classifier. The classifier
It analyses text using machine learning techniques and     can be applied to different domains and was trained on
assigns a sentiment score. IBM Watson may be               a labelled dataset of movie reviews.TextBlob also
tailored to particular domains and can evaluate text in    assigns a subjectivity and sentiment polarity score to
a variety of languages.Moreover, IBM Watson offers         each piece of text. The emotion of the text is indicated
a subjectivity score and a sentiment polarity score for    by the polarity score, which is a float value between -
text. The emotion of the text is indicated by the          1 and 1. (negative to positive).
polarity score, which is a float value between -1 and 1.
(negative to positive). The subjectivity score, which      C. linitation of Existing Systems/Research gap
ranges from 0 to 1, reflects how subjective                1) Sampling bias: Since Twitter users may not
(opinionated) or objective the text is (factual)           represent the entire population, the opinions posted

IJIRT 159154              INTERNATIONAL JOURNAL OF INNOVATIVE RESEARCH IN TECHNOLOGY                             496
Twitter Classification of Complaints
© April 2023| IJIRT | Volume 9 Issue 11 | ISSN: 2349-6002

there might not be reflective of the general public.        Data preprocessing and cleaning: After the data has
Compared to the average population, Twitter users           been gathered, it must be cleaned and preprocessed to
tend to be younger, more urban, and politically active.     weed out extraneous information and make it ready for
2) Limited context: Because Twitter tweets can only         analysis. There are multiple steps in this:
be 280 characters long, they frequently lack the            a. Eliminating duplicate tweets: It's critical to get rid
background information required to fully comprehend         of duplicate tweets before sentiment analysis because
the mood being expressed. An seeming negative tweet,        they can bias the results.
for instance, can actually be ironic or satirical.          b. Retweets can potentially bias the results of
3) Slang and colloquial language: Twitter users             sentiment analysis because they do not reflect the
frequently employ colloquial language, which can be         original sentiment.
challenging for sentiment analysis algorithms to            c. Removing spam: Spam tweets, which can be found
understand. This can result in incorrect sentiment          via automatic or manual approaches, can also skew the
analysis findings.                                          results.
4) Irony and sarcasm: Twitter users frequently employ       d. Eliminating non-textual content: To draw attention
irony and sarcasm to convey their thoughts, which           to text-only content, non-textual content such as
sentiment analysis algorithms may find challenging to       images, videos, and other media might be eliminated.
identify. Even though a tweet seems to be harsh, it         e. Text normalisation, often known as case-insensitive
could also be positive.                                     text or punctuation removal, is the process of
5) Sentiment polarity: Text is often categorised as         transforming text to a standard format.
either positive, negative, or neutral by algorithms that
analyse sentiment. But sentiment polarity is frequently     Classification of sentiment: Positive, negative, or
more nuanced than this and can change based on the          neutral sentiment is assigned to each tweet based on
situation and the speaker.                                  algorithms that analyse sentiment. There are various
6) Minimal training data: To learn how to classify text,    methods for sentiment analysis, such as:
sentiment analysis algorithms need training data.           a. Rule-based strategies: These strategies utilise a set
Nevertheless, there may not be enough training data         of pre-established rules to assess the tone of a tweet
for Twitter sentiment analysis, which could result in       based on its keywords or other characteristics.
unreliable findings.                                        b. Machine learning techniques: Using machine
7) Emojis and emoticons: Twitter users frequently           learning techniques, a model is trained on a labelled
utilise emojis and emoticons to convey emotion,             dataset to discover how to categorise text according to
which can be challenging for algorithms that analyse        its sentiment.
sentiment. Depending on the situation, the same emoji       c. Hybrid methods: To increase accuracy, hybrid
can imply something entirely different.                     approaches integrate rule-based and machine learning
8) Tweets in multiple languages are possible on             techniques.
Twitter because it is a multilingual platform. But, it's
possible that algorithms for sentiment analysis won't       Analyzing and visualising the results is the last step in
be able to effectively interpret tweets in languages        the Twitter sentiment analysis process. In order to do
they haven't been trained on.                               this, it may be necessary to discover trends and
                                                            patterns in the data, identify significant influencers or
                III. METHODOLOGY
                                                            subjects, and create graphs and charts to represent
Data Collection:Data gathering is the initial step in the   sentiment across time.
sentiment analysis process for Twitter. The Twitter
API, third-party programmes like Tweepy, or for-
profit services like Gnip or DataSift can all be used for
this. The information may be gathered in real-time or
over a predetermined period of time, and it may
comprise tweets on a certain subject or hashtag, as well
as tweets from particular individuals or places.

IJIRT 159154               INTERNATIONAL JOURNAL OF INNOVATIVE RESEARCH IN TECHNOLOGY                           497
Twitter Classification of Complaints
© April 2023| IJIRT | Volume 9 Issue 11 | ISSN: 2349-6002

                                                            —Fig 4:Live Tweet Sentiment Analysis of UEFA

             —Fig 2: System Architecture
Overall, there are a number of processes in the
approach for Twitter sentiment analysis, including
data collection, cleaning and preprocessing, sentiment
classification,     sentiment     aggregation,     and
visualisation and analysis. Depending on the study
issue and the objectives of the analysis, many
approaches may be employe.
                   IV.RESULTS
                                                                 —Fig 4: Entered Tweet Sentimental Analysis
                                                                              of UEFA
                                                                          V. CONCLUSION
                                                         In this paper we conclude that the proposed twitter
                                                         sentimental analysis addresses the critical issues of
                                                         getting the sentiments of a human from their text.The
                                                         use of tweepy api of python helps to fetch live data from
                                                         the twitter api and sentiment can be used for predictive
                                                         analysis: Twitter sentiment analysis can be used to
                                                         predict future trends or events. By analyzing the
                —Fig 3: Home Page
                                                         sentiment of tweets, researchers and businesses can
                                                         gain insight into the likelihood of a particular event
                                                         occurring or how consumers may react to a new
                                                         product or service. Machine learning algorithms can be
                                                         used to improve the accuracy of Twitter sentiment
                                                         analysis. By training the algorithms on large datasets,
                                                         they can learn to identify patterns in language and
                                                         sentiment that can help to improve the accuracy of the
                                                         analysis.
                                                                       ACKNOWLEDGMENT

                                                         We wish to express our sincere gratitude to Dr. Sanjay
                                                         U. Bokade, Principal and Prof. S. P. Khachane, Head
                                                         of Department of Computer Engineering at MCT's
             —Fig 3:Tweet Sentiment Analysis

IJIRT 159154             INTERNATIONAL JOURNAL OF INNOVATIVE RESEARCH IN TECHNOLOGY                          498
Twitter Classification of Complaints
© April 2023| IJIRT | Volume 9 Issue 11 | ISSN: 2349-6002

Rajiv Gandhi Institute of Technology, for providing us         Fukuoka, Japan, Cyber Science and Technology
with the opportunity to work on our project,                   Congress, 2019
"Crowdfunding Platform Using Blockchain." This            [8] M. Gamon, “Sentiment classification on customer
project would not have been possible without the               feedback data: noisydata, large feature vectors,
guidance and encouragement of our project guide,               and the role of linguistic analysis,” in COLING
Prof. Naina Kaushik. We would also like to thank our           2004: Proceedings of the 20th International
colleagues and friends who helped us in completing             Conference on Computational Linguistics, pp.
this project successfully.                                     841–847, 2004
                                                          [9] S. A. El Rahman, F. A. AlOtaibi, and W. A.
                    REFERENCES                                 AlShehri, “Sentiment analysis of twitter data,” in
[1] A. Hassan, A. Abbasi and D. Zeng, “Twitter                 2019 International Conference on Computer and
      sentiment analysis: A bootstrap ensemble                 Information Sciences (ICCIS), pp. 1–4, IEEE,
      framework,” in Proc. ASE/ IEEE Int. Conf. on             2019
      Social Computing, Alexandria, VA, USA, pp.          [10] A. Bermingham and A. F. Smeaton, “Classifying
      357–364, 2013.                                           sentiment in microblogs: is brevity an
[2]   N. Yadav, O. Kudale, S. Gupta, A. Rao, and A.            advantage?,” in Proceedings of the 19th ACM
      Shitole, “Twitter sentiment analysis using               international conference on Information and
      machine learning for product evaluation,”                knowledge management, pp. 1833–1836, 2010
      in2020 International Conference on Inventive        [11] B. Bhavitha, A. P. Rodrigues, and N. N.
      Computation Technologies (ICICT), pp. 181–               Chiplunkar, “Comparativestudy of machine
      185, IEEE, 2020                                          learning techniques in sentimental analysis,”
[3]   N. F. F. Da Silva and E. R. Hruschka, “Tweet             in2017 International Conference on Inventive
      sentiment analysis with classifier ensembles,”           Communication andComputational Technologies
      Decision Support Systems, vol. 66, no. C, pp.            (ICICCT), pp. 216–221, IEEE, 2017.
      170–179, 2014.                                      [12] L. M. Rojas-Barahona, “Deep learning for
[4]   C. Yang, K. H.-Y. Lin, and H.-H. Chen, “Emotion          sentiment analysis,”Language and Linguistics
      classification using web blog corpora,” in                Compass, vol. 10, no. 12, pp. 701–719,2016.
      IEEE/WIC/ACM         International     Conference   [13] A. Severyn and A. Moschitti, “Unitn: Training
      onWeb Intelligence (WI’07), pp. 275–278, IEEE,           deep convolutionalneural network for twitter
      2007.                                                    sentiment classification,” in Proceedingsof the 9th
[5]   A. P. Patel, A. V. Patel, S. G. Butani and P. B.         international workshop on semantic evaluation
      Sawant, "Literature Survey on Sentiment                  (SemEval2015), pp. 464–469, 2015.
      Analysis of Twitter Data using Machine              [14] M. Cliche, “Bb twtr at semeval-2017 task 4:
      Learning      Approaches", IJIRST-International          Twitter sentiment analysis with cnns and lstms,”
      journal for Innovative Reasearch in Scinece &            arXiv preprint arXiv:1704.06125, 2017.
      Technology, vol. 3, no. 10, 2017                    [15] C. Baziotis, N. Pelekis, and C. Doulkeridis,
[6]   N. F. Alshammari and A. A. AlMansour, “State-            “Datastories atsemeval-2017 task 4: Deep lstm
      of-the-art review on twitter sentiment analysis,”        with attention for message-leveland topic-based
      in 2019 2nd International Conference on                  sentiment analysis,” in Proceedings of the
      Computer Applications & Information Security             11thinternational workshop on semantic
      (ICCAIS), pp. 1–8,IEEE, 2019.                            evaluation (SemEval-2017),pp. 747–754, 2017.
[7]   J. Ho, D. Ondusko, B. Roy and D. F. Hsu,            [16] M.-H. Su, C.-H. Wu, K.-Y. Huang, and Q.-B.
      “Sentiment analysis on tweets using machine              Hong, “Lstm-based textemotion recognition
      learning and combinatorial fusion,” in Int. Conf.        using semantic and emotional word vectors,”
      on Dependable, Autonomic and Secure                      in2018 First Asian Conference on Affective
      Computing,      Pervasive     Intelligence    and        Computing and IntelligentInteraction (ACII
      Computing, Cloud and Big Data Computing,                 Asia), pp. 1–6, IEEE, 2018.[28] Z. Jianqiang, G.
                                                               Xiaolin, and Z. Xuejun, “Deep convolution

IJIRT 159154              INTERNATIONAL JOURNAL OF INNOVATIVE RESEARCH IN TECHNOLOGY                        499
© April 2023| IJIRT | Volume 9 Issue 11 | ISSN: 2349-6002

   neuralnetworks for twitter sentiment analysis,”
   IEEE Access, vol. 6,pp. 23253–23260, 2018.

IJIRT 159154           INTERNATIONAL JOURNAL OF INNOVATIVE RESEARCH IN TECHNOLOGY   500
You can also read