Preparation of Improved Turkish DataSet for Sentiment Analysis in Social Media

Semiha Makinist; İbrahim Rıza Hallaç; Betül Ay Karakuş; Galip Aydın

doi:10.1051/itmconf/20171301030

Open Access

Issue		ITM Web Conf. Volume 13, 2017 2^nd International Conference on Computational Mathematics and Engineering Sciences (CMES2017)


Article Number		01030
Number of page(s)		6
DOI		https://doi.org/10.1051/itmconf/20171301030
Published online		02 October 2017

ITM Web of Conferences 13, 01030 (2017)

Preparation of Improved Turkish DataSet for Sentiment Analysis in Social Media

Semiha Makinist¹^*, İbrahim Rıza Hallaç²^*, Betül Ay Karakuş²^* and Galip Aydın²^*

¹ Sentis Software
² Department of Computer Engineering, Firat University, Elazig, Turkey

^⁎ Semiha Makinist: This email address is being protected from spambots. You need JavaScript enabled to view it.
^⁎ İbrahim Rıza Hallaç: This email address is being protected from spambots. You need JavaScript enabled to view it.
^⁎ Beetül Ay Karakuş: This email address is being protected from spambots. You need JavaScript enabled to view it.
^⁎ Galip Aydın: This email address is being protected from spambots. You need JavaScript enabled to view it.

Abstract

A public dataset, with a variety of properties suitable for sentiment analysis [1], event prediction, trend detection and other text mining applications, is needed in order to be able to successfully perform analysis studies. The vast majority of data on social media is text-based and it is not possible to directly apply machine learning processes into these raw data, since several different processes are required to prepare the data before the implementation of the algorithms. For example, different misspellings of same word enlarge the word vector space unnecessarily, thereby it leads to reduce the success of the algorithm and increase the computational power requirement. This paper presents an improved Turkish dataset with an effective spelling correction algorithm based on Hadoop [2]. The collected data is recorded on the Hadoop Distributed File System and the text based data is processed by MapReduce programming model. This method is suitable for the storage and processing of large sized text based social media data. In this study, movie reviews have been automatically recorded with Apache ManifoldCF (MCF) [3] and data clusters have been created. Various methods compared such as Levenshtein and Fuzzy String Matching have been proposed to create a public dataset from collected data. Experimental results show that the proposed algorithm, which can be used as an open source dataset in sentiment analysis studies, have been performed successfully to the detection and correction of spelling errors.

Key words: Sentiment Analysis / Natural Language Processing / Hadoop / Turkish Data Set

This is an Open Access article distributed under the terms of the Creative Commons Attribution License 4.0, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.

Current usage metrics show cumulative count of Article Views (full-text article views including HTML views, PDF and ePub downloads, according to the available data) and Abstracts Views on Vision4Press platform.

Data correspond to usage on the plateform after 2015. The current usage metrics is available 48-96 hours after online publication and is updated daily on week days.

Initial download of the metrics may take a while.