2019 IEEE 4th International Conference on Computer and Communication Systems (ICCCS) · 2019 · 17 citations · 15 references
EngineeringCorpus LinguisticsText MiningWord EmbeddingsNatural Language ProcessingBook Genre ClassificationClassification MethodAlgorithmic ComparisonsInformation RetrievalData ScienceData MiningComputational LinguisticsDocument ClassificationLanguage StudiesContent AnalysisAutomatic ClassificationKnowledge DiscoveryTerminology ExtractionIntelligent ClassificationData ClassificationText ProcessingLinguistics
This paper presents algorithmic comparisons for producing a book's genre based on its title. While some titles are easy to interpret, some are irrelevant to the genre that they belong to. Henceforth, we seek to determine the optimal and most accurate method for accomplishing the task. Several data preprocessing steps were implemented, in which word embeddings were created to make the titles operable by the computer. Five different machine learning models were tested throughout the experiment. Each different algorithm was fine-tuned for attaining the best parameter values, while no modifications were conducted on the dataset. The results indicate that the Long Short-Term Memory (LSTM) with a dropout is the top performing architecture among the algorithms, with an accuracy of 65.58%. To the authors' knowledge, no prior study has been done about book genre classification by title, therefore the present study is the current best in the field.
15
Sepp Hochreiter, Jürgen Schmidhuber · Neural Computation · 1997 · 93.8K citations
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey E. Hinton, Alex Krizhevsky et al. · 2014 · 34.2K citations