Text Analytics Fundamentals – Introduction to Text Analytics with R Part 2

Text analytics fundamentals covers:

– The importance of splitting data in to training and test datasets
– Stratified sampling of imbalanced data using the caret package
– Representing text data for the purposes of machine learning
– Introduction to tokenization, stop words, and stemming
– The bag-of-words model
– Considerations for data pre-processing

Kaggle Dataset can be found here

The data and R code used in this series is available here


About The Author
- Data Science Dojo is a paradigm shift in data science learning. We enable all professionals (and students) to extract actionable insights from data.


