In Python tokenization basically refers to splitting up a larger body of text into smaller lines, words or even creating words for a non-English language.
How do you use Tokenize in Python?
The Natural Language Tool kit(NLTK) is a library used to achieve this. Install NLTK before proceeding with the python program for word tokenization. Next we use the word_tokenize method to split the paragraph into individual words. When we execute the above code, it produces the following result.
What does NLTK Tokenize do?
NLTK contains a module called tokenize() which further classifies into two sub-categories: Word tokenize: We use the word_tokenize() method to split a sentence into tokens or words. Sentence tokenize: We use the sent_tokenize() method to split a document or paragraph into sentences.