Skip to content

Word extraction

anyts.extractors.WordsExtractor

Description

A class for extracting words from a text. It allows using different tokenizers, filtering stop words, numbers and punctuation, lemmatizing, building N-grams, and setting the minimum and maximum length of extracted words.

Language hooks

The language enters the extractor through three hooks that a language library overrides in a subclass:

Hook Kind Default Used for
tokenize(text) method a word character \w followed by word characters, combining marks, zero-width joiners and non-joiners and soft hyphens, so that a word in NFD, in Devanagari or in Persian with U+200C stays whole the tokenizer when tokenizer is not given
lemmatize(word) method the word itself use_lexemes=True
number_pattern class attribute a signed number with separators and an optional percent sign: -5, 1990-1995, 1,500.50, 12/03/2020, 3:30, 10% filter_nums=True, matched against the whole lower-cased word

Example

import re

from anyts import WordsExtractor


class EnglishWords(WordsExtractor):
    number_pattern = re.compile(r"\d+(?:st|nd|rd|th)?")

    def tokenize(self, text):
        return text.split()


EnglishWords(filter_nums=True).extract("The 3rd time")

Result:

('The', 'time')

Parameters

Parameter Type Default Description
tokenizer Pattern/Callable None Tokenizer or regular expression; by default the tokenize method
filter_punct bool True Filter punctuation marks
filter_nums bool False Filter numbers matched by number_pattern
use_lexemes bool False Use word lemmas of the lemmatize method
stopwords Collection[str] None Stop words, compared case-insensitively
lowercase bool False Convert words to lower case
ngram_range Tuple[int, int] (1, 1) Lower and upper bound of the N-gram size
min_len int 0 Minimum length of an extracted word in characters, 0 for no bound
max_len int 0 Maximum length of an extracted word in characters, 0 for no bound

Note

A regular expression as the tokenizer is a separator: the text is split with re.split. The filters are applied in order: punctuation, numbers, lemmatization, lower case, stop words, word length. A lower-case stop word list also filters a capitalized word at the start of a sentence. A punctuation mark is a token consisting entirely of marks and symbols, including multi-character ones: ?!, --, …, € (see anyts.utils.is_punctuation). Empty and whitespace tokens are dropped before the filters. N-grams join the words with _.

Methods

extract

Extracts words from a text.

Parameter Type Default Description
text str - Text string

An example of word extraction with bigrams as tokens, after filtering numbers and stop words:

Example

Code:

from anyts import WordsExtractor

text = "Better 100 friends than 100 dollars"

we = WordsExtractor(lowercase=True, stopwords=["than"], filter_nums=True, ngram_range=(1, 2))
we.extract(text)

Result:

('better', 'friends', 'dollars', 'better_friends', 'friends_dollars')

get_most_common

Returns the top words of the text as a list of (word, frequency) pairs, the most frequent first.

Parameter Type Default Description
n int 10 Number of top words

Warning

The method must be called after words have been extracted with extract.

Example

Code:

from anyts import WordsExtractor

we = WordsExtractor(lowercase=True)
we.extract("The cat saw the dog and the bird")
we.get_most_common(2)

Result:

[('the', 3), ('cat', 1)]