Skip to content

Character N-gram extraction

anyts.extractors.CharNgramsExtractor

Description

A class for extracting character N-grams from a text - sequences of N characters taken with a sliding window over the string. Character N-grams are a standard feature of stylometry and authorship attribution (Stamatatos 2009) and can replace words as the units of a text in measures such as Burrows's Delta.

Whitespace runs are collapsed into a single space beforehand, punctuation marks are kept. With within_words=True N-grams do not cross word boundaries: the text is split into words by the tokenizer, punctuation is dropped, and words shorter than N yield no N-grams.

Language hooks

The default word tokenizer for within_words is the method tokenize(text), which a language library overrides in a subclass. Here it is the default tokenizer of WordsExtractor: a word character \w followed by word characters, combining marks, zero-width joiners and non-joiners and soft hyphens, so don't gives the words don and t.

Parameters

Parameter Type Default Description
n int 2 N-gram length in characters
lowercase bool False Convert the text to lower case
within_words bool False Take N-grams only inside words
tokenizer Pattern/Callable None Word tokenizer for within_words or a regular expression; by default the tokenize method

Methods

extract

Extracts N-grams from a text.

Parameter Type Default Description
text str - Text string

Example

Code:

from anyts import CharNgramsExtractor

text = "The cat slept  on the sill, and the dog - on the floor."

CharNgramsExtractor(n=3, lowercase=True).extract(text)[:6]

Result:

('the', 'he ', 'e c', ' ca', 'cat', 'at ')

N-grams inside words only:

Example

Code:

CharNgramsExtractor(n=4, lowercase=True, within_words=True).extract(text)

Result:

('slep', 'lept', 'sill', 'floo', 'loor')

get_most_common

Returns the top N-grams of the text as a list of (N-gram, frequency) pairs, the most frequent first.

Parameter Type Default Description
n int 10 Number of top N-grams

Warning

The method must be called after N-grams have been extracted with extract.

Example

Code:

ce = CharNgramsExtractor(n=3, lowercase=True)
ce.extract(text)
ce.get_most_common(2)

Result:

[('the', 4), ('he ', 4)]