Skip to content

Basic statistics

anyts.basic_stats.BasicStats, anyts.basic_stats.count_punctuations()

Description

The basic statistics of a text: the counts of its sentences, words, characters, letters, spaces, syllables and punctuation marks, the distributions of the words by letters and by syllables, and the long, complex, simple, monosyllabic and polysyllabic words. The data source can be either a text or a Doc object of spaCy. Letters are counted by str.isalpha, so digits, hyphens and marks inside a word are not letters.

The words of a string come from the word extractor and its sentences from the sentence extractor. The words of a Doc come from its tokens (punctuation marks and symbols such as € or % dropped) and its sentences from its boundaries, or from the sentence extractor when it has none; an extractor passed explicitly is used on the text of the Doc. A sentence of punctuation alone holds no word and is not counted.

Language hooks

A language library subclasses BasicStats:

Hook Kind Description
count_syllables(word) method The syllables of a word; an abstract method, since syllables depend on the language
sents_extractor_class, words_extractor_class class attributes The extractors used when none is given; SentsExtractor and WordsExtractor of the core by default
join_hyphens class attribute Join the parts of the hyphenated words of a Doc (iter_doc_words); False by default
count_punctuations(text) method The punctuation marks by type; count_punctuations of the module by default
stats_desc, stats_headers class attributes The descriptions of the statistics and the headers of the columns for print_stats

The thresholds of the complex and the long words are parameters; a library sets its own defaults in its __init__ and passes them on.

Parameters

Parameter Type Default Description
source str/Doc - Data source (a string or a Doc object)
sents_extractor SentsExtractor None Sentence extraction tool; used for a Doc too, on its text
words_extractor WordsExtractor None Word extraction tool; used for a Doc too, on its text
normalize bool False Compute normalized statistics
complex_syl_factor int COMPLEX_SYL_FACTOR Minimum number of syllables in a complex word
long_word_letter_factor int LONG_WORD_LETTER_FACTOR Minimum number of letters in a long word

A source that is neither a string nor a Doc and an extractor of another type raise SourceTypeError, a source without words SourceError, a threshold that is not an integer of at least one ParameterError.

The thresholds of the core are COMPLEX_SYL_FACTOR = 3 and LONG_WORD_LETTER_FACTOR = 7 of anyts.constants.

Attributes

Attribute Type Description
c_letters dict[int, int] Distribution of words by number of letters
c_syllables dict[int, int] Distribution of words by number of syllables
n_sents int Number of sentences containing words
n_words int Number of words
n_unique_words int Number of unique words, case ignored
n_long_words int Number of long words
n_complex_words int Number of complex words
n_simple_words int Number of simple words: with a syllable, below the complex ones
n_monosyllable_words int Number of monosyllabic words
n_polysyllable_words int Number of polysyllabic words
n_chars int Number of characters without the line breaks
n_letters int Number of letters
n_spaces int Number of spaces and tabs
n_syllables int Number of syllables
n_punctuations int Number of punctuation marks
c_punctuations dict[str, int] Distribution of punctuation marks by type
p_unique_words float Normalized number of unique words
p_long_words float Normalized number of long words
p_complex_words float Normalized number of complex words
p_simple_words float Normalized number of simple words
p_monosyllable_words float Normalized number of monosyllabic words
p_polysyllable_words float Normalized number of polysyllabic words
p_letters float Normalized number of letters
p_spaces float Normalized number of spaces
p_punctuations float Normalized number of punctuation marks

The counts of the words are normalized by the number of words, those of the characters by the number of characters; the normalized statistics are computed with normalize=True.

Methods

count_words_by_syllables, count_words_by_letters

count_words_by_syllables(min_syllables) and count_words_by_letters(min_letters) return the number of words with at least the given number of syllables or letters; a minimum that is not an integer raises ParameterError.

get_stats

Returns a dictionary with the computed statistics - a copy: editing it does not change the object.

Prints a table with the counts of the statistics, their descriptions from stats_desc and the headers from stats_headers.

Punctuation marks

count_punctuations(text, marks, dash_pattern) counts punctuation marks by their types - the same distribution lives in the c_punctuations attribute: commas, periods, question and exclamation marks, ellipses, colons, semicolons, dashes, hyphens, guillemets, quotation marks, parentheses and the other marks. The dashes typed with hyphens (dash_pattern) are counted first, then the ellipses - the … character, three or more periods, or two periods after ? and !, one mark whose periods do not count as periods (Who?.. is a question and an ellipsis) - and then every mark by its type in marks; any other character of the Unicode categories P and S is another mark (§, €, °). marks maps every mark, a single character, to one of the types, otherwise ParameterError, and dash_pattern is a compiled regular expression; a text that is not a string, marks that are not a mapping and dashes that are not a compiled expression raise SourceTypeError.

In the core the types are anyts.constants.PUNCTUATION_TYPES, and marks and dash_pattern are PUNCTUATION_MARKS and DASH_PATTERN of anyts.basic_stats. The marks count the inverted ¿ and ¡ as question and exclamation marks (¿So? carries two question marks), —, – and the horizontal bar ― as dashes, a hyphen inside a word, before a digit or at the end of a line inside a word as a hyphen (well-known, -5), «» as guillemets and "“”‘’ as quotation marks. A dash typed with hyphens is a run of two or more hyphens, a hyphen after whitespace, at the start of a line or after a closing mark, before a space, a tab or the end of the text, or between a letter and an opening or a closing mark (--Hello --said John, - They left - he said).

Usage example

A subclass with the syllables of a language - here the groups of vowels, a rough rule for English.

Example

import re

from anyts.basic_stats import BasicStats


class Stats(BasicStats):
    def count_syllables(self, word):
        return len(re.findall(r"[aeiouy]+", word, re.IGNORECASE))


bs = Stats("The cat sat on the mat. A beautiful day!")
bs.n_sents, bs.n_words, bs.n_syllables, bs.n_complex_words
# (2, 9, 11, 1)
bs.c_punctuations["exclamation"]
# 1