Skip to content

Sentence lengths

anyts.visualizers.sentence_lengths_plot(), anyts.visualizers.sentence_lengths()

Description

The curve of the lengths of the sentences - the rhythm of a text: the length of every sentence in words in order, the moving average over a window of window sentences and an inset with the histogram of the lengths. Short and long sentences in turn are an editorial sign of a lively text, a flat curve - of a monotonous one. The function takes the axes ax and returns Axes.

sentence_lengths(source) extracts the lengths: a string is split into sentences by the sentence extractor and every sentence into words by the word extractor; the sentences of a Doc come from its boundaries and its words from iter_doc_words with join_hyphens, while a Doc without boundaries is counted as its text, by the extractors; sentences without words are skipped. Ready lengths - a sequence or an iterator of integers that are not negative - are used as they are; a table, a set, a mapping, bytes or a length that is not an integer raise SourceTypeError, a negative length SourceError.

Parameters

Parameter Type Default Description
source str/Doc/Iterable[int] - Text, Doc object or lengths of the sentences (a list, a numpy array, a Series)
window int 10 Window of the moving average in sentences
inset bool True Show the inset with the histogram
ax Axes None Axes for the plot
labels dict[str, str] None Labels over the defaults: title, xlabel, ylabel, length (the curve), average (a format string with window), distribution (the inset)
sents_extractor SentsExtractor None Extractor of the sentences of a string; the sentence extractor of the library by default
words_extractor WordsExtractor None Extractor of the words of a sentence of a string; the word extractor of the library by default

The core functions also take join_hyphens, False by default: join the parts of hyphenated words of a Doc with sentence boundaries (iter_doc_words); a library whose tokenizer keeps hyphenated words whole sets it. The extractors of the core are SentsExtractor() and WordsExtractor().

Usage example

Example

from anyts.visualizers import sentence_lengths, sentence_lengths_plot

text = "The cat sleeps. Does the dog? It eats, she said. Then both of them went out."
sentence_lengths(text)
# [3, 3, 4, 6]
sentence_lengths_plot(text, window=2)