Metric functions¶
Classic nausea¶
ests.style_stats.calc_classic_nausea()
The classic nausea of Advego - the square root of the number of occurrences of the most frequent word. It measures how insistent one word is regardless of the length of the text, so it grows with the text. The norm of Advego is at most 7, 1-5 in practice.
Formula:
where \(n_k\) is the number of occurrences of the word \(k\).
| Parameter | Type | Default | Description |
|---|---|---|---|
text |
list[str] | - |
List of words |
Academic nausea¶
ests.style_stats.calc_academic_nausea()
The academic nausea of Advego - the share of the occurrences of the most frequent words of the text in percent. The exact formula of Advego is not published: the summed frequency of the top_n most frequent words is divided by the number of words. The norm of Advego is 5-15%.
Formula:
| Parameter | Type | Default | Description |
|---|---|---|---|
text |
list[str] | - |
List of words |
top_n |
int | 10 |
Number of the most frequent words |
Water content¶
ests.style_stats.calc_water()
The water content of Text.ru - the share of the words that carry no content in percent: the stopwords of is_stopword or of the list passed, in any case. The norms of Text.ru - up to 15% natural, 15-30% excessive, above 30% high - are set for Russian, which has no articles; a Spanish text has more water by its grammar alone.
Formula:
| Parameter | Type | Default | Description |
|---|---|---|---|
text |
list[str] | - |
List of words |
stopwords |
list[str]/set[str] | None |
List or set of stopwords; if not given, is_stopword is used |
Stopword¶
ests.style_stats.is_stopword()
Checks whether a word is a stopword, in any case. The stopwords are the 265 forms of ests.constants.STOPWORDS - the closed classes of the grammar: articles and the other determiners (este, cada, mucho), pronouns (él, cuyo, nadie), prepositions, conjunctions and interjections, with the adverbs that point, relate or ask (aquí, así, dónde) and the ones that negate, affirm or focus (no, solo, también) - and the one-word parenthetical expressions of PARENTHETICALS (finalmente, naturalmente). The forms of the old orthography are included (á, ó, tí).
| Parameter | Type | Default | Description |
|---|---|---|---|
word |
str | - |
Word |
Spam score¶
ests.style_stats.calc_spam()
The spam score of Text.ru - the share of the repeated words of the text in percent: every occurrence of a word but the first is a repetition, so the spam score is \(100 \times (1 - TTR)\). To compute it by lemmas, extract the words with lemmatization. The norms of Text.ru: up to 30% natural, 30-60% SEO-optimized, above 60% spammed.
Formula:
where \(N\) is the number of words and \(V\) the number of distinct words.
| Parameter | Type | Default | Description |
|---|---|---|---|
text |
list[str] | - |
List of words |
Naturalness by Zipf's law¶
ests.style_stats.calc_zipf_naturalness()
How well the frequencies of the most frequent words agree with the ideal distribution \(f_r = f_1 / r\) of Zipf's law, where \(f_1\) is the frequency of the most frequent word and \(r\) is the rank of a word (pr-cy, megaindex). It is 100 times one minus the mean relative deviation of the frequencies from the ideal ones over the ranks from 2 to \(R = \min(top\_n, V, f_1)\). Negative values are clipped to 0; the norm of the services is at least 50%. It is nan when there are no ranks to compare: every word is a hapax, the text has one word type or top_n is below 2.
Formula:
| Parameter | Type | Default | Description |
|---|---|---|---|
text |
list[str] | - |
List of words |
top_n |
int | 10 |
Number of the most frequent words |
Keyword density¶
ests.style_stats.calc_keyword_density()
The frequency of every keyword per 100 words of the text (Text.ru). A keyword of several words separated by spaces is looked for as a sequence of words, and its occurrences may overlap. The words are compared in any case; to compare by lemmas, extract the words with lemmatization and pass lemmas.
| Parameter | Type | Default | Description |
|---|---|---|---|
text |
list[str] | - |
List of words |
keywords |
list[str]/set[str] | - |
Keywords or phrases |
Verbal nouns¶
ests.style_stats.calc_verbal_nouns()
The share of the nouns derived from a verb among the lemmas of the nouns of a text in percent, nan for a text without nouns. A noun is derived from a verb by its suffix - -ción, -sión, -miento, -anza, -encia, -ancia, -aje, -dura, -azgo (revisión, nombramiento, aprendizaje) - or is one of the nouns whose derivation leaves no suffix behind (uso, pago, envío), as in the split predicates of SyntaxStats. The rule also catches the nouns of other origins with the same endings (ciencia, distancia). The function takes the lemmas of the tokens tagged NOUN, not the words.
Example
from ests.style_stats import calc_verbal_nouns
calc_verbal_nouns(["revisión", "proyecto", "nombramiento", "casa"])
# 50.0
| Parameter | Type | Default | Description |
|---|---|---|---|
nouns |
list[str] | - |
Lemmas of the nouns |
Phrase density¶
ests.style_stats.calc_phrase_density(), ests.style_stats.expand_phrases()
The number of occurrences of the phrases of a list per 100 words - the compound prepositions (COMPOUND_PREPOSITIONS), the parenthetical expressions (PARENTHETICALS) and the clichés (OFFICIALESE_CLICHES). At every position the longest phrase is taken, and the phrases found do not overlap. expand_phrases spells the phrases out in the forms of the text: a phrase ending in a or de also takes the contraction with the article (a efectos del, conforme al), and a phrase whose first word is an infinitive takes the forms of the text with that lemma (proceder a - procedió a, ser de aplicación - es de aplicación). The lemma is the one of simplemma, and a pronominal lemma counts for its verb (llévese - llevar); the forms simplemma does not lemmatize - the irregular participles (ha dado, ha hecho) and the imperative dese - are given by IRREGULAR_VERB_FORMS. The words after the verb rule out the readings as a noun (el hecho, el puesto). When two phrases spell out the same words, a phrase written so wins over the forms of another (a efectos del stays itself next to a efectos de), and then the first phrase of a list, or of a set in sorted order.
The compound prepositions (42) and the clichés (74) are the ones the Spanish guides to plain language and style manuals of the administrations flag:
Left out are the forms the guides recommend (sobre la base de) or accept (de acuerdo a) and the phrases with frequent neutral uses (en este sentido), but for proceder a and llevar a cabo. A few compound prepositions have neutral uses too (a través de, en caso de, con respecto a), so the density of a neutral text is not zero: compare texts with each other.
| Parameter | Type | Default | Description |
|---|---|---|---|
text |
list[str] | - |
List of words |
phrases |
list[str]/set[str] | - |
Phrases, words separated by spaces |
Parenthetical expressions¶
ests.style_stats.calc_parentheticals(), ests.style_stats.is_parenthetical()
The parenthetical expressions of PARENTHETICALS (sin embargo, es decir, por ejemplo, finalmente) per 100 words; is_parenthetical checks a word against the one-word expressions of the list. The punctuation is not looked at, so the list holds the expressions set off by commas or standing at the edge of a sentence in most of their occurrences in the corpus of literature, with the series of order (en primer lugar, en segundo lugar). The adverbs of doubt (tal vez, quizá), sobre todo, sin duda and además are not in the list.
| Parameter | Type | Default | Description |
|---|---|---|---|
text |
list[str] | - |
List of words |