Skip to content

Metric functions

Classic nausea

ests.style_stats.calc_classic_nausea()

The classic nausea of Advego - the square root of the number of occurrences of the most frequent word. It measures how insistent one word is regardless of the length of the text, so it grows with the text. The norm of Advego is at most 7, 1-5 in practice.

Formula:

\[ \sqrt{\max_k n_k} \]

where \(n_k\) is the number of occurrences of the word \(k\).

Parameter Type Default Description
text list[str] - List of words

Academic nausea

ests.style_stats.calc_academic_nausea()

The academic nausea of Advego - the share of the occurrences of the most frequent words of the text in percent. The exact formula of Advego is not published: the summed frequency of the top_n most frequent words is divided by the number of words. The norm of Advego is 5-15%.

Formula:

\[ 100\times\frac{\sum_{k \in \textrm{top}_n} n_k}{N} \]
Parameter Type Default Description
text list[str] - List of words
top_n int 10 Number of the most frequent words

Water content

ests.style_stats.calc_water()

The water content of Text.ru - the share of the words that carry no content in percent: the stopwords of is_stopword or of the list passed, in any case. The norms of Text.ru - up to 15% natural, 15-30% excessive, above 30% high - are set for Russian, which has no articles; a Spanish text has more water by its grammar alone.

Formula:

\[ 100\times\frac{\textrm{Number of stopwords}}{\textrm{Number of words}} \]
Parameter Type Default Description
text list[str] - List of words
stopwords list[str]/set[str] None List or set of stopwords; if not given, is_stopword is used

Stopword

ests.style_stats.is_stopword()

Checks whether a word is a stopword, in any case. The stopwords are the 265 forms of ests.constants.STOPWORDS - the closed classes of the grammar: articles and the other determiners (este, cada, mucho), pronouns (él, cuyo, nadie), prepositions, conjunctions and interjections, with the adverbs that point, relate or ask (aquí, así, dónde) and the ones that negate, affirm or focus (no, solo, también) - and the one-word parenthetical expressions of PARENTHETICALS (finalmente, naturalmente). The forms of the old orthography are included (á, ó, tí).

Parameter Type Default Description
word str - Word

Spam score

ests.style_stats.calc_spam()

The spam score of Text.ru - the share of the repeated words of the text in percent: every occurrence of a word but the first is a repetition, so the spam score is \(100 \times (1 - TTR)\). To compute it by lemmas, extract the words with lemmatization. The norms of Text.ru: up to 30% natural, 30-60% SEO-optimized, above 60% spammed.

Formula:

\[ 100\times\frac{N - V}{N} \]

where \(N\) is the number of words and \(V\) the number of distinct words.

Parameter Type Default Description
text list[str] - List of words

Naturalness by Zipf's law

ests.style_stats.calc_zipf_naturalness()

How well the frequencies of the most frequent words agree with the ideal distribution \(f_r = f_1 / r\) of Zipf's law, where \(f_1\) is the frequency of the most frequent word and \(r\) is the rank of a word (pr-cy, megaindex). It is 100 times one minus the mean relative deviation of the frequencies from the ideal ones over the ranks from 2 to \(R = \min(top\_n, V, f_1)\). Negative values are clipped to 0; the norm of the services is at least 50%. It is nan when there are no ranks to compare: every word is a hapax, the text has one word type or top_n is below 2.

Formula:

\[ \max\left(0,\ 100\times\left(1 - \frac{1}{R-1}\sum_{r=2}^{R}\frac{|f_r - f_1/r|}{f_1/r}\right)\right) \]
Parameter Type Default Description
text list[str] - List of words
top_n int 10 Number of the most frequent words

Keyword density

ests.style_stats.calc_keyword_density()

The frequency of every keyword per 100 words of the text (Text.ru). A keyword of several words separated by spaces is looked for as a sequence of words, and its occurrences may overlap. The words are compared in any case; to compare by lemmas, extract the words with lemmatization and pass lemmas.

Parameter Type Default Description
text list[str] - List of words
keywords list[str]/set[str] - Keywords or phrases

Verbal nouns

ests.style_stats.calc_verbal_nouns()

The share of the nouns derived from a verb among the lemmas of the nouns of a text in percent, nan for a text without nouns. A noun is derived from a verb by its suffix - -ción, -sión, -miento, -anza, -encia, -ancia, -aje, -dura, -azgo (revisión, nombramiento, aprendizaje) - or is one of the nouns whose derivation leaves no suffix behind (uso, pago, envío), as in the split predicates of SyntaxStats. The rule also catches the nouns of other origins with the same endings (ciencia, distancia). The function takes the lemmas of the tokens tagged NOUN, not the words.

Example

from ests.style_stats import calc_verbal_nouns

calc_verbal_nouns(["revisión", "proyecto", "nombramiento", "casa"])
# 50.0
Parameter Type Default Description
nouns list[str] - Lemmas of the nouns

Phrase density

ests.style_stats.calc_phrase_density(), ests.style_stats.expand_phrases()

The number of occurrences of the phrases of a list per 100 words - the compound prepositions (COMPOUND_PREPOSITIONS), the parenthetical expressions (PARENTHETICALS) and the clichés (OFFICIALESE_CLICHES). At every position the longest phrase is taken, and the phrases found do not overlap. expand_phrases spells the phrases out in the forms of the text: a phrase ending in a or de also takes the contraction with the article (a efectos del, conforme al), and a phrase whose first word is an infinitive takes the forms of the text with that lemma (proceder a - procedió a, ser de aplicación - es de aplicación). The lemma is the one of simplemma, and a pronominal lemma counts for its verb (llévese - llevar); the forms simplemma does not lemmatize - the irregular participles (ha dado, ha hecho) and the imperative dese - are given by IRREGULAR_VERB_FORMS. The words after the verb rule out the readings as a noun (el hecho, el puesto). When two phrases spell out the same words, a phrase written so wins over the forms of another (a efectos del stays itself next to a efectos de), and then the first phrase of a list, or of a set in sorted order.

The compound prepositions (42) and the clichés (74) are the ones the Spanish guides to plain language and style manuals of the administrations flag:

Source Author Year
Libro de estilo de la Justicia RAE, CGPJ 2017
Diccionario panhispánico de dudas, 2nd ed. RAE, ASALE online
Guía panhispánica de lenguaje claro y accesible RAE, ASALE 2024
Claridad y derecho a comprender Comisión de Modernización del Lenguaje Jurídico 2011
Cómo escribir con claridad European Commission 2015
Manual del Lenguaje Administrativo Ayuntamiento de Madrid about 2008
Comunicación Clara. Guía práctica Ayuntamiento de Madrid 2017
Guía de Comunicación Clara Comunidad de Madrid 2021
Manual de Lenguaje y Estilo Administrativo Región de Murcia 2008
Manual de Lenguaje Claro Secretaría de la Función Pública, Mexico 2007
Guía de lenguaje claro para servidores públicos Departamento Nacional de Planeación, Colombia 2015
Guía de lenguaje claro Colombia Compra Eficiente 2024
Manual de lenguaje claro Gobierno de la Ciudad de Buenos Aires 2024
Guía para el uso del Lenguaje Claro Legislatura de la Ciudad de Buenos Aires 2024
Manual de Estilo del Lenguaje para uso de la Administración Pública Provincial Provincia de Salta about 2007

Left out are the forms the guides recommend (sobre la base de) or accept (de acuerdo a) and the phrases with frequent neutral uses (en este sentido), but for proceder a and llevar a cabo. A few compound prepositions have neutral uses too (a través de, en caso de, con respecto a), so the density of a neutral text is not zero: compare texts with each other.

Parameter Type Default Description
text list[str] - List of words
phrases list[str]/set[str] - Phrases, words separated by spaces

Parenthetical expressions

ests.style_stats.calc_parentheticals(), ests.style_stats.is_parenthetical()

The parenthetical expressions of PARENTHETICALS (sin embargo, es decir, por ejemplo, finalmente) per 100 words; is_parenthetical checks a word against the one-word expressions of the list. The punctuation is not looked at, so the list holds the expressions set off by commas or standing at the edge of a sentence in most of their occurrences in the corpus of literature, with the series of order (en primer lugar, en segundo lugar). The adverbs of doubt (tal vez, quizá), sobre todo, sin duda and además are not in the list.

Parameter Type Default Description
text list[str] - List of words