Metric functions¶
The coefficients of the formulas are parameters; the defaults below are those of the original English formulas.
Flesch reading ease¶
anyts.readability_stats.calc_flesch_reading_easy()
Computation of the Flesch reading ease (Flesch, 1948).
The higher the value, the easier the text; the scale runs nominally from 0 to 100, though the simplest texts go above 100 and the hardest below 0.
Formula:
Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
n_syllables |
int | - |
Number of syllables |
n_words |
int | - |
Number of words |
n_sents |
int | - |
Number of sentences |
a |
float | 1.015 |
Coefficient a, at the mean sentence length |
b |
float | 84.6 |
Coefficient b, at the mean word length |
c |
float | 206.835 |
Coefficient c, the constant |
Flesch-Kincaid grade¶
anyts.readability_stats.calc_flesch_kincaid_grade()
Computation of the Flesch-Kincaid grade (Kincaid et al., 1975).
The years of schooling needed to read the text; the higher the value, the harder the text.
Formula:
Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
n_syllables |
int | - |
Number of syllables |
n_words |
int | - |
Number of words |
n_sents |
int | - |
Number of sentences |
a |
float | 0.39 |
Coefficient a, at the mean sentence length |
b |
float | 11.8 |
Coefficient b, at the mean word length |
c |
float | 15.59 |
Coefficient c, the constant |
Coleman-Liau index¶
anyts.readability_stats.calc_coleman_liau_index()
Computation of the Coleman-Liau index (Coleman and Liau, 1975).
The years of schooling needed to read the text from the letters and the sentences per 100 words; the higher the value, the harder the text.
Formula:
Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
n_letters |
int | - |
Number of letters of the words |
n_words |
int | - |
Number of words |
n_sents |
int | - |
Number of sentences |
a |
float | 0.0588 |
Coefficient a, at the letters per 100 words |
b |
float | 0.296 |
Coefficient b, at the sentences per 100 words |
c |
float | 15.8 |
Coefficient c, the constant |
Automated readability index (ARI)¶
anyts.readability_stats.calc_automated_readability_index()
Computation of the automated readability index (Smith and Senter, 1967).
The years of schooling needed to read the text from the letters per word and the words per sentence; the higher the value, the harder the text.
Formula:
Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
n_letters |
int | - |
Number of letters of the words |
n_words |
int | - |
Number of words |
n_sents |
int | - |
Number of sentences |
a |
float | 4.71 |
Coefficient a, at the letters per word |
b |
float | 0.5 |
Coefficient b, at the mean sentence length |
c |
float | 21.43 |
Coefficient c, the constant |
SMOG index¶
anyts.readability_stats.calc_smog_index()
Computation of the SMOG index (McLaughlin, 1969).
The years of schooling needed to read the text from the polysyllabic words per sentence; the higher the value, the harder the text. In the original formula the polysyllabic words have three or more syllables and b brings their number to a sample of 30 sentences.
Formula:
Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
n_complex |
int | - |
Number of polysyllabic words |
n_sents |
int | - |
Number of sentences |
a |
float | 1.043 |
Coefficient a, at the square root |
b |
float | 30 |
Coefficient b, the number of sentences of the sample |
c |
float | 3.1291 |
Coefficient c, the constant |
Gunning fog index¶
anyts.readability_stats.calc_gunning_fog_index()
Computation of the Gunning fog index (Gunning, 1952).
The years of schooling needed to read the text from the mean sentence length and the percentage of complex words, in the original formula the words of three or more syllables; the higher the value, the harder the text.
Formula:
Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
n_complex |
int | - |
Number of complex words |
n_words |
int | - |
Number of words |
n_sents |
int | - |
Number of sentences |
a |
float | 0.4 |
Coefficient a |
LIX readability index¶
anyts.readability_stats.calc_lix()
Computation of the LIX readability index (Björnsson, 1968).
The mean sentence length plus the percentage of long words, the words of more than six letters; there are no coefficients to fit. The higher the value, the harder the text:
| Value | Text |
|---|---|
0-30 |
very easy texts, children's books |
30-40 |
easy texts, fiction, newspaper articles |
40-50 |
texts of medium difficulty, magazine articles |
50-60 |
hard texts, popular science, official texts |
60-100 |
very hard texts, laws and bureaucratic language |
Formula:
Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
n_long_words |
int | - |
Number of long words |
n_words |
int | - |
Number of words |
n_sents |
int | - |
Number of sentences |
RIX readability index¶
anyts.readability_stats.calc_rix()
Computation of the RIX readability index (Anderson, 1983).
The long words, of more than six letters, per sentence; there are no coefficients to fit. The higher the value, the harder the text:
| Value | Grade |
|---|---|
< 0.2 |
1 |
0.2-0.5 |
2 |
0.5-0.8 |
3 |
0.8-1.3 |
4 |
1.3-1.8 |
5 |
1.8-2.4 |
6 |
2.4-3.0 |
7 |
3.0-3.7 |
8 |
3.7-4.5 |
9 |
4.5-5.3 |
10 |
5.3-6.2 |
11 |
6.2-7.2 |
12 |
> 7.2 |
college |
Formula:
Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
n_long_words |
int | - |
Number of long words |
n_sents |
int | - |
Number of sentences |
Legibilidad µ¶
anyts.readability_stats.calc_mu_index()
Computation of Legibilidad µ (Muñoz Baquedano and Muñoz Urra, 2006).
The mean number of letters per word divided by its sample variance, times 100; there are no coefficients to fit, and the higher the value, the easier the text. Words without letters (numbers) are left out; with fewer than two words or without variability the index is undefined (nan).
Formula:
where \(x_i\) is the number of letters of a word and \(n\) the number of words.
Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
c_letters |
dict[int, int] | - |
Distribution of words by number of letters |
Reading ease as a grade¶
anyts.readability_stats.flesch_reading_easy_to_grade()
Conversion of the Flesch reading ease into years of schooling for the consensus grade: the grade of the first band whose lower bound the value reaches, or below under the last bound; values above 100 belong to the first band.
The default bands are those of text_standard of textstat, which follow the table of Flesch (1948) down to 60 and split his lower bands, with 8.5 for its grades 8 and 9: 90-100 - 5, 80-90 - 6, 70-80 - 7, 60-70 - 8.5, 50-60 - 10, 40-50 - 11, 30-40 - 12, below 30 - 13 (anyts.constants.READING_EASE_GRADES).
Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
flesch_reading_easy |
float | - |
Value of the reading ease |
grades |
list[tuple[float, float]] | READING_EASE_GRADES |
Lower bounds and their grades in descending order of the bounds |
below |
float | 13 |
Grade below the last bound |
Consensus grade¶
anyts.readability_stats.calc_consensus_grade()
Computation of the consensus grade: the median of the values of the grade formulas rounded half up, together with the reading ease converted into years of schooling by to_grade and added without rounding - so a band of 8.5 votes for 8.5. No values at all and a grade that is not a finite number raise ParameterError, a to_grade that is not callable SourceTypeError.
Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
grades |
list[float] | - |
Values of the grade formulas |
flesch_reading_easy |
float | None |
Value of the reading ease |
to_grade |
Callable | flesch_reading_easy_to_grade |
Conversion of the reading ease into years of schooling |
School stage and age¶
anyts.readability_stats.grade_to_age()
The school stage and reader age by the value of a grade formula: the value is rounded half up and falls into the first stage whose last year it does not exceed, values below 1 into the first stage; above the last stage lies above. A grade that is not a finite number raises ParameterError.
The default stages are those of the United States (anyts.constants.GRADE_AGE_LEVELS and POSTGRADUATE_LEVEL):
| Grade | Stage | Age |
|---|---|---|
| 1-5 | elementary school, grades 1-5 | 6-11 years |
| 6-8 | middle school, grades 6-8 | 11-14 years |
| 9-12 | high school, grades 9-12 | 14-18 years |
| 13-16 | college | 18-22 years |
| above 16 | graduate school | over 22 years |
Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
grade |
float | - |
Value of a grade formula |
levels |
list[tuple[int, int, str, str]] | GRADE_AGE_LEVELS |
Stages as the first and the last year, the stage and the age, in ascending order |
above |
tuple[str, str] | POSTGRADUATE_LEVEL |
Stage and age above the last stage |
Reading time¶
anyts.readability_stats.calc_reading_time()
Computation of the reading time of a text in minutes: the number of words divided by the reading speed; a speed that is not a positive number raises ParameterError.
The default speed is the silent reading speed of adults in English, 238 words per minute (Brysbaert, 2019).
Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
n_words |
int | - |
Number of words |
wpm |
float | 238 |
Reading speed, words per minute |