Lexical diversity metrics¶
anyts.diversity_stats.DiversityStats
Description¶
A class for computing the main lexical diversity metrics of a text from its words.
Note
The metrics are computed by accessing the corresponding attribute or by calling the get_stats method of the DiversityStats object.
Language hooks¶
The class takes the words as given: a language library extracts them from its sources (a string, a Doc object of spaCy) with its own extractor and lower-cases them, since every metric counts lexemes. The printed table of print_stats takes its texts from two class attributes that a library overrides in a subclass: stats_desc, the descriptions of the metrics by their keys, and stats_headers, the headers of the two columns.
Parameters¶
| Parameter | Type | Default | Description |
|---|---|---|---|
words |
Iterable[str] | - |
Words of the text: a sequence or an iterator of strings |
window_len |
int | 50 |
Window size for MATTR and segment size for MSTTR |
mtld_threshold |
float | 0.72 |
TTR threshold for MTLD, MA-MTLD and MTLD-W |
mtld_min_len |
int | 10 |
Minimum factor length for MTLD, MA-MTLD and MTLD-W |
hdd_sample_size |
int | 42 |
Sample size for HD-D |
log_base |
float | 10 |
Logarithm base for the Summer, Maas and Dugast metrics |
The parameters are checked by check_params(window_len, mtld_threshold, mtld_min_len, hdd_sample_size, log_base), which raises ParameterError for a window, a sample size or a minimum factor length out of range, a threshold outside (0, 1) or a logarithm base not above 1. Called before the words are extracted, it reports a wrong parameter before an empty text.
Conventions¶
The values of some metrics depend on conventions that differ between libraries. All of them except the comparison with the MTLD threshold are class parameters:
| Parameter | Default | Other libraries |
|---|---|---|
| Logarithm base for Summer, Maas, Dugast's U and Dugast's k | 10 | LexicalRichness, textcomplexity and zipfR - natural |
| MATTR window and MSTTR segment | 50 | quanteda and koRpus - 100 |
| TTR threshold for MTLD | 0.72 | 0.66-0.75 in the literature |
| Comparison with the MTLD threshold | a factor closes at TTR ≤ 0.72 (McCarthy & Jarvis, 2010) | lexical-diversity and TAALED - strict <; the values differ when TTR hits the threshold exactly |
| Minimum MTLD factor length | 10 | koRpus applies it only to MTLD-MA and drops the shorter factors instead of extending them, LexicalRichness and textcomplexity do not apply it |
| HD-D sample size | 42 | 35-50 in the literature |
By Zenker and Kyle (2021) MATTR, MTLD and HD-D are stable on texts of 50-200 words and longer, MTLD-W, MA-MTLD and Maas are unstable on short texts, and the TTR family never stabilizes. To compare texts of different lengths use the windowed computation with confidence intervals.
Attributes¶
| Attribute | Type | Description |
|---|---|---|
words |
tuple[str] | Tuple of the words |
window_len, mtld_threshold, mtld_min_len, hdd_sample_size, log_base |
int/float | The parameters of the metrics; changing them on the object changes the metrics |
frequency_spectrum |
dict[int, int] | Frequency spectrum - the number of lexemes with a given frequency |
ttr |
float | Type-Token Ratio (TTR) |
rttr |
float | Root Type-Token Ratio (RTTR) |
cttr |
float | Corrected Type-Token Ratio (CTTR) |
httr |
float | Herdan Type-Token Ratio (HTTR) |
sttr |
float | Summer Type-Token Ratio (STTR) |
mttr |
float | Maas Type-Token Ratio (MTTR) |
dttr |
float | Dugast Type-Token Ratio (DTTR) |
mattr |
float | Moving Average Type-Token Ratio (MATTR) |
msttr |
float | Mean Segmental Type-Token Ratio (MSTTR) |
mtld |
float | Measure of Textual Lexical Diversity (MTLD) |
mamtld |
float | Moving Average Measure of Textual Lexical Diversity (MA-MTLD) |
mtldw |
float | MTLD with a moving window and text wrap (MTLD-W) |
hdd |
float | Hypergeometric Distribution D (HD-D) |
simpson_index |
float | Simpson's index (D) |
inverse_simpson_index |
float | Inverse Simpson's index (1/D) |
gini_simpson_index |
float | Gini-Simpson index (1-D) |
hapax_index |
float | Hapax index, a.k.a. Honoré's R |
honore_r |
float | Alias for the hapax index |
yule_k |
float | Yule's characteristic (Yule's K) |
yule_i |
float | Inverse Yule's characteristic (Yule's I) |
herdan_vm |
float | Herdan's Vm |
sichel_s |
float | Sichel's S |
michea_m |
float | Michéa's M |
brunet_w |
float | Brunet's W |
dugast_k |
float | Dugast's k |
baayen_p |
float | Baayen's P |
hapax_ratio |
float | Share of hapaxes among lexemes |
alpha2 |
float | The α₂ exponent |
entropy |
float | Shannon entropy in bits |
evenness |
float | Evenness - the ratio of entropy to its maximum |
perplexity |
float | Perplexity |
zipf_alpha |
float | Zipf's law slope |
heaps_beta |
float | Heaps' law exponent |
Note
Every metric can also be computed by its function; the metrics and their functions are described in the corresponding section.
Methods¶
windowed¶
Windowed computation of a metric by its name, as in calc_windowed: its value over consecutive text windows of equal length, the mean, the sample standard deviation and the confidence interval of the mean by Student's distribution. The edge cases (short texts, nan and infinite values) are described there.
Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
stat |
str | - |
Metric name from get_stats |
window_len |
int | 100 |
Window size |
step |
int | None |
Window step, by default equal to the window size (windows do not overlap) |
confidence |
float | 0.95 |
Confidence level |
Returns a WindowStats named tuple with the fields mean, std, lower, upper and n_windows. An unknown metric name raises UnknownStatError.
Example
from anyts import DiversityStats
ds = DiversityStats("a b c d e b c f g h g i g j k".split())
ds.windowed("ttr", window_len=5)
# WindowStats(mean=0.9333333333333332, std=0.11547005383792512, lower=0.6464898180167025, upper=1.220176848649964, n_windows=3)
ds.windowed("ttr", window_len=5, step=2).n_windows
# 6
get_stats¶
Returns a dictionary with the computed lexical diversity metrics.
Example
Code:
from anyts import DiversityStats
ds = DiversityStats("a b c d e b c f g h g i g j k".split())
ds.get_stats()
Result:
{'ttr': 0.7333333333333333,
'rttr': 2.840187787218772,
'cttr': 2.008316044185609,
'httr': 0.8854692840710255,
'sttr': 0.2500605793160848,
'mttr': 0.09738250756232528,
'dttr': 10.268784661968118,
'mattr': 0.7333333333333333,
'msttr': 0.7333333333333333,
'mtld': 15.0,
'mamtld': 12.0,
'mtldw': 13.25,
'hdd': nan,
'simpson_index': 0.047619047619047616,
'inverse_simpson_index': 21.0,
'gini_simpson_index': 0.9523809523809523,
'hapax_index': 992.9517404041437,
'yule_k': 444.44444444444446,
'yule_i': 8.642857142857142,
'herdan_vm': 0.1421338109037403,
'sichel_s': 0.18181818181818182,
'michea_m': 5.5,
'brunet_w': 6.00637847898991,
'dugast_k': 14.783895126869226,
'baayen_p': 0.5333333333333333,
'hapax_ratio': 0.7272727272727273,
'alpha2': 0.5,
'entropy': 3.3232314287976203,
'evenness': 0.9606293157795304,
'perplexity': 10.009038104159247,
'zipf_alpha': 0.4884512334695912,
'heaps_beta': 0.8366147342060046}
print_stats¶
Prints a table with the computed lexical diversity metrics, their descriptions from stats_desc and the headers from stats_headers.
The example continues the previous one:
Example
Code:
ds.print_stats()
Result:
Metric | Value
-------------------------------------------------------------------------------------
Type-Token Ratio (TTR) | 0.73
Root Type-Token Ratio (RTTR) | 2.84
Corrected Type-Token Ratio (CTTR) | 2.01
Herdan Type-Token Ratio (HTTR) | 0.89
Summer Type-Token Ratio (STTR) | 0.25
Maas Type-Token Ratio (MTTR) | 0.10
Dugast Type-Token Ratio (DTTR) | 10.27
Moving Average Type-Token Ratio (MATTR) | 0.73
Mean Segmental Type-Token Ratio (MSTTR) | 0.73
Measure of Textual Lexical Diversity (MTLD) | 15.00
Moving Average Measure of Textual Lexical Diversity (MA-MTLD) | 12.00
Moving Average Measure of Textual Lexical Diversity with Wrap (MTLD-W) | 13.25
Hypergeometric Distribution D (HD-D) | nan
Simpson's index (D) | 0.05
Inverse Simpson's index (1/D) | 21.00
Gini-Simpson index (1-D) | 0.95
Hapax index (Honoré's R) | 992.95
Yule's characteristic K | 444.44
Yule's inverse characteristic I | 8.64
Herdan's Vm | 0.14
Sichel's S | 0.18
Michéa's M | 5.50
Brunet's W | 6.01
Dugast's k | 14.78
Baayen's P | 0.53
Hapax ratio | 0.73
Exponent α₂ | 0.50
Shannon entropy (bits) | 3.32
Evenness | 0.96
Perplexity | 10.01
Zipf's law slope (α) | 0.49
Heaps' law exponent (β) | 0.84