Word dispersion¶
anyts.corpus.dispersion(), anyts.corpus.Dispersion
Description¶
The dispersion of a word is how evenly it is spread over the parts of a text or of a corpus. Frequency does not tell a word that occurs once in every chapter from a word gathered in one of them; the measures of dispersion (Gries 2008, 2020) complement frequency.
The text is split into parts: parts is the number of parts of about equal size or the sizes of the parts in order (sentences, paragraphs, chapters, documents of a corpus), which add up to the number of words. For every word its frequencies by part and six measures are computed; Gries recommends DP as the main one.
Words are compared as they are: case and lemmatization belong to the word extractor.
Measures¶
For \(n\) parts of shares \(s_i\) of the text, frequencies of the word by part \(v_i\) and a total frequency \(f = \sum v_i\); \(p_i = v_i / n_i\) is the relative frequency in a part of size \(n_i\):
| Measure | Field | Formula | Values |
|---|---|---|---|
| Deviation of proportions DP | dp |
\(\frac{1}{2} \sum \left\lvert \frac{v_i}{f} - s_i \right\rvert\) | 0 - in proportion to the sizes of the parts, \(1 - s_i\) - all in part \(i\), so at most \(1 - \min s_i\); Gries (2008) |
| Normalized DP | dp_norm |
\(\frac{DP}{1 - \min s_i}\) | the maximum is one whatever the split; Lijffijt and Gries (2012) |
| Juilland's D | juilland_d |
\(1 - \frac{V}{\sqrt{n - 1}}\), \(V = \frac{\sigma(p)}{\mu(p)}\) | 1 - even, 0 - in one part; Juilland and Chang-RodrÃguez (1964), with the population standard deviation \(\sigma\) as in Gries (2020); Gries (2008) takes the sample one |
| Carroll's D2 | carroll_d2 |
\(\frac{H(p)}{\log_2 n}\) | entropy of the distribution \(p_i\); 1 - even, 0 - in one part; Carroll (1970) |
| Rosengren's S | rosengren_s |
\(\frac{(\sum \sqrt{s_i v_i})^2}{f}\) | 1 - in proportion, tends to \(1/n\) when gathered in one of equal parts; Rosengren (1971) |
| Kullback-Leibler divergence | kl_divergence |
\(\sum \frac{v_i}{f} \log_2 \frac{v_i / f}{s_i}\) | in bits; 0 - in proportion, grows when gathered in small parts; Gries (2020) |
The measures are available as the functions calc_dp, calc_dp_norm, calc_juilland_d, calc_carroll_d2, calc_rosengren_s and calc_kl_divergence with the arguments (frequencies, sizes) - the frequencies of the word by part and the sizes of the parts - of the module anyts.corpus.dispersion; their names are in anyts.constants.DISPERSION_STATS_DESC. For a word of zero frequency every measure is nan.
Parameters¶
| Parameter | Type | Default | Description |
|---|---|---|---|
words |
list[str] | - |
Words of the text in order |
parts |
int/list[int] | 10 |
Number of parts (from 2 to the number of words) or sizes of the parts |
word |
str | None |
Word whose dispersion is needed; None - every word |
min_freq |
int | 1 |
Minimum frequency of a word |
Result¶
A list of Dispersion named tuples by descending frequency: word, freq and the six measures of the table; pd.DataFrame(result) gives a table.
Example¶
Example
from anyts import SentsExtractor, WordsExtractor
from anyts.corpus import dispersion
text = (
"The cat was at the window and watched the birds. The birds flew away and the cat "
"fell asleep at the window. Tomorrow the cat will be at the window and watch the birds."
)
we = WordsExtractor(lowercase=True)
words = we.extract(text)
sizes = [len(we.extract(sent)) for sent in SentsExtractor().extract(text)]
sizes
# [10, 12, 12]
dispersion(words, parts=sizes, word="cat")
# [Dispersion(word='cat', freq=3, dp=0.039215686274509776, dp_norm=0.05555555555555552,
# juilland_d=0.9374999999999999, carroll_d2=0.9965119062478288,
# rosengren_s=0.998213770592287, kl_divergence=0.005215975085958165)]
# Three equal parts: birds are gathered at the edges of the text
[(d.word, round(d.dp, 2)) for d in dispersion(words, parts=3, min_freq=3)]
# [('the', 0.1), ('cat', 0.02), ('at', 0.02), ('window', 0.02), ('and', 0.02), ('birds', 0.32)]