Keywords¶
anyts.corpus.keyness(), anyts.corpus.Keyword, anyts.corpus.FrequencyReference
Description¶
Keyword extraction (keyness) for a target corpus against a reference one: the words that occur significantly more often in the target corpus than in the reference.
For every word two values are computed that Gabrielatos and Marchi and Hardie recommend reading together: the log-likelihood \(G^2\) with its p-value (the significance of the difference - whether there is one) and Log Ratio (the size of the effect - how large it is). The chosen measure score is computed as well and used for the sorting. The measures of significance (\(G^2\), chi-square, BIC, ELL) are signed: negative when the word is more frequent in the reference; the measures of effect (%DIFF, Log Ratio, odds ratio) are directional by construction.
The reference may be a list of words, a mapping of frequencies (its size is the sum of the counts) or a FrequencyReference. Words are compared as they are: case, lemmatization and stop words belong to the word extractor, and both corpora have to be extracted the same way.
Measures¶
For a word of frequency \(a\) in a target corpus of size \(c\) and of frequency \(b\) in a reference corpus of size \(d\), \(N = c + d\):
| Measure | Key | Formula | Description |
|---|---|---|---|
| Log-likelihood | log_likelihood |
\(G^2 = 2\,(a \ln \frac{a}{E_1} + b \ln \frac{b}{E_2})\), \(E_1 = \frac{c\,(a+b)}{N}\), \(E_2 = \frac{d\,(a+b)}{N}\) | Rayson and Garside (2000); critical values anyts.constants.G2_CRITICAL_VALUES: 3.84 for p < 0.05, 6.63 for p < 0.01, 10.83 for p < 0.001, 15.13 for p < 0.0001 |
| Chi-square | chi2 |
\(\chi^2 = \frac{N\,\max(\lvert a(d-b) - b(c-a) \rvert - N/2,\ 0)^2}{(a+b)(N-a-b)\,c\,d}\) | with Yates's correction over the 2×2 contingency table; when the correction exceeds the difference, the statistic is zero |
| %DIFF | diff |
\(\frac{NF_a - NF_b}{NF_b} \cdot 100\) | Gabrielatos and Marchi (2011); \(NF\) - frequency per million words |
| Log Ratio | log_ratio |
\(\log_2 \frac{NF_a}{NF_b}\) | Hardie (2014); one means the word is twice as frequent in the target corpus |
| BIC | bic |
\(\operatorname{sign}(G^2) \cdot (\lvert G^2 \rvert - \ln N)\) | Wilson (2013); in absolute value above 2 - positive evidence of a difference, above 6 - strong, above 10 - very strong; a negative value with \(\lvert G^2 \rvert < \ln N\) means no evidence, not the opposite direction |
| ELL | ell |
\(\frac{G^2}{N \ln \min(E_1, E_2)}\) | Johnston, Berry and Mielke (2006); size of the effect of \(G^2\), the share of the greatest possible departure from the expected frequencies, from 0 to 1; nan when the least expected frequency is at most one - then the logarithm is zero or negative; just above one the measure grows without bound |
| Odds ratio | odds_ratio |
\(\frac{a / (c - a)}{b / (d - b)}\) | one means equal odds; inf if the word fills the whole target corpus, 0 - the whole reference |
A zero frequency in one of the corpora is replaced with 0.5 for %DIFF, Log Ratio and the odds ratio (Hardie 2014). The p-value of \(G^2\) comes from the chi-square distribution with one degree of freedom (calc_p_value). The measures are available as the functions calc_log_likelihood, calc_chi2, calc_diff, calc_log_ratio, calc_bic, calc_ell and calc_odds_ratio with the arguments (a, b, c, d) of the module anyts.corpus.keyness; their names and descriptions are in anyts.constants.KEYNESS_MEASURES.
Parameters¶
| Parameter | Type | Default | Description |
|---|---|---|---|
target |
list[str]/dict[str, int] | - |
Words of the target corpus or their frequencies |
reference |
list[str]/dict[str, float]/FrequencyReference | - |
Words of the reference corpus, their frequencies or a reference by frequencies |
measure |
str | log_likelihood |
Measure of anyts.constants.KEYNESS_MEASURES for score and the sorting |
min_freq |
int | 1 |
Minimum frequency of a keyword in its own corpus |
positive |
bool | True |
Positive keywords (more frequent in the target corpus) or negative ones (more frequent in the reference) |
top_n |
int | None |
Number of keywords; None - all of them |
Reference by frequencies¶
FrequencyReference(counts, size, missing=0.0, key=None, keep=None) describes a reference corpus by its frequencies alone, such as a frequency dictionary built from a corpus too large to pass as words:
| Field | Type | Default | Description |
|---|---|---|---|
counts |
dict[str, float] | - |
Frequencies of the keys in the reference corpus |
size |
float | - |
Size of the reference corpus in words |
missing |
float | 0.0 |
Frequency of a key missing from counts |
key |
callable | None |
Key of a word of the target corpus in counts, such as its lemma; None - the word itself |
keep |
callable | None |
Whether a word of the target corpus is counted; None - every word |
The words of the target that keep passes are counted under their key, and the words it leaves out do not count in the size of the target either; negative keywords come only from the keys of counts that keep passes. The frequency of a key missing from counts is missing: a dictionary gives its least frequency, an upper bound of the true one, so such a word may be a positive keyword but never a negative one.
Result¶
A list of Keyword named tuples by descending keyness (ties broken by descending frequency and alphabetically, words of an undefined measure last); pd.DataFrame(keywords) gives a table.
| Field | Type | Description |
|---|---|---|
word |
str | Word |
freq_target |
int | Frequency in the target corpus |
freq_reference |
float | Frequency in the reference corpus |
ipm_target |
float | Frequency in the target corpus per million words |
ipm_reference |
float | Frequency in the reference corpus per million words |
g2 |
float | Signed \(G^2\) |
p_value |
float | p-value of \(G^2\) |
log_ratio |
float | Log Ratio |
score |
float | Value of the chosen measure |
Example¶
Example
from anyts import WordsExtractor
from anyts.corpus import FrequencyReference, keyness
we = WordsExtractor(lowercase=True)
target = we.extract(
"The cat was at the window and watched the birds. The birds flew away and the cat "
"fell asleep at the window. Tomorrow the cat will be at the window and watch the birds."
)
reference = we.extract(
"The dog was on the floor and slept. Then the dog ate and slept again. "
"Tomorrow the dog will go for a walk."
)
[(k.word, k.freq_target, round(k.g2, 2)) for k in keyness(target, reference, top_n=3)]
# [('at', 3, 3.1), ('birds', 3, 3.1), ('cat', 3, 3.1)]
[(k.word, round(k.g2, 2)) for k in keyness(target, reference, positive=False, top_n=2)]
# [('dog', -5.45), ('slept', -3.63)]
# Frequencies of the lemmas in a dictionary of a million words
counts = {
"the": 60_000.0,
"and": 28_000.0,
"at": 5_000.0,
"be": 6_000.0,
"was": 9_000.0,
"will": 3_000.0,
"cat": 50.0,
"window": 60.0,
"bird": 40.0,
"watch": 90.0,
"away": 400.0,
"tomorrow": 100.0,
}
lemmas = {"birds": "bird", "watched": "watch"}
dictionary = FrequencyReference(
counts, size=1_000_000, missing=0.1, key=lambda word: lemmas.get(word, word)
)
[(k.word, round(k.g2, 1)) for k in keyness(target, dictionary, top_n=3)]
# [('bird', 40.0), ('cat', 38.7), ('window', 37.6)]