Skip to content

Collocations

anyts.corpus.collocations(), anyts.corpus.Collocation

Description

Collocation extraction - pairs of words that occur together more often than independence would give: fixed expressions, terminology, the combinatorics of a word. The measures of association are those of Sketch Engine and of nltk.metrics.association.

Pairs of words are ordered, as in NLTK: the right word occurs no more than window words after the left one, and every pair of positions is counted once; window=1 gives bigrams. Inside the measure the frequency of a pair is divided by the size of the window (Church and Hanks 1990, as in NLTK), so that the expected frequency does not depend on the window and Dice and the minimum sensitivity do not exceed one; the field freq_pair keeps the undivided frequency. With window > 1 the scale of the Dice measures therefore shifts: a pair always side by side gets a logDice of \(14 - \log_2 window\) (13 at a window of 2) and not 14 as in Sketch Engine, which uses the raw co-occurrence; 14 is reached only by a pair that occurs at every distance within the window. The parameter node keeps the pairs with the given word on the left or on the right - the combinatorics of one word.

Words are compared as they are: case, lemmatization and stop words belong to the word extractor; lemmas suit fixed expressions, word forms suit grammatical constructions.

Measures

For a pair of words of frequencies \(f_a\) and \(f_b\), a pair frequency \(f_{ab}\) and a number of words \(N\):

Measure Key Formula Description
Mutual information mi \(\log_2 \frac{f_{ab} N}{f_a f_b}\) Church and Hanks (1990); overrates rare pairs
MI³ mi3 \(\log_2 \frac{f_{ab}^3 N}{f_a f_b}\) Oakes (1998); favours frequent pairs
t-score t_score \(\frac{f_{ab} - f_a f_b / N}{\sqrt{f_{ab}}}\) Church et al. (1991); favours frequent pairs
Dice coefficient dice \(\frac{2 f_{ab}}{f_a + f_b}\) does not depend on the size of the text
logDice logdice \(14 + \log_2 \frac{2 f_{ab}}{f_a + f_b}\) Rychlý (2008); does not depend on the size of the text, at most 14, below zero - a weak link; the default measure, as in Sketch Engine
Log-likelihood log_likelihood \(G^2 = 2 \sum O \ln \frac{O}{E}\) Dunning (1993); over the 2×2 contingency table, nan if one of the words fills the whole text
NPMI npmi \(\frac{MI}{-\log_2 (f_{ab} / N)}\) Bouma (2009); from −1 to 1, one means the words occur only together
Minimum sensitivity min_sensitivity \(\min(\frac{f_{ab}}{f_a}, \frac{f_{ab}}{f_b})\) Pedersen (1998); from 0 to 1

The measures are available as the functions calc_mi, calc_mi3, calc_t_score, calc_dice, calc_logdice, calc_log_likelihood, calc_npmi and calc_min_sensitivity with the arguments (freq_a, freq_b, freq_ab, n) of the module anyts.corpus.collocations; their names and descriptions are in anyts.constants.COLLOCATION_MEASURES.

Parameters

Parameter Type Default Description
words list[str] - Words of the text in order
window int 5 Greatest distance between the words of a pair
measure str logdice Measure of anyts.constants.COLLOCATION_MEASURES
min_freq int 2 Minimum frequency of a pair
node str None Word whose combinatorics is needed; None - every pair
top_n int None Number of collocations; None - all of them

Result

A list of Collocation named tuples by descending measure and frequency of the pair (ties broken alphabetically); pd.DataFrame(found) gives a table.

Field Type Description
left str Left word
right str Right word, occurring within the window after the left one
freq_left int Frequency of the left word
freq_right int Frequency of the right word
freq_pair int Frequency of the co-occurrence
score float Value of the chosen measure

Example

Example

from anyts import WordsExtractor
from anyts.corpus import collocations

words = WordsExtractor(lowercase=True).extract(
    "The cat was at the window and watched the birds. The birds flew away and the cat "
    "fell asleep at the window. Tomorrow the cat will be at the window and watch the birds."
)

collocations(words, window=2, top_n=1)
# [Collocation(left='at', right='window', freq_left=3, freq_right=3, freq_pair=3, score=13.0)]

[
    (c.left, c.right, round(c.score, 2))
    for c in collocations(words, window=1, node="cat", min_freq=1, measure="mi")[:2]
]
# [('cat', 'fell', 3.5), ('cat', 'was', 3.5)]