Cohesion helpers¶
anyts.cohesion
Functions that measure how the sentences of a text are tied together by shared elements, in the manner of Coh-Metrix. A sentence is given as the collection of its elements - lemmas of nouns, arguments or content words, values of a feature of its verbs - so the functions see no language: a library chooses the elements.
A list of sentences that is a string, a Doc, an iterator, a set or a table raises SourceTypeError, and so does a sentence that is not a list of strings (check_words): a string would be read character by character, an iterator exhausted by the first pass, and spaCy tokens never match, since a token equals only itself. The overlaps (calc_overlap, calc_proportional_overlap, calc_overlaps) compare sets of words and take a sentence as a set too; count_given and calc_repetition count the words in order and refuse a set.
calc_overlap¶
The binary overlap of Coh-Metrix (CRFNO1, CRFAO1, CRFSO1 over the adjacent pairs, CRFNOa, CRFAOa, CRFSOa over all of them): the share of the pairs of sentences that share at least one element; nan for a text shorter than two sentences. All the pairs are counted as in calc_overlaps, without going through them.
| Parameter | Type | Default | Description |
|---|---|---|---|
sets |
list[set[str]] | - |
Elements of every sentence |
adjacent |
bool | True |
Count the adjacent pairs only, otherwise all the pairs |
calc_proportional_overlap¶
The proportional overlap of Coh-Metrix (CRFCWO1, CRFCWOa): the Dice coefficient of a pair of sentences averaged over the pairs; nan for a text shorter than two sentences. Over all the pairs the sum is computed as in calc_overlaps, without going through them.
| Parameter | Type | Default | Description |
|---|---|---|---|
sets |
list[set[str]] | - |
Elements of every sentence |
adjacent |
bool | True |
Count the adjacent pairs only, otherwise all the pairs |
dice¶
The Dice coefficient of two sets:
and 0 for two empty sets.
| Parameter | Type | Default | Description |
|---|---|---|---|
first |
frozenset[str] | - |
First set |
second |
frozenset[str] | - |
Second set |
calc_overlaps¶
The values of calc_overlap and calc_proportional_overlap over the adjacent and over all the pairs at once, as an Overlap named tuple with the fields adjacent, all, prop_adjacent and prop_all. The values over all the pairs are computed without going through every pair: the pairs sharing an element are counted over bit masks of the sentences, and the sum of the Dice coefficients over the histograms of the lengths of the sentences holding every element. That sum is added by math.fsum, so it does not depend on the order of the elements. Without proportional the fields prop_adjacent and prop_all are nan.
| Parameter | Type | Default | Description |
|---|---|---|---|
sets |
list[set[str]] | - |
Elements of every sentence |
proportional |
bool | True |
Compute the proportional overlap as well |
count_given¶
The number of given elements: an element is given when the same lemma was used earlier in any sentence, the current one included; the first occurrence is new.
| Parameter | Type | Default | Description |
|---|---|---|---|
sents |
list[list[str]] | - |
Lemmas of every sentence in the order of the text |
dominant¶
The dominant value: the most frequent one, the first of the equally frequent; None for an empty list.
| Parameter | Type | Default | Description |
|---|---|---|---|
values |
list[str] | - |
Values |
calc_repetition¶
The repetition of the tense and of the mood of Coh-Metrix (SMTEMP): the share of the adjacent pairs of sentences whose verbs have the same dominant value of the feature; a pair where one of the sentences has no verb with the feature is skipped, and without a single such pair the value is nan.
| Parameter | Type | Default | Description |
|---|---|---|---|
sents |
list[list[str]] | - |
Values of the feature of the verbs of every sentence |
Example
from anyts.cohesion import calc_overlaps, calc_repetition, count_given
calc_overlaps([{"cat", "house"}, {"cat", "dog"}, {"bird"}])
# Overlap(adjacent=0.5, all=0.3333333333333333, prop_adjacent=0.25, prop_all=0.16666666666666666)
count_given([["cat", "house"], ["cat", "dog"], ["house"]])
# 2
calc_repetition([["Pres"], ["Pres", "Past", "Pres"], ["Past"]])
# 0.5