Skip to content

Stylometry

anyts.corpus.delta(), anyts.corpus.delta_profiles(), anyts.corpus.frequency_table(), anyts.corpus.z_scores(), anyts.corpus.zeta(), anyts.corpus.kilgarriff_chi2(), anyts.corpus.mendenhall_curve(), anyts.corpus.mendenhall_distance()

Description

Measures of stylometry and authorship attribution: distances between texts by the frequencies of the most frequent words (Burrows's Delta and its variants, as in stylo), markers of preferred and avoided words (Zeta), Kilgarriff's chi-square distance between corpora and the Mendenhall curve. The functions work on lists of units of a text: lower-case word forms (the usual choice for Delta), lemmas or character N-grams - case and lemmatization belong to the extractors.

Burrows's Delta

A corpus is a dictionary "name of a text → units". frequency_table(corpus, n_mfw=100, culling=0.0) builds the table of relative frequencies: rows are the texts, columns the n_mfw most frequent units by descending mean relative frequency (alphabetically when equal); culling keeps the units that occur in at least the given share of the texts, as in stylo.

z_scores(table) standardizes the columns with the sample standard deviation, like scale() of R; a column with the same frequency in every text gives zeros.

delta computes a symmetric matrix of distances from the z-scores (a DataFrame with the names of the texts); at least three texts are needed - with two, the z-scores degenerate to ±1/√2 and the distances do not depend on the frequencies.

Variants (anyts.constants.DELTA_VARIANTS), with the formulas of the sources of stylo; \(n\) is the number of units, \(z_A\) and \(z_B\) the vectors of z-scores of the texts:

Variant Key Formula Source
Burrows's Delta burrows \(\frac{1}{n} \sum_i \lvert z_{A,i} - z_{B,i} \rvert\) Burrows (2002), dist.delta
Quadratic Delta quadratic \(\frac{1}{n} \sqrt{\sum_i (z_{A,i} - z_{B,i})^2}\) Argamon (2008), dist.argamon
Eder's Delta eder \(\sum_i \frac{n - i + 2}{n} \lvert z_{A,i} - z_{B,i} \rvert\), \(i\) - rank of the unit by frequency Eder, dist.eder
Cosine Delta cosine \(1 - \frac{z_A \cdot z_B}{\lVert z_A \rVert \lVert z_B \rVert}\) Smith and Aldridge (2011), Evert et al. (2015), dist.wurzburg

Cosine Delta clusters the texts by author best in the experiments of Evert et al. The usual number of units is 100 to 500 most frequent words, 100-200 for character N-grams.

Parameters of delta:

Parameter Type Default Description
corpus dict[str, list[str]] - Units of the texts by the names of the texts
n_mfw int 100 Number of the most frequent units; None - all of them
variant str burrows Variant of Delta of anyts.constants.DELTA_VARIANTS
culling float 0.0 Smallest share of the texts a unit occurs in

For authorship attribution there is delta_profiles(reference, samples, n_mfw, variant, culling, statistics): the most frequent units, the culling and the statistics of the z-scores come from the reference texts reference (the profiles of the authors) or from a separate set statistics - for instance, the training windows when the profiles are too few to estimate the spread of the frequencies; the texts under test samples are described by the same units and scaled by the same statistics. The result is the distances from the texts under test to the reference ones, and the nearest reference in a row is the presumed author. Unlike in delta, the texts under test affect neither the units nor the scaling, so the result for a text does not depend on the texts passed along with it.

Example

from anyts import WordsExtractor
from anyts.corpus import delta, delta_profiles, frequency_table

texts = {
    "A": (
        "The cat was at the window and watched the birds. "
        "The birds flew away and the cat slept at the window."
    ),
    "B": "The dog was on the floor and slept. Then the dog ate and slept again on the floor.",
    "C": "Tomorrow the cat will come back to the window and watch the birds, but the dog will sleep.",
}
we = WordsExtractor(lowercase=True)
corpus = {name: we.extract(text) for name, text in texts.items()}

frequency_table(corpus, n_mfw=5).round(3)
#      the    and    dog  slept  birds
# A  0.286  0.095  0.000  0.048  0.095
# B  0.222  0.111  0.111  0.111  0.000
# C  0.222  0.056  0.056  0.000  0.056

delta(corpus, n_mfw=5).round(3)
#        A      B      C
# A  0.000  1.483  1.161
# B  1.483  0.000  1.219
# C  1.161  1.219  0.000

delta(corpus, n_mfw=5, variant="cosine").round(3)
#        A      B      C
# A  0.000  1.676  1.273
# B  1.676  0.000  1.525
# C  1.273  1.525  0.000

sample = {"?": we.extract("The cat woke up at the window and watched the birds again.")}
delta_profiles(corpus, sample, n_mfw=5).round(3)
#        A      B      C
# ?  0.499  1.493  0.662

Zeta

Markers of preferred and avoided words after Burrows (2007) and Craig and Kinney (2009). Every text of both corpora is split into segments of about segment_size words (the number of segments is the ratio of the length to the size rounded half up, at least one), and for a word the share of the segments of each corpus where it occurs (\(DP\)) is computed. Zeta is the difference of the shares \(DP_{target} - DP_{comparison}\) from −1 to 1 (zeta.craig in the notation of stylo; the classic Zeta of Craig \(DP_{target} + (1 - DP_{comparison})\) is greater by one), the logarithmic Zeta is \(\log_2 \frac{DP_{target}}{DP_{comparison}}\) (Schöch et al. 2018), a zero share replaced by half a segment. The list starts with the words the target corpus prefers and ends with the avoided ones.

Parameter Type Default Description
target list[str]/list[list[str]] - Words of the target corpus - one text or a list of texts
comparison list[str]/list[list[str]] - Words of the comparison corpus
segment_size int 2000 Size of a segment in words
top_n int None Number of words from the start of the list; None - all of them

The result is a list of ZetaScore(word, dp_target, dp_comparison, zeta, log_zeta) named tuples by descending Zeta, by descending logarithmic Zeta and alphabetically when equal.

Example

from anyts.corpus import zeta

zeta(corpus["A"], corpus["B"], segment_size=5, top_n=2)
# [ZetaScore(word='at', dp_target=0.5, dp_comparison=0.0, zeta=0.5, log_zeta=2.0),
#  ZetaScore(word='birds', dp_target=0.5, dp_comparison=0.0, zeta=0.5, log_zeta=2.0)]

zeta(corpus["A"], corpus["B"], segment_size=5)[-1]
# ZetaScore(word='on', dp_target=0.0, dp_comparison=0.5, zeta=-0.5, log_zeta=-2.0)

Kilgarriff's chi-square

The distance between two corpora after Kilgarriff (2001): for the n_mfw most frequent words of the joint corpus the expected frequencies in the corpora are proportional to their sizes, \(\chi^2 = \sum (O - E)^2 / E\) over the words and both corpora. The greater the value, the more the corpora differ; the value grows with the size of the corpora, so pairs of corpora are comparable with each other at equal sizes.

Parameter Type Default Description
words_a list[str] - Words of the first corpus
words_b list[str] - Words of the second corpus
n_mfw int 500 Number of the most frequent words of the joint corpus

Example

from anyts.corpus import kilgarriff_chi2

round(kilgarriff_chi2(corpus["A"], corpus["B"], n_mfw=5), 3)
# 4.113

Mendenhall curve

mendenhall_curve(words) - the shares of the words of every length in characters (Mendenhall 1887), a profile of the author comparable between texts whatever their size.

mendenhall_distance(words_a, words_b) - the Jensen-Shannon distance with base 2 between the curves, from 0 (the distributions coincide) to 1.

Example

from anyts.corpus import mendenhall_curve, mendenhall_distance

{length: round(share, 3) for length, share in mendenhall_curve(corpus["A"]).items()}
# {2: 0.095, 3: 0.524, 4: 0.095, 5: 0.143, 6: 0.095, 7: 0.048}

round(mendenhall_distance(corpus["A"], corpus["B"]), 3)
# 0.303