Skip to content

Corpus comparison

anyts.corpus.compare_features(), anyts.corpus.check_comparison_params(), anyts.corpus.compare_values(), anyts.corpus.calc_cohen_d(), anyts.corpus.calc_cliff_delta(), anyts.corpus.bootstrap_median_diff(), anyts.corpus.holm_correction()

Description

Comparing two corpora feature by feature. Each corpus comes as a table of the features of its windows - texts split into parts of about the same size, so that the features do not depend on the length of the texts. compare_features(table_a, table_b, labels, n_bootstrap, seed) compares the tables column by column and returns one row per feature, sorted by descending absolute Cliff's delta, the features without statistics last. The bootstrap resamples whole texts by the level text of the index of a table; a table without that level has every row taken for a text of its own. A column missing from one of the tables gives nan.

Statistics

For a feature with the values \(x_1 \dots x_{n_A}\) in corpus A and \(y_1 \dots y_{n_B}\) in corpus B (undefined and infinite values dropped; with fewer than two values on a side - nan):

Column Description
mean_A, mean_B, median_A, median_B means and medians
median_diff, ci_low, ci_high the difference of the medians and its 95% percentile bootstrap interval: both sets are resampled n_bootstrap times (bootstrap_median_diff)
cohen_d \(d = (\bar{x} - \bar{y}) / s\), \(s\) - the pooled standard deviation; 0.2 - a small effect, 0.5 - medium, 0.8 - large (calc_cohen_d)
cliff_delta \(\delta = P(x > y) - P(x < y)\) from −1 to 1; \(\lvert\delta\rvert\) < 0.147 - a negligible effect, < 0.33 - small, < 0.474 - medium, large otherwise (Romano et al. 2006; calc_cliff_delta)
auc the feature as a classifier on its own: the share of the pairs of windows where the value in A is greater than in B, ties counted as half; \(\delta = 2 \cdot AUC - 1\), 0.5 - the feature does not tell the corpora apart
u, p_value the Mann-Whitney U statistic and the two-sided p-value (scipy.stats.mannwhitneyu)
p_holm the p-value with Holm's correction for the number of features (holm_correction)
n_A, n_B number of windows with a defined value
n_texts_A, n_texts_B number of texts behind those windows

The names of the columns are anyts.corpus.COMPARISON_COLUMNS, with A and B replaced by the labels of the corpora. Cliff's delta and the AUC come from the same U statistic and agree with each other; Cohen's d is sensitive to outliers and to departures from normality, so it is best read next to the delta.

Windows of one text are not independent

The test and the effect sizes take every window for an independent observation, and the windows of one text are not. With few texts in a corpus the p-values are too small and reflect the texts chosen as much as the corpora. The bootstrap resamples whole texts instead (a cluster bootstrap), so its interval accounts for the spread between the texts; it needs at least two texts on each side and is rough with only a few.

Parameters

Parameter Type Default Description
table_a DataFrame - Features of the windows of the first corpus, one row per window
table_b DataFrame - Features of the windows of the second corpus
labels tuple[str, str] ("A", "B") Names of the corpora for the columns, two strings that give distinct columns
n_bootstrap int 1000 Number of bootstrap samples
seed int/Generator 0 Seed of the random number generator, a non-negative integer, or a numpy Generator; None - a random one

check_comparison_params(labels=("A", "B"), n_bootstrap=1000, seed=0) checks the parameters of compare_features and raises ParameterError unless the labels are two strings that give distinct columns (("diff", "B") would repeat median_diff), the number of bootstrap samples is an integer of at least one and the seed is None, a non-negative integer or a numpy Generator. Called before the tables are built from texts, it reports a wrong parameter before the texts are processed.

Example

import pandas as pd

from anyts.corpus import compare_features

short = pd.DataFrame({"length": [14.0, 15.0, 13.0, 16.0]})
long = pd.DataFrame({"length": [33.0, 35.0, 31.0, 29.0]})
result = compare_features(short, long, labels=("short", "long"), n_bootstrap=200)
result.loc[
    "length", ["mean_short", "mean_long", "median_diff", "ci_low", "ci_high", "cliff_delta"]
].tolist()
# [14.5, 32.0, -17.5, -20.5, -14.5, -1.0]

Functions of the statistics

compare_values(values_a, values_b, n_bootstrap=1000, rng=None, texts_a=None, texts_b=None) compares two sets of values of one feature, undefined and infinite values - a missing value (None, NA) included - dropped together with their texts, and returns the values of one row in the order of anyts.corpus.COMPARISON_COLUMNS, p_holm left nan; texts_a and texts_b give the text of every value for the bootstrap.

calc_cohen_d(values_a, values_b) - Cohen's d with the pooled sample variances (ddof=1); nan with fewer than two values on a side or without spread.

calc_cliff_delta(values_a, values_b) - Cliff's delta, the share of the pairs where the first value is greater minus the share where it is smaller; nan for an empty set or an undefined value. It is computed by sorting, in \(O((n_A + n_B) \log n_B)\) time and linear memory.

bootstrap_median_diff(values_a, values_b, n_bootstrap=1000, rng=None, confidence=0.95, texts_a=None, texts_b=None) - the percentile bootstrap interval of the difference of the medians. With texts_a and texts_b whole texts are resampled, and the median of a draw is that of the values of the drawn texts put together; with fewer than two texts on a side the interval is nan.

holm_correction(p_values) - Holm's correction for multiple comparisons: the p-values are sorted in ascending order, the i-th is multiplied by (m − i + 1), where m is the number of defined values, then the running maximum is taken and capped at one; nan stays nan, and a p-value outside [0, 1] raises ParameterError.

Example

from anyts.corpus import calc_cliff_delta, calc_cohen_d, holm_correction

round(calc_cohen_d([2, 4, 6, 8], [1, 3, 5, 7]), 3)
# 0.387
calc_cliff_delta([2, 4, 6, 8], [1, 3, 5, 7])
# 0.25
holm_correction([0.01, 0.04, 0.03]).round(3).tolist()
# [0.03, 0.06, 0.06]