Phonetic helpers¶
anyts.phonetics
Functions over the sounds or the letters of the words of a text. A word is given as the collection of its features - the characters of a string, the sounds of a transcription - so the functions see no language: a library chooses the features.
calc_repetition_index¶
The repetition index: the number of windows of window_len neighbouring words where a feature - a letter or a sound, as a library chooses - occurs in two words or more, summed over the features, to the number expected if the words stood in random order. A window of shuffled words is a sample of them without replacement, so for a feature found in \(K\) of the \(N\) words of the text a window of \(w\) words holds it in two words or more with the hypergeometric probability
and the expected number is the sum of these probabilities over the features times the \(N - w + 1\) windows. The expectation comes from the text itself, so the index tells whether the repetitions gather in neighbouring words, not how frequent a feature is: it averages to 1 over the orders of the words, and is well above 1 when the repetitions come closer than chance. A word counts a feature once, however often it holds it; nan for a text shorter than the window and for a text where no feature is shared by two words.
| Parameter | Type | Default | Description |
|---|---|---|---|
words |
list[str]/list[list[str]] | - |
Features of every word: a string of its letters or a collection of its sounds |
window_len |
int | - |
Window in words, at least 2 |
features |
Collection[str] | None |
Features to count, all of them by default |
Example
from anyts.phonetics import calc_repetition_index
words = ["sea", "sun", "gas", "fog", "box", "hat"]
calc_repetition_index(words, 2, {"s"})
# 2.0