Skip to content

KWIC concordance

anyts.corpus.kwic(), anyts.corpus.format_kwic(), anyts.corpus.print_kwic(), anyts.corpus.Concordance

Description

A KWIC concordance (keyword in context) - every occurrence of a word or a phrase with its context on the left and on the right. The occurrences are looked for among the words of the text by the word form, ignoring case or respecting it (ignore_case=False), or by the lemma (by_lemma=True), where the words of the keyword are lemmatized as words of a string: with a lemmatizer a phrase may be written in any form, without one it is given by lemmas. The words of a string and of the keyword come from the tokenizer, those of a Doc from its tokens, so a tokenizer that splits a string as the pipeline does finds in a Doc what it finds in its text; punctuation and symbols are not words. The forms are compared in the composed form of Unicode (NFC), without soft hyphens. A phrase does not run across the end of a paragraph or of a sentence - a boundary of a Doc anywhere between its words, an opening mark included, or in a text without boundaries the start of a sentence of the sentence splitter, by default a final mark before whitespace - unless the keyword has a final mark or the start of a sentence in the same place (U.S. Army). The context is window words on each side as they are written in the text, with the punctuation between them; whitespace collapses into one space, and occurrences do not overlap.

Language hooks

The language enters through five parameters, which a language library fills in its own kwic:

Parameter Default Description
tokenize a run of word characters The words of a string and of the keyword as triples of the start, the end and the text
lemmatize the word and the lemma of the model The lemmas of a word, given its text and its tokens in a Doc (none for a string and the keyword), or one lemma as a string; a word matches when one of its lemmas is one of those of the keyword
fold str.lower The folding of a word form (with ignore_case) or a lemma before the comparison, such as the letters a language takes for the same
sentenize a final mark before whitespace The sentences of a string, of the keyword and of the text of a Doc without boundaries as triples of the start, the end and the text; a phrase does not run from one sentence into the next
join_hyphens False Join the parts of the hyphenated words of a Doc the tokenizer split (iter_doc_units); the default tokenizer then keeps them whole in a string and in the keyword too

The words of a Doc are its tokens, a byte order mark at the start of a word left out. By default the lemmas of a word are the word itself and, for a word of one token of a Doc with lemmas, the lemma of the model, so a word is found by its own form and a Doc of a pipeline with a lemmatizer is searched by its lemmas.

Parameters

Parameter Type Default Description
source str/Doc - Text or Doc object
keyword str - Word or phrase
window int 5 Number of words of context on each side
by_lemma bool False Compare lemmas instead of word forms
ignore_case bool True Ignore case when comparing word forms

A keyword without words and a negative window raise ParameterError; a source that is neither a string nor a Doc, a keyword that is not a string and a hook that is not callable raise SourceTypeError.

format_kwic(concordances, width=40) aligns the lines on the keyword: the left context is cut on the left and aligned to the right, the right one is cut on the right; width is the width of a context in characters, at least one. The lines may come from any iterable and are taken in the composed form of Unicode (NFC), so that the widths count the characters as they are shown.

print_kwic(concordances, width=40) prints the result of format_kwic.

Result

A list of Concordance named tuples in the order of the text.

Field Type Description
start int Position of the first character of the occurrence in the text
end int Position after the last character of the occurrence
left str Context on the left
keyword str Occurrence as written in the text
right str Context on the right

Example

Example

from anyts.corpus import kwic, print_kwic

text = (
    "The cat was at the window and watched the birds. The birds flew away and the cat "
    "fell asleep at the window. Tomorrow the cat will be at the window and watch the birds."
)
lines = kwic(text, "window", window=3)
lines[0]
# Concordance(start=19, end=25, left='was at the', keyword='window', right='and watched the')

print_kwic(lines, width=20)
#           was at the  window  and watched the
#        asleep at the  window  . Tomorrow the cat
#            be at the  window  and watch the

lemmas = {"watched": "watch", "birds": "bird"}
[
    line.keyword
    for line in kwic(
        text,
        "watch the bird",
        by_lemma=True,
        lemmatize=lambda word, tokens: [lemmas.get(word, word)],
    )
]
# ['watched the birds', 'watch the birds']