KWIC concordance¶
anyts.corpus.kwic(), anyts.corpus.format_kwic(), anyts.corpus.print_kwic(), anyts.corpus.Concordance
Description¶
A KWIC concordance (keyword in context) - every occurrence of a word or a phrase with its context on the left and on the right. The occurrences are looked for among the words of the text by the word form, ignoring case or respecting it (ignore_case=False), or by the lemma (by_lemma=True), where the words of the keyword are lemmatized as words of a string: with a lemmatizer a phrase may be written in any form, without one it is given by lemmas. The words of a string and of the keyword come from the tokenizer, those of a Doc from its tokens, so a tokenizer that splits a string as the pipeline does finds in a Doc what it finds in its text; punctuation and symbols are not words. The forms are compared in the composed form of Unicode (NFC), without soft hyphens. A phrase does not run across the end of a paragraph or of a sentence - a boundary of a Doc anywhere between its words, an opening mark included, or in a text without boundaries the start of a sentence of the sentence splitter, by default a final mark before whitespace - unless the keyword has a final mark or the start of a sentence in the same place (U.S. Army). The context is window words on each side as they are written in the text, with the punctuation between them; whitespace collapses into one space, and occurrences do not overlap.
Language hooks¶
The language enters through five parameters, which a language library fills in its own kwic:
| Parameter | Default | Description |
|---|---|---|
tokenize |
a run of word characters | The words of a string and of the keyword as triples of the start, the end and the text |
lemmatize |
the word and the lemma of the model | The lemmas of a word, given its text and its tokens in a Doc (none for a string and the keyword), or one lemma as a string; a word matches when one of its lemmas is one of those of the keyword |
fold |
str.lower |
The folding of a word form (with ignore_case) or a lemma before the comparison, such as the letters a language takes for the same |
sentenize |
a final mark before whitespace | The sentences of a string, of the keyword and of the text of a Doc without boundaries as triples of the start, the end and the text; a phrase does not run from one sentence into the next |
join_hyphens |
False |
Join the parts of the hyphenated words of a Doc the tokenizer split (iter_doc_units); the default tokenizer then keeps them whole in a string and in the keyword too |
The words of a Doc are its tokens, a byte order mark at the start of a word left out. By default the lemmas of a word are the word itself and, for a word of one token of a Doc with lemmas, the lemma of the model, so a word is found by its own form and a Doc of a pipeline with a lemmatizer is searched by its lemmas.
Parameters¶
| Parameter | Type | Default | Description |
|---|---|---|---|
source |
str/Doc | - |
Text or Doc object |
keyword |
str | - |
Word or phrase |
window |
int | 5 |
Number of words of context on each side |
by_lemma |
bool | False |
Compare lemmas instead of word forms |
ignore_case |
bool | True |
Ignore case when comparing word forms |
A keyword without words and a negative window raise ParameterError; a source that is neither a string nor a Doc, a keyword that is not a string and a hook that is not callable raise SourceTypeError.
format_kwic(concordances, width=40) aligns the lines on the keyword: the left context is cut on the left and aligned to the right, the right one is cut on the right; width is the width of a context in characters, at least one. The lines may come from any iterable and are taken in the composed form of Unicode (NFC), so that the widths count the characters as they are shown.
print_kwic(concordances, width=40) prints the result of format_kwic.
Result¶
A list of Concordance named tuples in the order of the text.
| Field | Type | Description |
|---|---|---|
start |
int | Position of the first character of the occurrence in the text |
end |
int | Position after the last character of the occurrence |
left |
str | Context on the left |
keyword |
str | Occurrence as written in the text |
right |
str | Context on the right |
Example¶
Example
from anyts.corpus import kwic, print_kwic
text = (
"The cat was at the window and watched the birds. The birds flew away and the cat "
"fell asleep at the window. Tomorrow the cat will be at the window and watch the birds."
)
lines = kwic(text, "window", window=3)
lines[0]
# Concordance(start=19, end=25, left='was at the', keyword='window', right='and watched the')
print_kwic(lines, width=20)
# was at the window and watched the
# asleep at the window . Tomorrow the cat
# be at the window and watch the
lemmas = {"watched": "watch", "birds": "bird"}
[
line.keyword
for line in kwic(
text,
"watch the bird",
by_lemma=True,
lemmatize=lambda word, tokens: [lemmas.get(word, word)],
)
]
# ['watched the birds', 'watch the birds']