Skip to content

Corpus plots

anyts.visualizers.dispersion_plot(), anyts.visualizers.keyness_plot(), anyts.visualizers.collocation_network()

Description

Plots for the corpus measures: the lexical dispersion - where in a text a word occurs, a chart of the keywords found by keyness and a network of the collocations found by collocations. The matplotlib functions take the axes ax and return Axes: without ax a new figure is created, with it the plot goes into a grid of one's own; the network of collocations is built by graphviz and returns a Graph, as the word tree returns a Digraph.

Lexical dispersion

A row for every word of targets and a tick at the position of each of its occurrences in the text. Words are compared as they are: case and lemmatization belong to the extraction of the words.

Parameter Type Default Description
words list[str] - Words of the text in order
targets list[str] - Words whose occurrences are shown
ax Axes None Axes for the plot
labels dict[str, str] None Labels over the defaults: title, xlabel

Chart of keywords

Diverging horizontal bars: the words of positive to the right, of negative - the result of keyness with positive=False - to the left, the length of a bar is the absolute value of the field (score, g2, log_ratio), so the side is set by the list and not by the sign of the measure; top_n words on each side, the words with an undefined or infinite value are skipped. For the odds ratio (score from 0 to infinity, one - equal odds) set log=True: the absolute \(\log_2\) of the value is plotted, symmetric around one.

Parameter Type Default Description
positive list[Keyword] - Positive keywords
negative list[Keyword] () Negative keywords
top_n int 20 Number of words on each side
labels dict[str, str]/tuple[str, str] None Labels over the defaults: title, xlabel and xlabel_log (format strings with field), target and reference (the legend); a pair of strings sets the legend alone
field str score Field of Keyword whose values are plotted
log bool False Plot \(\log_2\) of the value - for the odds ratio
ax Axes None Axes for the plot

Network of collocations

An undirected graph: the nodes are the words with the size of the font by frequency, the edges the pairs with the width and the label by the value of the measure; the neato layout. A pair of a word with itself - a word repeated within the window - would be a loop and is left out before top_n pairs are taken. Rendering needs the executables of Graphviz; in Jupyter the graph displays itself, and graph.render("network") saves a png file.

Parameter Type Default Description
collocations list[Collocation] - Collocations
top_n int None Number of pairs from the start of the list; None - all of them

Usage example

Example

from anyts.corpus import collocations, keyness
from anyts.visualizers import collocation_network, dispersion_plot, keyness_plot

target = "the cat sleeps and the cat eats and the cat purrs".split()
reference = "the dog sleeps and the dog eats and the dog barks".split()
dispersion_plot(target, ["cat", "eats"])
keyness_plot(keyness(target, reference), keyness(target, reference, positive=False))
print(collocation_network(collocations(target, window=1)).source)