Skip to content

Text highlighting

anyts.visualizers.HighlightedText, anyts.visualizers.Highlight; the helpers of the layers are in anyts.visualizers.highlight

Description

The machinery of the highlighting of a text by layers, in the manner of the style checkers: every layer marks the fragments of a text that a statistic counts - long sentences, complex words, the passive. The data source can be either a text or a Doc object of spaCy. The result shows in Jupyter as HTML with styles and a legend; the method to_html returns the same markup for documentation and web applications, and the fragments are kept in the attribute highlights for a rendering of one's own. Fragments of different layers may overlap.

Language hooks

The core has no layers of its own: a language library subclasses HighlightedText and sets its layers and their search.

Hook Kind Description
layers_desc class attribute The layers in the order of drawing with their names in the legend, dict[str, str]
default_layers class attribute The layers on by default, among those the source allows, tuple[str]
layer_annotations class attribute The annotations of a Doc a layer needs (DEP, POS, LEMMA...), dict[str, tuple[str]]; a layer without them is available for a string too
layer_styles class attribute The CSS declarations of every layer, dict[str, str]; the text of a layer with a background keeps a dark colour on it
css_prefix class attribute The prefix of the CSS classes, str
find(layer, words, sents, doc) method The fragments of a layer: a list of Highlight from the words and the sentences of the text and the Doc (None for a string)
iter_words(text), iter_sents(text) methods The words and the sentences of a string as triples of the start, the end and the text; by default those of the default tokenizers of WordsExtractor and SentsExtractor
doc_words(doc) method The words of a Doc; by default get_doc_words

The sentences of a Doc come from its boundaries, or from iter_sents over its text when it has none, and the words of a sentence are those of doc_words or iter_words that start in it. find runs in __init__, so a subclass with parameters of its own stores them before it calls the __init__ of the base; a fragment of another layer or one outside the text raises ParameterError.

In the core the layers are empty and css_prefix is anyts.

Parameters

Parameter Type Default Description
source str/Doc - Data source (a string or a Doc object)
layers list[str]/str None Layers of the highlighting; if not given, the default layers the source allows; "all" - every layer allowed

A source that is neither a string nor a Doc raises SourceTypeError, a source without words SourceError; layers that are not a name or a list of names, a layer that is unknown or one that needs an annotation the source lacks raise ParameterError.

Attributes

Attribute Type Description
text str Text of the data source
layers tuple[str] Layers turned on, in the order of drawing
highlights tuple[Highlight] Highlighted fragments, sorted by their start and then by descending end
counts dict[str, int] Number of fragments of every layer

A Highlight fragment is an immutable object with the fields start and end (positions in the text), layer (the layer) and note (the explanation for the tooltip).

Methods

to_html

Returns the HTML markup of the highlighted text: a div of the class <prefix>-highlight with the legend and its counts and the text, where the highlighted segments are wrapped in a span of the classes <prefix>-hl and <prefix>-hl-<layer>, the notes going to the attribute title; <prefix> is the prefix of the CSS classes. Overlapping fragments of different layers give segments with several classes, in the order of drawing. Line breaks (\n, \r\n, \r) are kept as character references, one per break, so the markup can be put into Markdown.

Parameter Type Default Description
legend bool True Add the legend with the counts of the fragments
css bool True Add the styles of the layers

css

The class method css() returns the styles to_html adds: the container, the legend and the text under the classes of css_prefix, the declarations of layer_styles for every layer, and a dark colour of the text on the layers with a background.

Helpers of the layers

Word(start, end, text, pos=None, lemma=None) - a word with its position in the text and, when known, its part of speech and lemma.

Sent(start, end, n_words) - a sentence with its position and number of words.

get_doc_words(doc) - the words of iter_doc_tokens with their positions, a byte order mark at the start of the text left out, with the part of speech and the lemma when the Doc has them.

iter_doc_sents(doc) - the sentences of the boundaries of a Doc as triples of the start, the end and the text; the whitespace tokens at the edges of a sentence are left out, and a sentence of whitespace alone is skipped.

get_text_sents(spans, words) - the sentences given by their positions with their numbers of words; a word belongs to the sentence of its first character.

group_words_by_sents(words, sents) - the words of every sentence; the words out of the sentences are skipped.

tokens_span(tokens) - the positions of the fragment that covers some tokens, without the punctuation and the whitespace at its edges.

split_segments(length, highlights) - the segments of a text with the same set of fragments: their bounds are the starts and the ends of all the fragments, and overlapping or nested fragments give segments with several layers.

Usage example

A layer of the long sentences in a subclass.

Example

from anyts.visualizers import Highlight, HighlightedText


class Highlighter(HighlightedText):
    layers_desc = {"long_sents": "long sentence"}
    default_layers = ("long_sents",)
    layer_styles = {"long_sents": "background: #fef9c3;"}
    css_prefix = "demo"

    def find(self, layer, words, sents, doc):
        return [
            Highlight(sent.start, sent.end, layer, f"{sent.n_words} words")
            for sent in sents
            if sent.n_words >= 8
        ]


ht = Highlighter("The cat sleeps. The dog barks at the cat that sleeps on the mat.")
ht.highlights
# (Highlight(start=16, end=64, layer='long_sents', note='11 words'),)
ht.to_html(legend=False, css=False)
# '<div class="demo-highlight"><div class="demo-highlight-text">The cat sleeps. <span class="demo-hl demo-hl-long_sents" title="11 words">The dog barks at the cat that sleeps on the mat.</span></div></div>'