Skip to content

Text highlighting

ests.visualizers.highlight(), ests.visualizers.HighlightedText, ests.visualizers.Highlight

Description

The machinery of the highlighting of a text by layers, in the manner of the style checkers: every layer marks the fragments of a text that a statistic counts - long sentences, complex words, the passive. The data source can be either a text or a Doc object of spaCy. The result shows in Jupyter as HTML with styles and a legend; the method to_html returns the same markup for documentation and web applications, and the fragments are kept in the attribute highlights for a rendering of one's own. Fragments of different layers may overlap.

The function highlight returns a HighlightedText object of esTS, which extends the HighlightedText of the anyTS core with the layers of the library, their search and their parameters, and the prefix ests of the CSS classes.

Layers of the highlighting:

Group Layer What it marks Statistic
Readability long_sents Sentences of long_sent_word_factor words or more BasicStats
complex_words Words of complex_syl_factor syllables or more BasicStats
rare_words Words with a lemma beyond the embedded top 10000; numbers, words with a hyphen or a digit and stopwords are not highlighted LexicalStats
Syntax passive Passive verb forms with their auxiliary ser or their se; the note tells the two apart and marks the passive with ser without an agent SyntaxStats
participle_clauses Participial clauses SyntaxStats
gerund_clauses Gerund clauses SyntaxStats
de_chains Chains of complements with de with their head SyntaxStats
split_predicates Split predicates from the verb to the noun SyntaxStats
Officialese verbal_nouns Nouns derived from a verb StyleStats
compound_prepositions Compound prepositions of COMPOUND_PREPOSITIONS StyleStats
cliches Clichés of OFFICIALESE_CLICHES or of the parameter cliches StyleStats
Style stopwords Stopwords of STOPWORDS or of the parameter stopwords, the water of the text StyleStats
parentheticals Parenthetical expressions StyleStats
connectors Discourse markers, the class and the kind in the note CohesionStats
Phonics alliteration Repetitions of a consonant sound in neighbouring words, unlikely by the frequencies of the Spanish sounds; the note gives the sound and the letters that write it PhonStats

The groups are defined in ests.constants.HIGHLIGHT_LAYER_GROUPS, the annotations of a Doc a layer needs in HIGHLIGHT_LAYER_ANNOTATIONS and the styles of the layers in HIGHLIGHT_LAYER_STYLES. The layers of the group "Syntax" need a Doc with a parse and a lemmatizer (the models es_core_news_sm, es_core_news_md, es_core_news_lg), as SyntaxStats does; verbal_nouns needs a Doc with the parts of speech and the lemmas. A Doc without the sentence boundaries (a blank pipeline, a pipeline without the parser) is split by the rules of SentsExtractor. The layers of HIGHLIGHT_DEFAULT_LAYERS the source allows are on by default - long sentences, complex words, passive, chains of de, split predicates, clichés; layers="all" turns on every layer allowed. The layers overlap (a compound preposition is made of stopwords), so pick the ones you need.

A sentence is long from 30 words, the bound of the Spanish guides to plain language (the Comunidad de Madrid and the Gobierno de la Ciudad de Buenos Aires), and a word is complex from four syllables; complex_syl_factor=3, the bound of BasicStats, shows the words of n_complex_words.

Note

An alliteration is looked for inside a sentence as a run of two neighbouring words or more with the same consonant sound in their stems. The sounds are the ones of the transcription, so casa and queso repeat k; the stem is the common start of the transcriptions of the word form and of its lemma (cantaban - k a n t a), which leaves out the endings that repeat by agreement (las casas blancas). The probability of a run under an independent spread of the sounds is the product over its words of the probability to meet the consonant among the sounds of the stem, \(1 - (1 - f)^n\), where \(f\) is the frequency of the sound in the corpus of literature and \(n\) the number of the sounds of the stem; a run is highlighted when the probability is below alliteration_threshold. A repetition of a rare sound shows in two or three words (deje la abeja), while the s of los suspiros se escapan is too frequent to count in two words. Words shorter than three letters (de, la, el) and stopwords (que, los, con) neither break nor continue a run. The alliteration index of PhonStats measures how the repetitions cluster over the whole text; the highlighting shows where they are.

Parameters

Parameter Type Default Description
source str/Doc - Data source (a string or a Doc object)
layers list[str]/str None Layers of the highlighting; if not given, the default layers the source allows; "all" - every layer allowed

A source that is neither a string nor a Doc raises SourceTypeError, a source without words SourceError; layers that are not a name or a list of names, a layer that is unknown or one that needs an annotation the source lacks raise ParameterError.

The parameters of the layers of esTS:

Parameter Type Default Description
long_sent_word_factor int 30 Minimum number of words of a long sentence
complex_syl_factor int 4 Minimum number of syllables of a complex word
stopwords list[str]/set[str] None List or set of stopwords; if not given, STOPWORDS and the one-word parenthetical expressions
cliches list[str]/set[str] None List or set of clichés; if not given, OFFICIALESE_CLICHES
alliteration_threshold float 0.001 Probability of a repetition of a consonant under an independent spread of the sounds, below which the repetition is alliteration

A threshold that is not an integer of at least one or a probability outside (0, 1] raises ParameterError, stopwords or clichés that are not strings SourceTypeError.

Attributes

Attribute Type Description
text str Text of the data source
layers tuple[str] Layers turned on, in the order of drawing
highlights tuple[Highlight] Highlighted fragments, sorted by their start and then by descending end
counts dict[str, int] Number of fragments of every layer

A Highlight fragment is an immutable object with the fields start and end (positions in the text), layer (the layer) and note (the explanation for the tooltip).

The note of a fragment gives the number of words of the sentence, of syllables of the word, the length of the chain or the phrase of the list.

Methods

to_html

Returns the HTML markup of the highlighted text: a div of the class <prefix>-highlight with the legend and its counts and the text, where the highlighted segments are wrapped in a span of the classes <prefix>-hl and <prefix>-hl-<layer>, the notes going to the attribute title; <prefix> is the prefix of the CSS classes. Overlapping fragments of different layers give segments with several classes, in the order of drawing. Line breaks (\n, \r\n, \r) are kept as character references, one per break, so the markup can be put into Markdown.

Parameter Type Default Description
legend bool True Add the legend with the counts of the fragments
css bool True Add the styles of the layers

css

The class method css() returns the styles to_html adds: the container, the legend and the text under the classes of css_prefix, the declarations of layer_styles for every layer, and a dark colour of the text on the layers with a background.

Usage example

Example

Code:

import spacy
from ests.visualizers import highlight

nlp = spacy.load("es_core_news_sm")
text = (
    "El proyecto, elaborado durante el verano, fue aprobado por el consejo sin debate. "
    "El aumento de la eficiencia del uso de los recursos públicos se analizó, siguiendo "
    "las normas del reglamento. "
    "Los representantes de los ministerios regionales no consiguieron hacer una revisión "
    "conjunta de las cuestiones de financiación y de reparto de responsabilidades entre los "
    "organismos, puesto que cada uno de ellos defendía su propia interpretación de las "
    "disposiciones del acuerdo. "
    "En el marco de la reunión se procedió a la votación, y la decisión quedó aplazada "
    "hasta la próxima sesión."
)

# Highlight the text with the default layers
ht = highlight(nlp(text))
ht.counts
# {'long_sents': 1, 'complex_words': 16, 'passive': 2, 'de_chains': 3,
#  'split_predicates': 2, 'cliches': 1}

ht.highlights[:2]
# (Highlight(start=13, end=22, layer='complex_words', note='complex word, 5 syllables'),
#  Highlight(start=42, end=54, layer='passive', note='passive with ser'))

# Every layer
highlight(nlp(text), layers="all").counts
# {'long_sents': 1, 'complex_words': 16, 'rare_words': 0, 'passive': 2,
#  'participle_clauses': 2, 'gerund_clauses': 1, 'de_chains': 3, 'split_predicates': 2,
#  'verbal_nouns': 9, 'compound_prepositions': 1, 'cliches': 1, 'stopwords': 48,
#  'parentheticals': 0, 'connectors': 3, 'alliteration': 0}

# Show in Jupyter or save the markup
ht
html = ht.to_html()

Result (hover over a fragment to see its note):

Long sentences1Complex words16Passive2Chains of de3Split predicates2Clichés1
El proyecto, elaborado durante el verano, fue aprobado por el consejo sin debate. El aumento de la eficiencia del uso de los recursos públicos se analizó, siguiendo las normas del reglamento. Los representantes de los ministerios regionales no consiguieron hacer una revisión conjunta de las cuestiones de financiación y de reparto de responsabilidades entre los organismos, puesto que cada uno de ellos defendía su propia interpretación de las disposiciones del acuerdo. En el marco de la reunión se procedió a la votación, y la decisión quedó aplazada hasta la próxima sesión.

A string without a model of spaCy allows every layer but the syntactic ones and verbal_nouns: long sentences, complex words and clichés by default, and the alliteration on request.

Example

from ests.visualizers import highlight

# Antonio Machado, Poesías completas (the corpus of literature)
text = "El cierzo corre por el campo yerto alborotando en blancos torbellinos la nieve silenciosa."
highlight(text, layers="alliteration").highlights
# (Highlight(start=35, end=78, layer='alliteration', note='alliteration on /b/ (b, v)'),)

Example

from ests.visualizers import highlight

text = "Se procedió a la revisión del expediente en el marco del plan a la mayor brevedad."
highlight(text, layers=["compound_prepositions", "cliches"]).highlights
# (Highlight(start=3, end=13, layer='cliches', note='cliché: «proceder a»'),
#  Highlight(start=41, end=56, layer='compound_prepositions', note='compound preposition: «en el marco de»'),
#  Highlight(start=62, end=81, layer='cliches', note='cliché: «a la mayor brevedad»'))