Text highlighting¶
ests.visualizers.highlight(), ests.visualizers.HighlightedText, ests.visualizers.Highlight
Description¶
The machinery of the highlighting of a text by layers, in the manner of the style checkers: every layer marks the fragments of a text that a statistic counts - long sentences, complex words, the passive. The data source can be either a text or a Doc object of spaCy. The result shows in Jupyter as HTML with styles and a legend; the method to_html returns the same markup for documentation and web applications, and the fragments are kept in the attribute highlights for a rendering of one's own. Fragments of different layers may overlap.
The function highlight returns a HighlightedText object of esTS, which extends the HighlightedText of the anyTS core with the layers of the library, their search and their parameters, and the prefix ests of the CSS classes.
Layers of the highlighting:
| Group | Layer | What it marks | Statistic |
|---|---|---|---|
| Readability | long_sents |
Sentences of long_sent_word_factor words or more |
BasicStats |
complex_words |
Words of complex_syl_factor syllables or more |
BasicStats | |
rare_words |
Words with a lemma beyond the embedded top 10000; numbers, words with a hyphen or a digit and stopwords are not highlighted | LexicalStats | |
| Syntax | passive |
Passive verb forms with their auxiliary ser or their se; the note tells the two apart and marks the passive with ser without an agent |
SyntaxStats |
participle_clauses |
Participial clauses | SyntaxStats | |
gerund_clauses |
Gerund clauses | SyntaxStats | |
de_chains |
Chains of complements with de with their head |
SyntaxStats | |
split_predicates |
Split predicates from the verb to the noun | SyntaxStats | |
| Officialese | verbal_nouns |
Nouns derived from a verb | StyleStats |
compound_prepositions |
Compound prepositions of COMPOUND_PREPOSITIONS |
StyleStats | |
cliches |
Clichés of OFFICIALESE_CLICHES or of the parameter cliches |
StyleStats | |
| Style | stopwords |
Stopwords of STOPWORDS or of the parameter stopwords, the water of the text |
StyleStats |
parentheticals |
Parenthetical expressions | StyleStats | |
connectors |
Discourse markers, the class and the kind in the note | CohesionStats | |
| Phonics | alliteration |
Repetitions of a consonant sound in neighbouring words, unlikely by the frequencies of the Spanish sounds; the note gives the sound and the letters that write it | PhonStats |
The groups are defined in ests.constants.HIGHLIGHT_LAYER_GROUPS, the annotations of a Doc a layer needs in HIGHLIGHT_LAYER_ANNOTATIONS and the styles of the layers in HIGHLIGHT_LAYER_STYLES. The layers of the group "Syntax" need a Doc with a parse and a lemmatizer (the models es_core_news_sm, es_core_news_md, es_core_news_lg), as SyntaxStats does; verbal_nouns needs a Doc with the parts of speech and the lemmas. A Doc without the sentence boundaries (a blank pipeline, a pipeline without the parser) is split by the rules of SentsExtractor. The layers of HIGHLIGHT_DEFAULT_LAYERS the source allows are on by default - long sentences, complex words, passive, chains of de, split predicates, clichés; layers="all" turns on every layer allowed. The layers overlap (a compound preposition is made of stopwords), so pick the ones you need.
A sentence is long from 30 words, the bound of the Spanish guides to plain language (the Comunidad de Madrid and the Gobierno de la Ciudad de Buenos Aires), and a word is complex from four syllables; complex_syl_factor=3, the bound of BasicStats, shows the words of n_complex_words.
Note
An alliteration is looked for inside a sentence as a run of two neighbouring words or more with the same consonant sound in their stems. The sounds are the ones of the transcription, so casa and queso repeat k; the stem is the common start of the transcriptions of the word form and of its lemma (cantaban - k a n t a), which leaves out the endings that repeat by agreement (las casas blancas). The probability of a run under an independent spread of the sounds is the product over its words of the probability to meet the consonant among the sounds of the stem, \(1 - (1 - f)^n\), where \(f\) is the frequency of the sound in the corpus of literature and \(n\) the number of the sounds of the stem; a run is highlighted when the probability is below alliteration_threshold. A repetition of a rare sound shows in two or three words (deje la abeja), while the s of los suspiros se escapan is too frequent to count in two words. Words shorter than three letters (de, la, el) and stopwords (que, los, con) neither break nor continue a run. The alliteration index of PhonStats measures how the repetitions cluster over the whole text; the highlighting shows where they are.
Parameters¶
| Parameter | Type | Default | Description |
|---|---|---|---|
source |
str/Doc | - |
Data source (a string or a Doc object) |
layers |
list[str]/str | None |
Layers of the highlighting; if not given, the default layers the source allows; "all" - every layer allowed |
A source that is neither a string nor a Doc raises SourceTypeError, a source without words SourceError; layers that are not a name or a list of names, a layer that is unknown or one that needs an annotation the source lacks raise ParameterError.
The parameters of the layers of esTS:
| Parameter | Type | Default | Description |
|---|---|---|---|
long_sent_word_factor |
int | 30 |
Minimum number of words of a long sentence |
complex_syl_factor |
int | 4 |
Minimum number of syllables of a complex word |
stopwords |
list[str]/set[str] | None |
List or set of stopwords; if not given, STOPWORDS and the one-word parenthetical expressions |
cliches |
list[str]/set[str] | None |
List or set of clichés; if not given, OFFICIALESE_CLICHES |
alliteration_threshold |
float | 0.001 |
Probability of a repetition of a consonant under an independent spread of the sounds, below which the repetition is alliteration |
A threshold that is not an integer of at least one or a probability outside (0, 1] raises ParameterError, stopwords or clichés that are not strings SourceTypeError.
Attributes¶
| Attribute | Type | Description |
|---|---|---|
text |
str | Text of the data source |
layers |
tuple[str] | Layers turned on, in the order of drawing |
highlights |
tuple[Highlight] | Highlighted fragments, sorted by their start and then by descending end |
counts |
dict[str, int] | Number of fragments of every layer |
A Highlight fragment is an immutable object with the fields start and end (positions in the text), layer (the layer) and note (the explanation for the tooltip).
The note of a fragment gives the number of words of the sentence, of syllables of the word, the length of the chain or the phrase of the list.
Methods¶
to_html¶
Returns the HTML markup of the highlighted text: a div of the class <prefix>-highlight with the legend and its counts and the text, where the highlighted segments are wrapped in a span of the classes <prefix>-hl and <prefix>-hl-<layer>, the notes going to the attribute title; <prefix> is the prefix of the CSS classes. Overlapping fragments of different layers give segments with several classes, in the order of drawing. Line breaks (\n, \r\n, \r) are kept as character references, one per break, so the markup can be put into Markdown.
| Parameter | Type | Default | Description |
|---|---|---|---|
legend |
bool | True |
Add the legend with the counts of the fragments |
css |
bool | True |
Add the styles of the layers |
css¶
The class method css() returns the styles to_html adds: the container, the legend and the text under the classes of css_prefix, the declarations of layer_styles for every layer, and a dark colour of the text on the layers with a background.
Usage example¶
Example
Code:
import spacy
from ests.visualizers import highlight
nlp = spacy.load("es_core_news_sm")
text = (
"El proyecto, elaborado durante el verano, fue aprobado por el consejo sin debate. "
"El aumento de la eficiencia del uso de los recursos públicos se analizó, siguiendo "
"las normas del reglamento. "
"Los representantes de los ministerios regionales no consiguieron hacer una revisión "
"conjunta de las cuestiones de financiación y de reparto de responsabilidades entre los "
"organismos, puesto que cada uno de ellos defendía su propia interpretación de las "
"disposiciones del acuerdo. "
"En el marco de la reunión se procedió a la votación, y la decisión quedó aplazada "
"hasta la próxima sesión."
)
# Highlight the text with the default layers
ht = highlight(nlp(text))
ht.counts
# {'long_sents': 1, 'complex_words': 16, 'passive': 2, 'de_chains': 3,
# 'split_predicates': 2, 'cliches': 1}
ht.highlights[:2]
# (Highlight(start=13, end=22, layer='complex_words', note='complex word, 5 syllables'),
# Highlight(start=42, end=54, layer='passive', note='passive with ser'))
# Every layer
highlight(nlp(text), layers="all").counts
# {'long_sents': 1, 'complex_words': 16, 'rare_words': 0, 'passive': 2,
# 'participle_clauses': 2, 'gerund_clauses': 1, 'de_chains': 3, 'split_predicates': 2,
# 'verbal_nouns': 9, 'compound_prepositions': 1, 'cliches': 1, 'stopwords': 48,
# 'parentheticals': 0, 'connectors': 3, 'alliteration': 0}
# Show in Jupyter or save the markup
ht
html = ht.to_html()
Result (hover over a fragment to see its note):
A string without a model of spaCy allows every layer but the syntactic ones and verbal_nouns: long sentences, complex words and clichés by default, and the alliteration on request.
Example
from ests.visualizers import highlight
# Antonio Machado, Poesías completas (the corpus of literature)
text = "El cierzo corre por el campo yerto alborotando en blancos torbellinos la nieve silenciosa."
highlight(text, layers="alliteration").highlights
# (Highlight(start=35, end=78, layer='alliteration', note='alliteration on /b/ (b, v)'),)
Example
from ests.visualizers import highlight
text = "Se procedió a la revisión del expediente en el marco del plan a la mayor brevedad."
highlight(text, layers=["compound_prepositions", "cliches"]).highlights
# (Highlight(start=3, end=13, layer='cliches', note='cliché: «proceder a»'),
# Highlight(start=41, end=56, layer='compound_prepositions', note='compound preposition: «en el marco de»'),
# Highlight(start=62, end=81, layer='cliches', note='cliché: «a la mayor brevedad»'))