Skip to content

Utilities

anyts.utils

Helper functions shared by the extractors and the statistics.

is_punctuation

Checks whether a token consists only of punctuation marks and symbols: the characters of the Unicode categories P (punctuation), S (symbols), M (combining marks) and Cf (invisible format characters such as the zero-width space U+200B, the byte order mark U+FEFF and the zero-width joiner U+200D). Multi-character tokens like ?! and --, symbols like € and № and a lone invisible character that a tokenizer splits off are punctuation too, while a token with a letter or a digit is not (e with a combining acute accent is a word); an empty token is punctuation as well.

Parameter Type Default Description
token str - Token

Example

from anyts.utils import is_punctuation

is_punctuation("?!"), is_punctuation("€"), is_punctuation("no.")
# (True, True, False)

count_letters

Counts the letters of a word: the characters of any alphabet (str.isalpha), without digits, hyphens and marks. The results are cached by word form, so the function is meant for words, not whole texts.

Parameter Type Default Description
word str - Word form

Example

from anyts.utils import count_letters

count_letters("well-known"), count_letters("3rd")
# (9, 2)

safe_divide

Divides two numbers and returns default for a zero denominator.

Parameter Type Default Description
num float/int - Numerator
den float/int - Denominator
default float/int 0 Value returned for a zero denominator

has_words

Checks whether a text, a Doc or a Span holds a word: an empty text or one of whitespace and the characters of is_punctuation alone holds none.

Parameter Type Default Description
source str/Doc/Span - Text, Doc or Span object

Example

from anyts.utils import has_words

has_words("The cat sleeps"), has_words("?!"), has_words("")
# (True, False, False)

check_sequence

Checks that an argument is a sequence and not a text or an iterator: a string (of characters or bytes), a Doc, a Span, an iterator or an object that cannot be iterated raises SourceTypeError, since a string would be iterated character by character and an iterator would be exhausted by the first pass over it; so does a table - a two-dimensional array or a DataFrame. A set or a mapping holds every item once, in no order of the text, and raises the error too, unless ordered=False - for a collection whose order and repeats do not matter, such as stop words.

Parameter Type Default Description
value object - Value to check
what str "words" What is expected, for the message of the error
ordered bool True Whether the order and the repeats of the items matter

check_words

Checks that an argument is a list of words: it passes check_sequence and every item is a string, so that a list of spaCy tokens, which would count every token as a lexeme of its own, raises SourceTypeError.

Parameter Type Default Description
value Iterable - Value to check
what str "words" What is expected, for the message of the error
ordered bool True Whether the order and the repeats of the words matter

check_integer

Checks that a parameter is an integer: a bool or a float, even a whole one like 5.0, raises ParameterError, since it would fail only later, as an index or a count.

Parameter Type Default Description
value object - Value to check
what str - Name of the parameter, for the message of the error

check_number

Checks that a parameter is a finite real number: a bool, a string, None, nan or an infinity raises ParameterError instead of failing later in a comparison.

Parameter Type Default Description
value object - Value to check
what str - Name of the parameter, for the message of the error

check_counts

Checks that an argument is a counter - a mapping of words to their frequencies, such as a Counter: a value that is not a mapping, a word that is not a string or a frequency that is not a number raises SourceTypeError, and a negative, nan or infinite frequency SourceError.

Parameter Type Default Description
value object - Value to check

merge_labels

Merges the labels given for a plot over its default ones: a label that is not given keeps its default, so a single label can be changed alone. A default with fields in braces, such as "Moving average ({window})", is a format string: a label given for it may use only those fields, and a literal brace in it is doubled ({{); the other labels are taken as they are. A format label is tried on sample values of its fields, of the types the plot gives them, so that a format spec the values do not take ({window:s} for a number) fails before a figure is created. Labels that are not a mapping of strings, have a key the plot does not know, a field their default does not have or a format spec their values do not take raise ParameterError.

Parameter Type Default Description
defaults dict[str, str] - Default labels by key
labels dict[str, str] - Labels given; None - the default ones
**samples object - Values of the fields of the format labels, for trying the labels given

The defaults of the visualizers of the core are anyts.constants.VISUALIZER_LABELS, by the name of the function.

Example

from anyts.utils import merge_labels

merge_labels({"title": "Plot", "xlabel": "x"}, {"title": "My plot"})
# {'title': 'My plot', 'xlabel': 'x'}

iter_doc_tokens

Yields the tokens of the words of a Doc or a Span: whitespace tokens and the tokens of is_punctuation, symbols like % and € included, are skipped.

Parameter Type Default Description
source Doc/Span - Doc or Span object

iter_doc_units

Yields the words of a Doc or a Span as lists of tokens: one token each, as in iter_doc_tokens. With join_hyphens=True a word that the tokenizer split at its hyphens (well-known into well, -, known) is joined back when no whitespace separates its parts and no sentence starts at a hyphen or a part; a language library whose own tokenizer keeps such words whole turns it on, so that a string and a Doc give the same words.

Parameter Type Default Description
source Doc/Span - Doc or Span object
join_hyphens bool False Join the parts of hyphenated words

iter_doc_words

Yields the words of iter_doc_units as tuples: the position of the first character, the position after the last character and the text of the word. A byte order mark glued to the start of a word, as in a file read with utf-8 instead of utf-8-sig, is dropped.

Parameter Type Default Description
source Doc/Span - Doc or Span object
join_hyphens bool False Join the parts of hyphenated words

Example

import spacy

from anyts.utils import iter_doc_words

doc = spacy.blank("xx")("A well-known cat")
[word for _, _, word in iter_doc_words(doc)]
# ['A', 'well', 'known', 'cat']
list(iter_doc_words(doc, join_hyphens=True))
# [(0, 1, 'A'), (2, 12, 'well-known'), (13, 16, 'cat')]