Skip to content

Sentence extraction

anyts.extractors.SentsExtractor

Description

A class for extracting sentences from a text. It allows using different tokenizers and setting the minimum and maximum length of extracted sentences.

Language hooks

The default tokenizer is the method sentenize(text), which a language library overrides in a subclass. Here it splits the text at whitespace after ., !, ? or …, optionally followed by a closing quote or bracket; it knows no abbreviations, so e.g. this is two sentences. A piece without words, such as the dots of a spaced ellipsis . . . or a lone !, stays with the sentence before it, or with the one after it at the start of the text.

Note

A spaCy pipeline can be passed as the tokenizer: tokenizer=lambda text: (sent.text for sent in nlp(text).sents).

Parameters

Parameter Type Default Description
tokenizer Pattern/Callable None Tokenizer or regular expression; by default the sentenize method
min_len int 0 Minimum length of an extracted sentence in characters, 0 for no bound
max_len int 0 Maximum length of an extracted sentence in characters, 0 for no bound

Note

A regular expression as the tokenizer is a separator: the text is split with re.split. The sentences of any tokenizer are stripped of whitespace at the edges, before the length bounds, and empty ones are dropped.

Methods

extract

Extracts sentences from a text.

Parameter Type Default Description
text str - Text string

An example of sentence extraction with the default tokenizer:

Example

Code:

from anyts import SentsExtractor

SentsExtractor().extract('It rains. Does it? He said "Yes!" Then he left.')

Result:

('It rains.', 'Does it?', 'He said "Yes!"', 'Then he left.')

An example of sentence extraction with a regular expression as the tokenizer:

Example

Code:

import re

from anyts import SentsExtractor

SentsExtractor(tokenizer=re.compile(r", ")).extract("Rain today, sun tomorrow")

Result:

('Rain today', 'sun tomorrow')