Sentence extraction¶
anyts.extractors.SentsExtractor
Description¶
A class for extracting sentences from a text. It allows using different tokenizers and setting the minimum and maximum length of extracted sentences.
Language hooks¶
The default tokenizer is the method sentenize(text), which a language library overrides in a subclass. Here it splits the text at whitespace after ., !, ? or …, optionally followed by a closing quote or bracket; it knows no abbreviations, so e.g. this is two sentences. A piece without words, such as the dots of a spaced ellipsis . . . or a lone !, stays with the sentence before it, or with the one after it at the start of the text.
Note
A spaCy pipeline can be passed as the tokenizer: tokenizer=lambda text: (sent.text for sent in nlp(text).sents).
Parameters¶
| Parameter | Type | Default | Description |
|---|---|---|---|
tokenizer |
Pattern/Callable | None |
Tokenizer or regular expression; by default the sentenize method |
min_len |
int | 0 |
Minimum length of an extracted sentence in characters, 0 for no bound |
max_len |
int | 0 |
Maximum length of an extracted sentence in characters, 0 for no bound |
Note
A regular expression as the tokenizer is a separator: the text is split with re.split. The sentences of any tokenizer are stripped of whitespace at the edges, before the length bounds, and empty ones are dropped.
Methods¶
extract¶
Extracts sentences from a text.
| Parameter | Type | Default | Description |
|---|---|---|---|
text |
str | - |
Text string |
An example of sentence extraction with the default tokenizer:
Example
Code:
from anyts import SentsExtractor
SentsExtractor().extract('It rains. Does it? He said "Yes!" Then he left.')
Result:
('It rains.', 'Does it?', 'He said "Yes!"', 'Then he left.')
An example of sentence extraction with a regular expression as the tokenizer:
Example
Code:
import re
from anyts import SentsExtractor
SentsExtractor(tokenizer=re.compile(r", ")).extract("Rain today, sun tomorrow")
Result:
('Rain today', 'sun tomorrow')