Skip to content

Components

anyts.components.StatsComponent

Description

The base of the components of spaCy that put the statistics of a text into an extension of Doc: a component registers the extension when it is added to a pipeline and, for every document, computes the statistics and puts them into doc._.<name>. Writing components in general is described in the documentation of spaCy.

Language hooks

A language library subclasses StatsComponent for every class of its statistics and registers the subclass as a factory under the prefix of the library - Language.factory("<prefix>_<statistics>") - so that the factories of several libraries live in one process. Declared as entry points of spacy_factories, the factories load with spacy.load without importing the library. The subclass checks its parameters in its __init__ before it calls the __init__ of the base, so a wrong parameter fails at add_pipe.

Hook Description
compute(doc) The statistics of a document; an abstract method, so a subclass without it fails at add_pipe
prepare(nlp) Preparing the pipeline when the component is added, such as the rules of the tokenizer; nothing by default
accepts(doc) Whether a document gets the statistics; by default whether it has a word (has_words)
from_extension(doc, name, stats_class, factory=None) The statistics another component put into the document, for reusing them instead of computing them again; an extension without statistics of the class raises SourceError, which names the factory to add. The other component must accept every document the reusing one does, since a document it passes leaves None

Names

The name of the pipe is the name of the extension: nlp.add_pipe(factory, name="basic") puts the statistics into doc._.basic. Without name the pipe and the extension keep the name of the factory. The extension follows the pipe when the pipeline renames it (nlp.rename_pipe), and a name taken by an extension of another package raises ParameterError instead of replacing it. A component taken from another pipeline (add_pipe(name, source=other)) is the same object in both, so under a new name it writes to that extension in the other pipeline too; a component of its own comes from the factory, add_pipe(factory, name=...). The same component can be added twice under different names, with different parameters. A document with no words - an empty string, whitespace, punctuation alone - passes through a component untouched, its extension left at None.

Serialization

A component keeps an object of statistics in doc._.<name>, and spaCy cannot serialize it: Doc.to_bytes(), DocBin(store_user_data=True) and nlp.pipe(..., n_process>1) fail with these components in the pipeline. To save a document, leave the user data out (doc.to_bytes(exclude=["user_data"])) or keep doc._.<name>.get_stats() on your own; for multiprocessing compute the statistics in the main process after nlp.pipe, without the components.

Parameters

Parameter Type Default Description
nlp Language - Pipeline the component is added to
name str - Name of the component in the pipeline and of the extension

Usage example

A component of the lexical diversity of the words of a Doc.

Example

import spacy
from spacy.language import Language

from anyts import DiversityStats
from anyts.components import StatsComponent
from anyts.diversity_stats import check_params
from anyts.utils import iter_doc_words


@Language.factory("demo_diversity")
class DiversityComponent(StatsComponent):
    def __init__(self, nlp, name="demo_diversity", window_len=50):
        check_params(window_len=window_len)
        self.window_len = window_len
        super().__init__(nlp, name)

    def compute(self, doc):
        words = [word.lower() for _, _, word in iter_doc_words(doc)]
        return DiversityStats(words, window_len=self.window_len)


nlp = spacy.blank("xx")
nlp.add_pipe("demo_diversity", name="diversity", config={"window_len": 3})
nlp("The cat and the dog")._.diversity.ttr
# 0.8
nlp("?!")._.diversity
# None