Skip to content

Russian Texts Statistics (ruTS)

ruts

ruTS computes for Russian texts what usually requires assembling several separate tools: from basic statistics, readability and lexical diversity to morphology, syntax, cohesion, stylometry and corpus measures - by published formulas adapted to the Russian language.

The library works both with raw strings and with Doc objects of spaCy - every statistic is available both as a standalone class and as a spaCy pipeline component.

Try it without installing in the demo on Hugging Face Spaces: paste a text and get the readability grade, metrics, plots and highlighted fragments.

Features

Installation

Requires Python 3.11 or newer.

pip install ruts

More on dependencies and installing from the repository - on the Installation page.

Quick start

>>> from ruts import BasicStats, DiversityStats, ReadabilityStats

>>> text = "Существуют три вида лжи: ложь, наглая ложь и статистика"

>>> BasicStats(text).get_stats()
{'c_letters': {1: 1, 3: 2, 4: 3, 6: 1, 10: 2},
 'c_syllables': {1: 5, 2: 1, 3: 1, 4: 2},
 'n_sents': 1,
 'n_words': 9,
 'n_unique_words': 8,
 'n_long_words': 3,
 'n_complex_words': 2,
 'n_simple_words': 7,
 'n_monosyllable_words': 5,
 'n_polysyllable_words': 4,
 'n_chars': 55,
 'n_letters': 45,
 'n_spaces': 8,
 'n_syllables': 18,
 'n_punctuations': 2,
 'c_punctuations': {'comma': 1, 'period': 0, 'question': 0, 'exclamation': 0,
                    'ellipsis': 0, 'colon': 1, 'semicolon': 0, 'dash': 0,
                    'hyphen': 0, 'angle_quotes': 0, 'straight_quotes': 0,
                    'parentheses': 0, 'other': 0}}

>>> ReadabilityStats(text).flesch_reading_easy
74.93500000000003

>>> DiversityStats(text).ttr
0.8888888888888888

Any statistic can be printed in a readable form:

>>> BasicStats(text).print_stats()
     Статистика     | Значение
------------------------------
Предложения         |    1
Слова               |    9
Уникальные слова    |    8
Длинные слова       |    3
Сложные слова       |    2
Простые слова       |    7
Односложные слова   |    5
Многосложные слова  |    4
Символы             |    55
Буквы               |    45
Пробелы             |    8
Слоги               |    18
Знаки препинания    |    2

A walkthrough of one short story with every tool of the library is in the notebook examples/01_text_walkthrough.ipynb, which opens in Colab; the other notebooks are on the examples page.

Project structure
  • docs - project documentation
  • ruts:
    • basic_stats.py - basic text statistics
    • cohesion_stats.py - cohesion statistics
    • components.py - spaCy components
    • constants.py - main constants
    • diversity_stats.py - lexical diversity metrics
    • exceptions.py - library exceptions
    • extractors.py - tools for object extraction from a text
    • lexical_stats.py - lexical sophistication statistics
    • morph_stats.py - morphological statistics
    • phon_stats.py - phonostatistics
    • readability_stats.py - readability metrics
    • style_stats.py - SEO style metrics
    • syntax_stats.py - syntactic statistics
    • utils.py - helper tools
    • verse_stats.py - verse statistics: stresses, meter, rhyme, stanzas
    • corpus - corpus measures:
      • collocations.py - collocations and association measures
      • compare.py - corpus comparison by text features
      • dispersion.py - word dispersion across text parts
      • keyness.py - keywords relative to a reference corpus
      • kwic.py - KWIC concordance
      • stylometry.py - Burrows's Delta, Zeta and other stylometry measures
    • datasets - datasets:
      • dataset.py - base class for working with datasets
      • freq2011.py - the Lyashevskaya-Sharoff frequency dictionary
      • poetry_corpus.py - Ilya Gusev's PoetryCorpus
      • russian_literature.py - the RusLit collection of Russian classics
      • sov_chrest_lit.py - soviet reading-books for literature classes
      • stalin_works.py - the collected works of Stalin
      • stress_dict.py - the Koziev stress dictionary
      • texts_by_grade.py - texts with grade labels from the Plain Russian Language project
    • resources - embedded lexical resources (the most frequent lemmas list, the connectives dictionary)
    • visualizers - tools for text visualization:
      • corpus.py - lexical dispersion, keyness chart, collocation network
      • fingerprinting.py - Literature Fingerprinting
      • highlight.py - text highlighting
      • sentences.py - sentence lengths
      • stylometry.py - dendrogram, principal components, scaling, Mendenhall curves
      • vocabulary.py - Heaps's law and frequency spectrum
      • word_tree.py - Word Tree
      • zipf.py - Zipf's law
  • tests - tests mirroring the package structure
  • examples - example notebooks
  • demo - Gradio demo for Hugging Face Spaces