{{Short description|Database of Russian texts}} The '''Russian National Corpus''' ({{langx|ru|Национальный корпус русского языка||National Corpus of the Russian Language}}) is a corpus of the Russian language that has been partially accessible through a query interface online since April 29, 2004. It is being created by the Institute of Russian language, Russian Academy of Sciences.

It currently contains more than 1 billion word forms<ref>{{cite web |url=http://ruscorpora.ru/ |title=Национальный корпус русского языка |language=Russian |website=Национальный корпус русского языка |access-date=August 28, 2022 |archive-url=https://web.archive.org/web/20220305105245/https://ruscorpora.ru/new/ |archive-date=March 5, 2022}}</ref> that are automatically lemmatized and POS-/grammeme-tagged, i.e. all the possible morphological analyses for each orthographic form are ascribed to it. Lemmata, POS, grammatical items, and their combinations are searchable. Additionally, 6 million word forms are in the subcorpus with manually resolved homonymy. <!---A statistical program that resolves homonymy using the manually tagged corpus as a machine-learning tool is already constructed and the results will be made available online soon.--->

The subcorpus with resolved morphological homonymy is also automatically accentuated. The whole corpus has a searchable tagging concerning lexical semantics (LS),<ref>{{cite conference | last1 = Apresjan | first1 = Ju. | last2 = Boguslavsky | first2 = I. | last3 = Iomdin | first3 = B. | last4 = Iomdin | first4 = L. | last5 = Sannikov | first5 = A. | last6 = Sizov | first6 = V. | year = 2006 | citeseerx = 10.1.1.111.8165 | title = A Syntactically and Semantically Tagged Corpus of Russian: State of the Art and Prospects | conference = Proceedings of LREC | location = Genova, Italy | pages = 1378–1381 }}</ref> including morphosemantic POS subclasses (proper noun, reflexive pronoun etc.), LS characteristics proper (thematic class, causativity, evaluation), derivation (diminutive, adverb formed from adjective etc.).

The RNC includes also the following subcorpora: *a treebank of syntactical dependencies (largely based on the Igor Mel'čuk's Meaning-Text Theory) *English⇔Russian, German⇒Russian, Ukrainian⇔Russian and Belorussian⇔Russian parallel corpora; *a large (100+ million words) separate corpus of modern newspapers (2001–2011); *a corpus of Russian poetry, where the rhyming words and poetic prosody (including meter, stanzas etc.) is additionally tagged; *a corpus of Russian dialects with specific dialect grammar tagging; *a multimedia corpus with searchable tagged fragments of Russian-language movies; *a corpus showing the history of Russian stress *an educational subcorpus reflecting school standards.

All the texts have tags bearing metatextual information - the author, his/her birth date, creation date, text size, text genres (general fiction, detective story, newspaper article etc.); all these categories are browsable and searchable separately. It is possible to define a user's subcorpus to search lemmata/POS-grammeme/semantic tags combinations only within this subset.

==See also== * General Internet Corpus of Russian

==References== {{reflist}}

==External links== * {{official website | 1 = ruscorpora.ru | 2 = Russian National corpus }}

{{Corpus linguistics}}

Category:Corpora Category:Russian language Category:Applied linguistics Category:Linguistic research

{{corpora-stub}} {{slavic-lang-stub}}