Similar Text Fragments Extraction for Identifying Common Wikipedia Communities

Similar text fragments extraction from weakly formalized data is the task of natural language processing and intelligent data analysis and is used for solving the problem of automatic identification of connected knowledge fields. In order to search such common communities in Wikipedia, we propose to use as an additional stage a logical-algebraic model for similar collocations extraction. With Stanford Part-Of-Speech tagger and Stanford Universal Dependencies parser, we identify the grammatical characteristics of collocation words. WithWordNet synsets, we choose their synonyms. Our dataset includes Wikipedia articles from different portals and projects. The experimental results show the frequencies of synonymous text fragments inWikipedia articles that form common information spaces. The number of highly frequented synonymous collocations can obtain an indication of key common up-to-date Wikipedia communities.

Ключові слова

information extraction, short text fragment similarity, Wikipedia communities, NLP

Бібліографічний опис

Similar Text Fragments Extraction for Identifying Common Wikipedia Communities / S. Petrasova [et al.] // Data. – 2018. – Vol. 3, iss. 4. – 9 p.

URI

https://repository.kpi.kharkov.ua/handle/KhPI-Press/46382

Колекції

Кафедра "Інтелектуальні комп'ютерні системи"

Повна інформація про документ
Google Scholar

Similar Text Fragments Extraction for Identifying Common Wikipedia Communities

Файли

Дата

Автори

ORCID

DOI

Науковий ступінь

Рівень дисертації

Шифр та назва спеціальності

Рада захисту

Установа захисту

Науковий керівник/консультант

Члени комітету

Назва журналу

Номер ISSN

Назва тому

Видавець

Анотація

Опис

Ключові слова

Бібліографічний опис

URI

Колекції

Підтвердження

Рецензія

Додано до

Згадується в