Тип публикации: доклад, тезисы доклада, статья из сборника материалов конференций
Конференция: International Workshop “Hybrid methods of modeling and optimization in complex systems” (HMMOCS 2022); Krasnoyarsk; Krasnoyarsk
Год издания: 2023
Идентификатор DOI: 10.15405/epct.23021.3
Ключевые слова: Text message clustering, semantic proximity, machine learning
Аннотация: At the present moment the relevance of natural data processing problem solving is rising. A massive data amount of text data has been accumulated in recent years. Classical analytical methods, such as machine learning methods, are not capable of dealing with raw text data, which complicates the analysis significantly. Therefore, a Показать полностьюmodern set of methods of text data vectorization has been developed, which gained massive popularity in the recent years for analyzing text data, specifically for solving text clustering problem, as one of the most relevant text data related analytical problems. In this paper, a few of these methods were researched; a new dictionary optimization approach has been proposed and tested on the real text datasets; a number of conclusions on the effectiveness and of the methods for the given tasks has been made. For the future work a more thorough research on the dictionary optimization scheme (genetic algorithm parameters) and vectorization method are planned. ]
Журнал: Hybrid methods of modeling and optimization in complex systems
Номера страниц: 19-31
Место издания: London, United Kingdom