МӘТІНДІ АВТОМАТТЫ ЖЕҢІЛДЕТУ: ЗЕРТТЕУЛЕР МЕН БАҒЫТТАР

The article provides a comprehensive survey on automatic text simplification (ATS) as an independent research area in the field of natural language processing. The study aims to systematically and critically present the development, current state, and key challenges of ATS. Based on an extensive lit...

Descripción completa

Detalles Bibliográficos
Publicado en:Eurasian Journal of Philology: Science & Education Vol. 202; no. 2; pp. 79 - 94
Autores principales: Карымхан, А. А., Мәмбетова, М. Қ., Нұрланғазықызы, Б.
Formato: Artículo
Publicado: Al-Farabi Kazakh National University 2026
Materias:
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:The article provides a comprehensive survey on automatic text simplification (ATS) as an independent research area in the field of natural language processing. The study aims to systematically and critically present the development, current state, and key challenges of ATS. Based on an extensive literature review in Scopus, Google Scholar and the ACL Anthology, research from 1998 to 2025 is analyzed. The article examines the evolution of ATS development from rule-based approaches to the use of large language models using statistical and neural models. It is shown that this process goes hand in hand with the gradual expansion of high-quality parallel corpora in several languages. A particular focus is placed on the analysis of lexical simplification, including the (1) identification of complex words, as well as (2) selection, (3) generation and (4) ranking of substitutions. The study shows that isolated rule-based, frequency-based, or purely data-driven approaches often reach their limits and that hybrid, linguistically grounded solutions deliver the best results. Key challenges remain the preservation of meaning and coherence, the strong dominance of English in research, and the lack of resources for typologically complex languages like Kazakh. The article notes that purely neural approaches to such languages are not enough. Instead, a stepby-step approach is proposed, based on linguistically sound reliable modeling, as well as complemented by automated and neural methods. The survey highlights the importance of improving evaluation procedures for the further development of automatic text simplification, as well as typologically oriented linguistically sound research.
В статье дается всесторонний обзор автоматического упрощения текста (ATS) как самостоя- тельного направления исследований в области обработки естественного языка. Цель исследова- ния - систематически и критически представить развитие, текущее состояние и ключевые про- блемы ATS. На основе обширного обзора литературы в Scopus, Google Scholar и ACL Anthology анализируются исследования с 1998 по 2025 год. В статье рассматривается эволюция развития ATS от подходов, основанных на правилах, до использования более крупных языковых моделей с помощью статистических и нейронных моделей. Показано, что этот процесс идет рука об руку с постепенным расширением высококачественных параллельных корпусов на нескольких языках. Особое внимание уделяется анализу процесса лексического упрощения, включая определе- ние (1) сложных слов, а также (2) отбор, (3) генерацию и (4) ранжирование замен. Исследование показывает, что изолированные подходы, основанные на правилах, частоте или исключительно на данных, часто достигают своих пределов, и что гибридные, лингвистически обоснованные решения дают наилучшие результаты. Ключевыми проблемами остаются сохранение смысла и связности, сильное доминирование английского языка в исследованиях и нехватка ресурсов для типологически сложных языков, таких как казахский. В статье отмечается, что чисто нейронных подходов к таким языкам недостаточно. Вместо этого предлагается поэтапный подход, основанный на лингвистически обоснованном надежном моделировании, а также дополненный автоматизированными и нейронными методами. Обзор подчеркивает важность совершенствования процедур оценки для дальнейшего развития авто- матического упрощения текста, а также типологически ориентированных лингвистически обо- снованных исследований.