Linguistic steganography with knowledge-poor paraphrase generation.

Paraphrasing is very useful for many applications that normally involve deep linguistic alterations of a sentence, like summarization, textual entailment and question answering, and usually require sophisticated external resources, pre-processing, and semantic thesauri. This article presents a metho...

Descripción completa

Detalles Bibliográficos
Publicado en:Literary & Linguistic Computing Vol. 26; no. 4; pp. 417 - 435
Autor principal: Kermanidis, Katia Lida
Formato: Artículo
Publicado: Oxford University Press / USA Dec2011
Materias:
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:Paraphrasing is very useful for many applications that normally involve deep linguistic alterations of a sentence, like summarization, textual entailment and question answering, and usually require sophisticated external resources, pre-processing, and semantic thesauri. This article presents a methodology for generating shallow linguistic alterations of Modern Greek sentences making use of only a low-resource chunker and taking advantage of the freedom in phrase ordering of the language. A statistical significance testing process is applied for extracting ‘swappable’ phrase bigrams. A supervised filtering phase follows, which helps to remove erroneous paraphrasing schemata, taking into account the context in which the alteration is to take place. Unlike most previous approaches to paraphrasing, the proposed process is knowledge-poor (and thus quite easily portable to other languages with a syntactic structure similar to Modern Greek), robust (applicable to any type of text), domain independent, and leads to the generation of a significant number of paraphrases, by allowing the application of more than one syntactic alterations per sentence. The significance of the automatically generated paraphrases is shown by their application in hiding secret information underneath a cover text in a steganographic communication channel. For this purpose, they need not be sophisticated linguistic alterations, but grammatically correct and significant in number, to ensure security. A steganographic security and capacity analysis of the presented implementation, as well as an explanatory description of the trade-off between them, is included to show the usefulness and the practical value of the methodology.