论文标题
猫是模糊的宠物:一种语料库和对潜在委婉语的分析
CATs are Fuzzy PETs: A Corpus and Analysis of Potentially Euphemistic Terms
论文作者
论文摘要
尽管是礼貌和象征性语言的重要因素,但委婉语在自然语言处理中并没有得到太多关注。委婉语被证明是一个困难的话题,不仅是因为它们会遭受语言的改变,还因为人类可能不同意什么是委婉语,什么不是同意。然而,解决该问题的第一步是收集和分析委婉语的例子。我们提出了潜在的委婉术语(PET)的语料库以及Glowbe语料库的示例文本。此外,我们提出了文本的子库,其中这些宠物没有被委婉地使用,这可能对将来的应用有用。我们还讨论了在语料库上进行的多个分析的结果。首先,我们发现对委婉文本的情感分析支持宠物通常会减少负面和进攻性情绪。其次,我们在注释任务中观察到分歧的案例,在我们的语料库文本示例的一个子集中,要求人类将宠物标记为委婉语。我们将分歧归因于各种潜在原因,包括宠物是普遍接受的术语(CAT)。
Euphemisms have not received much attention in natural language processing, despite being an important element of polite and figurative language. Euphemisms prove to be a difficult topic, not only because they are subject to language change, but also because humans may not agree on what is a euphemism and what is not. Nevertheless, the first step to tackling the issue is to collect and analyze examples of euphemisms. We present a corpus of potentially euphemistic terms (PETs) along with example texts from the GloWbE corpus. Additionally, we present a subcorpus of texts where these PETs are not being used euphemistically, which may be useful for future applications. We also discuss the results of multiple analyses run on the corpus. Firstly, we find that sentiment analysis on the euphemistic texts supports that PETs generally decrease negative and offensive sentiment. Secondly, we observe cases of disagreement in an annotation task, where humans are asked to label PETs as euphemistic or not in a subset of our corpus text examples. We attribute the disagreement to a variety of potential reasons, including if the PET was a commonly accepted term (CAT).