Evaluating Distributional Properties of Tagsets

被引:0
|
作者
Dickinson, Markus [1 ]
Jochim, Charles [2 ]
机构
[1] Indiana Univ, Bloomington, IN 47405 USA
[2] Univ Stuttgart, Stuttgart, Germany
关键词
FREQUENT FRAMES; CATEGORIES; CUE;
D O I
暂无
中图分类号
H [语言、文字];
学科分类号
05 ;
摘要
We investigate which distributional properties should be present in a tagset by examining different mappings of various current part-of-speech tagsets, looking at English, German, and Italian corpora. Given the importance of distributional information, we present a simple model for evaluating how a tagset mapping captures distribution, specifically by utilizing a notion of frames to capture the local context. In addition to an accuracy metric capturing the internal quality of a tagset, we introduce a way to evaluate the external quality of tagset mappings so that we can ensure that the mapping retains linguistically important information from the original tagset. Although most of the mappings we evaluate are motivated by linguistic concerns, we also explore an automatic, bottom-up way to define mappings, to illustrate that better distributional mappings are possible. Comparing our initial evaluations to POS tagging results, we find that more distributional tagsets can sometimes result in worse accuracy, underscring the need to carefully define the properties of a tagset.
引用
收藏
页码:2522 / 2529
页数:8
相关论文
共 50 条