Tagger models
Models
All models were trained with the ISO 8859-1 encoding (latin1). They were also trained with a sentence delimiter in the corpus (usually empty line, but it depends on the tagger), as this gives better results than without a delimiter.
Evaluation results for the evaluation models are reported in the format described in the result legend.
The accuracy and proportion results are computed with tnt-diff on 10 folds of SUC, using a TnT lexical model trained on the same material as the evaluated tagger to get information on which words are known or unknown, while the standard deviation, min and max results are computed with the Perl module Statistics::Descriptive.
The models here, however, are the final tag models, which also includes all of SUC.
