Language Testing · 2011 · 103 citations · 24 references
Generalizability TheoryWriting AssessmentEducationRating QualityLanguage ProficiencyJournalismLanguage Assessment (Second Language Acquisition)Foreign Language WritingLanguage TestingPerformance AssessmentLanguage Assessment (Speech Language Pathology)Quality ReviewLanguage StudiesReliabilityWriting InstructionNovice RatersEducational TestingEducational MeasurementExperienced RatersPerformance StudiesEvaluation MeasurePerformance MeasureRating PerformanceEducational AssessmentSurvey Methodology
Raters are central to writing performance assessment, and rater development – training, experience, and expertise – involves a temporal dimension. However, few studies have examined new and experienced raters’ rating performance longitudinally over multiple time points. This study uses operational data from the writing section of the MELAB (n = 20,662 ratings), an international exam of English proficiency, to investigate the rating quality of new and experienced raters over three time periods of 12 to 21 months. Rating quality was operationalized in terms of rater severity and consistency, and estimates of those modeled using multi-facet Rasch methodology. Results indicate that, within one particular rating context, (1) novice raters, where initially differing in performance, learn to rate appropriately relatively quickly, (2) raters are able to maintain rating quality over time, and (3) rating volume and rating quality may be related. Implications for rater preparation, rater certification, and the notion of expert rater are discussed.
24
Reasonable mean-square fit values
Benjamin G. Wright · Medical Entomology and Zoology · 1994 · 1.6K citations
Lloyd G. Humphreys, E. F. Lindquist · The American Journal of Psychology · 1952 · 1.5K citations