The Effect of Test Length and IRT Model on the Distribution and Stability of Three Appropriateness Indexes

Brian Noonan, Marvin W. Boss, Marc E. Gessaroli

Applied Psychological Measurement · 1992 · 18 citations · 13 references

DOIFull text

Open access

Abstract

The extent to which three appropriateness
\nindexes-Z, ECIZ₄, and W (a variation of Wright’s
\nperson-fit statistic)-are well-standardized was
\ninvestigated in a monte carlo study. To assess the
\neffects of the item response theory (IRT.) model and
\ntest length on the distribution of the indexes and
\ntheir cutoff values at three false positive rates,
\nnonaberrant response patterns were generated.
\nECIZ₄ most closely approximated a normal
\ndistribution, showing less skewness and kurtosis
\nthan Z, and W. The ECIZ₄ cutoff values were
\naffected less by test length and the IRT model than
\nwere Z, and W. In contrast, the distribution of W
\nwas the least stable over replications, and its cutoff
\nvalues varied greatly depending on the IRT model
\nand test length. Index terms: appropriateness
\nmeasurement, caution index, item response theory
\n(person fit), person-fit statistics, unusual response
\npatterns.

References

13