An upper bound on software testing effectiveness

Tsong Yueh Chen, Robert Merkel

ACM Transactions on Software Engineering and Methodology · 2008 · 98 citations · 16 references

Concepts

TL;DR

Failure patterns are typically clustered in contiguous input regions, prompting the development of several debug testing methods. The study investigates the theoretical upper bound of debug testing effectiveness achievable by assuming specific shapes, sizes, and orientations of failure patterns. The authors analyze bounds for testing strategies that minimize the F‑measure, maximize the P‑measure, and maximize the E‑measure, and examine the assumptions and implications underlying the upper bound. Existing methods not based on these assumptions achieve empirically measured effectiveness close to the theoretical upper bound.

Abstract

Failure patterns describe typical ways in which inputs revealing program failure are distributed across the input domain—in many cases, clustered together in contiguous regions. Based on these observations several debug testing methods have been developed. We examine the upper bound of debug testing effectiveness improvements possible through making assumptions about the shape, size and orientation of failure patterns. We consider the bounds for testing strategies with respect to minimizing the F-measure, maximizing the P-measure, and maximizing the E-measure. Surprisingly, we find that the empirically measured effectiveness of some existing methods that are not based on these assumptions is close to the theoretical upper bound of these strategies. The assumptions made to obtain the upper bound, and its further implications, are also examined.

References

16