NivaarExam PrepOfficial exam papers ↗

25-Comp-A6 Software Engineering · May 2013

Question 5 of 9: Software Testing

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

National Exams — May 2013 — 98-Comp-A6 Software Engineering. Three-hour, closed-book exam, no calculator permitted. Format: nine questions, candidates answer any five of the nine (all questions equal weight — each of the five counted questions is worth 20%; only the first five questions as they appear in the answer book are marked). All nine questions are solved below for completeness.

Reference texts: Sommerville, Software Engineering (10th ed., Pearson) — software process models, object-oriented and function-oriented design, software testing, dependability and critical systems, reliability metrics, configuration management, real-time software engineering; Pressman, Software Engineering: A Practitioner's Approach (9th ed.) — supplementary process and testing coverage; Gamma, Helm, Johnson & Vlissides, Design Patterns — object-oriented design vocabulary.

Question 5: Software Testing (20 marks)

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

(a) Testing Detects Presence, Not Absence, of Errors

A test run exercises the program on one specific input (or a small finite set of inputs) and observes whether the output matches what was expected; if it does not, an error has been demonstrably found. But the space of possible inputs to any non-trivial program is astronomically large or infinite, and exhaustively trying every one is not feasible in any practical amount of time. A test suite can therefore only ever sample that space. Passing every test in the suite proves that the program behaves correctly on the particular inputs tried — it says nothing about the vastly larger set of inputs that were never tried, any one of which could still trigger a latent defect. This is Dijkstra's well-known observation: testing can conclusively show a bug exists (by triggering it) but can never conclusively show that no bug exists, because "no bug found yet" and "no bug present" are not the same statement given an untested remainder of the input space.

(b) Fitness for Purpose Without Zero Defects

Removing every last defect from a non-trivial system has a cost that rises steeply as the remaining defect density falls — the easy, frequently-triggered defects are found early and cheaply, while the last few require disproportionate testing and debugging effort to locate, because by definition they occur only under rare or obscure conditions. Beyond some point, the cost of continuing to search for ever-rarer defects exceeds the cost their occasional occurrence would actually impose on the customer, especially when the consequence of a rare failure is minor (a cosmetic glitch, an inconvenience that a restart fixes) rather than safety- or business-critical. It is therefore economically rational, for most systems, to ship once the program is "good enough" — reliable enough, for the intended operating profile, that the expected cost of remaining defects is lower than the cost of continuing to test for them. Testing supports this release decision as a validation activity rather than a proof of correctness: by exercising the system against realistic usage scenarios and its stated requirements, testing builds evidence that the delivered system does what the customer actually needs under the conditions it will actually encounter, which is a different and more achievable question than "does this program contain zero defects of any kind." The acceptable residual defect rate is not fixed, however — a word processor can ship with cosmetic defects a safety-critical system could never tolerate, so how far testing must be pushed to establish fitness for purpose depends directly on the criticality of the system.

(c) Functional vs. Structural Testing

Functional (black-box) testing derives test cases purely from the requirements or specification, without reference to how the code is written internally; techniques such as equivalence partitioning and boundary-value analysis are used to choose a manageable set of representative inputs that should, according to the spec, produce known outputs. Structural (white-box or glass-box) testing instead derives test cases from the code's internal structure — its statements, branches, and paths — with the specific goal of exercising code that has not yet been executed by any existing test, measured via coverage criteria such as statement or branch coverage.

The two are complementary rather than substitutable. Functional testing alone can achieve full requirements coverage yet still leave large parts of the actual code never executed by any test, silently hiding defects in that unexercised code (including code that should not even exist, such as an unremoved debug branch). Structural testing alone can achieve 100% code coverage yet still miss a defect of omission — functionality the specification requires that was simply never implemented, since there is no code path to cover for something that isn't there. The standard defect-testing process therefore starts from functional test cases derived directly from the requirements to validate behaviour against what the customer asked for, then uses a coverage tool to measure how much of the actual code those tests exercised, and adds targeted structural tests to reach the coverage target the project has set for cases the functional tests happened not to reach. Used this way, functional testing anchors correctness to the specification while structural testing closes the gap between "the spec is satisfied" and "the code that implements it has actually all been run."