22-Mec-B5 Product Design and Development · May 2016
Question 6 of 7: Enhancing Reliability and Robustness
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
Notes on this paper
National Exams, May 2016 — 07-Mec-B5 Product Design and Development. Three hours. Open book; no calculator is permitted. Question 1 must be completed and is worth 40 marks; four of the six remaining questions are chosen, each worth 15 marks, for 100 marks in total. Only the first five questions as they appear in the answer book are marked. The paper states that most questions require an answer in essay format or the use of tables, figures and charts, and that clarity and organisation of the answer are important.
The paper prints 40 + 6 × 15 = 130 marks and a candidate attempts 40 + 4 × 15 = 100 of them. All seven questions are answered below, because this set is a study resource rather than an examination script. The marking scheme printed on the last source page splits Question 1 as 9 / 9 / 4 / 9 / 9 and gives the part weights for every 15-mark question; the answers here are proportioned to that split. Because no calculator is permitted, every calculation is arranged so that it can be carried out on paper in one or two lines — ratios of round numbers, never a logarithm that has to be evaluated.
Reference texts for this subject
K. T. Ulrich and S. D. Eppinger, Product Design and Development — the framework text for this exam code: identifying customer needs, target and final specifications, concept generation and selection, design for manufacturing, and product development economics.
G. E. Dieter and L. C. Schmidt, Engineering Design — problem definition, decision matrices, design for the environment, standards and codes, and reliability and robustness in design.
G. Pahl and W. Beitz, Engineering Design: A Systematic Approach — the systematic conceptual / embodiment / detail sequence and function structures.
G. Boothroyd, P. Dewhurst and W. Knight, Product Design for Manufacture and Assembly — part-count reduction, handling and insertion time, and the manual-versus-automated assembly comparison used in Question 2.
M. F. Ashby, Materials Selection in Mechanical Design — translate, screen, rank and document; the material indices used in Question 7.
S. Kalpakjian and S. R. Schmid, Manufacturing Engineering and Technology — the process routes and their tooling economics.
Centre for Universal Design (North Carolina State University), The Principles of Universal Design, and CSA B651 Accessible design for the built environment — the accessibility framework and operating-force limits used in Question 1.
M. J. Kirwan (ed.) is not required here; for Question 4 the Canadian statutory frame is the Patent Act, Industrial Design Act, Trade-marks Act and Copyright Act as administered by the Canadian Intellectual Property Office (CIPO).
Question 6: Enhancing Reliability and Robustness (15 marks)
Reliability is the probability that the product performs its function for a stated time under stated conditions; robustness is insensitivity of that performance to variation. They are related but not identical, and the process below improves both because it treats variation, rather than average performance, as the design variable. It is a sequence, and the ordering is the substance of the answer.
Define the mission profile and the reliability target quantitatively. Nothing downstream is meaningful until someone states what “reliable enough” means: the duty cycle, the environment, the required life, and a target expressed as a reliability at a time with a confidence — for example, 95 per cent surviving three years at a stated duty and confidence of 90 per cent — or as a B10 life, or as failures per thousand units in the first year of service. The target is then allocated down the architecture to subsystems, so each design group owns a number.
Decompose the function and identify how it can fail. Build the function structure, then work it with a design FMEA: for each function, the failure modes, their effects, their causes, the existing controls, and a risk priority from severity, occurrence and detection. Fault-tree analysis complements it top-down for the high-severity effects. The output is a ranked list of what actually threatens the target, which is what stops the effort being spent on the parts that are easiest to analyse rather than the parts that fail.
Identify the noise factors. This is the step that distinguishes robustness work from reliability accounting. Taguchi’s P-diagram sorts the inputs into the signal, the control factors the designer may set, and the noises: piece-to-piece variation, change over time through wear and drift, customer usage and duty, the external environment, and interactions with neighbouring subsystems. Naming the noises explicitly is what makes them testable, and most field failures trace to a noise nobody wrote down.
Do robust parameter design before tolerance design. Use designed experiments — an orthogonal array with the control factors in the inner array and the noises deliberately applied in the outer array — to find the combination of control-factor settings at which the response is least sensitive to the noises, maximising the signal-to-noise ratio, and only then adjust the mean onto target with a factor that shifts it without affecting variance. This exploits nonlinearity in the transfer function and it is free: it changes nominal settings, not tolerances. The order matters enormously, because the alternative — reaching for tighter tolerances first — buys the same variance reduction with permanently higher manufacturing cost.
Then do tolerance design and derating. Allocate tolerances by worst-case or root-sum-square stack-up against the functional requirement, and check each against the process capability that will actually be available, $C_p = T/6\sigma$, so that the tolerance is one the plant can hold rather than one the drawing asserts. In parallel, derate stressed components against their ratings, apply safety factors where the load or strength distribution is uncertain, and design out single points of failure or add redundancy where the severity justifies it.
Grow the reliability by test-analyse-and-fix. Build hardware, stress it until it breaks, find the root cause, fix it, and repeat. Highly accelerated life testing steps temperature, vibration and combined stress well past specification to find the operating and destruct margins, which tells you not just that the design passes but by how much. Each failure is driven to root cause and closed, and the failure modes found are fed back into the FMEA. Reliability growth is tracked against a Duane or Crow-AMSAA model so that the programme can tell whether it is on course.
Close the loop into production and the field. Robustness achieved in design is lost if the process drifts, so the critical characteristics identified above are put under statistical process control with capability requirements, and field returns are analysed and fed back into the next design and into the FMEA library.
Part B — How this is validated (3 marks)
Validation is a planned programme, not a final test, and it is written down in a design verification plan and report that ties every requirement to the test that demonstrates it. Its elements are accelerated life testing to demonstrate the life requirement in an achievable calendar time; HALT to establish margin and HASS on the production line to screen for process escapes; a reliability demonstration test that establishes the target statistically; environmental and abuse testing to the applicable standards; and finally pilot builds and a field trial, which is the only test conducted with real users and real installation quality. Two calculations do most of the work, and both are arranged below so they can be checked on paper.
Given. A connected thermostat is to demonstrate a reliability of 0.95 at three years of continuous service, with 90 per cent confidence, using a zero-failure test. Its dominant wear-out mechanism is thermally activated with an activation energy of 0.70 eV; the use temperature is 40 °C and the chamber is set at 85 °C. Boltzmann’s constant is 8.617 × 10−5 eV/K.
Find. The sample size for the zero-failure demonstration, and the chamber hours that represent three years of service.
Size the zero-failure demonstration test. If $n$ units are tested with no failures, the lower confidence bound on reliability follows from the binomial: the probability of seeing zero failures when the true reliability is $R$ is $R^n$, so demanding $R^n \le 1-C$ and solving gives
$$n = \frac{\ln(1-C)}{\ln R} = \frac{\ln 0.10}{\ln 0.95} = \frac{-2.3026}{-0.05129} = 44.9 \rightarrow \boxed{45\ \text{units}}$$
tested for the full equivalent life with no failures. The shape of this result is the lesson: demonstrating high reliability with zero failures is expensive in samples, and demonstrating 0.99 at the same confidence would take 230 units — which is why accelerated and degradation methods exist.
Compute the acceleration factor and convert calendar life to chamber hours. For a thermally activated mechanism the Arrhenius model gives the ratio of life at use temperature to life at stress temperature as
$$AF = \exp\!\left[\frac{E_a}{k}\left(\frac{1}{T_u}-\frac{1}{T_s}\right)\right] = \exp\!\left[\frac{0.70}{8.617\times10^{-5}}\left(\frac{1}{313}-\frac{1}{358}\right)\right] = e^{3.262} = 26.1$$
so one chamber hour at 85 °C stands for 26 hours of service at 40 °C. Three years of continuous operation is $3 \times 8766 = 26{,}298$ hours, hence
$$t_{stress} = \frac{26{,}298}{26.1} = \boxed{1{,}007\ \text{hours}}$$
— about six weeks in the chamber, which is a schedule a programme can actually accommodate. The model must be justified, not assumed: the activation energy has to correspond to the mechanism that actually limits life, and the stress must not be so high that it activates a mechanism which never occurs in service, or the test demonstrates the wrong thing.
Result
Value
Zero-failure sample size for R = 0.95 at 90 per cent confidence
45 units
Same test for R = 0.99 at 90 per cent confidence
230 units
Arrhenius acceleration factor, 85 °C stress versus 40 °C use, Ea = 0.70 eV
26.1
Service life represented by 1,000 chamber hours
26,100 h ≈ 2.98 years
Chamber hours to represent a 3-year life
1,007 h ≈ 6 weeks
Part C — How success is assessed (3 marks)
Success is assessed against the target set in step 1, using field evidence rather than laboratory evidence, because the laboratory tested the noises that were anticipated and the field tests the ones that were not. The measures fall into three groups.
Field reliability measures. Failures per thousand units at twelve months in service is the standard consumer-durable metric and is compared directly with the allocated target; warranty claim rate and warranty cost per unit convert the same information into money; B10 life and mean time between failures are used where the product is repairable. Because these accrue over time, an early read is taken from units shipped in the first months and projected.
Distributional analysis, which is where the diagnosis lives. Field returns are fitted to a Weibull distribution and the shape parameter is read: a shape below one means the hazard rate is falling, so the failures are infant mortality and the cause is a process or screening problem; a shape near one means random failures and points at overstress or a design margin problem; a shape above one means wear-out, and if it appears inside the design life the cause is a design or material problem. That single parameter directs the corrective action to the right department, and it is the most useful thing on the list.
Process and programme measures. Capability indices on the critical characteristics, defective parts per million at final test, first-pass yield, and the no-fault-found rate on returns, which is a measure of robustness in the usability sense — a high no-fault-found rate means the product is failing to communicate rather than failing to work. Alongside these sit the leading indicators from the programme itself: the proportion of FMEA high-risk items closed, the HALT margins achieved against the specification, and whether the reliability growth curve reached the target before launch.
The assessment is only meaningful if it is compared against something — the allocated target, the previous generation, and the competition — and if the loop is closed, so that what is learned is written back into the design rules and the FMEA library for the next product.