Question 2 of 9: Chi-Squared Goodness-of-Fit and Generator Diagnosis
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
Notes on this paper
National Exams — May 2018 — 17-Ind-A6 Systems Simulation. Three-hour, closed-book exam; one of two permitted calculators (Sharp or Casio), one 8.5″×11.0″ aid sheet (both sides). Format: three sections — Section A (four concept questions, candidates choose any two, 10 marks each, 20 marks total), Section B (three methods questions built around one continuing warehouse-simulation case study, candidates choose any two, 15 marks each, 30 marks total), Section C (two applications questions, candidates choose any one, 20 marks each). All nine questions are solved below for completeness. Two source anomalies are flagged where they occur: the front-page summary table states Section A is "Do 2 of 3," while Section A's own instructions and its four printed question sets read "two of the following four" — the printed four-question section is answered in full here; and Part C Question 2's sub-parts (a)–(c) are never printed anywhere in the paper, even though the results text for (d)–(f) explicitly refers back to "the factorial design matrix in (a)." Two-page Normal-distribution tables were supplied with the exam; the values below are the same table values obtained by direct computation.
Reference texts: Banks, Carson, Nelson & Nicol, Discrete-Event System Simulation (5th ed., Pearson) — random-number/random-variate generation, input modeling and goodness-of-fit testing, output analysis (warm-up, replication length, batch means vs. replication/deletion), and comparing alternative systems; Montgomery, Design and Analysis of Experiments (current ed., Wiley) — single-factor ANOVA, multiple comparisons, and 2k factorial designs with interaction analysis (Part C).
Given. A 10-bin histogram of "100" generated data points on $[0,5]$ (the printed bin counts actually total 101 — a one-count data-entry discrepancy noted below); target pdf $f(x)=2x/25$; MCG stated as $a=11,\ m=100,\ X_0=17$, described as "the same MCG from Q1" even though Q1 itself used $a=23$ — a genuine inconsistency in the source narrative, flagged rather than silently resolved.
Find. (a) a Chi-Squared goodness-of-fit test at $\alpha=0.05$; (b) the reason the data fails to fit $f(x)$, and how to correct it.
Approach. Convert $f(x)$ to per-bin probabilities via its CDF, form $\chi^2=\sum(O-E)^2/E$, compare to the critical value; then diagnose the generator itself via its period (multiplicative order of $a$ modulo $m$).
Expected bin probabilities and counts. For a bin $[a_i,b_i)$, $p_i=F(b_i)-F(a_i)=(b_i^2-a_i^2)/25$; with the stated sample size $N=100$, $E_i=100\,p_i$:
Chi-squared goodness-of-fit table
Bin
Observed $O_i$
Expected $E_i$
$(O_i-E_i)^2/E_i$
[0.0, 0.5)
21
1.00
400.00
[0.5, 1.0)
10
3.00
16.33
[1.0, 1.5)
10
5.00
5.00
[1.5, 2.0)
10
7.00
1.29
[2.0, 2.5)
10
9.00
0.11
[2.5, 3.0)
5
11.00
3.27
[3.0, 3.5)
10
13.00
0.69
[3.5, 4.0)
5
15.00
6.67
[4.0, 4.5)
10
17.00
2.88
[4.5, 5.0)
10
19.00
4.26
Chi-squared statistic. Summing the last column:
$$\chi^2=\sum_{i=1}^{10}\frac{(O_i-E_i)^2}{E_i}=\boxed{440.51}.$$
Compare to the critical value and conclude. The pdf is fully specified (no parameters were estimated from the sample), so $df=k-1=9$ and $\chi^2_{0.05,9}=16.92$. Since $440.51\gg16.92$ (effectively $p\approx0$), reject $H_0$: $\boxed{\text{the data does not come from } f(x)=2x/25}$. The very first bin alone contributes $400/440.51\approx91\%$ of the statistic — the fit fails almost entirely because far too many low-$x$ values were generated, the opposite of what an increasing density like $f(x)$ should produce.
(b) Diagnose the cause. Two issues are worth separating. First, the question's own claim that this is "the same MCG from Q1" is internally inconsistent: Q1 used $a=23$, this part uses $a=11$ — a genuine discrepancy in the source, not something this solution can silently reconcile. Second, and more fundamentally: a pure multiplicative congruential generator with a composite, non-prime modulus $m=100$ cannot reach the full period $m$; its actual period equals the multiplicative order of $a$ modulo $m$. For $a=11$, direct computation gives order $10$ ($11^{10}\equiv1\ (\mathrm{mod}\ 100)$, and no smaller power does) — so the "100 data points" this generator can produce are really only $10$ distinct pseudo-random values, each repeated $10$ times, nowhere near enough diversity to populate a smooth 10-bin histogram matching a continuous target density. (Even Q1's $a=23$ stream only reaches order $20$ — better, but still a small fraction of the $m=100$ ceiling.) Correction: choose generator parameters that satisfy the Hull–Dobell full-period conditions — as the mixed congruential generator used in Questions 3 and 4 does (verified full period $100$ there, see Q3(c)) — or replace the generator altogether with an industrial-strength one (a combined multiple-recursive generator or Mersenne Twister) whose period vastly exceeds the required sample size.
Item
Result
(a) $\chi^2$ statistic vs. critical value
$440.51 \gg \chi^2_{0.05,9}=16.92$ — reject $H_0$
(b) root cause
MCG with $a=11,\ m=100$ has period (mult. order) $=10$ — far too few distinct values
(b) fix
use a Hull–Dobell full-period (or industrial-strength) generator