Basic Mathematical Statistics
The chi-square, Student’s \(t\), and \(F\) distributions all arise from normal random variables. This appendix explains the constructions, proves the distributional results, and connects them to the general linear hypothesis in Chapter 3 — Distribution Theory of OLS and Inference, Section 3.8: General Linear Hypotheses.
The logical order is normal variables \(\longrightarrow\) sums of squares \(\longrightarrow\) standardized ratios. Degrees of freedom count independent normal components. Independence is essential when forming the ratios.
Chi-square: a sum of squared standard normals
Definition 1 (Chi-square distribution) If \(Z_1,\ldots,Z_r\) are independent \(N(0,1)\) variables, then
\[ U=\sum_{j=1}^{r}Z_j^2\sim\chi^2_r. \]
Each squared standardized component contributes one degree of freedom. Thus \(U\) measures a total squared distance from zero. Its density is
\[ \boxed{f_U(u)=\frac{u^{r/2-1}e^{-u/2}}{2^{r/2}\Gamma(r/2)},\qquad u>0.} \]
This is a gamma density with shape \(r/2\) and scale \(2\). Here \(\Gamma(a)=\int_0^\infty x^{a-1}e^{-x}\,dx\) for \(a>0\). For the sum-of-squares construction, \(r\) is a positive integer; the density also defines \(\chi^2_r\) for any real \(r>0\).
Why normal quadratic forms are chi-square
Theorem 1 (Normal quadratic forms) Let \(\mathbf Z\sim N_n(\mathbf0,\mathbf I_n)\), and let \(\mathbf A\) be symmetric and idempotent, with rank \(r\geq1\). Then
\[ \boxed{\mathbf Z^\top\mathbf A\mathbf Z\sim\chi^2_r.} \]
Corollary 1 (Residual variation in the normal linear model) Throughout the regression applications, assume
\[ \mathbf Y=\mathbf X\boldsymbol\beta+\boldsymbol\varepsilon, \quad \boldsymbol\varepsilon\sim N_n(\mathbf0,\sigma^2\mathbf I_n), \quad \sigma^2>0,\quad \operatorname{rank}(\mathbf X)=p<n, \]
with fixed \(\mathbf X\). Write
\[ \mathbf V=(\mathbf X^\top\mathbf X)^{-1},\quad \mathbf H=\mathbf X\mathbf V\mathbf X^\top,\quad \mathbf M=\mathbf I_n-\mathbf H,\quad \nu=n-p. \]
Since \(\mathbf M\) is symmetric and idempotent with rank \(\nu\), and \(\mathbf M\mathbf X=\mathbf0\),
\[ \frac{\mathrm{SSE}}{\sigma^2} =\left(\frac{\boldsymbol\varepsilon}{\sigma}\right)^\top \mathbf M\left(\frac{\boldsymbol\varepsilon}{\sigma}\right) \sim\chi^2_\nu. \]
Consequently, \(\hat\sigma^2=\mathrm{SSE}/\nu\) is unbiased for \(\sigma^2\).
Proposition 1 (Independence of OLS estimates and residual variation) Both \(\hat{\boldsymbol\beta}-\boldsymbol\beta=\mathbf V\mathbf X^\top\boldsymbol\varepsilon\) and \(\mathbf e=\mathbf M\boldsymbol\varepsilon\) are jointly normal, and
\[ \operatorname{Cov}(\hat{\boldsymbol\beta},\mathbf e) =\sigma^2\mathbf V\mathbf X^\top\mathbf M=\mathbf0. \]
A jointly normal vector with zero cross-covariance has a characteristic function that factors into the two marginal characteristic functions. Thus these vectors are independent, even though the residual covariance is singular. Since \(\mathrm{SSE}=\mathbf e^\top\mathbf e\), \(\hat{\boldsymbol\beta}\) and \(\mathrm{SSE}\) are independent. Zero covariance alone would not suffice without joint normality.
Student’s t: estimating an unknown scale
Definition 2 (Student’s t distribution) Let \(Z\sim N(0,1)\) and \(U\sim\chi^2_\nu\) be independent, with \(\nu>0\). Then
\[ \boxed{T=\frac{Z}{\sqrt{U/\nu}}\sim t_\nu.} \]
The denominator is a random estimate of a standard deviation. It can be small, producing larger standardized values than a normal denominator fixed at \(1\) would produce. The result is a symmetric density with heavier tails than the standard normal density.
Corollary 2 (Student t inference for a regression contrast) For any fixed nonzero \(\mathbf a\in\mathbb R^p\), let \(v_a=\mathbf a^\top\mathbf V\mathbf a>0\). Then
\[ Z=\frac{\mathbf a^\top\hat{\boldsymbol\beta}-\mathbf a^\top\boldsymbol\beta} {\sigma\sqrt{v_a}}\sim N(0,1),\qquad U=\frac{\mathrm{SSE}}{\sigma^2}\sim\chi^2_\nu. \]
The normal quadratic-form section proves their independence. Therefore
\[ \frac{\mathbf a^\top\hat{\boldsymbol\beta}-\mathbf a^\top\boldsymbol\beta} {\hat\sigma\sqrt{v_a}}\sim t_\nu. \]
Under \(H_0:\mathbf a^\top\boldsymbol\beta=d\), replace the true contrast in the numerator by \(d\). Choosing a coordinate vector for \(\mathbf a\) gives the usual coefficient \(t\) test.
F: comparing two independent sums of squares
Definition 3 (F distribution) If \(U_1\sim\chi^2_r\) and \(U_2\sim\chi^2_\nu\) are independent, with \(r,\nu>0\), then
\[ \boxed{F=\frac{U_1/r}{U_2/\nu}\sim F_{r,\nu}.} \]
The two degrees of freedom have different roles: \(r\) belongs to the numerator and \(\nu\) to the denominator. Both normalized sums of squares have expectation \(1\). Their ratio is positive and can be large when the numerator is unusually large or the denominator is small. The ratio itself does not have mean \(1\): its mean is \(\nu/(\nu-2)\) when \(\nu>2\).
Proof of the general linear hypothesis F test
Theorem 2 (F test for the general linear hypothesis) Consider \(H_0:\mathbf C\boldsymbol\beta=\mathbf d\), where \(\mathbf C\) is \(r\times p\) with row rank \(r\), \(1\leq r\leq p\), and \(\mathbf d\in\mathbb R^r\). Keep the normal-model assumptions from the normal quadratic-form section, and write
\[ \mathbf D=\mathbf C\mathbf V\mathbf C^\top, \qquad \mathbf L=\mathbf C\hat{\boldsymbol\beta}-\mathbf d. \]
Under \(H_0\), the statistic
\[ F=\frac{\mathbf L^\top\mathbf D^{-1}\mathbf L}{r\hat\sigma^2} \sim F_{r,n-p}. \]
Corollary 3 (The identity \(T^2=F\) for one restriction) When \(r=1\) and \(\mathbf C=\mathbf a^\top\),
\[ F=\frac{(\mathbf a^\top\hat{\boldsymbol\beta}-d)^2} {\hat\sigma^2\mathbf a^\top\mathbf V\mathbf a}=T^2. \]
Equivalently, squaring the defining ratio for \(T\) gives \(T^2=(Z^2/1)/(U/\nu)\), and \(Z^2\sim\chi^2_1\) independently of \(U\). Hence \(T^2\sim F_{1,\nu}\). The upper-tail \(F\) test matches the two-sided \(t\) test, not a one-sided test.
A compact regression example and reading guide
Example 1 (Regression with 20 observations and 3 coefficients) Suppose the normal linear model has \(n=20\) observations and \(p=3\) coefficients, including an intercept. Then \(\nu=17\). If \(\mathrm{SSE}=34\), the observed variance estimate is \(\hat\sigma^2=34/17=2\).
- The sampling law of \(\mathrm{SSE}/\sigma^2\) is \(\chi^2_{17}\); the observed value \(34\) does not identify the unknown \(\sigma^2\).
- For \(H_0:\beta_1=0\), suppose \(\hat\beta_1=1.5\) and \(V_{22}=0.05\) (the intercept is the first coordinate). Then
\[ T=\frac{1.5}{\sqrt{2(0.05)}}\approx4.743, \qquad F=T^2=22.5. \]
The null reference distributions are \(t_{17}\) and \(F_{1,17}\), respectively.
- For a joint test of two independent restrictions, \(r=2\). If the observed unscaled numerator quadratic form \(\mathbf L^\top\mathbf D^{-1}\mathbf L\) is \(12\), then \(F=12/(2\times2)=3\), compared with \(F_{2,17}\).
Reference books.
- Seber, G. A. F., and Lee, A. J. (2003). Linear Regression Analysis, 2nd ed. Wiley. Read Chapter 2, Multivariate Normal Distribution, for the normal-vector background; Chapter 3, Linear Regression: Estimation and Distribution Theory, for regression sampling distributions; and Chapter 4, Hypothesis Testing, for linear restrictions. This is the main reference for the regression applications above. Publisher and contents.
- Casella, G., and Berger, R. (2024 reprint). Statistical Inference, 2nd ed. Chapman & Hall/CRC. Chapters 3, Common Families of Distributions, and 5, Properties of a Random Sample, provide the probability background for chi-square, \(t\), and \(F\) distributions. The publisher identifies this as a reprint of the earlier second edition. Publisher and contents.
The derivations here are written out for these notes; the books provide further development and exercises. This appendix supports the Chapter 3 — Distribution Theory of OLS and Inference, Section 3.8: General Linear Hypotheses.