Basic Mathematical Statistics

The chi-square, Student’s \(t\), and \(F\) distributions all arise from normal random variables. This appendix explains the constructions, proves the distributional results, and connects them to the general linear hypothesis in Chapter 3 — Distribution Theory of OLS and Inference, Section 3.8: General Linear Hypotheses.

The logical order is normal variables \(\longrightarrow\) sums of squares \(\longrightarrow\) standardized ratios. Degrees of freedom count independent normal components. Independence is essential when forming the ratios.

Chi-square: a sum of squared standard normals

Definition 1 (Chi-square distribution) If \(Z_1,\ldots,Z_r\) are independent \(N(0,1)\) variables, then

\[ U=\sum_{j=1}^{r}Z_j^2\sim\chi^2_r. \]

Each squared standardized component contributes one degree of freedom. Thus \(U\) measures a total squared distance from zero. Its density is

\[ \boxed{f_U(u)=\frac{u^{r/2-1}e^{-u/2}}{2^{r/2}\Gamma(r/2)},\qquad u>0.} \]

This is a gamma density with shape \(r/2\) and scale \(2\). Here \(\Gamma(a)=\int_0^\infty x^{a-1}e^{-x}\,dx\) for \(a>0\). For the sum-of-squares construction, \(r\) is a positive integer; the density also defines \(\chi^2_r\) for any real \(r>0\).

NoteProof: chi-square density and moments

For \(s<1/2\),

\[ \begin{aligned} \mathbb E(e^{sZ_j^2}) &=\frac{1}{\sqrt{2\pi}}\int_{-\infty}^{\infty} \exp\!\left\{-\frac{1-2s}{2}z^2\right\}\,dz\\ &=(1-2s)^{-1/2}. \end{aligned} \]

The substitution \(w=\sqrt{1-2s}\,z\) gives the second equality. Independence therefore gives

\[ M_U(s)=\prod_{j=1}^{r}\mathbb E(e^{sZ_j^2})=(1-2s)^{-r/2}. \]

To identify the law, integrate the displayed gamma density against \(e^{su}\):

\[ \begin{aligned} \int_0^\infty e^{su}\frac{u^{r/2-1}e^{-u/2}}{2^{r/2}\Gamma(r/2)}\,du &=\frac{\Gamma(r/2)(1/2-s)^{-r/2}}{2^{r/2}\Gamma(r/2)}\\ &=(1-2s)^{-r/2}. \end{aligned} \]

Moment-generating functions finite in a neighborhood of zero uniquely determine a probability law, so \(U\) has the claimed density. Differentiating at zero gives

\[ \mathbb E(U)=r,\qquad \operatorname{Var}(U)=2r. \]

In particular, \(U/r\) has mean \(1\) and variance \(2/r\): it estimates a unit variance more stably as the degrees of freedom increase.

Why normal quadratic forms are chi-square

Theorem 1 (Normal quadratic forms) Let \(\mathbf Z\sim N_n(\mathbf0,\mathbf I_n)\), and let \(\mathbf A\) be symmetric and idempotent, with rank \(r\geq1\). Then

\[ \boxed{\mathbf Z^\top\mathbf A\mathbf Z\sim\chi^2_r.} \]

NoteProof: normal quadratic forms

The spectral theorem gives an orthogonal matrix \(\mathbf Q\) such that

\[ \mathbf A=\mathbf Q\operatorname{diag}(\mathbf I_r,\mathbf0)\mathbf Q^\top. \]

Indeed, idempotence implies \(\lambda^2=\lambda\), so every eigenvalue is \(0\) or \(1\); exactly \(r\) eigenvalues are \(1\). Set \(\mathbf W=\mathbf Q^\top\mathbf Z\). An orthogonal transformation preserves the standard multivariate normal law, because its mean is zero and its covariance is \(\mathbf Q^\top\mathbf Q=\mathbf I_n\). Hence the \(W_j\) are independent \(N(0,1)\) variables, and

\[ \mathbf Z^\top\mathbf A\mathbf Z =\mathbf W^\top\operatorname{diag}(\mathbf I_r,\mathbf0)\mathbf W =\sum_{j=1}^{r}W_j^2\sim\chi^2_r. \]

Geometrically, \(\mathbf A\) projects onto an \(r\)-dimensional subspace. The quadratic form is the squared length of the projected normal vector. If \(r=0\), it is identically zero rather than having a positive-degree chi-square density.

Corollary 1 (Residual variation in the normal linear model) Throughout the regression applications, assume

\[ \mathbf Y=\mathbf X\boldsymbol\beta+\boldsymbol\varepsilon, \quad \boldsymbol\varepsilon\sim N_n(\mathbf0,\sigma^2\mathbf I_n), \quad \sigma^2>0,\quad \operatorname{rank}(\mathbf X)=p<n, \]

with fixed \(\mathbf X\). Write

\[ \mathbf V=(\mathbf X^\top\mathbf X)^{-1},\quad \mathbf H=\mathbf X\mathbf V\mathbf X^\top,\quad \mathbf M=\mathbf I_n-\mathbf H,\quad \nu=n-p. \]

Since \(\mathbf M\) is symmetric and idempotent with rank \(\nu\), and \(\mathbf M\mathbf X=\mathbf0\),

\[ \frac{\mathrm{SSE}}{\sigma^2} =\left(\frac{\boldsymbol\varepsilon}{\sigma}\right)^\top \mathbf M\left(\frac{\boldsymbol\varepsilon}{\sigma}\right) \sim\chi^2_\nu. \]

Consequently, \(\hat\sigma^2=\mathrm{SSE}/\nu\) is unbiased for \(\sigma^2\).

Proposition 1 (Independence of OLS estimates and residual variation) Both \(\hat{\boldsymbol\beta}-\boldsymbol\beta=\mathbf V\mathbf X^\top\boldsymbol\varepsilon\) and \(\mathbf e=\mathbf M\boldsymbol\varepsilon\) are jointly normal, and

\[ \operatorname{Cov}(\hat{\boldsymbol\beta},\mathbf e) =\sigma^2\mathbf V\mathbf X^\top\mathbf M=\mathbf0. \]

A jointly normal vector with zero cross-covariance has a characteristic function that factors into the two marginal characteristic functions. Thus these vectors are independent, even though the residual covariance is singular. Since \(\mathrm{SSE}=\mathbf e^\top\mathbf e\), \(\hat{\boldsymbol\beta}\) and \(\mathrm{SSE}\) are independent. Zero covariance alone would not suffice without joint normality.

Student’s t: estimating an unknown scale

Definition 2 (Student’s t distribution) Let \(Z\sim N(0,1)\) and \(U\sim\chi^2_\nu\) be independent, with \(\nu>0\). Then

\[ \boxed{T=\frac{Z}{\sqrt{U/\nu}}\sim t_\nu.} \]

The denominator is a random estimate of a standard deviation. It can be small, producing larger standardized values than a normal denominator fixed at \(1\) would produce. The result is a symmetric density with heavier tails than the standard normal density.

NoteProof: Student t density

Use \(t=z/\sqrt{u/\nu}\) and retain \(u\). The inverse map and absolute Jacobian determinant are

\[ z=t\sqrt{u/\nu},\qquad \left|\frac{\partial(z,u)}{\partial(t,u)}\right|=\sqrt{u/\nu}. \]

Independence gives a product joint density. Integrating out \(u\) yields

\[ \begin{aligned} f_T(t) &=\int_0^\infty \frac{e^{-t^2u/(2\nu)}}{\sqrt{2\pi}} \frac{u^{\nu/2-1}e^{-u/2}}{2^{\nu/2}\Gamma(\nu/2)} \sqrt{\frac{u}{\nu}}\,du\\ &=\frac{1}{\sqrt{2\pi\nu}\,2^{\nu/2}\Gamma(\nu/2)} \int_0^\infty u^{(\nu+1)/2-1} e^{-(1+t^2/\nu)u/2}\,du. \end{aligned} \]

Using \(\int_0^\infty u^{a-1}e^{-bu}\,du=\Gamma(a)b^{-a}\) for \(a,b>0\), we obtain

\[ \boxed{f_T(t)=\frac{\Gamma((\nu+1)/2)}{\sqrt{\nu\pi}\,\Gamma(\nu/2)} \left(1+\frac{t^2}{\nu}\right)^{-(\nu+1)/2},\qquad t\in\mathbb R.} \]

This is the Student \(t\) density. Its dependence on \(t^2\) proves symmetry. Its tails decay polynomially, as \(|t|^{-(\nu+1)}\), rather than exponentially as for a normal density. Also, \(\operatorname{Var}(U/\nu)=2/\nu\to0\), so \(U/\nu\to1\) in probability and \(t_\nu\) approaches \(N(0,1)\) as \(\nu\to\infty\).

Corollary 2 (Student t inference for a regression contrast) For any fixed nonzero \(\mathbf a\in\mathbb R^p\), let \(v_a=\mathbf a^\top\mathbf V\mathbf a>0\). Then

\[ Z=\frac{\mathbf a^\top\hat{\boldsymbol\beta}-\mathbf a^\top\boldsymbol\beta} {\sigma\sqrt{v_a}}\sim N(0,1),\qquad U=\frac{\mathrm{SSE}}{\sigma^2}\sim\chi^2_\nu. \]

The normal quadratic-form section proves their independence. Therefore

\[ \frac{\mathbf a^\top\hat{\boldsymbol\beta}-\mathbf a^\top\boldsymbol\beta} {\hat\sigma\sqrt{v_a}}\sim t_\nu. \]

Under \(H_0:\mathbf a^\top\boldsymbol\beta=d\), replace the true contrast in the numerator by \(d\). Choosing a coordinate vector for \(\mathbf a\) gives the usual coefficient \(t\) test.

F: comparing two independent sums of squares

Definition 3 (F distribution) If \(U_1\sim\chi^2_r\) and \(U_2\sim\chi^2_\nu\) are independent, with \(r,\nu>0\), then

\[ \boxed{F=\frac{U_1/r}{U_2/\nu}\sim F_{r,\nu}.} \]

The two degrees of freedom have different roles: \(r\) belongs to the numerator and \(\nu\) to the denominator. Both normalized sums of squares have expectation \(1\). Their ratio is positive and can be large when the numerator is unusually large or the denominator is small. The ratio itself does not have mean \(1\): its mean is \(\nu/(\nu-2)\) when \(\nu>2\).

NoteProof: F density

Put \(a=r/2\), \(b=\nu/2\), and \(c=r/\nu\). Transform \((U_1,U_2)\) to \((F,V)\), where \(V=U_2\). The inverse map is

\[ u_1=cfv,\qquad u_2=v,\qquad f>0,\ v>0, \]

and its absolute Jacobian determinant is

\[ \left|\det\begin{pmatrix}cv&cf\\0&1\end{pmatrix}\right|=cv. \]

Using independence and the two chi-square densities,

\[ \begin{aligned} f_F(f) &=\int_0^\infty \frac{(cfv)^{a-1}e^{-cfv/2}}{2^a\Gamma(a)} \frac{v^{b-1}e^{-v/2}}{2^b\Gamma(b)}\,cv\,dv\\ &=\frac{c^a f^{a-1}}{2^{a+b}\Gamma(a)\Gamma(b)} \int_0^\infty v^{a+b-1}e^{-(1+cf)v/2}\,dv\\ &=\frac{\Gamma(a+b)}{\Gamma(a)\Gamma(b)} c^a f^{a-1}(1+cf)^{-(a+b)}. \end{aligned} \]

Substituting the definitions of \(a,b,c\) gives the \(F\) density:

\[ \boxed{f_F(f)=\frac{\Gamma((r+\nu)/2)}{\Gamma(r/2)\Gamma(\nu/2)} \left(\frac r\nu\right)^{r/2} f^{r/2-1}\left(1+\frac r\nu f\right)^{-(r+\nu)/2},\quad f>0.} \]

NoteRemark: normalization and order of the degrees of freedom

The raw ratio \(U_1/U_2\) equals \((r/\nu)F\), and hence is not generally \(F_{r,\nu}\). Reversing the numerator and denominator gives \(1/F\sim F_{\nu,r}\).

Proof of the general linear hypothesis F test

Theorem 2 (F test for the general linear hypothesis) Consider \(H_0:\mathbf C\boldsymbol\beta=\mathbf d\), where \(\mathbf C\) is \(r\times p\) with row rank \(r\), \(1\leq r\leq p\), and \(\mathbf d\in\mathbb R^r\). Keep the normal-model assumptions from the normal quadratic-form section, and write

\[ \mathbf D=\mathbf C\mathbf V\mathbf C^\top, \qquad \mathbf L=\mathbf C\hat{\boldsymbol\beta}-\mathbf d. \]

Under \(H_0\), the statistic

\[ F=\frac{\mathbf L^\top\mathbf D^{-1}\mathbf L}{r\hat\sigma^2} \sim F_{r,n-p}. \]

NoteProof: general linear hypothesis F test

Step 1: standardize the restrictions. The matrix \(\mathbf D\) is positive definite: for nonzero \(\mathbf b\), full row rank implies \(\mathbf C^\top\mathbf b\ne\mathbf0\), so \(\mathbf b^\top\mathbf D\mathbf b=(\mathbf C^\top\mathbf b)^\top\mathbf V(\mathbf C^\top\mathbf b)>0\). Under \(H_0\), \(\mathbf L\sim N_r(\mathbf0,\sigma^2\mathbf D)\). Consequently,

\[ \mathbf W=\sigma^{-1}\mathbf D^{-1/2}\mathbf L\sim N_r(\mathbf0,\mathbf I_r), \]

where \(\mathbf D^{-1/2}\) is the symmetric positive definite inverse square root. The numerator quadratic form therefore satisfies

\[ Q=\frac{\mathbf L^\top\mathbf D^{-1}\mathbf L}{\sigma^2} =\mathbf W^\top\mathbf W\sim\chi^2_r. \]

Step 2: identify the independent denominator. We already proved \(U=\mathrm{SSE}/\sigma^2\sim\chi^2_\nu\), with \(\nu=n-p\). The quantity \(Q\) is a function of \(\hat{\boldsymbol\beta}\), whereas \(U\) is a function of \(\mathrm{SSE}\). Their independence follows from the normal-model independence proved in the normal quadratic-form section.

Step 3: form the ratio. Apply the F ratio result and cancel \(\sigma^2\):

\[ \boxed{F=\frac{Q/r}{U/\nu} =\frac{\mathbf L^\top\mathbf D^{-1}\mathbf L}{r\hat\sigma^2} \sim F_{r,n-p}\quad\text{under }H_0.} \]

Substituting \(\mathbf L=\mathbf C\hat{\boldsymbol\beta}-\mathbf d\) and \(\mathbf D=\mathbf C(\mathbf X^\top\mathbf X)^{-1}\mathbf C^\top\) recovers the statistic in Chapter 3 — Distribution Theory of OLS and Inference, Section 3.8: General Linear Hypotheses. Thus \(r\) counts the independent restrictions and \(n-p\) counts residual degrees of freedom. Large values indicate a departure from \(H_0\); a level-\(\alpha\) test rejects when \(F>F_{1-\alpha;r,n-p}\), where the right-hand side is an upper quantile. The central \(F\) law above is a null distribution; under an alternative the numerator is generally noncentral chi-square.

Corollary 3 (The identity \(T^2=F\) for one restriction) When \(r=1\) and \(\mathbf C=\mathbf a^\top\),

\[ F=\frac{(\mathbf a^\top\hat{\boldsymbol\beta}-d)^2} {\hat\sigma^2\mathbf a^\top\mathbf V\mathbf a}=T^2. \]

Equivalently, squaring the defining ratio for \(T\) gives \(T^2=(Z^2/1)/(U/\nu)\), and \(Z^2\sim\chi^2_1\) independently of \(U\). Hence \(T^2\sim F_{1,\nu}\). The upper-tail \(F\) test matches the two-sided \(t\) test, not a one-sided test.

A compact regression example and reading guide

Example 1 (Regression with 20 observations and 3 coefficients) Suppose the normal linear model has \(n=20\) observations and \(p=3\) coefficients, including an intercept. Then \(\nu=17\). If \(\mathrm{SSE}=34\), the observed variance estimate is \(\hat\sigma^2=34/17=2\).

  • The sampling law of \(\mathrm{SSE}/\sigma^2\) is \(\chi^2_{17}\); the observed value \(34\) does not identify the unknown \(\sigma^2\).
  • For \(H_0:\beta_1=0\), suppose \(\hat\beta_1=1.5\) and \(V_{22}=0.05\) (the intercept is the first coordinate). Then

\[ T=\frac{1.5}{\sqrt{2(0.05)}}\approx4.743, \qquad F=T^2=22.5. \]

The null reference distributions are \(t_{17}\) and \(F_{1,17}\), respectively.

  • For a joint test of two independent restrictions, \(r=2\). If the observed unscaled numerator quadratic form \(\mathbf L^\top\mathbf D^{-1}\mathbf L\) is \(12\), then \(F=12/(2\times2)=3\), compared with \(F_{2,17}\).
ImportantWhat to check before using these results

Confirm the normal-model assumptions, the rank of the projection or restriction matrix, the residual degrees of freedom, and the required independence. A symmetric matrix alone is not enough for the quadratic-form chi-square result: idempotence is also needed. Likewise, correct marginal distributions alone do not establish a \(t\) or \(F\) ratio law.

Reference books.

  1. Seber, G. A. F., and Lee, A. J. (2003). Linear Regression Analysis, 2nd ed. Wiley. Read Chapter 2, Multivariate Normal Distribution, for the normal-vector background; Chapter 3, Linear Regression: Estimation and Distribution Theory, for regression sampling distributions; and Chapter 4, Hypothesis Testing, for linear restrictions. This is the main reference for the regression applications above. Publisher and contents.
  2. Casella, G., and Berger, R. (2024 reprint). Statistical Inference, 2nd ed. Chapman & Hall/CRC. Chapters 3, Common Families of Distributions, and 5, Properties of a Random Sample, provide the probability background for chi-square, \(t\), and \(F\) distributions. The publisher identifies this as a reprint of the earlier second edition. Publisher and contents.

The derivations here are written out for these notes; the books provide further development and exercises. This appendix supports the Chapter 3 — Distribution Theory of OLS and Inference, Section 3.8: General Linear Hypotheses.