Assignment 2

STAT 8561 | Fall 2026
Due date: September 30, 2026
Total: 40 points
Coverage: Chapter 2 — Least Squares Estimation and Chapter 3 — Distribution Theory of OLS and Inference.

This assignment emphasizes conceptual understanding of least squares, projection, and inference. Read the assumptions in each question carefully.

Question Topic Points
1 Least squares: assumptions and consequences 6
2 Projection geometry and rank 8
3 Sampling properties and the role of normality 8
4 Interpreting confidence and prediction intervals 6
5 Distribution and inference for OLS 12
Total 40
NoteInstructions

Use Quarto or R Markdown. Upload both files to the Assignment 2 submission folder on iCollege:

  1. Your rendered PDF, named A2_LastName_FirstName.pdf.
  2. Its matching source file, named A2_LastName_FirstName.qmd or A2_LastName_FirstName.Rmd.

The PDF must be generated from the submitted source file. HTML alone does not satisfy the submission requirement.

  • Put your name and course number at the beginning of the PDF.
  • Clearly label every question and subpart.
  • For each multiple-choice item, select exactly one answer. Enter only the letter.
  • For True/False items, enter only T or F. For matching items, enter one letter per row; letters may be reused.
  • Questions 1-4 require no written explanations, derivations, or R output. Work through the reasoning before selecting your answers.
  • Question 5 requires only the requested choices and numerical answers. Round numerical answers to three decimal places, retaining full precision during calculations.
  • R is optional for checking calculations. If you use R, include your code after your answers in executable chunks.
  • Your source must render from beginning to end without errors. Type mathematical expressions using LaTeX; do not submit screenshots of answers or code.

All necessary information is supplied in the questions. No external data or additional R packages are required.

Question 1: Least Squares: Assumptions and Consequences [6]

Consider \(\mathbf{Y}=\mathbf{X}\boldsymbol{\beta}+\boldsymbol{\varepsilon}\), where \(\mathbf{X}\) is fixed, has full column rank, and \(n>p\). Select one answer per part.

(a) Changing the distributional assumption [2]

Two students use the same observed \(\mathbf{X}\) and \(\mathbf{Y}\). One assumes normal errors; the other does not specify an error distribution. Both define their estimate by minimizing

\[ S(\boldsymbol{\beta})=\|\mathbf{Y}-\mathbf{X}\boldsymbol{\beta}\|^2. \]

Which statement is correct?

  1. The second student cannot compute a least squares estimate.

  2. The estimates must differ because the distributional assumptions differ.

  3. The estimates are the same because they minimize the same criterion for the same data.

  4. The estimates are the same only if the observed residuals are all zero.

C. Both students minimize the same function of the same observed data. Normality is not needed to compute the least squares estimate; it is used for exact distributional inference.

(b) What makes the minimizer unique? [2]

Which statement explains why solving the normal equations gives a unique least squares minimizer under the stated assumptions?

  1. Full column rank makes \(\mathbf{X}^\top\mathbf{X}\) positive definite.

  2. Normal errors make every system of normal equations invertible.

  3. Having more observations than parameters guarantees full column rank for any design matrix.

  4. Minimizing the sum of squared residuals forces every residual to be zero.

A. Full column rank makes \(\mathbf{X}^\top\mathbf{X}\) positive definite because

\[ \mathbf{a}^\top\mathbf{X}^\top\mathbf{X}\mathbf{a} =\|\mathbf{X}\mathbf{a}\|^2>0\quad\text{for every }\mathbf{a}\ne\mathbf{0}. \]

The least squares criterion therefore has a unique minimizer. The condition \(n>p\) alone does not rule out linearly dependent columns.

(c) Shifting every response [2]

Now suppose the model includes an intercept. Add 5 to every observed response, keep all predictors unchanged, and refit the same model by least squares. What happens?

  1. Every slope increases by 5, while the intercept is unchanged.

  2. The coefficients are unchanged, and every residual increases by 5.

  3. The intercept increases by 5, and every residual increases by 5.

  4. The intercept increases by 5, the slopes are unchanged, and the residuals are unchanged.

D. Adding 5 to the intercept adds 5 to every fitted response. The slopes are unchanged, and

\[ (\mathbf{Y}+5\mathbf{1}_n)-(\hat{\mathbf{Y}}+5\mathbf{1}_n)=\mathbf{e}, \]

so the residuals are unchanged.

Question 2: Projection Geometry and Rank [8]

For parts (a) and (c), let \(\mathbf{X}\) have full column rank. Define

\[ \mathbf{H}=\mathbf{X}(\mathbf{X}^\top\mathbf{X})^{-1}\mathbf{X}^\top, \quad \mathbf{M}=\mathbf{I}_n-\mathbf{H}, \quad \hat{\mathbf{Y}}=\mathbf{H}\mathbf{Y}, \quad \mathbf{e}=\mathbf{M}\mathbf{Y}. \]

(a) Projecting a second time [2]

A student applies \(\mathbf{H}\) to the fitted vector and to the residual vector. Which pair of results must hold?

  1. \(\mathbf{H}\hat{\mathbf{Y}}=\mathbf{0}\) and \(\mathbf{H}\mathbf{e}=\mathbf{e}\).

  2. \(\mathbf{H}\hat{\mathbf{Y}}=\hat{\mathbf{Y}}\) and \(\mathbf{H}\mathbf{e}=\mathbf{0}\).

  3. \(\mathbf{H}\hat{\mathbf{Y}}=\mathbf{Y}\) and \(\mathbf{H}\mathbf{e}=\mathbf{e}\).

  4. \(\mathbf{H}\hat{\mathbf{Y}}=\hat{\mathbf{Y}}\) and \(\mathbf{H}\mathbf{e}=\mathbf{e}\).

B. Since \(\mathbf{H}^2=\mathbf{H}\),

\[ \mathbf{H}\hat{\mathbf{Y}}=\mathbf{H}^2\mathbf{Y}=\hat{\mathbf{Y}}, \qquad \mathbf{H}\mathbf{e}=\mathbf{H}(\mathbf{I}_n-\mathbf{H})\mathbf{Y}=\mathbf{0}. \]

The fitted vector lies in the model space, and the residual vector is orthogonal to it.

(b) Adding a redundant predictor [2]

Suppose \(\mathbf{x}\) is not constant. Compare the least squares fits using

\[ \mathbf{X}=[\mathbf{1}_n,\mathbf{x}] \quad\text{and}\quad \mathbf{W}=[\mathbf{1}_n,\mathbf{x},2\mathbf{x}] \]

for the same response vector. Which statement is correct?

  1. The extra column must strictly reduce SSE.

  2. The fit using \(\mathbf{W}\) does not exist because \(\mathbf{W}^\top\mathbf{W}\) is singular.

  3. The fit using \(\mathbf{W}\) has nonunique coefficients, but its fitted values and SSE equal those from \(\mathbf{X}\).

  4. The fit using \(\mathbf{W}\) has nonunique coefficients and therefore nonunique fitted values.

C. Adding \(2\mathbf{x}\) does not change the column space. Both models project the response onto the same space, giving identical fitted values, residuals, and SSE. The coefficients for \(\mathbf{W}\) are nonunique because its columns are linearly dependent.

(c) Geometric consequences [4]

Enter T or F. Each statement is worth 1 point. Here \(\bar Y=n^{-1}\sum_iY_i\).

Statement T/F
(i) The uncentered identity \(\mathbf{Y}^\top\mathbf{Y}=\hat{\mathbf{Y}}^\top\hat{\mathbf{Y}}+\mathbf{e}^\top\mathbf{e}\) holds for the least squares fit. ___
(ii) The centered identity \(\sum_i(Y_i-\bar Y)^2=\sum_i(\hat Y_i-\bar Y)^2+\mathrm{SSE}\) is guaranteed for every least squares model, even one without an intercept. ___
(iii) If the model includes an intercept, the average fitted value equals the average observed response. ___
(iv) The residual vector must be orthogonal to the observed response vector \(\mathbf{Y}\) itself. ___

(i) T; (ii) F; (iii) T; (iv) F.

  • (i) \(\hat{\mathbf{Y}}^\top\mathbf{e}=0\) gives the uncentered sum-of-squares decomposition.
  • (ii) The centered decomposition holds when the model includes an intercept, but is not guaranteed without one.
  • (iii) An intercept gives \(\mathbf{1}_n^\top\mathbf{e}=0\), so the observed and fitted means agree.
  • (iv) In fact, \(\mathbf{e}^\top\mathbf{Y}=\mathbf{e}^\top(\hat{\mathbf{Y}}+\mathbf{e})=\mathrm{SSE}\), which need not be zero.

Question 3: Sampling Properties and the Role of Normality [8]

Throughout this question, \(\mathbf{X}\) is fixed with full column rank, \(n>p\), and \(\sigma^2>0\).

(a) Match the result to its assumptions [4]

Choose the weakest sufficient set of assumptions among A, B, and C for each result. Each row is worth 1 point. Letters may be reused.

  1. The least squares criterion and full column rank alone; no error moment or distribution assumptions.

  2. \(\mathbb{E}(\boldsymbol{\varepsilon})=\mathbf{0}\) and \(\operatorname{Var}(\boldsymbol{\varepsilon})=\sigma^2\mathbf{I}_n\), without assuming normality.

  3. The normal linear model, \(\boldsymbol{\varepsilon}\sim N_n(\mathbf{0},\sigma^2\mathbf{I}_n)\).

Result Letter
(i) The least squares coefficient vector is the unique minimizer of \(S(\boldsymbol{\beta})\). ___
(ii) \(\mathbb{E}(\hat{\boldsymbol{\beta}})=\boldsymbol{\beta}\). ___
(iii) \(\mathrm{SSE}/\sigma^2\sim\chi^2_{n-p}\). ___
(iv) \(\operatorname{Var}(\hat{\boldsymbol{\beta}})=\sigma^2(\mathbf{X}^\top\mathbf{X})^{-1}\). ___

(i) A; (ii) B; (iii) C; (iv) B.

  • (i) Uniqueness of the minimizer is an algebraic consequence of full column rank.
  • (ii) Zero-mean errors give \(\mathbb{E}(\hat{\boldsymbol{\beta}})=\boldsymbol{\beta}\).
  • (iii) The exact chi-square distribution requires the normal linear model among the listed assumptions.
  • (iv) The covariance formula follows from \(\operatorname{Var}(\boldsymbol{\varepsilon})=\sigma^2\mathbf{I}_n\) without normality.

For (ii), zero-mean errors alone suffice; B is the weakest sufficient set among the choices provided.

(b) Why does independence follow? [2]

Under the normal linear model, which argument correctly explains why \(\hat{\boldsymbol{\beta}}\) and \(\mathrm{SSE}\) are independent?

  1. \(\hat{\boldsymbol{\beta}}\) and \(\mathbf{e}\) are jointly normal with zero covariance, and SSE is a function of \(\mathbf{e}\).

  2. Any two uncorrelated random vectors are independent, regardless of their distribution.

  3. They estimate different parameters, so they must be independent.

  4. Each observed residual is independent of every other observed residual in every normal linear model.

A. Both \(\hat{\boldsymbol{\beta}}\) and \(\mathbf{e}\) are linear transformations of the normal response vector, so they are jointly normal. Also,

\[ \operatorname{Cov}(\hat{\boldsymbol{\beta}},\mathbf{e}) =\sigma^2(\mathbf{X}^\top\mathbf{X})^{-1}\mathbf{X}^\top\mathbf{M} =\mathbf{0}. \]

Joint normality turns zero covariance into independence. Since \(\mathrm{SSE}=\mathbf{e}^\top\mathbf{e}\) is a function of \(\mathbf{e}\), it is independent of \(\hat{\boldsymbol{\beta}}\). Uncorrelatedness alone would not suffice without joint normality.

(c) Increasing the error variance [2]

Keep \(\mathbf{X}\) and \(\boldsymbol{\beta}\) fixed. In a second model, the errors still have mean zero, but their covariance is \(4\sigma^2\mathbf{I}_n\) instead of \(\sigma^2\mathbf{I}_n\). Which statement describes the sampling distribution’s mean and covariance for the OLS estimator?

  1. Its mean and covariance are both multiplied by 4.

  2. Its mean is unchanged, and its covariance is multiplied by 16.

  3. Its mean and covariance are both unchanged.

  4. Its mean is unchanged, and its covariance is multiplied by 4; each coefficient’s standard deviation doubles.

D. The estimator remains unbiased, while its covariance becomes

\[ 4\sigma^2(\mathbf{X}^\top\mathbf{X})^{-1}. \]

Each coefficient’s variance is multiplied by 4, so its standard deviation is multiplied by \(\sqrt{4}=2\).

Question 4: Interpreting Confidence and Prediction Intervals [6]

Assume a correctly specified normal linear model. For prediction, a future observation has an error independent of the original data and with the same error variance.

(a) Mean response versus one future response [2]

At the same predictor value, software reports two intervals at the 95% level:

  • Interval I: \((2.493,3.507)\);
  • Interval II: \((1.659,4.341)\).

One is a confidence interval for the mean response; the other is a prediction interval for one future response. Which interpretation is correct?

  1. I is the prediction interval because predicting one observation involves less uncertainty than estimating a mean.

  2. Both intervals describe the same target; their different widths are only rounding error.

  3. I is the mean-response interval; II is the prediction interval because it also accounts for the new observation’s error.

  4. II is the mean-response interval because averaging adds the new observation’s error variance.

C. At the same predictor value, the mean-response interval uses standard error \(\hat\sigma\sqrt{h_0}\), whereas the prediction interval uses \(\hat\sigma\sqrt{1+h_0}\), where \(h_0=\mathbf{x}_0^\top(\mathbf{X}^\top\mathbf{X})^{-1}\mathbf{x}_0\). The extra 1 accounts for the new observation’s error variance. Thus the narrower interval I concerns the mean and the wider interval II concerns one future response.

(b) Connecting a confidence interval and a test [2]

A 95% confidence interval for \(\beta_1\) is \((0.4,1.6)\). Consider the corresponding two-sided \(t\) test of

\[ H_0:\beta_1=1\qquad\text{versus}\qquad H_1:\beta_1\neq1 \]

at \(\alpha=0.05\). Which conclusion is correct?

  1. Reject \(H_0\) because zero is outside the interval.

  2. Fail to reject \(H_0\) because 1 is inside the interval; this does not prove that \(\beta_1=1\).

  3. Accept \(H_0\) as proven true because 1 is inside the interval.

  4. No conclusion is possible because a \(t\) test can only test a coefficient against zero.

B. The hypothesized value 1 lies inside \((0.4,1.6)\), so the corresponding two-sided test fails to reject \(H_0:\beta_1=1\) at the 5% level. This does not prove the null hypothesis. Whether zero lies in the interval is irrelevant to this stated null value.

(c) What an interval means [2]

Enter T or F. Each statement is worth 1 point. Assume a positive estimated error variance and \(\mathbf{x}_0^\top(\mathbf{X}^\top\mathbf{X})^{-1}\mathbf{x}_0>0\).

Statement T/F
(i) Holding the data and predictor value fixed, changing a mean-response confidence level from 95% to 99% makes the interval wider. ___
(ii) A 95% confidence interval for the mean response is designed to contain 95% of individual future responses at that predictor value. ___

(i) T; (ii) F.

  • (i) Holding the data fixed, a higher confidence level uses a larger critical value and produces a wider interval.
  • (ii) A confidence interval for the mean concerns the unknown mean response. It does not describe a range containing 95% of individual future responses; those require a prediction interval.

Question 5: Distribution and Inference [12]

Suppose a multiple linear regression model has \(n=20\) observations and \(p=3\) parameters: an intercept and two slopes. The first column of \(\mathbf{X}\) is \(\mathbf{1}_{20}\). Assume

\[ \mathbf{Y}\sim N_{20}\!\left(\mathbf{X}\boldsymbol{\beta},\sigma^2\mathbf{I}_{20}\right). \]

Suppose

\[ \hat{\boldsymbol{\beta}}=\begin{pmatrix}2.0\\1.5\\-0.5\end{pmatrix}, \qquad (\mathbf{X}^\top\mathbf{X})^{-1} =\begin{pmatrix}0.05&0&0\\0&0.05&0\\0&0&0.08\end{pmatrix}, \]

and \(\mathrm{SSE}=34\). Use \(t_{0.975,17}=2.110\).

(a) [3]

Choose the correct distributions.

For \(\hat{\boldsymbol{\beta}}\):

  1. \(N_3\!\left(\mathbf{0},\sigma^2\mathbf{I}_3\right)\)

  2. \(N_3\!\left(\boldsymbol{\beta},\sigma^2(\mathbf{X}^\top\mathbf{X})^{-1}\right)\)

  3. \(t_{17}\)

  4. \(\chi^2_{17}\)

For \(\mathrm{SSE}/\sigma^2\):

  1. \(N(0,1)\)

  2. \(t_{17}\)

  3. \(\chi^2_{17}\)

  4. \(F_{3,17}\)

B for \(\hat{\boldsymbol{\beta}}\) and C for \(\mathrm{SSE}/\sigma^2\):

\[ \hat{\boldsymbol{\beta}}\sim N_3\!\left(\boldsymbol{\beta},\sigma^2(\mathbf{X}^\top\mathbf{X})^{-1}\right), \qquad \frac{\mathrm{SSE}}{\sigma^2}\sim\chi^2_{17}. \]

The residual degrees of freedom are \(n-p=20-3=17\).

(b) [2]

Compute

\[ \hat{\sigma}^2 = \frac{\mathrm{SSE}}{n-p} = \underline{\hspace{2cm}}. \]

\[ \hat\sigma^2=\frac{34}{20-3}=\boxed{2.000}. \]

(c) [4]

For testing

\[ H_0:\beta_1=0 \qquad\text{versus}\qquad H_1:\beta_1\neq0, \]

where \(\beta_1\) is the second entry of \(\boldsymbol{\beta}\), fill in the following. Use significance level \(\alpha=0.05\).

\[ \mathrm{SE}(\hat{\beta}_1)=\underline{\hspace{2cm}}, \qquad t=\underline{\hspace{2cm}}, \]

and choose the decision:

  1. Reject \(H_0\)

  2. Fail to reject \(H_0\)

Use the second diagonal entry of \((\mathbf{X}^\top\mathbf{X})^{-1}\), which is \(0.05\):

\[ \mathrm{SE}(\hat\beta_1)=\sqrt{2(0.05)}=0.3162278\approx\boxed{0.316}, \]

\[ t=\frac{1.5-0}{\sqrt{0.1}}=4.7434165\approx\boxed{4.743}. \]

Since \(|t|>2.110\), choose A: Reject \(H_0\) at the 5% significance level.

(d) [3]

Which of the following is the 95% confidence interval for \(\beta_1\)?

  1. \((0.8329,\;2.1671)\)

  2. \((1.1838,\;1.8162)\)

  3. \((-0.6671,\;0.6671)\)

  4. \((0,\;3.0000)\)

A. Using the supplied critical value,

\[ 1.5\pm2.110\sqrt{0.1} =(0.832759,\;2.167241) \approx\boxed{(0.833,\;2.167)}. \]

These endpoints agree with option A at the requested three decimal places.