data(cars)
fit <- lm(dist ~ speed, data = cars)
coef(fit)(Intercept) speed
-17.579095 3.932409
STAT 8561 | Fall 2026
Total: 40 points
Coverage: Chapter 1 — Matrix Representation and Notation and Chapter 2 — Least Squares Estimation.
This assignment focuses on matrix representation, least squares estimation, projection geometry, and linear regression in R.
| Question | Topic | Points |
|---|---|---|
| 1 | Matrix representation | 8 |
| 2 | Least squares and projection | 12 |
| 3 | Consequences and geometry of least squares | 12 |
| 4 | Regression in R | 8 |
| Total | 40 |
Use Quarto or R Markdown. Upload both files to the Assignment 1 submission folder on iCollege:
A1_LastName_FirstName.pdf.A1_LastName_FirstName.qmd or A1_LastName_FirstName.Rmd.The PDF must be generated from the submitted source file. HTML alone does not satisfy the submission requirement.
T or F.Consider the simple linear regression model
\[ Y_i = \beta_0 + \beta_1 x_i + \varepsilon_i, \qquad i=1,\ldots,5, \]
with the following observed data:
| \(i\) | \(x_i\) | \(Y_i\) |
|---|---|---|
| 1 | 0 | 2 |
| 2 | 1 | 4 |
| 3 | 2 | 5 |
| 4 | 3 | 8 |
| 5 | 4 | 9 |
Which of the following is the correct design matrix \(\mathbf{X}\)?
\[ \begin{pmatrix} 0 & 2\\ 1 & 4\\ 2 & 5\\ 3 & 8\\ 4 & 9 \end{pmatrix} \]
\[ \begin{pmatrix} 1 & 0\\ 1 & 1\\ 1 & 2\\ 1 & 3\\ 1 & 4 \end{pmatrix} \]
\[ \begin{pmatrix} 0 & 1\\ 1 & 1\\ 2 & 1\\ 3 & 1\\ 4 & 1 \end{pmatrix} \]
\[ \begin{pmatrix} 1 & 2\\ 1 & 4\\ 1 & 5\\ 1 & 8\\ 1 & 9 \end{pmatrix} \]
B. Each row is \((1,x_i)\): the first column represents the intercept, and the second contains the predictor values.
Fill in the following table.
| Quantity | Dimension / Rank |
|---|---|
| \(\mathbf{Y}\) | ___ |
| \(\mathbf{X}\) | ___ |
| \(\boldsymbol{\beta}\) | ___ |
| \(\operatorname{rank}(\mathbf{X})\) | ___ |
| Quantity | Dimension / Rank |
|---|---|
| \(\mathbf{Y}\) | \(5\times 1\) |
| \(\mathbf{X}\) | \(5\times 2\) |
| \(\boldsymbol{\beta}\) | \(2\times 1\) |
| \(\operatorname{rank}(\mathbf{X})\) | \(2\) |
The two columns of \(\mathbf{X}\) are linearly independent because the predictor is not constant.
Which pair is correct?
\[ \mathbf{X}^\top\mathbf{X} = \begin{pmatrix} 5 & 10\\ 10 & 30 \end{pmatrix}, \qquad \mathbf{X}^\top\mathbf{Y} = \begin{pmatrix} 28\\ 74 \end{pmatrix} \]
\[ \mathbf{X}^\top\mathbf{X} = \begin{pmatrix} 5 & 5\\ 5 & 30 \end{pmatrix}, \qquad \mathbf{X}^\top\mathbf{Y} = \begin{pmatrix} 28\\ 74 \end{pmatrix} \]
\[ \mathbf{X}^\top\mathbf{X} = \begin{pmatrix} 30 & 10\\ 10 & 5 \end{pmatrix}, \qquad \mathbf{X}^\top\mathbf{Y} = \begin{pmatrix} 77\\ 28 \end{pmatrix} \]
\[ \mathbf{X}^\top\mathbf{X} = \begin{pmatrix} 5 & 10\\ 10 & 20 \end{pmatrix}, \qquad \mathbf{X}^\top\mathbf{Y} = \begin{pmatrix} 28\\ 77 \end{pmatrix} \]
A. Here \(\sum_i x_i=10\), \(\sum_i x_i^2=30\), \(\sum_i Y_i=28\), and \(\sum_i x_iY_i=74\). Thus
\[ \mathbf{X}^\top\mathbf{X}=\begin{pmatrix}5&10\\10&30\end{pmatrix}, \qquad \mathbf{X}^\top\mathbf{Y}=\begin{pmatrix}28\\74\end{pmatrix}. \]
Assume
\[ \mathbb{E}[\boldsymbol{\varepsilon}] = \mathbf{0}, \qquad \operatorname{Var}(\boldsymbol{\varepsilon})=\sigma^2\mathbf{I}_5. \]
For each statement, enter T or F.
| Statement | T/F |
|---|---|
| \(\mathbb{E}[\mathbf{Y}] = \mathbf{X}\boldsymbol{\beta}\) | ___ |
| \(\operatorname{Var}(\mathbf{Y}) = \sigma^2\mathbf{I}_5\) | ___ |
| \(\mathbb{E}[\mathbf{Y}] = \mathbf{0}\) | ___ |
| The five errors all have variance \(\sigma^2\) | ___ |
| Statement | T/F |
|---|---|
| \(\mathbb{E}[\mathbf{Y}] = \mathbf{X}\boldsymbol{\beta}\) | T |
| \(\operatorname{Var}(\mathbf{Y}) = \sigma^2\mathbf{I}_5\) | T |
| \(\mathbb{E}[\mathbf{Y}] = \mathbf{0}\) | F |
| The five errors all have variance \(\sigma^2\) | T |
Continue using the data and design matrix from Question 1.
Starting from
\[ S(\boldsymbol{\beta}) = (\mathbf{Y}-\mathbf{X}\boldsymbol{\beta})^\top (\mathbf{Y}-\mathbf{X}\boldsymbol{\beta}), \]
derive the normal equations
\[ \mathbf{X}^\top\mathbf{X}\hat{\boldsymbol{\beta}} = \mathbf{X}^\top\mathbf{Y}, \]
and then state the OLS estimator when \(\mathbf{X}\) has full column rank.
Keep your derivation concise.
Expanding,
\[ S(\boldsymbol{\beta}) = \mathbf{Y}^\top\mathbf{Y} - 2\boldsymbol{\beta}^\top\mathbf{X}^\top\mathbf{Y} + \boldsymbol{\beta}^\top\mathbf{X}^\top\mathbf{X}\boldsymbol{\beta}. \]
Differentiating,
\[ \frac{\partial S(\boldsymbol{\beta})}{\partial\boldsymbol{\beta}} = -2\mathbf{X}^\top\mathbf{Y} + 2\mathbf{X}^\top\mathbf{X}\boldsymbol{\beta}. \]
Setting the derivative equal to zero gives
\[ \mathbf{X}^\top\mathbf{X}\hat{\boldsymbol{\beta}} = \mathbf{X}^\top\mathbf{Y}. \]
Hence,
\[ \boxed{ \hat{\boldsymbol{\beta}} = (\mathbf{X}^\top\mathbf{X})^{-1} \mathbf{X}^\top\mathbf{Y} }. \]
Using the quantities from Question 1, compute
\[ \hat{\boldsymbol{\beta}} = \begin{pmatrix} \hat{\beta}_0\\ \hat{\beta}_1 \end{pmatrix}. \]
Fill in the blanks:
\[ \hat{\beta}_0 = \underline{\hspace{2cm}}, \qquad \hat{\beta}_1 = \underline{\hspace{2cm}}. \]
Hence,
\[ \hat{Y} = \underline{\hspace{2cm}} + \underline{\hspace{2cm}}x. \]
Using the data in Question 1,
\[ \hat{\boldsymbol{\beta}} =\frac{1}{50}\begin{pmatrix}30&-10\\-10&5\end{pmatrix} \begin{pmatrix}28\\74\end{pmatrix} =\begin{pmatrix}2.0\\1.8\end{pmatrix}. \]
Therefore \(\hat{\beta}_0=2.000\), \(\hat{\beta}_1=1.800\), and
\[ \boxed{\hat Y=2.0+1.8x}. \]
These are the coefficients used in Question 3.
For the least squares residual vector
\[ \mathbf{e} = \mathbf{Y}-\hat{\mathbf{Y}}, \]
which statement must be true?
\(\mathbf{X}\mathbf{e}=\mathbf{0}\)
\(\mathbf{X}^\top\mathbf{e}=\mathbf{0}\)
\(\mathbf{e}^\top\mathbf{e}=0\)
\(\mathbf{H}\mathbf{e}=\mathbf{e}\)
In one sentence, explain the geometric meaning of your choice.
Your explanation:
____________________________________________________________
B. The residual vector is orthogonal to the column space of \(\mathbf{X}\).
For the hat matrix
\[ \mathbf{H} = \mathbf{X} (\mathbf{X}^\top\mathbf{X})^{-1} \mathbf{X}^\top, \]
enter T or F.
| Statement | T/F |
|---|---|
| \(\mathbf{H}^\top=\mathbf{H}\) | ___ |
| \(\mathbf{H}^2=\mathbf{H}\) | ___ |
| \(\hat{\mathbf{Y}}=\mathbf{H}\mathbf{Y}\) | ___ |
| \(\mathbf{H}\) is an orthogonal projection matrix | ___ |
| Statement | T/F |
|---|---|
| \(\mathbf{H}^\top=\mathbf{H}\) | T |
| \(\mathbf{H}^2=\mathbf{H}\) | T |
| \(\hat{\mathbf{Y}}=\mathbf{H}\mathbf{Y}\) | T |
| \(\mathbf{H}\) is an orthogonal projection matrix | T |
Continue using the data and fitted regression model from Questions 1 and 2.
Recall that
\[ \hat{\beta}_0=2.0, \qquad \hat{\beta}_1=1.8. \]
Compute the fitted values
\[ \hat{\mathbf{Y}} = \mathbf{X}\hat{\boldsymbol{\beta}} \]
and the residual vector
\[ \mathbf{e} = \mathbf{Y}-\hat{\mathbf{Y}}. \]
Fill in the blanks:
\[ \hat{\mathbf{Y}} = \begin{pmatrix} \underline{\hspace{1cm}}\\ \underline{\hspace{1cm}}\\ \underline{\hspace{1cm}}\\ \underline{\hspace{1cm}}\\ \underline{\hspace{1cm}} \end{pmatrix}, \qquad \mathbf{e} = \begin{pmatrix} \underline{\hspace{1cm}}\\ \underline{\hspace{1cm}}\\ \underline{\hspace{1cm}}\\ \underline{\hspace{1cm}}\\ \underline{\hspace{1cm}} \end{pmatrix}. \]
\[ \hat{\mathbf{Y}} = \begin{pmatrix} 2.0\\ 3.8\\ 5.6\\ 7.4\\ 9.2 \end{pmatrix}, \qquad \mathbf{e} = \begin{pmatrix} 0\\ 0.2\\ -0.6\\ 0.6\\ -0.2 \end{pmatrix}. \]
For each statement, enter T or F.
| Statement | T/F |
|---|---|
| The residuals satisfy \(\sum_{i=1}^n e_i=0\). | ___ |
| The fitted values satisfy \(\bar{\hat{Y}}=\bar{Y}\). | ___ |
| The vector \(\mathbf{1}_n\) belongs to \(\mathcal{C}(\mathbf{X})\). | ___ |
| These properties necessarily hold for every regression model without an intercept. | ___ |
| Statement | T/F |
|---|---|
| The residuals satisfy \(\sum_{i=1}^n e_i=0\). | T |
| The fitted values satisfy \(\bar{\hat{Y}}=\bar{Y}\). | T |
| The vector \(\mathbf{1}_n\) belongs to \(\mathcal{C}(\mathbf{X})\). | T |
| These properties necessarily hold for every regression model without an intercept. | F |
Because the model contains an intercept, \(\mathbf{1}_n\) is a column of \(\mathbf{X}\). Since \(\mathbf{X}^\top\mathbf{e}=\mathbf{0}\),
\[ \mathbf{1}_n^\top\mathbf{e}=0, \]
which implies
\[ \sum_{i=1}^n e_i=0 \]
and therefore
\[ \bar{\hat{Y}}=\bar{Y}. \]
Define
\[ \mathbf{M} = \mathbf{I}_n-\mathbf{H}. \]
For each statement, enter T or F.
| Statement | T/F |
|---|---|
| \(\mathbf{e}=\mathbf{M}\mathbf{Y}\) | ___ |
| \(\mathbf{M}^\top=\mathbf{M}\) | ___ |
| \(\mathbf{M}^2=\mathbf{M}\) | ___ |
| \(\mathbf{H}\mathbf{M}=\mathbf{0}\) | ___ |
| Statement | T/F |
|---|---|
| \(\mathbf{e}=\mathbf{M}\mathbf{Y}\) | T |
| \(\mathbf{M}^\top=\mathbf{M}\) | T |
| \(\mathbf{M}^2=\mathbf{M}\) | T |
| \(\mathbf{H}\mathbf{M}=\mathbf{0}\) | T |
The matrix \(\mathbf{H}\) projects onto \(\mathcal{C}(\mathbf{X})\), whereas \(\mathbf{M}\) projects onto \(\mathcal{C}(\mathbf{X})^\perp\).
Suppose now that a design matrix \(\mathbf{X}\) does not have full column rank.
For each statement, enter T or F.
| Statement | T/F |
|---|---|
| \((\mathbf{X}^\top\mathbf{X})^{-1}\) necessarily exists. | ___ |
| The least squares coefficient vector \(\hat{\boldsymbol{\beta}}\) may not be unique. | ___ |
| The fitted vector \(\hat{\mathbf{Y}}\) is still unique. | ___ |
| The residual vector \(\mathbf{e}\) is still unique. | ___ |
| Statement | T/F |
|---|---|
| \((\mathbf{X}^\top\mathbf{X})^{-1}\) necessarily exists. | F |
| The least squares coefficient vector \(\hat{\boldsymbol{\beta}}\) may not be unique. | T |
| The fitted vector \(\hat{\mathbf{Y}}\) is still unique. | T |
| The residual vector \(\mathbf{e}\) is still unique. | T |
Even when several coefficient vectors produce the same least squares solution, the orthogonal projection of \(\mathbf{Y}\) onto \(\mathcal{C}(\mathbf{X})\) is unique.
For this question, use the built-in R dataset cars.
The variables are:
speed: speed of the car in miles per hour;dist: stopping distance in feet.Consider the model
\[ \mathrm{dist}_i = \beta_0 + \beta_1\mathrm{speed}_i + \varepsilon_i. \]
Load the data using
Fit the model using lm() and complete the table.
| Quantity | Your answer |
|---|---|
| \(\hat{\beta}_0\) | ___ |
| \(\hat{\beta}_1\) | ___ |
Which interpretation of \(\hat{\beta}_1\) is correct?
For each 1 mph increase in speed, the stopping distance increases by exactly 3.9324 feet for every car.
For each 1 mph increase in speed, the fitted mean stopping distance increases by approximately 3.9324 feet.
The average speed increases by 3.9324 mph for every extra foot of stopping distance.
A car travelling at 0 mph has a stopping distance of 3.9324 feet.
Construct \(\mathbf{Y}\) and \(\mathbf{X}\) manually and reproduce the OLS coefficients using matrix operations.
Complete the following code:
Then report
\[ \hat{\boldsymbol{\beta}} = \begin{pmatrix} \underline{\hspace{2cm}}\\ \underline{\hspace{2cm}} \end{pmatrix}. \]
[,1]
[1,] -17.579095
[2,] 3.932409
The result agrees with lm():
\[ \hat{\boldsymbol{\beta}} =\begin{pmatrix}-17.579095\\3.932409\end{pmatrix} \approx\begin{pmatrix}-17.579\\3.932\end{pmatrix}. \]