Instruction

STAT 8561 | Fall 2026
Due on December 13, 2026
Total: 100 points
Course-grade weight: 20% (see Grade Distribution)
Group size: 2-3 students
Project: Linear Regression Data Analysis and Machine Learning Comparison

Analyze a real dataset with a continuous response using multiple linear regression. Explain the findings in context, and compare predictions with one machine learning regression algorithm of your choice.

Criterion Topic Points
1 Question and data 10
2 Linear model and interpretation 25
3 Model comparison and diagnostics 20
4 Machine learning method 10
5 Fair predictive evaluation 15
6 Findings and critical discussion 10
7 Communication and reproducibility 10
Total 100
NoteInstruction

Use Quarto or R Markdown. Upload both files to the project submission area on iCollege:

  1. Your rendered PDF, named Project_Report.pdf.
  2. Its matching source file, named Project_Report.qmd or Project_Report.Rmd.

Write 8-10 pages of main text, excluding the title page, references, and appendices. List all group members and their contributions.

The submitted source and supporting files must reproduce the PDF and results. HTML alone does not satisfy the requirement.

Python or Jupyter Notebook: Discuss the format with the instructor in advance; a PDF and reproducible source are still required.

See Report requirements and submission for the full instructions. The submission folder will be announced on iCollege.

Assessment criteria

The project totals 100 points and is worth 20% of the course grade (Grade Distribution). Each weight below is a percentage of the project grade: 25% equals 25 points. A score of 90/100 contributes 18 percentage points to the course grade.

Criterion Weight Full-credit evidence (points)
Question and data 10% Clear question, observational unit, and outcome units (3); cited source and variable definitions (3); justified cleaning, missing-data handling, and dependence assessment (4).
Linear model and interpretation 25% Correct mean model and predictor coding (5); correct interpretations of at least two coefficients or contrasts and the intercept, including units and adjustment (12); correct 95% confidence intervals and uncertainty discussion (8).
Model comparison and diagnostics 20% Valid nested-model restriction and test, with assumptions addressed (6); explanation of retained coefficients across models (4); informative residual, Q-Q, and influence checks with justified responses to problems (10).
Machine learning method 10% Appropriate regression method and accurate explanation (4); documented implementation, preprocessing, and final settings (6).
Fair predictive evaluation 15% Appropriate common test set and no data leakage (6); documented training-only validation and tuning (4); correct baseline, test RMSE/MAE, and prediction plots for both models (5).
Findings and critical discussion 10% Contextual conclusions about effect size and uncertainty (4); reasoned prediction comparison and model recommendation (3); meaningful limitations and next step (3).
Communication and reproducibility 10% Organized report and readable, discussed figures and tables (3); runnable source that reproduces results with documented data, seed, and environment (5); citations and each member’s contribution (2).
Total 100% 100 points

Subitem points appear in parentheses. Partial credit reflects accuracy, completeness, and justification; omissions earn zero. Full credit does not require machine learning to outperform linear regression.

How to get started

  1. Choose a question and data. Identify a continuous outcome, its units, the observational unit, and plausible predictors. Cite the original data source.
  2. Plan the comparison. Choose one machine learning method for regression, such as a regression tree, random forest, gradient boosting, or support vector regression. Decide how to reserve test observations before fitting or tuning models.
  3. Analyze and explain. Fit and diagnose the linear regression, interpret its coefficients and uncertainty, and compare a scientifically motivated pair of nested linear models.
  4. Evaluate and communicate. Compare both methods on the same held-out observations using RMSE and MAE. Explain what each model reveals, where it performs poorly, and which you would use for your stated purpose.

A few possible questions

  • How is a property’s sale price associated with floor area and other recorded characteristics, and does a random forest improve prediction?
  • How is a building’s energy use associated with its characteristics, and does gradient boosting capture patterns missed by the linear model?
  • How is a measured environmental outcome associated with recorded conditions, and how does a regression tree compare with linear regression?

These are examples of scope; you may choose another application and another suitable regression algorithm.

NoteWhat makes a strong project?

A strong project gives a defensible linear regression analysis, clear interpretations, a fair prediction comparison, and reproducible evidence. A machine learning model does not have to outperform linear regression. Explain the result you obtain, including its limitations.

Use existing software to fit the models, and explain the method, important settings, and analysis decisions. Record 8561 as the seed for randomized procedures. In R, use set.seed(8561) and prefer ggplot2 for figures.

Next: Report requirements, assessment criteria, and submission.