Skip to the content.

What the notebooks fit

Sam Castillo wrote this on 13 November 2018 for the Kaggle House Prices competition, while he was a mathematics student at UMass Amherst. The tables are the Ames, Iowa sales from that contest: 1,460 homes in train.csv and 1,459 homes in test.csv. The training file has 81 columns. The modeling notebook counts 43 categorical fields and 38 numeric ones, covering size, quality, neighborhood, and the terms of the sale. Neighborhood names in the codebook sit inside the Ames city limits. The long R Markdown file is the modeling argument. The rendered HTML notebook is the exploratory pass over the same files.

Both notebooks model log(SalePrice + 1) and score root mean squared error on that scale. Raw sale price is right-skewed; the log pulls the histogram and the normal quantile plot into a shape a linear model can use. Numeric predictors with absolute skew above 0.75 are shifted by one and Box–Cox transformed with lambda 0.15. Quality fields that arrive unordered — fireplace, basement, kitchen, exterior, garage, pool, heating — are releveled None, Po, Fa, TA, Gd, Ex and then label-encoded. The garage-quality boxplots are the check: after releveling, the boxes step up with price.

Most missing cells mean the feature is absent. Pool, alley, fence, fireplace, basement, and garage quality become the level "None", and the matching areas and counts become 0. Lot frontage is filled with the median frontage in that neighborhood. Utilities is dropped because training has a single level. TotalSF adds basement, first-floor, and second-floor area. Two training sales with living area above 4,000 square feet and price under $400,000 are removed, along with one home of overall quality below 5 that sold above $200,000. The test file still carries every id, because the Kaggle submission has to. Remaining factors are dummy-coded, near-zero-variance columns are cut with caret::nearZeroVar (freqCut = 95/10), and numeric columns are centered at the median and divided by the interquartile range. When that range is 0, the scale falls back to the standard deviation.

The first fit is a linear model of log price on lot area and overall quality, with 5-fold cross-validation. The notebook puts that RMSE near 0.21. Lasso, glmnet with alpha fixed at 1, is the next real model: the notes record RMSE 0.1088 before those near-zero columns are removed and 0.0978 after. An elastic net allowed to choose alpha walks the choice back to 1, so it repeats the lasso. A caret gradient boosting machine is tuned on interaction depth, tree count, shrinkage, and minimum node size. XGBoost is then tuned in stages. The grid’s lowest training error, learning rate 0.05 with 1,000 rounds, bends against the learning curves, so the notebook sets it aside and keeps learning rate 0.01, 2,500 rounds, depth 2, minimum child weight 3, a column sample of 0.4, and every row. The reported RMSE for that booster is 0.0887, lower than the straight average of the lasso, the elastic net, and the caret GBM.

The first file sent to Kaggle, a lasso on 14 November 2018, landed near 4,000th on the public leaderboard. The XGBoost submission written in the notebook is dated 17 November 2018. A caretEnsemble stack is drafted in the source and left commented out; the blend written to disk is the unweighted average of those three base models. label_encoding.R is the small closure that maps a factor onto the sorted levels it was built with, then round-trips through saveRDS. train_final.RDS and test_final.RDS are the matrices after the pipeline.

Notebooks R

The modeling write-up, the rendered exploratory pass, and the field list. Training and test CSVs stay in the repository root next to the RDS matrices.

Modeling notebook

R Markdown

Skewness, ordered quality grades, missingness, outliers, lasso, elastic net, GBM, and the XGBoost grid. Dated 13 November 2018.

Kaggle - Advanced Regression for Housing Prices.Rmd

Exploratory notebook

HTML

Rendered

The earlier pass, already knitted. Same cleaning, with sketches of a random forest and a neural net beside the regression path.

house prices - EDA.nb.html

Encoder and codebook

R

label_encoding.R builds an integer map from the levels it has seen. data_description.txt is the Kaggle field list, including the Ames neighborhoods.

Also in the root: train.csv, test.csv, sample_submission.csv