Skip to the content.

What the script fits

This repository is the working script for the Zillow Prize, the Kaggle contest that asked for the error in Zillow's own home-value estimate. If Z is the Zestimate and S is the sale price, the target is logerror = log(Z) - log(S) = log(Z / S). A prediction of 0 says the Zestimate already matches the sale. The submission is one row per parcel and six columns: October, November, and December of 2016, then the same three months in 2017. The contest discussion lives at the Zillow Prize forum.

Sales in train_2016_v2.csv are left-joined to properties_2016.csv on parcelid, so every training sale keeps its property attributes. Identifier fields that arrive as numbers are recoded as factors: building quality type, FIPS, heating system, property land use, census tract and block, city, county, and unit count. The sale month is taken from transactiondate with lubridate::month and stored as a factor. The date column is then dropped. The script's own note is that month of the year has predictive power, which is why the six contest months are not one shared forecast.

make_features builds the columns that sit on top of the raw assessor file. N_value_ratio is taxvaluedollarcnt / taxamount, assessed value per dollar of tax. N_living_area_prop is calculatedfinishedsquarefeet / lotsizesquarefeet. N_tax_score is the product taxvaluedollarcnt * taxamount. N_zip_count is the number of parcels sharing regionidzip. A structure-tax deviation, abs(structuretaxvaluedollarcnt - N_Avg_structuretaxvalue) / N_Avg_structuretaxvalue, is still written in the function. The group_by(regionidcity) that would create N_Avg_structuretaxvalue is commented out, so that deviation is not an input the current script can compute. Raw census tract, zoning description, census tract and block, and assessment year are dropped inside the same function.

Columns with more than 80 percent missing values are removed from the training frame, and the same names are removed from the property file used at score time. Training rows are kept when logerror is between -0.39 and 0.4. On the multiplicative scale those bounds are exp(-0.39) ≈ 0.677 and exp(0.4) ≈ 1.492, so the fit sees sales where the Zestimate runs from about 0.68 times the sale price to about 1.49 times the sale price. The model is gbm::gbm with distribution = "gaussian", squared error on logerror, using every remaining column. The source script sets 600 trees, interaction.depth = 5, shrinkage = 0.0033, and bag.fraction = 0.8, and it uses half of detectCores(). Depth 1 would be an additive model. Depth 5 allows a tree whose splits combine up to five variables. xgboost is loaded with the other packages. The call that is fit is gbm.

Scoring builds a month proxy so the property file has the same factor the training rows used. predict.gbm is called with all 600 trees after the month is set to "10", "11", and "12", and those three vectors are written as 201610, 201611, and 201612. The 2017 columns 201710, 201711, and 201712 are filled with 0, the forecast that the log residual is zero a year later. The script writes submission_with_new_features.csv with scipen = 999 so the parcel ids stay in decimal form. That CSV is not in the repository, and neither is a public leaderboard score. The knitted notebook is an earlier pass of the same file: 200 trees, shrinkage = 0.033, and the knit stops with transactiondate not found and then an invalid -drop.column, before a model is saved.

Notebook R

The source script is the 600-tree model. The HTML file is the earlier knit, errors included.

Gradient boosted machine

R Markdown

Join, factor recodes, the four engineered columns, the 80 percent missingness screen, the log-error filter, and the 600-tree Gaussian GBM.

zillow_gbm.Rmd

Knitted notebook

HTML

Earlier run

200 trees and shrinkage 0.033. The knit reports transactiondate not found, then fails on drop.column, and does not save a model.

zillow_gbm.nb.html