Skip to the content.

A range, not a single guess

This is Sam Castillo’s notebook for the Society of Actuaries Predictive Analytics and Futurism Section Jupyter contest: Predicting Uncertainty: Prediction Intervals from Gradient Boosting Quantile Regression. It is student-era work in mathematics and actuarial science from the University of Massachusetts Amherst. The proofs and the fitted models stay in the notebook. This page is the front door.

A point prediction hides how much the data can support it. The notebook’s line is that fortune tellers hand you one number, and actuaries hand you an interval. Quantile regression changes the loss so a gradient boosting model can aim at a percentile. Squared error targets the mean. The pinball loss, with quantile level τ, targets any percentile. Here the band runs from the 5th percentile to the 95th.

The demonstration uses the public medical-cost file that accompanies Brett Lantz’s Machine Learning with R (also on Kaggle): 1,338 patients. Age, sex, BMI, number of children, smoker, and region predict annual charges. Those charges are right-skewed, with a mean of $13,270. One third of the rows are held out, with the random seed fixed at 42. A linear regression baseline reaches an R² of 0.75 and a mean absolute error of about $4,243. The gradient boosting model of the mean improves that to an R² of 0.809 and a mean absolute error of about $3,161.

The same quantiles do two jobs. First they put an interval on a health-insurance risk score, so two members with the same average cost can still carry very different uncertainty. Actuarial Standard of Practice No. 23 is the hook: when the data are thin, the actuary should say so, and the width of the interval is one way to say it. Second, they audit the file. Of 442 patients in the holdout set, 10 fall outside the 5th-to-95th band. Several of those rows are implausible on their face: impossible body-mass indexes, and a 12-year-old listed with three children.

Percentiles do not add the way averages do. A mean can be split by group and reassembled; a 95th percentile cannot. The notebook shows that algebra, and it keeps the two short proofs — mean squared error, then the quantile loss — next to the code that uses them.

Read the submission

The contest entry is the notebook. The rendered HTML and the .ipynb in this repository are unchanged.

Source

scikit-learn gradient boosting, a linear baseline, and the 5th and 95th percentile fits. Random seed 42.

Where to start in the notebook

Jump straight to a section of the rendered submission.

Also collected on SamWiki, the index of Sam Castillo’s public math, machine learning, and actuarial projects.