Skip to the content.

Exam PA, worked in the open

These notes are an open-source record of preparing for the Society of Actuaries Predictive Analytics exam, written for the June 2019 sitting. The written notes are a PDF and a OneNote package dated June 30, 2019, covering exam content, strategy, modeling methods, and code snippets. Exam PA was a five-hour computer-based project. A candidate received a business problem and a data set, then had to explore the data, fit and compare models in R, and write a report a non-modeler could use. The modules behind that sitting covered data types and graphics, principal components and clustering, generalized linear models, trees and random forests, and how to validate a model and explain it. The notes come from a first attempt that scored 7, about 100 hours of study, taken in the same window as a first sitting of STAM.

The code uses the exam stack. dplyr and ggplot2 handle cleaning and graphics. caret builds stratified train and test splits. rpart fits trees. glm, and glmnet on the student-success solution, fit generalized linear models. Factor coding is part of the model. Reference levels move to the most common category so the intercept is a real baseline. Thin levels are collapsed before they become dummy variables. caret::dummyVars is run at full rank when a factor is expanded, so the design matrix does not pick up a redundant column.

Hospital readmissions is a binary classification. The target is Readmission.Status. Three submissions clean length of stay, age, emergency-room use, and HCC risk score, then fit binomial GLMs. The last submission compares logit, probit, cauchit, and complementary log-log links, with a Gender-by-Race interaction as a worked example. Discrimination is read from an ROC curve and AUC. A later task prices errors at a probability cutoff of 0.075 and builds a confusion matrix, which is the right comparison when a missed readmission costs more than a false alarm.

The December 2018 Miners Union practice predicts injury counts, with employee hours as exposure. That sitting used a less structured project format than June 2019. Poisson trees go through rpart’s poisson method, with cbind(EMP_HRS_TOTAL/2000, NUM_INJURIES) on the left-hand side, then pruning on cross-validated error. The Poisson GLM uses a log link and an offset of log(hours/2000), so the coefficients are log injury rates and a response-scale prediction is an expected count. Fit is scored with RMSE, MAE, and a Poisson loglikelihood that replaces nonpositive tree predictions before the log is taken.

Student academic success is the first form of the June 2019 sample project, before the report structure was tightened. About 585 students are flagged pass or fail from final grade G3, with pass at 10 or above. Period grades G1 and G2, and absences, are dropped so they cannot leak the outcome. The comparison is a classification tree, a caret random forest, and a binomial logit GLM. The solution file also fits an L1-penalized logistic regression in glmnet (alpha = 1), with the penalty chosen by cross-validation. The syllabus has moved since 2019. Check the SOA Exam PA page before using these notes for a current sitting.