Skip to the content.

Level 3Advanced

Exam PA datasets, in a Shiny app

This repository is an R Shiny explorer for the datasets in the ExamPAData package, written as practice for the Society of Actuaries Predictive Analytics exam. The app asks for a dataset, looks up whether the problem is regression or classification, and trains a short H2O AutoML run limited to distributed random forest and gradient boosting. The screen then shows performance metrics and a plot. Beside that app, the same repository holds a modeling report that predicts a numeric Amount on a table of unnamed features and scores a holdout file with mean absolute error.

It is for Exam PA candidates who want the dataset list, the target variable, and a train action in one place, and for R users who have already fit a regression or a classifier and want that workflow inside Shiny. The install notes cover Java, the h2o package, and the other libraries the app loads: shiny, ExamPAData, caret, dplyr, and ggplot2. The report uses a wider stack, including xgboost, glmnet, and caretEnsemble.

Difficulty: advanced. On the SamWiki ladder this is Level 3, advanced. The app starts an H2O cluster, converts the selected frame, and runs AutoML with a small model cap and two-fold cross-validation. The modeling report goes further along the same problem: a log transform of a right-skewed Amount, indicators for point masses in the target, a baseline linear model, generalized linear models, and a tuned xgboost scored on MAE. Following either piece takes a course in supervised learning and a working R session with Java available for H2O.

The write-up below is the app as it was recorded: the Java and h2o install, ui.R, server.R, and the steps to run it. Fourteen ExamPAData sets are wired to a target, from customer_phone_calls (purchase) and exam_pa_titanic (survived) through health_insurance (charges) and boston (medv). The longer fit is Modeling Project.Rmd, with a rendered Modeling Project.pdf. Saved pieces of that fit are on GitHub as train_y.RDS and xgb_test.RDS, next to test.txt and sample.txt. Ode to Generalized Linear Models is a short verse on the same family of models the report fits.

What you run

The app, as written

The notes below are the original write-up for the Shiny app, kept in full.

Modeling-Project

Below is a complete solution for creating an R Shiny app that allows users to explore and model datasets from the ExamPAData package. The app will enable users to select a dataset, choose an appropriate machine learning model (for regression or classification), and view performance metrics and visualizations. I’ll provide the code for the two main Shiny files: ui.R and server.R, along with instructions on how to set up and run the app.

🚀 How to Install h2o and Java to Run Your Shiny App

If you want to run a Shiny app that uses h2o, you must install Java and h2o on your system. Follow these steps based on your operating system:


1️⃣ Install Java (Required for h2o)

h2o requires Java 8 or later. Here’s how to install it:

🔵 Windows

  1. Download Java:
    • Go to: Java Downloads
    • Download and install the latest version of Java.
  2. Check if Java is installed:
    Open Command Prompt (Win + R, type cmd, and press Enter), then run:
    java -version
    
    • ✅ If you see something like java version "1.8.0_301", Java is installed!
    • ❌ If you get “Java not recognized”, restart your computer or reinstall Java.

🟢 Mac (MacOS)

  1. Install Java using Homebrew (recommended):
    brew install openjdk
    
  2. Check if Java is installed:
    java -version
    

    If it shows a Java version, you’re good to go!


🟠 Linux (Ubuntu/Debian)

  1. Install Java:
    sudo apt update
    sudo apt install default-jre
    
  2. Check if Java is installed:
    java -version
    

2️⃣ Install h2o in R

Once Java is installed, install h2o in R:

# Install dependencies
install.packages(c("RCurl", "jsonlite"))

# Install h2o
install.packages("h2o")

# Load h2o
library(h2o)

# Start h2o
h2o.init()

Common Issues & Fixes


🎯 Summary

  1. Install Java (java -version should work).
  2. Install h2o in R (install.packages("h2o")).
  3. Run h2o.init() to start the h2o engine.

R Shiny App Files

ui.R

This file defines the user interface of the Shiny app.

library(shiny)

fluidPage(
  titlePanel("ExamPAData Explorer: Predictive Analytics Tool"),
  
  sidebarLayout(
    sidebarPanel(
      selectInput("dataset", "Choose a dataset:", 
                  choices = c("customer_phone_calls", "patient_length_of_stay", "patient_num_labs", 
                              "actuary_salaries", "june_pa", "customer_value", "exam_pa_titanic", 
                              "apartment_apps", "health_insurance", "student_success", 
                              "readmission", "auto_claim", "boston", "bank_loans")),
      
      selectInput("model", "Choose a model:", choices = NULL)  # Dynamically updated based on dataset
    ),
    
    mainPanel(
      verbatimTextOutput("metrics"),
      plotOutput("plot")
    )
  )
)

server.R

This file contains the server logic to handle user inputs, train models, and generate outputs.

library(shiny)
library(ExamPAData)
library(caret)
library(dplyr)
library(ggplot2)
library(h2o)

# Map each dataset to its target variable
dataset_info <- list(
  customer_phone_calls  = list(type = "classification", target = "purchase"),
  patient_length_of_stay= list(type = "regression",    target = "days"),
  patient_num_labs      = list(type = "regression",    target = "num_labs"),
  actuary_salaries      = list(type = "regression",    target = "salary"),
  june_pa               = list(type = "regression",    target = "CLM_AMT"),
  customer_value        = list(type = "regression",    target = "score"),
  exam_pa_titanic       = list(type = "classification",target = "survived"),
  apartment_apps        = list(type = "classification",target = "apartment_apps"),
  health_insurance      = list(type = "regression",    target = "charges"),
  student_success       = list(type = "classification",target = "G3"),
  readmission           = list(type = "classification",target = "Readmission.Status"),
  auto_claim            = list(type = "regression",    target = "CLM_AMT"),
  boston                = list(type = "regression",    target = "medv"),
  bank_loans            = list(type = "classification",target = "y")
)

shinyServer(function(input, output, session) {
  
  # Initialize H2O cluster
  h2o.init(nthreads = -1, enable_assertions = FALSE)
  
  # When session ends, shut down H2O
  session$onSessionEnded(function() {
    h2o.shutdown(prompt = FALSE)
  })
  
  # Reactive expression to load the selected dataset
  selected_data <- reactive({
    req(input$dataset)
    # Load dataset from ExamPAData
    data(list = input$dataset, package = "ExamPAData", envir = .GlobalEnv)
    dataset <- get(input$dataset, envir = .GlobalEnv)
    
    # Validate dataset
    validate(
      need(!is.null(dataset) && nrow(dataset) > 0,
           paste("Error: Dataset", input$dataset, "is empty or invalid."))
    )
    
    # Make sure target variable exists
    target <- dataset_info[[input$dataset]]$target
    validate(
      need(target %in% names(dataset),
           paste("Error: Target variable", target, "not found in dataset."))
    )
    
    # Convert classification target to factor
    if (dataset_info[[input$dataset]]$type == "classification") {
      dataset[[target]] <- as.factor(dataset[[target]])
    }
    
    dataset
  })
  
  # Show target var info
  output$targetInfo <- renderUI({
    req(input$dataset)
    target_name <- dataset_info[[input$dataset]]$target
    problem_type <- dataset_info[[input$dataset]]$type
    strong(paste("Target variable:", target_name, "(",
                 ifelse(problem_type=="classification","classification","regression"), ")"))
  })
  
  # Data preview
  output$dataPreview <- renderTable({
    head(selected_data(), 10)
  })
  
  # Data summary
  output$dataSummary <- renderPrint({
    summary(selected_data())
  })
  
  # Update predictor choices whenever a dataset is selected
  observeEvent(input$dataset, {
    df <- selected_data()
    target <- dataset_info[[input$dataset]]$target
    # By default, set all columns except the target as predictors
    predictor_choices <- setdiff(names(df), target)
    updateSelectInput(session, "predictors",
                      choices = predictor_choices,
                      selected = predictor_choices)
  })
  
  # Compute baseline (benchmark) metric
  # Classification -> majority class accuracy
  # Regression -> RMSE using mean of target
  output$baselineMetric <- renderPrint({
    req(selected_data())
    df <- selected_data()
    target <- dataset_info[[input$dataset]]$target
    problem_type <- dataset_info[[input$dataset]]$type
    
    if (problem_type == "classification") {
      # majority class accuracy
      tbl <- table(df[[target]])
      majority_class <- names(tbl)[which.max(tbl)]
      baseline_preds <- rep(majority_class, nrow(df))
      actual <- as.character(df[[target]])
      acc <- mean(baseline_preds == actual)
      cat("Baseline (majority class) accuracy =", round(acc, 4))
    } else {
      # regression -> mean of target
      actual <- df[[target]]
      mu <- mean(actual, na.rm = TRUE)
      baseline_preds <- rep(mu, length(actual))
      rmse <- sqrt(mean((actual - baseline_preds)^2, na.rm = TRUE))
      cat("Baseline (mean) RMSE =", round(rmse, 4))
    }
  })
  
  # Train H2O AutoML with limited algorithms
  observeEvent(input$trainModel, {
    # Show immediate message
    output$trainLog <- renderText("Starting H2O AutoML training...")
    
    df <- selected_data()
    target <- dataset_info[[input$dataset]]$target
    problem_type <- dataset_info[[input$dataset]]$type
    
    # partition data into train/test or just use all for demonstration
    # For speed, let's train on all data, but normally you'd do a split
    # ...
    
    # convert to H2O frame
    h2o_df <- as.h2o(df)
    
    # ensure the target is factor for classification
    if (problem_type == "classification") {
      h2o_df[[target]] <- h2o.asfactor(h2o_df[[target]])
    }
    
    # set predictor variables
    predictors <- input$predictors
    if (length(predictors) < 1) {
      # fallback if none selected
      predictors <- setdiff(names(df), target)
    }
    
    # For speed, let's exclude most algorithms
    # We'll only allow DRF (Random Forest) and GBM
    # Also reduce cross-validation folds and set a small max_models
    captured_log <- capture.output({
      aml <- h2o.automl(
        x = predictors,
        y = target,
        training_frame = h2o_df,
        include_algos = c("DRF","GBM"),    # Limit to 2 algorithms
        max_models = 5,                   # limit number of models
        nfolds = 2,                       # fewer folds for speed
        seed = 42,
        sort_metric = ifelse(problem_type=="classification","AUC","RMSE"),
        keep_cross_validation_predictions = FALSE,
        keep_cross_validation_models = FALSE,
        keep_cross_validation_fold_assignment = FALSE
      )
      
      # store the leader for performance
      leader <- aml@leader
      perf <- h2o.performance(leader, h2o_df)  # evaluate on same data (demo)
      
      # Print summary so it appears in the captured log
      print(aml@leaderboard)
      cat("\n--- Leader Model Summary ---\n")
      print(leader)
      cat("\n--- Performance on training data: ---\n")
      print(perf)
    })
    
    # Render progress/log in UI
    output$trainLog <- renderText(paste(captured_log, collapse = "\n"))
    
    # Render final performance metrics in a separate output for clarity
    output$perfMetrics <- renderPrint({
      # optional final summary
      cat("Final Model Performance (see log above for details).")
    })
  })
})



How to Set Up and Run the App

1. Install Required Packages

Before running the app, ensure you have the necessary R packages installed. Run the following command in your R console:

install.packages(c("shiny", "ExamPAData", "caret", "dplyr", "ggplot2"))

2. Create the App Files

3. Run the App


App Functionality

Features

How It Works

  1. The app loads the selected dataset from ExamPAData.
  2. Based on the dataset’s problem type (regression or classification), it offers appropriate model options.
  3. It trains the chosen model on 80% of the data and tests it on the remaining 20%.
  4. The app then calculates performance metrics and generates a plot to visualize the results.

Notes

This app is a valuable tool for exploring predictive analytics and practicing for the Society of Actuaries’ Predictive Analytics Exam (Exam PA). Enjoy experimenting with the datasets and models!

Contribute

A correction to the dataset map, a clearer install note, or a result from the modeling report belongs in a pull request on this repository.

Contribute / Open a PR Source