---
title: 'Car Insurance'
subtitle: 'Data Preparatation'
author: "Philippe De Brouwer"
date: "`r Sys.Date()`"
bibliography: 
  - /home/philippe/Documents/science/latex.inc/library.bib
  - /home/philippe/Documents/science/latex.inc/lib_sociology.bib
nocite: |
  @coeckelbergh2020ai, @hardt2018fairness, @dubber2020oxford, @agarwal2021responsible
output:
  html_document:
    theme: cerulean
    highlight: haddock
    smart: true
    number_sections: true
    toc: true
    toc_depth: 3
    toc_float: true
    collapsed: false
    smooth_scroll: true
    fig_width: 7
    fig_height: 6
    fig_caption: true
#    df_print: paged  # this is not compatible with dfSummary, so use each time kable instead
toc-title: "Table of Contents"
---



```{r setup, include=FALSE}
knitr::opts_chunk$set(echo = TRUE, message=FALSE, warning=FALSE)

library(tidyverse)    # loads automatically with RMD 
library(dplyr)        # data wrangling
library(tidyr)        # tidy modelling
library(readr)        # read in data and parse functions
library(ggplot2)      # beautiful plots

library(caret)        # preProcess (normalising data)
library(DataExplorer) # visualization of missing values
library(naniar)       # visualization of missing values

library(forcats)      # working with factors in the tidyverse
library(gridExtra)    # put multiple ggplot graphs alongside in one box
#library(missForest)  # impute missing values with estimates
library(mice)
#library(InformationValue) # WOE and IV
library(scorecard)    # informationvalue, WOEbinning, scorecard, etc.
library(ROCR)         # AUC
library(rpart)        # fit the decision tree
library(knitr)        # include beautiful tables
library(summarytools) # frequency tables, cross-tabulations, descriptions, etc.

# set the options for {summarytools}
st_options(
  plain.ascii = FALSE, 
  style = "rmarkdown",
  dfSummary.style = "grid",
  dfSummary.valid.col = FALSE,
  dfSummary.graph.magnif = .52,
  tmp.img.dir = "/tmp",
  headings = FALSE
)

opts_chunk$set(comment=NA, 
               prompt=FALSE,
               cache=FALSE,
               echo=TRUE,
               results='asis')
```

# Executive Summary

This is part 1 on the analysis that demonstrate bias in data and the possibilities to de-bias data. In this document we show how to clean the data and prepare it in order to create models. We will prepare two data-sets:

* one with binned data -- for generalized linear models
* one with un-binned data -- for decision trees and other machine learning techniques that have their own mechanism

Then those data-frames are saved to disk, in order to be used by the next file to build to models.

# Introduction

The car insurance data is provides a wealth of information and points to discuss. What is fair? What is ethical. We will explore questions like:

 - Is it fair to use gender?
 - No? Well can we use age then? That is a good indicator proxy for driving experience, maturity, etc. and indeed has a high information value.
 - Also not? Well, let's just look at past track record then (did the person have an accident in the last 5 years)
 - But wait, that is somehow discriminatory for people that drive at least 5 years, because only they could build up such track record?
 - That leaves us wealth related parameters such as income, value of the car, etc. They do the same ... but isn't that discriminatory for poor people?
 - So, we have left things related to car use (private or commercial use, daily average travel time)
 - This model won't be that strong. The insurance company now need to increase prices ... which will impact the poor ...
 
 The dilemma is that when we believe that there is a protected feature $S$ (say income), then there are at least two groups for which using the feature will results in more desirable outcome or less desirable outcome. Leaving out the feature will make the model weaker and increase prices for all ... impacting the hardest the group that we set out to protect.
 
**Workflow**

The work-flow is inspired by: [@debrouwer2020rbook] and is prepared by Philippe De Brouwer to illustrate the concepts explained in the book.

We will approach the problem in a few phases

1. Prepare the data: this file
2. Build a naive model
  1. Build a classification model based with `CLAIM_FLAG` as dependent variable (this variable is 0 of no claim was filed)
  2. Now we need to find a price for the insurance policy. We can use a simple average or build a model. We will build a regression model to determine the fee that the customer should pay based on `CLM_AMT` (the amount claimed) of those customers that are expected to have an accident
3. Decide on Protected Features
  1. Remove protected features and proxy features
  2. repeat both models 
4. Draw conclusions


# Loading the Data

First we need to read in data and some custom functions.

```{r dataImport}
setwd("/home/philippe/Documents/courses/lectures/bias_data")

# Read in the data:
d0 = read_csv("./car_insurance_claim_data.csv") 

# Read in the functions:
source("ethics_functions.R")

# List the functions defined in this file:
tmt.env <- new.env()
sys.source("ethics_functions.R", envir = tmt.env)
utils::lsf.str(envir=tmt.env)

# Remove the temporary environment:
rm(tmt.env)
```

The data is imported from [kaggle](https://www.kaggle.com/datasets/xiaomengsun/car-insurance-claim-data), the licence is unknown and it is believed to be "public domain."

## Data Dictionary

The data contains `r nrow(d0)` rows with `r ncol(d0)` columns.

Each record (row) represents a set of attributes of an insurance company individual customer that are related to their socio-demographic profile and the insured vehicle.
The binary response variable `TARGET_FLAG` is 1 if the customer's car was in a crash, and 0 if not.
The continuous response variable `TARGET_AMT` defines the cost related to the car crash.
  
The **variables** of the data are: `r colnames(d0)`.

The **data dictionary** is as follows:

| variable name | definition | prejudice or expectation|
|---------------+------------|-------------------------|
| ID            | unique identifier | none             |
| KIDSDRIV      | number of teenager drivers| teenagers cause more accidents|
| BIRTH       | birth date   | used to derive age      |
| AGE         | age in years | young people cause more accidents, and older people too |
| HOMEKIDS    | nbr kids at home | more kids might be more distraction in the car |
| YOJ         | years on job | people that job-hop might be more prone to risk taking and accidents |
| INCOME      | annual income in USD | income correlates to reliability and responsibility |
| PARENT1     | yes or no single parent | not clear |
| HOME_VAL    | home value in USD if owner | similar to INCOME |
| MSTATUS     | marital status "yes" if married | marriage might be a sign of stability and risk aversion(?)|
| GENDER      | sex          | men cause more and more expensive accidents |
| EDUCATION   | level of the diploma | higher education correlates to safer driving |
| OCCUPATION  | categories of employment | white color workers might drive safer |
| TRAVTIME | distance to work (probably in minutes) | longer distance translates in more probability to be involved in an accident | 
| CAR_USE | commercial or private use | commercial use might be more risky|
| BLUEBOOK | resale value of the car | not clear |
| TIF      | time in force = the time with the same insurer (numeric, probably in years, minimum is "1") | longer should be better | 
| CAR_TYPE | categorical | sports cars are more prone to accidents than minivans |
| RED_CAR | "yes" if the car is red | urban legend says that red cars are more prone to accidents | 
| OLDCLAIM | total amount in USD of claims in the last 5 years | your past performance might be indicative, but note how this should interact with TIF | 
| CLM_FREQ | not claim frequency, but the number of claims in the last 5 years | past performance might be indicative
| REVOKED | "yes" if main driver's licence was revoked during the last 7 years | licence being revoked should be indicative for driving style | 
| MVR_PTS | motor vehicle record points (number) | traffic tickets should be indicative for the driving style and hence propensity to be involved in an accident | 
| CLM_AMT | if last year in a car accident, the dollar amount of the claim paid by insurer | target variable | 
| CAR_AGE | age of car in years | one might assume that drivers of older cars are less prudent | 
| CLAIM_FLAG | 0 = NO, 1 = YES | target variable |
| URBANICITY | categorical, 2 options | cities should be more dangerous |


All Variables are relevant.
Therefore we do not delete any columns at this point.


## Data quality

We were not able to contact the data provider, and close observation made us assume that some variables have been manipulated to eliminate missing values. The missing values have then been assigned to some "assumption" for example in the variable `EDUCATION` we have "`<High School`" (lower than highs school) and "`z_High School`". The latter seems to be the collection of missing values.


```{r tblhighschool, results="asis"}
EDUCATION <- factor(d0$EDUCATION);
CLAIM_FLAG <- factor(d0$CLAIM_FLAG);
summarytools::ctable(EDUCATION, CLAIM_FLAG) 
```

The table above of the `EDUCATION` variable shows that somehow the probability of being involved in an accident increases with education. Therefore we must conclude that `z_High_School` and `<High_School` are not really logical here. The people with lower than high school are for 32.4% involved in accidents and the `z_High_School` have a higher probability to be in accidents (34.6%). 

For this particular variable we might take the two classes with highest probability to be involved in an accident together as `<=High_School`. This is logically consistent, and leaves us with similar buckets sizes.

Therefore we assume that this unfortunate data manipulation is also done in other variables. For example the variable `GENDER` seems to have mixed the missing values with females. Conclusions based on small differences hence will not be reliable.

```{r tblgender, results='asis'}
GENDER <- factor(d0$GENDER)
summarytools::ctable(GENDER, CLAIM_FLAG)
```

These considerations place serious footnotes at the reliability of the data.

**Issues to consider for data quality**:

 - the partial treatment of missing values (see above)
 - the huge amount of accidents. This data set is probably a subset that has been designed to have more accidents (one would expect one in thousand and not 26% accidents)
 - the relative large amount of red cars indicates the a similar issue

## First Data Cleanup


```{r}
# Preprocess the variables into numeric / factors as necessary

# dollar values are in the format "$20,540" (string) -> convert this:
df = as.tbl(d0) %>% 
  mutate_at(c("INCOME","HOME_VAL","BLUEBOOK","OLDCLAIM", "CLM_AMT"),
            parse_number) %>% 
  mutate_at(c("EDUCATION","OCCUPATION","CAR_TYPE","URBANICITY"),
            space_to_underscore) %>% 
  mutate_at(c("PARENT1", "MSTATUS", "GENDER", "EDUCATION","OCCUPATION", "CAR_USE", "CAR_TYPE", "RED_CAR", "REVOKED", "URBANICITY"),
            as.factor) %>% 
  mutate(CLAIM_FLAG = as.factor(CLAIM_FLAG))
```

We also notice that 

 - the `ID` column will never yield useful information 
 - the `BIRTH` data is the same as `AGE` but in a less usable format.
 
Therefore, we will remove those columns before moving forward.

```{r removeInvalidCols}
df <- df %>% dplyr::select(-c(ID, BIRTH))
```


# Exploring the Data (Univariate Analysis)

The univariate summaries for the individual variables (after some cleaning) are provided below.  

```{r dfSummary1, results="asis"}
summarytools::dfSummary(df, 
          plain.ascii  = FALSE,
          style        = 'grid',
          graph.magnif = 0.85,
          varnumbers = FALSE,
          valid.col    = FALSE,
          tmp.img.dir  = "/tmp")
```

# Missing Values


## Exploring the Structure of Missing Values

The packages `DataExplorer` and `naniar` provide tools to investigate missing values.

```{r, fig.cap=c('`plot_missing()` from `DataExplorer` can show the percentaages of missing values.', '`gg_miss_upset` shows patterns in missing data by exploring joint missing values.\\label{fig:missval1}')}
DataExplorer::plot_missing(d0, title = "Percentage of missing data per variable")

naniar::gg_miss_upset(d0)
```


The correlation structure and amplitude of the missing values are illustrated in Figure \xref{fig:missval1}.

 The variable `OCCUPATION` has the highest number of missing values: `r sum(is.na(df$OCCUPATION))`. So, `r round(sum(is.na(df$OCCUPATION))/nrow(df)*100, 2)`% is missing. 
 
 

 
```{r, freqOccupation}
# To illustrate the relevance of the NA with summarytools:
freq(d0$OCCUPATION) %>%
  kable(caption='The number of missing values is almost the same as the amount of Home makers or Students.')
```

Figure \xref{fig:missval} indicates also that the data is not missing at random (MAR)^[We already argued that the missing values of education might indicate lower education, missing values in job category might be indicative for the risk too.]. 

We can consider the following solutions for missing values:

1. Do not bother if the missing value is in a column that we will not use -- in our case that could work for `RED_CAR` that is not very useful, but that one has no missing values. Most missing values are in `OCCUPATION`
1. Remove rows with missing values (in useful columns) -- we have `r round(sum(is.na(df$OCCUPATION))/nrow(df)*100, 2)`% missing in `OCCUPATION`
2. If the data is categorical, we can consider to add a specific category `missing` and study the correlations structure of the missing values.
3. input a variable
    1. if missing at random, we can use `mice`^[Multiple Imputation by Chained Equations (this method assumes missing at random).]
    2. if not missing at random, we can consider random forest methods to input values for the missing data. Good choices are: `missForest`, `missRanger` and `Amelia`.


## Imputing Missing values

To make the computations lighter we cn remove rows that miss more than one value first (these are small numbers in our case). There are `r sum(rowSums(is.na(df)) > 1)/nrow(df)*100`% rows that have more than 2 data points missing (total rows that have missing data is `r sum(rowSums(is.na(df)) >= 1)/nrow(df)*100`%).

```{r missRanger, echo=TRUE, message=FALSE, warning=FALSE, output='hide', results='hide'}
# Delete the rows that have more than once field missing:
df <- df[rowSums(is.na(df)) <= 1, ]

# Use missRanger to impute missing values:
df <- missRanger::missRanger(df)
```

**Here is how we could use mice:** Note that this code chunks is not executed (`eval=FALSE`)

```{r echo=TRUE, message=FALSE, warning=FALSE, output='hide', results='hide',  eval=FALSE}
# Impute the columns with missing values

# install.packages("mice")
#library(mice)
df_missing_cols = df[,c("AGE","INCOME","YOJ","HOME_VAL","CAR_AGE")]
df_imp_cols_tmp = mice::mice(data = df_missing_cols, m = 1, method = "pmm", maxit = 50, seed = 500)
df_imp_cols = mice::complete(df_imp_cols_tmp)
df_imp = bind_cols(df %>% select(-AGE, -INCOME, -YOJ, -HOME_VAL, -CAR_AGE), df_imp_cols)

# For the rest of the analysis we use the imputed data
df <- df_imp[complete.cases(df_imp),]
```

**Here is how we could use missForest:** Note that this code chunks is not executed (`eval=FALSE`)

```{r missingImpute, eval=FALSE}
# Input missing values via random forest estimates:
# ---
library(missForest)# impute missing values

# Note: this fails:
# d1_imp <- missForest(df, )$ximp
# While everything that worked for a data-frame will also work
# for a tibble, this is an exception, so we have 2 solutions:

# 1. coerce to data-frame
d_imp <- missForest(as.data.frame(df), )$ximp

# 2. Manual fit the randomforest and impute the missing data:
x_cc <- df %>% filter(!is.na(df$OCCUPATION))
x_nc <- df %>% filter(is.na(df$OCCUPATION))
rf <- randomForest::randomForest(OCCUPATION ~ ., data = x_cc)
x_nc$OCCUPATION <- predict(rf, x_nc)
x <- rbind(x_cc, x_nc)
```



# Further Data Exploration

## Numeric Variables

```{r contVars, results='asis'}
d_cont <- data.frame()[1:nrow(df), ]
n <- 1
for (col in colnames(df)) {
  if (!(is.factor(df[[col]]))) {
    d_cont <- cbind(d_cont, df[[col]])
    colnames(d_cont)[n] <- gsub("\"", "", deparse(col))
    n <- n + 1
  }
}
d_cont <- mutate(d_cont, CLAIM_FLAG = CLAIM_FLAG <- as.numeric(paste(df$CLAIM_FLAG)))
d_cont <- d_cont[complete.cases(d_cont),]
```


```{r correlation1, fig.cap=c("The correlation structure of the data", "The correlations of the normalised data")}
# Use the library GGaly to explore the correlation structure
# ---
GGally::ggcorr(d_cont,  method = c("everything", "pearson"),
               label = TRUE, label_round = 2, label_size = 3,
               size = 2.5,  geom="circle")
```


### Variables Related to Age

```{r dependCont}
## Visualizing the dependency for the 
p1 <- make_dependency_plot(d_cont, "AGE",      "CLM_AMT", "age",              
                           method = 'gam', q_ymax = 0.95)
p2 <- make_dependency_plot(d_cont, "HOMEKIDS", "CLM_AMT", "nbr kids at home", 
                           method = 'glm', q_ymax = 0.95)
p3 <- make_dependency_plot(d_cont, "KIDSDRIV", "CLM_AMT", "number of teenage drivers", 
                           method = 'lm', q_ymax = 0.95)
p4 <- make_dependency_plot(d_cont, "YOJ",      "CLM_AMT", "years on job",     
                           method = 'gam', q_ymax = 0.95)
grid.arrange(p1, p2, p3, p4, ncol=2)

```

```{r dep2}

p1 <- make_dependency_plot(d_cont, "AGE",      "CLAIM_FLAG", "age",              
                           method = 'gam', q_ymax = 1, q_xmax = 0.975)
p2 <- make_dependency_plot(d_cont, "HOMEKIDS", "CLAIM_FLAG", "nbr kids at home", 
                           method = 'glm', q_ymax = 1)
p3 <- make_dependency_plot(d_cont, "KIDSDRIV", "CLAIM_FLAG", "number of teenage drivers", 
                           method = 'glm', q_ymax = 1)
p4 <- make_dependency_plot(d_cont, "YOJ",      "CLAIM_FLAG", "years on job",     
                           method = 'gam', q_ymax = 1)
grid.arrange(p1, p2, p3, p4, ncol=2)
```

### Variables Related to Financial Situation

```{r dep2Wealth}
p1 <- make_dependency_plot(d_cont, "INCOME",      "CLAIM_FLAG", "income",              
                           method = 'gam', q_ymax = 1, q_xmax = 0.95)
p2 <- make_dependency_plot(d_cont, "HOME_VAL", "CLAIM_FLAG", "home value", 
                           method = 'gam', q_ymax = 1, q_xmax = 0.95)
p3 <- make_dependency_plot(d_cont, "BLUEBOOK", "CLAIM_FLAG", "value of car", 
                           method = 'gam', q_ymax = 1, q_xmax = 0.95)
p4 <- make_dependency_plot(d_cont, "TRAVTIME",      "CLAIM_FLAG", "travel time",     
                           method = 'gam', q_ymax = 1, q_xmax = 0.95)
grid.arrange(p1, p2, p3, p4, ncol=2)
```

### Driving History Related Variables

```{r dep2Hist}
p1 <- make_dependency_plot(d_cont, "CLM_FREQ",      "CLAIM_FLAG", "Claim Frequency (5yrs)",              
                           method = 'glm', q_ymax = 1, q_xmax = 0.95)
p2 <- make_dependency_plot(d_cont, "OLDCLAIM", "CLAIM_FLAG", "amount of claims in last 5 years", 
                           method = 'gam', q_ymax = 1, q_xmax = 0.95)
p3 <- make_dependency_plot(d_cont, "MVR_PTS", "CLAIM_FLAG", "motor vehicle record points ", 
                           method = 'lm', q_ymax = 1, q_xmax = 0.95)
p4 <- make_dependency_plot(d_cont, "CAR_AGE",      "CLAIM_FLAG", "car age",     
                           method = 'gam', q_ymax = 1, q_xmax = 0.95)
grid.arrange(p1, p2, p3, p4, ncol=2)
```


## Factorial Variables

```{r infValues}

df %>% 
  select_if(~class(.) == 'factor') %>% 
  iv(y = 'CLAIM_FLAG') %>%
  kable(
     caption = 'The table of all information values 
              for each categorical variable ordered in 
              decreasing order.')


```

[@debrouwer2020rbook] (pg. 359) suggests the following rule of thumb for those information values:

|IV          | predictability
|------------+---------------
|< 0.02      | not predictable
|0.02 -- 0.1 | weak
|0.1 -- 0.3  | medium
|0.13 -- 0.5 | strong
|> 0.5       | suspicious

This implies that `CARUSE`, `MSTATUS`, and `GENDER` probably are not very predictable. At least we don't expect a linear relationship.

# Data Binning

## Introduction to binning

Data binning is a data preprocessing technique that groups continuous numerical values into a smaller number of discrete "bins" or "buckets" (or groups values of categorical variables). It simplifies data, reduces noise, and can improve the performance of analytical models by converting a continuous variable into a categorical one. For example, ages can be binned into groups like "0-17," "18-64," and "65+," instead of using each individual age. 

**How data binning works**

* **Define bins**: Intervals are created to group the data. These can be of equal size, like equal age groups, or based on other criteria, such as quantiles.
* **Assign data points**: Each individual data point is placed into the bin that corresponds to its value. For example, a person age 25 would fit in the 20 to 30 bin.
* **Represent the bin**: All values within a bin are then replaced by a single representative value for that bin, such as the midpoint or mean -- unless we use the data as categorical.

**Why data binning is used**

* **Data smoothing**: It reduces the impact of minor observation errors and noise by grouping values.
* **Outlier mitigation**: Extreme values are grouped into bins, which lessens their influence on the analysis.
* **Simplifying analysis**: It might make complex data easier to analyze, visualize, and understand by simplifying the number of unique values.
* **Feature engineering**: It can convert continuous variables into categorical ones, which can be more useful for certain machine learning algorithms.
* **Improving model performance**: By reducing noise and handling outliers, binning can sometimes improve the accuracy of predictive models. 

**Considerations**

The number and size of the bins can significantly affect the results. Too few bins can oversimplify the data, while too many can make it cluttered and fail to smooth out noise.
Choosing the right binning method is important to avoid hiding important trends or creating misleading representations of the data.

**Binning in R**
Note that the package `{scorecard}` provides a function `woebin()` that will automatically propose bins for all variables. Then it can be fine-tuned with `woe_adj()` and visualized with `woe_plot()`.

## Information Value and Weight of evidence

Weight of Evidence (WoE) is a fundamental concept used primarily in credit scoring and binary classification modeling to quantify the strength and direction of the relationship between a predictor variable and the target outcome. WoE transforms categorical or binned continuous variables into continuous numeric values based on the distribution of "events" (e.g., defaults) and "non-events" (e.g., non-defaults). Mathematically, for a given bin \(i\), WoE is defined as:

\[
\text{WoE}_i = \log \left( \frac{\text{Distribution of non-events in bin } i}{\text{Distribution of events in bin } i} \right) = \log \left( \frac{P(\text{non-event}_i)}{P(\text{event}_i)} \right)
\]

where \(P(\text{event}_i)\) and \(P(\text{non-event}_i)\) represent the proportions of events and non-events in bin \(i\) relative to the total events and non-events, respectively. A positive WoE indicates that the bin has a higher concentration of non-events than events (predicting a "good" outcome), while a negative WoE suggests the opposite.

One possible way to modify WOE so that it takes values between $0$ and $1$, is Information Value (IV) for bin $i$:
\[
\text{IV_i} =  \left( P(\text{non-event}_i) - P(\text{event}_i) \right) \times \text{WoE}_i
\]


The advantage of IV is that it can meaningfully be added over all bins of a given variable. That summary measure quantifies the variable's overall predictive power and is calculated as:

\[
\text{IV} = \sum_{i} \left( P(\text{non-event}_i) - P(\text{event}_i) \right) \times \text{WoE}_i
\]

Intuitively, IV measures how well the variable's categories or bins differentiate between the two outcome groups. Variables with higher IV values are more predictive. Typical IV interpretation thresholds are:

- IV < 0.02: Not predictive
- 0.02 ≤ IV < 0.1: Weak predictive power
- 0.1 ≤ IV < 0.3: Medium predictive power
- 0.3 ≤ IV < 0.5: Strong predictive power
- IV ≥ 0.5: Suspiciously high, possibly indicating data issues or over-fitting


## Automated binning

The package `scorecard` provides tools to investigate existing bins, propose optimal bins, implement or change them,

```{r woebin, results='markup'}
scorecard::woebin(df, 'CLAIM_FLAG')
```

**2 data-sets**
For our purpose, we will create 2 datasets

* `d_fact`: a fully categorical dataset (to be used for generalised linear models)
* `d`: a dataset that holds both numerical and categorical data -- to be used for decision trees, random forest, etc.

```{r dbin}
# We create a copy of the dataset that will hold the a fully categorical
# data-set:
d_fact <- df
```


## KIDSDRIV

The variable `KIDSDRV` has an information value of `r scorecard::iv(df, "CLAIM_FLAG", x = "KIDSDRIV")[1,2]` with the existing binning.
We use the function `ctable()` from `{summarytools}` to get more insight in the variable.

```{r, results='asis'}
# KiDSDRIV is <dbl> and needs to be made categorical first
# Since there are only a limited number of possibilities, we can do this:
d_fact$KIDSDRIV <- factor(d_fact$KIDSDRIV)

# show the cross-tabulation:
summarytools::ctable(d_fact$KIDSDRIV, d_fact$CLAIM_FLAG)
```

```{r, results='markup'}
bins1 = woebin(df, y="CLAIM_FLAG", x="KIDSDRIV")
p = woebin_plot(bins1, line_value = 'woe') #+ 
  #theme(axis.text.x = element_text(angle = 90, vjust = 0.5, hjust=1))
print(p)
```

This allows us to make a decision on what bins to use. We create the bins so that 
 
 - no bin is too small (e.g. lower than 10% of the data);
 - there is a meaningful difference between the weight of evidence of the bins; and
 - if we combine variables, we decide so to do this meaningful (e.g. larger cars vs. small cars, or age bins that follow each other instead of grouping the youngest and oldest bin, etc.).
 
 **Note** that in the following sections we will leave out this exploratory phase and refer to the previous chapters about data exploration instead.

```{r binKIDSDRV}
d_fact$KIDSDRV <- if_else(df$KIDSDRIV == 0, 'none', 
                        if_else(df$KIDSDRIV == 1, '1', '>=2')) %>%
factor(levels = c('none', '1', '>=2'))
f_comment_iv(d_fact, y = 'CLAIM_FLAG', x = 'KIDSDRIV')
```
```{r, results='asis'}
summarytools::ctable(d_fact$KIDSDRV, df$CLAIM_FLAG)
```


## AGE

```{r, results='markup'}
# Manual binning can look like this:
#d_fact$AGE <- if_else(df$AGE < 30 , '<30', 
#                        if_else(df$AGE < 40, '30--40', 
#                                if_else(df$AGE < 55, '40--55', '>55'))) %>%
#    factor(level = c('<30', '30--40', '40--55', '>55'))
# f_comment_iv(d_fact, x = 'AGE', y = 'CLAIM_FLAG')

# we propose automated binning:
bins1 = woebin(d_fact, y="CLAIM_FLAG", x="AGE")
p = woebin_plot(bins1, line_value = 'woe') 
print(p)

d_fact <- woebin_ply(d_fact, bins1, to = 'bin') %>%
  dplyr::rename('AGE' = 'AGE_bin') %>%  # set the name back to what it was.
  dplyr::mutate(AGE = factor(AGE))      # make sure factor (required for automation)
```

```{r, results='asis'}
summarytools::ctable(d_fact$AGE, df$CLAIM_FLAG)
```

## HOMEKIDS


```{r, results='markup', fig.cap='The WOE per suggested bin for HOMEKIDS.'}

bins1 = woebin(d_fact, y="CLAIM_FLAG", x="HOMEKIDS")
p = woebin_plot(bins1, line_value = 'woe') 
print(p)

# woebin() suggests the same bins as our untuition guided us to
# We can now choose how to implement these bins

# (1) automatically via woebin_ply():
d_fact <- woebin_ply(d_fact, bins1, to = 'bin') %>%
  dplyr::rename('HOMEKIDS' = 'HOMEKIDS_bin')    %>%  # set the name back 
  dplyr::mutate(HOMEKIDS = factor(HOMEKIDS))      

# or (2) manually
d_fact$HOMEKIDS <- if_else(df$HOMEKIDS == 0 , '0', 
                        if_else(df$HOMEKIDS == 1, '1', '>=2')) %>%
  factor(level = c('0', '1', '>=2'))


bins1 = woebin(d_fact, y="CLAIM_FLAG", x="HOMEKIDS")
p = woebin_plot(bins1, line_value = 'woe') 
print(p)


# eventually check the IV of the binning:
f_comment_iv(d_fact, x = 'HOMEKIDS', y = 'CLAIM_FLAG')
```

```{r, results='asis'}
summarytools::ctable(d_fact$HOMEKIDS, df$CLAIM_FLAG)
```


## YOJ

"Years on job" refers to the number of years that the customer is employed in the same company.

```{r, results='markup', fig.cap=c('The suggested binning for YOJ by {scorecard}')}
# Investigate the suggested binning:
bins1 = woebin(d_fact, y="CLAIM_FLAG", x="YOJ")
p = woebin_plot(bins1, line_value = 'woe') 
print(p)
```

We notice that the pattern is a little erratic and not completely in line with intuition. There is a deep decline of claim probability for someone who is one year in the same job, but then it increases again. This is most likely a pattern that is only relevant in our own data. Therefore we will make our own binning.

```{r binYOJself, results='markup', fig.cap='Our own binning for YOJ'}
d_fact$YOJ <- if_else(df$YOJ <= 1 , '<=1', '>1') %>%
    factor(level = c('<=1', '>1'))

bins1 = woebin(d_fact, y="CLAIM_FLAG", x="YOJ") # extract the bin information
p = woebin_plot(bins1, line_value = 'woe')      # prepare the plot
print(p)                                        # shot it on screen
```

The information value of the our own binning is `r scorecard::iv(d_fact, x = 'YOJ', y = 'CLAIM_FLAG')[1,2]`.


The cross-table for `YOJ` looks now as follows:
```{r, results='asis'}
summarytools::ctable(d_fact$YOJ, df$CLAIM_FLAG)
```

## INCOME

Income is one of those variable that could have some significance. Once could expect that people who were able to study well, work hard and work themselves towards a good income are people that understand the risk of driving a car and drive more carefully. 
 
```{r binIncome1, results='markup'}
# Investigate the suggested binning:
bins1 = woebin(d_fact, y="CLAIM_FLAG", x="INCOME")
p = woebin_plot(bins1, line_value = 'woe') 
print(p)
```
While the proposed bins are rather small, the relationships with the dependent variable is almost perfectly linear and in line with intuition. Therefore, we will accept the binning as proposed by the package `scorecard`.
```{r binIncome2}
# Manual binning could look like this:
#d_fact$INCOME <- if_else(df$INCOME < 25000 , '<25K', 
#                       if_else(df$INCOME < 55000, '25--55K',
#                               if_else(df$INCOME < 75000, '55--75K', '>75K'))) %>%
#  factor(level = c('<25K', '25--55K', '55--75K', '>75K'))

d_fact <- woebin_ply(d_fact, bins1, to = 'bin') %>%
  dplyr::rename('INCOME' = 'INCOME_bin') %>%  # set the name back to what it was.
  dplyr::mutate(INCOME = factor(INCOME))  %>%    # create factor object
  dplyr::mutate(INCOME = fct_reorder(INCOME, df$INCOME)) # re-order the bins in order of increasing income

```

The information value of the `INCOME` variable is now: `r scorecard::iv(d_fact, x = 'INCOME', y = 'CLAIM_FLAG')[1,2]`; and the cross-table with `CLAIM_FLAG` is as follows.



```{r, results='asis'}
summarytools::ctable(d_fact$INCOME, df$CLAIM_FLAG)
```


## PARENT1

This variable is binary, and hence binning is obvious. We expect to see that people who are single parents have more accidents, as they will have more often distracting passengers in the car.

```{r, results='markup'}
d_fact$PARENT1 <- df$PARENT1 %>%
  factor(level = c('Yes', 'No'))
f_comment_iv(d_fact, x = 'PARENT1', y = 'CLAIM_FLAG')
```



```{r binPARENT1}
bins1 = woebin(d_fact, y="CLAIM_FLAG", x="PARENT1")
p = woebin_plot(bins1, line_value = 'woe') 
print(p)
```

```{r, results='asis'}
summarytools::ctable(d_fact$PARENT1, df$CLAIM_FLAG)
```
 
 
 
## HOME_VAL
 
The value of the house in which people live is a proxy for wealth, foresight, planning, and might be predictable for careful driving to some extent.  The value is numeric.

```{r, results='markup'}
# First try our own approach:
d_fact$HOME_VALtmp <- if_else(df$HOME_VAL < 100000 , '<100K', 
                       if_else(df$HOME_VAL < 200000, '100--200K',
                            if_else(df$HOME_VAL < 300000, '200--300K', '>300K'))) %>%
  factor(level = c('<100K', '100--200K', '200--300K', '>300K'))
f_comment_iv(d_fact, x = 'HOME_VALtmp', y = 'CLAIM_FLAG')
```

```{r homeVAL2}
# Try now the automated binning via woebin:
bins1 = woebin(d_fact, y="CLAIM_FLAG", x="HOME_VAL")
p = woebin_plot(bins1, line_value = 'woe') 
print(p)

# The IV of this binning is higher, the smallest bin is larger -> accept this one:
d_fact <- woebin_ply(d_fact, bins1, to = 'bin') %>%
  dplyr::rename('HOME_VAL' = 'HOME_VAL_bin') %>%  # set the name back to what it was.
  dplyr::mutate(HOME_VAL = factor(HOME_VAL))  %>%    # create factor object
  dplyr::mutate(HOME_VAL = fct_reorder(HOME_VAL, df$HOME_VAL)) %>% # re-order the bins in order of increasing income
  dplyr::select (-c("HOME_VALtmp"))  # remove the previously created column
```


```{r ctableHOMEVAL, results='asis'}
summarytools::ctable(d_fact$HOME_VAL, df$CLAIM_FLAG)
```

 
## MSTATUS

Is the insured person married or not. While obviously correlated to age, typically people in marriage tend to be more risk-averse and plan better. The variable is binary, hence binning is straightforward.

```{r, results='markup'}
d_fact$MSTATUS <- if_else(df$MSTATUS == 'Yes', 'Yes', 'No') %>%
  factor(level = c('Yes', 'No'))
f_comment_iv(d_fact, x = 'MSTATUS', y = 'CLAIM_FLAG')
```

```{r MSTATUSvis}
# Visualise the bins for MSTATUS:
bins1 = woebin(d_fact, y="CLAIM_FLAG", x="MSTATUS")
woebin_plot(bins1, line_value = 'woe') %>% print()
```

```{r, results='asis'}
summarytools::ctable(d_fact$MSTATUS, df$CLAIM_FLAG)
```
 
 
## GENDER

Would gender influence the propensity to claim insurance for a car accident?

```{r, results='markup'}
d_fact$GENDER <- if_else(df$GENDER == 'M', 'M', 'F') %>%
  factor(level = c('M', 'F'))
f_comment_iv(d_fact, x = 'GENDER', y = 'CLAIM_FLAG')
```

We find that in this dataset, women claim slightly more insurance, however the difference is not useful for modelling -- it is probably an example of spurious correlation. 

```{r, results='asis'}
summarytools::ctable(d_fact$GENDER, df$CLAIM_FLAG)
```

Gender is not predictive for the probability to be involved in an accident.


## EDUCATION

```{r, results='markup'}
d_fact$EDUCATION <- factor(df$EDUCATION, 
                         levels = c('<High_School', 'z_High_School', 'Bachelors', 'Masters', 'PhD'))

f_comment_iv(d_fact, x = 'EDUCATION', y = 'CLAIM_FLAG')
```

```{r visEDUCATION}
bins1 = woebin(d_fact, y="CLAIM_FLAG", x="EDUCATION")
p = woebin_plot(bins1, line_value = 'woe') 
print(p)
```

People with a PhD are a smaller group, but since the relationship is stable and logical, we keep all groups as they are.

```{r, results='asis'}
summarytools::ctable(d_fact$EDUCATION, df$CLAIM_FLAG)
```



## OCCUPATION

`OCCUPATION` is most likely predictive for the propensity to claim car insurance. It is nominal data (labels without specific order), and hence we can take together those lavels that have the most similar WOE. This allows us to trust the suggestion of `woebin()`.

```{r binOCCUPATION, results='markup'}
bins1 = woebin(d_fact, y="CLAIM_FLAG", x="OCCUPATION")
p = woebin_plot(bins1, line_value = 'woe') 

# in order to turn the x-axis labels, we need to do this for each object in the list:
for (varname in names(p)) {
  print(p[[varname]] + 
        theme(axis.text.x = element_text(angle = 45, vjust = 1.0, hjust = 1)))
}

d_fact <- woebin_ply(d_fact, bins1, to = 'bin') %>%
  dplyr::rename('OCCUPATION' = 'OCCUPATION_bin') %>%  # set the name back to what it was.
  dplyr::mutate(OCCUPATION = factor(OCCUPATION))      # make sure factor (required for automation)

f_comment_iv(d_fact, x = 'OCCUPATION', y = 'CLAIM_FLAG')
```

```{r, results='asis'}
summarytools::ctable(d_fact$OCCUPATION, df$CLAIM_FLAG)
```



## TRAVTIME

`TRAVTIME` is numeric. One can expect a linear relationship between the commute time and the probability to claim insurance.

```{r binTRAVTIME}
bins1 = woebin(d_fact, y="CLAIM_FLAG", x="TRAVTIME")
p = woebin_plot(bins1, line_value = 'woe') %>% print

d_fact <- woebin_ply(d_fact, bins1, to = 'bin') %>%
  dplyr::rename('TRAVTIME' = 'TRAVTIME_bin') %>%  # set the name back to what it was.
  dplyr::mutate(TRAVTIME = factor(TRAVTIME))      # make sure factor (required for automation)

f_comment_iv(d_fact, x = 'TRAVTIME', y = 'CLAIM_FLAG')
```


```{r, results='markup'}
d_fact$TRAVTIME <- if_else(df$TRAVTIME < 25, '<25',  
                           if_else(df$TRAVTIME < 40, '25--40', '>40')) %>%
     factor(levels = c('<25', '25--40', '>40'))
f_comment_iv(d_fact, x = 'EDUCATION', y = 'CLAIM_FLAG')
```

```{r, results='asis'}
summarytools::ctable(d_fact$TRAVTIME, df$CLAIM_FLAG)
```


## CAR_USE

This is a binary variable.

```{r, results='markup'}
d_fact$CAR_USE <- paste(df$CAR_USE) %>%
     factor()
f_comment_iv(d_fact, x = 'CAR_USE', y = 'CLAIM_FLAG')
```

```{r, results='asis'}
summarytools::ctable(d_fact$CAR_USE, df$CLAIM_FLAG)
```


## BLUEBOOK

A numeric variable. We first build a challenger binning based on data exploration.

```{r, results='markup'}


d_fact$BLUEBOOK <- if_else(df$BLUEBOOK < 10000, '<10K',  
                           if_else(df$BLUEBOOK < 15000, '10--15K', 
                                   if_else(df$BLUEBOOK < 20000, '15--20K', '>20K'))) %>%
     factor(levels = c('<10K', '10--15K', '15--20K', '>20K'))
f_comment_iv(d_fact, x = 'BLUEBOOK', y = 'CLAIM_FLAG')
```

```{r binBLUE}
d_fact$BLUEBOOK <- df$BLUEBOOK # reset back to the numeric variable
# /or/ d_fact <- d_fact %>% dplyr::mutate(BLUEBOOK = df$BLUEBOOK) 

bins = woebin(d_fact, y="CLAIM_FLAG", x="BLUEBOOK")
p = woebin_plot(bins, line_value = 'woe') %>% print
# Note that the 4th bin has an increase of WOE compared to the 3rd - that is not logical.

# We decide on a bin less and all bins wider:
custom_breaks <- list(BLUEBOOK = c(7000, 12000, 20000)) # define custom breaks

# Generate bins using custom breakpoints:
bins <- woebin(d_fact, y = "CLAIM_FLAG", x="BLUEBOOK", breaks_list = custom_breaks)

# The IV of this new binning is about 35% higher, hence we keep the new one:
d_fact <- d_fact                             %>%
  scorecard::woebin_ply(bins, to = 'bin')    %>% # re-apply the binning
  dplyr::rename('BLUEBOOK' = 'BLUEBOOK_bin') %>% # set the name back to what it was.
  dplyr::mutate(BLUEBOOK = factor(BLUEBOOK))     # coerce to factor

p = woebin_plot(bins, line_value = 'woe') %>% print # plot the new bins
f_comment_iv(d_fact, x = 'BLUEBOOK', y = 'CLAIM_FLAG')
```
While the IV of this new binning is lower, it does display a more intuitive pattern.

```{r, results='asis'}
summarytools::ctable(d_fact$BLUEBOOK, df$CLAIM_FLAG)
```

## TIF

Time in force, this is the time that the customer is on the books of our insurance company. Typically longer periods correlate to more stable behaviour, hence a better insurance risk. This phenomenon will allow us to reward loyal customers.

```{r, results='markup'}
bins = woebin(d_fact, y="CLAIM_FLAG", x="TIF")
p = woebin_plot(bins, line_value = 'woe') %>% print

# We will use our own binning for more equal bins
d_fact$TIF <- if_else(df$TIF <= 1, '1',  
                           if_else(df$TIF <= 5, '2--5', 
                                   if_else(df$TIF <= 8, '6--8', '>8'))) %>%
     factor(levels = c('1', '2--5', '6--8', '>8'))
f_comment_iv(d_fact, x = 'TIF', y = 'CLAIM_FLAG')
```

```{r, results='asis'}
summarytools::ctable(d_fact$TIF, df$CLAIM_FLAG)
```

## CAR_TYPE

```{r, results='markup'}
x <- paste(df$CAR_TYPE)
# We only merge Panel Truck and Van (too small bins + similar values)
d_fact$CAR_TYPE <- if_else(x == 'Panel_Truck' | x == 'Van' , 
                           'PTruck_Van', x) %>%
  factor()
f_comment_iv(d_fact, x = 'CAR_TYPE', y = 'CLAIM_FLAG')
```

```{r, results='asis'}
summarytools::ctable(d_fact$CAR_TYPE, df$CLAIM_FLAG)
```

## RED_CAR

```{r, results='markup'}
d_fact$RED_CAR <- paste(df$RED_CAR) %>%
     factor()
f_comment_iv(d_fact, x = 'RED_CAR', y = 'CLAIM_FLAG')
```

The urban legend that red cars would be bought by less careful drivers seems to be wrong. The IV is not significant (`r scorecard::iv(d_fact, x = 'RED_CAR', y = 'CLAIM_FLAG')[1,2]``). In our data, red cars have even every so slightly less insurance claims.

```{r, results='asis'}
summarytools::ctable(d_fact$RED_CAR, df$CLAIM_FLAG)
```

We will not use the variable `RED_CAR`. This variable might make sense for the first buyers, people who buy a red care second hand might have a very different behaviour. We don't have this information and hence cannot use the colour of the car as variable.

## OLDCLAIM

The amount of claims over the last five years, could together with the next variable (number of claims in the last five years).

```{r, results='markup'}
d_fact$OLDCLAIM <- if_else(df$OLDCLAIM == 0, '0', '>0') %>%
     factor(levels = c('0', '>0'))
f_comment_iv(d_fact, x = 'OLDCLAIM', y = 'CLAIM_FLAG')
```



```{r, results='asis'}
summarytools::ctable(d_fact$OLDCLAIM, df$CLAIM_FLAG)
```

The cumulative amount of claims in the last 5 years does not show a trend. Probability of having a claim is around 40% for people how did claim something regardless the amount (for those that didn't claim is it is less only 17%). So, we can only create two meaningful bins here (at least assuming that we are targeting the variable `CLAIM_FLAG` -- not the amount)

## CLM_FREQ
```{r clmfreq1}
bins = woebin(d_fact, y="CLAIM_FLAG", x="CLM_FREQ")
p = woebin_plot(bins, line_value = 'woe') %>% print
```

We prefer to keep people appart that did not claim anything and try to make the bins more of similar size.
```{r, results='markup'}
d_fact$CLM_FREQ <- if_else(df$CLM_FREQ == 0, '0',  
                           if_else(df$CLM_FREQ == 1, '1', 
                                   if_else(df$CLM_FREQ == 2, '2', '>2'))) %>%
     factor(levels = c('0', '1', '2', '>2'))
f_comment_iv(d_fact, x = 'CLM_FREQ', y = 'CLAIM_FLAG')
```

```{r, results='asis'}
summarytools::ctable(d_fact$CLM_FREQ, df$CLAIM_FLAG)
```
The information value is nearly the same as for the binning proposed by `summarytools::woebin()`, hence we keep our binning.

## REVOKED

Has the driver's licence been revoked? This is a binary variable and hence binning is obvious.

```{r, results='markup'}
d_fact$REVOKED <- paste(df$REVOKED) %>%
     factor()
f_comment_iv(d_fact, x = 'REVOKED', y = 'CLAIM_FLAG')
```

```{r, results='asis'}
summarytools::ctable(d_fact$REVOKED, df$CLAIM_FLAG)
```

People that have their licence revoked are about twice as likely to file a claim.

## MVR_PTS

```{r mvrpts1, results='markup'}
bins = woebin(d_fact, y="CLAIM_FLAG", x="MVR_PTS")
p = woebin_plot(bins, line_value = 'woe') %>% print
```

We try our own bins, keeping the bins of more similar size and using one bin less.
```{r mvrpts2}
d_fact$MVR_PTS <- if_else(df$MVR_PTS == 0, '0',  
                           if_else(df$MVR_PTS == 1, '1', 
                                   if_else(df$MVR_PTS == 2, '2', 
                                           if_else(df$MVR_PTS <= 4, '3--4', '>4')))) %>%
     factor(levels = c('0', '1', '2', '3--4', '>4'))
f_comment_iv(d_fact, x = 'MVR_PTS', y = 'CLAIM_FLAG')
```

```{r, results='asis'}
summarytools::ctable(d_fact$MVR_PTS, df$CLAIM_FLAG)
```

## CAR_AGE

There is one negative car age, we will simply bin it in the bucket of cars less than one year.


```{r, results='markup'}
d_fact$CAR_AGE <- if_else(df$CAR_AGE <= 1, '<=1',  
                           if_else(df$CAR_AGE <= 7, '2--7', 
                                   if_else(df$CAR_AGE <= 11, '8--11', 
                                           if_else(df$CAR_AGE <= 15, '12--15', '>=16')))) %>%
     factor(levels = c('<=1', '2--7', '8--11', '12--15', '>=16'))
f_comment_iv(d_fact, x = 'CAR_AGE', y = 'CLAIM_FLAG')
```

```{r, results='asis'}
summarytools::ctable(d_fact$CAR_AGE, df$CLAIM_FLAG)
```


## URBANICITY

Binary variable, indicating where the primary use of the car is.

```{r, results='markup'}
d_fact$URBANICITY <- paste(df$URBANICITY) %>%
     factor()
f_comment_iv(d_fact, x = 'URBANICITY', y = 'CLAIM_FLAG')
```

```{r, results='asis'}
summarytools::ctable(d_fact$URBANICITY, df$CLAIM_FLAG)
```



## Combination of Variables

We illustrate the lack of differentiation between genders for the claim amount.
```{r genderCLMAMNT}
p <- ggplot(df, aes(x=OLDCLAIM, color = GENDER)) + 
  #geom_histogram(aes(y=..density..), position="identity", alpha=0.5) +
  geom_density(alpha=0.6) + 
  scale_y_continuous(trans='log2')
p
```

###  `GENDER` and `AGE`

```{r}
d_fact$GENDER.AGE <- paste0(d_fact$GENDER, ".", d_fact$AGE) %>% factor()
f_comment_iv(d_fact, x = 'GENDER.AGE', y = 'CLAIM_FLAG')
```

```{r, results='asis'}
summarytools::ctable(d_fact$GENDER.AGE, df$CLAIM_FLAG)
```

**Average cost of accidents:**

```{r}
tibble(GENDER_AGE = paste(d_fact$GENDER.AGE), CLM_AMT = as.numeric(df$CLM_AMT)) %>%
  filter(CLM_AMT > 0) %>%
  group_by(GENDER_AGE) %>%
  summarise_at(vars(CLM_AMT), funs(mean(., na.rm=TRUE))) %>%
  kable
```

```{r avgGenderGG}
# prepare the data
data <- tibble(matrix(nrow = nrow(df), ncol = 0))
data$claim_amount <- df$CLM_AMT
data$gender       <- d_fact$GENDER
data$age          <- d_fact$AGE
```

We can first prepare the data with `group-by()` and then plot (option 1):
```{r genderggavg1, eval=FALSE}
# -- option 1:
mean_data <- group_by(data, gender, age) %>%
             summarise(claim_amount = mean(claim_amount, na.rm = TRUE))
ggplot(na.omit(mean_data), aes(x = age, y = claim_amount, colour = gender)) +
  geom_point() + geom_line()
```

Or use `stat_summary()`:
```{r avgGenderSS}
# -- option 2:
ggplot(data = data, aes(x = age, y = claim_amount, group = gender, color = gender)) + 
   stat_summary(geom = "line", fun.y = mean, size = 3)
```

Overall averages per gender:
```{r}
tibble(GENDER = paste(d_fact$GENDER), CLM_AMT = as.numeric(df$CLM_AMT)) %>%
  filter(CLM_AMT > 0) %>%
  group_by(GENDER)    %>%
  summarise_at(vars(CLM_AMT), funs(mean(., na.rm=TRUE))) %>%
  kable
```

While both males and females have higher claims for younger people, this effect is more pronounced in males. The average claim amount for a young male is `r round((3309.060/2599.308-1)*100,2 )`% higher than for the same age group in females.


So, while the young female has about 8% more probability to be involved in an accident, the claims of the male are much higher (this effect will be stronger when we would consider only the group that was involved in an accident.). 



# Output Data 

## A Completely Binned Dataset (only categorical data)

This dataset can be used for logistic regression, and some machine learning techniques.

This is `d_fact`, and we save it for later use.
```{r saveBin1}
# leave out CLM_AMNt because it is part of the target variable:
saveRDS(d_fact %>% select(-c(CLM_AMT)), './d_fact.R')
```

## Numerical values normalised

We choose to normalize so that all variables range from 0 to 1 (this is method `range` in the formula below)
```{r normNum}
# remove CLM_AMNT (because it is a 100% predictor)
df <- df %>% select (-c('CLM_AMT'))

# Normalise all numerical data
preProc_norm <- caret::preProcess(as.data.frame(df), method="range")
d_norm <- predict(preProc_norm, as.data.frame(df)) %>%
  as.data.frame() %>%   # coerces text to factor
  as.tibble()           # actually we prefer tibbles, but want the factors here

# Now replace all factorial columns with the binned ones:
for (col in colnames(d_norm)) {
    if (is.factor(d_norm[[col]])) {
      d_norm[[col]] <- d_fact[[col]]
      }
}
saveRDS(preProc_norm, './preProc_norm.R') # to convert data back and forth:
saveRDS(d_norm,       './d_norm.R')       # the data
```

## A Binary DataSet based on the factorial one

Some methods such as linear regressions, generalized linear regressions, neural networks, etc. will require all data to be numerical. So, we prepare an equivalent data set for those models.

```{r binaryDataSet1}
d_bin = data.frame(matrix(nrow = nrow(d_fact), ncol = 0)) 

for (col in colnames(d_fact)) {
    if (is.factor(d_fact[[col]])) {
      x <- expand_factor2bin(d_fact, col)  # get the expanded columns
      
      # find the largest bin and leave it out
      sums <- colSums(x)
      theCol <- which(sums == max(sums))
      x <- x[,-theCol]
      
      d_bin <- bind_cols(d_bin, x)     # add them to d_bin
    } 
}
n <- which(colnames(d_bin) == "CLAIM_FLAG.1") # this process has renamed the target variable
colnames(d_bin)[n] <- "CLAIM_FLAG"            # set back to original name
saveRDS(d_bin, './d_bin.R')                   # save the binary data-set
```


# References

<!--
https://cran.r-project.org/web/packages/DataExplorer/vignettes/dataexplorer-intro.html 
https://www.statmethods.net/advgraphs/trellis.html
http://topepo.github.io/caret/visualizations.html
https://www.analyticsvidhya.com/blog/2016/03/tutorial-powerful-packages-imputing-missing-values/ 
https://web.stanford.edu/~hastie/glmnet/glmnet_alpha.html#lin
https://stackoverflow.com/questions/35247522/error-in-cross-validation-in-glmnet-package-r-for-binomial-target-variable 
-->

## Acknowledgement

In this document we used workflow and code from: De Brouwer, Philippe J. S. 2020. The Big r-Book: From Data Science to Learning Machines and Big Data. New York: John Wiley & Sons, Ltd.

In this document we used mainly:

* Part III (p. 213--256): **Data Import**
* Part IV (p. 257--372): **Data Wrangling**

## Bibliography
