---
title: "The non-parametric Bootstrap"
canonical: "https://modelassist.epixanalytics.com/space/EA/26575405/The%20non-parametric%20Bootstrap"
format: markdown
---
The non-parametric Bootstrap is used to estimate a parameter or parameters of a population or probability distribution from a set of observations {x<sub>i</sub>} where we don't wish to make a guess of the distributional form (e.g. Normal, Gamma, lognormal). The non-parametric Bootstrap has three stages:

 

1. Estimate the population (or probability) distribution from the data set;
2. Simulate the sampling from the population distribution that led to the set of observations {x<sub>i</sub>};
3. For each sampling, calculate the sample statistic of interest.

 

### 1. Estimate the distribution from the data

<span style="color: #000000">For the non-parametric Bootstrap, we simply use the frequency distribution of the </span><span style="color: #000000">*n*</span><span style="color: #000000"> data values as our best guess of the population or probability) distribution. In other words the list of values {x</span><sub><span style="color: #000000">i</span></sub><span style="color: #000000">} is assumed to be our population distribution. Clearly this will be an increasingly poor estimation the fewer the observations.</span>

 

### 2. Simulate the data collection

<span style="color: #000000">We now perform a Monte Carlo experiment to replicate the process via which we acquired the data set. In other words, we take a sample at random from the estimated population distribution. For the non-parametric Bootstrap this equates to randomly picking any of {x</span><sub><span style="color: #000000">i</span></sub><span style="color: #000000">} as a possible value for </span><span style="color: #000000">*each*</span><span style="color: #000000"> of the data points we actually observed. A simulated set of </span><span style="color: #000000">*n*</span><span style="color: #000000"> observations is called a Bootstrap replicate.</span>

 

### 3. Calculate the sample statistic

We now run a large number of iterations, each one generating a new Bootstrap replicate, and for each Bootstrap replicate we calculate the sample estimate of the statistic in question.  If we are interested in the standard deviation, for example, of the population, we simply calculate the sample standard deviation ([STDEV](https://epixanalytics.atlassian.net/wiki/spaces/EA/pages/26575570/) in Excel) of the Bootstrap replicates.

 

In summary, the non-parametric Bootstrap proceeds as follows:

 

- Collect the data set of *n* samples {*x*<sub>*1*</sub>, …*x*<sub>*n*</sub>}
- Create *B* Bootstrap samples {*x*<sub>*1*</sub>*, …*x*<sub>*n*</sub>*} where each *x*<sub>*i*</sub>* is a random sample with replacement from {*x*<sub>*1*</sub>, …*x*<sub>*n*</sub>}
- For each Bootstrap replicate {*x*<sub>*1*</sub>*, …*x*<sub>*n*</sub>*} calculate the required statistic
  > Macro (mathinline)

. The distribution of these *B* estimates of *q* represents the Bootstrap estimate of uncertainty about the true value of *q*.

 

 

<span style="color: #000000">**Example**</span><span style="color: #000000">:</span>

[Estimation population mean, standard deviation and percentile range for a continuous variable](https://epixanalytics.atlassian.net/wiki/spaces/EA/pages/26575409/)

 

 

 

---