---
title: "Kolmogorov-Smirnoff (K-S) Statistic"
canonical: "https://modelassist.epixanalytics.com/space/EA/26575334/Kolmogorov-Smirnoff%20(K-S)%20Statistic"
format: markdown
---
The K-S statistic *Dn* is defined as:

 

                *D*<sub>*n*</sub> = max [ | *F*<sub>*n*</sub>*(x)* - *F(x)* | ]

 

where     *D*<sub>*n*</sub> is know as the K-S distance

             *n* = total number of data points

             *F(x)* = distribution function of the fitted distribution

             *F*<sub>*n*</sub>*(x)* = *i*/*n*

             *i* = the cumulative rank of the data point

 

The K-S statistic is thus only concerned with the maximum vertical distance between the cumulative distribution function of the fitted distribution and the cumulative distribution of the data. The figure below illustrates the concept for data fitted to a Uniform(0,1) distribution.

 

![image](media://7b6c54ae-996f-4eef-92f3-d46e30e26a5a)

**Figure 1**: Illustration of a K-S distance determination for a Uniform distribution

 

### Method

·       The data are ranked in ascending order

 

·       The upper *F*<sub>*U*</sub>*(i)* and lower *F*<sub>*L*</sub>*(i)* cumulative percentiles are calculated as follows:

 

![image](media://dc599fc8-6a35-4ea6-b319-9abdaddccd2b)

 

> Macro (mathblock)

![image](media://6caf0c13-614d-479d-aced-76d1bdfa1d1b)


> Macro (mathblock)

 

where *i* = the rank of the data point and *n* = the total number of data points.

 

·       *F(x)* is calculated for the Uniform distribution (in this case *F(x)* = *x*)

 

·       The maximum distance *Di* between *F(i)* and *F(x)* is calculated for each *i*:

 

      *Di* = MAX ( ABS ( *F(x)* - *F*<sub>*L*</sub>*(i)*), ABS ( *F(x)* - *F*<sub>*U*</sub>*(i)* ))

 

where ABS (...) finds the absolute value

 

·       The maximum value of the *D*<sub>*i*</sub> distances is then the K-S distance *D*<sub>*n*</sub>:

 

      *D*<sub>*n*</sub> = MAX ( {*D*<sub>*i*</sub>} )

 

The K-S statistic is generally more useful than the [c](https://epixanalytics.atlassian.net/wiki/spaces/EA/pages/26575333/)<sup>[2](https://epixanalytics.atlassian.net/wiki/spaces/EA/pages/26575333/)</sup>[ statistic](https://epixanalytics.atlassian.net/wiki/spaces/EA/pages/26575333/) in that the data are assessed at all data points and avoids the problem of determining the number of bands to split the data into. However, its value is only determined by the one largest discrepancy and takes no account of the lack of fit across the rest of the distribution.

 

The difference between the observed distribution *F*<sub>*n*</sub>*(x)* and the theoretical fitted distribution *F(x)* at any point, say *x*<sub>*0*</sub>, itself has a distribution with a mean of zero and a standard deviation *sK-S* given by binomial theory:

 

![image](media://b0c9f272-4eae-4bf9-8cb0-0f1741c56370)

                sK-S > Macro (mathinline)



 

The size of the K-S standard deviation *s*<sub>*K-S*</sub> varies considerably over the *x*-range of a fitted distribution, and depends greatly on the type of distribution being fitted, as shown in the graphs below:

:

![image](media://5252714a-4d89-4403-a02b-7d18f14a1d96)

**Figure 2**: Standard deviation sK-S for a number of distribution types with *n* = 100

 

The position of *Dn* along the *x*-axis is more likely to occur where s<sub>K-S</sub> is greatest which, Figure 2 shows, will generally be away from the low probability tails. This insensitivity of the K-S statistic to lack of fit at the extremes of the distributions is corrected for in the [Anderson-Darling statistic](https://epixanalytics.atlassian.net/wiki/spaces/EA/pages/26575335/).

 

 

 

---