Outliers are data points that significantly differ from the rest of a dataset. They may be unusually high or low values that lie far outside the expected range of variation. In statistics and Lean Six Sigma, identifying outliers is essential for understanding process stability, detecting measurement errors, and improving data quality.
The concept of outliers dates back to early statistical analysis, with Sir Francis Galton and Karl Pearson developing methods to detect abnormal values. Outliers can arise from random variation, special causes, or data recording errors.
In Lean Six Sigma, detecting outliers helps distinguish between common cause variation (normal process fluctuation) and special cause variation (specific, correctable issues). Recognising this distinction supports effective problem-solving in the Analyse phase of DMAIC.
Impact: Outliers can distort averages, standard deviations, and regression models if not properly handled.
Rule of Thumb (Normal Distribution):
\(
|z| = \left|\frac{x – \bar{x}}{s}\right| > 3
\)
Where:
A value with a z-score greater than 3 (or less than -3) is often considered an outlier.
Example:
If the average part weight is 50 g with a standard deviation of 2 g, any measurement above 56 g or below 44 g may be treated as an outlier and investigated for root cause.
Identifying and managing outliers is essential for maintaining data accuracy and process control. In Lean Six Sigma, understanding whether an outlier reflects random noise or a true process issue ensures that improvement efforts are focused and effective.
Removing invalid outliers improves the reliability of statistical analyses, while investigating valid ones often leads to breakthrough insights.