Unit 6: The Normal Distribution
I. Orientation: The probability model behind psychological measurement
The normal distribution is a continuous probability model used to describe how quantitative observations cluster around a central value. In psychological research, it is especially important because many measurements—such as test scores, reaction times after suitable transformation, and standardized abilities—are often approximately normally distributed, while sampling distributions frequently become approximately normal as sample size increases.
A. Defining properties and conventions
- Central tendency: The mean, median, and mode coincide in a perfectly normal distribution, so the centre represents the most typical score.
- Symmetry: The curve has identical left and right halves around the mean; deviations of equal size in opposite directions have equal probabilities.
- Parameters: The distribution is determined by its mean and standard deviation.
- ( \mu ): population mean.
- ( \sigma ): population standard deviation.
- Continuous measurement: A score can take any value within an interval; probability is represented by area under the curve rather than by the height of a single point.
- Standardization: Raw scores can be converted to standard scores using the (z)-score:
z = (X − μ) / σHere, (X) is an observed score, ( \mu ) is the mean, and ( \sigma ) is the standard deviation.
- Area convention: The entire area under the normal curve equals 1.00, representing total probability, or 100% of observations.
II. Introduction to normal curve — the basic probability shape
The normal curve is a smooth, bell-shaped graph showing the probability distribution of a continuous variable. Its height indicates relative density, while the area between two values indicates the probability of observations falling in that interval.
A. Introduction to normal curve
This subsection introduces the mathematical and visual meaning of the normal curve as a probability model.
- Probability density function: A normal variable (X) with mean ( \mu ) and standard deviation ( \sigma ) is represented by:
f(X) = [1 / (σ√(2π))] exp[−(X − μ)² / (2σ²)]Here, (f(X)) is density, (e) is the base of natural logarithms, and ( \pi ) is approximately 3.14159.
- Peak at the mean: The curve reaches its greatest height at (X=\mu), because values near the mean have the greatest density.
- Tails: The two tails extend indefinitely toward negative and positive values but approach the horizontal axis without touching it.
- Area and probability: The area between (X=\mu-\sigma) and (X=\mu+\sigma) is approximately 0.6826, or 68.26% of observations.
- Standard normal form: When ( \mu=0 ) and ( \sigma=1 ), the distribution is called the standard normal distribution. Its values are expressed as (z)-scores.
B. The nature of the normal curve
The nature of the curve refers to its shape, balance, continuity, and relationship with measures of location and spread.
- Bell-shaped form: Most cases gather near the centre, while increasingly extreme scores become less common; this produces a high middle and gradually declining sides.
- Unimodality: A normal curve has one mode. For example, a distribution of standardized intelligence scores has one central peak rather than separate clusters.
- Mean as balance point: The mean is the point at which the total signed deviations equal zero:
Σ(X − μ) = 0(X) denotes each observation and ( \mu ) denotes the mean.
- Inflection points: The curve changes from concave upward to concave downward at ( \mu-\sigma ) and ( \mu+\sigma ), approximately the boundaries of the central 68% region.
- Scale dependence: Changing ( \mu ) shifts the curve horizontally; changing ( \sigma ) changes its width. A larger ( \sigma ) produces a flatter, wider curve.
C. Characteristics of NPC
NPC means Normal Probability Curve, the graphical representation of the normal probability distribution.
- Asymptotic tails: The curve approaches, but never intersects, the baseline. Thus, mathematically, extreme values remain possible even though their probabilities are extremely small.
- Equal halves: The area on either side of ( \mu ) is 0.50. Therefore, 50% of cases lie below the mean and 50% lie above it.
- Fixed area divisions: The empirical rule gives useful approximate boundaries:
- Between ( \mu-\sigma ) and ( \mu+\sigma ): about 68.26%.
- Between ( \mu-2\sigma ) and ( \mu+2\sigma ): about 95.44%.
- Between ( \mu-3\sigma ) and ( \mu+3\sigma ): about 99.74%.
- Area outside two standard deviations: Approximately 4.56% lies outside ( \mu\pm2\sigma ), divided equally into about 2.28% in each tail.
- Unit-free comparison: A (z)-score expresses distance from the mean in standard-deviation units, allowing scores from different tests to be compared.
III. The normal curve as a model for sampling distributions — from populations to samples
A sampling distribution is the distribution of a statistic calculated from every possible sample of a fixed size drawn from a population. The normal curve can model such distributions, particularly through the Central Limit Theorem.
A. The normal curve as a model for sampling distributions
This subsection explains why sample statistics can be treated as approximately normal even when individual observations are not perfectly normal.
- Statistic versus parameter: A population mean ( \mu ) is a parameter, whereas a sample mean ( \bar X ) is a statistic. Repeated samples produce many different ( \bar X ) values.
- Central Limit Theorem: For independent observations with finite variance, the sampling distribution of ( \bar X ) approaches normality as sample size (n) increases, regardless of the population’s exact shape.
- Mean of the sampling distribution: The expected value of the sample mean equals the population mean:
E( X̄ ) = μ(E(\bar X)) is the expected sample mean.
- Standard error: The spread of sample means is measured by the standard error:
SE_X̄ = σ / √n(SE_{\bar X}) is the standard error, ( \sigma ) is population standard deviation, and (n) is sample size.
- Effect of sample size: Increasing (n) reduces sampling variability. If (n) changes from 25 to 100, the denominator changes from 5 to 10, so the standard error is halved.
- Standardizing a sample mean: A sample mean can be located on the sampling-distribution curve with:
z = (X̄ − μ) / (σ / √n)( \bar X ) is the observed sample mean.
B. Conditions and interpretation
The normal model is useful only when its assumptions and interpretation are appropriate.
- Independence: Observations should not influence one another. Repeated measures from the same participant require special methods because they are correlated.
- Random or representative sampling: A probability calculation applies most directly when samples are selected in a way that supports generalization to the population.
- Adequate sample size: A large sample usually improves normal approximation, but the required size depends on skewness, outliers, and the research design.
- Known versus estimated spread: The formula using ( \sigma ) applies when the population standard deviation is known. When ( \sigma ) is estimated by (s), inferential procedures commonly use the (t)-distribution.
- Meaning of a tail probability: A small tail area means that a result would be unusual if the assumed population model were true; it does not automatically prove that a psychological theory is true or false.
IV. Skewness and types — departures from symmetry
Skewness describes the direction and degree to which a distribution is asymmetrical. It is assessed by examining whether one tail extends farther or contains more extreme observations than the other.
A. Skewness
This subsection defines skewness and connects it with the relative positions of the mean, median, and mode.
- Symmetric distribution: In a normal distribution, skewness is zero and:
Mean = Median = Mode- Positive skewness: A few unusually high scores stretch the right tail. The typical ordering is:
Mode < Median < MeanFor example, income or response latency may be positively skewed when a small number of very large values occur.
- Negative skewness: A few unusually low scores stretch the left tail. The typical ordering is:
Mean < Median < ModeA difficult test may show negative skewness if most participants score high but a few score very low.
- Moment-based coefficient: One common population measure is:
γ₁ = μ₃ / σ³( \gamma_1 ) is the standardized third central moment, ( \mu_3 ) is the third central moment, and ( \sigma ) is standard deviation. Positive values indicate right skew; negative values indicate left skew.
- Interpretive caution: A small numerical skewness value does not by itself establish normality. Histograms, Q–Q plots, sample size, and the research purpose should also be considered.
B. Types of skewness
The principal types are distinguished by the direction and strength of the longer tail.
- 1. Zero skewness: The distribution is balanced around its centre, as in an ideal normal curve with ( \gamma_1=0 ).
- 2. Positive skewness: The right tail is longer; high extreme scores pull the mean upward. The mean is therefore usually greater than the median.
- 3. Negative skewness: The left tail is longer; low extreme scores pull the mean downward. The mean is therefore usually less than the median.
- Degree of skewness: Slight, moderate, or severe skewness refers to magnitude, not a different mathematical direction. Severe skewness makes normal-based procedures more sensitive to outliers and non-normal residuals.
V. Kurtosis and types — concentration and tail weight
Kurtosis describes the extent to which a distribution differs from the normal curve in its central concentration and tails. In practical data analysis, it is especially associated with tail heaviness and the likelihood of extreme observations.
A. Kurtosis
This subsection explains the measure and clarifies the distinction between ordinary and excess kurtosis.
- Moment coefficient: Kurtosis is commonly defined as:
β₂ = μ₄ / σ⁴( \beta_2 ) is kurtosis, ( \mu_4 ) is the fourth central moment, and ( \sigma ) is standard deviation.
- Normal reference: For a normal distribution, ( \beta_2=3 ). Excess kurtosis subtracts 3:
γ₂ = β₂ − 3( \gamma_2 ) is excess kurtosis, with a normal reference value of 0.
- Tail interpretation: The fourth power in ( \mu_4 ) gives extreme deviations considerable influence. A few very distant observations can greatly increase kurtosis.
- Not simply peakedness: A high peak may occur with heavy tails, but visual height alone is not a reliable measure. Kurtosis concerns how probability is distributed across the centre and tails relative to the normal model.
B. Types of kurtosis
The types compare a distribution with the normal curve using excess kurtosis.
- 1. Mesokurtic: A normal distribution is mesokurtic, with ( \beta_2=3 ) and ( \gamma_2=0 ). Its tail weight provides the reference point.
- 2. Leptokurtic: This distribution has positive excess kurtosis, ( \gamma_2>0 ), and relatively heavy tails. Psychological data may become leptokurtic when most scores cluster near the centre but a few extreme responses remain.
- 3. Platykurtic: This distribution has negative excess kurtosis, ( \gamma_2<0 ), and relatively light tails with a flatter overall form.
- Analytical consequence: Heavy tails increase the chance of outliers and can affect standard errors, significance tests, and confidence intervals. Robust methods or transformations may be appropriate when departures are substantial.
VI. Applications and limitations — using the model responsibly
The normal distribution supports probability estimation, standardization, confidence intervals, and hypothesis testing, but it is an approximation rather than a guarantee about every psychological variable.
A. Applications and limitations
This subsection identifies appropriate uses and the main boundaries of interpretation.
- Applications: Researchers use normal models to convert raw scores into (z)-scores, estimate percentile locations, model sampling means, and construct many confidence intervals.
- Measurement interpretation: A score of (z=2) lies two standard deviations above the mean; under a standard normal model, approximately 2.28% of observations lie above it.
- Limitations of bounded scales: Variables restricted to a narrow range, such as a 0–10 rating scale, may not achieve a true continuous normal shape.
- Outliers: A single extreme score can distort the mean, standard deviation, skewness, and kurtosis. Data screening should distinguish recording errors from genuine psychological variation.
- Non-normal populations: The Central Limit Theorem helps sampling means, but small samples from strongly skewed or heavy-tailed populations may still have non-normal sampling distributions.
- Substantive judgment: Statistical normality should be evaluated alongside theory, measurement design, sample size, graphical evidence, and the consequences of using a particular inferential procedure.
Did this save you a night before the exam?
LPU Notes is free, and it stays free. Ads cover part of the server bill. The rest comes out of a student's own pocket: the domain, the storage, and keeping the site up through the weeks everyone needs it at once.
The payment button didn't load. An ad blocker or a filtered network is the usual reason. to try again.
Nothing here is ever locked, and nothing unlocks. Chip in only if it was worth it. What it pays for →