Unit 8: Descriptive Statistics and Time Series

DEMGN832 — Research Methodology 9 min read

I. Orientation: Describing Data Across Units and Time

Descriptive statistics organize, summarize, and present observed data without making probabilistic generalizations beyond those observations. Central tendency, dispersion, and distribution describe a dataset; index numbers compare related values; time-series methods analyze observations arranged chronologically.

  • Purpose: Convert raw observations into meaningful summaries for comparison, interpretation, planning, and research reporting.
  • Types of data:
    • Ungrouped data: Individual observations, such as 12, 15, 18, 20.
    • Grouped data: Observations arranged into classes with corresponding frequencies.
    • Cross-sectional data: Measurements on several units at one time.
    • Time-series data: Measurements of one variable over successive periods.
  • Measurement conventions:
    • Central tendency and dispersion must be appropriate to the nominal, ordinal, interval, or ratio scale used.
    • Statistics should be reported with units; variance has squared units, while mean and standard deviation retain the original unit.
  • Interpretive principle: No single descriptive measure is sufficient; an average should normally be considered together with variability and distributional shape.
  • Data-quality assumption: Results depend on accurate observations, consistent definitions, comparable periods, and suitable classification.

II. Measures of Central Tendency — Locating the Typical Value

A measure of central tendency identifies a central or representative value around which observations tend to cluster. The principal measures are arithmetic mean, median, and mode.

A. Measures of central tendency for ungrouped and ungrouped data

Central tendency is calculated differently for individual observations and frequency-grouped observations.

  • Arithmetic mean for ungrouped data: The sum of all values divided by the number of observations.
TEXT
x̄ = Σx / n

Here, is the sample mean, Σx is the sum of observations, and n is the number of observations.

  • Grouped-data mean: Each value or class midpoint is weighted by its frequency.
TEXT
x̄ = Σ(fx) / Σf

Here, f is frequency, x is a value or class midpoint, and Σf = N is total frequency. For continuous classes, the midpoint is (lower limit + upper limit)/2.

  • Weighted mean: Used when observations have unequal importance.
TEXT
x̄w = Σ(wx) / Σw

Here, w is the weight and x̄w is the weighted mean.

  • Median for ungrouped data: The middle value after arranging observations in ascending order.

    • If n is odd, its position is (n + 1)/2.
    • If n is even, it is the mean of the values in positions n/2 and (n/2) + 1.
  • Median for continuous grouped data:

TEXT
Median = L + [(N/2 − C) / f]h

Here, L is the lower class boundary of the median class, C is cumulative frequency before it, f is its frequency, and h is class width.

  • Mode: The most frequently occurring value. For grouped data, the modal class has the highest frequency.
TEXT
Mode = L + [(f1 − f0) / (2f1 − f0 − f2)]h

Here, f1 is modal-class frequency, while f0 and f2 are preceding and succeeding class frequencies.

  • Worked example: For 2, 3, 3, 5, 7, the mean is 20/5 = 4, the median is 3, and the mode is 3.

B. Selection, Applications, and Limitations

The appropriate average depends on data type, distributional shape, and research purpose.

  1. Mean versus median:

    • Mean: Uses every observation and supports algebraic analysis, but is strongly affected by extreme values.
    • Median: Resists outliers and suits skewed income or property-price data, but ignores the magnitudes of most observations.
  2. Mode versus other averages:

    • Mode: Works with nominal data and identifies the most common size or category.
    • Limitation: It may be absent, unstable, or non-unique in bimodal and multimodal distributions.
  • Empirical relationship: In a moderately skewed distribution:
TEXT
Mode ≈ 3(Median) − 2(Mean)

This is an approximation, not a universal identity.

III. Dispersion — Measuring Variability

Dispersion describes how far observations spread around a central value. Two datasets can have equal means but substantially different degrees of consistency.

A. Dispersion

Measures of dispersion may be absolute, expressed in data units, or relative, expressed as ratios or percentages.

  • Range: Difference between the largest and smallest observations.
TEXT
Range = Xmax − Xmin
Coefficient of range = (Xmax − Xmin) / (Xmax + Xmin)

It is simple but depends only on two extreme values.

  • Quartile deviation: Half the distance between the first and third quartiles.
TEXT
QD = (Q3 − Q1) / 2

It is useful for skewed or open-ended distributions but excludes half the observations.

  • Mean deviation: Mean of the absolute deviations from a selected central value A.
TEXT
MD = Σ|x − A| / n

For grouped data, the numerator becomes Σf|x − A|.

  • Variance and standard deviation:
TEXT
Population variance: σ² = Σ(x − μ)² / N
Sample variance: s² = Σ(x − x̄)² / (n − 1)
Standard deviation: σ = √σ² or s = √s²

Here, μ is population mean, N is population size, and n − 1 provides the sample degrees-of-freedom correction.

  • Coefficient of variation:
TEXT
CV = (s / x̄) × 100

CV compares relative variability across datasets with different means or units; a lower CV indicates greater relative consistency.

  • Worked example: For 2, 4, 6, the population mean is 4, variance is (4 + 0 + 4)/3 = 2.67, and standard deviation is approximately 1.63.

B. Applications and Limitations

Dispersion supports judgments about reliability, uniformity, risk, and representativeness.

  • Applications: Standard deviation is used in quality control, financial volatility, normal-distribution analysis, and inferential statistics.
  • Interpretation: A small standard deviation indicates clustering near the mean; a large value indicates wider spread.
  • Comparability: Absolute measures should compare datasets in the same units, whereas CV is suitable for proportional comparison.
  • Limitation: Variability measures can conceal clusters, gaps, or asymmetry; graphical and distributional analysis remains necessary.

IV. Distribution — Frequency, Shape, and Position

A statistical distribution shows how observations or frequencies are arranged across possible values or class intervals.

A. Distribution

Distributional analysis examines frequency patterns, symmetry, concentration, and departures from a typical shape.

  • Frequency distribution: Lists each value or class together with its frequency; relative frequency is f/N, while cumulative frequency is the running total.
  • Discrete distribution: Contains separate countable values, such as number of children.
  • Continuous distribution: Uses intervals for measurements such as height or income; classes should be mutually exclusive and exhaustive.
  • Symmetry and skewness:
    • Symmetric: Mean, median, and mode generally coincide.
    • Positive skew: Long right tail, commonly giving mean > median > mode.
    • Negative skew: Long left tail, commonly giving mean < median < mode.
TEXT
Pearson’s skewness = 3(x̄ − Median) / s

Positive values indicate right skew and negative values indicate left skew.

  • Kurtosis: Describes tail weight and peak characteristics relative to the normal distribution; distributions may be leptokurtic, mesokurtic, or platykurtic.
  • Normal distribution: A symmetric, bell-shaped theoretical distribution determined by μ and σ; approximately 68%, 95%, and 99.7% of observations lie within one, two, and three standard deviations of μ.

B. Presentation and Interpretation

A distribution should be presented in a form that reveals patterns without distorting frequencies.

  • Tabular presentation: Frequency tables efficiently summarize large datasets, although broad classes may hide detail.
  • Graphical presentation: Histograms represent continuous classes; frequency polygons show shape; ogives display cumulative frequencies and locate medians or quartiles.
  • Research significance: Shape affects the choice of average, variability measure, transformation, and statistical procedure.
  • Limitation: Similar means and standard deviations do not guarantee identical distributions.

V. Index Number — Measuring Relative Change

An index number expresses the relative change in a variable or group of variables, usually taking a selected base period as 100.

A. Index number

Index numbers summarize changes in prices, quantities, values, production, or living costs.

  • Price relative:
TEXT
Price relative = (p1 / p0) × 100

Here, p0 is base-period price and p1 is current-period price.

  • Simple aggregative price index:
TEXT
P01 = (Σp1 / Σp0) × 100
  • Weighted indices:
TEXT
Laspeyres: PL = [Σ(p1q0) / Σ(p0q0)] × 100
Paasche: PP = [Σ(p1q1) / Σ(p0q1)] × 100
Fisher: PF = √(PL × PP)

Here, q0 and q1 are base- and current-period quantities. Laspeyres uses base weights; Paasche uses current weights.

  • Value index:
TEXT
V01 = [Σ(p1q1) / Σ(p0q0)] × 100
  • Worked example: If a price rises from ₹50 to ₹60, its price relative is (60/50) × 100 = 120, indicating a 20% rise.

B. Construction, Uses, and Limitations

A valid index requires a representative base year, suitable items, reliable quotations, and defensible weights.

  • Uses: Index numbers measure inflation, deflate monetary series, adjust wages, compare output, and support policy decisions.
  • Consistency tests: Fisher’s index satisfies the time-reversal and factor-reversal tests under consistent ratio notation.
  • Base-year issue: An abnormal or outdated base produces misleading comparisons; chain-base indices can incorporate changing structures.
  • Limitations: Substitution, quality change, new products, sampling choices, and regional differences can bias an index.

VI. Time Series Analysis — Studying Change Through Time

A time series is a sequence of observations recorded at regular chronological intervals. Analysis seeks to identify systematic movements and separate them from irregular variation.

A. Time series analysis

Classical time-series analysis decomposes observed values into trend, seasonal, cyclical, and irregular components.

  • Trend (T): Long-term upward or downward movement caused by population, technology, productivity, or structural change.
  • Seasonal variation (S): Regular within-year movement associated with climate, customs, holidays, or institutional schedules.
  • Cyclical variation (C): Multi-year expansion and contraction associated with business or economic cycles.
  • Irregular variation (I): Unpredictable effects of strikes, disasters, policy shocks, or other exceptional events.
  • Decomposition models:
TEXT
Additive model: Y = T + S + C + I
Multiplicative model: Y = T × S × C × I

Here, Y is the observed value. The additive model assumes constant component magnitudes; the multiplicative model assumes proportional effects.

  • Trend measurement:

    • Moving averages: Smooth short-term fluctuations by averaging a fixed number of consecutive observations.
    • Least squares: Fits a trend such as Ŷ = a + bt, where a is the estimated level, b is change per period, and t denotes time.
  • Seasonal indices: Estimated through methods such as seasonal averages or ratio-to-moving-average; quarterly indices average 100 and total 400.

  • Forecasting principle: Future values are projected from established patterns, provided the underlying structure remains reasonably stable.

B. Applications and Limitations

Time-series methods help explain past movement and produce conditional forecasts for planning.

  • Applications: Sales forecasting, budgeting, inventory control, economic monitoring, staffing, and evaluation of policy effects.
  • Chronological dependence: Observations are not freely interchangeable because order and autocorrelation carry information.
  • Comparability: Definitions, price levels, geographic coverage, and observation intervals must remain consistent over time.
  • Limitations: Structural breaks, missing values, short series, changing seasonality, and unprecedented shocks reduce forecast accuracy; trend projection alone does not establish causation.