Unit 2: Graphical Representations and Frequency Distribution

PSY115 — Statistical Methods For Psychological Research 9 min read

I. Orientation: Organising Psychological Data

Statistical methods convert observations such as reaction times, test scores, anxiety ratings, or memory errors into organised numerical and visual forms. This unit is based on the principle that raw data become interpretable when they are classified, counted, displayed, and located relative to the whole distribution. Frequency distributions show how often values occur, while graphs reveal patterns, concentration, spread, and unusual observations.

  • Data: Observations collected for analysis; for example, scores of 40 students on a memory test.
  • Variable: A characteristic that can take different values, such as age, intelligence score, or number of errors.
  • Frequency: The number of observations having a particular value or belonging to a particular interval.
  • Class interval: A range used to group scores, such as 10–19 or 20–29.
  • Central principle: A table or graph must preserve the relationship between values and their frequencies without misleading scale, spacing, or classification.
  • Position measures: Quartiles, deciles, and percentiles locate an individual score within an ordered distribution.
  • Conventions: Graphs require a title, labelled axes, suitable units, an appropriate scale, and a clearly identified population or sample.

II. Graphical Representation of Data: Visual Communication

Graphical representation is the presentation of numerical information through visual forms such as bar diagrams, histograms, frequency polygons, and ogives. Its purpose is not decoration but efficient interpretation of the distribution.

A. importance of graphical representation of data

The importance of graphical representation lies in making large or complex datasets easier to understand, compare, and interpret.

  • Condensation: A graph summarises many observations in a compact form; the distribution of 500 psychological test scores can be viewed more quickly than in a 500-row data table.
  • Pattern detection: A histogram may immediately show symmetry, positive skew, negative skew, or two peaks that are difficult to see in raw scores.
  • Comparison: Two groups can be compared using adjacent bar charts or frequency polygons; for example, treatment and control groups may display different levels of depressive symptoms.
  • Communication: A labelled graph communicates findings to researchers, clinicians, and readers who may not inspect every numerical calculation.
  • Relationship identification: A scatterplot can show whether higher stress scores tend to accompany poorer sleep scores; this is a visual indication of association, not proof of causation.
  • Outlier detection: A single reaction time far above the remaining observations may indicate an unusual participant, recording error, or meaningful exceptional response.
  • Decision support: Researchers can judge whether data are suitable for methods such as the mean and standard deviation by examining distributional shape.
  • Accuracy requirement: A truncated vertical axis can exaggerate differences; a graph comparing mean scores should use a scale that does not distort the magnitude of change.

B. graphical representation of grouped and ungrouped data

The appropriate graph depends on whether individual values are displayed separately or have been placed into intervals.

  • Ungrouped data: Each distinct value is retained. A bar graph is suitable for discrete categories or counts, such as the number of participants choosing “yes,” “no,” or “uncertain.”
  • Grouped data: Values are combined into class intervals. A histogram is suitable for continuous measurements such as reaction time grouped into 100–199 ms, 200–299 ms, and 300–399 ms.
  • Bar graph: Bars are separated by spaces because categories are distinct; bar height represents frequency.
  • Histogram: Bars touch because adjacent intervals represent continuous values; area and height represent frequency when class widths are equal.
  • Frequency polygon: Class midpoints are joined by straight lines, allowing distributions from two groups to be compared on one set of axes.
  • Ogive: Cumulative frequencies are plotted against upper class boundaries, producing a curve useful for locating medians, quartiles, and percentiles.
  • Pie chart: A circle displays parts of a total; a category with frequency (f) in a total of (N) has angle:
TEXT
Angle = (f / N) × 360°

Here, (f) is the category frequency and (N) is the total frequency. Pie charts are most effective when there are few non-overlapping categories.

III. Frequency Distributions: Systematic Classification

A frequency distribution is a table that lists scores, categories, or class intervals together with the number of observations associated with each. It transforms raw data into an organised description of occurrence.

A. frequency distributions

A frequency distribution identifies how observations are distributed across values or intervals.

  • Simple frequency distribution: Each score and its frequency are listed; if scores 2, 3, and 4 occur 5, 8, and 2 times, the table records those exact counts.
  • Grouped frequency distribution: Scores are placed into intervals, such as 0–9, 10–19, and 20–29, when many distinct values make an ungrouped table unwieldy.
  • Frequency symbol: (f) represents the frequency of a score or class; the total is:
TEXT
N = Σf

Here, (N) is the total number of observations and (\Sigma f) means the sum of all frequencies.

  • Relative frequency: The proportion in a category is:
TEXT
Relative frequency = f / N

Multiplication by 100 gives the percentage frequency.

  • Cumulative frequency: Frequencies are progressively added. “Less-than” cumulative frequency counts observations below an upper boundary; “more-than” cumulative frequency counts observations above a lower boundary.
  • Class width: For equal intervals, width may be found from successive lower limits, such as (20-10=10). Consistent widths make histogram comparison easier.
  • Class boundaries: For whole-number scores in 10–19 and 20–29, continuous boundaries are commonly 9.5–19.5 and 19.5–29.5.
  • Class midpoint: The representative value of an interval is:
TEXT
Midpoint = (Lower limit + Upper limit) / 2

For 20–29, the midpoint is (24.5), which is used in a frequency polygon or approximate grouped-data calculation.

  • Good construction: Classes should be mutually exclusive, collectively exhaustive, clearly labelled, and preferably equal in width.

B. Worked example: constructing and interpreting a distribution

A researcher records anxiety scores and groups them into four intervals: 0–9 ((f=3)), 10–19 ((f=7)), 20–29 ((f=6)), and 30–39 ((f=4)).

  • Total sample: (N=3+7+6+4=20).
  • Relative frequency: The 10–19 interval represents (7/20=0.35), or 35%.
  • Cumulative frequency: The less-than cumulative frequency through 19 is (3+7=10), meaning 10 participants scored below the next interval.
  • Interpretive value: The table shows concentration in the lower and middle ranges, while the graph can reveal whether the distribution is approximately balanced or skewed.

IV. Measures of Positional Location: Deciles, Quartiles, and Percentiles

Positional measures divide an ordered dataset into equal portions. They are especially useful when psychological scores are interpreted relative to a norm group rather than only through an average.

A. docile

In statistical terminology, the intended concept is generally decile, not “docile.” Deciles divide ordered data into ten equal parts and indicate the percentage of observations below a position.

  • Definition: The (D_k) decile is the point below which approximately (10k)% of observations lie, where (k=1,\ldots,9).
  • Position for ungrouped data: A common convention is:
TEXT
Position of Dₖ = k(N + 1) / 10

Here, (k) is the decile number and (N) is the number of ordered observations. Some textbooks use (kN/10), so the selected convention should remain consistent.

  • Interpretation: (D_1) is the 10th percentile, (D_5) is the 50th percentile or median, and (D_9) is the 90th percentile.
  • Grouped-data estimate: For a decile located in a class interval:
TEXT
Dₖ = L + [(kN/10 − CF) / f] × h

Here, (L) is the lower class boundary, (CF) is cumulative frequency before the decile class, (f) is the decile-class frequency, and (h) is class width.

  • Use in psychology: A child at (D_8) on an anxiety measure scores above approximately 80% of the norm group, subject to the quality and representativeness of that norm group.

B. quartile

Quartiles divide ordered data into four equal sections, each containing approximately 25% of observations.

  • Definitions: (Q_1) is the 25th percentile, (Q_2) is the median or 50th percentile, and (Q_3) is the 75th percentile.
  • Ungrouped positions: A common rule is:
TEXT
Position of Q₁ = (N + 1) / 4
Position of Q₂ = (N + 1) / 2
Position of Q₃ = 3(N + 1) / 4
  • Grouped-data formula: For quartile (Q_k), where (k=1,2,3):
TEXT
Qₖ = L + [(kN/4 − CF) / f] × h

The symbols have the same meanings as in the grouped decile formula.

  • Interquartile range: The middle 50% spread is:
TEXT
IQR = Q₃ − Q₁

A large IQR indicates greater dispersion among the central half of scores.

  • Robustness: Quartiles and the IQR are less affected by extreme observations than the range, making them useful for skewed clinical data such as hospital waiting times.

C. percentile

A percentile expresses the relative standing of a score by dividing an ordered distribution into 100 parts.

  • Definition: The (Pk) percentile is the value below which approximately (k)% of observations fall; (P{50}) is the median.
  • Ungrouped position: One common formula is:
TEXT
Position of Pₖ = k(N + 1) / 100

Here, (k) is the desired percentile from 1 to 99.

  • Grouped-data estimate: The percentile within a class is estimated by:
TEXT
Pₖ = L + [(kN/100 − CF) / f] × h

Here, (L), (CF), (f), and (h) identify the relevant class and its boundaries.

  • Worked example: If a student is at the 85th percentile on a reasoning test, approximately 85% of the norm group scored at or below that student, while approximately 15% scored higher.
  • Important distinction: A percentile rank is not the same as a percentage score. A score of 85% correct concerns items answered correctly; the 85th percentile concerns comparison with other people.
  • Interpretive caution: Percentile ranks are ordinal: the difference between the 50th and 60th percentiles need not represent the same score difference as the difference between the 80th and 90th percentiles.

D. Applications and limitations of positional measures

Positional measures are valuable for comparison, but their meaning depends on the distribution and reference group.

  • Norm-referenced interpretation: Percentiles are meaningful only relative to a specified population, such as a national age group or a defined clinical sample.
  • Unequal intervals: Percentile units are not equal measurement units; a movement from the 90th to 95th percentile may involve fewer raw-score points than a movement from the 50th to 55th.
  • Grouped approximation: Quartiles, deciles, and percentiles calculated from grouped data estimate locations because every score within an interval is treated as if it were distributed progressively across that interval.
  • Graphical location: On an ogive, a desired cumulative frequency is located on the vertical axis, projected to the curve, and then dropped to the horizontal score axis.
  • Research usefulness: These measures support screening, selection, clinical classification, and identification of unusually low or high performance without assuming that scores are normally distributed.