Unit 3: Measures of Central Tendency and variability - Subjective Questions
PSY115 — Statistical Methods For Psychological Research • Practice Questions with Detailed Answers
20 questions
Define the median and explain the procedure for finding it in an ungrouped data set. Illustrate your answer with an example.
Definition: The median is the middle value of an ordered data set. It divides the observations into two equal parts: 50% of the values lie below it and 50% lie above it.
Procedure:
- Arrange the observations in ascending or descending order.
- If the number of observations, , is odd, the median is the value at position .
- If is even, the median is the average of the values at positions and .
Example: For the data , . Therefore, the median is at position , so the median is .
Explain how the median is calculated for a discrete frequency distribution and a continuous grouped frequency distribution.
For a discrete frequency distribution:
- Find the total frequency, .
- Compute the cumulative frequencies.
- Locate the value corresponding to the th observation, or the observation that crosses the middle position.
For a continuous grouped distribution: the median is calculated using
where:
- is the lower class boundary of the median class.
- is the total frequency.
- is the cumulative frequency before the median class.
- is the frequency of the median class.
- is the class width.
The median class is the class whose cumulative frequency first exceeds .
Define the mode and discuss its usefulness in psychological research.
Definition: The mode is the value or category that occurs with the highest frequency in a data set.
Usefulness in psychological research:
- It identifies the most common response, behavior, or score.
- It can be used with nominal, ordinal, interval, and ratio data, although it is especially appropriate for nominal data.
- It is not affected by extremely high or low values.
- It is useful for identifying the most frequently selected category in questionnaires or diagnostic classifications.
A data set may be unimodal, bimodal, or multimodal, depending on whether it has one, two, or more modes.
Distinguish between the mean, median, and mode as measures of central tendency.
| Measure | Meaning | Main advantage | Main limitation |
|---|---|---|---|
| Mean | Sum of all values divided by their number | Uses every observation and is suitable for further statistical analysis | Strongly affected by extreme scores |
| Median | Middle value in an ordered distribution | Appropriate for skewed data and ordinal data | Does not use the exact value of every observation |
| Mode | Most frequently occurring value | Can be used with nominal data and is easy to identify | May not exist or may be non-unique |
The mean is generally preferred for symmetrical quantitative data, the median for skewed distributions, and the mode for categorical or frequency-based data.
Calculate the arithmetic mean for the scores and explain its interpretation.
The arithmetic mean is calculated as
For the given scores,
and . Therefore,
Interpretation: The average score of the group is . This does not mean that every participant scored $13`; rather, the total of all scores is equivalent to five scores of $13$ each.
Derive the formula for the arithmetic mean in a frequency distribution and explain the meaning of each symbol.
Suppose a value occurs with frequency . Instead of adding each observation separately, we can multiply each value by its frequency. Thus, the total of all observations is
The total number of observations is
Therefore, the arithmetic mean is
where:
- represents a score or class midpoint.
- represents the frequency associated with .
- is the weighted total.
- is the total number of observations.
For grouped continuous data, the class midpoint is usually used as .
Explain the effect of extreme scores on the mean, median, and mode using a suitable example.
Consider the data set . Its mean is , median is , and there is no unique mode. If the largest value changes from to , the data become .
- New mean: .
- New median: .
- The mode remains unchanged because no value occurs more frequently than another.
Thus, the mean is highly sensitive to extreme values, whereas the median and mode are relatively resistant. For highly skewed psychological data, such as reaction times or clinical symptom scores, the median may provide a more representative measure of central tendency.
Define the range and explain its advantages and limitations as a measure of variability.
Definition: The range is the difference between the largest and smallest observations.
Advantages:
- It is simple and quick to calculate.
- It gives an immediate indication of the total spread of scores.
- It is useful for preliminary comparisons.
Limitations:
- It depends only on two observations.
- It is strongly affected by extreme scores.
- It is unstable from sample to sample.
- It does not describe how the remaining observations are distributed.
Therefore, the range is useful as a basic measure but is generally less reliable than the standard deviation.
Calculate the range for the data set and interpret the result.
First identify the largest and smallest values:
- Largest value: .
- Smallest value: .
The range is
Interpretation: The scores extend across a total interval of units, from to . The range does not indicate whether the scores are evenly distributed within this interval.
Explain quartiles and define the interquartile range.
Quartiles divide an ordered data set into four equal parts:
- , or the first quartile, is the value below which 25% of observations lie.
- is the median, with 50% of observations below it.
- , or the third quartile, is the value below which 75% of observations lie.
The interquartile range, or IQR, is
It measures the spread of the middle 50% of observations. Since it ignores the lowest and highest quarters of the data, it is less affected by extreme values than the range.
For the ordered data , find , , and the interquartile range.
There are observations. Divide the data into two equal halves:
- Lower half: .
- Upper half: .
The median of the lower half is
The median of the upper half is
Therefore,
Thus, the middle 50% of the observations are spread across an interval of units.
Define the semi-interquartile range and explain its relationship with the interquartile range.
The semi-interquartile range, also called the quartile deviation, is half the interquartile range.
Since
we can write
It measures the average distance of the first and third quartiles from the median and describes the dispersion of the central 50% of observations. It is particularly useful when the distribution is skewed or contains extreme scores.
Compare the range, interquartile range, and semi-interquartile range as measures of variability.
- Range: . It uses only the extreme values and is highly affected by outliers.
- Interquartile range: . It measures the spread of the middle 50% and is less affected by extreme values.
- Semi-interquartile range: . It is half the IQR and provides a more compact measure of central dispersion.
The range is simplest but least stable. The IQR and semi-IQR are more suitable for skewed distributions and ordinal data. However, none of these measures uses every observation in the data set.
Define average deviation and derive its formula for individual observations.
Average deviation, also called mean deviation, is the arithmetic mean of the absolute deviations of observations from a selected central value, usually the mean or median.
For deviations from the mean, the formula is
For deviations from the median, it is
where:
- is an individual observation.
- is the mean.
- is the median.
- is the number of observations.
- represents the absolute deviation from the chosen central value .
Absolute values are used so that positive and negative deviations do not cancel one another.
Calculate the average deviation from the mean for the scores .
First calculate the mean:
Now calculate the absolute deviations:
The sum of absolute deviations is . Therefore,
The average deviation from the mean is score units.
Define variance and standard deviation, and explain the relationship between them.
Variance is the average of the squared deviations of observations from their mean. For a population,
For a sample, the unbiased estimate is
Standard deviation is the positive square root of variance:
or, for a sample,
Variance is expressed in squared units, while standard deviation is expressed in the original units of measurement. Consequently, standard deviation is usually easier to interpret.
Derive the computational formula for the population variance and standard deviation.
The deviation of each observation from the mean is . Therefore, population variance is
Expanding the squared term,
Since ,
Hence,
The computational formula for standard deviation is therefore
This form is convenient when the sum of scores and the sum of squared scores are known.
Calculate the variance and standard deviation of the population scores .
First calculate the mean:
Calculate the deviations and squared deviations:
Thus,
Population variance:
Population standard deviation:
Therefore, the variance is and the standard deviation is approximately .
Explain why the standard deviation is considered a superior measure of variability compared with the range and average deviation.
The standard deviation is often preferred because:
- It uses every observation in the data set.
- It is mathematically well defined and suitable for advanced statistical procedures.
- It gives greater weight to large deviations because deviations are squared.
- It is expressed in the same units as the original scores.
- It is useful in correlation, regression, analysis of variance, and standard-score calculations.
Compared with the range, it is less dependent on only two extreme scores. Compared with average deviation, it has stronger mathematical properties because squared deviations can be manipulated algebraically. However, it is more sensitive to outliers than the average deviation and interquartile range.
Discuss the major characteristics of the arithmetic mean as a measure of central tendency.
The arithmetic mean has the following characteristics:
- It is obtained by dividing the sum of all observations by their number.
- It is based on every score in the distribution.
- The sum of deviations from the mean is zero:
- The sum of squared deviations is minimum when deviations are calculated from the mean.
- It is uniquely defined for a given numerical data set.
- It is affected by extreme values.
- It is appropriate mainly for interval and ratio-level data.
- It is widely used in further statistical analysis.
These properties make the mean powerful, although the median may be preferable for highly skewed distributions.
Define the median and explain the procedure for finding it in an ungrouped data set. Illustrate your answer with an example.
Definition: The median is the middle value of an ordered data set. It divides the observations into two equal parts: 50% of the values lie below it and 50% lie above it.
Procedure:
- Arrange the observations in ascending or descending order.
- If the number of observations, , is odd, the median is the value at position .
- If is even, the median is the average of the values at positions and .
Example: For the data , . Therefore, the median is at position , so the median is .
Did this save you a night before the exam?
LPU Notes is free, and it stays free. Ads cover part of the server bill. The rest comes out of a student's own pocket: the domain, the storage, and keeping the site up through the weeks everyone needs it at once.
The payment button didn't load. An ad blocker or a filtered network is the usual reason. to try again.
Nothing here is ever locked, and nothing unlocks. Chip in only if it was worth it. What it pays for →