Unit 8: Descriptive Statistics and Time Series - Subjective Questions
DEMGN832 — Research Methodology • Practice Questions with Detailed Answers
20 questions
Define arithmetic mean. Explain how it is calculated for ungrouped data, and state its main merits and limitations.
Arithmetic mean is the sum of all observations divided by the total number of observations.
For observations , the arithmetic mean is:
Steps:
- Add all the observations.
- Count the number of observations.
- Divide the total by the number of observations.
Merits:
- It is simple to understand and calculate.
- It uses every observation in the data set.
- It is suitable for further algebraic analysis.
- It is comparatively stable across different samples.
Limitations:
- It is highly affected by extreme values.
- It may not represent a highly skewed distribution.
- It cannot be determined reliably for open-ended class intervals.
- It may produce a value that is not present in the original data.
Derive and explain the formulas used to calculate the arithmetic mean of grouped data by the direct, assumed mean, and step-deviation methods.
For grouped data, let denote the class midpoint, the corresponding frequency, and .
1. Direct method:
Each class midpoint is multiplied by its frequency, and the products are added.
2. Assumed mean method:
Choose a convenient assumed mean and calculate . Then:
This follows because .
3. Step-deviation method:
When class intervals have equal width , define:
The mean is:
The direct method is straightforward, while the assumed mean and step-deviation methods reduce computational effort when values are large or class intervals are uniform.
What is the median? Describe the procedure for finding the median of ungrouped and grouped data.
The median is the value that divides an ordered data set into two equal parts. Half of the observations lie below it and half lie above it.
For ungrouped data:
- Arrange the observations in ascending order.
- If is odd, the median is the value of the th observation.
- If is even, the median is the average of the th and th observations.
For grouped continuous data:
- Calculate cumulative frequencies.
- Find , where .
- Identify the median class whose cumulative frequency first exceeds .
- Apply:
where is the lower boundary of the median class, is the cumulative frequency before it, is its frequency, and is the class width.
The median is particularly useful for skewed data and open-ended distributions.
Define mode and explain how it is determined for individual observations and grouped frequency distributions.
The mode is the value that occurs most frequently in a data set. It indicates the most common or typical observation.
For individual observations:
- Count how many times each value occurs.
- The value with the highest frequency is the mode.
- A distribution may be unimodal, bimodal, multimodal, or have no mode.
For grouped data:
- Identify the modal class, which has the highest frequency.
- Use the formula:
where:
- is the lower boundary of the modal class.
- is the frequency of the modal class.
- is the frequency of the preceding class.
- is the frequency of the succeeding class.
- is the class width.
Mode is useful for nominal data and for identifying the most popular size, category, or choice.
Compare the mean, median, and mode as measures of central tendency. Under what circumstances should each measure be used?
Mean, median, and mode describe the central position of data, but they differ in their properties.
Mean:
- Uses all observations.
- Is suitable for quantitative and approximately symmetric data.
- Is affected by extreme values.
- Supports algebraic and statistical calculations.
Median:
- Is based on the position of observations.
- Is suitable for skewed distributions, ordinal data, and open-ended classes.
- Is not greatly affected by extreme values.
- Does not use the exact magnitude of every observation.
Mode:
- Represents the most frequent value.
- Can be used with nominal, ordinal, or quantitative data.
- Is useful in market research and production planning.
- May be unstable or may not exist uniquely.
For a symmetric unimodal distribution, . In a moderately skewed distribution, the empirical relationship is:
What is dispersion? Explain why measures of dispersion are necessary in statistical analysis.
Dispersion refers to the extent to which observations are scattered around a central value. It shows the degree of variability or consistency in a data set.
Measures of central tendency alone may be misleading because two distributions can have the same mean but very different spreads.
Importance of dispersion:
- It indicates the reliability and representativeness of an average.
- It enables comparison of variability between data sets.
- It helps assess consistency, stability, and risk.
- It supports quality control and decision-making.
- It provides a basis for advanced methods such as correlation, regression, and hypothesis testing.
Absolute measures include range, quartile deviation, mean deviation, variance, and standard deviation. Relative measures, such as the coefficient of range and coefficient of variation, are useful when comparing series measured in different units or having different averages.
Explain range and quartile deviation as measures of dispersion. State their formulas, merits, and limitations.
Range is the difference between the largest and smallest observations:
Its relative measure is the coefficient of range:
Range is simple and useful for quick comparisons, but it depends only on two extreme values and is highly unstable.
Quartile deviation, also called the semi-interquartile range, measures the spread of the middle of observations:
Its coefficient is:
Merits of quartile deviation:
- It is less affected by extreme observations.
- It is suitable for skewed and open-ended distributions.
Limitations:
- It ignores half of the observations.
- It is not suitable for extensive algebraic treatment.
- It is less stable than standard deviation.
Define variance and standard deviation. Derive their formulas for ungrouped and grouped data and explain their interpretation.
Variance is the average of the squared deviations of observations from their arithmetic mean. Standard deviation is the positive square root of variance.
For ungrouped population data:
For grouped data:
A useful computational formula is:
For a sample, the unbiased estimate usually uses :
A small standard deviation indicates that observations are concentrated near the mean, while a large standard deviation indicates greater variability. Unlike variance, standard deviation is expressed in the original unit of measurement.
What is the coefficient of variation? Explain how it is used to compare the consistency of two or more data sets.
The coefficient of variation, abbreviated as , is a relative measure of dispersion. It expresses standard deviation as a percentage of the arithmetic mean:
For sample data, it may be written as:
Interpretation:
- A lower indicates greater consistency or less relative variability.
- A higher indicates lower consistency or greater relative variability.
For example, if Series A has and Series B has , Series A is relatively more consistent.
The coefficient is especially useful for comparing data sets that have different means or different units. However, it should not be used when the mean is zero or very close to zero because the resulting value becomes undefined or misleading.
Distinguish between symmetrical, positively skewed, and negatively skewed distributions. Explain the relationship among mean, median, and mode in each case.
A distribution is described by its shape and the direction in which its tail extends.
Symmetrical distribution:
- The two sides are approximately mirror images.
- The frequencies are balanced around the center.
- The relationship is .
Positively skewed distribution:
- The longer tail extends toward larger values.
- A few high observations pull the mean to the right.
- The usual relationship is .
Negatively skewed distribution:
- The longer tail extends toward smaller values.
- A few low observations pull the mean to the left.
- The usual relationship is .
Skewness is important because it influences the choice of average and statistical technique. The median is generally more representative than the mean in a highly skewed distribution.
Explain the concepts of skewness and kurtosis. How do they help describe the shape of a distribution?
Skewness measures the degree and direction of asymmetry in a distribution.
Karl Pearson's coefficient of skewness is:
If the mode is not clearly defined, it can be estimated by:
A positive value indicates right skewness, a negative value indicates left skewness, and a value near zero indicates approximate symmetry.
Kurtosis measures the relative peakedness or flatness of a distribution. Using central moments:
The excess kurtosis is .
Types of kurtosis:
- Mesokurtic: , similar to a normal distribution.
- Leptokurtic: , more peaked with heavier tails.
- Platykurtic: , flatter with lighter tails.
Together, skewness and kurtosis provide more information about distributional shape than central tendency and dispersion alone.
Define an index number. Describe its characteristics, uses, and limitations in research and economic analysis.
An index number is a statistical measure that expresses the relative change in a variable or group of related variables over time, place, or other conditions. The base-period value is generally taken as .
Characteristics:
- It measures relative rather than absolute change.
- It summarizes changes in one or several variables.
- It requires a suitable base period.
- It may be simple or weighted.
Uses:
- Measuring changes in prices, quantities, and values.
- Estimating inflation and changes in the cost of living.
- Deflating monetary values to obtain real values.
- Supporting wage revision, policy formulation, and business decisions.
- Comparing economic conditions across periods.
Limitations:
- Results depend on the selected base year and commodities.
- Data may be inaccurate or unrepresentative.
- Consumer preferences and product quality may change.
- Different formulas can produce different results.
- An index is an approximate measure and may conceal variations among individual items.
Explain the construction of simple aggregative and simple average of price relatives index numbers.
A simple aggregative price index compares the total prices of selected commodities in the current period with their total prices in the base period:
where represents base-period prices and represents current-period prices.
Its main limitation is that commodities with high absolute prices receive greater implicit importance, even when they are not economically more important.
Under the simple average of price relatives method, first calculate the price relative for each commodity:
The arithmetic mean of price relatives is:
Alternatively, the geometric mean may be used:
This method gives equal importance to every commodity. However, equal importance may also be unrealistic when commodities differ substantially in consumption or economic significance.
Derive the Laspeyres, Paasche, and Fisher price index numbers. Compare their weighting systems and explain why Fisher's index is called an ideal index.
Let and denote base-period prices and quantities, while and denote current-period prices and quantities.
Laspeyres price index:
It uses base-period quantities as weights. It may overstate price increases because it ignores substitution toward relatively cheaper goods.
Paasche price index:
It uses current-period quantities as weights. It may understate price increases because current quantities reflect substitution.
Fisher price index:
Fisher's index is the geometric mean of the Laspeyres and Paasche indices. It balances the influence of base-period and current-period weights.
It is called an ideal index because it satisfies important consistency tests, particularly the time-reversal test and factor-reversal test. Its main practical limitation is that it requires quantity data for both periods.
Describe the time-reversal test and factor-reversal test used to evaluate index numbers.
Time-reversal test: An index should give a consistent result when the time subscripts are reversed. In ratio form, it requires:
If indices are expressed with base , the equivalent condition is:
This test ensures that reversing the base and current periods exactly reverses the measured change.
Factor-reversal test: The product of the price index and quantity index should equal the value index:
where:
This test ensures consistency among changes in price, quantity, and total value.
Fisher's ideal index satisfies both tests. Laspeyres and Paasche indices generally do not satisfy both tests individually.
What is a time series? Explain the four major components of a time series with suitable examples.
A time series is a set of observations recorded in chronological order, usually at equal time intervals such as days, months, quarters, or years.
The four major components are:
- Trend (): The long-term upward or downward movement in the series. For example, population may show a long-term increasing trend.
- Seasonal variation (): Regular and recurring movements within a year. For example, ice cream sales may rise every summer.
- Cyclical variation (): Wave-like movements lasting more than one year, often associated with business cycles such as expansion and recession.
- Irregular variation (): Unpredictable movements caused by events such as natural disasters, strikes, wars, or sudden policy changes.
In an additive model:
In a multiplicative model:
The appropriate model depends on whether component effects are approximately constant or proportional to the level of the series.
Distinguish between the additive and multiplicative models of time series. State the conditions under which each model is appropriate.
The additive and multiplicative models describe how the components of a time series combine.
Additive model:
- Component effects are measured in the same units as the original series.
- Seasonal and irregular fluctuations remain approximately constant as the level of the series changes.
- It is appropriate when variation around the trend does not increase with the trend level.
- Seasonal effects over a complete cycle generally sum to zero.
Multiplicative model:
- Seasonal, cyclical, and irregular effects are expressed as ratios or percentages.
- Fluctuations increase or decrease in proportion to the level of the series.
- It is common in economic and business data.
- Seasonal indices over a complete cycle average , or when expressed as percentages.
A plot of the series helps determine the appropriate model. Roughly constant fluctuations suggest an additive model, while expanding or contracting fluctuations suggest a multiplicative model.
Explain the moving average method of measuring trend in a time series. Discuss its procedure, advantages, and limitations.
The moving average method measures trend by smoothing short-term fluctuations. Each trend value is calculated as the average of a fixed number of consecutive observations.
For a -period moving average:
Procedure:
- Select a moving-average period based on the seasonal cycle, such as quarters or months.
- Calculate the average of the first observations.
- Drop the first observation, include the next observation, and calculate another average.
- Place each average at the center of its period.
- For an even period, calculate centered moving averages by averaging consecutive moving averages.
Advantages:
- It is simple and reduces seasonal and irregular variation.
- It does not require an assumed mathematical trend equation.
Limitations:
- Trend values are unavailable at both ends of the series.
- The choice of period can affect the result.
- It does not provide an equation for forecasting beyond the observed period.
- It may conceal turning points.
Derive the least-squares linear trend equation for a time series and explain how it is used for forecasting.
A linear trend assumes that the time series changes at an approximately constant rate. The trend equation is:
where is the estimated trend value, is the intercept, is the change per time period, and represents time.
The least-squares method minimizes:
Differentiating with respect to and gives the normal equations:
If time is coded so that , the equations simplify to:
After estimating and , trend values are obtained by substituting the relevant . A future value can be forecast by inserting a future time code into the equation. The method is objective and provides a forecasting equation, but its accuracy declines when the actual trend is nonlinear or structural changes occur.
Describe the ratio-to-moving-average method for measuring seasonal variation and explain how seasonal indices are interpreted.
The ratio-to-moving-average method estimates seasonal variation after removing the trend-cycle component from a time series. It is mainly used with the multiplicative model.
Procedure:
- Calculate moving averages equal to the seasonal period, such as quarters or months.
- Center the moving averages when the seasonal period is even.
- Divide each actual observation by the corresponding centered moving average and multiply by :
- Group the seasonal relatives by month or quarter.
- Calculate the average or median for each seasonal group.
- Adjust the resulting indices so that their total is for monthly data or for quarterly data.
Interpretation:
- An index of means that the period is typically above the normal trend-cycle level.
- An index of means that the period is typically below that level.
Seasonal indices can be used to deseasonalize data through in ratio form and to incorporate seasonal effects into forecasts.
Define arithmetic mean. Explain how it is calculated for ungrouped data, and state its main merits and limitations.
Arithmetic mean is the sum of all observations divided by the total number of observations.
For observations , the arithmetic mean is:
Steps:
- Add all the observations.
- Count the number of observations.
- Divide the total by the number of observations.
Merits:
- It is simple to understand and calculate.
- It uses every observation in the data set.
- It is suitable for further algebraic analysis.
- It is comparatively stable across different samples.
Limitations:
- It is highly affected by extreme values.
- It may not represent a highly skewed distribution.
- It cannot be determined reliably for open-ended class intervals.
- It may produce a value that is not present in the original data.
Did this save you a night before the exam?
LPU Notes is free, and it stays free. Ads cover part of the server bill. The rest comes out of a student's own pocket: the domain, the storage, and keeping the site up through the weeks everyone needs it at once.
The payment button didn't load. An ad blocker or a filtered network is the usual reason. to try again.
Nothing here is ever locked, and nothing unlocks. Chip in only if it was worth it. What it pays for →