Unit 5: Correlation and Regression - Subjective Questions
MGN206 — Research Methodology • Practice Questions with Detailed Answers
20 questions
Define correlation and explain its main characteristics.
Correlation is a statistical measure that indicates the degree and direction of association between two or more variables.
Main characteristics:
- The correlation coefficient is generally denoted by .
- Its value lies between and , that is, .
- A positive value indicates that the variables move in the same direction.
- A negative value indicates that the variables move in opposite directions.
- A value of zero indicates the absence of a linear relationship.
- Correlation is a unit-free measure and is unaffected by changes in origin and scale.
- Correlation indicates association, but it does not by itself establish causation.
Explain the major assumptions that should be satisfied when applying Pearson's correlation coefficient.
Pearson's correlation coefficient is based on the following assumptions:
- Level of measurement: Both variables should be measured on an interval or ratio scale.
- Linearity: The relationship between the variables should be approximately linear.
- Normality: Both variables should be approximately normally distributed, especially when significance tests are performed.
- Homoscedasticity: The variability of one variable should remain reasonably constant across the values of the other variable.
- Independence: Each pair of observations should be independent of other pairs.
- Absence of influential outliers: Extreme observations should not unduly determine the value of .
- Paired observations: Each value of one variable must correspond to a value of the other variable for the same unit of analysis.
These assumptions should be examined through study design, descriptive statistics, and graphical tools such as scatter plots.
Describe the different types of correlation on the basis of direction, number of variables, and form of relationship.
Correlation can be classified in several ways:
1. On the basis of direction:
- Positive correlation: Both variables increase or decrease together, so .
- Negative correlation: One variable increases while the other decreases, so .
- Zero correlation: No systematic linear association exists, so .
2. On the basis of the number of variables:
- Simple correlation: Association between two variables.
- Partial correlation: Association between two variables while controlling one or more additional variables.
- Multiple correlation: Association of one variable with two or more other variables jointly.
3. On the basis of form:
- Linear correlation: The relationship can be represented approximately by a straight line.
- Nonlinear or curvilinear correlation: The rate or direction of change varies across the values of the variables.
What is a scatter diagram? Explain how it is used to study correlation.
A scatter diagram is a graph in which each paired observation is represented by a point with coordinates .
Interpretation:
- Points rising from left to right indicate a positive correlation.
- Points falling from left to right indicate a negative correlation.
- Points lying close to a straight line indicate a strong linear correlation.
- Widely scattered points indicate a weak correlation.
- Points with no recognizable pattern suggest little or no association.
- A curved pattern indicates a possible nonlinear relationship.
- Isolated points may reveal outliers.
A scatter diagram is useful for checking linearity, unusual observations, and changes in variability before calculating a correlation coefficient. However, it provides a visual assessment rather than an exact numerical measure.
Define Pearson's product-moment correlation coefficient and explain the meaning of its formula.
Pearson's product-moment correlation coefficient measures the strength and direction of a linear relationship between two quantitative variables.
Its sample formula is:
Here:
- and are paired observations.
- and are their sample means.
- The numerator measures the joint variation of and .
- The denominator standardizes the joint variation using the variability of both variables.
Equivalently, Pearson's coefficient may be expressed as:
A value near indicates a strong positive linear relationship, a value near indicates a strong negative linear relationship, and a value near zero indicates a weak linear relationship.
Derive the computational formula for Pearson's correlation coefficient from the deviation-score formula.
The deviation-score formula is:
Expanding the numerator gives:
Similarly:
and
Substituting these expressions and multiplying the numerator and denominator terms by produces the computational formula:
This form is convenient for direct calculation from raw scores because it uses sums of values, squares, and cross-products.
Explain how the magnitude and direction of Pearson's correlation coefficient should be interpreted.
The sign and absolute magnitude of have different meanings:
- The sign indicates direction: represents a positive relationship, while represents a negative relationship.
- The magnitude indicates the strength of the linear relationship.
- represents a perfect positive linear relationship.
- represents a perfect negative linear relationship.
- indicates no linear relationship, although a nonlinear relationship may still exist.
A common descriptive guide is:
- to : very weak
- to : weak
- to : moderate
- to : strong
- to : very strong
These boundaries are contextual rather than universal. Interpretation should also consider sample size, measurement reliability, research design, outliers, and practical significance.
Discuss the properties, advantages, and limitations of Pearson's correlation coefficient.
Properties:
- .
- It is symmetric: .
- It is dimensionless and unaffected by linear changes of origin and positive scale.
- It measures linear association.
- Its sign is the same as that of the covariance.
Advantages:
- It provides a precise and widely understood measure of linear association.
- It uses the actual magnitudes of all observations.
- It supports inferential procedures such as confidence intervals and significance tests.
- It is closely connected to simple linear regression.
Limitations:
- It can be severely affected by outliers.
- A value near zero may conceal a strong nonlinear relationship.
- Restricted ranges can reduce the observed correlation.
- Mixing distinct subgroups can create a misleading correlation.
- It does not prove that one variable causes the other.
- Measurement error may weaken the observed association.
Define Spearman's rank correlation coefficient and state when it is appropriate.
Spearman's rank correlation coefficient, denoted by or , is a nonparametric measure of the strength and direction of a monotonic relationship between two variables. It is calculated by applying Pearson's correlation to the ranked values.
When there are no tied ranks, it is calculated as:
where is the difference between the two ranks for observation , and is the number of paired observations.
It is appropriate when:
- The data are ordinal or naturally expressed as ranks.
- The relationship is monotonic but not necessarily linear.
- Pearson's assumptions are seriously violated.
- Quantitative data contain influential outliers.
- The sample is small and normality is doubtful.
Its value ranges from to .
Describe the procedure for calculating Spearman's rank correlation coefficient, including the treatment of tied observations.
The calculation involves the following steps:
- Arrange or inspect the paired observations for the two variables.
- Assign ranks separately to the values of and using the same ranking direction.
- When values are tied, assign each tied value the average of the ranks that they would have occupied.
- Compute the rank difference for every observation.
- Square each difference and obtain .
- If there are no ties, use:
When ties occur, the most reliable general method is to calculate Pearson's correlation between the average-rank variables:
The resulting sign gives the direction, while the absolute value gives the strength of the monotonic association.
Compare Pearson's correlation coefficient with Spearman's rank correlation coefficient.
Pearson's correlation:
- Uses actual numerical values.
- Usually requires interval- or ratio-scale variables.
- Measures a linear relationship.
- Is more sensitive to outliers and non-normal distributions.
- Is appropriate when its parametric assumptions are reasonably satisfied.
Spearman's correlation:
- Uses ranks rather than original magnitudes.
- Can be applied to ordinal data.
- Measures a monotonic relationship, which may be curved.
- Is generally less affected by extreme values.
- Does not require normality in the same way as Pearson's method.
Similarity: Both coefficients range from to and describe the direction and strength of association.
Pearson's coefficient is usually more informative and statistically efficient for a genuinely linear relationship with suitable quantitative data. Spearman's coefficient is preferable for ranked data, monotonic nonlinear patterns, or data containing serious outliers.
Explain why correlation does not imply causation. Illustrate your answer with possible sources of a misleading correlation.
A correlation shows that variables vary together, but it does not establish that changes in one variable produce changes in the other.
A correlation may be misleading because of:
- Third-variable influence: A separate variable may affect both observed variables. For example, temperature may increase both cold-drink sales and electricity use.
- Reverse causality: may influence , may influence , or both processes may occur.
- Coincidence: A sample may show an association purely by chance.
- Selection bias: The observed sample may not represent the target population.
- Common trends: Two unrelated time-series variables may rise over time and appear correlated.
- Aggregation: Combined subgroup data may show a pattern that differs from the patterns within the subgroups.
A causal conclusion generally requires temporal precedence, control of alternative explanations, a plausible mechanism, and preferably experimental or strong quasi-experimental evidence.
Define regression and explain the basic concepts involved in regression analysis.
Regression analysis is a statistical method used to model the relationship between a dependent variable and one or more independent variables, often for explanation or prediction.
Basic concepts:
- Dependent variable: The outcome to be explained or predicted, commonly denoted by .
- Independent variable: A predictor or explanatory variable, commonly denoted by .
- Regression function: The expected value of for specified predictor values.
- Regression coefficient: A parameter that measures the expected change in associated with a change in a predictor.
- Predicted value: The fitted outcome, denoted by .
- Residual: The difference between an observed and fitted outcome, .
- Error term: Unobserved influences on the population outcome.
In simple linear regression, the model is:
where is the intercept, is the slope, and is the random error.
Distinguish between correlation and regression analysis.
Correlation and regression both study relationships, but their purposes differ:
- Correlation measures the strength and direction of association; regression specifies an equation for explaining or predicting an outcome.
- Correlation treats the two variables symmetrically; regression distinguishes the dependent variable from the independent variable.
- Pearson's is unit-free; regression coefficients generally have measurement units.
- Interchanging and does not change , but it usually changes the regression equation.
- Correlation is bounded between and ; a regression slope is not subject to these bounds.
- Correlation provides a single coefficient; regression provides an intercept, slope, fitted values, and residuals.
Neither method alone proves causality. A regression coefficient has a causal interpretation only when the research design and assumptions justify it.
State and explain the assumptions of the classical simple linear regression model.
The principal assumptions are:
- Linearity: The conditional mean of is linear in : .
- Independence: Observations or error terms are independent.
- Zero conditional mean: , meaning omitted influences are not systematically related to .
- Homoscedasticity: The conditional error variance is constant: .
- Normality of errors: Errors are normally distributed for exact small-sample tests and confidence intervals.
- Predictor variability: must vary across observations so that a slope can be estimated.
- Absence of influential observations: No single observation should exert excessive influence on the fitted equation.
Violations can produce biased estimates, inefficient estimates, incorrect standard errors, or unreliable predictions, depending on the particular assumption violated.
Derive the least-squares estimates of the intercept and slope in simple linear regression.
For the fitted equation
ordinary least squares minimizes the residual sum of squares:
Differentiate with respect to and and set both derivatives equal to zero:
These yield the normal equations. Solving them gives:
and
Therefore, the fitted regression equation is:
The result also shows that the least-squares line passes through the point .
Explain the interpretation of the intercept, slope, predicted value, and residual in a simple linear regression equation.
Consider the fitted equation:
- Intercept : The predicted value of when . It may lack a meaningful practical interpretation if zero is outside the observed range of .
- Slope : The expected change in the predicted value of for a one-unit increase in . A positive slope indicates an increasing relationship, while a negative slope indicates a decreasing relationship.
- Predicted value : The value of the outcome estimated by the fitted line for observation .
- Residual : The vertical difference between the observed and predicted values:
A positive residual means that the observed value lies above the fitted line; a negative residual means that it lies below the line. Residuals are used to assess model fit and regression assumptions.
Explain the coefficient of determination and its relationship with correlation in simple linear regression.
The coefficient of determination, denoted by , measures the proportion of sample variation in explained by the fitted regression model.
The total variation is decomposed as:
where:
- is the total sum of squares.
- is the regression sum of squares.
- is the error sum of squares.
Thus:
With an intercept, . For example, means that 64 percent of the observed sample variation in is explained by the model.
In simple linear regression with one predictor and an intercept:
does not establish causality and does not, by itself, prove that the model is correctly specified.
Describe how residual analysis is used to evaluate a linear regression model.
Residual analysis examines the values to determine whether the fitted model and its assumptions are reasonable.
Important diagnostic checks:
- A residual-versus-fitted plot should show a random horizontal band around zero.
- A curved pattern suggests that the assumed linear relationship is inadequate.
- A funnel-shaped pattern indicates possible heteroscedasticity.
- A residual-versus-order or residual-versus-time plot can reveal dependence or serial correlation.
- A histogram or normal probability plot can be used to assess approximate normality.
- Large standardized residuals may indicate outliers.
- Leverage and influence measures, such as Cook's distance, help identify observations that strongly affect the fitted equation.
When a violation is detected, possible responses include transforming variables, adding relevant predictors or nonlinear terms, using robust standard errors, investigating data quality, or selecting a more appropriate model.
Explain the relationship between the regression coefficients and Pearson's correlation coefficient. Also distinguish interpolation from extrapolation.
For the regression of on , the slope is:
For the regression of on , the slope is:
Therefore:
and
Both regression coefficients have the same sign as . Unlike , their magnitudes depend on the units of measurement and may exceed one.
Interpolation means using the regression equation to predict for an value within the observed data range. It is generally more reliable because it uses the model where data provide direct support.
Extrapolation means predicting for an value outside the observed range. It is risky because the relationship may change beyond the available data, even when the fitted line describes the observed range well.
Define correlation and explain its main characteristics.
Correlation is a statistical measure that indicates the degree and direction of association between two or more variables.
Main characteristics:
- The correlation coefficient is generally denoted by .
- Its value lies between and , that is, .
- A positive value indicates that the variables move in the same direction.
- A negative value indicates that the variables move in opposite directions.
- A value of zero indicates the absence of a linear relationship.
- Correlation is a unit-free measure and is unaffected by changes in origin and scale.
- Correlation indicates association, but it does not by itself establish causation.
Did this save you a night before the exam?
LPU Notes is free, and it stays free. Ads cover part of the server bill. The rest comes out of a student's own pocket: the domain, the storage, and keeping the site up through the weeks everyone needs it at once.
The payment button didn't load. An ad blocker or a filtered network is the usual reason. to try again.
Nothing here is ever locked, and nothing unlocks. Chip in only if it was worth it. What it pays for →