Unit 10: Test of Association - Subjective Questions
DEMGN832 — Research Methodology • Practice Questions with Detailed Answers
20 questions
Define the correlation coefficient. Explain its purpose, range, and interpretation in research methodology.
Correlation coefficient is a statistical measure that describes the strength and direction of the relationship between two quantitative variables.
- It is commonly denoted by .
- Its value ranges from to .
- indicates a perfect positive correlation.
- indicates a perfect negative correlation.
- indicates no linear correlation.
A positive value means that both variables tend to increase or decrease together. A negative value means that one variable tends to increase when the other decreases. The closer the absolute value of is to , the stronger the relationship. Correlation indicates association, but it does not by itself prove causation.
Explain the difference between positive, negative, and zero correlation with suitable examples.
Positive correlation: Two variables move in the same direction. For example, study time and examination scores may have a positive correlation.
Negative correlation: Two variables move in opposite directions. For example, price and demand for a product may have a negative correlation.
Zero correlation: There is no systematic linear relationship between the variables. For example, shoe size and examination performance may show approximately zero correlation in a group of students.
The direction is identified by the sign of , while the strength is judged by the magnitude of .
State and explain the main assumptions of correlation analysis.
The important assumptions of correlation analysis are:
- Quantitative variables: The variables should generally be measured on an interval or ratio scale.
- Linearity: The relationship between the variables should be approximately linear.
- Independence: Each pair of observations should be independent of other pairs.
- Absence of influential outliers: Extreme observations can substantially change the correlation coefficient.
- Approximate normality: For significance testing using Pearson correlation, the variables should be approximately normally distributed, especially in small samples.
If these assumptions are violated, the researcher may use a transformation, a graphical investigation, or a non-parametric method such as Spearman rank correlation.
Describe Spearman rank correlation and explain when it should be used.
Spearman rank correlation is a non-parametric measure of the strength and direction of a monotonic relationship between two variables. It is denoted by or .
It is appropriate when:
- Data are ordinal or can be converted into ranks.
- The sample is small or the normality assumption is doubtful.
- The relationship is monotonic but not necessarily linear.
- The data contain outliers that may affect Pearson correlation.
For paired observations with no tied ranks, it is calculated as:
where is the difference between the two ranks for observation . Its value ranges from to .
Derive and explain the formula for Spearman rank correlation coefficient for data without tied ranks.
Let and be the ranks of the two variables for the th observation. Define the difference between the ranks as:
The squared differences are summed as . Spearman rank correlation is based on the extent to which corresponding ranks differ. The formula is:
where:
- is the number of paired observations.
- is the difference between the paired ranks.
- is the sum of squared rank differences.
If all ranks are identical, and . If the rankings are exactly reversed, . The formula assumes that there are no tied ranks. When ties occur, average ranks are assigned and a standard correlation calculation on the ranks is generally used.
Calculate Spearman rank correlation for the following paired ranks: X = 1, 2, 3, 4, 5 and Y = 2, 1, 4, 3, 5. Interpret the result.
The calculation is shown below:
| X rank | Y rank | ||
|---|---|---|---|
| 1 | 2 | -1 | 1 |
| 2 | 1 | 1 | 1 |
| 3 | 4 | -1 | 1 |
| 4 | 3 | 1 | 1 |
| 5 | 5 | 0 | 0 |
Thus, and .
The result, , indicates a strong positive monotonic association between the two rankings.
Explain how tied ranks are handled in Spearman rank correlation.
Tied ranks occur when two or more observations have the same value. The usual procedure is:
- Arrange the observations in order.
- Assign the average of the ranks that would have been occupied by the tied observations.
- Repeat this procedure for both variables.
- Calculate the correlation between the resulting rank values.
For example, if two observations occupy ranks and but have equal values, each receives the average rank:
The simple formula is exact when there are no ties. With ties, calculating Pearson correlation on the assigned ranks is preferred because it automatically accounts for the tied values.
Define Karl Pearson's product-moment correlation coefficient and write its formula.
Karl Pearson's product-moment correlation coefficient measures the degree and direction of the linear relationship between two quantitative variables. It is usually denoted by .
The computational formula is:
An alternative form is:
where is the covariance and and are the standard deviations of and . The coefficient ranges from to .
Derive Karl Pearson's correlation coefficient from the deviation form of covariance and standard deviation.
Let and be the means of variables and . Define deviations as:
The covariance between and is proportional to , while the standard deviations are proportional to and . Therefore, Pearson's correlation coefficient is:
Using and , this becomes:
This standardizes covariance by the variability of both variables, producing a dimensionless measure between and .
Compare Pearson correlation and Spearman rank correlation.
| Basis | Pearson correlation | Spearman rank correlation |
|---|---|---|
| Type | Parametric measure | Non-parametric measure |
| Data | Interval or ratio data | Ordinal data or ranked data |
| Relationship | Measures a linear relationship | Measures a monotonic relationship |
| Assumptions | More sensitive to normality and outliers | Fewer distributional assumptions |
| Calculation | Uses actual numerical values | Uses ranks or ranked values |
| Symbol | or |
Pearson correlation is preferred when the data satisfy its assumptions and the relationship is linear. Spearman correlation is useful for ordinal data, non-normal data, monotonic relationships, or data affected by outliers.
Explain the relationship between covariance and Pearson correlation coefficient.
Covariance indicates whether two variables tend to vary together. It is positive when high values of one variable are associated with high values of the other and negative when they move in opposite directions.
However, covariance depends on the units of measurement and cannot be easily compared across studies. Pearson correlation standardizes covariance by dividing it by the product of the standard deviations:
Consequently:
- The sign of is the same as the sign of covariance.
- The value of is unit-free.
- The magnitude of is bounded between and .
Thus, Pearson correlation is a standardized form of covariance.
Describe the interpretation of the magnitude and sign of Pearson's correlation coefficient.
The interpretation of Pearson's involves both its sign and magnitude.
- Sign: A positive sign indicates that the variables move in the same direction, while a negative sign indicates movement in opposite directions.
- Magnitude: Values close to indicate a weak linear association, whereas values close to or indicate a strong linear association.
- Perfect association: or represents a perfect linear relationship.
- No linear association: indicates no linear relationship, although a non-linear relationship may still exist.
The exact labels such as weak, moderate, or strong should be interpreted in the context of the subject area, sample size, measurement quality, and research purpose.
Distinguish between correlation and causation. Why should researchers avoid interpreting a correlation as proof of cause and effect?
Correlation describes an association between variables, whereas causation means that changes in one variable produce changes in another.
A correlation does not establish causation because:
- Third-variable problem: An unmeasured variable may influence both observed variables.
- Reverse causality: The direction of influence may be opposite to the assumed direction.
- Chance association: The observed relationship may occur due to sampling variation.
- Selection bias: The sample may not represent the target population.
For example, ice-cream sales and cases of sunburn may be positively correlated because both increase during warm weather. Temperature is a third variable. Causal conclusions require stronger designs, such as randomized experiments, or careful longitudinal and statistical analysis.
What is meant by an association between nominal data? Explain with an example.
Nominal data consist of categories that have names or labels but no meaningful numerical order. An association between two nominal variables exists when the distribution of one variable differs across the categories of the other variable.
For example, a researcher may examine the association between gender and preferred learning mode. If the proportions preferring online, face-to-face, and blended learning are different across gender categories, the variables may be associated.
The data are summarized in a contingency table. Association is commonly tested using the chi-square test of independence. The null hypothesis states that the two nominal variables are independent, while the alternative hypothesis states that they are associated.
Explain the construction and components of a contingency table for testing association between nominal variables.
A contingency table cross-classifies observations according to the categories of two nominal variables.
Its main components are:
- Rows: Categories of the first variable.
- Columns: Categories of the second variable.
- Observed frequency (): The actual count in each cell.
- Row total: The total frequency across a row.
- Column total: The total frequency down a column.
- Grand total (): The total number of observations.
For a cell in row and column , the expected frequency under independence is:
Comparing observed and expected frequencies provides the basis for the chi-square test of association.
State the null and alternative hypotheses for a chi-square test of association and explain their meaning.
For a chi-square test of association between two categorical variables:
- Null hypothesis (): The two variables are independent; there is no association between them in the population.
- Alternative hypothesis (): The two variables are not independent; there is an association between them in the population.
For example, when testing the relationship between smoking status and respiratory disease, states that smoking status and respiratory disease are independent. states that the distribution of respiratory disease differs according to smoking status.
The test does not identify the direction or causal mechanism of the association. It only evaluates whether the observed differences are larger than would reasonably be expected by chance.
Derive the chi-square test statistic for testing independence between two categorical variables.
Suppose a contingency table has observed frequency in row and column . Under the null hypothesis of independence, the expected frequency is:
For each cell, the difference between observed and expected frequencies is assessed relative to the expected frequency. The chi-square statistic is therefore:
where is the number of rows and is the number of columns.
The degrees of freedom are:
A large value of indicates that the observed frequencies differ substantially from the frequencies expected under independence. The calculated statistic is compared with a critical value, or its -value is compared with the chosen significance level.
Calculate the expected frequencies for a contingency table with row totals 40 and 60, column totals 50 and 50, and grand total 100.
Expected frequency is calculated using:
The expected frequencies are:
- First row, first column:
- First row, second column:
- Second row, first column:
- Second row, second column:
Therefore, the expected-frequency table is:
| Column 1 | Column 2 | Total | |
|---|---|---|---|
| Row 1 | 20 | 20 | 40 |
| Row 2 | 30 | 30 | 60 |
| Total | 50 | 50 | 100 |
Explain the assumptions and conditions required for applying the chi-square test of association.
The main conditions for the chi-square test are:
- The variables must be categorical.
- Each observation must contribute to only one cell of the table.
- Observations should be independent.
- The categories should be mutually exclusive and collectively exhaustive.
- Expected frequencies should generally be sufficiently large. A common guideline is that no expected frequency should be less than , and at least of expected frequencies should be or more.
- The sample should be obtained using an appropriate sampling method.
When expected frequencies are too small, the researcher may combine substantively similar categories, increase the sample size, or use an exact test such as Fisher's exact test for a small table.
Differentiate between the chi-square goodness-of-fit test and the chi-square test of association.
| Feature | Goodness-of-fit test | Test of association or independence |
|---|---|---|
| Purpose | Compares one categorical variable with a specified distribution | Examines the relationship between two categorical variables |
| Data display | One-way frequency table | Two-way contingency table |
| Expected frequencies | Based on hypothesized proportions | Based on row and column totals |
| Degrees of freedom | Usually | |
| Example | Testing whether a die is fair | Testing whether gender and product preference are associated |
Both tests use the general statistic:
The difference lies in the research question and the method used to obtain expected frequencies.
Define the correlation coefficient. Explain its purpose, range, and interpretation in research methodology.
Correlation coefficient is a statistical measure that describes the strength and direction of the relationship between two quantitative variables.
- It is commonly denoted by .
- Its value ranges from to .
- indicates a perfect positive correlation.
- indicates a perfect negative correlation.
- indicates no linear correlation.
A positive value means that both variables tend to increase or decrease together. A negative value means that one variable tends to increase when the other decreases. The closer the absolute value of is to , the stronger the relationship. Correlation indicates association, but it does not by itself prove causation.
Did this save you a night before the exam?
LPU Notes is free, and it stays free. Ads cover part of the server bill. The rest comes out of a student's own pocket: the domain, the storage, and keeping the site up through the weeks everyone needs it at once.
The payment button didn't load. An ad blocker or a filtered network is the usual reason. to try again.
Nothing here is ever locked, and nothing unlocks. Chip in only if it was worth it. What it pays for →