Unit 12: Multivariate Analysis - Subjective Questions
DEMGN832 — Research Methodology • Practice Questions with Detailed Answers
20 questions
Define multivariate analysis and explain how multivariate techniques may be classified.
Multivariate analysis refers to a collection of statistical techniques used to analyze three or more variables simultaneously. It helps researchers understand complex relationships, reduce data, classify observations, and test dependence among variables.
Multivariate techniques can be classified as follows:
- Dependence techniques: One or more variables are treated as dependent variables and are predicted or explained by independent variables. Examples include discriminant analysis, multivariate regression, and conjoint analysis.
- Interdependence techniques: No variable is designated as dependent. The objective is to discover an underlying structure among variables or objects. Examples include factor analysis, cluster analysis, and multidimensional scaling.
- Metric-data techniques: These are applied when variables are measured on interval or ratio scales.
- Non-metric-data techniques: These are used with nominal or ordinal variables.
The selection of a technique depends on the research objective, measurement scale, number of dependent variables, and assumptions about the data.
Distinguish between dependence methods and interdependence methods of multivariate analysis.
The distinction is based on whether variables are assigned dependent and independent roles.
Dependence methods:
- One or more variables are explicitly identified as dependent variables.
- Other variables are used to explain, predict, or classify the dependent variables.
- A theoretical or causal relationship is normally specified in advance.
- Examples include discriminant analysis, multivariate analysis of variance, multiple regression, and conjoint analysis.
Interdependence methods:
- No variable is designated as dependent or independent.
- The aim is to identify patterns, dimensions, or groups within the data.
- These methods are generally exploratory.
- Examples include factor analysis, cluster analysis, and multidimensional scaling.
Thus, dependence methods test or estimate specified relationships, whereas interdependence methods discover structures that are not defined beforehand.
What is factor analysis? Explain its major objectives and applications in research.
Factor analysis is an interdependence technique that explains correlations among a large number of observed variables through a smaller number of unobserved variables called factors.
A common factor model is:
where is an observed variable, is a factor loading, is a common factor, and is the unique component.
Major objectives:
- Data reduction: Replace many correlated variables with fewer factors.
- Structure identification: Detect underlying dimensions represented by observed variables.
- Scale development: Group related items and assess construct structure.
- Multicollinearity control: Create relatively independent factor scores for further analysis.
- Variable screening: Select representative variables from large datasets.
Factor analysis is widely applied in attitude measurement, consumer research, psychology, education, organizational studies, and index construction.
Explain the important methods used to extract factors in factor analysis.
Important factor-extraction methods include:
- Principal component analysis: Extracts components that account for the maximum total variance. The first component explains the largest possible variance, and each subsequent component explains the largest remaining variance while remaining uncorrelated with earlier components.
- Principal axis factoring: Extracts factors from common variance shared among variables. Initial communalities are estimated and progressively refined.
- Maximum likelihood method: Estimates factor loadings that maximize the likelihood of obtaining the observed correlation matrix. It permits statistical tests and confidence intervals when multivariate normality is reasonable.
- Centroid method: Uses an approximate procedure to obtain factors without repeatedly solving complex equations. It is historically important but less precise than modern methods.
- Image factoring: Uses the portion of each variable predictable from the remaining variables and separates common from unique variance.
- Alpha factoring: Extracts factors to maximize the reliability or generalizability of the factors.
Principal component analysis is primarily a data-reduction method, whereas common factor methods seek to identify latent constructs.
Compare principal component analysis and common factor analysis.
Principal component analysis (PCA) and common factor analysis (CFA) both reduce a large set of variables, but their purposes and variance assumptions differ.
Principal component analysis:
- Uses the total variance of each observed variable.
- Represents a component as a weighted combination of observed variables.
- Includes common, unique, and error variance.
- Is mainly used for data reduction and creation of composite scores.
- Initially assigns each variable a communality of .
Common factor analysis:
- Uses only the common variance shared among variables.
- Assumes that latent factors produce the correlations among observed variables.
- Separates common variance from unique and error variance.
- Is mainly used to identify underlying theoretical constructs.
- Uses estimated communalities that are generally less than .
PCA is suitable when efficient summarization is the main aim, while common factor analysis is preferable when the researcher seeks latent dimensions.
Describe the complete procedure for conducting factor analysis.
A systematic factor-analysis procedure includes the following steps:
- Define the objective: Decide whether the purpose is data reduction, construct identification, or scale development.
- Select variables and sample: Include theoretically relevant metric variables and obtain an adequate sample.
- Prepare the correlation matrix: Examine whether meaningful correlations exist among variables.
- Assess suitability: Use the Kaiser-Meyer-Olkin measure and Bartlett's test of sphericity.
- Choose the extraction method: Select PCA, principal axis factoring, maximum likelihood, or another appropriate method.
- Determine the number of factors: Apply eigenvalues, a scree plot, cumulative variance, parallel analysis, and theoretical interpretability.
- Rotate the solution: Use orthogonal or oblique rotation to obtain a simpler and more interpretable structure.
- Interpret and name factors: Examine the magnitude and direction of factor loadings.
- Compute factor scores when required: Use regression, Bartlett, or summated-score methods.
- Validate the solution: Check reliability, replicate the structure, or perform confirmatory analysis on another sample.
Each decision should be supported by both statistical evidence and subject knowledge.
Explain how the suitability of data for factor analysis is assessed using the KMO measure and Bartlett's test of sphericity.
The suitability of data is commonly evaluated through the following tests:
Kaiser-Meyer-Olkin measure:
- The KMO statistic compares observed correlations with partial correlations.
- It ranges from to .
- A value closer to indicates compact correlation patterns and greater suitability for factor analysis.
- As a practical guideline, values of at least are generally acceptable, while values above are considered strong.
Bartlett's test of sphericity:
- It tests the null hypothesis that the population correlation matrix is an identity matrix.
- The hypotheses are and .
- A statistically significant result indicates that the variables possess sufficient correlations for factor analysis.
Researchers should also inspect sample size, variable distributions, correlations, partial correlations, outliers, and variables with very low communalities.
What are eigenvalues, communalities, and factor loadings? Explain their role in factor analysis.
Eigenvalue: An eigenvalue indicates the amount of variance explained by a factor. For standardized variables, the total variance equals the number of variables. A larger eigenvalue indicates a more important factor.
Communality: The communality of variable is the proportion of its variance explained by retained common factors:
A low communality means that the retained factor solution does not represent the variable well.
Factor loading: A factor loading expresses the strength and direction of the relationship between variable and factor . Large absolute loadings indicate that a variable is important in interpreting a factor.
Together, these measures help researchers determine factor importance, evaluate how well variables are represented, interpret the extracted dimensions, and decide whether variables or factors should be retained.
Discuss the criteria used to determine the appropriate number of factors to retain.
The number of factors should be determined by combining statistical criteria with theoretical judgment.
- Eigenvalue criterion: Retain factors with eigenvalues greater than when standardized variables are analyzed. This rule is simple but can overestimate or underestimate the number of factors.
- Scree plot: Plot eigenvalues against factor numbers and retain factors above the point at which the curve begins to level off.
- Percentage of variance: Retain enough factors to explain an acceptable cumulative proportion of total variance. The acceptable level depends on the research field.
- Parallel analysis: Compare observed eigenvalues with eigenvalues from randomly generated data. Retain factors whose observed eigenvalues exceed the random-data values.
- Interpretability: Retained factors should have meaningful variable loadings and a defensible theoretical explanation.
- Residual correlations: A good solution should leave relatively small unexplained correlations.
Parallel analysis and interpretability generally provide stronger evidence than reliance on the eigenvalue-greater-than-one rule alone.
Why is rotation performed in factor analysis? Distinguish between orthogonal and oblique rotation.
Rotation is performed after factor extraction to obtain a simpler and more interpretable loading pattern. It redistributes explained variance among retained factors but does not normally change the total variance explained by the complete solution.
Orthogonal rotation:
- Keeps factors uncorrelated at an angle of .
- Produces a single loading matrix that is relatively easy to interpret.
- Common methods include Varimax, Quartimax, and Equamax.
- It is appropriate when theory suggests independent dimensions.
Oblique rotation:
- Allows factors to be correlated.
- Produces a pattern matrix, a structure matrix, and a factor-correlation matrix.
- Common methods include Direct Oblimin and Promax.
- It is appropriate when underlying constructs are expected to be related.
The choice should follow theoretical expectations. Imposing orthogonality on naturally related constructs can produce an unrealistic representation.
Compare Varimax, Quartimax, and Equamax methods of orthogonal factor rotation.
All three methods maintain uncorrelated factors but optimize different aspects of the loading matrix.
- Varimax rotation: Simplifies the columns of the loading matrix. It attempts to produce factors with a few high loadings and many near-zero loadings. It is the most widely used method because it usually provides clearly differentiated factors.
- Quartimax rotation: Simplifies the rows of the loading matrix. It seeks to make each variable load highly on as few factors as possible. It may produce one broad general factor on which many variables load.
- Equamax rotation: Combines the objectives of Varimax and Quartimax. It attempts to simplify both variables and factors simultaneously.
Varimax is often selected when the aim is to interpret distinct dimensions, Quartimax when a general underlying factor is plausible, and Equamax when a compromise between row and column simplification is desired.
Define cluster analysis and explain its objectives, applications, and major stages.
Cluster analysis is an interdependence technique that divides objects or cases into relatively homogeneous groups called clusters. Objects within a cluster should be similar, while objects belonging to different clusters should be dissimilar.
Objectives and applications:
- Identify naturally occurring groups in data.
- Segment consumers, markets, products, regions, or organizations.
- Develop taxonomies and classify research observations.
- Summarize complex data and support targeted decisions.
Major stages:
- Define the clustering objective and select relevant variables.
- Standardize variables when their scales differ substantially.
- Choose a similarity or distance measure.
- Select a hierarchical or non-hierarchical clustering method.
- Decide the number of clusters.
- Interpret and profile each cluster.
- Validate stability by using alternative methods, split samples, or external variables.
Cluster analysis is exploratory, so its results depend strongly on variable selection, scaling, distance measures, and clustering algorithms.
Distinguish between hierarchical and non-hierarchical methods of cluster analysis.
Hierarchical clustering:
- Builds a nested sequence of clusters.
- Agglomerative methods begin with each object as a separate cluster and progressively merge clusters.
- Divisive methods begin with one cluster and progressively divide it.
- Results are commonly displayed using a dendrogram.
- The number of clusters need not be specified initially.
- Early merging or splitting decisions generally cannot be reversed.
Non-hierarchical clustering:
- Divides observations directly into a predetermined number of clusters.
- The K-means algorithm is a common example.
- Observations may be reassigned during the process, improving within-cluster homogeneity.
- It is computationally efficient for large datasets.
- Results can depend on starting seeds and the chosen value of .
A practical strategy is to use hierarchical clustering to identify a reasonable number of clusters and starting centers, followed by K-means to refine the solution.
Explain important distance measures and linkage methods used in cluster analysis.
A distance measure quantifies dissimilarity between observations. Common measures include:
- Euclidean distance:
- Squared Euclidean distance: Gives greater weight to large differences.
- Manhattan distance:
- Mahalanobis distance: Adjusts for differences in variance and correlations among variables.
Important hierarchical linkage methods are:
- Single linkage: Uses the minimum distance between members of two clusters; it can produce chaining.
- Complete linkage: Uses the maximum distance and tends to create compact clusters.
- Average linkage: Uses the average pairwise distance between clusters.
- Centroid linkage: Uses the distance between cluster centroids.
- Ward's method: Merges clusters that produce the smallest increase in within-cluster variance.
Variables should often be standardized before distance calculation because variables with larger scales can dominate the result.
What is discriminant analysis? State its objectives and major assumptions.
Discriminant analysis is a dependence technique used to distinguish between two or more predefined groups and classify observations into those groups using a set of predictor variables.
Its main objectives are to:
- Determine whether groups differ significantly on a combination of predictors.
- Identify variables that contribute most to group separation.
- Develop discriminant functions.
- Predict the group membership of new observations.
- Evaluate the accuracy of classification.
Major assumptions:
- The dependent grouping variable is categorical and groups are mutually exclusive.
- Predictor variables are metric or suitably coded.
- Predictors follow multivariate normal distributions within groups.
- Group covariance matrices are approximately equal for linear discriminant analysis.
- Observations are independent.
- Predictors do not exhibit severe multicollinearity.
- Extreme outliers are absent or appropriately treated.
Violations may reduce classification accuracy and make significance tests unreliable.
Derive the basic form of a discriminant function and explain how its performance is evaluated.
A linear discriminant function combines predictor variables into a score that maximizes separation among predefined groups:
where is the discriminant score, is a constant, is a discriminant coefficient, and is a predictor.
For two groups, the coefficients are selected to maximize the ratio:
An observation's score is compared with a cutoff score or group centroids to assign group membership.
Performance evaluation includes:
- Wilks' lambda: A smaller value indicates stronger group separation.
- Canonical correlation: Measures the association between discriminant scores and groups.
- Group centroids: Show the mean discriminant score of each group.
- Classification matrix: Compares actual and predicted membership.
- Hit ratio: Gives the percentage classified correctly.
- Cross-validation: Estimates performance on observations not used to fit the function.
- Comparison with chance accuracy: Determines whether the model improves meaningfully over random classification.
Validation on an independent sample provides the strongest assessment of predictive usefulness.
Define multidimensional scaling and describe its procedure and interpretation.
Multidimensional scaling (MDS) is an interdependence technique that represents perceived similarities or dissimilarities among objects as distances between points in a low-dimensional spatial map.
Procedure:
- Select the objects to be studied.
- Obtain similarity, dissimilarity, distance, or preference judgments.
- Decide whether metric or non-metric MDS is appropriate.
- Select the number of dimensions.
- Estimate object coordinates so that spatial distances reproduce the input proximities as closely as possible.
- Evaluate model fit using stress and related measures.
- Interpret and label dimensions using object positions or external attributes.
Objects positioned close together are perceived as similar, whereas distant objects are perceived as dissimilar. The orientation and direction of axes are arbitrary; therefore, dimensions must be interpreted substantively rather than from axis direction alone. MDS is commonly used to create perceptual maps of brands, products, institutions, or preferences.
Distinguish between metric and non-metric multidimensional scaling and explain the role of stress.
Metric MDS:
- Assumes proximity data are measured on an interval or ratio scale.
- Attempts to preserve the actual numerical distances or a specified metric relationship.
- Is suitable when quantitative distances have meaningful magnitudes.
Non-metric MDS:
- Requires only ordinal proximity information.
- Attempts to preserve the rank order of dissimilarities.
- Uses a monotonic transformation between observed dissimilarities and fitted distances.
- Is useful for ranked similarity judgments.
Stress measures the disagreement between fitted spatial distances and disparities derived from observed dissimilarities. A common form is:
where represents fitted distances and represents disparities. Lower stress indicates a better fit. However, increasing the number of dimensions usually reduces stress, so fit must be balanced against interpretability and parsimony.
What is conjoint analysis? Explain its basic concepts and research applications.
Conjoint analysis is a multivariate dependence technique used to estimate how respondents value the attributes and levels that jointly define a product, service, or policy alternative.
Basic concepts:
- Attribute: A characteristic such as price, brand, size, or delivery time.
- Level: A specific value of an attribute, such as a low, medium, or high price.
- Profile: A combination of one level from each attribute.
- Part-worth utility: The estimated contribution of a particular attribute level to overall preference.
- Attribute importance: The relative influence of an attribute, usually derived from the range of its part-worth utilities.
Applications include:
- Product and service design.
- Pricing decisions.
- Market simulation and demand estimation.
- Feature prioritization.
- Brand-positioning research.
- Estimation of trade-offs, such as how much price increase respondents will accept for improved quality.
Its central advantage is that it approximates real decisions by requiring respondents to evaluate complete alternatives rather than isolated attributes.
Describe the complete procedure of conjoint analysis and explain how part-worth utilities and attribute importance are calculated.
A conjoint study generally follows these steps:
- Define the decision problem: Specify the product, service, or policy choice to be studied.
- Select attributes and levels: Choose relevant, actionable, and realistic levels while avoiding excessive overlap.
- Choose a conjoint approach: Options include full-profile, choice-based, adaptive, and ratings- or rankings-based conjoint analysis.
- Construct profiles: Use a full factorial or an efficient fractional-factorial design to create alternatives.
- Collect preferences: Ask respondents to rate, rank, or choose among profiles.
- Estimate utilities: Apply regression, logit, hierarchical Bayes, or another model suited to the response format.
- Assess model fit: Compare predicted and observed preferences and use holdout profiles where possible.
- Conduct simulations: Estimate preference shares for alternative product configurations.
In an additive model, total utility is:
where is the part-worth utility of level of attribute .
The importance of attribute is calculated from its utility range:
Large utility ranges indicate that changes in the attribute exert a strong influence on preference.
Define multivariate analysis and explain how multivariate techniques may be classified.
Multivariate analysis refers to a collection of statistical techniques used to analyze three or more variables simultaneously. It helps researchers understand complex relationships, reduce data, classify observations, and test dependence among variables.
Multivariate techniques can be classified as follows:
- Dependence techniques: One or more variables are treated as dependent variables and are predicted or explained by independent variables. Examples include discriminant analysis, multivariate regression, and conjoint analysis.
- Interdependence techniques: No variable is designated as dependent. The objective is to discover an underlying structure among variables or objects. Examples include factor analysis, cluster analysis, and multidimensional scaling.
- Metric-data techniques: These are applied when variables are measured on interval or ratio scales.
- Non-metric-data techniques: These are used with nominal or ordinal variables.
The selection of a technique depends on the research objective, measurement scale, number of dependent variables, and assumptions about the data.
Did this save you a night before the exam?
LPU Notes is free, and it stays free. Ads cover part of the server bill. The rest comes out of a student's own pocket: the domain, the storage, and keeping the site up through the weeks everyone needs it at once.
The payment button didn't load. An ad blocker or a filtered network is the usual reason. to try again.
Nothing here is ever locked, and nothing unlocks. Chip in only if it was worth it. What it pays for →