Unit 4: UNSUPERVISED LEARNING: CLUSTERING AND PATTERN DETECTION - Practice Quiz

INT234 — Predictive Analytics 60 Questions
0 Correct 0 Wrong 60 Left
0/60

1 What is the main purpose of K-Means clustering?

K-Means Clustering Easy
A. To arrange every data point into a fixed sequence based on its original row number
B. To predict continuous values
C. To label training examples
D. To group similar data points

2 In K-Means clustering, what does represent?

K-Means Clustering Easy
A. The number of features
B. The number of iterations
C. The number of labeled examples used to train the clustering model
D. The number of clusters

3 What is a centroid in K-Means clustering?

K-Means Clustering Easy
A. The mean position of a cluster
B. The boundary that separates every pair of clusters in the complete dataset
C. The label assigned to a cluster
D. The most distant point in a cluster

4 How does K-Means usually assign a data point to a cluster?

K-means clustering intuition Easy
A. By choosing the largest cluster
B. By choosing a random label
C. By choosing the centroid with the greatest number of features and observations
D. By choosing the nearest centroid

5 What happens after data points are assigned to their nearest centroids?

K-means clustering intuition Easy
A. The data points are permanently removed from clusters that contain fewer observations
B. The features are deleted
C. The centroids are recomputed
D. The cluster count is doubled

6 What is the K-Means random initialization trap?

K-means random initialization trap Easy
A. Random initialization automatically finds the globally best clustering in every dataset
B. Random data always creates one cluster
C. Different initial centroids can produce different clusters
D. Centroids cannot move after initialization

7 Which method is commonly used to choose better initial centroids for K-Means?

K-means random initialization trap Easy
A. Single linkage
B. Ordinary least squares regression with randomly generated target values
C. K-Means++
D. Apriori

8 Which method plots clustering error against different values of ?

K-means selecting number of clusters Easy
A. Elbow method
B. Apriori method
C. Linkage method
D. A supervised classification method that compares predicted labels with known labels

9 In the elbow method, which value of is usually selected?

K-means selecting number of clusters Easy
A. The number of data features
B. The point where improvement begins to slow
C. The point where the clustering error becomes larger than the original dataset size
D. The largest possible value

10 How does agglomerative hierarchical clustering begin?

Hierarchical Clustering (Agglomerative) Easy
A. With all points in one cluster
B. With two randomly labeled clusters
C. With one cluster for every feature and one additional cluster for every target variable
D. With each point in its own cluster

11 What is a dendrogram?

Hierarchical Clustering (Agglomerative) Easy
A. A chart of regression errors
B. A table of class labels
C. A diagram that displays only the final centroid coordinates produced by K-Means
D. A tree diagram of cluster merges

12 How is the distance between two clusters defined in single linkage?

Types of Linkages (Single, Complete, Average, Centroid) Easy
A. The average pairwise distance
B. The sum of all distances from every observation to the overall dataset mean
C. The minimum pairwise distance
D. The maximum pairwise distance

13 How is the distance between two clusters defined in complete linkage?

Types of Linkages (Single, Complete, Average, Centroid) Easy
A. The distance between the first observations recorded in each of the two clusters
B. The median pairwise distance
C. The maximum pairwise distance
D. The minimum pairwise distance

14 What does average linkage use to compare two clusters?

Types of Linkages (Single, Complete, Average, Centroid) Easy
A. The mean of all pairwise distances
B. The smallest pairwise distance
C. The mean distance between each observation and every centroid in the full dataset
D. The largest pairwise distance

15 What does centroid linkage compare?

Types of Linkages (Single, Complete, Average, Centroid) Easy
A. The complete ordered lists of all observations and features in both clusters
B. The nearest points of two clusters
C. The centroids of two clusters
D. The farthest points of two clusters

16 Which form represents a typical association rule?

Association Rules Easy
A. and are assigned to separate supervised learning target classes
B.
C.
D.

17 What does the support of an itemset measure?

Association Rules Easy
A. How many clusters are generated after every iteration of a hierarchical algorithm
B. How far an item is from a centroid
C. How frequently the itemset occurs
D. How accurately a class is predicted

18 What does the confidence of the rule measure?

Association Rules Easy
A. How often occurs when occurs
B. How often occurs without
C. How many items appear in
D. How closely the items in are located to the centroid of itemset

19 In unsupervised pattern detection, what is the algorithm generally trying to discover?

Finding Patterns Easy
A. A fixed mathematical formula supplied in advance for every possible observation
B. Only missing values in a table
C. Recurring structures in unlabeled data
D. Known labels in training data

20 What is the main goal of market basket analysis?

Market Basket Analysis Using Association Rules Easy
A. To predict exact product prices
B. To divide customers by age only
C. To identify products often purchased together
D. To calculate the physical distance between products placed on different store shelves

21 A K-means model has centroids and . Using Euclidean distance, to which cluster should the point be assigned?

K-Means Clustering Medium
A. Cluster
B. Both clusters equally
C. Neither cluster
D. Cluster

22 After an assignment step, one cluster contains the points , , and . What centroid will be used in the next iteration?

K-Means Clustering Medium
A.
B.
C.
D.

23 During K-means training, several points change cluster assignments after the centroids are updated. What is the most appropriate next action?

K-means clustering intuition Medium
A. Remove the reassigned observations
B. Increase by one immediately
C. Repeat assignment and centroid updates
D. Stop because centroids were updated once

24 A dataset contains two compact circular groups and one long curved group. Why might K-means represent the curved group poorly?

K-means clustering intuition Medium
A. It favors centroid-based convex clusters
B. It ignores distances between observations
C. It requires all features to be categorical
D. It merges points using maximum distance

25 Two K-means runs with the same value of produce within-cluster sums of squares of and . Assuming both runs converged, which result should generally be preferred?

K-means random initialization trap Medium
A. The run with value
B. The run with value
C. Neither run without increasing
D. Either run because is equal

26 A K-means solution changes substantially across repeated runs on the same normalized dataset. Which strategy most directly reduces this initialization problem?

K-means random initialization trap Medium
A. Use K-means++ with multiple restarts
B. Stop after the first assignment step
C. Select all centroids from one region
D. Replace distances with cluster labels

27 The within-cluster sums of squares for are . According to the elbow method, which value of is most reasonable?

K-means selecting number of clusters Medium
A.
B.
C.
D.

28 The average silhouette scores for are . Which cluster count is supported by this criterion?

K-means selecting number of clusters Medium
A.
B.
C.
D.

29 Four observations have pairwise distances , , , , , and . Which pair is merged first in agglomerative clustering?

Hierarchical Clustering (Agglomerative) Medium
A. and
B. and
C. and
D. and

30 A dendrogram has a large jump in merge height from to . What is the usual interpretation when choosing a cut?

Hierarchical Clustering (Agglomerative) Medium
A. Cut between and
B. Cut exactly at height
C. Cut above the final merge
D. Cut below the first merge

31 For clusters and , the cross-cluster distances are and . What are the single-linkage and complete-linkage distances, respectively?

Types of Linkages (Single, Complete, Average, Centroid) Medium
A. and
B. and
C. and
D. and

32 The cross-cluster distances between two clusters are and . What distance does average linkage use?

Types of Linkages (Single, Complete, Average, Centroid) Medium
A.
B.
C.
D.

33 Cluster has centroid and cluster has centroid . What is their centroid-linkage distance using Euclidean distance?

Types of Linkages (Single, Complete, Average, Centroid) Medium
A.
B.
C.
D.

34 A clustering result produced by single linkage contains long chains connecting otherwise distinct dense groups. Which linkage would most directly encourage more compact clusters?

Types of Linkages (Single, Complete, Average, Centroid) Medium
A. Nearest linkage
B. Single linkage
C. Complete linkage
D. Minimum linkage

35 In transactions, contain coffee, contain cookies, and contain both. What is the confidence of the rule ?

Association Rules Medium
A.
B.
C.
D.

36 A rule has confidence , while the overall support of is . What is the lift of the rule?

Association Rules Medium
A.
B.
C.
D.

37 An itemset appears in of transactions. If the minimum support threshold is , how should the itemset be classified?

Finding Patterns Medium
A. Frequent with support
B. Infrequent with support
C. Frequent with support
D. Infrequent with support

38 The itemset is found to be infrequent under a fixed minimum support. According to the Apriori principle, which conclusion is valid?

Finding Patterns Medium
A. Every superset containing is frequent
B. Every subset containing two items is infrequent
C. Every superset containing is infrequent
D. Every subset containing one item is infrequent

39 A supermarket finds that the rule has support , confidence , and lift . Which interpretation is most appropriate?

Market Basket Analysis Using Association Rules Medium
A. Eight percent of pasta baskets exclude sauce
B. Pasta and sauce occur independently in all baskets
C. Sauce buyers purchase pasta with probability
D. Pasta buyers purchase sauce more often than average

40 A retailer wants a cross-selling rule that applies to many baskets and strongly increases the likelihood of the consequent. Which rule is the best choice?

Market Basket Analysis Using Association Rules Medium
A. Support , confidence , lift
B. Support , confidence , lift
C. Support , confidence , lift
D. Support , confidence , lift

41 For the one-dimensional data , consider all possible contiguous two-cluster partitions. Which partition globally minimizes the K-means within-cluster sum of squares?

K-Means Clustering Hard
A. and
B. and
C. and
D. and

42 A dataset has two variables: annual spending measured in dollars and satisfaction measured on a - scale. Without standardization, K-means produces clusters almost entirely separated by spending. Which explanation is most accurate?

K-means clustering intuition Hard
A. Euclidean distances are dominated by the variable with the larger numerical scale
B. K-means ignores variables whose values are bounded above and below
C. Centroid updates automatically assign greater weight to continuous variables
D. Satisfaction contributes only when its distribution is approximately Gaussian

43 Standard K-means is run repeatedly with the same but different random initial centroids. Run X ends with objective , while Run Y ends with . Both runs have converged. What can be concluded?

K-means random initialization trap Hard
A. Run Y is globally optimal because its objective is lower
B. Run Y found a better local solution, but global optimality is not established
C. Run X and Run Y must represent the same partition up to label permutation
D. Run X has better-separated clusters because it converged sooner

44 In K-means++ initialization, the first centroid has been selected. Three remaining observations have distances , , and from their nearest selected centroid. What is the probability that the observation at distance is selected next?

K-means random initialization trap Hard
A.
B.
C.
D.

45 A K-means run initializes two centroids at identical coordinates. After the first assignment step, one cluster receives no observations because all ties are resolved in favor of the other cluster. Which response best preserves the intended value of ?

K-means random initialization trap Hard
A. Leave the empty centroid unchanged until assignments eventually change
B. Delete the empty cluster and continue with clusters
C. Reinitialize the empty centroid using a poorly represented observation
D. Set the empty centroid equal to the global feature mean

46 For an observation, the mean distance to other members of its cluster is , and the smallest mean distance to any other cluster is . What is its silhouette coefficient?

K-means selecting number of clusters Hard
A.
B.
C.
D.

47 The K-means objective decreases sharply through and only slightly afterward, but the average silhouette score is highest at . Which conclusion is most defensible?

K-means selecting number of clusters Hard
A. is certainly correct because silhouette scores identify the true classes
B. The criteria reflect different cluster properties, so domain validation is needed
C. is certainly correct because the elbow overrides all other criteria
D. Neither criterion is usable because the K-means objective always decreases with

48 Four one-dimensional observations are located at , , , and . Under single-linkage agglomerative clustering, at which heights do the three merges occur, assuming Euclidean distance?

Hierarchical Clustering (Agglomerative) Hard
A.
B.
C.
D.

49 An agglomerative dendrogram is cut horizontally at height . A merge occurring exactly at is treated as already completed. How many clusters remain if the merge heights are , , , , and , and ?

Hierarchical Clustering (Agglomerative) Hard
A.
B.
C.
D.

50 Clusters and contain and observations, respectively. Under average linkage, and . After merging and , what is ?

Types of Linkages (Single, Complete, Average, Centroid) Hard
A.
B.
C.
D.

51 Two compact, dense groups are connected by a sparse sequence of intermediate observations. Which linkage is most likely to merge the groups prematurely because of this structure?

Types of Linkages (Single, Complete, Average, Centroid) Hard
A. Complete linkage
B. Single linkage
C. Average linkage
D. Centroid linkage

52 Which linkage can produce a dendrogram inversion, where a later merge is recorded at a lower dissimilarity than an earlier merge?

Types of Linkages (Single, Complete, Average, Centroid) Hard
A. Unweighted average linkage
B. Single linkage
C. Centroid linkage
D. Complete linkage

53 For clusters and on the real line, which tuple gives their single-linkage, complete-linkage, average-linkage, and centroid-linkage distances, respectively?

Types of Linkages (Single, Complete, Average, Centroid) Hard
A.
B.
C.
D.

54 In transactions, itemset appears in , itemset appears in , and both appear in . What are the confidence and lift of the rule ?

Association Rules Hard
A. Confidence and lift
B. Confidence and lift
C. Confidence and lift
D. Confidence and lift

55 Suppose , , and . Which interpretation of the rule is correct?

Association Rules Hard
A. The rule is invalid because the consequent is too frequent
B. The rule is positively associated because its confidence is
C. The rule is negatively associated because its support is only
D. The rule shows independence because its lift equals

56 A rule has support , confidence , and . What is its leverage?

Association Rules Hard
A.
B.
C.
D.

57 During Apriori mining, the itemset is found to be infrequent. Which candidate can be pruned immediately without counting its support?

Finding Patterns Hard
A.
B.
C.
D.

58 An itemset is frequent. It has a frequent proper superset, but every proper superset has strictly lower support than . How should be classified?

Finding Patterns Hard
A. Closed but not maximal
B. Neither closed nor maximal
C. Maximal but not closed
D. Both closed and maximal

59 A retailer finds that the rule has confidence , while milk appears in of all baskets. Which conclusion is justified?

Market Basket Analysis Using Association Rules Hard
A. Bread slightly reduces the relative likelihood of milk because lift is below one
B. Bread strongly promotes milk because confidence is close to one
C. Bread and milk are positively associated because support must exceed
D. Bread and milk are independent because both items occur in most baskets

60 A store has baskets. Item occurs in , item in , and both in . The thresholds are minimum support , minimum confidence , and minimum lift . Does qualify?

Market Basket Analysis Using Association Rules Hard
A. No, because only its support meets the thresholds
B. Yes, because support, confidence, and lift all pass
C. No, because confidence and lift pass but support fails
D. No, because support and confidence pass but lift fails