Unit 5: Role of Statistics in Data Science - Practice Quiz

ECAP792 60 Questions
0 Correct 0 Wrong 60 Left
0/60

1 What is the main purpose of hypothesis testing?

Hypothesis testing Easy
A. To create charts from raw data
B. To evaluate a claim using sample data
C. To organize records in a database
D. To remove all variation from data

2 What does the null hypothesis usually state?

Null hypothesis Easy
A. Every observation is an outlier
B. There is no effect or difference
C. There is a strong positive effect
D. Every variable has equal values

3 What does the alternative hypothesis generally propose?

Alternative hypothesis Easy
A. An effect or difference exists
B. The population size is unknown
C. No relationship can ever exist
D. The sample has no observations

4 A result is commonly considered statistically significant when which condition is met?

Statistical significance Easy
A. The data contains no missing values
B. The sample size equals the population size
C. The mean equals the median exactly
D. The p-value is below the significance level

5 What is a Type 1 error?

Type 1 and type 2 errors Easy
A. Recording an incorrect sample size
B. Accepting a true null hypothesis
C. Rejecting a true null hypothesis
D. Rejecting a false null hypothesis

6 What is a Type 2 error?

Type 1 and type 2 errors Easy
A. Rejecting a true null hypothesis
B. Selecting an unsuitable chart type
C. Failing to reject a false null hypothesis
D. Rejecting a false null hypothesis

7 What does a p-value measure?

p-value Easy
A. The number of variables collected
B. Evidence against the null hypothesis
C. The exact size of the population
D. The accuracy of every observation

8 If and , what is the usual decision?

p-value Easy
A. Increase the p-value to
B. Conclude that no data exists
C. Reject the null hypothesis
D. Retain every possible hypothesis

9 What is ANOVA mainly used to compare?

ANOVA Easy
A. Means of three or more groups
B. Sizes of two database tables
C. Names of three or more variables
D. Colors of two different charts

10 What does ANOVA stand for?

ANOVA Easy
A. Assessment of Variable Accuracy
B. Analysis of Visual Attributes
C. Analysis of Variance
D. Arrangement of Variable Averages

11 What does the null hypothesis in a basic one-way ANOVA state?

ANOVA Easy
A. All observations are identical
B. All group sizes are equal
C. All variables are categorical
D. All group means are equal

12 Which type of data is commonly analyzed using a chi-square test?

Chi-square test Easy
A. Unstructured image data
B. Continuous measurement data
C. Ordered time-series data
D. Categorical frequency data

13 A chi-square test of independence examines whether two categorical variables are what?

Chi-square test Easy
A. Associated with each other
B. Normally distributed together
C. Equal in numerical value
D. Measured in identical units

14 How does statistics support data science?

Data science Easy
A. It guarantees every prediction is correct
B. It replaces the need to collect data
C. It converts all data into images
D. It helps analyze data and draw conclusions

15 What symbol commonly represents the significance level?

Statistical significance Easy
A.
B.
C.
D.

16 Which two statements are directly compared in hypothesis testing?

Hypothesis testing Easy
A. Mean and median values
B. Sample and chart titles
C. Rows and column labels
D. Null and alternative hypotheses

17 Which everyday term best describes a Type 1 error?

Type 1 and type 2 errors Easy
A. True negative
B. True positive
C. False negative
D. False positive

18 Which everyday term best describes a Type 2 error?

Type 1 and type 2 errors Easy
A. False negative
B. True negative
C. True positive
D. False positive

19 Which p-value provides stronger evidence against the null hypothesis?

p-value Easy
A.
B.
C.
D.

20 What does the null hypothesis in a chi-square test of independence usually state?

Chi-square test Easy
A. The numerical variables are normally distributed
B. The categorical variables have equal names
C. The numerical variables have equal means
D. The categorical variables are independent

21 A data scientist measures the prediction latency of the same 30 servers before and after an optimization. Which test is most appropriate for comparing the mean latencies?

Hypothesis testing Medium
A. Chi-square independence test
B. One-way ANOVA
C. Paired-samples t-test
D. Independent-samples t-test

22 An online retailer tests whether a redesigned page changes the conversion rate from its current value of 8%. Which null hypothesis is appropriate?

Null hypothesis Medium
A.
B.
C.
D.

23 A company will adopt a new recommendation algorithm only if it increases mean revenue above the current value of per user. Which alternative hypothesis should be used?

Alternative hypothesis Medium
A.
B.
C.
D.

24 A hypothesis test produces a p-value of with significance level . What is the appropriate decision?

Statistical significance Medium
A. Retain as statistically significant
B. Reject as statistically significant
C. Increase before deciding
D. Reject as practically significant

25 In a fraud-detection test, states that a transaction is legitimate. What is a Type I error?

Type 1 and type 2 errors Medium
A. Approving a fraudulent transaction as legitimate
B. Approving a legitimate transaction as legitimate
C. Flagging a legitimate transaction as fraud
D. Flagging a fraudulent transaction as fraud

26 A medical screening study uses : the patient does not have the condition. Which outcome represents a Type II error?

Type 1 and type 2 errors Medium
A. Missing a patient who has the condition
B. Correctly clearing a patient without the condition
C. Flagging a patient without the condition
D. Correctly identifying a patient with the condition

27 A data scientist wants to reduce the probability of a Type II error while keeping fixed. Which change is generally most effective?

Type 1 and type 2 errors Medium
A. Increase measurement noise
B. Increase the sample size
C. Use fewer observations
D. Decrease the sample size

28 A model comparison produces a p-value of . Which interpretation is correct?

p-value Medium
A. Assuming is true, equally or more extreme data have probability 0.02
B. The observed effect has a 0.02 probability of being useful
C. The probability that is true is exactly 0.02
D. The probability that is false is exactly 0.02

29 A test returns . At which significance levels would the result be statistically significant?

p-value Medium
A. and
B. and
C. and
D. and

30 A data scientist wants to compare mean customer spending across four membership tiers. Which test is the most appropriate initial analysis?

ANOVA Medium
A. Paired-samples t-test
B. One-way ANOVA
C. Pearson correlation test
D. Chi-square goodness-of-fit test

31 Which set of conditions best matches the standard assumptions of a one-way ANOVA?

ANOVA Medium
A. Paired observations, uniform residuals, and unequal medians
B. Independent categories, binary residuals, and equal frequencies
C. Independent observations, normal residuals, and similar variances
D. Repeated observations, skewed residuals, and unequal ranges

32 An ANOVA comparing five model configurations is statistically significant. What should the data scientist do next to identify which means differ?

ANOVA Medium
A. Accept that all five means differ
B. Replace every mean with its median
C. Apply a post-hoc multiple-comparison test
D. Run one test on the combined sample

33 A researcher examines whether preferred device type (mobile, tablet, desktop) is associated with subscription status (free, paid). Which test is appropriate?

Chi-square test Medium
A. One-way repeated-measures ANOVA
B. Chi-square test of independence
C. One-sample t-test
D. Paired-samples t-test

34 In a contingency table, a row total is 120, a column total is 50, and the grand total is 200. What is the expected count for their intersecting cell?

Chi-square test Medium
A.
B.
C.
D.

35 A die is rolled 120 times to determine whether its outcomes follow a uniform distribution. Which statistical test should be used?

Chi-square test Medium
A. One-way repeated-measures ANOVA
B. Chi-square goodness-of-fit test
C. Two-sample t-test
D. Chi-square test of independence

36 What are the degrees of freedom for a chi-square test of independence using a contingency table?

Chi-square test Medium
A.
B.
C.
D.

37 A machine-learning model has 92% validation accuracy, while another has 91.8%. Why should a data scientist perform a statistical comparison?

Data science Medium
A. To guarantee the first model will perform best in production
B. To prove that validation accuracy measures business value
C. To remove the need for an independent test dataset
D. To assess whether the difference may reflect sampling variation

38 A team independently tests 10 model features and wants a family-wise significance level of 0.05 using the Bonferroni correction. What threshold should each test use?

Hypothesis testing Medium
A.
B.
C.
D.

39 With a very large dataset, a new model reduces mean error by 0.001 units and produces . What is the best conclusion?

Statistical significance Medium
A. The reduction is statistically significant but may lack practical importance
B. The reduction is not statistically significant because it is small
C. The reduction is practically important because the p-value is small
D. The reduction proves the model will improve every prediction

40 In a one-way ANOVA, what does a large -statistic generally indicate?

ANOVA Medium
A. The response and group variables are both categorical
B. Between-group variation is large relative to within-group variation
C. All group means and variances are necessarily equal
D. Within-group variation is large relative to between-group variation

41 A researcher observes a symmetric test statistic and chooses the one-sided alternative only after seeing whether is positive or negative. They then reject whenever the selected one-sided p-value is below . Under , what is the actual Type I error rate?

Hypothesis testing Hard
A. Approximately
B. Approximately
C. Approximately
D. Approximately

42 For a composite null hypothesis , which p-value construction most directly guarantees that a level- test controls Type I error for every ?

Null hypothesis Hard
A. for unrestricted
B.
C. for null
D.

43 To establish that a mean difference is practically equivalent to zero within a margin , which hypotheses are appropriate for the two one-sided tests procedure?

Alternative hypothesis Hard
A. versus
B. versus
C. versus
D. versus

44 Twenty independent parameters each have a pointwise confidence interval. If all model assumptions hold, what is the probability that all 20 intervals simultaneously contain their parameters?

Statistical significance Hard
A.
B.
C.
D.

45 For a fixed nonzero effect size and correctly specified model, sample size is increased while the test threshold is recalibrated to keep . What generally happens?

Type 1 and type 2 errors Hard
A. Type I error decreases and Type II error stays fixed
B. Type I error stays fixed and Type II error increases
C. Type I error increases and Type II error decreases
D. Type I error stays fixed and Type II error decreases

46 A data scientist performs 20 independent hypothesis tests, each at significance level , and every null hypothesis is true. What is the probability of at least one Type I error?

Type 1 and type 2 errors Hard
A.
B.
C.
D.

47 A valid test produces . Which interpretation is correct?

p-value Hard
A. The probability that the observed result arose by chance equals
B. The probability that is true equals
C. If repeated, the probability of replication equals
D. Under , results at least as extreme have probability

48 Under , a one-sided test rejects for unusually large . If is observed, what is the exact p-value?

p-value Hard
A.
B.
C.
D.

49 In a balanced ANOVA, treatment B exceeds treatment A by units at the low level of a second factor, but treatment B is units below treatment A at the high level. Which conclusion is most defensible?

ANOVA Hard
A. The interaction must vanish because effects have equal magnitude
B. The treatment must have a negative main effect
C. The treatment must have a positive main effect
D. The treatment main effect can vanish despite an interaction

50 Three groups each contain four observations. Their sample means are , , and , while the pooled within-group sum of squares is . What is the one-way ANOVA statistic ?

ANOVA Hard
A.
B.
C.
D.

51 A study records five repeated measurements from each participant but analyzes all measurements using ordinary one-way ANOVA as if they came from different people. With positive within-participant correlation, what is the main inferential risk?

ANOVA Hard
A. Group means are necessarily biased toward zero
B. Standard errors are underestimated and Type I error is inflated
C. Residual variance is necessarily estimated without bias
D. Standard errors are overestimated and Type I error is inflated

52 A one-way comparison has strongly unequal group variances and unequal sample sizes, with the largest variance occurring in the smallest group. Which omnibus procedure is generally most appropriate?

ANOVA Hard
A. Repeated-measures ANOVA using participant-level blocks
B. Classical ANOVA using the pooled within-group variance
C. Welch's ANOVA using variance-adjusted group weights
D. Two-way ANOVA using variance as another factor

53 A chi-square goodness-of-fit test uses six categories. Two distribution parameters are estimated from the same observations before expected counts are calculated. What are the test's degrees of freedom?

Chi-square test Hard
A.
B.
C.
D.

54 A table has observed counts . What is Pearson's chi-square statistic without continuity correction?

Chi-square test Hard
A.
B.
C.
D.

55 Binary outcomes are measured on the same individuals before and after an intervention. Which test appropriately assesses marginal change while accounting for pairing?

Chi-square test Hard
A. Pearson's chi-square independence test on all four cells
B. Chi-square goodness-of-fit test using diagonal pairs
C. One-way ANOVA applied to the four cell frequencies
D. McNemar's chi-square test using discordant pairs

56 A feature-selection algorithm is run once on the complete labeled dataset, after which cross-validation evaluates a model using only the selected features. Why is the estimated performance generally optimistic?

Data science Hard
A. Validation labels influenced selection before the folds were evaluated
B. Cross-validation requires every available feature in every model
C. The folds become unequal because selected features change sample counts
D. Feature selection always increases the model's irreducible error

57 Suppose outcomes are missing at random conditional on observed covariates , and every observation has a positive response probability. Which method can consistently estimate the population outcome mean if its required model is correctly specified?

Data science Hard
A. Complete-case averaging without using
B. Mean imputation using only observed outcomes
C. Inverse-probability weighting using modeled response probabilities
D. Deleting covariates associated with the missingness indicator

58 Five ordered p-values are . Using the Benjamini–Hochberg procedure with false discovery rate , how many hypotheses are rejected?

Statistical significance Hard
A. Four hypotheses
B. Two hypotheses
C. Five hypotheses
D. Three hypotheses

59 For four equally sized groups, which pair of contrasts is orthogonal?

ANOVA Hard
A. and
B. and
C. and
D. and

60 An aggregated chi-square test shows a strong association between treatment and recovery, but the association largely disappears within every hospital. Hospital is related to both treatment assignment and recovery. What is the best next analysis?

Chi-square test Hard
A. Combine sparse hospitals until the aggregate association becomes stronger
B. Repeat the aggregated test with a smaller significance level
C. Apply a goodness-of-fit test only to the treatment totals
D. Use a stratified method such as the Cochran–Mantel–Haenszel test