Unit 2: Data analysis - Practice Quiz

BTY587 — Data Analysis And Simulations 60 Questions
0 Correct 0 Wrong 60 Left
0/60

1 What does descriptive statistics primarily aim to do?

statistical analysis Easy
A. Generate random numbers for simulations
B. Encrypt data for secure storage
C. Summarize and describe the main features of a dataset
D. Prove that a theory is universally true

2 Which measure of central tendency is most affected by extreme outliers?

statistical analysis Easy
A. Median
B. Mode
C. Mean
D. Range

3 The standard deviation is a measure of:

statistical analysis Easy
A. The most frequent value in the data
B. The middle value of the data
C. The total sum of all data points
D. The spread or dispersion of data around the mean

4 In hypothesis testing, the null hypothesis () typically represents:

hypothesis testing Easy
A. The alternative claim being tested
B. A guaranteed positive result
C. No effect or no difference
D. The largest possible effect size

5 A Type I error occurs when we:

hypothesis testing Easy
A. Fail to collect enough data
B. Accept a true null hypothesis
C. Reject a true null hypothesis
D. Reject a false null hypothesis

6 The significance level of a test is usually denoted by which symbol?

hypothesis testing Easy
A.
B.
C.
D.

7 A p-value represents the probability of:

significance of p-value Easy
A. Observing results as extreme as the data, assuming is true
B. The alternative hypothesis being true
C. The null hypothesis being true
D. Making a correct decision in the test

8 If a p-value is less than the significance level , we typically:

significance of p-value Easy
A. Accept the null hypothesis
B. Change the alternative hypothesis
C. Reject the null hypothesis
D. Increase the sample size

9 A smaller p-value indicates:

significance of p-value Easy
A. Stronger evidence for the null hypothesis
B. A larger sample was required
C. Stronger evidence against the null hypothesis
D. A guaranteed real effect exists

10 The chi-square test is most commonly used to analyze:

chi-square Easy
A. Correlation between two continuous variables
B. Time-series trends
C. Categorical (frequency) data
D. Continuous numeric averages

11 A chi-square test of independence checks whether:

chi-square Easy
A. Two means are equal
B. Two categorical variables are associated
C. A sample follows a normal distribution
D. Three or more group variances differ

12 In the chi-square statistic , what does represent?

chi-square Easy
A. The expected frequency
B. The number of outcomes
C. The observed frequency
D. The overall mean

13 A t-test is generally used to compare:

t-test Easy
A. The variances of three or more groups
B. The correlation of two datasets
C. The frequencies of categorical data
D. The means of one or two groups

14 Which type of t-test compares the means of two related measurements from the same subjects?

t-test Easy
A. Welch's t-test
B. Independent samples t-test
C. One-sample t-test
D. Paired t-test

15 The t-test is preferred over the z-test mainly when:

t-test Easy
A. The population is infinitely large
B. More than three groups are compared
C. The sample size is small and population variance is unknown
D. The data is purely categorical

16 ANOVA is primarily used to compare:

ANOVA Easy
A. Two categorical variables
B. The spread of a single dataset
C. The median of two groups
D. Means of three or more groups

17 The test statistic used in ANOVA is the:

ANOVA Easy
A. z-statistic
B. Chi-square statistic
C. t-statistic
D. F-statistic

18 What does the null hypothesis of a one-way ANOVA state?

ANOVA Easy
A. All group variances are zero
B. All group means are equal
C. The groups are categorical
D. At least one mean is different

19 Bayes' theorem is used to update the probability of a hypothesis based on:

bayesian probability Easy
A. The sample size only
B. New evidence or data
C. The number of groups compared
D. The significance level

20 In Bayesian probability, the probability assigned before observing any new data is called the:

bayesian probability Easy
A. Likelihood
B. Prior probability
C. Posterior probability
D. Marginal probability

21 A researcher sets a significance level of and obtains a test statistic that leads to a p-value of . What is the correct decision?

hypothesis testing Medium
A. Increase the sample size before deciding
B. Accept the null hypothesis as true
C. Fail to reject the null hypothesis
D. Reject the null hypothesis

22 Which statement best describes the meaning of a p-value in hypothesis testing?

significance of p-value Medium
A. The probability that the alternative hypothesis is correct
B. The probability that the null hypothesis is true given the data
C. The probability of making a Type II error in the test
D. The probability of observing data as extreme as the sample, assuming the null hypothesis is true

23 You want to compare the mean test scores of two independent groups with small samples and unknown, unequal variances. Which test is most appropriate?

t-test Medium
A. One-sample z-test
B. Chi-square test of independence
C. Welch's two-sample t-test
D. Paired t-test

24 A chi-square goodness-of-fit test compares observed frequencies with expected frequencies . What is the test statistic formula?

chi-square Medium
A.
B.
C.
D.

25 When comparing the means of four different groups simultaneously, why is one-way ANOVA preferred over multiple t-tests?

ANOVA Medium
A. It works only when variances are unequal
B. It controls the inflation of the overall Type I error rate
C. It eliminates the need for a null hypothesis
D. It requires fewer assumptions about the data

26 A disease affects of a population. A test is sensitive and specific. Using Bayes' theorem, which quantity is the posterior probability of disease given a positive test?

bayesian probability Medium
A.
B.
C.
D.

27 A test fails to reject the null hypothesis even though the alternative is actually true. What type of error has occurred?

hypothesis testing Medium
A. Sampling error
B. Type I error
C. Standard error
D. Type II error

28 In a study measuring blood pressure of the same patients before and after treatment, which t-test correctly accounts for the data structure?

t-test Medium
A. One-way ANOVA
B. Independent two-sample t-test
C. Welch's t-test
D. Paired t-test

29 A result yields with . Which interpretation is most accurate?

significance of p-value Medium
A. The result proves the alternative hypothesis is true
B. The effect size is guaranteed to be large
C. The result is statistically significant but the effect may still be small
D. There is a chance the alternative is correct

30 A chi-square test of independence uses a contingency table. How many degrees of freedom does it have?

chi-square Medium
A.
B.
C.
D.

31 In one-way ANOVA, what does a large F-statistic generally indicate?

ANOVA Medium
A. Within-group variance is large relative to between-group variance
B. All group means are exactly equal
C. The sample sizes are too small to analyze
D. Between-group variance is large relative to within-group variance

32 In Bayesian inference, how is the posterior distribution related to the prior and the likelihood?

bayesian probability Medium
A. Posterior Likelihood Prior
B. Posterior Prior Likelihood
C. Posterior Prior Likelihood
D. Posterior Likelihood Prior

33 A dataset is strongly right-skewed with several large outliers. Which measure of central tendency best represents its typical value?

statistical analysis Medium
A. Standard deviation
B. Range
C. Mean
D. Median

34 Increasing the sample size in a hypothesis test, while keeping other factors fixed, generally has what effect on statistical power?

hypothesis testing Medium
A. Power decreases
B. Power increases
C. Power stays exactly the same
D. Power becomes undefined

35 A one-sample t-test gives with degrees of freedom and a two-tailed critical value of at . What is the conclusion?

t-test Medium
A. Switch to a chi-square test
B. The test is inconclusive
C. Reject the null hypothesis
D. Fail to reject the null hypothesis

36 In a chi-square test, expected cell counts are recommended to be at least . What is the main reason for this guideline?

chi-square Medium
A. To make the degrees of freedom larger
B. To guarantee the observed counts are whole numbers
C. To ensure the chi-square approximation to the sampling distribution is reliable
D. To force the null hypothesis to be true

37 After a significant one-way ANOVA result, why are post-hoc tests such as Tukey's HSD performed?

ANOVA Medium
A. To increase the between-group variance
B. To identify which specific group means differ from each other
C. To reduce the number of groups compared
D. To confirm the ANOVA assumptions were met

38 A prior probability of rain is . New forecast evidence has likelihood under rain and under no rain. Which value is the denominator (evidence) in Bayes' theorem for updating?

bayesian probability Medium
A.
B.
C.
D.

39 Two variables have a Pearson correlation of . Which conclusion is justified?

statistical analysis Medium
A. There is a strong positive linear association between the variables
B. The relationship is strongly negative
C. The variables are unrelated
D. One variable definitely causes the other

40 A researcher lowers the significance threshold from to . What is the direct effect on the two error types?

significance of p-value Medium
A. Both error probabilities decrease equally
B. Neither error probability is affected
C. Type I error probability increases while Type II decreases
D. Type I error probability decreases while Type II error probability tends to increase

41 A researcher obtains in a two-sided test with and declares the effect "real and important." Which statement is the most accurate critique?

significance of p-value Hard
A. The result cannot be significant because is too close to the threshold
B. A -value this small guarantees the finding will replicate
C. A -value just below neither quantifies effect size nor the probability the null is true; significance and importance are distinct
D. A means there is a probability the null hypothesis is true

42 In a fixed-sample test, holding effect size and sample size constant, what happens to the Type II error rate as you lower from to ?

hypothesis testing Hard
A. is unaffected by
B. increases (power decreases)
C. becomes equal to
D. decreases (power increases)

43 Two groups have equal sample sizes but markedly unequal variances. Using a Student's pooled-variance -test instead of Welch's -test primarily risks:

t-test Hard
A. Always inflating power regardless of variance ratio
B. Making the test invalid only when exceeds 30
C. Distorting the Type I error rate because the pooled variance misestimates the standard error
D. Producing an identical result to Welch's test in all cases

44 A one-way ANOVA across 4 groups yields a significant . What is the correct next step and why?

ANOVA Hard
A. Run post-hoc pairwise comparisons with a correction (e.g., Tukey HSD) because ANOVA only shows that at least one mean differs
B. Conclude only the largest and smallest means differ
C. Conclude all four means are pairwise different
D. Run 6 independent -tests at each with no correction

45 In a chi-square test of independence on a table, several expected cell counts fall below 5. The most appropriate response is to:

chi-square Hard
A. Use Fisher's exact test, since the chi-square approximation is unreliable with small expected counts
B. Report the chi-square statistic anyway with no adjustment
C. Convert counts to proportions and use a -test on raw percentages
D. Increase to compensate for the small counts

46 A disease affects of a population. A test has sensitivity and specificity. Given a positive result, the posterior probability of disease is approximately:

bayesian probability Hard
A.
B.
C.
D.

47 A dataset is strongly right-skewed with several extreme high values. Which pair of summary statistics best describes its center and spread?

statistical analysis Hard
A. Mean and standard deviation
B. Mean and variance
C. Mode and range
D. Median and interquartile range (IQR)

48 Researchers test 100 independent true null hypotheses each at . What is the expected number of false positives, and what does this illustrate?

hypothesis testing Hard
A. Exactly 1 false positive, fixed by
B. About 50 false positives, because compounds multiplicatively
C. About 5 false positives, illustrating the multiple-comparisons problem
D. Zero false positives, because all nulls are true

49 A paired-samples design (before/after on the same subjects) is incorrectly analyzed with an independent two-sample -test. The typical consequence is:

t-test Hard
A. The test becomes invalid only for negative correlations
B. An identical -statistic to the paired test
C. Automatic Type I error inflation regardless of correlation
D. Loss of power because the analysis ignores within-subject correlation and inflates the standard error

50 In a one-way ANOVA, if the between-group variability is small relative to within-group variability, the -statistic will be:

ANOVA Hard
A. Very large, indicating strong group differences
B. Negative, indicating error in computation
C. Close to 1, suggesting group means do not differ beyond chance
D. Exactly 0 whenever means are unequal

51 Which interpretation of a confidence interval that excludes 0 for a mean difference is correct?

significance of p-value Hard
A. The -value must be exactly
B. The corresponding two-sided test rejects at
C. There is a probability the true difference lies in this specific interval
D. The true difference is definitely nonzero

52 A chi-square goodness-of-fit test compares observed counts across 6 categories to a theoretical distribution with no estimated parameters. The degrees of freedom are:

chi-square Hard
A. 1
B. 6
C. 5
D. 4

53 In Bayesian inference, as the sample size grows very large, the influence of a proper prior on the posterior typically:

bayesian probability Hard
A. Becomes undefined for continuous parameters
B. Diminishes, so the posterior is dominated by the likelihood
C. Stays constant regardless of data
D. Grows, so the prior overwhelms the data

54 A strong positive linear correlation () is found between ice-cream sales and drowning incidents. The best explanation is:

statistical analysis Hard
A. Drownings cause increased ice-cream sales
B. The correlation must be a computational artifact
C. A likely confounder (e.g., temperature) drives both; correlation does not imply causation
D. Ice-cream consumption causes drowning

55 For a one-sample -test with , the degrees of freedom and the reason are:

t-test Hard
A. , because two parameters are estimated
B. , because equals the sample size
C. , because
D. , because one parameter (the sample mean) is estimated from the data

56 A study is 'underpowered.' What is the most direct risk this poses to interpreting a non-significant result?

hypothesis testing Hard
A. The -value is guaranteed to be below
B. The Type I error rate is inflated above
C. The confidence interval will be too narrow
D. A true effect may exist but go undetected (high Type II error)

57 A two-way ANOVA shows a significant interaction between Factor A and Factor B. How should the main effects be interpreted?

ANOVA Hard
A. As fully independent and directly interpretable
B. As identical to the interaction effect
C. With caution, because the effect of one factor depends on the level of the other
D. As automatically non-significant

58 Which practice most directly leads to 'p-hacking' and inflated false-positive rates?

significance of p-value Hard
A. Trying many analyses and reporting only those with
B. Pre-registering a single hypothesis before data collection
C. Correcting for multiple comparisons
D. Reporting effect sizes with confidence intervals

59 For a contingency table with rows and columns, the degrees of freedom for the chi-square test of independence are:

chi-square Hard
A.
B.
C.
D.

60 You flip a coin 10 times and get 8 heads. Using a uniform prior on the head probability , the posterior distribution is:

bayesian probability Hard
A.
B.
C.
D.