Unit 5: Sampling Design

DEMGN832 — Research Methodology 9 min read

I. Foundations of Sampling Design

Sampling design is the definite plan used to select a subset of units from a population so that information from the sample can support conclusions about that population. It specifies who or what may be selected, how selection occurs, and how many units are included.

  • Population or universe: The complete set of elements relevant to the research problem—for example, all 4,000 students enrolled at a university.
  • Target population: The population to which the researcher intends to generalize findings, such as all undergraduate students in a state.
  • Accessible population: The portion that can practically be reached, such as students listed in participating institutions.
  • Element: The basic unit about which information is collected—for example, a person, household, firm, transaction, or document.
  • Sampling unit: The unit selected at a particular stage; it may be an individual, school, village, or geographic block.
  • Sampling frame: The operational list or representation of sampling units, such as an electoral roll or employee register.
  • Sample: The units actually selected from the population.
  • Parameter and statistic:
    • A parameter, such as population mean (\mu), describes the population.
    • A statistic, such as sample mean (\bar{x}), is calculated from sample observations.
  • Sampling error: Random difference between a sample estimate and the true population parameter.
  • Non-sampling error: Error caused by non-response, inaccurate measurement, recording mistakes, frame defects, or data-processing failures.
  • Central principle: A sample should provide sufficient information about the population at acceptable cost, time, and precision.

II. Planning the Sample — From Population to Selection

A. Sampling design process

The sampling design process converts a research objective into an operational procedure for selecting observable units.

  • Define the research problem: Identify the information required and the population parameter to be estimated; for example, estimating the average monthly expenditure of urban households.
  • Define the target population: Specify four boundaries:
    • Element: Urban household.
    • Geographical scope: A named city or region.
    • Time period: The current financial year.
    • Relevant characteristics: Occupied residential households.
  • Choose the sampling unit: Decide whether selection will involve individuals, households, institutions, areas, or clusters. In a school survey, schools may be first-stage units and students second-stage units.
  • Construct or identify the sampling frame: Evaluate the list for omissions, duplication, outdated entries, and inclusion of ineligible units. A defective frame creates coverage error even when selection is random.
  • Select the sampling approach: Choose probability or non-probability sampling according to the study’s purpose, availability of a frame, desired generalizability, and resources.
  • Determine sample size: For estimating a population proportion under simple random sampling, an initial size may be calculated as:
TEXT
n₀ = z²p(1 − p) / e²
  • (n_0) = initial sample size for a large population.
  • (z) = standard-normal value for the chosen confidence level.
  • (p) = anticipated population proportion.
  • (e) = permitted margin of error.
  • At 95% confidence, (z \approx 1.96); if (p) is unknown, (p=0.50) gives the most conservative size.
    • Specify the selection procedure: Record how random numbers, intervals, strata, quotas, or referral chains will be used so that selection can be implemented consistently.
    • Anticipate non-response: If 400 completed responses are required and an 80% response rate is expected:
TEXT
Adjusted sample = 400 / 0.80 = 500
  • Conduct a pilot study: Test the frame, contact procedure, questionnaire, response rate, and field workload on a small scale.
  • Implement and monitor selection: Prevent unauthorized substitutions and document refusals, inaccessible units, and deviations from the plan.

B. Practical controls and limitations

A sampling plan must remain methodologically sound during field implementation.

  • Cost–precision balance: Larger samples usually reduce sampling error but increase field, supervision, and processing costs.
  • Design effect: Cluster sampling often produces less precision than simple random sampling because members of one cluster may resemble each other.
  • Ethical control: Selection and recruitment should protect privacy, obtain informed consent, and avoid coercion.
  • Documentation: Eligibility rules, selection probabilities, substitutions, and non-response must be recorded to permit evaluation and replication.

III. Standards of Sample Quality

A. Characteristics of a good sample

A good sample represents the relevant population accurately enough to satisfy the study’s objectives without unnecessary expenditure.

  • Representativeness: Important population groups appear in appropriate proportions. A workforce sample should not exclude night-shift employees if conclusions concern all employees.
  • Adequate size: The sample contains enough observations to achieve the required precision, but size alone cannot correct biased selection.
  • Unbiased selection: No group receives systematic preference. In probability sampling, each unit has a known, non-zero selection probability.
  • Precision: Repeated samples would produce estimates close to one another. The standard error of a sample mean is approximately:
TEXT
SE(x̄) = s / √n
  • (SE(\bar{x})) = estimated standard error of the sample mean.
  • (s) = sample standard deviation.
  • (n) = sample size.
    • Accurate frame coverage: The frame includes eligible units once and excludes ineligible units.
    • Independence: Selection of one unit should not improperly determine another, except where the design deliberately uses clusters or linked units.
    • Low non-response: High or unequal non-response can make respondents systematically different from non-respondents.
    • Suitability to heterogeneity: Diverse populations may require stratification, while geographically dispersed populations may justify clustering.
    • Economy and feasibility: The design should be executable with available time, personnel, money, and access.
    • Measurability of reliability: Probability samples permit estimation of sampling error and confidence intervals.

B. Evaluation of sample quality

Sample quality must be assessed through both statistical evidence and field records.

  • Sampling error check: Examine standard errors, confidence intervals, and design effects rather than relying only on sample size.
  • Coverage check: Compare frame totals with reliable population totals by region, age, sex, or another relevant characteristic.
  • Non-response analysis: Compare respondents and non-respondents using available frame data.
  • Weighting: Apply sampling weights when selection probabilities differ:
TEXT
wᵢ = 1 / πᵢ
  • (w_i) = sampling weight for unit (i).
  • (\pi_i) = probability that unit (i) is included.

IV. Classification of Sampling Designs

A. Types of sampling design

Sampling designs are primarily classified according to whether selection probabilities are known.

  1. Probability sampling

    • Principle: Selection is governed by a random mechanism, and every unit has a known, non-zero probability of selection.
    • Inference: Sampling error can be estimated, supporting confidence intervals and population generalization.
    • Requirements: Usually needs an adequate frame, clear selection rules, and tighter field control.
  2. Non-probability sampling

    • Principle: Selection depends on accessibility, researcher judgment, participant choice, quotas, or referrals.
    • Inference: Selection probabilities are unknown, so conventional sampling-error estimates are not justified.
    • Requirements: Useful when no frame exists, the population is hidden, or exploratory understanding is more important than statistical generalization.
  • Explicit contrast: Probability designs prioritize measurable representativeness; non-probability designs prioritize feasibility, access, speed, or information-rich cases.
  • Mixed use: A multistage study may randomly select institutions but purposively select key informants within them; conclusions must reflect each stage’s limitations.

B. Basis for choosing a design

The appropriate design follows from the research purpose and operational setting.

  • Descriptive surveys: Probability methods are preferred when estimating prevalence, averages, or totals.
  • Exploratory studies: Purposive or snowball methods may reveal concepts, experiences, and hard-to-observe relationships.
  • Population structure: Stratification suits identifiable subgroups; clustering suits naturally grouped and geographically dispersed units.
  • Available resources: A statistically efficient design may be impractical if the frame is costly or field travel is extensive.

V. Probability-Based Selection

A. Random sampling techniques

Random sampling techniques use chance procedures to reduce selection bias and support design-based statistical inference.

  • Simple random sampling (SRS): Every possible sample of size (n) has an equal chance of selection. Units are chosen through random numbers or computerized generation; selection may be with or without replacement.
  • Systematic sampling: From (N) ordered units, calculate interval (k=N/n), choose a random start between 1 and (k), and select every (k)th unit. Hidden periodicity in the list can bias results.
  • Stratified random sampling: Divide the population into internally similar strata—such as departments or age groups—and randomly sample within each.
    • Proportionate allocation: Each stratum’s sample reflects its population share.
    • Disproportionate allocation: Small or analytically important strata are oversampled and later weighted.
  • Cluster sampling: Divide the population into natural groups, randomly select clusters, and observe all units or a sample within selected clusters. It reduces travel costs but may increase standard errors.
  • Multistage sampling: Selection occurs sequentially—for example, districts, then schools, then students. Each stage must have a defined random procedure.
  • Probability proportional to size (PPS): Larger clusters receive greater selection probability, commonly based on household or enrolment counts.
  • Multiphase sampling: Collect basic information from a large first-phase sample and detailed information from a smaller subsample.

B. Applications and limitations

Random techniques are strongest when population inference and measurable precision are required.

  • Applications: National surveys, opinion polls, public-health estimates, market studies, and institutional assessments.
  • Advantages: Reduced subjective selection, calculable sampling error, and defensible confidence intervals.
  • Limitations: Frame construction may be expensive; dispersed samples increase travel; non-response can still create bias; complex designs require weights and specialized analysis.

VI. Non-Probability-Based Selection

A. Non-random sampling techniques

Non-random sampling techniques select units without known probabilities and are mainly used for accessibility, exploration, or specialized insight.

  • Convenience sampling: Select readily available units, such as visitors present at one shopping centre. It is fast but highly vulnerable to coverage and self-selection bias.
  • Purposive or judgment sampling: Deliberately choose information-rich cases that meet defined criteria, such as experienced disaster-response officers.
    • Common forms include typical-case, extreme-case, critical-case, and maximum-variation sampling.
  • Quota sampling: Set category targets—for example, 60 women and 40 men—but allow non-random selection within each quota. It resembles stratification without random selection.
  • Snowball sampling: Initial participants refer other eligible participants. It is useful for hidden populations, but network structure may overrepresent closely connected groups.
  • Volunteer or self-selection sampling: Individuals choose to participate, as in an open online poll; strong opinions may be disproportionately represented.
  • Consecutive sampling: Include every eligible accessible case over a specified period until the desired size is reached, commonly in clinical settings.

B. Applications and limitations

Non-random methods are valuable for particular research purposes but require restrained interpretation.

  • Applications: Pilot studies, qualitative research, case studies, expert consultation, rare populations, and early instrument development.
  • Advantages: Low cost, rapid recruitment, flexibility, and access to specialized or hidden groups.
  • Limitations: Unknown selection probabilities, unmeasurable sampling error, weak population generalization, and dependence on researcher or participant choices.
  • Quality safeguards: State eligibility criteria, diversify recruitment sites, document referral chains, compare sample characteristics with known population information, and avoid presenting findings as statistically representative.