Unit 5: Sampling Design
I. Foundations of Sampling Design
Sampling design is the definite plan used to select a subset of units from a population so that information from the sample can support conclusions about that population. It specifies who or what may be selected, how selection occurs, and how many units are included.
- Population or universe: The complete set of elements relevant to the research problem—for example, all 4,000 students enrolled at a university.
- Target population: The population to which the researcher intends to generalize findings, such as all undergraduate students in a state.
- Accessible population: The portion that can practically be reached, such as students listed in participating institutions.
- Element: The basic unit about which information is collected—for example, a person, household, firm, transaction, or document.
- Sampling unit: The unit selected at a particular stage; it may be an individual, school, village, or geographic block.
- Sampling frame: The operational list or representation of sampling units, such as an electoral roll or employee register.
- Sample: The units actually selected from the population.
- Parameter and statistic:
- A parameter, such as population mean (\mu), describes the population.
- A statistic, such as sample mean (\bar{x}), is calculated from sample observations.
- Sampling error: Random difference between a sample estimate and the true population parameter.
- Non-sampling error: Error caused by non-response, inaccurate measurement, recording mistakes, frame defects, or data-processing failures.
- Central principle: A sample should provide sufficient information about the population at acceptable cost, time, and precision.
II. Planning the Sample — From Population to Selection
A. Sampling design process
The sampling design process converts a research objective into an operational procedure for selecting observable units.
- Define the research problem: Identify the information required and the population parameter to be estimated; for example, estimating the average monthly expenditure of urban households.
- Define the target population: Specify four boundaries:
- Element: Urban household.
- Geographical scope: A named city or region.
- Time period: The current financial year.
- Relevant characteristics: Occupied residential households.
- Choose the sampling unit: Decide whether selection will involve individuals, households, institutions, areas, or clusters. In a school survey, schools may be first-stage units and students second-stage units.
- Construct or identify the sampling frame: Evaluate the list for omissions, duplication, outdated entries, and inclusion of ineligible units. A defective frame creates coverage error even when selection is random.
- Select the sampling approach: Choose probability or non-probability sampling according to the study’s purpose, availability of a frame, desired generalizability, and resources.
- Determine sample size: For estimating a population proportion under simple random sampling, an initial size may be calculated as:
n₀ = z²p(1 − p) / e²- (n_0) = initial sample size for a large population.
- (z) = standard-normal value for the chosen confidence level.
- (p) = anticipated population proportion.
- (e) = permitted margin of error.
- At 95% confidence, (z \approx 1.96); if (p) is unknown, (p=0.50) gives the most conservative size.
- Specify the selection procedure: Record how random numbers, intervals, strata, quotas, or referral chains will be used so that selection can be implemented consistently.
- Anticipate non-response: If 400 completed responses are required and an 80% response rate is expected:
Adjusted sample = 400 / 0.80 = 500- Conduct a pilot study: Test the frame, contact procedure, questionnaire, response rate, and field workload on a small scale.
- Implement and monitor selection: Prevent unauthorized substitutions and document refusals, inaccessible units, and deviations from the plan.
B. Practical controls and limitations
A sampling plan must remain methodologically sound during field implementation.
- Cost–precision balance: Larger samples usually reduce sampling error but increase field, supervision, and processing costs.
- Design effect: Cluster sampling often produces less precision than simple random sampling because members of one cluster may resemble each other.
- Ethical control: Selection and recruitment should protect privacy, obtain informed consent, and avoid coercion.
- Documentation: Eligibility rules, selection probabilities, substitutions, and non-response must be recorded to permit evaluation and replication.
III. Standards of Sample Quality
A. Characteristics of a good sample
A good sample represents the relevant population accurately enough to satisfy the study’s objectives without unnecessary expenditure.
- Representativeness: Important population groups appear in appropriate proportions. A workforce sample should not exclude night-shift employees if conclusions concern all employees.
- Adequate size: The sample contains enough observations to achieve the required precision, but size alone cannot correct biased selection.
- Unbiased selection: No group receives systematic preference. In probability sampling, each unit has a known, non-zero selection probability.
- Precision: Repeated samples would produce estimates close to one another. The standard error of a sample mean is approximately:
SE(x̄) = s / √n- (SE(\bar{x})) = estimated standard error of the sample mean.
- (s) = sample standard deviation.
- (n) = sample size.
- Accurate frame coverage: The frame includes eligible units once and excludes ineligible units.
- Independence: Selection of one unit should not improperly determine another, except where the design deliberately uses clusters or linked units.
- Low non-response: High or unequal non-response can make respondents systematically different from non-respondents.
- Suitability to heterogeneity: Diverse populations may require stratification, while geographically dispersed populations may justify clustering.
- Economy and feasibility: The design should be executable with available time, personnel, money, and access.
- Measurability of reliability: Probability samples permit estimation of sampling error and confidence intervals.
B. Evaluation of sample quality
Sample quality must be assessed through both statistical evidence and field records.
- Sampling error check: Examine standard errors, confidence intervals, and design effects rather than relying only on sample size.
- Coverage check: Compare frame totals with reliable population totals by region, age, sex, or another relevant characteristic.
- Non-response analysis: Compare respondents and non-respondents using available frame data.
- Weighting: Apply sampling weights when selection probabilities differ:
wᵢ = 1 / πᵢ- (w_i) = sampling weight for unit (i).
- (\pi_i) = probability that unit (i) is included.
IV. Classification of Sampling Designs
A. Types of sampling design
Sampling designs are primarily classified according to whether selection probabilities are known.
-
Probability sampling
- Principle: Selection is governed by a random mechanism, and every unit has a known, non-zero probability of selection.
- Inference: Sampling error can be estimated, supporting confidence intervals and population generalization.
- Requirements: Usually needs an adequate frame, clear selection rules, and tighter field control.
-
Non-probability sampling
- Principle: Selection depends on accessibility, researcher judgment, participant choice, quotas, or referrals.
- Inference: Selection probabilities are unknown, so conventional sampling-error estimates are not justified.
- Requirements: Useful when no frame exists, the population is hidden, or exploratory understanding is more important than statistical generalization.
- Explicit contrast: Probability designs prioritize measurable representativeness; non-probability designs prioritize feasibility, access, speed, or information-rich cases.
- Mixed use: A multistage study may randomly select institutions but purposively select key informants within them; conclusions must reflect each stage’s limitations.
B. Basis for choosing a design
The appropriate design follows from the research purpose and operational setting.
- Descriptive surveys: Probability methods are preferred when estimating prevalence, averages, or totals.
- Exploratory studies: Purposive or snowball methods may reveal concepts, experiences, and hard-to-observe relationships.
- Population structure: Stratification suits identifiable subgroups; clustering suits naturally grouped and geographically dispersed units.
- Available resources: A statistically efficient design may be impractical if the frame is costly or field travel is extensive.
V. Probability-Based Selection
A. Random sampling techniques
Random sampling techniques use chance procedures to reduce selection bias and support design-based statistical inference.
- Simple random sampling (SRS): Every possible sample of size (n) has an equal chance of selection. Units are chosen through random numbers or computerized generation; selection may be with or without replacement.
- Systematic sampling: From (N) ordered units, calculate interval (k=N/n), choose a random start between 1 and (k), and select every (k)th unit. Hidden periodicity in the list can bias results.
- Stratified random sampling: Divide the population into internally similar strata—such as departments or age groups—and randomly sample within each.
- Proportionate allocation: Each stratum’s sample reflects its population share.
- Disproportionate allocation: Small or analytically important strata are oversampled and later weighted.
- Cluster sampling: Divide the population into natural groups, randomly select clusters, and observe all units or a sample within selected clusters. It reduces travel costs but may increase standard errors.
- Multistage sampling: Selection occurs sequentially—for example, districts, then schools, then students. Each stage must have a defined random procedure.
- Probability proportional to size (PPS): Larger clusters receive greater selection probability, commonly based on household or enrolment counts.
- Multiphase sampling: Collect basic information from a large first-phase sample and detailed information from a smaller subsample.
B. Applications and limitations
Random techniques are strongest when population inference and measurable precision are required.
- Applications: National surveys, opinion polls, public-health estimates, market studies, and institutional assessments.
- Advantages: Reduced subjective selection, calculable sampling error, and defensible confidence intervals.
- Limitations: Frame construction may be expensive; dispersed samples increase travel; non-response can still create bias; complex designs require weights and specialized analysis.
VI. Non-Probability-Based Selection
A. Non-random sampling techniques
Non-random sampling techniques select units without known probabilities and are mainly used for accessibility, exploration, or specialized insight.
- Convenience sampling: Select readily available units, such as visitors present at one shopping centre. It is fast but highly vulnerable to coverage and self-selection bias.
- Purposive or judgment sampling: Deliberately choose information-rich cases that meet defined criteria, such as experienced disaster-response officers.
- Common forms include typical-case, extreme-case, critical-case, and maximum-variation sampling.
- Quota sampling: Set category targets—for example, 60 women and 40 men—but allow non-random selection within each quota. It resembles stratification without random selection.
- Snowball sampling: Initial participants refer other eligible participants. It is useful for hidden populations, but network structure may overrepresent closely connected groups.
- Volunteer or self-selection sampling: Individuals choose to participate, as in an open online poll; strong opinions may be disproportionately represented.
- Consecutive sampling: Include every eligible accessible case over a specified period until the desired size is reached, commonly in clinical settings.
B. Applications and limitations
Non-random methods are valuable for particular research purposes but require restrained interpretation.
- Applications: Pilot studies, qualitative research, case studies, expert consultation, rare populations, and early instrument development.
- Advantages: Low cost, rapid recruitment, flexibility, and access to specialized or hidden groups.
- Limitations: Unknown selection probabilities, unmeasurable sampling error, weak population generalization, and dependence on researcher or participant choices.
- Quality safeguards: State eligibility criteria, diversify recruitment sites, document referral chains, compare sample characteristics with known population information, and avoid presenting findings as statistically representative.
Did this save you a night before the exam?
LPU Notes is free, and it stays free. Ads cover part of the server bill. The rest comes out of a student's own pocket: the domain, the storage, and keeping the site up through the weeks everyone needs it at once.
The payment button didn't load. An ad blocker or a filtered network is the usual reason. to try again.
Nothing here is ever locked, and nothing unlocks. Chip in only if it was worth it. What it pays for →