Sampling methods in research: a glossary
The words for deciding who or what a study will look at: the population and the list it is drawn from, the ways of choosing a sample, how many to include, and what goes wrong when a sample differs from its population. Every definition was checked against the sources below.
- A priori power analysis
- Accessible population
- Address-based sampling
- Area sampling
- Attrition
- Attrition allowance
- Attrition bias
- Berkson's bias
- Break-off
- Census
- Cluster sampling
- Complex sample design
- Compromise power analysis
- Confidence level
- Confirming and disconfirming case sampling
- Consecutive sampling
- Contact rate
- Convenience sampling
- Cooperation rate
- Coverage error
- Criterion sampling
- Critical case sampling
- Design effect
- Design weight
- Disproportionate stratified sampling
- Effective sample size
- Element
- Eligibility criteria
- Eligibility screening
- Equal probability of selection method
- Estimation domain
- Exclusion criteria
- Expert sampling
- Extreme case sampling
- Finite population correction
- Gatekeeper
- Generalisability
- Hidden population
- Homogeneous sampling
- Implicit stratification
- Inclusion criteria
- Information power
- Information-rich case
- Intensity sampling
- Intracluster correlation coefficient
- Item non-response
- Judgement sampling
- Loss to follow-up
- Margin of error
- Master sample
- Maximum variation sampling
- Measure of size
- Minimal statistically detectable effect
- Minimum detectable effect
- Multistage sampling
- Non-contact
- Non-probability sampling
- Non-proportional quota sampling
- Non-response bias
- Non-response weighting
- Non-sampling error
- Opportunistic sampling
- Opt-in panel
- Optimal allocation
- Overcoverage
- Oversampling
- Periodicity
- Planning for precision
- Population
- Post hoc power analysis
- Post-stratification
- Power analysis
- Primary sampling unit
- Probability of selection
- Probability proportional to size sampling
- Probability sampling
- Probability-based panel
- Propensity weighting
- Proportional quota sampling
- Proportionate stratified sampling
- Proximal similarity model
- Purposeful random sampling
- Purposive sampling
- Quota sampling
- Raking
- Random digit dialling
- Random selection
- Random walk sampling
- Recruitment
- Recruitment rate
- Refusal rate
- Representativeness
- Respondent-driven sampling
- Response rate
- Retention
- Sample
- Sample size
- Sample size calculation
- Sample size heuristic
- Sample size justification
- Sampling
- Sampling bias
- Sampling design
- Sampling error
- Sampling fraction
- Sampling frame
- Sampling interval
- Sampling unit
- Sampling with replacement
- Sampling without replacement
- Secondary sampling unit
- Seed
- Selection bias
- Self-selection bias
- Self-weighting sample
- Sensitivity power analysis
- Simple random sampling
- Smallest effect size of interest
- Snowball sampling
- Stratified purposeful sampling
- Stratified sampling
- Stratum
- Study population
- Substitution
- Survivorship bias
- Systematic sampling
- Target difference
- Target population
- Theoretical sampling
- Theory-based sampling
- Total survey error
- Two-phase sampling
- Typical case sampling
- Undercoverage
- Underpowered study
- Unit non-response
- Volunteer bias
- Volunteer sampling
- Weighting
- WEIRD sample
A priori power analysis
Also called: a-priori power analysis, prospective power analysis, power calculation
An a priori power analysis is a calculation, made before any data are collected, of how many participants a study needs to detect an effect of a stated size with a chosen power and significance level. It is the usual basis for the sample size of a trial or experiment that tests a hypothesis, and its answer depends heavily on the effect size assumed, which must itself be justified.
Accessible population
The accessible population is the part of the target population that a researcher can actually reach and sample from, such as people with a condition who attend two named hospitals rather than everyone who has it. Findings apply most directly to the accessible population, and any claim about the wider target population rests on how closely the two resemble each other.
Address-based sampling
Also called: ABS, address based sampling
Address-based sampling is the random selection of households from a list of postal addresses, after which people at the chosen addresses are invited to take part by post or online. Pew Research Center has recruited its probability-based panel this way since 2018, using the US Postal Service's master list, in place of the telephone sampling it used before.
Area sampling
Also called: area probability sampling, area frame sampling
Area sampling is probability sampling in which geographic areas, such as census enumeration areas or mapped blocks, are selected first and the dwellings or people within them are then listed and sampled. It is used when no suitable list of people or addresses exists, and the European Social Survey permits it, with field enumeration of dwellings, only when no better method is possible.
Attrition
Also called: participant attrition, dropout, drop-out
Attrition is the loss of participants from a study after they have joined, whether they withdraw, stop responding or can no longer be traced. It matters most in trials, cohort studies and panels that follow people over time, because it shrinks the sample and, if those who leave differ from those who stay, it biases the results.
Attrition allowance
Also called: allowance for attrition, dropout allowance, allowance for loss to follow-up
An attrition allowance is the extra number of participants recruited so that a study still has the sample size it needs after the expected dropouts, usually found by dividing the required number by one minus the expected dropout rate. Surveys make the same adjustment for expected non-response. The allowance restores numbers but cannot remove bias if those who leave differ from those who stay.
Attrition bias
Also called: loss to follow-up bias, dropout bias, drop-out bias
Attrition bias is a distortion in a study's results that arises when the participants who drop out differ systematically from those who remain, or when groups being compared lose participants at different rates. In a trial, for example, people whose symptoms worsen may leave one arm more often than the other, changing the groups in ways that have nothing to do with the treatment.
Berkson's bias
Also called: admission rate bias, Berkson bias, Berkson's paradox
Berkson's bias is a selection bias that distorts the apparent link between an exposure and a disease when participants are drawn from hospital patients, because having both conditions makes admission more likely. Berkson described it in 1946 using clinic patients with gallbladder disease and diabetes, and it chiefly threatens case-control studies that recruit their cases in hospital.
Break-off
Also called: breakoff, survey break-off
A break-off is a survey case in which a respondent starts the questionnaire but stops before giving enough answers for it to count as a partial interview. The American Association for Public Opinion Research treats break-offs as a kind of refusal when outcome rates are calculated, and asks researchers to define in advance where a partial interview ends and a break-off begins.
Census
Also called: complete enumeration, full enumeration, total population sampling
A census is the collection of data from every member of a population instead of from a sample. National population censuses are the best-known kind, and they often supply the sampling frames for later surveys. In a research project, studying every member of a small, finite population, such as all employees of one firm, is also a census and makes the sample size easy to justify.
Cluster sampling
Also called: cluster sample, cluster random sampling, one-stage cluster sampling, single-stage cluster sampling
Cluster sampling is probability sampling in which the population is divided into natural groups, such as schools or villages, some groups are chosen at random, and everyone in them is studied. It saves travel and listing costs when no list of individuals exists, but members of a cluster tend to be alike, so it is usually less precise than a simple random sample of equal size.
Complex sample design
Also called: complex survey design, complex sampling design
A complex sample design is a survey design that combines features such as stratification, several stages of selection, clustering and unequal selection probabilities instead of taking a simple random sample. Most large household surveys are designed this way, and their data must be analysed with methods that take the design into account, or the precision reported for the results will be misleading.
Compromise power analysis
A compromise power analysis is a power analysis that fixes the sample size and the effect size of interest and then sets the significance level and power so that the two kinds of error stand in a chosen ratio. Lakens suggests it for very large samples, where the usual five per cent level makes false positives needlessly likely next to false negatives, and for very small ones.
Confidence level
Also called: level of confidence, confidence coefficient
The confidence level, in sample size planning, is the proportion of repeated samples whose margin of error would be expected to capture the true population value, conventionally 95 per cent. Raising it to 99 per cent requires a larger sample to keep the same margin of error, so the level must be chosen before the sample size can be calculated.
Confirming and disconfirming case sampling
Also called: confirming and disconfirming cases
Confirming and disconfirming case sampling is a purposive strategy, used in the later stages of qualitative fieldwork, of seeking further cases that fit an emerging pattern and cases that run against it. Confirming cases add depth and credibility, while disconfirming ones offer rival explanations and mark the limits of what the findings can claim.
Consecutive sampling
Also called: consecutive recruitment, consecutive series, consecutive enrolment
Consecutive sampling is the recruitment of every eligible person who presents during a set period, such as all patients meeting the inclusion criteria at a clinic over six months, until the target number is reached. It is common in clinical research and is often regarded as the least biased non-probability method, because the researcher does not pick and choose among those available.
Contact rate
The contact rate is the proportion of sampled cases in which a survey reached someone, such as a responsible member of a selected household, whatever the outcome of that contact. Read beside the refusal rate, it shows whether non-response came mainly from failing to reach people or from people declining once they were reached.
Convenience sampling
Also called: convenience sample, availability sampling, accidental sampling, haphazard sampling, opportunity sampling
Convenience sampling is the selection of whoever is easiest to reach, such as students in the researcher's own class or shoppers passing a stall. It is quick and cheap and suits pilot and exploratory work, but nothing ensures the sample resembles the population. Psychology courses in Britain often call it opportunity sampling, which differs from opportunistic sampling in qualitative fieldwork.
Cooperation rate
The cooperation rate is the proportion of eligible people who were actually contacted and then completed the survey. It differs from the response rate by leaving out the people who were never reached, so it measures willingness to take part rather than overall success in obtaining interviews, and the two should not be confused when surveys are compared.
Coverage error
Also called: frame error, coverage bias
Coverage error is the error that arises when a sampling frame does not match the target population, so that some members have no chance of selection or some non-members and duplicates are included. It is hard to measure because it concerns people the frame misses, and unlike sampling error it does not shrink as the sample grows.
Criterion sampling
Also called: criterion-based sampling, criterion-i sampling
Criterion sampling is a purposive strategy of selecting all the cases that meet a predetermined criterion of importance, such as every trainer who delivered a new programme. Palinkas and colleagues found it the commonest purposive strategy in implementation research, and they also describe a variant that selects cases falling outside a criterion, such as services that missed a deadline.
Critical case sampling
Critical case sampling is a purposive strategy of choosing a case that allows a logical generalisation: if something holds there, it is likely to hold in other cases too. It depends on knowing which features make a case critical, and it is especially useful when resources allow only one site, programme or community to be studied.
Design effect
Also called: deff, design effects
The design effect is the ratio of the variance of an estimate from a complex sample to the variance that a simple random sample of the same size would give. Clustering and unequal selection probabilities usually push it above one. For clustering it is often estimated as one plus the intracluster correlation times one less than the average cluster size, and sample sizes are multiplied by it.
Design weight
Also called: base weight, design weights
A design weight is a number attached to each sampled case equal to the inverse of its probability of selection, so a person with a one-in-a-hundred chance counts for a hundred people. It corrects for deliberate differences in selection chances, such as oversampling, and later adjustments for non-response or to match population totals are applied on top of it.
Disproportionate stratified sampling
Also called: disproportional stratified sampling, disproportionate stratification, non-proportional stratified sampling, disproportionate allocation
Disproportionate stratified sampling is stratified sampling in which different strata are sampled at different rates, so some groups make up a larger share of the sample than of the population. It is used to obtain enough cases from small groups for separate analysis or to reduce variance, and weighting is then needed to restore each stratum's true share in overall estimates.
Effective sample size
Also called: neff
The effective sample size is the size of a simple random sample that would give the same precision as the complex sample actually achieved, found by dividing the number of completed interviews by the design effect. The European Social Survey, for instance, requires an effective sample size of at least 1,500 in most countries, so countries with clustered designs must interview more people.
Element
Also called: population element, sampling element
An element is a single member of a population, the unit about which information is collected, such as a person, a household or a patient record. In simple designs the elements are selected directly, while in cluster sampling groups of elements are selected first, so the sampling unit and the element are different things.
Eligibility criteria
Also called: selection criteria, eligibility requirements
Eligibility criteria are the rules that decide who may take part in a study, made up of the inclusion criteria and exclusion criteria together. Reporting guidelines such as STROBE ask authors to state them with the sources and methods of selecting participants, because they fix the population to which the findings can be applied.
Eligibility screening
Also called: screening for eligibility, eligibility assessment, pre-screening
Eligibility screening is the step of checking potential participants against a study's inclusion and exclusion criteria before they are enrolled, often with a short set of questions or a review of records. Reporting guidelines ask how many people were examined for eligibility, how many were confirmed eligible and how many were then included, so that losses at each step are visible.
Equal probability of selection method
Also called: EPSEM, epsem sample, equal probability sampling
An equal probability of selection method is any sampling design in which every element of the population has the same chance of being selected, even when selection happens in several stages. Such samples are self-weighting. Selecting clusters with probability proportional to size and then taking the same number of households in each is a common way to achieve one.
Estimation domain
Also called: survey domain, domain of estimation, reporting domain
An estimation domain is a part of the population, such as a region, for which a survey is designed to give separate estimates of acceptable precision. Planning for domains usually means a larger total sample or oversampling of the smaller ones, because each domain needs enough cases of its own.
Exclusion criteria
Also called: exclusion criterion
Exclusion criteria are characteristics that rule out people who otherwise meet the inclusion criteria, usually because they could put the person at extra risk, make follow-up unlikely or distort the results, such as a second serious illness. Restating an inclusion criterion in reverse, for example excluding women from a study that includes only men, is a common error.
Expert sampling
Expert sampling is a non-probability method in which the sample is made up of people with known or demonstrable expertise in the subject, such as a panel of specialists asked to judge a questionnaire. It suits questions that only experienced people can answer, but experts can be wrong, and a panel's views reflect whoever was invited to sit on it.
Extreme case sampling
Also called: deviant case sampling, extreme or deviant case sampling, outlier sampling
Extreme case sampling is a purposive strategy of selecting unusual cases, such as the best and worst performers, to learn from striking successes or failures. Readers may dismiss such cases as too unusual to be useful, which is why researchers sometimes choose intensity sampling instead, taking strong but not extreme examples.
Finite population correction
Also called: FPC, finite population correction factor, finite population correction fraction
The finite population correction is a factor, one minus the sampling fraction, that reduces the variance of an estimate when a sample drawn without replacement makes up a sizeable part of a finite population. When the sample is a tiny share of the population the factor is close to one and can be ignored, which is why precision depends on the sample's size far more than the population's.
Gatekeeper
Also called: gatekeepers, research gatekeeper
A gatekeeper is a person who can grant or refuse a researcher access to a setting or group, such as a head teacher, a ward manager or a community leader. Gatekeepers often decide which potential participants hear about a study at all, so their decisions can shape who is invited and therefore who ends up in the sample.
Generalisability
Also called: generalizability, generalisation, generalization, generalisability of findings
Generalisability is the extent to which a study's results hold for people, places and times beyond the sample studied. In quantitative research it rests mainly on how the sample was drawn and who took part, which is why reporting guidelines such as STROBE ask authors to discuss it. Where no random sample was drawn, researchers argue instead from how similar other settings are to the one studied.
Hidden population
Also called: hard-to-reach population, hard to reach population, hard-to-reach groups
A hidden population is a group for which no list exists and whose members may not wish to be identified, often because membership is stigmatised or illegal, such as people who inject drugs or children living on the street. Ordinary probability sampling is impractical for such groups, so researchers use referral-based methods such as snowball or respondent-driven sampling.
Homogeneous sampling
Also called: homogenous sampling, homogeneous purposive sampling
Homogeneous sampling is a purposive strategy of choosing participants who share a key characteristic or experience, so that one subgroup can be described in depth. It reduces variation and simplifies analysis, and it is often used to form focus groups, such as a group made up only of managers of one kind of service.
Implicit stratification
Implicit stratification is the practice of sorting a sampling frame by characteristics such as region or urban and rural area and then taking a systematic sample through the sorted list. Because the selections are spread evenly along the list, the sample mirrors the population on those characteristics without separate strata being drawn.
Inclusion criteria
Also called: inclusion criterion
Inclusion criteria are the characteristics a person must have to take part in a study, such as an age range, a diagnosis or residence in a given area, chosen because they define the population the research question is about. They should be set with an eye on external validity, since every criterion narrows the group to which the results apply.
Information power
Information power is a way of judging qualitative sample size, proposed by Malterud, Siersma and Guassora in 2016, holding that the more relevant information a sample holds, the fewer participants are needed. It depends on the study's aim, how specific the sample is, the use of established theory, the quality of the interview dialogue and the analysis strategy, and it was offered as an alternative to saturation.
Information-rich case
Also called: information-rich cases, information rich case
An information-rich case is a person, group or setting that can tell a researcher a great deal about the issue under study, usually because of direct knowledge or experience of it. Purposive sampling in qualitative research is built around finding such cases, and a participant's willingness and ability to reflect aloud on the experience also count.
Intensity sampling
Intensity sampling is a purposive strategy of choosing cases that show the phenomenon of interest strongly but not in its most extreme form. It shares the purpose of extreme case sampling while avoiding cases that readers might dismiss as too unusual, and it usually requires some exploratory work first to learn how the phenomenon varies.
Intracluster correlation coefficient
Also called: intra-cluster correlation, intracluster correlation, intra-cluster correlation coefficient, ICC
The intracluster correlation coefficient is the share of the total variation in an outcome that lies between clusters rather than within them, and so measures how alike members of the same cluster are. Even small values matter: combined with large clusters they raise the design effect and the sample needed, which is why cluster studies must estimate it in advance.
Item non-response
Also called: item nonresponse, item non response
Item non-response is the failure of a respondent who takes part in a survey to answer particular questions, for example by skipping a question about income. It leaves gaps in otherwise usable records and differs from unit non-response, in which no answers at all are obtained from a sampled person or household.
Judgement sampling
Also called: judgment sampling, judgemental sampling, judgmental sampling
Judgement sampling is non-probability sampling in which the researcher or a group of experts chooses the units they believe will serve the study, often those they consider typical. Many textbooks treat it as another name for purposive sampling. Survey statisticians warn that what counts as typical is subjective and that nobody can calculate the chance each unit had of being chosen.
Loss to follow-up
Also called: lost to follow-up
Loss to follow-up is the situation in which participants in a cohort study or trial can no longer be contacted or assessed before the study ends, so their later outcomes are unknown. It is one part of attrition, beside withdrawal, and reporting guidelines ask authors to give the numbers completing follow-up and the reasons others did not.
Margin of error
Also called: MOE, sampling margin of error, margin of sampling error
The margin of error is the amount by which a survey estimate could differ from the true population value through sampling error alone, at a stated confidence level, usually 95 per cent. A simple random sample of a little over 1,000 gives about plus or minus three percentage points, subgroups have wider margins, and it says nothing about coverage, non-response or measurement errors.
Master sample
A master sample is a large probability sample of areas or dwellings selected once and then used as the source for many surveys or survey rounds, often over about ten years. It spares a national statistics office from drawing a fresh frame and first-stage sample for every survey it runs.
Maximum variation sampling
Also called: heterogeneous sampling, heterogeneity sampling, sampling for diversity, maximum variation purposive sampling
Maximum variation sampling is a purposive strategy of choosing cases that differ as widely as possible on dimensions that matter to the study, such as urban and rural services in different regions. It aims both to document the variety of experience and to find patterns that hold across it, which carry weight precisely because they emerged from such different cases.
Measure of size
Also called: MOS, size measure
A measure of size is a count or estimate of how big each unit in a sampling frame is, such as the number of households in a village, used to give larger units a greater chance of selection. In household surveys it usually comes from the most recent census, and it drives probability proportional to size sampling.
Minimal statistically detectable effect
Also called: critical effect size
The minimal statistically detectable effect is the smallest observed effect size that would give a statistically significant result, given the test, the significance level and the sample size. Lakens recommends computing it when no power analysis was done: with 15 participants per group in a t test, only observed effects larger than d = 0.75 would be significant.
Minimum detectable effect
Also called: MDE, minimum detectable effect size, MDES
The minimum detectable effect is the smallest true effect that a study, as designed, has a stated chance, usually 80 per cent, of detecting as statistically significant. A larger sample, lower variance in the outcome, a less strict significance level and lower intracluster correlation all make it smaller, and evaluators use it to judge whether a design is worth running.
Multistage sampling
Also called: multi-stage sampling, multistage sample, multistage cluster sampling, multi-stage cluster sampling
Multistage sampling is probability sampling carried out in successive stages, for example selecting districts, then villages within the chosen districts, then households within the chosen villages. Unlike one-stage cluster sampling, only some units within each chosen cluster are taken. It makes national surveys affordable at some cost in precision, and some texts use the name cluster sampling for it.
Non-contact
Also called: noncontact, non-contacts
A non-contact is a sampled case that a survey could not reach at all, for example because nobody was ever at home when an interviewer called. Non-contacts are one of the main kinds of unit non-response, beside refusals and other reasons, and repeated call-backs at different times are the usual remedy.
Non-probability sampling
Also called: nonprobability sampling, non-random sampling, non-probability sample, nonprobability sample
Non-probability sampling is any method of choosing a sample in which each member's chance of selection is unknown, because people are chosen by judgement, availability, referral or self-selection rather than by a random procedure. It is standard in qualitative research, where depth matters more than representativeness, but it does not support the usual calculations of sampling error.
Non-proportional quota sampling
Also called: nonproportional quota sampling, disproportionate quota sampling
Non-proportional quota sampling is quota sampling in which the researcher sets a minimum number of participants for each category instead of matching the categories' shares in the population. It makes sure small groups are heard, much as disproportionate stratification does for probability samples, but without random selection.
Non-response bias
Also called: nonresponse bias, non response bias
Non-response bias is the bias that arises when the people who do not respond to a survey or decline to take part differ from those who do, in ways that matter to the results. A low response rate raises the risk but does not prove bias, and a high one does not rule it out, so researchers compare respondents and non-respondents where they can.
Non-response weighting
Also called: nonresponse weighting, non-response adjustment, nonresponse adjustment
Non-response weighting is the adjustment of survey weights so that respondents also stand in for similar people who were sampled but did not respond. Typically the sample is divided into groups and each respondent's weight is multiplied by the inverse of the response rate in their group. It removes bias only to the extent that non-respondents resemble respondents within those groups.
Non-sampling error
Also called: nonsampling error, non sampling error
Non-sampling error is any error in survey results that does not come from studying a sample instead of the whole population, including coverage error, non-response, measurement error and mistakes in processing. Unlike sampling error, it is not reduced by taking a bigger sample, and a census, which has no sampling error, can still suffer from it.
Opportunistic sampling
Also called: emergent sampling, opportunistic or emergent sampling
Opportunistic sampling is a qualitative strategy of adding participants or settings as chances arise during fieldwork, such as an unexpected event or a newly met informant, instead of fixing the whole sample at the start. It is used mainly in ethnography, when the population to be sampled cannot be identified in advance, and should not be confused with opportunity sampling.
Opt-in panel
Also called: opt-in online panel, opt-in sample, nonprobability online panel, non-probability online panel, online opt-in sample
An opt-in panel is a pool of people who have signed up to take surveys, usually online in response to advertisements or pop-ups, so nobody knows their chance of selection. Such panels are cheap and fast, but Pew Research Center found the accuracy of samples from them varied widely between vendors, with especially large errors for some minority groups.
Optimal allocation
Also called: optimum allocation, Neyman allocation
Optimal allocation is a rule for dividing a stratified sample among strata so that the estimate has the smallest variance for its cost, giving more cases to strata that are larger, more variable or cheaper to sample. When costs are the same in every stratum it is usually called Neyman allocation, a name some texts use for optimal allocation in general.
Overcoverage
Also called: over-coverage
Overcoverage is the inclusion in a sampling frame of units that do not belong to the target population, or of the same unit more than once. Examples include people on a population register who now live abroad and households listed twice, and if it goes undetected it inflates estimates of totals.
Oversampling
Also called: over-sampling, oversample
Oversampling is the deliberate selection of more members of a group than its share of the population would give, so that the group can be analysed on its own with adequate precision. Pew Research Center, for example, has oversampled Hispanic voters in election surveys. Weighting then brings the group back to its true share whenever figures for the whole population are produced.
Periodicity
Periodicity is a repeating pattern in a sampling frame that can bias a systematic sample when the pattern's cycle lines up with the sampling interval. If a list of flats puts a corner flat at every tenth place and the interval is ten, the sample could contain only corner flats. Checking or shuffling the order of the list beforehand guards against it.
Planning for precision
Also called: planning for accuracy, precision-based sample size, sample size for precision
Planning for precision is choosing a sample size so that an estimate will have a confidence interval no wider than a set amount, instead of so that a hypothesis test will have enough power. It suits studies that aim to measure something, such as a prevalence or a mean, though Lakens notes there is little guidance on how narrow the interval should be.
Population
Also called: statistical population, universe, population of interest
A population, in research, is the whole group of people, organisations, events or objects that a study wants to draw conclusions about, such as all undergraduates in a country. It is defined by the research question and need not be made up of people, and because it is rarely possible to study all of it, researchers usually study a sample.
Post hoc power analysis
Also called: post-hoc power, post hoc power, observed power, retrospective power analysis
A post hoc power analysis is a calculation of a study's power made after the event, treating the effect size observed in the data as if it were the true effect. Methodologists such as Lakens advise against it because it adds nothing to the p value already reported, and they recommend a sensitivity power analysis across effect sizes of interest for interpreting a non-significant result.
Post-stratification
Also called: poststratification, post-stratification weighting
Post-stratification is the grouping of respondents into strata after the sample has been collected, followed by weighting so that each group's share matches its known share of the population. It is used when the stratifying characteristic, such as sex or age, is not on the sampling frame, and it corrects imbalances only on the characteristics used.
Power analysis
Also called: statistical power analysis
Power analysis is the set of calculations linking sample size, effect size, significance level and statistical power, so that any one can be found from the other three. Done before a study as an a priori analysis it sets the sample size, a sensitivity analysis asks what a fixed sample can detect, and power calculated afterwards from the observed effect is discouraged.
Primary sampling unit
Also called: PSU, first-stage unit, first-stage sampling unit, primary unit
A primary sampling unit is a unit selected at the first stage of a multistage sample, such as a census enumeration area, a district or a school, within which later stages select smaller units. Survey data files usually identify each case's primary sampling unit so that analysts can allow for clustering when they calculate standard errors.
Probability of selection
Also called: selection probability, inclusion probability, probability of inclusion, chance of selection
The probability of selection is the chance that a particular member of a population ends up in the sample under a given design. Probability sampling requires it to be known and above zero for everyone, though not necessarily equal, and its inverse becomes the design weight. Statisticians distinguish the chance per draw, used when sampling with replacement, from the overall chance of inclusion.
Probability proportional to size sampling
Also called: PPS sampling, PPS, probability proportionate to size, probability-proportional-to-size sampling, PPS selection
Probability proportional to size sampling is a method of selecting clusters in which each cluster's chance of selection is proportional to a measure of its size, such as its number of households. Combined with taking the same number of households from each selected cluster, it gives every household roughly the same chance of selection while fixing the sample size and interviewer workloads in advance.
Probability sampling
Also called: probability sample, probability-based sampling, random probability sampling
Probability sampling is any method of choosing a sample in which every member of the population has a known, non-zero chance of selection, decided by a random process rather than anyone's judgement. It allows sampling error to be estimated and supports inference to the population. Some introductory texts say every member must have an equal chance, but that holds only for certain designs.
Probability-based panel
Also called: probability panel, probability-based online panel, probability-based survey panel
A probability-based panel is a group of people recruited by random sampling, for example from lists of addresses, and then surveyed repeatedly, usually online. Only people who were selected can join, which is what separates it from an opt-in panel, and members who stop taking part are replaced through fresh random recruitment.
Propensity weighting
Also called: propensity score weighting
Propensity weighting is a weighting method that models each respondent's probability of being in the sample, often by comparing an opt-in sample with a reference probability sample, and gives larger weights to the kinds of people who were less likely to be included. Its drawback is that very uneven weights can make estimates much less precise.
Proportional quota sampling
Also called: proportionate quota sampling
Proportional quota sampling is quota sampling in which the number recruited in each category matches that category's share of the population, for example 40 per cent women if women make up 40 per cent. Once a category's quota is full, further volunteers from it are turned away, but the people within each quota are still chosen by convenience or judgement.
Proportionate stratified sampling
Also called: proportional stratified sampling, proportionate stratification, proportional allocation, proportionate stratified random sampling
Proportionate stratified sampling is stratified sampling in which every stratum is sampled at the same rate, so each makes up the same share of the sample as of the population. It keeps the sample self-weighting and can improve precision when the strata differ on the outcome, but it may leave small strata with too few cases for separate analysis.
Proximal similarity model
Also called: proximal similarity
The proximal similarity model is a way of reasoning about generalisation that asks how closely other people, places and times resemble those of the study, instead of relying on a random sample from a defined population. Trochim describes findings as carrying most confidently to the settings that are most similar, and never with certainty.
Purposeful random sampling
Also called: purposive random sampling
Purposeful random sampling is the random selection of a small number of cases from a purposively defined group, such as interviewing a randomly chosen handful of the providers in a programme. It is used in qualitative research to add credibility and to reduce suspicion that cases were hand-picked, though the sample is too small to be representative in the statistical sense.
Purposive sampling
Also called: purposeful sampling, purposive sample
Purposive sampling is the deliberate selection of participants or cases because they have characteristics or experience relevant to the research question, rather than at random. It is the main approach in qualitative research, where the aim is information-rich cases, and it covers many named strategies, among them maximum variation, homogeneous, typical, extreme, critical and criterion sampling.
Quota sampling
Also called: quota sample
Quota sampling is non-probability sampling in which the researcher decides how many people to recruit in each category, such as age group and sex, and interviewers fill those quotas with whoever they can find. It matches the population on the quota characteristics, but selection within quotas is subjective, and a quota poll in Washington State overstated Dewey's vote in the 1948 US presidential election.
Raking
Also called: iterative proportional fitting, raking weights
Raking is a weighting method that adjusts survey weights one variable at a time, over and over, until the weighted sample matches known population figures for each variable, such as age, sex and region. It is the most common weighting method in public polling because it needs only each variable's population distribution, not the full cross-classification of all of them.
Random digit dialling
Also called: random digit dialing, RDD, random-digit dialling, random-digit-dial sampling
Random digit dialling is a telephone sampling method in which phone numbers are generated at random and called, so the sampling frame is a procedure rather than a printed list. Pew Research Center recruited its panel this way from 2014 to 2017, calling both landlines and mobile phones, before moving to address-based sampling.
Random selection
Also called: random sampling, random sample
Random selection is the use of a chance process, such as a random number generator, to decide which members of a population enter a sample. It is not the same as random assignment, which uses chance to decide which condition each participant in a study receives: random selection supports generalising to a population, while random assignment controls other influences on the comparison.
Random walk sampling
Also called: random route sampling, random route technique, random walk
Random walk sampling is a field method in which interviewers start at a chosen point and follow set rules of travel, such as calling at every fifth house, to pick households instead of selecting them from a list. It saves listing costs, but because interviewers control selection the probabilities cannot be calculated, and the European Social Survey does not permit it.
Recruitment
Also called: participant recruitment, recruiting participants
Recruitment is the process of informing eligible people about a study and inviting them to take part, through letters, posters, emails, clinicians or other routes. It decides who actually enters the sample, and poor recruitment is a common reason trials end up underpowered or fail, which is why methods of recruiting are now tested in trials of their own.
Recruitment rate
The recruitment rate is the proportion of eligible people approached who agree to join a study. Cochrane reviews of recruitment strategies use the proportion of eligible participants recruited as their main outcome, and a low rate signals both a risk of an underpowered study and a possibly unrepresentative sample.
Refusal rate
The refusal rate is the proportion of all potentially eligible cases in which the selected person or household declines to be interviewed or breaks off the interview. The American Association for Public Opinion Research gives three versions, which differ in how they treat cases whose eligibility is unknown.
Representativeness
Also called: representative sample, sample representativeness
Representativeness is the degree to which a sample resembles its population in all the ways that matter for the research, so that results from the sample can stand for the population. It depends on the sample actually obtained, not only on the method used to draw it, since non-response and attrition can make even a random sample unrepresentative.
Respondent-driven sampling
Also called: RDS, respondent driven sampling
Respondent-driven sampling is a form of chain-referral sampling for hidden populations, developed by Douglas Heckathorn in 1997, in which participants recruit their peers with a limited number of coded coupons and are paid both for taking part and for each successful recruit. Weights based on each person's network size aim to correct recruitment bias, though how well they do so is disputed.
Response rate
Also called: survey response rate, return rate
The response rate is the number of completed interviews or questionnaires divided by the number of eligible people or households in the sample. The American Association for Public Opinion Research defines six versions, depending on how partial interviews and cases of unknown eligibility are counted, so a study should say which it used. A low rate does not in itself prove non-response bias.
Retention
Also called: participant retention, retention of participants
Retention is the keeping of participants in a study until its final data collection, the opposite of attrition. Poor retention leaves outcome data missing, which can bias results and reduce power, so studies use a range of retention strategies, though a Cochrane review found that few of them had been formally evaluated.
Sample
Also called: research sample, study sample
A sample is the subset of a population from which a study collects its data, such as 500 students chosen from all the students at a university. Its value lies in what it can say about the population, so how it was chosen, and who in the end took part, matter as much as its size.
Sample size
Also called: n, study size, number of participants
Sample size is the number of units, usually participants, from which a study collects data, written n. It should be decided before data collection and justified in the write-up, since too small a sample may miss real effects or give imprecise estimates, and too large a one wastes resources and exposes more people than necessary to the burdens of research.
Sample size calculation
Also called: sample size determination, sample size estimation, sample size computation
A sample size calculation is a formal computation, made before a study begins, of how many participants are needed to meet its aim, whether that is enough power to detect a target difference or enough precision to estimate a value. Its inputs include the significance level, the power or confidence level, the variability of the outcome, and allowances for dropout and clustering.
Sample size heuristic
Also called: sample size rule of thumb, rules of thumb for sample size
A sample size heuristic is a rule of thumb, such as a fixed number of participants per group or per predictor, used instead of a calculation for the study in hand. Lakens counts heuristics among the six common kinds of sample size justification but warns that many popular rules rest on weak foundations and some were misread from the papers that proposed them.
Sample size justification
A sample size justification is the explanation of why a study collected the amount of data it did, given what it set out to learn. Lakens identifies six kinds: measuring almost the whole population, resource limits, an a priori power analysis, planning for precision, a heuristic, and an honest statement that there was no justification.
Sampling
Also called: sampling method, sampling technique, sampling procedure, sampling process
Sampling is the process of selecting the people, cases, places or other units a study will examine from the larger group it is interested in. Methods divide into probability sampling, which uses random selection and supports statistical inference, and non-probability sampling, which relies on judgement, availability or referral and is usual in qualitative work.
Sampling bias
Also called: sample bias, biased sample
Sampling bias is a systematic difference between a sample and its population caused by how the sample was chosen, when some members were less likely to be selected than others. Unlike sampling error, it does not shrink as the sample grows. Many writers treat it as one kind of selection bias, the kind that arises when the sample is drawn.
Sampling design
Also called: sample design, sampling plan, sampling scheme, sampling strategy
A sampling design is the full plan for selecting a sample, including the population and frame, the method of selection, any stratification or stages, and the sample size. The design determines how the data must be analysed, since estimates and their precision depend on how units were chosen, and it should be reported so that readers can judge the sample.
Sampling error
Also called: random sampling error
Sampling error is the difference between a result from a sample and the true value in the population that arises because only a sample, not the whole population, was studied. It is random, can be estimated for probability samples and gets smaller as the sample grows. The margin of error reported for polls describes this error only.
Sampling fraction
Also called: sampling rate
The sampling fraction is the proportion of the population included in the sample, the sample size divided by the population size, often written f or n/N. It appears in the finite population correction, and in stratified designs it is kept equal across strata under proportionate allocation and varied under disproportionate allocation.
Sampling frame
Also called: sample frame, sampling frames
A sampling frame is the list, or the procedure, from which a sample is actually drawn, such as a patient register, a list of addresses or a method of generating telephone numbers. The more completely it covers the target population, the better the sample can be, and gaps or extra entries in it cause coverage error.
Sampling interval
Also called: selection interval
The sampling interval is the fixed gap between selections in systematic sampling, found by dividing the population size by the sample size. With 1,000 names on a list and a sample of 100, the interval is ten, so after a random start between one and ten every tenth name is taken.
Sampling unit
Also called: unit of selection, selection unit
A sampling unit is the unit actually selected at a given stage of sampling, which may be a single element, such as a person, or a group of elements, such as a household, a school or a village. Multistage designs have sampling units at each stage, from the primary sampling units down to the final ones.
Sampling with replacement
Sampling with replacement is selection in which each chosen unit is returned to the population before the next draw, so the same unit can be selected more than once. It is rare in practice, but its formulas are simpler, and when the sample is small relative to the population the results barely differ from sampling without replacement.
Sampling without replacement
Sampling without replacement is selection in which a unit, once chosen, cannot be chosen again, which is how almost all real samples of people are drawn. When the sample is a sizeable share of a small population, variance calculations should include the finite population correction to reflect it.
Secondary sampling unit
Also called: SSU, second-stage unit, secondary unit
A secondary sampling unit is a unit selected at the second stage of a multistage sample from within a primary sampling unit already chosen, such as households within a selected village. In one-stage cluster sampling all the secondary units in each chosen cluster are included, while two-stage sampling takes only some of them.
Seed
Also called: seeds, seed participant, seed respondent
A seed is one of the first participants in a chain-referral study, chosen by the researchers to start recruitment through their own networks. In respondent-driven sampling a small number, often between three and fifteen, are picked to be varied and well connected, and they do not have to be chosen at random.
Selection bias
Selection bias is a systematic error that arises when the people included in a study, or in the groups being compared, differ from the population of interest because of how they were selected or came to take part. It can limit generalisability or distort comparisons, and in trials the term also covers differences between arms that proper randomisation and allocation concealment are meant to prevent.
Self-selection bias
Also called: self selection bias
Self-selection bias is the bias that arises when people decide for themselves whether to be in a sample, as in call-in polls or online surveys open to anyone, so those who choose to take part differ from those who do not. Weighting can reduce it only on characteristics that are measured, which is why self-selected results are often unreliable.
Self-weighting sample
Also called: self-weighting design, self-weighting
A self-weighting sample is one in which every case carries the same weight because every element had the same overall chance of selection. Its estimates of proportions and averages can be calculated without design weights, which is convenient, though later adjustments for non-response may still give cases different weights.
Sensitivity power analysis
A sensitivity power analysis is a power analysis that fixes the sample size, significance level and desired power, and calculates the smallest effect size the design could detect with that power. It is used when the sample size is already set, for example by existing data or a limited budget, to show whether the study can detect effects that are plausible and of interest.
Simple random sampling
Also called: SRS, simple random sample, simple random selection
Simple random sampling is selection in which every possible sample of the chosen size is equally likely, so each member of the population has the same chance of inclusion, usually by drawing random numbers from a numbered list. It needs a complete sampling frame, is the benchmark against which other designs are compared, and is rarely used on its own in large household surveys.
Smallest effect size of interest
Also called: SESOI, smallest effect of interest
The smallest effect size of interest is the smallest effect that a researcher would consider theoretically or practically worth detecting, stated before the study. Lakens calls it the strongest basis for a power analysis, because it ties the sample size to what matters rather than to a guess at the true effect, which is unknown.
Snowball sampling
Also called: snowballing, chain-referral sampling, chain referral sampling, referral sampling, snowball sample
Snowball sampling is a non-probability method in which the first participants refer others they know who also meet the criteria, and those people refer more in turn. It suits hard-to-reach or stigmatised groups with no list, such as homeless young people, but the sample reflects the social networks of the first recruits and cannot be assumed representative.
Stratified purposeful sampling
Also called: stratified purposive sampling
Stratified purposeful sampling is a qualitative strategy that divides potential cases into strata, such as above-average, average and below-average performers, and selects a few purposively from each. It aims to capture the main variations rather than a common core, taking in more than typical case sampling but less than a full maximum variation sample.
Stratified sampling
Also called: stratified random sampling, stratified sample, stratified random sample
Stratified sampling is probability sampling in which the population is first divided into non-overlapping groups called strata, such as regions or year groups, and a separate random sample is drawn from each. It guarantees that every stratum is represented, allows estimates for each, and improves precision when members of a stratum resemble one another on the outcome.
Stratum
Also called: strata
A stratum is one of the non-overlapping groups into which a population is divided before a stratified sample is drawn, such as one region, one age band or one type of school. Good strata are internally alike and differ from one another on what the study measures, and every member of the population must belong to exactly one.
Study population
The study population is the group from which a study actually draws or recruits its participants, though authors use the term in more than one sense. Some methods texts treat it as another name for the accessible population, while clinical papers often use it for the people actually recruited, as against the target population. A report should make its meaning plain.
Substitution
Also called: substitution of non-respondents, replacement of non-respondents
Substitution is the replacement of a sampled person or household that cannot be contacted or refuses with another, more available one. It keeps interview numbers up but breaks probability sampling, because the substitute's chance of selection is unknown, and it carries the same bias as the non-response it hides, so the European Social Survey forbids it.
Survivorship bias
Also called: survivor bias
Survivorship bias is a form of selection bias that arises when only the people, firms or cases that came through some selection process are studied, while those that dropped out or failed are overlooked. In a longitudinal survey, for example, analysing only people who answered every wave can give too hopeful a picture if those who left were faring worse.
Systematic sampling
Also called: systematic random sampling, systematic sample, interval sampling
Systematic sampling is the selection of every kth unit from a list after a random start within the first k units, where k is the sampling interval. It is easy to carry out, even in the field without a full list, and often at least as precise as simple random sampling, but a pattern in the list that matches the interval can bias it.
Target difference
Also called: target effect size
The target difference is the difference between treatments that a trial's sample size calculation is built to detect, chosen because it is realistic, important, or both. The DELTA2 guidance says it should matter to at least one key group, such as patients or clinicians, and warns that pilot trials are usually too small to show what difference is realistic.
Target population
Also called: theoretical population, coverage universe
The target population is the whole group to which a study intends its findings to apply, defined by the research question, such as all adults with type 2 diabetes in a country. It is usually larger than the accessible population the researcher can reach, and the gap between the two limits how far the findings can be carried.
Theoretical sampling
Theoretical sampling is the grounded theory practice of deciding what data to collect next, and from whom, on the basis of the concepts emerging from analysis of the data already gathered. It usually follows an initial purposive sample and continues, alongside constant comparison, until theoretical saturation, with each choice aimed at filling gaps or testing hunches in the developing theory.
Theory-based sampling
Also called: operational construct sampling
Theory-based sampling is a purposive strategy of choosing cases because they show a theoretical construct of interest, so that the construct and its variations can be examined. It differs from grounded theory's theoretical sampling in that the construct is usually set in advance, although some descriptions blur the two by letting emerging concepts guide further selection.
Total survey error
Also called: TSE
Total survey error is a framework that considers every source of error in a survey estimate together, including sampling error, coverage error, non-response error and measurement error, instead of the margin of sampling error alone. It reminds readers that a small margin of error does not guarantee an accurate result.
Two-phase sampling
Also called: double sampling, two phase sampling
Two-phase sampling is a design in which a first, larger sample is drawn and a second sample is then taken from within it, often using information gathered in the first phase to screen or stratify. It is used, for example, to find members of a rare group cheaply before interviewing them at length.
Typical case sampling
Also called: modal instance sampling
Typical case sampling is a purposive strategy of choosing cases that are normal or average for the setting, to show what is typical to readers who do not know it. It describes rather than generalises, and deciding what is typical can be harder than it looks when people differ on several characteristics at once.
Undercoverage
Also called: under-coverage, non-coverage, noncoverage
Undercoverage is the omission from a sampling frame of members of the target population, who then have no chance of being selected, such as recent immigrants missing from a population register. It is a frequent problem in large household surveys and biases results whenever the missing people differ from those listed.
Underpowered study
Also called: underpowered, underpowered trial
An underpowered study is one whose sample is too small to give a good chance of detecting an effect of the size that matters, so a non-significant result tells readers little. Poor recruitment is a common cause in trials, and when underpowered studies do find significant effects, the estimates tend to be exaggerated.
Unit non-response
Also called: unit nonresponse, unit non response
Unit non-response is the failure to obtain any survey data at all from a sampled person or household, through refusal, non-contact or other reasons. It lowers the response rate and can cause non-response bias, and surveys compensate for it with call-backs during fieldwork and weighting afterwards.
Volunteer bias
Volunteer bias is the distortion that arises because people who volunteer for research differ from the population of interest, not only in demographics but in motivation, attitudes and health. The risk grows as more people decline to volunteer, and it can persist through a study because those who stay to the end also differ from those who leave.
Volunteer sampling
Also called: self-selected sampling, self-selection sampling, volunteer sample, self-selected sample, self-referral
Volunteer sampling is a non-probability method in which participants put themselves forward, for example by answering an advertisement, a poster or an open online survey. It can recruit quickly and reaches people motivated to take part, but volunteers tend to differ from non-volunteers, so the results may not hold for the wider population.
Weighting
Also called: survey weighting, sample weighting, survey weights, sampling weights
Weighting is the adjustment of survey data so that each case counts for more or less than one, making the weighted sample match the population. It combines design weights for unequal chances of selection with adjustments for non-response and to known population totals, and while it can remove bias on the characteristics used, it also widens margins of error.
WEIRD sample
Also called: WEIRD, WEIRD samples, WEIRD populations
A WEIRD sample is one drawn from Western, educated, industrialised, rich and democratic societies, most often university students in the United States. Critics point out that much behavioural research rests on such samples yet is reported as if it described people in general, which risks overgeneralising from a narrow slice of humanity.
Where these definitions were checked
Barbara Illowsky and Susan Dean, OpenStax, Introductory Statistics 2e, 1.2: Data, Sampling, and Variation in Data and Sampling
Matthew DeCarlo (open textbook), Scientific Inquiry in Social Work, 10.1: Basic concepts of sampling
Matthew DeCarlo (open textbook), Scientific Inquiry in Social Work, 10.2: Sampling in qualitative research
Matthew DeCarlo (open textbook), Scientific Inquiry in Social Work, 10.3: Sampling in quantitative research
Matthew DeCarlo (open textbook), Scientific Inquiry in Social Work, 10.4: A word of caution, questions to ask about samples
Rajiv S. Jhangiani, I-Chant A. Chiang and colleagues (open textbook), Research Methods in Psychology: Experimental Design (random assignment and random sampling)
William M.K. Trochim, Research Methods Knowledge Base: Sampling Terminology
William M.K. Trochim, Research Methods Knowledge Base: Probability Sampling
William M.K. Trochim, Research Methods Knowledge Base: Nonprobability Sampling
William M.K. Trochim, Research Methods Knowledge Base: External Validity
Department of Statistics, Penn State Eberly College of Science, STAT 506 Sampling Theory and Methods, Lesson 1: Estimating Population Mean and Total under SRS
Department of Statistics, Penn State Eberly College of Science, STAT 506 Sampling Theory and Methods, Lesson 3: Unequal Probability Sampling
Department of Statistics, Penn State Eberly College of Science, STAT 506 Sampling Theory and Methods, Lesson 6: Stratified Sampling
Department of Statistics, Penn State Eberly College of Science, STAT 506 Sampling Theory and Methods, Lesson 7: Cluster and Systematic Sampling
Department of Statistics, Penn State Eberly College of Science, STAT 506 Sampling Theory and Methods, Lesson 9: Multi-Stage Designs
United Nations Statistics Division, Designing Household Survey Samples: Practical Guidelines (Studies in Methods, Series F No. 98)
European Social Survey ERIC, European Social Survey Round 11 Sampling Guidelines: Principles and Implementation
American Association for Public Opinion Research, Standard Definitions: Final Dispositions of Case Codes and Outcome Rates for Surveys, 10th edition (2023)
American Association for Public Opinion Research, Response Rates: An Overview
Pew Research Center, Understanding the margin of error in election polls (2016)
Pew Research Center, Oversampling is used to study small groups, not bias poll results (2016)
Pew Research Center, How different weighting methods work (2018)
Pew Research Center, The American Trends Panel
Pew Research Center, Evaluating Online Nonprobability Surveys (2016)
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Selection bias
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Non-response bias
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Volunteer bias
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Admission rate bias
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Attrition bias
Lawrence A. Palinkas, Sarah M. Horwitz, Carla A. Green, Jennifer P. Wisdom, Naihua Duan and Kimberly Hoagwood, Purposeful Sampling for Qualitative Data Collection and Analysis in Mixed Method Implementation Research (Administration and Policy in Mental Health, 2015)
K. Malterud, V.D. Siersma and A.D. Guassora, Sample Size in Qualitative Interview Studies: Guided by Information Power (Qualitative Health Research, 2016), abstract
Ylona Chun Tie, Melanie Birks and Karen Francis, Grounded theory research: A design framework for novice researchers (SAGE Open Medicine, 2019)
Daniel Lakens (open textbook, largely identical to his 2022 article in Collabra: Psychology), Improving Your Statistical Inferences, chapter 8: Sample Size Justification
Jonathan A. Cook, Steven A. Julious, William Sones and colleagues, DELTA2 guidance on choosing the target difference and undertaking and reporting the sample size calculation for a randomised controlled trial (BMJ, 2018)
Sabhya Gupta and Sarah Kopper, Abdul Latif Jameel Poverty Action Lab (J-PAL), Power calculations
K.P. Suresh and S. Chandrashekara, Sample size estimation and power analysis for clinical research studies (Journal of Human Reproductive Sciences, 2012)
Yousef Alimohamadi and Mojtaba Sepandi, Considering the design effect in cluster sampling (Journal of Cardiovascular and Thoracic Research, 2019)
Mohamed Elfil and Ahmed Negida, Sampling methods in Clinical Research; an Educational Review (Emergency, 2017)
Cecilia Maria Patino and Juliana Carvalho Ferreira, Inclusion and exclusion criteria in research studies: definitions and why they matter (Jornal Brasileiro de Pneumologia, 2018)
Belinda Thewes, Judith A.C. Rietjens, Sanne W. van den Berg and colleagues, One way or another: The opportunities and pitfalls of self-referral and consecutive sampling as recruitment strategies for psycho-oncology intervention trials (Psycho-Oncology, 2018)
Mark É. Czeisler, Joshua F. Wiley, Charles A. Czeisler and colleagues, Uncovering survivorship bias in longitudinal mental health surveys during the COVID-19 pandemic (Epidemiology and Psychiatric Sciences, 2021)
A. Parker, S. Treweek and colleagues, Strategies to improve recruitment to randomised trials (Cochrane Database of Systematic Reviews, 2026), abstract
K. Gillies, S. Treweek and colleagues, Strategies to improve retention in randomised trials (Cochrane Database of Systematic Reviews, 2021), abstract
Population Health Methods, Columbia University Mailman School of Public Health, Respondent-Driven Sampling
Lee Harvey, Quality Research International, Social Research Glossary: Gatekeeper
Allison Tong, Peter Sainsbury and Jonathan Craig (checklist as published by F1000Research), COREQ: Consolidated criteria for reporting qualitative studies, 32-item checklist
STROBE Initiative, STROBE Statement: checklist for cross-sectional studies
Saul McLeod, Simply Psychology, Sampling Methods in Research