Validity, reliability and bias in research: a glossary
The words for what a study measures and how far to trust it: kinds of variable, levels of measurement, reliability and validity, the accuracy of tests, and the biases that distort results. Every definition was checked against the sources listed below.
- Accuracy (measurement)
- Allocation bias
- Apprehension bias
- Area under the ROC curve
- Ascertainment bias
- Bias
- Binary variable
- Bland-Altman plot
- Calibration (measurement)
- Categorical variable
- Ceiling effect
- Classical test theory
- Cohen's kappa
- Collider bias
- Common method bias
- Compensatory equalisation of treatments
- Compensatory rivalry
- Concept
- Conceptual definition
- Conceptualisation
- Concurrent validity
- Confirmation bias
- Confounding
- Confounding by indication
- Confounding variable
- Construct
- Construct validity
- Content validity
- Continuous variable
- Control variable
- Convergent validity
- Converging operations
- Correction for attenuation
- Covariate
- Criterion validity
- Criterion variable
- Criterion-referenced score
- Cronbach's alpha
- Cross-cultural validity
- Cut-off point
- Demand characteristics
- Dependent variable
- Detection bias
- Diagnostic accuracy
- Differential item functioning
- Differential misclassification
- Differential verification bias
- Diffusion of treatment
- Discrete variable
- Discriminant validity
- Ecological validity
- Effect modification
- End-digit preference
- Evaluation apprehension
- Experimental mortality
- Experimenter expectancy effect
- Explanatory variable
- Exposure
- External validity
- Extraneous variable
- Face validity
- False negative
- False positive
- Floor effect
- Generalisability theory
- Gold standard
- Halo effect
- Hawthorne effect
- Healthy user bias
- Healthy worker effect
- History (threat to validity)
- Hypothesis guessing
- Immortal time bias
- Incorporation bias
- Incremental validity
- Independent variable
- Index (composite measure)
- Index test
- Indicator
- Information bias
- Insensitive measure bias
- Instrumentation (threat to validity)
- Inter-rater reliability
- Interaction of testing and treatment
- Internal consistency
- Internal validity
- Interpretability
- Interval scale
- Interviewer bias
- Intra-rater reliability
- Intraclass correlation coefficient
- Item analysis
- Item difficulty
- Item discrimination
- Item response theory
- Known-groups validity
- Kuder-Richardson Formula 20
- Latent variable
- Lead time bias
- Length time bias
- Levels of measurement
- Likelihood ratio
- Limits of agreement
- Lurking variable
- Manipulation check
- Maturation (threat to validity)
- McDonald's omega
- Measurement
- Measurement error
- Measurement invariance
- Mediator
- Minimal clinically important difference
- Misclassification
- Moderator
- Mono-method bias
- Mono-operation bias
- Multitrait-multimethod matrix
- Mundane realism
- Negative predictive value
- Nocebo effect
- Nominal scale
- Nomological network
- Non-differential misclassification
- Norm-referenced score
- Norms
- Novelty bias
- Nuisance variable
- Observed variable
- Observer bias
- Operational definition
- Operationalisation
- Ordinal scale
- Outcome variable
- Parallel-forms reliability
- Partial verification bias
- Percentage agreement
- Performance bias
- Placebo effect
- Positive predictive value
- Precision (measurement)
- Predictive validity
- Predictor variable
- Primary outcome
- Psychological realism
- Psychometric properties
- Psychometrics
- Quantitative variable
- Random error
- Ratio scale
- Reactivity
- Recall bias
- Reference standard
- Regression dilution bias
- Regression to the mean
- Reification
- Reliability
- Repeatability
- Reporting bias
- Reproducibility (measurement)
- Resentful demoralisation
- Residual confounding
- Response bias
- Response variable
- Responsiveness
- ROC curve
- Scale (multi-item measure)
- Secondary outcome
- Selection (threat to validity)
- Selection interaction
- Sensitivity (of a test)
- Spearman-Brown prophecy formula
- Specificity (of a test)
- Spectrum bias
- Split-half reliability
- Spontaneous remission
- Spurious relationship
- Standard error of measurement
- Standardised instrument
- Statistical conclusion validity
- Structural validity
- Surrogate endpoint
- Systematic error
- Target condition
- Test-retest reliability
- Testing (threat to validity)
- Threats to internal validity
- Translation validity
- True negative
- True positive
- True score
- Trueness
- Typology
- Unidimensionality
- Validity
- Variable
- Verification bias
- Youden index
Accuracy (measurement)
Also called: measurement accuracy, accuracy of measurement
Accuracy, in measurement, is how close a measured value comes to the true value of what is being measured. The international vocabulary of metrology treats it as a quality rather than a number and separates it from precision: a scale that always reads two kilograms heavy gives precise but inaccurate readings. Accuracy depends on both trueness and precision.
Allocation bias
Allocation bias is a systematic difference between the groups in a controlled trial that arises from how participants were assigned to them, for example when whoever enrols patients can foresee the next assignment. Randomisation with concealed allocation is the defence, and Cochrane's 2011 tool assessed it under selection bias.
Apprehension bias
Apprehension bias is a distortion in a measurement that occurs because a participant is anxious about being observed or assessed, often without being aware of it. The standard example is white coat hypertension, blood pressure that reads higher in the clinic than at home. It differs from the Hawthorne effect, which is a more deliberate change of behaviour under observation.
Area under the ROC curve
Also called: AUC, AUROC, area under the curve, ROC AUC, c-statistic, C statistic, concordance statistic
The area under the ROC curve is a single summary of a test's ability to discriminate, equal to the probability that a randomly chosen person with the condition is ranked as more likely to have it than a randomly chosen person without it. It runs from 0.5, no better than chance, to 1, perfect discrimination.
Ascertainment bias
Ascertainment bias is a distortion that arises when the way cases or outcomes are identified makes some people more likely than others to be counted in a study's results. The Catalogue of Bias links it to sampling, selection, detection and observer bias, and because writers use it in more than one of these senses it is worth defining whenever it is used.
Bias
Also called: research bias, study bias, systematic bias
Bias is a systematic error, or deviation from the truth, in a study's results or in the inferences drawn from them, which can push an estimate too high or too low. Unlike random error, which averages out over repeated studies, it is not reduced by making a study larger. Cochrane assesses each study's risk of bias rather than its general quality.
Binary variable
Also called: dichotomous variable, dichotomous data, binary data, binary outcome, dichotomous outcome
A binary variable is a categorical variable with exactly two possible values, such as yes or no, alive or dead, or passed or failed, often coded 0 and 1 for analysis. It is the simplest kind of nominal variable, and many trial outcomes, such as whether a patient recovered, are recorded this way.
Bland-Altman plot
Also called: Bland-Altman analysis, Bland Altman plot, Bland and Altman plot, Bland-Altman method
A Bland-Altman plot is a graph for judging whether two methods of measuring the same thing agree, plotting the difference between each pair of readings against their mean. Introduced by Bland and Altman in the 1980s, it shows any systematic difference and the limits of agreement, which a correlation coefficient cannot, since two methods can correlate closely yet disagree.
Calibration (measurement)
Also called: calibration, instrument calibration, measurement calibration
Calibration is the procedure of comparing an instrument's readings with measurement standards of known value and using that relationship to obtain correct results from future readings. The international vocabulary of metrology distinguishes it from adjusting the instrument and from checking a calibration, and in everyday use the first step, the comparison, is often all that is meant.
Categorical variable
Also called: categorical data, qualitative variable, categorical variables
A categorical variable is a variable whose values are categories or labels rather than measured quantities, such as blood type, nationality or level of education. It can be nominal, with no order, or ordinal, with an order. OpenStax calls such data qualitative, a sense that has nothing to do with qualitative research.
Ceiling effect
Also called: ceiling effects
A ceiling effect occurs when a considerable share of participants score at or near the highest possible score on a measure, so it cannot tell them apart or show further improvement. A reading test far too easy for a class is an example. In orthopaedic outcome research a ceiling effect is usually declared when 15% or more of a sample reach the top score.
Classical test theory
Also called: CTT, true score theory, true-score theory, classical true score theory
Classical test theory is the measurement model that treats every observed score as the sum of a true score and random error, written X = T + E. Reliability follows from it as the share of observed-score variance that is true-score variance. It works with total scores, whereas item response theory models each item.
Cohen's kappa
Also called: kappa, kappa statistic, kappa coefficient, Cohen kappa, κ
Cohen's kappa is a statistic for agreement between two raters sorting the same cases into categories, which corrects the observed agreement for the agreement expected by chance. It runs from minus one to one, where zero means chance-level agreement. Thresholds for a good kappa differ, and McHugh argued that health research should demand more than Cohen's original bands suggest.
Collider bias
Also called: collider stratification bias, collider-stratification bias, conditioning on a collider
Collider bias is a distortion of the association between an exposure and an outcome that arises when a study selects on, stratifies by or adjusts for a variable that both of them cause, called a collider. Restricting a study to hospital patients, for instance, can create a link between two diseases that both lead to admission.
Common method bias
Also called: common method variance, CMV, common method biases, shared method variance, method bias
Common method bias is distortion in the relationships between measures that comes from measuring them in the same way, for example with one self-report questionnaire completed at one sitting, rather than from the constructs themselves. Shared method can inflate the correlations between measures of different traits. Podsakoff and colleagues' 2003 review set out its sources and procedural and statistical remedies.
Compensatory equalisation of treatments
Also called: compensatory equalization of treatments, compensatory equalisation, compensatory equalization, compensatory equalization of treatment
Compensatory equalisation of treatments is a threat to internal validity in which administrators or staff, feeling the comparison group has been treated unfairly, give it extra resources or services to make up. This narrows the difference between the groups, so a programme that works can look as though it does not.
Compensatory rivalry
Also called: John Henry effect
Compensatory rivalry is a threat to internal validity in which a comparison group, learning what the intervention group receives, works harder to keep up or prove a point, narrowing the difference between them. It is also called the John Henry effect, after Saretsky's 1972 study of schools. How often it matters in practice is disputed.
Concept
Also called: concepts
A concept is the idea or mental image people form from a cluster of related observations, such as masculinity, poverty or stress, which a study must define before it can measure. Concepts are abstract and can have several dimensions. In quantitative research, a concept that cannot be observed directly and is studied as a variable is usually called a construct.
Conceptual definition
Also called: conceptual definitions
A conceptual definition is a statement of what a concept or construct means, describing its parts and how it relates to other variables, written before deciding how to measure it. For neuroticism it would describe a tendency to experience negative emotions. The operational definition then says exactly how that meaning will be turned into a score.
Conceptualisation
Also called: conceptualization, conceptualising, conceptualizing
Conceptualisation is the process of writing clear, precise definitions of a study's key concepts, drawing on theory and earlier research. In quantitative research it is settled before data are collected, while qualitative researchers often let a working definition develop from what participants say. It comes before operationalisation, which decides how each concept will be measured.
Concurrent validity
Concurrent validity is the kind of criterion validity shown when scores on a measure correlate with a criterion measured at about the same time, such as a new depression questionnaire checked against a clinician's diagnosis made that week. The Research Methods Knowledge Base frames it differently, as the measure's ability to tell apart groups it should distinguish.
Confirmation bias
Confirmation bias is the tendency to look for, favour and use information that supports what one already believes or hopes to find, and to discount what contradicts it. In research it can shape which data are collected, how ambiguous results are read and when a search for evidence stops. Francis Bacon described the tendency in 1620.
Confounding
Also called: confounding bias, confounding effect, third variable problem, third-variable problem
Confounding is a distortion of the association between an exposure and an outcome caused by a third factor linked to both, which can create an apparent effect or hide a real one. Randomisation is the strongest protection at the design stage. In analysis, researchers adjust for measured confounders by stratification or regression, but unmeasured confounders remain.
Confounding by indication
Confounding by indication is confounding in which the reason a treatment was given, such as a fever or a more severe illness, is itself the true cause of an outcome blamed on the treatment. It is common in observational studies of medicines: children given paracetamol seemed more likely to develop asthma, but the infections that prompted the paracetamol may explain it.
Confounding variable
Also called: confounder, confounders, confounding factor, confound
A confounding variable, or confounder, is a variable associated with both the exposure or independent variable and the outcome, and not on the causal path between them, so it offers a rival explanation for their association. Smoking confounds a link between alcohol and heart disease. In an experiment, a confound is an extraneous variable that differs across conditions.
Construct
Also called: constructs, psychological construct, hypothetical construct, theoretical construct
A construct is an attribute that cannot be observed directly, such as intelligence, self-esteem or anxiety, which researchers infer from behaviour, answers or physiological signs and measure through indicators. Because a construct is never seen directly, its measures need evidence of validity. Constructs belong to the theoretical level of a study and their indicators to the observable level.
Construct validity
Construct validity is the extent to which a measure, or a study's manipulation, reflects the construct it is meant to represent, judged by whether its scores relate to other variables as theory predicts. Modern test theory treats it as the overarching form of validity, with content, criterion, convergent and discriminant evidence all contributing to one argument.
Content validity
Content validity is the extent to which a measure covers the whole of the construct it is meant to capture, with nothing important missing and nothing irrelevant added. It is usually judged by experts comparing the items with a clear definition of the construct or a test blueprint, rather than calculated. COSMIN treats face validity as part of it.
Continuous variable
Also called: continuous data, continuous measure
A continuous variable is a quantitative variable that can take any value within a range, including fractions and decimals, because it is measured rather than counted. Height, weight, time and blood pressure are examples. It contrasts with a discrete variable, which can take only certain separate values, usually counts.
Control variable
Also called: controlled variable, control variables
A control variable is a variable whose influence a researcher removes so that it cannot distort the relationship of interest, either by holding it constant in the design, such as testing everyone in the same room, or by adjusting for it in the analysis. DeCarlo describes the second sense as controlling for a variable mathematically to rule out a third-variable explanation.
Convergent validity
Convergent validity is evidence that scores on a measure correlate strongly with scores on other measures of the same or closely related constructs, as theory says they should. A new self-esteem scale, for example, should correlate with an established one. It is judged together with discriminant validity, and both count as evidence of construct validity.
Converging operations
Converging operations is the practice of measuring the same construct in several different ways, such as by self-report, behaviour and physiology, and checking that the results agree. When different operational definitions of a construct give similar results, that is good evidence that the construct is being measured well.
Correction for attenuation
Also called: disattenuation, disattenuated correlation, attenuation correction
Correction for attenuation is a calculation that estimates what the correlation between two measures would be if neither contained random measurement error, using the reliability of each. It is used in validity studies to show how far unreliability may be hiding a relationship, and the corrected figure is an estimate about true scores, not an observed result.
Covariate
Also called: covariates, concomitant variable, co-variate
A covariate is a variable, other than the treatment or exposure of main interest, that is measured and included in an analysis because it helps explain variation in the outcome. In analysis of covariance it is a continuous variable adjusted for so that groups are compared fairly and error variance is reduced. Older texts called covariates concomitant or nuisance variables.
Criterion validity
Also called: criterion-related validity, criterion related validity
Criterion validity is the extent to which scores on a measure correlate with an outside measure, the criterion, that they should relate to, often an accepted gold standard. Its two forms, concurrent and predictive, differ in whether the criterion is measured at the same time or later. It can be no better than the criterion, and unreliability in either measure weakens it.
Criterion variable
Also called: criterion, criterion measure
A criterion variable is the variable a researcher is trying to predict, the outcome in a prediction or regression study and the counterpart of a predictor variable. Baron and Kenny, for example, write of predictor and criterion where others would say independent and dependent variable. In validity studies, the criterion is the standard a new measure is checked against.
Criterion-referenced score
Also called: criterion-referenced test, criterion-referenced assessment, criterion referencing, criterion-referenced, criterion referenced test, content-referenced test
A criterion-referenced score interprets a result against a fixed standard of what someone should know or be able to do, such as a pass mark or a mastery level, regardless of how others performed. A driving test works this way. It contrasts with a norm-referenced score, which ranks a person against a norm group.
Cronbach's alpha
Also called: Cronbach alpha, Cronbachs alpha, coefficient alpha, alpha coefficient, Cronbach's α, α
Cronbach's alpha is a coefficient, usually between 0 and 1, that estimates the internal consistency reliability of a multi-item scale from one administration, and equals the mean of all possible split-half coefficients. Values from 0.70 to 0.95 are often called acceptable. A high alpha does not show that the scale measures one thing, and alpha rises with the number of items.
Cross-cultural validity
Also called: cross cultural validity
Cross-cultural validity is the degree to which the items of a translated or culturally adapted instrument perform in the same way as the items of the original version. COSMIN counts it as part of construct validity and pairs it with measurement invariance, which tests whether a measure works the same way across language or cultural groups.
Cut-off point
Also called: cut-off, cutoff, cut-point, cutpoint, threshold, positivity threshold, cut-off score, cutoff score, criterion of positivity
A cut-off point is the value of a test or scale above or below which a result is classed as positive, such as the score at which a screening questionnaire flags possible depression. Moving it trades sensitivity against specificity, so the choice depends on whether missed cases or false alarms matter more. STARD asks studies to define and justify their cut-offs.
Demand characteristics
Also called: demand characteristic, demand effects, demand effect
Demand characteristics are the cues in a study that let participants work out what the researcher is investigating or expects, which can lead them to behave as they think they should rather than naturally. Orne, who named them in 1962, described the good subject who tries to help the research succeed. Evidence for the effect outside laboratories is thin.
Dependent variable
Also called: DV, dependent variables
A dependent variable is the variable a researcher measures to see whether it changes when the independent variable changes, the presumed effect in a cause-and-effect relationship. In an experiment on noise and memory, the number of words recalled is the dependent variable. Statisticians often call it the response or outcome variable.
Detection bias
Detection bias is a systematic difference between the groups in a study in how outcomes are looked for, measured or confirmed, for instance when assessors who know who received the treatment judge a subjective outcome differently. Cochrane's 2011 tool linked it to blinding of outcome assessment, and its RoB 2 tool calls it bias in measurement of the outcome.
Diagnostic accuracy
Also called: diagnostic test accuracy, DTA, test accuracy, diagnostic performance
Diagnostic accuracy is the ability of a test to distinguish people who have a target condition from those who do not, measured against a reference standard and expressed as sensitivity, specificity, predictive values, likelihood ratios or the area under the ROC curve. A diagnostic accuracy study compares an index test with a reference standard in the same people, and STARD guides its reporting.
Differential item functioning
Also called: DIF
Differential item functioning is present when people from different groups, such as men and women or speakers of different languages, who stand at the same level of the underlying trait have different chances of giving a particular answer to an item. Some of it reflects real group differences, but DIF caused by a different understanding of a word is a form of measurement bias.
Differential misclassification
Also called: differential misclassification bias, differential information bias
Differential misclassification is misclassification of exposure or outcome whose rate differs between the groups being compared, for example when cases recall past exposures more completely than controls. Because the errors are unequal, it can make an association look stronger or weaker than it really is, and the direction is hard to predict.
Differential verification bias
Also called: differential reference bias, differential verification
Differential verification bias arises in a diagnostic accuracy study when participants have their true status confirmed by different reference standards, for example a definitive test for some and a less exact one for others. Combining results from references of different quality distorts the index test's apparent accuracy, in a direction that is hard to predict.
Diffusion of treatment
Also called: diffusion or imitation of treatment, diffusion of treatments, treatment diffusion, imitation of treatment
Diffusion of treatment is a threat to internal validity in which people in the comparison group learn about or copy the intervention from those receiving it, so the groups become more alike and a real effect is harder to detect. It is closely related to what trialists call contamination. Keeping groups apart, for example at different sites, reduces it.
Discrete variable
Also called: discrete data
A discrete variable is a quantitative variable that can take only certain separate values, usually whole numbers produced by counting, such as the number of children in a family or of phone calls in a day. It differs from a continuous variable, which can take any value within a range.
Discriminant validity
Also called: divergent validity
Discriminant validity is evidence that scores on a measure do not correlate strongly with measures of constructs that theory says are distinct from it. A test of mood, for instance, should not simply track intelligence. Together with convergent validity it supports construct validity, and the multitrait-multimethod matrix was designed to examine both at once.
Ecological validity
Ecological validity is usually taken to mean the extent to which a study's tasks, settings and findings resemble and apply to everyday life outside the laboratory. The term has drifted from Brunswik's original 1949 sense, the statistical link between a sensory cue and the object it signals, and Holleman and colleagues argue it is now used too loosely to be useful.
Effect modification
Also called: effect modifier, effect modifiers
Effect modification is present when the effect of an exposure on an outcome differs in size or direction across levels of a third variable, the effect modifier, as when the effect of smoking on lung cancer differs with radon exposure. Unlike a confounder, an effect modifier is a real feature to report by subgroup, not a bias to adjust away.
End-digit preference
Also called: end-digit preference bias, end digit preference, terminal digit preference, digit preference
End-digit preference is a recording error in which people reading or writing down measurements round to favoured final digits, usually zero or five, or to a value just on one side of a treatment threshold. It is well documented in manual blood pressure readings and is reduced by automatic digital devices and standard measurement protocols.
Evaluation apprehension
Evaluation apprehension is participants' anxiety about being judged in a study, which can make some perform worse from nerves and others better because they want to look good, so the anxiety rather than the treatment shapes the results. The Research Methods Knowledge Base lists it as a threat to construct validity.
Experimental mortality
Also called: mortality threat, mortality (threat to validity), attrition threat
Experimental mortality is the name research design texts give to attrition as a threat to internal validity: participants drop out between pre-test and post-test, and if those who leave differ from those who stay, the change in the group reflects who remains, not the treatment. When dropout differs between groups it becomes a selection interaction.
Experimenter expectancy effect
Also called: experimenter expectancy, experimenter expectancy effects, experimenter expectancies, observer-expectancy effect, expectancy effect
The experimenter expectancy effect is an unintended influence of a researcher's expectations on how participants behave or how their responses are recorded, usually through small differences in manner or procedure. Rosenthal's studies, in which students' beliefs about their rats changed the rats' measured performance, made it famous. Double-blind procedures and standardised scripts reduce it.
Explanatory variable
Also called: explanatory variables
An explanatory variable is a variable thought to cause or account for changes in another variable, called the response variable. In the OpenStax account of experiments it is the variable the researcher manipulates, and its different values are the treatments, so in that setting it is the independent variable under another name.
Exposure
Also called: exposure variable, exposures
An exposure is the factor whose possible effect on an outcome a study examines, such as smoking, a medicine, an occupation or a policy, a term used especially in epidemiology and observational research. It plays the part of the independent variable but is usually observed rather than assigned. Studies compare exposed with unexposed people, or different levels of exposure.
External validity
External validity is the degree to which a study's conclusions would hold for other people, in other places and at other times. The Research Methods Knowledge Base names these three as its main threats and sees replication across varied settings as the strongest support. Sampling texts discuss the same question as generalisability.
Extraneous variable
Also called: extraneous variables, extraneous factor
An extraneous variable is any variable other than the independent and dependent variables that could affect the outcome of a study, such as participants' mood, the time of day or the room used. Researchers control extraneous variables by holding them constant or by random assignment. An extraneous variable that differs systematically between conditions becomes a confounding variable.
Face validity
Face validity is the extent to which a measure appears, on the face of it, to measure what it is meant to, judged informally by researchers, participants or experts. Jhangiani and colleagues rate it as very weak evidence, since intuitions are often wrong and some well-established measures work without it. COSMIN treats it as one aspect of content validity.
False negative
Also called: false negatives, false-negative result, FN
A false negative is a test result that says a condition or feature is absent when it is in fact present, such as a screening test that misses a disease. The false negative rate is one minus the sensitivity. In hypothesis testing the matching error, failing to detect an effect that exists, is called a Type II error.
False positive
Also called: false positives, false-positive result, FP, false alarm
A false positive is a test result that says a condition or feature is present when it is in fact absent, such as a screening test that flags a healthy person. The false positive rate is one minus the specificity. When a condition is rare, even a test with high specificity can produce more false positives than true positives.
Floor effect
Also called: floor effects
A floor effect occurs when a considerable share of participants score at or near the lowest possible score on a measure, so it cannot distinguish among them or register further decline. A test far too hard for a group shows one. COSMIN notes that floor and ceiling effects can signal weak content validity and can reduce a measure's reliability.
Generalisability theory
Also called: generalizability theory, G theory, G-theory
Generalisability theory is an extension of reliability theory, introduced by Cronbach and colleagues in 1971, that estimates how much of the variation in scores comes from each source of error, such as raters, occasions or test forms, instead of treating all error as one. Cronbach later preferred it to coefficient alpha, though it never became as widely used.
Gold standard
Also called: gold-standard test, gold standard test
A gold standard is a method of establishing whether a condition is present that is treated as error-free, the ideal against which a new test is judged. STARD 2015 describes a gold standard as an error-free reference standard, and because few methods are error-free, diagnostic researchers increasingly speak of the reference standard, the best method available, instead.
Halo effect
Also called: halo error, halo bias
The halo effect is a rating error in which a rater's overall impression of a person or thing colours judgements of its separate qualities, so that ratings of distinct traits come out too high, too even and too closely correlated. Thorndike described it in 1920. It threatens any study that relies on observers' or examiners' ratings of several attributes.
Hawthorne effect
Also called: Hawthorne effects
The Hawthorne effect is a change in people's behaviour caused by knowing they are being observed or studied, as when hand hygiene improves while staff are watched. It is named after studies at the Hawthorne Works factory near Chicago between 1924 and 1933. A 2014 systematic review found no single Hawthorne effect and proposed speaking of research participation effects instead.
Healthy user bias
Also called: healthy user effect
Healthy user bias arises in observational studies when people who use a treatment or preventive measure are healthier, more health-conscious or more engaged with care than those who do not, so part or all of the apparent benefit reflects who they are. It has affected studies of hormone therapy, statins and vaccines, and much of it survives adjustment.
Healthy worker effect
Also called: healthy worker bias
The healthy worker effect is a selection bias in occupational studies that arises because people in work are, on average, healthier than the general population, which includes those too ill to work. Comparing a workforce with the general population can therefore make a hazardous job look safer than it is.
History (threat to validity)
Also called: history threat, history effect
History is a threat to internal validity in which an event outside the study, occurring between the pre-test and the post-test, causes the change that is credited to the intervention. A television programme on the same topic broadcast during an anti-drug course is an example. A comparison group exposed to the same events is the usual protection.
Hypothesis guessing
Hypothesis guessing is a threat to construct validity in which participants try to work out what a study is testing and change their behaviour to fit, or to defy, their guess, so the outcome reflects the guess rather than the treatment. The Research Methods Knowledge Base lists it beside evaluation apprehension and experimenter expectancies.
Immortal time bias
Also called: immortal-time bias
Immortal time bias arises in cohort studies when follow-up includes a period during which people in the exposed group could not have had the outcome, usually because they were classed as exposed only after surviving long enough to start treatment. Crediting that time to the treated group makes the treatment look more protective than it really is.
Incorporation bias
Incorporation bias is a distortion in a diagnostic accuracy study that occurs when the result of the test being evaluated forms part of the reference standard it is compared with, typically when the reference combines several tests. Because the two are no longer independent, the index test's measured accuracy is distorted. The Catalogue of Bias counts it as a kind of verification bias.
Incremental validity
Incremental validity is the extent to which a measure improves the prediction of a criterion beyond what other measures or information already available can predict. It is usually tested by adding the new measure last in a hierarchical regression and seeing how much it adds. Hunsley and Meyer argued in 2003 that applied psychology had studied it too little.
Independent variable
Also called: IV, independent variables, manipulated variable, treatment variable
An independent variable is the variable a researcher manipulates or compares, the presumed cause whose effect on the dependent variable a study tests. In an experiment its values are the conditions or treatments, such as different doses of a drug. Outside experiments the term is also used for presumed causes that are only measured, such as age or income.
Index (composite measure)
Also called: index, composite index, indices, indexes
An index, in measurement, is a composite score that combines several indicators of a broader concept into one number, such as a deprivation index built from income, employment and housing. In DeCarlo's usage it differs from a scale, whose items vary in intensity, but other textbooks draw the line between index and scale differently.
Index test
Also called: index tests, test under evaluation
An index test is the test under evaluation in a diagnostic accuracy study, whose results are compared with those of a reference standard in the same people. It can be any way of gathering information about health, from a blood test or scan to a questionnaire or clinical sign, and STARD asks authors to state its intended use and role.
Indicator
Also called: indicators, observable indicator
An indicator is an observable sign, such as an answer, a behaviour or a recorded fact, taken to show the presence or level of a concept that cannot itself be observed. Several indicators are often combined into an index or scale. In factor analysis and structural equation models, the observed variables that reflect a latent variable are called its indicators.
Information bias
Also called: information biases
Information bias is any systematic error that arises in how information is collected, recalled, recorded or handled in a study, including how missing data are dealt with. Recall bias, observer bias, reporting bias and misclassification are its main forms. It can be differential between groups, distorting an association either way, or non-differential.
Insensitive measure bias
Also called: insensitive measure
Insensitive measure bias arises when an outcome is measured with a method too crude to detect differences that matter, so a real effect goes unnoticed and a study wrongly concludes there is none. Sackett listed it among the biases of research in 1979, and it is a form of information bias. A measure with shown responsiveness helps avoid it.
Instrumentation (threat to validity)
Also called: instrumentation threat, instrumentation effect
Instrumentation is a threat to internal validity in which the measuring instrument or procedure changes between measurements, so scores shift for reasons that have nothing to do with the intervention. A revised test at post-test, or observers who grow tired, more skilled or less strict over a study, are common causes.
Inter-rater reliability
Also called: interrater reliability, inter rater reliability, inter-observer reliability, interobserver reliability, inter-judge reliability
Inter-rater reliability is the extent to which different observers or raters, judging the same cases, give consistent scores or classifications. It is estimated with Cohen's kappa for categories, the intraclass correlation coefficient for continuous ratings, or percentage agreement, which ignores chance. In qualitative research the parallel idea for coders is intercoder reliability.
Interaction of testing and treatment
Also called: testing-treatment interaction, pretest sensitisation, pretest sensitization, pre-test sensitisation
Interaction of testing and treatment is a validity threat in which taking a pre-test makes participants more receptive to the intervention, so the effect found may hold only for people who were pre-tested. The Research Methods Knowledge Base lists it as a threat to construct validity, because the test becomes part of the treatment.
Internal consistency
Also called: internal consistency reliability, internal reliability, inter-item consistency
Internal consistency is the extent to which the items of a multi-item measure give consistent results because they reflect the same underlying construct, assessed from a single administration. It is estimated with the average inter-item correlation, split-half reliability or, most often, Cronbach's alpha. High internal consistency does not by itself show that a scale measures only one construct.
Internal validity
Internal validity is the extent to which a study's design and conduct support the conclusion that the independent variable caused the observed change in the dependent variable, rather than something else. Experiments with random assignment and control of extraneous variables are strongest on it. History, maturation, testing, instrumentation, regression to the mean and selection are classic threats.
Interpretability
Also called: interpretability of scores
Interpretability is the degree to which qualitative meaning, such as what counts as a small or an important change, can be attached to an instrument's scores or changes in score. COSMIN treats it as an important characteristic rather than a measurement property, shown by reporting score distributions, floor and ceiling effects and the minimal important change.
Interval scale
Also called: interval level of measurement, interval level, interval data, interval variable, interval measurement
An interval scale is the level of measurement at which values are ordered and the distances between them are equal and meaningful, but zero is arbitrary rather than an absence of the quantity. Temperature in Celsius is the standard example: 20 degrees is not twice as hot as 10. Differences and means make sense, but ratios do not.
Interviewer bias
Interviewer bias is a systematic error introduced when an interviewer's knowledge, expectations or manner affects how questions are asked or answers recorded, for example probing cases more thoroughly than controls about past exposures. It is a form of information bias, reduced by structured questions, training and keeping interviewers unaware of each participant's group.
Intra-rater reliability
Also called: intrarater reliability, intra rater reliability, intra-observer reliability, intraobserver reliability
Intra-rater reliability is the consistency of one rater's judgements when the same person scores the same cases on different occasions, for example a physiotherapist measuring the same joint angles a week apart. It is usually estimated with an intraclass correlation coefficient or kappa. It differs from inter-rater reliability, which compares different raters.
Intraclass correlation coefficient
Also called: intraclass correlation, intra-class correlation coefficient, intra-class correlation, ICC
The intraclass correlation coefficient is a reliability statistic for continuous measurements that reflects both correlation and agreement between repeated measurements, used for test-retest, inter-rater and intra-rater reliability. It comes in ten forms, and Koo and Li read values below 0.5 as poor and above 0.9 as excellent. The intracluster correlation used in cluster sampling shares the abbreviation ICC.
Item analysis
Also called: item analyses, item statistics
Item analysis is the examination of the individual items in a test or questionnaire to judge how well each contributes to the whole, using statistics such as item difficulty, item discrimination and how reliability would change if the item were removed. Weak items are revised or dropped before a measure is finalised.
Item difficulty
Also called: item difficulty index, item p-value, item facility
Item difficulty is the average score on a test item, which for items scored right or wrong is the proportion of people who answer it correctly, often called the item's p-value. In item response theory, difficulty is instead the trait level at which a person has an even chance of giving the keyed response.
Item discrimination
Also called: discrimination index, item discrimination index
Item discrimination is how well a test item separates people with higher levels of the measured trait from those with lower levels. In classical test theory it is usually estimated with the item-total correlation, often corrected by leaving the item out of the total. In item response theory it is a parameter setting how steeply the chance of success rises with the trait.
Item response theory
Also called: IRT, item response theory models
Item response theory is a family of statistical models that link a person's position on a latent trait to the probability of each response to each item, using item properties such as difficulty and discrimination. Unlike classical test theory it models items rather than total scores, which lets measurement precision vary along the trait and supports adaptive testing.
Known-groups validity
Also called: known groups validity, known-group validity, discriminative validity
Known-groups validity is evidence that a measure gives different scores to groups that, on theory or prior evidence, should differ on the construct, such as patients and healthy volunteers on a symptom scale. COSMIN treats it as one form of hypothesis testing for construct validity. The Research Methods Knowledge Base describes the same idea under concurrent validity.
Kuder-Richardson Formula 20
Also called: KR-20, KR20, Kuder-Richardson 20, Kuder Richardson formula 20, Kuder-Richardson coefficient
Kuder-Richardson Formula 20 is a coefficient of internal consistency for tests whose items are scored right or wrong, published by Kuder and Richardson in 1937. Cronbach's 1951 paper presented coefficient alpha as the general formula of which it is a special case, so for items scored in two categories the two give the same value.
Latent variable
Also called: latent variables, latent construct, latent trait, latent factor, unobserved variable
A latent variable is a variable that cannot be observed or measured directly, such as depression or reading ability, and is inferred from several observed variables that it is assumed to influence. Factor analysis, structural equation modelling and item response theory are built around latent variables, treating differences in observed scores as due to the latent variable plus measurement error.
Lead time bias
Also called: lead-time bias
Lead time bias is an apparent gain in survival among people whose disease is found by screening, which arises because diagnosis is moved earlier, not because death is delayed. If screening finds a cancer two years sooner but the person dies at the same age, survival from diagnosis looks two years longer. The Catalogue of Bias calls it the rule, not the exception, in early detection.
Length time bias
Also called: length-time bias, length bias
Length time bias is a distortion in screening studies that arises because screening is more likely to find slowly progressing disease, which spends longer in a detectable early stage, than aggressive disease that appears between screens. Screen-detected cases then seem to survive longer even if screening changes nothing. Analysing by invitation to screening rather than by how cases were found reduces it.
Levels of measurement
Also called: level of measurement, scales of measurement, measurement scales, measurement levels, Stevens' levels of measurement
Levels of measurement are the four kinds of scale, nominal, ordinal, interval and ratio, that describe what the numbers or categories given to a variable can tell you, set out by Stevens in 1946. Each level has the properties of the one below it and more, and the level decides which summaries and statistical analyses make sense.
Likelihood ratio
Also called: likelihood ratios, LR, positive likelihood ratio, negative likelihood ratio, LR+, LR-
A likelihood ratio is a measure of diagnostic accuracy that says how much more likely a given test result is in people with the condition than in people without it. The positive ratio is sensitivity divided by one minus specificity, and the negative ratio is one minus sensitivity divided by specificity. Unlike predictive values, likelihood ratios do not depend on prevalence.
Limits of agreement
Also called: 95% limits of agreement, LoA, Bland-Altman limits of agreement
Limits of agreement are the range within which about 95% of the differences between two methods of measuring the same thing are expected to fall, calculated as the mean difference plus or minus 1.96 standard deviations of the differences. They are the usual summary of a Bland-Altman analysis and assume the differences are roughly normally distributed.
Lurking variable
Also called: lurking variables
A lurking variable is a variable that is not part of a study's design or analysis but influences the variables being studied, so it can create or hide an apparent relationship. OpenStax notes that random assignment spreads lurking variables evenly across groups. In observational research a lurking variable is, in effect, an unmeasured confounder.
Manipulation check
Also called: manipulation checks
A manipulation check is a separate measure of the construct an experimenter is trying to manipulate, used to confirm that the independent variable really changed as intended, for example asking participants how anxious they felt after an anxiety induction. Manipulation checks matter most when an experiment finds no effect, since a failed manipulation is one explanation.
Maturation (threat to validity)
Also called: maturation threat, maturation effect
Maturation is a threat to internal validity in which participants change between pre-test and post-test through growth, learning, ageing, fatigue or recovery that would have happened anyway, and the change is wrongly credited to the intervention. It threatens one-group designs most, especially with children, and a comparison group that matures at the same rate controls it.
McDonald's omega
Also called: omega, coefficient omega, omega total, McDonald omega, McDonald's ω, ω
McDonald's omega is a reliability coefficient for a scale, calculated from a factor analysis, that allows items to relate to the underlying construct with different strengths. Where items are unequally good measures, it avoids the underestimation that affects Cronbach's alpha, and it is increasingly recommended in alpha's place, though it too is unreliable with strongly skewed data.
Measurement
Also called: measuring
Measurement is the assignment of numbers or categories to people, objects or events according to rules, so that the values represent some characteristic of them. In research it runs from defining a concept to choosing an instrument and scoring it. Stevens warned that a measurement can be no better than the procedures used to obtain it.
Measurement error
Also called: error of measurement, observational error
Measurement error is the difference between a measured value and the true value of what is being measured. It has two components: random error, which scatters readings unpredictably, and systematic error, which pushes them consistently one way. No measure is free of it, and COSMIN counts the size of a measure's error as a measurement property in its own right.
Measurement invariance
Also called: measurement equivalence, factorial invariance, invariance testing
Measurement invariance is the property of a measure working the same way across groups or occasions, so that equal scores mean the same thing in each. It is tested in steps with factor analysis, configural (the same structure), metric (equal loadings) and scalar (equal intercepts), and group means can be compared fairly only once scalar invariance holds.
Mediator
Also called: mediating variable, mediator variable, intervening variable, intermediate variable, mediators
A mediator is a variable through which an independent variable produces its effect on a dependent variable, the mechanism in the middle of a causal chain, as when exercise might lift mood by improving sleep. Baron and Kenny set it apart from a moderator in 1986: a mediator explains how or why an effect occurs, a moderator when it holds.
Minimal clinically important difference
Also called: MCID, minimal important difference, minimally important difference, MID, minimal important change, MIC, minimum clinically important difference
The minimal clinically important difference is the smallest change or difference in a score that patients would regard as meaningful, as opposed to one that is merely statistically significant. Jaeschke, Singer and Guyatt introduced it in 1989 by comparing questionnaire changes with patients' own ratings of change. COSMIN recommends anchor-based methods of this kind for setting it.
Misclassification
Also called: misclassification bias, misclassification error
Misclassification is the placing of participants in the wrong category of exposure, outcome or another variable, such as recording a smoker as a non-smoker, which distorts the association a study finds. It is non-differential when the chance of error is the same in every group and differential when it differs between groups.
Moderator
Also called: moderator variable, moderating variable, moderators
A moderator is a variable that changes the direction or strength of the relationship between an independent or predictor variable and a dependent or criterion variable, as when a treatment helps younger patients more than older ones. It can be categorical, such as sex, or quantitative. Moderation is tested as a statistical interaction, and epidemiologists describe the same pattern as effect modification.
Mono-method bias
Also called: monomethod bias
Mono-method bias is a threat to construct validity that arises when a study measures its outcome in only one way, so the results may reflect that method rather than the construct. Using several kinds of measure, such as a questionnaire and observed behaviour, gives stronger evidence that the intended construct is being measured.
Mono-operation bias
Also called: monooperation bias
Mono-operation bias is a threat to construct validity that arises when a study uses a single version of its intervention, in one place at one time, so the findings may reflect that particular version rather than the broader construct it was meant to represent. Testing several versions or doses of the treatment reduces it.
Multitrait-multimethod matrix
Also called: MTMM, MTMM matrix, multi-trait multi-method matrix, multitrait multimethod matrix
The multitrait-multimethod matrix is a table of correlations among several traits each measured by several methods, introduced by Campbell and Fiske in 1959 to assess construct validity. Measures of the same trait by different methods should correlate highly, showing convergent validity, while different traits should correlate less, showing discriminant validity, even when measured the same way.
Mundane realism
Mundane realism is the degree to which the situation in an experiment resembles situations participants meet in everyday life. It is one route to external validity, but an experiment can lack it and still generalise if it has psychological realism, engaging the same mental processes that operate in real life.
Negative predictive value
Also called: NPV, negative predictive values
Negative predictive value is the proportion of people with a negative test result who truly do not have the condition, answering the question a person with a negative result asks. Unlike sensitivity and specificity, it depends on how common the condition is in the group tested: the rarer the condition, the higher the negative predictive value.
Nocebo effect
Also called: nocebo, nocebo response, nocebo effects
The nocebo effect is the occurrence of unpleasant symptoms or adverse effects after an intervention that are caused not by its physical action but by negative expectations, information or experience. About a quarter of patients taking placebo in trials report an adverse event, and what participants are told about possible side effects can change how many they report.
Nominal scale
Also called: nominal level of measurement, nominal level, nominal data, nominal variable, nominal measurement
A nominal scale is the level of measurement at which values are simply names or labels for categories with no order, such as blood type, marital status or smartphone brand. Numbers given to nominal categories, like the numbers on players' shirts, are labels only. Counts and percentages are meaningful, but averages are not.
Nomological network
Also called: nomological net
A nomological network is the set of lawful relationships a theory predicts between a construct, other constructs and observable measures, proposed by Cronbach and Meehl in 1955 as the basis for construct validity. A measure has construct validity to the extent that its scores behave as the network predicts, relating to other measures as the theory says.
Non-differential misclassification
Also called: nondifferential misclassification, non-differential misclassification bias, random misclassification
Non-differential misclassification is misclassification whose rate is the same in all the groups being compared, such as an exposure measured equally badly in cases and controls. With two categories it usually biases an association towards no effect, so a real link may be missed, but with more categories it can push estimates in either direction.
Norm-referenced score
Also called: norm-referenced test, norm-referenced assessment, norm referencing, norm-referenced, norm referenced test
A norm-referenced score gives meaning to a result by comparing it with the scores of a specific norm group, as percentiles and age or grade equivalents do. It shows how someone performed relative to others, not what they can do. A criterion-referenced score instead compares performance with a fixed standard of content or skill.
Norms
Also called: normative data, norm group, normative sample, test norms
Norms are the distribution of scores on a standardised test in a defined reference group, such as children of a given age or school year, used to interpret an individual's score by comparison with that group. Percentile ranks and age equivalents are based on norms. They are only as useful as the norm group is relevant to the person tested.
Novelty bias
Novelty bias is the tendency for an intervention to look better when it is new than it later proves to be. It is thought to combine several other biases, such as careful selection of early participants and quicker reporting of good results, and studies have estimated that it inflates the apparent benefit of new treatments by between 2% and 27%.
Nuisance variable
Also called: nuisance factor, nuisance factors, nuisance variables
A nuisance variable is a source of variation in the outcome that is not of interest in itself but can obscure the effect being studied, such as batches of material, operators or days in an experiment. Designers handle known nuisance factors by blocking, adjust for measured continuous ones as covariates, and use random assignment to spread unknown ones across groups.
Observed variable
Also called: manifest variable, manifest variables, measured variable, observed variables
An observed variable is a variable measured directly, such as an answer to a questionnaire item, a test score or a recorded behaviour, as opposed to a latent variable that is inferred. In factor analysis and structural equation modelling, observed variables serve as indicators of latent variables, and their variation is attributed to those latent variables and to measurement error.
Observer bias
Also called: observer biases
Observer bias is a systematic difference between the true value and the value recorded, caused by the person observing or measuring, through their expectations, judgement or habits such as rounding. It matters most for subjective outcomes, such as reading a scan, and least for objective ones, such as death. Training observers and keeping them unaware of group assignment reduce it.
Operational definition
Also called: operational definitions, operationalised definition
An operational definition is a definition of a variable in terms of exactly how it will be measured or manipulated in a particular study, such as defining stress as a score on a named questionnaire. It should say what is measured, with which instrument and how scores will be interpreted. Self-report, behavioural and physiological measures are the three broad kinds.
Operationalisation
Also called: operationalization, operationalising, operationalizing, operationalise, operationalize
Operationalisation is the process of turning an abstract concept into something that can be measured, by choosing the indicators, instrument and scoring rules that will stand for it in a study. It follows conceptualisation and produces an operational definition. How well it has been done is the question construct validity asks.
Ordinal scale
Also called: ordinal level of measurement, ordinal level, ordinal data, ordinal variable, ordinal measurement
An ordinal scale is the level of measurement at which categories can be put in order but the distances between them are not known to be equal, such as satisfaction ratings from very dissatisfied to very satisfied, or finishing positions in a race. Ranks and medians are meaningful, but differences between categories cannot be measured.
Outcome variable
Also called: outcome, outcome measure, outcomes, endpoint, end point, outcome variables
An outcome variable is the variable a study measures to judge the effect of an intervention or exposure, such as blood pressure, exam results or survival. In trials and health research, outcome is the usual word where experimental psychologists would say dependent variable. A trial normally names one primary outcome and several secondary ones.
Parallel-forms reliability
Also called: parallel forms reliability, alternate-forms reliability, alternate forms reliability, equivalent-forms reliability, equivalent forms reliability, coefficient of equivalence
Parallel-forms reliability is the consistency of scores across two versions of a test built to measure the same construct in the same way, estimated by giving both forms to the same people and correlating the results. It suits situations where a test must be repeated without people meeting the same items twice, but writing two truly equivalent forms takes many items.
Partial verification bias
Also called: partial reference bias, partial verification
Partial verification bias arises in a diagnostic accuracy study when only some participants, often those who tested positive, go on to receive the reference standard, so true status is unknown for the rest. It is common because reference tests such as biopsies and angiography are invasive, costly or risky, and leaving out the unverified distorts the accuracy estimates.
Percentage agreement
Also called: percent agreement, per cent agreement, proportion agreement, simple agreement, observed agreement
Percentage agreement is the share of cases on which two or more raters give the same rating, the simplest measure of inter-rater reliability. It is easy to understand but ignores the agreement raters would reach by chance alone, which can make reliability look better than it is, so a kappa statistic is usually reported beside it or instead of it.
Performance bias
Performance bias is a systematic difference between the groups in a trial in the care or attention they receive, other than the intervention being tested, usually because participants or staff know who is getting what. Control participants might seek other treatment, for example. Blinding participants and personnel prevents it, and RoB 2 calls it bias due to deviations from intended interventions.
Placebo effect
Also called: placebo effects
The placebo effect is an improvement in symptoms or outcomes that follows a treatment with no active ingredient, or the part of any treatment's benefit that comes from expectation and the context of care rather than its specific action. Trials compare an intervention with a placebo so that this effect is present in both groups and can be separated from the treatment's own effect.
Positive predictive value
Also called: PPV, positive predictive values
Positive predictive value is the proportion of people with a positive test result who truly have the condition, answering the question a person with a positive result asks. It depends heavily on how common the condition is in the group tested, so a test with good sensitivity and specificity can still have a low positive predictive value when the condition is rare.
Precision (measurement)
Also called: measurement precision, precision of measurement
Precision, in measurement, is how closely repeated measurements of the same thing under specified conditions agree with one another, whatever the true value. It is expressed through measures of spread such as the standard deviation and reflects random error, whereas accuracy also depends on systematic error. Repeatability and reproducibility are precision under narrower and wider conditions.
Predictive validity
Predictive validity is the kind of criterion validity shown when scores on a measure predict a criterion measured later, such as an admissions test predicting first-year grades. It differs from concurrent validity only in timing. It is judged by the correlation between the measure and the later criterion, which unreliability in either can weaken.
Predictor variable
Also called: predictor, predictors, predictor variables, regressor
A predictor variable is a variable used to predict the value of another, the criterion or outcome variable, most often in regression. Unlike an independent variable in an experiment, a predictor need not be manipulated or even be a cause: a school test score can predict later grades without causing them. Baron and Kenny pair predictor with criterion in this way.
Primary outcome
Also called: primary outcome measure, primary endpoint, primary end point, primary outcomes
A primary outcome is the outcome a study names in advance as the most important measure of whether an intervention works, and usually the one its sample size is calculated for. CONSORT asks trials to define it fully, including how and when it is measured. Changing it after the data are in is a form of outcome reporting bias.
Psychological realism
Psychological realism is the extent to which an experiment engages the same mental processes that operate in the real world, even if its setting looks artificial. A laboratory task can have psychological realism without mundane realism, and Jhangiani and colleagues present it as a reason why artificial experiments can still generalise.
Psychometric properties
Also called: measurement properties, psychometric qualities
Psychometric properties are the qualities of a measurement instrument that show how well it measures, chiefly its reliability, validity and responsiveness. COSMIN, an international initiative on health measurement, set out consensus definitions of nine such properties, and a validation study reports evidence on them for a particular population and purpose.
Psychometrics
Also called: psychometric, psychometric theory
Psychometrics is the field concerned with the theory and methods of measuring psychological and educational attributes, such as ability, personality and attitudes, including how tests are built, scored and evaluated. Its main models are classical test theory, item response theory and factor analysis, and its core questions are about reliability and validity.
Quantitative variable
Also called: quantitative data, numerical variable, numeric variable, numerical data
A quantitative variable is a variable whose values are numbers obtained by counting or measuring, such as weight, pulse rate or number of siblings, so arithmetic on them is meaningful. It can be discrete, taking only certain values, or continuous, taking any value within a range, and it is measured at the interval or ratio level.
Random error
Also called: random measurement error, random error of measurement, noise
Random error is the component of measurement error that varies unpredictably from one measurement to the next, pushing some readings up and others down. It adds variability, or noise, but tends to cancel out across many measurements, so it does not shift the average. It lowers reliability and can weaken the relationships a study observes.
Ratio scale
Also called: ratio level of measurement, ratio level, ratio data, ratio variable, ratio measurement
A ratio scale is the level of measurement with equal intervals and a true zero that means none of the quantity, so ratios between values are meaningful: 20 kilograms is twice 10 kilograms. Height, weight, time, money and counts are examples. Everything that can be done with interval data can be done with ratio data, and ratios too.
Reactivity
Also called: measurement reactivity, reactivity of measurement, reactive effect, reactive measurement, assessment reactivity
Reactivity is the change in people's thoughts, feelings or behaviour that is caused by the act of measuring them, as when filling in a questionnaire about drinking leads someone to drink less. Several studies have found such effects, though it is not yet possible to predict when they will occur. The Hawthorne effect and the testing threat are related ideas.
Recall bias
Recall bias is systematic error caused by differences in how accurately or completely participants remember and report past events or experiences. It threatens case-control and retrospective studies most, because people with a disease may search their memories harder: parents of children with cancer may recall earlier infections more fully than other parents.
Reference standard
Also called: reference test, reference standard test
A reference standard is the best available method for establishing whether a target condition is present or absent, against which an index test's accuracy is measured in a diagnostic accuracy study. STARD prefers the term to gold standard because a reference standard can itself make mistakes, and problems in applying it, such as partial verification, distort the index test's apparent accuracy.
Regression dilution bias
Also called: regression dilution, attenuation, attenuation bias, attenuation due to unreliability
Regression dilution bias is the weakening of an estimated relationship that occurs when a predictor is measured with random error, so a regression slope or correlation comes out smaller than it would with the true values. Psychometricians call it attenuation. It holds for a single predictor, but when several variables, including confounders, carry error, the bias can run either way.
Regression to the mean
Also called: regression toward the mean, regression towards the mean, statistical regression, regression artefact, regression artifact, regression threat, RTM
Regression to the mean is the tendency for people who score unusually high or low on one measurement to score closer to the average when measured again, because part of any extreme score is chance. It threatens internal validity when participants are chosen for extreme scores, since their later improvement can be mistaken for a treatment effect.
Reification
Also called: reify, reified
Reification is the error of treating an abstract concept, such as intelligence or social class, as if it were a concrete thing that exists independently of how researchers have defined it. DeCarlo warns that definitions are agreements people make, not discoveries, so a measured score should not be mistaken for the concept itself.
Reliability
Also called: measurement reliability, reliability of measurement, reliable
Reliability is the consistency of a measure, the extent to which it gives the same results when nothing being measured has changed, across time, across items and across raters. In classical test theory it is the share of score variance due to true differences between people. A measure can be reliable without being valid, but it cannot be valid without being reliable.
Repeatability
Also called: measurement repeatability
Repeatability is the precision of measurements made under the same conditions, meaning the same procedure, operator, instrument and place, on the same or similar things over a short period. It is contrasted with reproducibility, which allows those conditions to change. In clinical research it overlaps with test-retest and intra-rater reliability.
Reporting bias
Also called: reporting biases
Reporting bias is the selective disclosure or withholding of information, either by researchers deciding which findings to publish or report, as in publication and outcome reporting bias, or by participants deciding what to tell researchers, such as under-reporting smoking. The Catalogue of Bias treats the first sense as a family of reporting biases and the second as information bias.
Reproducibility (measurement)
Also called: measurement reproducibility, reproducibility of measurement
Reproducibility, in measurement, is the precision of measurements of the same thing made under changed conditions, such as different operators, instruments, laboratories or times. It usually shows more variation than repeatability, which holds the conditions fixed. It is a different idea from the reproducibility of a study's findings, which the research design glossary covers.
Resentful demoralisation
Also called: resentful demoralization
Resentful demoralisation is a threat to internal validity in which a comparison group, learning that others are receiving something it is not, becomes discouraged and gives up, so its outcomes worsen. Unlike compensatory rivalry, which narrows differences, it widens them and can make an intervention look more effective than it is.
Residual confounding
Also called: residual confounders
Residual confounding is the confounding that remains after a study has adjusted for confounders, because some were measured inaccurately or not measured at all. It is why observational studies can rarely rule out confounding completely, and it explains why healthy user bias can persist even in carefully adjusted analyses.
Response bias
Also called: response biases, survey response bias
Response bias is any systematic tendency of participants to answer questions inaccurately or in a way shaped by something other than the content asked, such as a wish to please the researcher or to agree with every statement. Social desirability bias and acquiescence bias are its best-known forms, and leading question wording can provoke it.
Response variable
Also called: response, responses, response variables
A response variable is the variable whose changes a study measures in answer to changes in the explanatory variable, the outcome in an experiment or statistical model. OpenStax and the NIST handbook of statistical methods use the term, and NIST notes that responses are sometimes called dependent variables. In designed experiments the response is the output measured under each setting of the factors.
Responsiveness
Also called: sensitivity to change, responsiveness to change
Responsiveness is the ability of an instrument to detect change over time in the construct it measures, such as improvement after treatment. COSMIN treats it as a measurement property in its own right, alongside reliability and validity. A measure with a strong floor or ceiling effect has little room to show change at that end of the scale.
ROC curve
Also called: receiver operating characteristic curve, receiver operating characteristic, ROC, ROC plot, ROC analysis, receiver operator characteristic curve
A ROC curve, or receiver operating characteristic curve, is a graph of a test's sensitivity against one minus its specificity across every possible cut-off, showing the trade-off between catching true cases and raising false alarms. A curve bowing towards the top left corner means good discrimination, and a diagonal line means none. It is used to compare tests and choose cut-offs.
Scale (multi-item measure)
Also called: multi-item scale, summated scale, composite scale
A scale, as a measure, is a set of items that together measure one construct and are combined into a single score, such as a ten-item self-esteem scale. DeCarlo reserves the word for composite measures whose items differ in intensity, setting it apart from an index. The same word also names a variable's level of measurement, as in interval scale.
Secondary outcome
Also called: secondary outcome measure, secondary endpoint, secondary end point, secondary outcomes
A secondary outcome is an additional outcome a study measures besides its primary outcome, often to explore other benefits or to look for harms and unintended effects. Findings on secondary outcomes are usually read as supporting or exploratory, because the study was not sized for them and analysing many outcomes raises the problems of multiplicity.
Selection (threat to validity)
Also called: selection threat, selection effect, differential selection
Selection is a threat to internal validity in which the groups being compared differ before the study begins, so any difference at the end may reflect who was in each group rather than the intervention. It is the main weakness of non-equivalent groups designs, and random assignment is the standard remedy. It is one form of selection bias.
Selection interaction
Also called: selection interactions, interaction with selection
A selection interaction is a threat to internal validity in which another threat, such as history, maturation, testing, instrumentation, mortality or regression, affects the groups being compared differently because the groups differ. In a selection-maturation interaction, for instance, one group is simply developing faster than the other, which can look like a treatment effect.
Sensitivity (of a test)
Also called: sensitivity, diagnostic sensitivity, clinical sensitivity, true positive rate, TPR
Sensitivity is the proportion of people who truly have a condition whom a test correctly identifies as positive, also called the true positive rate. A highly sensitive test misses few cases. Sensitivity is not a fixed property of a test: it changes with the cut-off chosen and with the kind of people tested, and it trades off against specificity.
Spearman-Brown prophecy formula
Also called: Spearman-Brown formula, Spearman-Brown prediction formula, Spearman-Brown prophecy, Spearman Brown formula, Spearman-Brown correction
The Spearman-Brown prophecy formula predicts how a test's reliability will change if the test is lengthened or shortened with comparable items. Its best-known use is to step a split-half correlation, which describes a half-length test, up to an estimate for the whole test. Spearman and Brown published it independently in 1910.
Specificity (of a test)
Also called: specificity, diagnostic specificity, clinical specificity, true negative rate, TNR
Specificity is the proportion of people who truly do not have a condition whom a test correctly identifies as negative, also called the true negative rate. A highly specific test gives few false positives. Like sensitivity, it changes with the cut-off chosen and the population tested, and moving the cut-off to raise one usually lowers the other.
Spectrum bias
Spectrum bias occurs when a diagnostic test is studied in a range of people different from those it is meant for, such as very sick patients and healthy volunteers rather than the uncertain cases seen in practice. Because sensitivity and specificity vary with the mix of patients, which is called the spectrum effect, accuracy found in such a study can mislead.
Split-half reliability
Also called: split-half correlation, split half reliability, split-half method, split-half coefficient
Split-half reliability is an estimate of internal consistency made by dividing a test's items into two halves, often odd and even items, scoring each half and correlating the two scores. Because each half is only half the length of the test, the correlation is stepped up with the Spearman-Brown formula. Different splits give different values, which coefficient alpha resolves by averaging them.
Spontaneous remission
Spontaneous remission is the tendency of many medical and psychological problems to improve over time without any treatment. It threatens internal validity in one-group studies of treatments for conditions such as depression, because improvement that would have happened anyway can be credited to the intervention. A control group followed over the same period reveals it.
Spurious relationship
Also called: spurious association, spurious correlation, spuriousness
A spurious relationship is an association between two variables that looks causal but is really explained by a third variable affecting both, as when ice cream sales and drownings rise together because of summer weather. Showing that a relationship is not spurious is one of the standard criteria for claiming cause and effect.
Standard error of measurement
Also called: SEM
The standard error of measurement is the typical amount by which a person's observed score differs from their true score because of random error, calculated from the spread of scores and the test's reliability. It is used to put a confidence band around an individual's score, and the more reliable the test, the smaller it is.
Standardised instrument
Also called: standardized instrument, standardised test, standardized test, standardised measure, standardized measure, standardised questionnaire, standardized questionnaire, standardised assessment, standardized assessment
A standardised instrument is a test, questionnaire or observation schedule that is given, scored and interpreted in the same way every time, with fixed items, instructions and scoring rules, and often with published norms and evidence of reliability and validity. Jhangiani and colleagues advise choosing an existing measure over writing a new one unless there is a clear reason to do otherwise.
Statistical conclusion validity
Also called: statistical validity, conclusion validity
Statistical conclusion validity is the extent to which conclusions about whether, and how strongly, variables are related are justified by the data and the statistics used, including whether a test's assumptions are met. Its failures are concluding there is a relationship when there is none, or missing one that exists. The Research Methods Knowledge Base now calls it conclusion validity and applies it to qualitative analysis too.
Structural validity
Also called: factorial validity
Structural validity is the degree to which a measure's scores reflect the dimensional structure of the construct it measures, for example whether a questionnaire meant to have three subscales really does. It is usually tested with factor analysis. COSMIN treats it as part of construct validity, relevant only to measures whose items all reflect an underlying construct.
Surrogate endpoint
Also called: surrogate outcome, surrogate end point, surrogate marker, surrogate outcome measure
A surrogate endpoint is a laboratory measure or physical sign used in place of an outcome that directly matters to patients, because it is expected to predict that outcome, as eye pressure stands in for loss of vision in glaucoma. It does not measure the benefit itself, so its value rests entirely on how well it predicts the outcome that matters.
Systematic error
Also called: systematic measurement error, measurement bias, bias in measurement, non-random error
Systematic error is the component of measurement error that stays constant or varies predictably across repeated measurements, so readings are consistently pushed in one direction, as with a scale that always reads high. Unlike random error it does not average out, and it shifts the mean. A known systematic error can be corrected, and an estimate of it is called measurement bias.
Target condition
A target condition is the disease or condition that an index test is expected to detect in a diagnostic accuracy study, defined so that the reference standard can say whether each participant has it. Defining it precisely, including its stage or severity, matters because a test's sensitivity is often higher when more participants have advanced disease.
Test-retest reliability
Also called: test retest reliability, retest reliability, test-retest, coefficient of stability, stability reliability
Test-retest reliability is the consistency of scores when the same measure is given to the same people on two occasions, for a construct expected to stay stable between them. It is estimated by correlating the two sets of scores or with an intraclass correlation coefficient. The interval matters: too short and people remember their answers, too long and real change creeps in.
Testing (threat to validity)
Also called: testing threat, testing effect, pretest effect, pre-test effect
Testing is a threat to internal validity in which taking a pre-test changes how people score on the post-test, through practice, familiarity or thinking prompted by the questions, so the change is wrongly credited to the intervention. It is related to the practice effect in within-subjects designs.
Threats to internal validity
Also called: threat to internal validity, internal validity threats
Threats to internal validity are alternative explanations, other than the intervention, for a change or difference observed in a study. The classic list includes history, maturation, testing, instrumentation, regression to the mean, selection, experimental mortality and social threats such as diffusion of treatment, and a design is judged by how many of these it rules out.
Translation validity
Translation validity is the Research Methods Knowledge Base's umbrella term for face and content validity, the kinds of evidence that ask whether an operationalisation is a good translation of the construct's definition. It is contrasted there with criterion-related validity, which checks how the measure performs against other measures or outcomes.
True negative
Also called: true negatives, TN
A true negative is a test result that correctly says a condition or feature is absent in someone who does not have it. Specificity is the number of true negatives divided by everyone without the condition, and negative predictive value is the number of true negatives divided by everyone who tested negative.
True positive
Also called: true positives, TP
A true positive is a test result that correctly says a condition or feature is present in someone who has it. True positives form one cell of the two-by-two table that compares a test with a reference standard, and sensitivity is the number of true positives divided by everyone who has the condition.
True score
Also called: true scores
A true score, in classical test theory, is the score a person would obtain if a measure had no random error, defined as the average of their scores over an infinite number of administrations. It cannot be observed. It concerns consistency, not accuracy, so a biased instrument has a biased true score. Observed score equals true score plus error.
Trueness
Also called: measurement trueness, trueness of measurement
Trueness is how close the average of a very large number of repeated measurements comes to a reference value. It reflects systematic error only, not random scatter, so a biased instrument lacks trueness however consistent its readings are. The international vocabulary of metrology warns against calling it accuracy, which depends on precision as well.
Typology
Also called: typologies
A typology is a measure that sorts cases into types by combining their positions on two or more variables according to clear rules, such as classifying political views by attitudes to both economic and social questions. It produces categories rather than scores, unlike an index or scale. DeCarlo lists it with indices and scales as ways of measuring complex concepts.
Unidimensionality
Also called: unidimensional, unidimensional scale
Unidimensionality is the property of a set of items that all measure a single underlying construct, which is what justifies adding them into one total score. It is checked with factor analysis, not with Cronbach's alpha: a high alpha can occur in a scale that measures several related things, especially a long one.
Validity
Also called: measurement validity, validity of measurement, test validity, valid
Validity is the extent to which the scores from a measure represent what they are intended to, or more broadly the extent to which a study's conclusions are sound. Current testing standards treat it as a property of the interpretation and use of scores, supported by accumulated evidence, rather than a fixed property of a test, so a measure is valid for particular uses.
Variable
Also called: variables, research variable, study variable
A variable is any characteristic that can take different values across the people, objects or events in a study, such as age, blood type, income or a test score. Each possible value is called an attribute or category. Variables are described by their role in a study, such as independent or dependent, and by their level of measurement.
Verification bias
Also called: work-up bias, workup bias
Verification bias is a distortion in a diagnostic accuracy study that arises when not everyone given the index test has their true status confirmed by the same reference standard, often because the reference test is invasive or costly. Its two forms are partial verification and differential verification, and both can shift accuracy estimates in directions that are hard to predict.
Youden index
Also called: Youden's index, Youden's J, Youden J statistic, J statistic
The Youden index is a single summary of a test at a given cut-off, calculated as sensitivity plus specificity minus one, so it is zero for a test no better than chance and one for a perfect test. The cut-off with the highest Youden index is often chosen as the best, which treats false positives and false negatives as equally costly.
Where these definitions were checked
Rajiv Jhangiani, I-Chant Chiang, Carrie Cuttler and Dana Leighton (open textbook), Research Methods in Psychology, 4.2: Understanding Psychological Measurement
Rajiv Jhangiani, I-Chant Chiang, Carrie Cuttler and Dana Leighton (open textbook), Research Methods in Psychology, 4.3: Reliability and Validity of Measurement
Rajiv Jhangiani, I-Chant Chiang, Carrie Cuttler and Dana Leighton (open textbook), Research Methods in Psychology, 4.4: Practical Strategies for Psychological Measurement
Rajiv Jhangiani, I-Chant Chiang, Carrie Cuttler and Dana Leighton (open textbook), Research Methods in Psychology, 5.2: Experiment Basics
Rajiv Jhangiani, I-Chant Chiang, Carrie Cuttler and Dana Leighton (open textbook), Research Methods in Psychology, 5.4: Experimentation and Validity
Rajiv Jhangiani, I-Chant Chiang, Carrie Cuttler and Dana Leighton (open textbook), Research Methods in Psychology, 5.5: Practical Considerations
Rajiv Jhangiani, I-Chant Chiang, Carrie Cuttler and Dana Leighton (open textbook), Research Methods in Psychology, 8.2: One-Group Designs
Barbara Illowsky and Susan Dean, OpenStax, Introductory Statistics 2e, 1.2: Data, Sampling, and Variation in Data and Sampling
Barbara Illowsky and Susan Dean, OpenStax, Introductory Statistics 2e, 1.3: Frequency, Frequency Tables, and Levels of Measurement
Barbara Illowsky and Susan Dean, OpenStax, Introductory Statistics 2e, 1.4: Experimental Design and Ethics
Matthew DeCarlo (open textbook), Scientific Inquiry in Social Work, 7.2: Causal relationships
Matthew DeCarlo (open textbook), Scientific Inquiry in Social Work, 9.2: Conceptualization
Matthew DeCarlo (open textbook), Scientific Inquiry in Social Work, 9.3: Operationalization
Matthew DeCarlo (open textbook), Scientific Inquiry in Social Work, 9.4: Measurement quality
Matthew DeCarlo (open textbook), Scientific Inquiry in Social Work, 9.5: Complexities in quantitative measurement
William M. K. Trochim, Research Methods Knowledge Base (Conjointly), Research Methods Knowledge Base: Levels of Measurement
William M. K. Trochim, Research Methods Knowledge Base (Conjointly), Research Methods Knowledge Base: True Score Theory
William M. K. Trochim, Research Methods Knowledge Base (Conjointly), Research Methods Knowledge Base: Measurement Error
William M. K. Trochim, Research Methods Knowledge Base (Conjointly), Research Methods Knowledge Base: Types of Reliability
William M. K. Trochim, Research Methods Knowledge Base (Conjointly), Research Methods Knowledge Base: Types of Measurement Validity
William M. K. Trochim, Research Methods Knowledge Base (Conjointly), Research Methods Knowledge Base: The Nomological Network
William M. K. Trochim, Research Methods Knowledge Base (Conjointly), Research Methods Knowledge Base: Multitrait-Multimethod Matrix
William M. K. Trochim, Research Methods Knowledge Base (Conjointly), Research Methods Knowledge Base: Threats to Construct Validity
William M. K. Trochim, Research Methods Knowledge Base (Conjointly), Research Methods Knowledge Base: Conclusion Validity
William M. K. Trochim, Research Methods Knowledge Base (Conjointly), Research Methods Knowledge Base: External Validity
William M. K. Trochim, Research Methods Knowledge Base (Conjointly), Research Methods Knowledge Base: Single Group Threats
William M. K. Trochim, Research Methods Knowledge Base (Conjointly), Research Methods Knowledge Base: Multiple Group Threats
William M. K. Trochim, Research Methods Knowledge Base (Conjointly), Research Methods Knowledge Base: Social Interaction Threats
Tony Albano, Introduction to Educational and Psychological Measurement Using R (open textbook), Measurement, Scales, and Scoring (chapter 2)
Tony Albano, Introduction to Educational and Psychological Measurement Using R (open textbook), Reliability (chapter 5)
Tony Albano, Introduction to Educational and Psychological Measurement Using R (open textbook), Item Analysis (chapter 6)
Tony Albano, Introduction to Educational and Psychological Measurement Using R (open textbook), Item Response Theory (chapter 7)
Tony Albano, Introduction to Educational and Psychological Measurement Using R (open textbook), Validity (chapter 9)
Lidwine B. Mokkink, Cecilia A. C. Prinsen, Donald L. Patrick, Jordi Alonso, Lex M. Bouter, Henrica C. W. de Vet and Caroline B. Terwee, COSMIN, COSMIN methodology for systematic reviews of Patient-Reported Outcome Measures (PROMs): user manual, version 1.0 (Table 1, COSMIN definitions of measurement properties)
COSMIN initiative, COSMIN Taxonomy of Measurement Properties
Joint Committee for Guides in Metrology, International Vocabulary of Metrology (VIM3), JCGM 200:2012, Measurement accuracy, trueness, precision, error, systematic and random error, measurement bias, repeatability, reproducibility and calibration (entries 2.13 to 2.25 and 2.39)
NIST/SEMATECH e-Handbook of Statistical Methods, National Institute of Standards and Technology, 5.7. A Glossary of DOE Terminology
Penn State STAT 502, adapted on Statistics LibreTexts, Analysis of Variance and Design of Experiments, 9: ANCOVA Part I
Julian P. T. Higgins, Douglas G. Altman, Peter C. Gøtzsche and others, The Cochrane Collaboration's tool for assessing risk of bias in randomised trials (BMJ, 2011)
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Ascertainment bias
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Apprehension bias
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Collider bias
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Confirmation bias
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Confounding
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Confounding by indication
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Detection bias
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Differential reference bias
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: End-digit preference bias
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Hawthorne effect
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Healthy user bias
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Immortal time bias
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Incorporation bias
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Information bias
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Insensitive measure bias
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Lead time bias
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Length time bias
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Misclassification bias
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Novelty bias
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Observer bias
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Partial reference bias
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Performance bias
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Recall bias
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Reporting biases
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Spectrum bias
Catalogue of Bias Collaboration, Centre for Evidence-Based Medicine, University of Oxford, Catalogue of Bias: Verification bias
Health Knowledge (public health textbook), UK, Biases and Confounding
Health Knowledge (public health textbook), UK, Confounding in epidemiological studies
Reuben M. Baron and David A. Kenny, The Moderator-Mediator Variable Distinction in Social Psychological Research (Journal of Personality and Social Psychology, 1986)
Columbia University Mailman School of Public Health, Item Response Theory
Columbia University Mailman School of Public Health, Differential Item Functioning
Columbia University Mailman School of Public Health, Exploratory Factor Analysis
Sean T. H. Lee, Association for Psychological Science, Testing for Measurement Invariance: Does your measure mean the same thing for different participants? (APS Observer, 2018)
Mohsen Tavakol and Reg Dennick, Making sense of Cronbach's alpha (International Journal of Medical Education, 2011)
Italo Trizano-Hermosilla and Jesús M. Alvarado, Best Alternatives to Cronbach's Alpha Reliability in Realistic Conditions: Congeneric and Asymmetrical Measurements (Frontiers in Psychology, 2016)
Klaas Sijtsma, Psychometric Society, Invited discussion of Cronbach (1951), Coefficient alpha and the internal structure of tests
Mary L. McHugh, Interrater reliability: the kappa statistic (Biochemia Medica, 2012)
Terry K. Koo and Mae Y. Li, A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research (Journal of Chiropractic Medicine, 2016)
Davide Giavarina, Understanding Bland Altman analysis (Biochemia Medica, 2015)
Timo B. Brakenhoff, Maarten van Smeden, Frank L. J. Visseren and Rolf H. H. Groenwold, Random measurement error: Why worry? An example of cardiovascular risk factors (PLOS ONE, 2018)
Christopher R. Lim, Kristina Harris, Jill Dawson, David J. Beard, Ray Fitzpatrick and Andrew J. Price, Floor and ceiling effects in the OHS: an analysis of the NHS PROMs data set (BMJ Open, 2015)
Roman Jaeschke, Joel Singer and Gordon H. Guyatt, Measurement of health status: ascertaining the minimal clinically important difference (Controlled Clinical Trials, 1989), abstract
John Hunsley and Gregory J. Meyer, The incremental validity of psychological testing and assessment (Psychological Assessment, 2003), abstract
Philip M. Podsakoff, Scott B. MacKenzie, Jeong-Yeon Lee and Nathan P. Podsakoff, Common method biases in behavioral research: a critical review of the literature and recommended remedies (Journal of Applied Psychology, 2003), abstract
Gijs A. Holleman, Ignace T. C. Hooge, Chantal Kemner and Roy S. Hessels, The 'Real-World Approach' and Its Problems: A Critique of the Term Ecological Validity (Frontiers in Psychology, 2020)
Jim McCambridge, John Witton and Diana R. Elbourne, Systematic review of the Hawthorne effect: New concepts are needed to study research participation effects (Journal of Clinical Epidemiology, 2014)
Jim McCambridge, Marijn de Bruin and John Witton, The Effects of Demand Characteristics on Research Participant Behaviours in Non-Laboratory Settings: A Systematic Review (PLoS ONE, 2012)
David P. French and Stephen Sutton, Reactivity of measurement in health psychology: How much of a problem is it? What can be done about it? (British Journal of Health Psychology, 2010), abstract
Karolina Wartolowska, The nocebo effect as a source of bias in the assessment of treatment effects (F1000Research, 2019)
Chris Westbury and Daniel King, A Constant Error, Revisited: A New Explanation of the Halo Effect (Cognitive Science, 2024)
David McKenzie, World Bank, Are John Henry Effects as Apocryphal as their Eponym? (Development Impact blog, 2013)
David Moher, Sally Hopewell, Kenneth F. Schulz and others, CONSORT 2010 Explanation and Elaboration: updated guidelines for reporting parallel group randomised trials (BMJ, 2010)
National Center for Advancing Translational Sciences, US National Institutes of Health, citing the US Food and Drug Administration, Surrogate endpoint (Patient-Focused Drug Development Glossary)
Jérémie F. Cohen, Daniël A. Korevaar, Douglas G. Altman and others, STARD 2015 guidelines for reporting diagnostic accuracy studies: explanation and elaboration (BMJ Open, 2016)
Robert Trevethan, Sensitivity, Specificity, and Predictive Values: Foundations, Pliabilities, and Pitfalls in Research and Practice (Frontiers in Public Health, 2017)
Gilat Grunau and Shai Linn, Commentary: Sensitivity, Specificity, and Predictive Values: Foundations, Pliabilities, and Pitfalls in Research and Practice (Frontiers in Public Health, 2018)
Jonathan J. Deeks and Douglas G. Altman, Diagnostic tests 4: likelihood ratios (BMJ, 2004)
Karimollah Hajian-Tilaki, Receiver Operating Characteristic (ROC) Curve Analysis for Medical Diagnostic Test Evaluation (Caspian Journal of Internal Medicine, 2013)