Statistics and data analysis: a glossary of terms
The words for describing data and drawing conclusions from them: averages and spread, distributions, significance tests and their misreadings, effect sizes, the tests and models most often taught, Bayesian ideas and missing data. Every definition was checked against the sources below.
- Absolute risk
- Adjusted R squared
- Akaike information criterion
- Analysis of covariance
- Analysis of variance
- ARIMA model
- Autocorrelation
- Bayes factor
- Bayes' theorem
- Bayesian information criterion
- Bayesian statistics
- Benjamini-Hochberg procedure
- Bimodal distribution
- Binomial distribution
- Bonferroni correction
- Bootstrapping
- Box plot
- Censoring
- Central limit theorem
- Central tendency
- Chi-square distribution
- Chi-square goodness-of-fit test
- Chi-square test of independence
- Cluster analysis
- Cochran's Q test
- Coefficient of variation
- Cohen's d
- Conditional probability
- Confidence interval
- Confirmatory factor analysis
- Contingency table
- Cook's distance
- Correlation
- Covariance
- Cox proportional hazards model
- Cramér's V
- Credible interval
- Critical value
- Cumulative frequency
- Data coding
- Data transformation
- Degrees of freedom
- Descriptive statistics
- Direct effect
- Discriminant analysis
- Dummy coding
- Dunn's test
- Dunnett's test
- Effect size
- Eigenvalue
- Equivalence test
- Estimation
- Estimator
- Eta squared
- Expected value
- Exploratory data analysis
- Exploratory factor analysis
- Extrapolation
- F distribution
- F-ratio
- Factor analysis
- Factor loading
- Factor rotation
- Factorial ANOVA
- False discovery rate
- Familywise error rate
- Fisher's exact test
- Fisher's least significant difference
- Fit indices
- Five-number summary
- Fixed and random effects
- Frequency
- Frequency distribution
- Frequentist statistics
- Friedman test
- Full information maximum likelihood
- Games-Howell test
- Garden of forking paths
- General linear model
- Generalised estimating equations
- Generalised linear model
- Geometric mean
- Glass's delta
- HARKing
- Harmonic mean
- Hazard
- Hazard ratio
- Hedges' g
- Heteroscedasticity
- Hierarchical cluster analysis
- Hierarchical regression
- Histogram
- Hochberg procedure
- Holm-Bonferroni method
- Homogeneity of variance
- Hypothesis testing
- Imputation
- Independence of observations
- Independent events
- Independent-samples t-test
- Indirect effect
- Inferential statistics
- Influential observation
- Instrumental variable analysis
- Interaction effect
- Intercept
- Interquartile range
- Interval estimate
- K-means clustering
- Kaplan-Meier estimator
- Kendall's tau
- Kolmogorov-Smirnov test
- Kruskal-Wallis test
- Kurtosis
- Last observation carried forward
- Latent class analysis
- Law of large numbers
- Levene's test
- Leverage
- Likelihood
- Likelihood ratio test
- Listwise deletion
- Little's MCAR test
- Log-rank test
- Logistic regression
- Main effect
- Mann-Whitney U test
- Markov chain Monte Carlo
- Maximum likelihood estimation
- McNemar's test
- Mean
- Mean difference
- Mean imputation
- Median
- Mediation analysis
- Missing at random
- Missing completely at random
- Missing data
- Missing not at random
- Mixed ANOVA
- Mode
- Moderation analysis
- Multicollinearity
- Multilevel model
- Multinomial logistic regression
- Multiple comparisons problem
- Multiple imputation
- Multiple linear regression
- Multivariate analysis of variance
- Multiverse analysis
- Negative binomial regression
- Non-parametric test
- Normal distribution
- Normality assumption
- Number needed to harm
- Number needed to treat
- Odds
- Odds ratio
- Omega squared
- Omnibus test
- One-sample t-test
- One-tailed test
- One-way ANOVA
- Optional stopping
- Ordinal logistic regression
- Ordinary least squares
- Outlier
- Overdispersion
- Overfitting
- P-hacking
- P-value
- Paired t-test
- Pairwise deletion
- Parameter
- Parametric test
- Partial correlation
- Partial eta squared
- Path analysis
- Pearson correlation coefficient
- Percentile
- Permutation test
- Phi coefficient
- Planned comparison
- Point estimate
- Point-biserial correlation
- Poisson distribution
- Poisson regression
- Polynomial regression
- Post hoc test
- Posterior distribution
- Practical significance
- Principal component analysis
- Prior distribution
- Probability
- Probability distribution
- Propensity score
- Proportional hazards assumption
- Q-Q plot
- Quartile
- R squared
- Random variable
- Range
- Regression analysis
- Regression coefficient
- Relative frequency
- Relative risk reduction
- Repeated-measures ANOVA
- Researcher degrees of freedom
- Residual
- Risk difference
- Risk ratio
- Robust statistics
- Sampling distribution
- Scatter plot
- Scheffé test
- Scree plot
- Seasonality
- Segmented regression
- Sensitivity analysis
- Shapiro-Wilk test
- Sign test
- Significance level
- Simple effect
- Simple linear regression
- Skewness
- Sobel test
- Spearman's rank correlation
- Sphericity
- Standard deviation
- Standard error
- Standard normal distribution
- Standardised mean difference
- Standardised regression coefficient
- Stationarity
- Statistic
- Statistical model
- Statistical power
- Statistical significance
- Statistical software
- Stem-and-leaf plot
- Stepwise regression
- Structural equation modelling
- Sum of squares
- Survival analysis
- Symmetrical distribution
- t distribution
- t-test
- Test statistic
- Time series analysis
- Trimmed mean
- Tukey's HSD
- Two-tailed test
- Two-way ANOVA
- Type I error
- Type II error
- Uniform distribution
- Variability
- Variance
- Variance inflation factor
- Weighted mean
- Welch's ANOVA
- Welch's t-test
- Wilcoxon signed-rank test
- Yates's continuity correction
- Z-score
- Z-test
Absolute risk
Also called: risk
Absolute risk is the probability that an event happens to people in a given group over a stated period, estimated as the number with the event divided by the number at risk. Knowing the absolute risk in the comparison group is what turns a relative effect, such as halving the risk, into a meaningful number.
Adjusted R squared
Also called: adjusted R², adjusted R-squared, adjusted R2, R-squared adjusted
Adjusted R squared is R squared corrected for the number of predictors, so it falls when a new predictor adds less than chance would. It is used to compare models with different numbers of predictors, where the plain R squared always favours the larger model.
Akaike information criterion
Also called: AIC, Akaike's information criterion
The Akaike information criterion is a measure for comparing statistical models fitted to the same data that rewards good fit and penalises the number of parameters, lower values being better. It favours models that predict well, and it tends to choose larger models than the Bayesian information criterion.
Analysis of covariance
Also called: ANCOVA
Analysis of covariance is analysis of variance with one or more continuous covariates added, such as a baseline score, so that group means are compared after adjusting for them. It can remove noise and so add power. It assumes the covariate relates to the outcome in the same way in every group, known as homogeneity of regression slopes.
Analysis of variance
Also called: ANOVA, AOV
Analysis of variance is a family of methods that compare the means of several groups at once by splitting the total variation into parts due to the factors studied and a remainder due to error. The F-ratio compares those parts. It assumes normality, equal variances and independent observations, and it avoids the inflated false positives of running many t-tests.
ARIMA model
Also called: ARIMA, Box-Jenkins model, autoregressive integrated moving average model
An ARIMA model, short for autoregressive integrated moving average, is a model for a time series that predicts each value from past values, past random shocks and, where needed, differencing to remove a trend. Seasonal versions add terms for patterns that repeat each year, and these models are widely used for forecasting.
Autocorrelation
Also called: serial correlation, autocorrelation function, ACF
Autocorrelation is the correlation of a series with itself at an earlier time, so that each value tends to resemble the values just before it. It is common in time series and repeated measurements, breaks the independence assumption of ordinary regression, and makes standard errors too small if it is ignored.
Bayes factor
Also called: BF, BF10, BF01
A Bayes factor is the ratio of how well two competing hypotheses, such as an effect and no effect, predicted the observed data. A BF10 of 5 means the data were five times more likely under the alternative than under the null, and values below 1 favour the null, which a p-value can never do. Labels such as moderate or strong for ranges of values are conventions.
Bayes' theorem
Also called: Bayes theorem, Bayes' rule, Bayes rule, Bayes's theorem
Bayes' theorem is the rule for updating a probability in the light of new evidence, combining the prior probability of an event with how likely the evidence would be if it were true to give a posterior probability. It explains why a positive result on an accurate test for a rare condition can still leave the condition unlikely.
Bayesian information criterion
Also called: BIC, Schwarz criterion, Schwarz Bayesian criterion, SBC
The Bayesian information criterion compares models fitted to the same data, as the Akaike criterion does, but penalises extra parameters more heavily, so it favours simpler models. Lower values are better, and it is widely used to choose the number of classes in latent class analysis.
Bayesian statistics
Also called: Bayesian inference, Bayesian analysis, Bayesian methods, Bayesian approach
Bayesian statistics is an approach to inference that treats unknown quantities as having probability distributions and uses Bayes' theorem to update prior beliefs about them with data, giving a posterior distribution. It answers questions such as how probable it is that an effect is positive, which a p-value cannot, at the cost of having to state and defend a prior.
Benjamini-Hochberg procedure
Also called: Benjamini–Hochberg procedure, BH procedure, Benjamini-Hochberg correction, BH correction
The Benjamini-Hochberg procedure is the standard method for controlling the false discovery rate, published by Yoav Benjamini and Yosef Hochberg in 1995. The p-values are ranked, each is compared with a threshold that rises with its rank, and every result up to the largest one that passes is declared a discovery.
Bimodal distribution
Also called: bimodal, two-peaked distribution
A bimodal distribution is one with two distinct peaks. It often signals that two different groups or processes have been mixed together in one set of data, and a single mean or standard deviation then describes the data poorly, so the groups may need to be looked at separately.
Binomial distribution
Also called: binomial probability distribution
The binomial distribution gives the probability of each possible number of successes in a fixed number of independent trials, each with two outcomes and the same probability of success. The number of patients responding out of 20 treated, when each has a 30% chance of responding, follows a binomial distribution.
Bonferroni correction
Also called: Bonferroni adjustment, Bonferroni method, Bonferroni procedure
The Bonferroni correction controls the familywise error rate by dividing alpha by the number of tests, or equivalently multiplying each p-value by it, so five tests at 0.05 are each judged at 0.01. It is simple and assumes nothing about how the tests are related, but it becomes very conservative, and loses power, as the number of tests grows.
Bootstrapping
Also called: bootstrap, bootstrap resampling, bootstrap method
Bootstrapping is a resampling method that estimates the sampling distribution of a statistic by drawing many samples, with replacement, from the observed data and recalculating the statistic in each. The spread of the results gives a standard error or confidence interval without a formula or an assumption of normality, and it is the recommended test of an indirect effect in mediation.
Box plot
Also called: box-and-whisker plot, boxplot, box and whisker diagram
A box plot is a graph that draws the median and quartiles of a variable as a box, with whiskers that in the usual convention reach to the furthest values within 1.5 interquartile ranges of the box, and any values beyond marked as possible outliers. Box plots side by side compare the centre, spread and skew of groups at a glance.
Censoring
Also called: censored data, right censoring, censored observation
Censoring is incomplete information about the time to an event, most often because a participant had not had the event when the study ended or left it early, known as right censoring. Survival methods use these partial times correctly, provided censoring is unrelated to the chance of the event.
Central limit theorem
Also called: CLT
The central limit theorem states that the sampling distribution of a mean becomes approximately normal as the sample size grows, whatever the shape of the population, provided its variance is finite. It is why tests of means work with non-normal data in large samples. How large is large enough depends on how skewed the population is.
Central tendency
Also called: measure of central tendency, measures of central tendency, measure of location, measure of centre
Central tendency is the typical or central value of a set of data, and a measure of central tendency is a single number that stands for it. The mean, median and mode are the three usually taught. They are equal in a symmetrical distribution with one peak and drift apart as it becomes skewed.
Chi-square distribution
Also called: chi-squared distribution, χ² distribution, chi square distribution
The chi-square distribution is the distribution of a sum of squared independent standard normal values, a right-skewed curve that takes only values of zero or more. Its shape depends on its degrees of freedom, and it supplies the p-values for chi-square tests of independence and goodness of fit and for likelihood ratio tests.
Chi-square goodness-of-fit test
Also called: chi-squared goodness-of-fit test, chi-square goodness of fit test, one-way chi-square test
A chi-square goodness-of-fit test checks whether the counts in the categories of a single variable match a distribution stated in advance, such as equal numbers choosing each of four options, or the proportions in a census. Large gaps between observed and expected counts give a large chi-square and a small p-value.
Chi-square test of independence
Also called: chi-squared test of independence, chi-square test of association, Pearson's chi-square test, χ² test of independence
A chi-square test of independence checks whether two categorical variables in a contingency table are associated, by comparing the observed counts with the counts expected if they were unrelated. It needs independent observations and adequate expected counts, a common rule being at least 5 in most cells, and otherwise Fisher's exact test is used.
Cluster analysis
Also called: clustering
Cluster analysis is a family of exploratory methods that sort cases into groups, or clusters, so that cases in the same group are more alike than cases in different groups, without the groups being known beforehand. Hierarchical and k-means clustering are the common forms, and the clusters found depend on the variables and the measure of distance chosen.
Cochran's Q test
Also called: Cochran's Q, Cochran Q test
Cochran's Q test extends McNemar's test to three or more related yes-or-no measurements, such as the same students passing or failing three tasks, testing whether the proportion of successes is the same across them. With two measurements it matches McNemar's test. It is different from the Q statistic for heterogeneity in meta-analysis, also named after Cochran.
Coefficient of variation
Also called: CV, relative standard deviation, RSD
The coefficient of variation is the standard deviation divided by the mean, often given as a percentage. It compares spread across measures with different units or very different means, but it only makes sense for ratio-scale data with a true zero and becomes unstable when the mean is close to zero.
Cohen's d
Also called: Cohen d, Cohens d, d
Cohen's d is an effect size for the difference between two means, expressed as that difference divided by a standard deviation, usually the pooled one. A d of 0.5 means the groups differ by half a standard deviation. Cohen offered 0.2, 0.5 and 0.8 as small, medium and large, benchmarks that methodologists warn should not be applied rigidly. Paired designs have several versions, such as dz.
Conditional probability
Also called: P(A|B)
A conditional probability is the probability of one event given that another has occurred, written P(A|B). The order matters: the probability of a positive test given a disease is not the probability of the disease given a positive test, and Bayes' theorem is the rule linking the two.
Confidence interval
Also called: CI, 95% confidence interval, 95% CI, confidence limits
A confidence interval is a range calculated from sample data by a method that, over many repeated samples, would capture the true parameter a stated percentage of the time, usually 95%. A single interval does not have a 95% chance of containing the true value. Its width shows the precision of the estimate, and it says more than a p-value alone.
Confirmatory factor analysis
Also called: CFA
Confirmatory factor analysis tests a factor structure specified in advance, stating which variables measure which factors, and judges how well that model fits the data using fit indices. It is used to check that a questionnaire measures what its developers claimed, often in a new sample after an exploratory analysis.
Contingency table
Also called: cross-tabulation, crosstab, cross-tab, two-way table
A contingency table is a table of counts that cross-classifies observations by two or more categorical variables, such as treatment group by outcome. It shows whether the distribution of one variable depends on the other and is the starting point for a chi-square test, Fisher's exact test, an odds ratio or a risk ratio.
Cook's distance
Also called: Cook's D, Cooks distance, Cook distance
Cook's distance measures how much all of a regression's predicted values would change if one observation were left out, combining the size of its residual with its leverage. Cases with unusually large values are checked as possibly influential, using rules of thumb that vary between textbooks.
Correlation
Also called: correlation coefficient, correlation analysis
A correlation is a measure of how closely two variables move together, summarised by a coefficient running from −1, a perfect negative relationship, through 0, none, to +1, a perfect positive one. Pearson's r measures linear relationships, and Spearman's rho and Kendall's tau ranked ones. A correlation, however strong, does not show that one variable causes the other.
Covariance
Covariance is a measure of how two variables vary together, positive when high values of one go with high values of the other and negative when they go with low values. Its size depends on the units of both variables, which makes it hard to interpret directly, and dividing it by the two standard deviations gives Pearson's correlation.
Cox proportional hazards model
Also called: Cox regression, Cox model, proportional hazards regression, Cox PH model
The Cox proportional hazards model is a regression model for time-to-event data that relates the hazard of an event to several predictors at once, giving an adjusted hazard ratio for each. It leaves the baseline hazard over time unspecified but assumes each predictor multiplies it by a constant amount, the proportional hazards assumption, which should be checked.
Cramér's V
Also called: Cramer's V, Cramér V, Cramer V
Cramér's V is an effect size for the association between two categorical variables in a contingency table, derived from the chi-square statistic and scaled from 0 for no association to 1 for perfect association. It works for tables of any size, and for a two-by-two table it matches the phi coefficient.
Credible interval
Also called: Bayesian credible interval, credibility interval
A credible interval is a range that contains a stated share, such as 95%, of a posterior distribution, so that, given the model and the prior, there is a 95% probability the parameter lies inside it. That direct reading is the one often wrongly given to a frequentist confidence interval.
Critical value
Also called: cut-off value, critical values
A critical value is the cut-off on a test statistic's distribution beyond which the null hypothesis is rejected at the chosen significance level, such as 1.96 for a two-tailed z-test at 0.05. The values beyond it form the rejection region. Comparing a test statistic with a critical value leads to the same decision as comparing its p-value with alpha.
Cumulative frequency
Also called: cumulative frequency distribution, cumulative count
A cumulative frequency is the running total of frequencies up to and including a given value or class. Divided by the total, it becomes the cumulative relative frequency, the proportion of observations at or below that value, which is how percentiles and medians are read from a grouped table.
Data coding
Also called: numerical coding, variable coding
Data coding is the assignment of numbers or short codes to responses so they can be analysed, such as 1 for yes and 2 for no, or a code for each occupation. Codes, including those kept for different kinds of missing value, should be applied consistently and recorded in the codebook. In qualitative research coding means labelling passages of text by meaning.
Data transformation
Also called: transformation, variable transformation, transforming data
A data transformation applies the same mathematical function to every value of a variable, such as taking logarithms or square roots, usually to reduce skew, stabilise variance or make a relationship linear. Results are then on the transformed scale, so they need back-transforming or careful wording when they are reported.
Degrees of freedom
Also called: df, d.f.
Degrees of freedom are the number of independent pieces of information left to estimate something once other quantities have been estimated from the same data. A sample variance has n minus 1, because the sample mean has already been used. The degrees of freedom set the exact shape of the t, chi-square and F distributions.
Descriptive statistics
Also called: descriptive statistic, summary statistics
Descriptive statistics are numbers and graphs that summarise the data actually collected, such as a mean, a standard deviation, a frequency table or a histogram. They describe the sample in hand and make no claim beyond it, which is the job of inferential statistics.
Direct effect
Also called: c prime, c′ path
A direct effect is the part of a predictor's effect on an outcome that does not pass through the mediators in a model. When it falls to zero once the mediator is included, mediation is described as complete, and otherwise as partial.
Discriminant analysis
Also called: linear discriminant analysis, LDA, discriminant function analysis
Discriminant analysis is a method for classifying cases into groups known in advance, using the combination of measured variables that best separates the groups, such as deciding which of two diagnoses fits a patient from a set of test results. Where the aim is to discover groups that are not yet known, cluster analysis is used instead.
Dummy coding
Also called: dummy variable, indicator variable, indicator coding, dummy variable coding
Dummy coding represents a categorical variable in a regression by a set of variables coded 0 or 1, one fewer than the number of categories. The category left out is the reference group, and each dummy variable's coefficient is the difference between its category and that reference.
Dunn's test
Also called: Dunn test, Dunn's multiple comparison test
Dunn's test is the most common post hoc procedure after a significant Kruskal-Wallis test, comparing groups pair by pair on their mean ranks. Its p-values are usually adjusted for the number of comparisons, for example with the Bonferroni or Holm method.
Dunnett's test
Also called: Dunnett test, Dunnett's multiple comparison test
Dunnett's test is a post hoc procedure that compares each of several treatment groups with a single control group, rather than every group with every other. Because it makes fewer comparisons than an all-pairs method, it has more power for that purpose.
Effect size
Also called: ES, magnitude of effect, size of effect
An effect size is a number that expresses how large a difference or relationship is, such as a mean difference, Cohen's d, a correlation or an odds ratio. Unlike a p-value it does not grow with sample size, and reporting it with a confidence interval is what lets readers judge importance and combine studies in a meta-analysis.
Eigenvalue
Also called: eigenvalues, latent root, characteristic root
In factor and principal component analysis, an eigenvalue is the amount of the total variance that a factor or component accounts for. The Kaiser criterion keeps every factor with an eigenvalue above 1 and is the default in much software, but it often keeps too many, so scree plots and parallel analysis are recommended instead or as well.
Equivalence test
Also called: equivalence testing, TOST, two one-sided tests
An equivalence test checks whether an effect is small enough to be treated as negligible, by testing it against upper and lower bounds set in advance, usually the smallest effect size of interest. The two one-sided tests procedure, TOST, is the usual method. It can support the absence of a meaningful effect, which a non-significant ordinary test cannot.
Estimation
Also called: statistical estimation, parameter estimation
Estimation is the part of statistical inference that puts a value, or a range of values, on an unknown population parameter using sample data, rather than testing a yes-or-no hypothesis about it. Many methodologists, and the American Statistical Association's statement on p-values, urge reporting estimates with intervals, not p-values alone.
Estimator
Also called: unbiased estimator
An estimator is a rule or formula for calculating an estimate of a parameter from sample data, such as the sample mean for the population mean. It is unbiased if its average over all possible samples equals the parameter, and a good estimator also varies little from sample to sample.
Eta squared
Also called: η², eta-squared
Eta squared is an effect size for analysis of variance that gives the proportion of the total variation in the outcome accounted for by a factor, its sum of squares divided by the total sum of squares. It overestimates the population value, especially in small samples, which is why omega squared is sometimes preferred.
Expected value
Also called: expectation, E(X)
The expected value of a random variable is its long-run average, the probability-weighted mean of all its possible values. A fair six-sided die has an expected value of 3.5, even though 3.5 can never be rolled. An estimator whose expected value equals the parameter it estimates is called unbiased.
Exploratory data analysis
Also called: EDA
Exploratory data analysis is an approach to a data set that uses mostly graphs and simple summaries to find its structure, outliers and oddities before any formal model is fitted. John Tukey set it out in a 1977 book, and it contrasts with confirmatory analysis, which tests hypotheses fixed in advance.
Exploratory factor analysis
Also called: EFA
Exploratory factor analysis looks for the underlying factor structure of a set of variables without specifying it in advance, letting every variable load on every factor. The researcher chooses how many factors to keep and how to rotate them, decisions with few absolute rules that call for judgement and are often debated.
Extrapolation
Extrapolation is using a model to predict outcomes for predictor values outside the range covered by the data. The relationship may change beyond that range, so extrapolated predictions can be badly wrong, as when a trend fitted over five years is projected fifty years ahead.
F distribution
Also called: F-distribution, Snedecor's F distribution, Fisher-Snedecor distribution
The F distribution is the distribution of a ratio of two variances, a right-skewed curve with one number of degrees of freedom for the numerator and another for the denominator. It supplies the p-values for the F-ratio in analysis of variance and for tests comparing two variances or two regression models.
F-ratio
Also called: F statistic, F value, F-statistic, F-test
The F-ratio is the test statistic of analysis of variance: the variance between group means divided by the variance within groups, each expressed as a mean square. A value near 1 suggests no real difference, and larger values are judged against an F distribution with two sets of degrees of freedom, reported in a form such as F(2, 57) = 4.31.
Factor analysis
Also called: common factor analysis
Factor analysis is a set of methods that explain the correlations among many observed variables, such as questionnaire items, by a smaller number of unobserved factors. It is used to build and check scales. It models the variables as expressions of the factors, which is what separates it from principal component analysis.
Factor loading
Also called: loading, factor loadings, component loading
A factor loading is the strength of the relationship between an observed variable and a factor, read much like a correlation in a standardised solution. The variables that load highly on a factor define what it means, and the minimum loading worth counting, often put at about 0.30 to 0.32, is a rule of thumb.
Factor rotation
Also called: rotation
Factor rotation is a step in factor analysis that turns the factors so that each variable loads strongly on as few factors as possible, making the solution easier to interpret without changing how much variance is explained. Orthogonal rotations such as varimax keep factors uncorrelated, while oblique rotations such as promax or oblimin let them correlate, which is usually more realistic in the social sciences.
Factorial ANOVA
Also called: factorial analysis of variance, multi-factor ANOVA, multifactor ANOVA, n-way ANOVA
A factorial analysis of variance tests two or more factors whose levels are fully crossed, so every combination is observed. It estimates a main effect for each factor and interaction effects for combinations of factors, and two-way and three-way analyses are named after the number of factors.
False discovery rate
Also called: FDR
The false discovery rate is the expected proportion of false positives among all the results declared significant. Controlling it, rather than the chance of any false positive at all, accepts a few errors in exchange for much more power, which suits studies that test hundreds or thousands of hypotheses at once, as in genetics.
Familywise error rate
Also called: FWER, family-wise error rate, experimentwise error rate, experiment-wise error rate
The familywise error rate is the probability of making at least one Type I error across a whole set, or family, of tests. Methods such as Bonferroni and Holm hold it at the chosen alpha. That is strict, and suits a small number of tests where any single false positive would be costly.
Fisher's exact test
Also called: Fisher exact test, Fisher's exact probability test, Fisher-Irwin test
Fisher's exact test calculates the exact probability of a contingency table at least as extreme as the one observed, given its row and column totals, instead of relying on the chi-square approximation. It is used for two-by-two tables, and some larger ones, when samples are small or expected counts fall below about 5.
Fisher's least significant difference
Also called: Fisher's LSD, LSD test, least significant difference test, protected LSD
Fisher's least significant difference test compares pairs of group means with ordinary t-tests after a significant analysis of variance, without further adjustment. It is the most liberal of the common post hoc methods, finding the most differences, and gives weak protection against false positives when there are many groups.
Fit indices
Also called: model fit indices, goodness-of-fit indices, fit statistics
Fit indices are statistics that summarise how well a structural equation or factor model reproduces the observed data, such as the chi-square test, the comparative fit index (CFI), the root mean square error of approximation (RMSEA) and the standardised root mean square residual (SRMR). Published cut-offs for good fit differ between authors and are debated, so several indices are reported together.
Five-number summary
Also called: five number summary
The five-number summary describes a set of data with its minimum, first quartile, median, third quartile and maximum. It conveys the centre, the spread and a hint of skew in five values, and it is the information a basic box plot draws.
Fixed and random effects
Also called: fixed effects, random effects, fixed effect, random effect
In a mixed or multilevel model, a fixed effect is estimated for specific levels of interest, such as a treatment, while a random effect treats its levels, such as the schools sampled, as a draw from a wider population and estimates their variance. Econometrics and meta-analysis use the terms differently, so their meaning depends on the field.
Frequency
Also called: absolute frequency, count
A frequency is the number of times a value, or a value within a class interval, occurs in a set of data. Frequencies are the raw material of frequency tables, bar charts and histograms, and the counts that chi-square tests work on.
Frequency distribution
Also called: frequency table, frequency distribution table, grouped frequency table
A frequency distribution is a table or graph showing how many observations fall into each category or class interval of a variable. Grouping continuous data into intervals makes a large data set readable, and the result is usually drawn as a histogram or bar chart.
Frequentist statistics
Also called: frequentist inference, classical statistics, frequentism
Frequentist statistics is the approach that treats probability as long-run frequency and judges methods by how they would perform over many repeated samples. P-values, significance tests and confidence intervals belong to it. Parameters are treated as fixed unknowns, whereas Bayesian statistics gives them probability distributions.
Friedman test
Also called: Friedman's test, Friedman two-way ANOVA by ranks, Friedman's ANOVA
The Friedman test is a non-parametric test for three or more related measurements, such as the same participants rating three products, which ranks the values within each participant or block. It is the rank-based alternative to the one-way repeated-measures analysis of variance and extends the sign test to more than two conditions.
Full information maximum likelihood
Also called: FIML, direct maximum likelihood, raw maximum likelihood
Full information maximum likelihood estimates a model's parameters directly from all the available data, including incomplete cases, without filling anything in. Like multiple imputation it is valid when data are missing at random, and structural equation modelling programs commonly offer it.
Games-Howell test
Also called: Games-Howell post hoc test, Games and Howell test
The Games-Howell test is a post hoc procedure for comparing every pair of group means when the groups have unequal variances, and often unequal sizes. Each comparison uses only the two groups' own standard deviations and sample sizes, in the manner of Welch's t-test, which makes it a common follow-up to Welch's analysis of variance.
Garden of forking paths
Also called: forking paths
The garden of forking paths is Andrew Gelman and Eric Loken's name for the problem that even when only one analysis is run, the choices behind it were shaped by the data, and different data would have led to different but equally reasonable choices. The p-value then overstates the evidence, without any deliberate fishing by the researcher.
General linear model
Also called: linear model
The general linear model is the framework that treats t-tests, analysis of variance, analysis of covariance and linear regression as special cases of one equation, in which a continuous outcome is a weighted sum of predictors plus normally distributed error. It should not be confused with the generalised linear model, which extends it to other kinds of outcome.
Generalised estimating equations
Also called: generalized estimating equations, GEE
Generalised estimating equations are a method for correlated data, such as repeated measures or clustered observations, that extends generalised linear models to estimate average effects across a population while correcting standard errors for the correlation. Unlike a mixed-effects model, they give population-average rather than subject-specific effects.
Generalised linear model
Also called: generalized linear model, GLM, GLiM
A generalised linear model extends linear regression to outcomes that are not normally distributed, such as binary outcomes or counts, by linking the mean of the outcome to the predictors through a function such as the log or the logit. Logistic and Poisson regression are the best-known examples, and such models are fitted by maximum likelihood.
Geometric mean
The geometric mean is the nth root of the product of n positive values, which is the same as the back-transformed mean of their logarithms. It suits quantities that grow by multiplication, such as rates of growth, and positive data with a long right tail.
Glass's delta
Also called: Glass's Δ, Glass delta, Glass' delta
Glass's delta is a standardised mean difference that divides the difference between two group means by the standard deviation of the control or comparison group only. It is preferred when the intervention may itself change how spread out scores are, so that pooling the two standard deviations would mislead.
HARKing
Also called: hypothesizing after the results are known, hypothesising after the results are known, post hoc hypothesising
HARKing, short for hypothesising after the results are known, is presenting a hypothesis suggested by the results as if it had been set out before the study. Norbert Kerr named it in 1998. It turns an exploratory finding into what looks like a confirmed prediction and hides how many other patterns could have been reported instead.
Harmonic mean
The harmonic mean is the number of values divided by the sum of their reciprocals. It is the right average for rates such as speeds over equal distances, and in statistics it is often used to average unequal group sizes, for example in some post hoc tests.
Hazard
Also called: hazard rate, hazard function, instantaneous risk
The hazard is the instantaneous rate at which an event occurs at a given time among those who have not yet had it. Unlike cumulative risk it can rise or fall during follow-up, and survival models such as Cox regression are built on it.
Hazard ratio
Also called: HR
A hazard ratio compares the rate at which an event such as death or relapse occurs over time in one group with the rate in another, where 1 means no difference. It comes from survival analysis, usually a Cox model, and assumes the ratio stays roughly constant over follow-up. A ratio of 0.7 means a 30% lower event rate at any moment, not 30% fewer events overall.
Hedges' g
Also called: Hedges g, Hedges's g, corrected effect size
Hedges' g is a standardised mean difference like Cohen's d with a correction that removes d's slight overestimate in small samples. The two are almost identical once groups have more than about 20 people each. The standardised mean difference used in Cochrane reviews is Hedges' adjusted g.
Heteroscedasticity
Also called: heteroskedasticity, unequal variances, non-constant variance, heterogeneity of variance
Heteroscedasticity is unequal spread: residuals in a regression, or scores in different groups, that vary more at some levels than at others, as when errors grow with the predicted value. It does not bias regression coefficients but makes their standard errors and p-values unreliable. A funnel shape in a plot of residuals against fitted values is the classic sign.
Hierarchical cluster analysis
Also called: hierarchical clustering, agglomerative clustering, agglomerative hierarchical clustering
Hierarchical cluster analysis builds clusters step by step, usually starting with every case on its own and repeatedly merging the two closest clusters until all are joined. The result is drawn as a tree, or dendrogram, and the researcher chooses where to cut it to decide how many clusters there are.
Hierarchical regression
Also called: hierarchical multiple regression, sequential regression, blockwise regression
Hierarchical regression enters predictors into a multiple regression in blocks chosen by the researcher, such as background variables first and the variables of interest second, and tests how much R squared rises at each step. Despite the name it is unrelated to hierarchical linear models, which are multilevel models for nested data.
Histogram
A histogram is a graph of the distribution of a numeric variable in which the range of values is split into adjacent intervals and the height of each bar shows how many observations fall in it. Unlike a bar chart its bars touch, and its apparent shape depends partly on the width chosen for the intervals.
Hochberg procedure
Also called: Hochberg's step-up procedure, Hochberg method, Hochberg correction
The Hochberg procedure is a step-up correction for multiple tests that uses the same thresholds as Holm's method but starts from the largest p-value and works down. It has more power than Holm's method but relies on an assumption about how the tests are related, typically that they are independent or positively correlated.
Holm-Bonferroni method
Also called: Holm method, Holm's method, Holm procedure, Holm correction, sequential Bonferroni, Holm's step-down procedure
The Holm-Bonferroni method is a step-down correction that orders p-values from smallest to largest and tests each against a gradually less strict threshold, alpha divided by the number of tests still remaining. It controls the familywise error rate as Bonferroni does but always has at least as much power, so it is usually preferred.
Homogeneity of variance
Also called: homoscedasticity, homoskedasticity, equal variances, equality of variances, homogeneity of variances
Homogeneity of variance is the assumption that the groups being compared, or the residuals at every level of a predictor, have the same variance. Student's t-test, classical analysis of variance and linear regression assume it. When it fails, especially with unequal group sizes, Welch's versions of the t-test and analysis of variance are safer.
Hypothesis testing
Also called: null hypothesis significance testing, NHST, significance testing, statistical hypothesis testing, hypothesis test
Hypothesis testing is a procedure for deciding whether data are inconsistent enough with a null hypothesis to reject it: state the hypotheses, choose a significance level, calculate a test statistic and its p-value, then reject or fail to reject the null. Failing to reject is not the same as accepting the null. Critics argue the procedure encourages yes-or-no thinking about evidence.
Imputation
Also called: single imputation, data imputation
Imputation is replacing missing values with estimated ones so that methods for complete data can be used. Single imputation, such as filling in the mean or a regression prediction, treats guesses as if they were real data, which makes standard errors too small, the problem that multiple imputation was designed to solve.
Independence of observations
Also called: independence assumption, independent observations, assumption of independence
Independence of observations is the assumption that each data point carries no information about any other, as when each participant is measured once and does not influence the others. Repeated measurements or pupils in the same class break it, and ignoring the dependence understates standard errors, so paired, repeated-measures or multilevel methods are used instead.
Independent events
Also called: statistical independence, independence of events
Two events are independent when the occurrence of one does not change the probability of the other, so the probability of both is the product of their separate probabilities. Successive tosses of a fair coin are independent: a run of heads does not make tails any more likely next time.
Independent-samples t-test
Also called: independent samples t-test, independent t-test, two-sample t-test, unpaired t-test, between-subjects t-test
An independent-samples t-test compares the means of two separate groups, such as a treatment group and a control group with different people in each. The classical Student version pools the two variances and assumes they are equal. Welch's version does not, and many methodologists now recommend it as the default.
Indirect effect
Also called: mediated effect, ab path
An indirect effect is the part of a predictor's effect on an outcome that passes through a mediator, estimated as the product of the path from predictor to mediator and the path from mediator to outcome. In a simple linear mediation model the total effect equals the direct effect plus the indirect effect.
Inferential statistics
Also called: statistical inference, inductive statistics, inference
Inferential statistics are methods for drawing conclusions about a population from a sample while allowing for the chance variation that comes with sampling. Estimation with confidence intervals and hypothesis testing are its two main forms, and both depend on the sample having been drawn in a way that lets it stand for the population.
Influential observation
Also called: influential point, influential case, influential data point
An influential observation is a single case whose removal would noticeably change a regression's results. Influence combines being an outlier on the outcome with having high leverage, an unusual combination of predictor values, and it is measured with statistics such as Cook's distance.
Instrumental variable analysis
Also called: instrumental variables, IV analysis, instrumental variable estimation
Instrumental variable analysis estimates a causal effect from observational data by using an instrument, a variable that influences which treatment people receive but affects the outcome only through that treatment and shares no cause with it. It can deal with unmeasured confounding, but its key assumptions cannot be fully tested from the data.
Interaction effect
Also called: interaction, statistical interaction, interaction term
An interaction effect occurs when the effect of one factor on an outcome depends on the level of another, such as a teaching method that helps beginners but not advanced students. It shows as non-parallel lines on an interaction plot and is tested with a product term in analysis of variance or regression. Moderation analysis frames the same idea around a moderator.
Intercept
Also called: constant, y-intercept, regression constant, b0
The intercept is the predicted value of the outcome when every predictor in a regression equals zero, the point where the line crosses the vertical axis. It is often meaningless on its own, such as a predicted salary at age zero, unless predictors are centred so that zero stands for a sensible value such as the average.
Interquartile range
Also called: IQR, inter-quartile range, midspread
The interquartile range is the distance between the first and third quartiles, so it covers the middle 50% of the data. Because it ignores the top and bottom quarters it is not thrown by extreme values, and it is the usual measure of spread to report with a median.
Interval estimate
Also called: interval estimation
An interval estimate is a range of plausible values for a population parameter calculated from sample data, rather than a single number. A confidence interval is the frequentist form and a credible interval the Bayesian form, and the two are built and interpreted differently.
K-means clustering
Also called: k-means, k-means cluster analysis, K means clustering
K-means clustering divides cases into a number of clusters, k, chosen in advance, by repeatedly assigning each case to the cluster with the nearest mean and recalculating the means until assignments stop changing. It is efficient with large data sets, but its results depend on the starting points and on the choice of k.
Kaplan-Meier estimator
Also called: Kaplan-Meier method, Kaplan-Meier curve, Kaplan–Meier estimator, product-limit estimator, KM curve, Kaplan Meier
The Kaplan-Meier estimator is a method of estimating the proportion of people still free of an event at each point in time, using both complete and censored times. It is drawn as a step-shaped survival curve that drops at each event, and curves for different groups are compared with the log-rank test.
Kendall's tau
Also called: Kendall's τ, Kendall rank correlation, Kendall's tau-b, Kendall tau
Kendall's tau is a rank correlation based on counting pairs of observations: concordant pairs, ordered the same way on both variables, against discordant pairs, ordered in opposite ways. Like Spearman's rho it measures monotonic association, and its tau-b version adjusts for the tied values common in ordinal data.
Kolmogorov-Smirnov test
Also called: K-S test, KS test, Kolmogorov–Smirnov test
The Kolmogorov-Smirnov test compares a sample's distribution with a reference distribution, such as the normal, or with another sample, using the largest gap between their cumulative distributions. As a test of normality it has low power, and the Lilliefors correction is needed when the mean and standard deviation come from the data, so Shapiro-Wilk is generally preferred.
Kruskal-Wallis test
Also called: Kruskal–Wallis test, Kruskal-Wallis H test, H test, Kruskal-Wallis one-way ANOVA on ranks
The Kruskal-Wallis test is a non-parametric test comparing three or more independent groups on ranked data, the rank-based alternative to the one-way analysis of variance. It tests whether values in some groups tend to be higher than in others, and it compares medians only if the group distributions share the same shape. Dunn's test usually follows a significant result.
Kurtosis
Also called: excess kurtosis
Kurtosis describes how heavy the tails of a distribution are compared with a normal distribution, that is, how prone it is to extreme values. It is often taught as peakedness, but it mainly reflects the tails. A normal distribution has kurtosis 3, which many programs report as excess kurtosis of 0. Heavy tails are called leptokurtic and light tails platykurtic.
Last observation carried forward
Also called: LOCF
Last observation carried forward fills in a participant's missing later measurements with their last recorded value, a method once common in clinical trials. It assumes nothing would have changed after drop-out, can bias results in either direction even when data are missing completely at random, and understates uncertainty.
Latent class analysis
Also called: LCA, latent class model, latent class modelling
Latent class analysis is a method that identifies hidden subgroups, or classes, in a population from people's patterns of answers on categorical indicators, such as symptoms present or absent. It is person-centred, grouping people rather than variables, and the number of classes is chosen with fit statistics such as BIC. With continuous indicators the equivalent method is latent profile analysis.
Law of large numbers
Also called: LLN
The law of large numbers states that as a sample grows, its mean tends to get closer to the population mean. It underlies the reading of probability as a long-run proportion and the greater stability of estimates from larger samples, though it promises nothing about any short run of results.
Levene's test
Also called: Levene test, Levene's test for equality of variances
Levene's test checks whether several groups have equal variances, the homogeneity assumption of analysis of variance, and is less upset by non-normal data than Bartlett's test. Its median-based version, the Brown-Forsythe test, is more robust still. Using it to choose between Student's and Welch's t-test is now discouraged, because with small samples it often lacks power.
Leverage
Also called: hat value, leverage value
Leverage is how far an observation's predictor values lie from the centre of the other observations' predictor values. A high-leverage point has the potential to pull a regression line towards itself, and whether it actually does depends on whether its outcome also departs from the trend.
Likelihood
Also called: likelihood function
The likelihood is the probability, or probability density, of the observed data under each possible value of a parameter, viewed as a function of the parameter. It is the part of a Bayesian analysis that carries the evidence from the data, and maximising it gives maximum likelihood estimates. It is not the probability that the parameter takes a given value.
Likelihood ratio test
Also called: LR test, LRT, likelihood-ratio test
A likelihood ratio test compares two nested models, one a simpler version of the other, by the ratio of their maximised likelihoods. If the extra terms in the larger model improve the fit no more than chance would, the simpler model is kept. It is widely used with logistic and other generalised linear models.
Listwise deletion
Also called: complete-case analysis, complete case analysis, casewise deletion, listwise exclusion
Listwise deletion analyses only the cases with no missing values on any variable in the analysis, dropping everyone else. It is the default in most software. It is unbiased only when data are missing completely at random, and even then it wastes information and loses power.
Little's MCAR test
Also called: Little's test, Little's missing completely at random test
Little's MCAR test checks whether the pattern of missing values in a data set is consistent with data missing completely at random, a significant result suggesting it is not. A non-significant result does not prove it, and no test on the observed data can tell missing at random from missing not at random.
Log-rank test
Also called: logrank test, log rank test, Mantel-Cox test
The log-rank test compares the survival curves of two or more groups across the whole follow-up period, testing whether events happen at the same rate in each. It is the most widely used test in survival analysis and is usually reported beside Kaplan-Meier curves.
Logistic regression
Also called: binary logistic regression, logit regression, logit model, logistic model
Logistic regression models a binary outcome, such as pass or fail, by predicting the log odds of the event from one or more predictors. Exponentiated, each coefficient becomes an odds ratio: the factor by which the odds change for a one-unit increase in that predictor, adjusted for the others. It is fitted by maximum likelihood.
Main effect
A main effect is the effect of one factor on the outcome averaged across the levels of the other factors in a design. In a study of drug and dose, the main effect of drug is the overall difference between drugs. When there is an interaction, a main effect can mislead, because the factor's effect is not the same everywhere.
Mann-Whitney U test
Also called: Mann–Whitney U test, Mann-Whitney test, Wilcoxon rank-sum test, Wilcoxon-Mann-Whitney test, U test
The Mann-Whitney U test is a non-parametric test comparing two independent groups by ranking all the values together and asking whether one group tends to have higher values than the other. It is often described as a test of medians, but it is one only when the two distributions have the same shape and spread. It is the rank-based alternative to the independent-samples t-test.
Markov chain Monte Carlo
Also called: MCMC, Markov chain Monte Carlo sampling
Markov chain Monte Carlo is a family of computer methods that approximate a posterior distribution too complex to calculate directly by drawing a long chain of random samples from it. Early samples, the burn-in, are discarded, and several chains are run from different starting points to check that they have settled on the same distribution.
Maximum likelihood estimation
Also called: MLE, maximum likelihood, ML estimation
Maximum likelihood estimation is a method of estimating parameters by choosing the values that make the observed data most probable under the model. It is how logistic, Poisson and many other models are fitted, in place of the least squares method used for ordinary linear regression.
McNemar's test
Also called: McNemar test, McNemar's chi-square test
McNemar's test compares two proportions from paired data, such as the same people answering yes or no before and after an intervention. It uses only the discordant pairs, those who changed in one direction or the other, and tests whether change was equally likely both ways.
Mean
Also called: arithmetic mean, average, sample mean
The mean is the sum of all the values divided by how many values there are, and is what most people mean by the average. It uses every value, which makes it sensitive to outliers and skew, so the median is often reported alongside it, or instead of it, when a distribution is lopsided.
Mean difference
Also called: MD, difference in means, raw mean difference, unstandardised mean difference
A mean difference is simply one group's mean minus another's, in the original units of measurement, such as 4.2 points on a depression scale. It is the most interpretable effect size when everyone used the same measure, and it needs standardising only when scales differ.
Mean imputation
Also called: mean substitution, mean replacement, mean value imputation
Mean imputation fills each missing value with the mean of the observed values of that variable. It is simple, but it shrinks the variance, weakens correlations and biases almost every estimate except the mean itself, which is biased too unless data are missing completely at random, so it is not recommended.
Median
Also called: 50th percentile, Mdn
The median is the middle value when the data are put in order, so that half the values lie at or below it and half at or above. With an even number of values it is the mean of the middle two. It is barely moved by extreme values, which is why it is preferred for skewed data such as incomes.
Mediation analysis
Also called: mediation, mediation model, mediator analysis, mediational analysis
Mediation analysis tests whether a predictor affects an outcome through an intervening variable, the mediator, splitting the total effect into an indirect effect that runs through the mediator and a direct effect that does not. The older Baron and Kenny steps have largely given way to testing the indirect effect itself, usually with bootstrapping. It assumes a causal order the data alone cannot prove.
Missing at random
Also called: MAR
Data are missing at random when the chance of a value being missing depends only on information that was observed, as when older participants skip an online follow-up more often and age is recorded. The name misleads, since the missingness is not random in the everyday sense. Multiple imputation and maximum likelihood methods give valid results under it.
Missing completely at random
Also called: MCAR
Data are missing completely at random when the chance of a value being missing is the same for every case and unrelated to anything, observed or not, as when a sample tube is dropped by accident. It is the only mechanism under which analysing only the complete cases gives unbiased results, and it is rarely realistic.
Missing data
Also called: missing values, missingness, incomplete data
Missing data are values that were meant to be collected but were not, through non-response, drop-out, skipped items or errors. How they should be handled depends on why they are missing, the missing data mechanism, and simply analysing the complete cases can bias results and waste information.
Missing not at random
Also called: MNAR, not missing at random, NMAR, non-ignorable missingness, nonignorable missing data
Data are missing not at random when the chance of a value being missing depends on the missing value itself or on something unobserved, as when the heaviest drinkers are least likely to report their drinking. No standard method fully corrects for it, so results are checked under different assumptions in a sensitivity analysis.
Mixed ANOVA
Also called: mixed-design ANOVA, split-plot ANOVA, mixed between-within ANOVA, SPANOVA
A mixed analysis of variance combines at least one between-subjects factor, such as group, with at least one within-subjects factor, such as time. Its key result is often the group-by-time interaction, which asks whether the groups changed differently. It is not the same thing as a mixed-effects model, although the names are easily confused.
Mode
Also called: modal value
The mode is the value that occurs most often in a set of data. It is the only average that works for categories with no order, such as blood group, and a distribution can have one mode, several or none.
Moderation analysis
Also called: moderation, moderator analysis, moderated regression, moderated multiple regression
Moderation analysis tests whether the strength or direction of a relationship depends on a third variable, the moderator, such as a revision method that helps weaker students more than stronger ones. It is usually done by adding the product of predictor and moderator to a regression, and it answers when or for whom an effect holds, where mediation answers how.
Multicollinearity
Also called: collinearity, multi-collinearity
Multicollinearity is strong correlation among the predictors in a regression. It does little harm to the model's overall predictions, but it inflates the standard errors of individual coefficients, making them unstable and hard to interpret, and it can even flip their signs. It is detected with variance inflation factors.
Multilevel model
Also called: multilevel modelling, multilevel modeling, hierarchical linear model, HLM, mixed-effects model, mixed model, linear mixed model, LMM, random coefficient model
A multilevel model is a regression model for data with a nested structure, such as pupils within schools or repeated measurements within people, that estimates variation at each level. Treating nested observations as independent understates standard errors and overstates significance. It is also called a hierarchical linear or mixed-effects model, names that stress different features of the same approach.
Multinomial logistic regression
Also called: multinomial regression, multinomial logit model, polytomous logistic regression, baseline-category logit model
Multinomial logistic regression extends logistic regression to an outcome with three or more unordered categories, such as the choice between bus, car and bicycle, comparing each category with a chosen reference category. Its coefficients exponentiate to odds ratios for being in one category rather than the reference.
Multiple comparisons problem
Also called: multiple testing, multiplicity, multiple comparisons, multiple testing problem
The multiple comparisons problem is the rise in false positives when many hypothesis tests are run: at an alpha of 0.05, twenty independent tests of true null hypotheses give at least one significant result about 64% of the time. Corrections such as Bonferroni, Holm or false discovery rate methods are used to control it.
Multiple imputation
Also called: MI
Multiple imputation replaces each missing value with several plausible values drawn from a model of the data, creating several complete data sets. Each is analysed separately and the results are pooled using Rubin's rules, so the final standard errors include the uncertainty about the missing values. It gives valid results when data are missing at random and the imputation model is sound.
Multiple linear regression
Also called: multiple regression, MLR, multivariable linear regression, multivariable regression
Multiple linear regression models a continuous outcome from two or more predictors, so each coefficient shows the change in the outcome for a one-unit change in that predictor with the others held constant. It is how researchers adjust for other variables, though it cannot adjust for anything left out. Calling it multivariate regression is common but loose, since that strictly means several outcomes.
Multivariate analysis of variance
Also called: MANOVA
Multivariate analysis of variance tests group differences on several related outcome variables at once, such as three subscales of one questionnaire, taking their correlations into account. Wilks' lambda and Pillai's trace are its usual test statistics, and adding covariates gives multivariate analysis of covariance, MANCOVA.
Multiverse analysis
Also called: multiverse approach
A multiverse analysis runs and reports every reasonable combination of data processing and analysis choices for a single hypothesis, rather than one chosen path. Showing the whole spread of results reveals how much a conclusion depends on arbitrary decisions, and so how robust it is.
Negative binomial regression
Also called: negative binomial model, NB regression
Negative binomial regression is a model for count outcomes that allows the variance to exceed the mean, the overdispersion that makes Poisson regression understate its standard errors. Its coefficients are read as rate ratios in the same way as in Poisson regression.
Non-parametric test
Also called: nonparametric test, non-parametric statistics, nonparametric statistics, distribution-free test, distribution-free method
A non-parametric test is one that makes few or no assumptions about the shape of the population distribution, often by working with ranks instead of raw values. The Mann-Whitney, Wilcoxon, Kruskal-Wallis and Friedman tests are examples. They suit ordinal data, small samples and outliers, at the cost of some power when parametric assumptions do hold.
Normal distribution
Also called: Gaussian distribution, bell curve, normal curve
The normal distribution is a symmetrical, bell-shaped continuous probability distribution fixed by its mean and standard deviation. About 68% of values lie within one standard deviation of the mean and about 95% within two. Many tests assume it, and by the central limit theorem sample means tend towards it even when the data do not.
Normality assumption
Also called: assumption of normality, normality, normal distribution assumption
The normality assumption is the requirement of many parametric tests that the data, or more exactly the residuals or the sampling distribution of the statistic, be approximately normal. It matters most in small samples, since large samples are protected by the central limit theorem. It is checked with Q-Q plots, histograms and tests such as Shapiro-Wilk.
Number needed to harm
Also called: NNH, NNTH, number needed to treat for an additional harmful outcome
The number needed to harm is how many people must receive an intervention for one more person to suffer a particular harm, calculated like the number needed to treat from the increase in risk. Cochrane prefers the phrase number needed to treat for an additional harmful outcome, NNTH, which keeps the direction plain.
Number needed to treat
Also called: NNT, NNTB, number needed to treat for an additional beneficial outcome, number needed to benefit
The number needed to treat is how many people must receive an intervention rather than the comparison for one more person to benefit, calculated as one divided by the absolute risk reduction. An NNT of 10 means ten treated for one extra good outcome. It is meaningful only alongside the outcome, comparison and time period it refers to.
Odds
Odds are the probability that an event happens divided by the probability that it does not, so a 20% risk is odds of 0.25, or 1 to 4. Odds and risk are close when events are rare but diverge as events become common, which is why an odds ratio should not be read as a risk ratio.
Odds ratio
Also called: OR
An odds ratio is the odds of an event in one group divided by the odds in another, where 1 means no difference. It is the natural effect measure of case-control studies and logistic regression. When events are common, with risks above roughly 20%, it lies noticeably further from 1 than the risk ratio and, if misread as one, exaggerates the effect.
Omega squared
Also called: ω², omega-squared
Omega squared is an effect size for analysis of variance that estimates the proportion of variance a factor explains in the population, correcting the upward bias of eta squared. For the same data it is always a little smaller than eta squared.
Omnibus test
Also called: overall test, global test
An omnibus test checks whether there is any difference or effect at all across several groups or terms at once, without saying where it lies. The F-test in a one-way analysis of variance is the usual example: a significant result says the group means are not all equal, and post hoc tests then locate the differences.
One-sample t-test
Also called: one sample t-test, single-sample t-test, one-sample t test
A one-sample t-test checks whether the mean of one sample differs from a fixed value, such as a published norm or the midpoint of a rating scale. It divides the difference by the standard error of the mean and compares the result with a t distribution on n minus 1 degrees of freedom.
One-tailed test
Also called: one-sided test, directional test, one tailed test, one-tailed
A one-tailed test looks for an effect in one stated direction only, placing the whole rejection region in one tail of the distribution. It has more power in that direction but cannot detect an effect the other way, and it must be chosen before the data are seen or it inflates the false positive rate.
One-way ANOVA
Also called: one-way analysis of variance, single-factor ANOVA, one way ANOVA
A one-way analysis of variance compares the means of three or more independent groups defined by a single factor, such as three teaching methods. A significant F-ratio says the means are not all equal but not which ones differ, so it is followed by post hoc tests or planned comparisons.
Optional stopping
Also called: data peeking, peeking at the data
Optional stopping is checking the results repeatedly while data are being collected and stopping as soon as they become statistically significant. Without a correction planned in advance, such as the rules used for interim analyses in sequential trial designs, it inflates the false positive rate well above the stated alpha.
Ordinal logistic regression
Also called: ordinal regression, proportional odds model, ordered logistic regression, ordered logit model, cumulative logit model
Ordinal logistic regression models an outcome with ordered categories, such as a rating from poor to excellent, using the order without treating the gaps between categories as equal. Its usual form, the proportional odds model, assumes each predictor has the same effect at every cut-point between categories, an assumption that should be checked.
Ordinary least squares
Also called: OLS, least squares, method of least squares, least squares regression
Ordinary least squares is the standard method of fitting a linear regression, choosing the intercept and slopes that make the sum of the squared residuals as small as possible. Squaring makes large errors count heavily, which is why a few outliers can pull the fitted line a long way.
Outlier
Also called: outlying value, extreme value, anomalous value
An outlier is an observation that lies unusually far from the rest of the data. There is no universally agreed criterion, though a common rule of thumb flags values more than 1.5 interquartile ranges beyond the quartiles. An outlier should be investigated rather than deleted automatically, because it may be an error or a genuine and important case.
Overdispersion
Also called: over-dispersion, extra-Poisson variation
Overdispersion is more variability in count or binary data than a model assumes, most often counts whose variance is larger than their mean under a Poisson model. Left alone it makes standard errors too small and p-values too optimistic, and it is handled with a scale adjustment, a quasi-Poisson model or negative binomial regression.
Overfitting
Also called: over-fitting, overfit model
Overfitting is building a model so closely around the quirks of one data set that it describes noise as well as signal, so it fits the sample well but predicts new data poorly. Too many predictors for the sample size is the usual cause, and testing the model on new data, or by cross-validation, reveals it.
P-hacking
Also called: p hacking, data dredging, fishing, fishing expedition, data fishing, inflation bias
P-hacking is analysing data in many ways, or collecting and excluding data, until a statistically significant result appears, then reporting that result as if it were the only analysis. Examples include adding participants until p falls below 0.05, dropping outliers after seeing their effect and testing many outcomes but reporting one. It makes false positives far more likely than alpha suggests.
P-value
Also called: p value, probability value, p, significance probability
A p-value is the probability of getting a test statistic at least as extreme as the one observed if the null hypothesis and every assumption of the test were true. It measures how incompatible the data are with that model. It is not the probability that the null hypothesis is true or that chance alone produced the result, and it does not measure the size or importance of an effect.
Paired t-test
Also called: paired samples t-test, paired-samples t-test, dependent t-test, repeated-measures t-test, related t-test, matched-pairs t-test, correlated pairs t-test
A paired t-test compares two means from the same people or from matched pairs, such as scores before and after a course, by testing whether the mean of the differences is zero. Because each person serves as their own control, it removes stable differences between people and usually has more power than an independent-samples test.
Pairwise deletion
Also called: available-case analysis, available case analysis, pairwise exclusion
Pairwise deletion uses every case that has data for each particular calculation, so different correlations in the same analysis can rest on different subsets of people. It keeps more data than listwise deletion, but it is valid only when data are missing completely at random and leaves it unclear what sample size the standard errors should use.
Parameter
Also called: population parameter
A parameter is a numerical characteristic of a whole population, such as its true mean or proportion, usually written with a Greek letter such as μ or σ. It is almost never known and is estimated from a sample statistic.
Parametric test
Also called: parametric statistics, parametric method
A parametric test is one that assumes the data come from a distribution of known form, usually normal, described by parameters such as a mean and a variance. The t-test, analysis of variance and Pearson correlation are parametric. When their assumptions hold they have more power than non-parametric alternatives.
Partial correlation
A partial correlation is the correlation between two variables after the linear influence of one or more other variables has been removed from both. The correlation between reading and maths scores with age held constant is an example, used to see whether a relationship survives once a likely common cause is accounted for.
Partial eta squared
Also called: ηp², partial η², partial eta-squared
Partial eta squared is an effect size for analysis of variance that gives a factor's sum of squares as a proportion of that factor plus error, leaving out variation explained by the other factors. It is what many programs print by default. Because it depends on which other factors were in the design, values from different designs are not directly comparable.
Path analysis
Also called: path model
Path analysis is a form of structural equation modelling with observed variables only, which estimates a set of linked regression equations, such as a chain from parents' education to reading to exam results. Drawn as a path diagram, it separates direct from indirect effects but, like regression, cannot establish causation on its own.
Pearson correlation coefficient
Also called: Pearson's r, Pearson correlation, Pearson product-moment correlation, product-moment correlation coefficient, PPMCC, r
The Pearson correlation coefficient, r, measures the strength and direction of the linear relationship between two continuous variables, from −1 to +1, and is also used as an effect size. It is sensitive to outliers and can be near zero for a strong curved relationship, so a scatter plot should always be checked. Squared, it gives the proportion of variance the two variables share.
Percentile
Also called: centile
A percentile is the value below which a given percentage of the data fall, so the 90th percentile is exceeded by only about a tenth of the values. There is no single agreed rule for calculating one, and different statistical programs can give slightly different answers for the same small data set.
Permutation test
Also called: randomisation test, randomization test, exact permutation test
A permutation test works out a p-value by shuffling the data between groups many times, as if the null hypothesis of no difference were true, and counting how often a shuffled result is as extreme as the observed one. It makes few assumptions about distributions, though it handles only fairly simple designs.
Phi coefficient
Also called: phi, φ coefficient, phi correlation
The phi coefficient is a measure of association between two binary variables, calculated from a two-by-two table, and equal to a Pearson correlation between two variables coded 0 and 1. It runs from −1 to 1 and is commonly reported as the effect size for a two-by-two chi-square test.
Planned comparison
Also called: a priori comparison, planned contrast, a priori contrast, contrast
A planned comparison is a specific comparison between group means chosen before the data are seen, usually to test a hypothesis directly. Because only a few such comparisons are made, they need less correction than post hoc tests, and methodologists differ on how much correction, if any, a small set of independent, or orthogonal, contrasts requires.
Point estimate
Also called: point estimation
A point estimate is a single number calculated from a sample as the best guess of a population parameter, such as a sample mean of 24.3 for the population mean. On its own it says nothing about precision, which is why it is reported with a standard error or a confidence interval.
Point-biserial correlation
Also called: point biserial correlation, point-biserial correlation coefficient
The point-biserial correlation measures the relationship between a binary variable, such as treatment or control, and a continuous one. It is simply Pearson's r with the binary variable coded 0 and 1, and testing it is equivalent to an independent-samples t-test on the same data.
Poisson distribution
Also called: Poisson probability distribution
The Poisson distribution gives the probability of a given number of events in a fixed interval of time or space when events occur independently at a constant average rate. Its mean and variance are equal, and it is the usual starting model for counts such as admissions per night or typing errors per page.
Poisson regression
Also called: Poisson model
Poisson regression models a count outcome, such as the number of hospital visits, from predictors, with coefficients that exponentiate to rate ratios. An offset allows for different lengths of follow-up or sizes of population. It assumes the variance equals the mean, and when counts vary more than that, overdispersion must be dealt with.
Polynomial regression
Also called: curvilinear regression
Polynomial regression models a curved relationship by adding powers of a predictor, such as its square or cube, to a linear regression. A squared term captures a single bend, such as performance that rises and then falls as anxiety increases. Centring the predictor first reduces the strong correlation between it and its powers.
Post hoc test
Also called: post-hoc test, post hoc comparison, unplanned comparison, post hoc analysis, multiple comparison test
A post hoc test compares groups after an omnibus test, typically every pair of means after an analysis of variance, while controlling the error rate across all the comparisons. Tukey's HSD, Bonferroni, Scheffé, Dunnett and Games-Howell are common choices. Many textbooks advise running them only after a significant omnibus result, though that is a convention rather than a requirement of every procedure.
Posterior distribution
Also called: posterior, posterior probability
A posterior distribution describes how plausible each value of a parameter is after the prior distribution has been updated with the data. It is the main result of a Bayesian analysis and is summarised by its mean or median and a credible interval.
Practical significance
Also called: clinical significance, substantive significance, practical importance
Practical significance is whether an effect is large enough to matter in the real world, as distinct from whether it is statistically significant. With a big enough sample a trivial difference becomes statistically significant, so effect sizes and confidence intervals, judged against what matters in the field, are needed to decide importance.
Principal component analysis
Also called: PCA, principal components analysis
Principal component analysis is a data-reduction method that replaces many correlated variables with a few uncorrelated components, each a weighted combination of the originals capturing as much of their variance as possible. It is the default extraction in some software and often called factor analysis, but many methodologists say it is not a true factor analysis, since it summarises variance rather than modelling hidden factors.
Prior distribution
Also called: prior, prior probability
A prior distribution states how plausible each possible value of a parameter is before the data are seen, based on earlier evidence or a deliberately vague default. Informative priors carry real prior knowledge, and weakly informative or uninformative ones very little. With small samples the choice can change the result, so checking results under other priors is recommended.
Probability
Also called: chance
Probability is a number from 0 to 1 that expresses how likely an event is, where 0 means it cannot happen and 1 means it is certain. Frequentist statistics reads it as the long-run proportion of times the event would occur, while Bayesian statistics also uses it for degrees of belief.
Probability distribution
Also called: probability density function, probability mass function, PDF
A probability distribution gives every possible value of a random variable with its probability or, for a continuous variable, a density curve whose area over an interval is the probability of falling in it. The normal, binomial, Poisson, t, chi-square and F distributions are the ones students meet most.
Propensity score
Also called: propensity score analysis, propensity score methods
A propensity score is the probability that a person received a treatment given their measured characteristics, estimated from observational data, usually with logistic regression. Matching, stratifying, weighting or adjusting on it balances those measured characteristics between treated and untreated groups, but it cannot balance anything that was not measured.
Proportional hazards assumption
Also called: PH assumption
The proportional hazards assumption is the requirement of the Cox model that the hazard ratio between groups stays constant over the whole follow-up, so that the hazard in one group is a fixed multiple of the hazard in another. Survival curves that cross suggest it fails, and it is checked with plots or tests of residuals.
Q-Q plot
Also called: quantile-quantile plot, QQ plot, normal Q-Q plot
A Q-Q plot is a graph that plots the quantiles of one distribution against the quantiles of another, so that the points fall along a straight line when the two have the same shape. Plotted against a theoretical normal distribution, as a normal probability plot, it is the usual visual check of the normality assumption.
Quartile
Also called: quartiles, first quartile, third quartile, Q1, Q3, lower quartile, upper quartile
Quartiles are the three values that cut ordered data into four equal parts: the first quartile, which is the 25th percentile, the median, and the third quartile, which is the 75th percentile. They form the box of a box plot and give the interquartile range.
R squared
Also called: R², R-squared, R2, coefficient of determination, variance explained
R squared, the coefficient of determination, is the proportion of the variation in the outcome that a regression model accounts for, from 0 to 1. It never falls when a predictor is added, even a useless one. A high value does not mean the model is correct or predicts well, and a low value can still come with important effects.
Random variable
Also called: RV, discrete random variable, continuous random variable
A random variable is a quantity whose value is determined by the outcome of a chance process, such as the number of heads in ten tosses. It is discrete when it takes countable values and continuous when it can take any value in a range, and its behaviour is described by a probability distribution.
Range
The range is the largest value minus the smallest value in a set of data. It is quick to find but rests on only two observations, both of them extremes, so a single outlier can distort it and it tends to grow as more data are collected.
Regression analysis
Also called: regression, regression model
Regression analysis is a family of methods that model how an outcome variable changes with one or more predictor variables, estimating the size of each relationship. It is used to predict, to describe associations and to adjust for other variables. Its coefficients show association, and reading them as causes needs a design that supports it.
Regression coefficient
Also called: slope, b, unstandardised coefficient, unstandardized coefficient, regression weight, partial slope
A regression coefficient is the estimated change in the outcome for a one-unit increase in a predictor, holding the model's other predictors constant. In simple regression it is the slope of the line. It is in the units of the variables, so its size depends on how each variable was measured.
Relative frequency
Also called: percentage frequency, proportion of observations
A relative frequency is a frequency divided by the total number of observations, giving the proportion or percentage of the data that falls in each category. Relative frequencies let groups of different sizes be compared fairly, such as two classes of 18 and 31 students.
Relative risk reduction
Also called: RRR
The relative risk reduction is the proportion by which an intervention lowers risk compared with the control group, calculated as one minus the risk ratio. A drug that cuts risk from 2% to 1% gives a 50% relative risk reduction but only a one percentage point absolute reduction, which is why both should be reported.
Repeated-measures ANOVA
Also called: repeated measures ANOVA, within-subjects ANOVA, RM-ANOVA, repeated-measures analysis of variance
A repeated-measures analysis of variance compares means when the same participants are measured under every condition or at several time points. Removing stable differences between people adds power, but the method assumes sphericity and drops anyone with a missing measurement, which is one reason mixed-effects models are often preferred for longitudinal data.
Researcher degrees of freedom
Also called: analytic flexibility, analytical flexibility
Researcher degrees of freedom are the many defensible choices made in collecting, cleaning and analysing data, such as which outcomes, covariates, exclusions and tests to use. The flexibility is not wrong in itself, but left undisclosed it can be exploited, deliberately or not, to reach statistical significance. Preregistration and multiverse analysis are two responses.
Residual
Also called: residual error, prediction error
A residual is the difference between an observed value and the value a model predicts for it. Residuals are what a regression minimises and what its assumptions are checked on: plotted against fitted values they should show no pattern and an even spread. Strictly, residuals are the sample estimates of the model's unobservable errors.
Risk difference
Also called: absolute risk reduction, ARR, RD, absolute risk difference
The risk difference is the risk in one group minus the risk in another, the absolute change in the chance of an event. When an intervention lowers risk it is also called the absolute risk reduction. It shows the real size of a benefit or harm, and its reciprocal gives the number needed to treat.
Risk ratio
Also called: relative risk, RR
A risk ratio, or relative risk, is the risk of an event in one group divided by the risk in another, so 0.75 means the event is a quarter less likely. It is easy to read but hides the baseline: the same ratio could describe a fall from 40% to 30% or from 2% to 1.5%, which matter very differently.
Robust statistics
Also called: robust methods, robust statistical methods, robust estimation
Robust statistics are methods that still give sensible results when assumptions are moderately violated or a few extreme values are present. The median and the trimmed mean are robust averages, the interquartile range is a robust measure of spread, and Welch's t-test is robust to unequal variances where Student's is not.
Sampling distribution
Also called: sampling distribution of the mean
A sampling distribution is the distribution a statistic, such as a mean or a proportion, would have across all possible samples of the same size from a population. Its standard deviation is the standard error, and it is what confidence intervals and p-values are calculated from.
Scatter plot
Also called: scatterplot, scatter graph, scatter diagram, scattergram
A scatter plot is a graph that shows the relationship between two numeric variables by plotting each case as a point, one variable on each axis. It reveals the direction, form and strength of an association and any outliers, and should be looked at before any correlation or regression is calculated.
Scheffé test
Also called: Scheffé's method, Scheffe test, Scheffé's test
The Scheffé test is a post hoc procedure that controls the error rate across every possible contrast among group means, including complex ones such as the average of two groups against a third. That breadth makes it the most conservative of the common methods, so it is poorly suited to simple pairwise comparisons.
Scree plot
Also called: scree test
A scree plot is a line graph of the eigenvalues of successive factors or components, from largest to smallest. The point where the steep slope levels off into rubble, the scree, suggests how many factors are worth keeping, though reading the bend is partly a matter of judgement.
Seasonality
Also called: seasonal variation, seasonal effect, seasonal pattern
Seasonality is a pattern in a time series that repeats at a fixed calendar interval, such as higher admissions every winter or lower attendance every August. It must be modelled or removed before a trend or the effect of an intervention is judged, or it may be mistaken for one.
Segmented regression
Also called: segmented regression analysis, piecewise regression
Segmented regression is the usual analysis of an interrupted time series, fitting separate lines before and after an intervention to estimate the existing trend, any immediate change in level and any change in slope. The form of the expected change should be specified in advance, and the model must allow for autocorrelation, seasonality and other events at the same time.
Sensitivity analysis
A sensitivity analysis repeats an analysis under different reasonable assumptions or decisions, such as another way of handling missing data or outliers, to see whether the conclusions change. Results that survive are called robust. It is recommended whenever an analysis rests on assumptions that cannot be checked, such as how missing data arose.
Shapiro-Wilk test
Also called: Shapiro–Wilk test, Shapiro Wilk test, S-W test
The Shapiro-Wilk test checks whether a sample could have come from a normal distribution, a significant result suggesting that it did not. It has more power than the Kolmogorov-Smirnov test and is widely recommended. In large samples it flags trivial departures from normality, so plots should be read alongside it.
Sign test
The sign test is a simple non-parametric test for paired data, or for one sample against a set median, that counts how many differences are positive and how many negative and tests the split with the binomial distribution. It ignores the size of the differences, so it has less power than the Wilcoxon signed-rank test but needs almost no assumptions.
Significance level
Also called: alpha, alpha level, α, level of significance
The significance level, alpha, is the threshold set before a test for rejecting the null hypothesis, and it is the Type I error rate the researcher accepts: the probability of rejecting a null hypothesis that is in fact true. Five per cent, or 0.05, is the convention, but it is a custom rather than a law, and 0.01 and 0.10 are also used.
Simple effect
Also called: simple main effect, simple effects analysis
A simple effect is the effect of one factor at a single level of another factor, such as the effect of a drug among women only. Simple effects are examined to unpack a significant interaction, showing where one factor matters and where it does not.
Simple linear regression
Also called: SLR, bivariate regression, simple regression
Simple linear regression fits a straight line through the data to describe how a continuous outcome changes with a single predictor, estimating an intercept and a slope. It assumes a linear relationship and errors that are independent, normally distributed and of constant variance, and it is usually fitted by ordinary least squares.
Skewness
Also called: skew, skewed distribution, positive skew, negative skew, right skew, left skew
Skewness is the lack of symmetry in a distribution, where one tail stretches further than the other. In positive or right skew the long tail points towards high values and the mean is usually pulled above the median, as with incomes. In negative or left skew the long tail points towards low values, as with scores on an easy test.
Sobel test
Also called: Sobel's test, delta method test of the indirect effect
The Sobel test is a test of an indirect effect in mediation analysis that divides the product of the two paths by an approximate standard error and compares the result with the normal distribution. Because that product is rarely normally distributed, the test is conservative, and bootstrapped confidence intervals are now generally preferred.
Spearman's rank correlation
Also called: Spearman's rho, Spearman correlation, Spearman's rank correlation coefficient, Spearman's ρ
Spearman's rank correlation, rho, is Pearson's correlation calculated on the ranks of the values instead of the values themselves. It measures how consistently one variable rises or falls with another, a monotonic relationship, whether or not the relationship is a straight line, and it suits ordinal data, skewed data and outliers.
Sphericity
Also called: sphericity assumption, assumption of sphericity
Sphericity is the assumption in repeated-measures analysis of variance that the variances of the differences between every pair of conditions are equal. Mauchly's test is often used to check it, though it performs poorly with non-normal data. When sphericity fails, the Greenhouse-Geisser or Huynh-Feldt correction reduces the degrees of freedom to keep the Type I error rate honest.
Standard deviation
Also called: SD, std dev, sample standard deviation
The standard deviation is the square root of the variance and shows how far values typically lie from their mean, in the data's own units. In a roughly normal distribution about two-thirds of values fall within one standard deviation of the mean. It describes the spread of data, unlike the standard error, which describes the precision of an estimate.
Standard error
Also called: SE, standard error of the mean, S.E.
The standard error is the standard deviation of a statistic's sampling distribution, and so measures how much an estimate such as a sample mean would vary from sample to sample. The standard error of the mean is the standard deviation divided by the square root of the sample size, so it shrinks as samples grow. Report it for precision and the standard deviation for spread.
Standard normal distribution
Also called: z distribution, unit normal distribution
The standard normal distribution is the normal distribution with a mean of 0 and a standard deviation of 1. Converting values to z-scores places them on it, which is how probabilities and critical values such as 1.96 are read from a single table.
Standardised mean difference
Also called: standardized mean difference, SMD
A standardised mean difference is a difference between two group means divided by a standard deviation, which puts results measured on different scales into the same units. Cohen's d, Hedges' g and Glass's delta are its main forms, and it is how meta-analyses combine studies that measured one outcome with different instruments.
Standardised regression coefficient
Also called: standardized regression coefficient, beta weight, beta coefficient, standardised beta, β
A standardised regression coefficient is the change in the outcome, in standard deviations, for a one standard deviation increase in a predictor, with the others held constant. Being free of units, it lets predictors measured on different scales be compared within one model, although its size also depends on how varied the sample happened to be.
Stationarity
Also called: stationary series, weak stationarity, stationary time series
Stationarity is the property of a time series whose mean, variance and autocorrelation stay the same over time. Many time series models require it, and a series with a trend is usually made stationary by differencing, that is, analysing the change from one period to the next.
Statistic
Also called: sample statistic
A statistic is a number calculated from a sample, such as a sample mean or a correlation, and used to describe the sample or to estimate a population parameter. Different samples give different values of it, and that variation is what its sampling distribution and standard error describe.
Statistical model
Also called: model
A statistical model is a mathematical description of how data are assumed to have been generated, combining a systematic part, such as a regression equation, with a random part describing variation around it. Every test and estimate rests on one, and a p-value or interval is only as trustworthy as the model's assumptions.
Statistical power
Also called: power, power of a test, 1 − β
Statistical power is the probability that a test will detect an effect of a given size when that effect really exists, that is, correctly reject a false null hypothesis. It depends on sample size, effect size, alpha and variability, and 80% is a common target. Studies with low power miss real effects, and their significant results are less trustworthy.
Statistical significance
Also called: statistically significant, significant result, significance
Statistical significance is the label given to a result whose p-value falls at or below the chosen significance level. It says the data would be unusual if the null hypothesis were true, not that the effect is large, important or even real. A non-significant result is not evidence of no effect, and a result just either side of 0.05 carries almost the same evidence.
Statistical software
Also called: statistics software, statistical package, statistics package, stats package
Statistical software is a program for managing, analysing and graphing data, used through menus, written commands or both. SPSS, Stata, SAS, R, JASP and jamovi are common examples. Writing and saving commands, often called syntax, instead of only clicking through menus makes an analysis reproducible and every step checkable.
Stem-and-leaf plot
Also called: stem and leaf display, stemplot, stem-and-leaf diagram
A stem-and-leaf plot is a text-based display that splits each number into a stem, usually its leading digits, and a leaf, usually its last digit, and lists the leaves beside their stems. It shows the shape of a small data set as a histogram does while keeping every original value.
Stepwise regression
Also called: stepwise multiple regression, stepwise selection
Stepwise regression lets software add or remove predictors one at a time, by forward selection, backward elimination or both, according to statistical criteria such as p-values. It is widely criticised because it capitalises on chance, gives overly optimistic results and produces models that often fail to replicate, so theory-driven model building is usually preferred.
Structural equation modelling
Also called: structural equation modeling, SEM, structural equation model, covariance structure modelling
Structural equation modelling is a family of methods that tests a whole network of relationships at once, combining a measurement model, in which latent variables are measured by observed indicators, with a structural model of the paths between them. Because it models measurement error, it can estimate relationships between constructs more accurately than regression on raw scores.
Sum of squares
Also called: SS, sum of squared deviations, sums of squares
A sum of squares is the total of the squared deviations of values from a mean or from a model's predictions. Analysis of variance and regression split the total sum of squares into a part explained by the model and a residual part, and dividing a sum of squares by its degrees of freedom gives a mean square.
Survival analysis
Also called: time-to-event analysis, event history analysis, duration analysis, failure time analysis
Survival analysis is a set of methods for data in which the outcome is the time until an event, such as death, relapse or leaving a course. Its special feature is handling censored cases, people whose event had not happened by the end of follow-up, whose information ordinary methods would lose or distort.
Symmetrical distribution
Also called: symmetric distribution
A symmetrical distribution is one whose two halves are mirror images about its centre, so its mean and median are equal. The normal and uniform distributions are symmetrical, and any distribution with one tail longer than the other is skewed.
t distribution
Also called: Student's t distribution, Student's t-distribution, t-distribution
The t distribution is a symmetrical, bell-shaped distribution with heavier tails than the normal, used when a mean is tested or estimated with a standard deviation taken from the sample. Its exact shape depends on the degrees of freedom, and it approaches the normal as samples grow. William Gosset published it in 1908 under the pen name Student.
t-test
Also called: t test, Student's t-test, Student t-test
A t-test is a test that compares a mean with a set value, or two means with each other, using the t distribution. The one-sample, independent-samples and paired t-tests cover the three common designs. It assumes roughly normal data or a large enough sample, and the independent version in its original Student form also assumes equal variances.
Test statistic
A test statistic is a number calculated from sample data that measures how far the data depart from what the null hypothesis predicts, usually scaled by a standard error. Examples are t, z, F and chi-square, and comparing it with its known distribution under the null hypothesis gives the p-value.
Time series analysis
Also called: time-series analysis, time series
Time series analysis is the study of a sequence of measurements of the same variable taken at regular intervals, such as monthly hospital admissions, in which the order matters and neighbouring values are usually correlated. It describes trend, seasonality and autocorrelation and is used to forecast or to judge the effect of an event.
Trimmed mean
Also called: truncated mean
A trimmed mean is the mean worked out after removing a fixed percentage of the lowest and highest values, for example 10% from each end. It is a robust average that usually falls between the mean and the median, keeping more information than the median while resisting outliers.
Tukey's HSD
Also called: Tukey's honestly significant difference, Tukey HSD, Tukey test, Tukey's range test, Tukey-Kramer test
Tukey's honestly significant difference test compares every pair of group means after an analysis of variance while holding the familywise error rate at alpha. Based on the studentised range distribution, it gives the narrowest intervals when all pairwise comparisons are wanted, and it assumes equal variances.
Two-tailed test
Also called: two-sided test, non-directional test, two tailed test, two-tailed
A two-tailed test looks for an effect in either direction, splitting the rejection region between both tails of the distribution. It is the default in most research because it can detect an unexpected effect as well as the predicted one, and it matches a non-directional hypothesis.
Two-way ANOVA
Also called: two-way analysis of variance, two-factor ANOVA, two way ANOVA
A two-way analysis of variance tests the effects of two factors on an outcome at once, such as teaching method and class size, giving a main effect for each and a test of their interaction. It is the simplest factorial analysis of variance, and a significant interaction means neither main effect can be read on its own.
Type I error
Also called: false positive, alpha error, type 1 error, error of the first kind
A Type I error is rejecting a null hypothesis that is actually true, a false positive such as concluding that a treatment works when it does not. Its probability is set by the significance level, and the chance of at least one rises quickly when many tests are run without correction.
Type II error
Also called: false negative, beta error, type 2 error, error of the second kind
A Type II error is failing to reject a null hypothesis that is actually false, a false negative such as missing a treatment's real benefit. Its probability, beta, falls as the sample size, the true effect or alpha increase, and one minus beta is the test's power.
Uniform distribution
Also called: rectangular distribution, flat distribution
A uniform distribution is one in which every value in a range is equally likely, so its graph is flat. The roll of a fair die follows a discrete uniform distribution, and a random number between 0 and 1 generated by software follows a continuous one.
Variability
Also called: dispersion, spread, measure of dispersion, measure of variability, measure of spread, scatter
Variability is how spread out the values in a set of data are, and a measure of variability summarises it in one number. The range, the interquartile range, the variance and the standard deviation are the usual measures, each suited to different kinds of data and distributions.
Variance
Also called: sample variance, population variance, s squared
The variance is the average squared distance of the values from their mean. A sample variance divides the sum of squared deviations by n minus 1 rather than n, which makes it an unbiased estimate of the population variance. Because it is in squared units it is usually reported through its square root, the standard deviation.
Variance inflation factor
Also called: VIF
The variance inflation factor measures how much the variance of a regression coefficient is inflated because its predictor correlates with the other predictors. A value of 1 means no inflation. A common rule of thumb treats values above 4 as worth investigating and above 10 as serious multicollinearity, though these cut-offs are conventions.
Weighted mean
Also called: weighted average
A weighted mean is an average in which some values count more than others: each value is multiplied by a weight and the total is divided by the sum of the weights. A module mark built from coursework worth 40% and an exam worth 60% is a weighted mean.
Welch's ANOVA
Also called: Welch's F-test, Welch's F test, Welch ANOVA, Welch's W test
Welch's analysis of variance is a version of the one-way analysis of variance that does not assume the groups have equal variances. Simulations show the classical F-test can give badly wrong error rates when variances differ, and Welch's version has been argued for as the default, with Games-Howell comparisons to follow.
Welch's t-test
Also called: Welch t-test, Welch's test, unequal variances t-test, Aspin-Welch t-test, Welch-Satterthwaite t-test
Welch's t-test compares the means of two independent groups without assuming equal variances, adjusting the degrees of freedom to suit. Student's t-test gives misleading error rates when variances differ, especially if the smaller group has the larger variance, while Welch's loses little power when they are equal, so it is widely recommended as the default.
Wilcoxon signed-rank test
Also called: Wilcoxon signed rank test, Wilcoxon matched-pairs signed-rank test, signed-rank test
The Wilcoxon signed-rank test is a non-parametric test for paired data, or for one sample against a set value, that ranks the sizes of the differences and compares the ranks of the positive and negative ones. It is the rank-based alternative to the paired t-test and uses more information than the sign test. It tests the median difference only if the differences are symmetrical.
Yates's continuity correction
Also called: Yates correction, Yates' correction, continuity correction, Yates's correction for continuity
Yates's continuity correction is an adjustment to the chi-square test for a two-by-two table that subtracts 0.5 from each absolute difference between observed and expected counts before squaring. It makes the test more conservative when counts are small, and statistical software often prints it beside the uncorrected result.
Z-score
Also called: z score, standard score, standardised score, standardized score
A z-score is the number of standard deviations a value lies above or below the mean, found by subtracting the mean and dividing by the standard deviation. Standardising puts measures on different scales onto a common one, so a pupil's marks in two differently scored tests can be compared, but it does not make a skewed variable normal.
Z-test
Also called: z test
A z-test is a hypothesis test whose test statistic follows the standard normal distribution when the null hypothesis is true. It is used for means when the population standard deviation is known, which is rare, and more often for proportions in large samples, such as comparing two response rates.
Where these definitions were checked
Barbara Illowsky and Susan Dean, OpenStax, Introductory Statistics 2e, Chapter 2 key terms (Descriptive Statistics)
Barbara Illowsky and Susan Dean, OpenStax, Introductory Statistics 2e, Chapter 4 key terms (Discrete Random Variables)
Barbara Illowsky and Susan Dean, OpenStax, Introductory Statistics 2e, Chapter 6 key terms (The Normal Distribution)
Barbara Illowsky and Susan Dean, OpenStax, Introductory Statistics 2e, Chapter 7 key terms (The Central Limit Theorem)
Barbara Illowsky and Susan Dean, OpenStax, Introductory Statistics 2e, 7.3 Using the Central Limit Theorem
Barbara Illowsky and Susan Dean, OpenStax, Introductory Statistics 2e, Chapter 8 key terms (Confidence Intervals)
Barbara Illowsky and Susan Dean, OpenStax, Introductory Statistics 2e, Chapter 9 key terms (Hypothesis Testing with One Sample)
Barbara Illowsky and Susan Dean, OpenStax, Introductory Statistics 2e, 10.1 Two Population Means with Unknown Standard Deviations
Barbara Illowsky and Susan Dean, OpenStax, Introductory Statistics 2e, 11.1 Facts About the Chi-Square Distribution
Barbara Illowsky and Susan Dean, OpenStax, Introductory Statistics 2e, Chapter 13 key terms (F Distribution and One-Way ANOVA)
Barbara Illowsky and Susan Dean, OpenStax, Introductory Statistics 2e, 13.3 Facts About the F Distribution
National Institute of Standards and Technology (NIST) and SEMATECH, NIST/SEMATECH e-Handbook of Statistical Methods, 1.1.1 What is EDA?
National Institute of Standards and Technology (NIST) and SEMATECH, NIST/SEMATECH e-Handbook of Statistical Methods, 1.3.3.24 Quantile-Quantile Plot
National Institute of Standards and Technology (NIST) and SEMATECH, NIST/SEMATECH e-Handbook of Statistical Methods, 1.3.5.6 Measures of Scale
National Institute of Standards and Technology (NIST) and SEMATECH, NIST/SEMATECH e-Handbook of Statistical Methods, 1.3.5.10 Levene Test for Equality of Variances
National Institute of Standards and Technology (NIST) and SEMATECH, NIST/SEMATECH e-Handbook of Statistical Methods, 1.3.5.11 Measures of Skewness and Kurtosis
National Institute of Standards and Technology (NIST) and SEMATECH, NIST/SEMATECH e-Handbook of Statistical Methods, 7.4.1 The Kruskal-Wallis test
National Institute of Standards and Technology (NIST) and SEMATECH, NIST/SEMATECH e-Handbook of Statistical Methods, 7.4.7 How can we make multiple comparisons?
National Institute of Standards and Technology (NIST), Dataplot Reference Manual: Friedman Test
National Institute of Standards and Technology (NIST), Dataplot Reference Manual: Coefficient of Variation
David M. Lane (project leader), Rice University, Online Statistics Education: A Multimedia Course of Study, Glossary of Statistical Terms
Department of Statistics, Penn State Eberly College of Science, STAT 500 Applied Statistics, Lesson 6: Hypothesis Testing
Department of Statistics, Penn State Eberly College of Science, STAT 500 Applied Statistics, Lesson 8: Chi-Square Test for Independence
Department of Statistics, Penn State Eberly College of Science, STAT 500 Applied Statistics, Lesson 10: Introduction to ANOVA
Department of Statistics, Penn State Eberly College of Science, STAT 500 Applied Statistics, Lesson 11: Introduction to Nonparametric Tests and Bootstrap
Department of Statistics, Penn State Eberly College of Science, STAT 501 Regression Methods, Lesson 1: Simple Linear Regression
Department of Statistics, Penn State Eberly College of Science, STAT 501 Regression Methods, Lesson 4: SLR Model Assumptions
Department of Statistics, Penn State Eberly College of Science, STAT 501 Regression Methods, Lesson 8: Categorical Predictors
Department of Statistics, Penn State Eberly College of Science, STAT 501 Regression Methods, Lesson 9: Data Transformations
Department of Statistics, Penn State Eberly College of Science, STAT 501 Regression Methods, Lesson 10: Model Building
Department of Statistics, Penn State Eberly College of Science, STAT 501 Regression Methods, Lesson 11: Influential Points
Department of Statistics, Penn State Eberly College of Science, STAT 501 Regression Methods, Lesson 12: Multicollinearity & Other Regression Pitfalls
Department of Statistics, Penn State Eberly College of Science, STAT 502 Analysis of Variance and Design of Experiments, Lesson 2: ANOVA Foundations
Department of Statistics, Penn State Eberly College of Science, STAT 502 Analysis of Variance and Design of Experiments, Lesson 5: Multi-Factor ANOVA
Department of Statistics, Penn State Eberly College of Science, STAT 502 Analysis of Variance and Design of Experiments, Lesson 6: Random Effects and Introduction to Mixed Models
Department of Statistics, Penn State Eberly College of Science, STAT 502 Analysis of Variance and Design of Experiments, Lesson 9: ANCOVA Part I
Department of Statistics, Penn State Eberly College of Science, STAT 502 Analysis of Variance and Design of Experiments, Lesson 11: Introduction to Repeated Measures
Department of Statistics, Penn State Eberly College of Science, STAT 504 Analysis of Discrete Data, Lesson 3: Two-Way Tables: Independence and Association
Department of Statistics, Penn State Eberly College of Science, STAT 504 Analysis of Discrete Data, Lesson 4: Tests for Ordinal Data and Small Samples
Department of Statistics, Penn State Eberly College of Science, STAT 504 Analysis of Discrete Data, Lesson 6: Binary Logistic Regression
Department of Statistics, Penn State Eberly College of Science, STAT 504 Analysis of Discrete Data, Lesson 8: Multinomial Logistic Regression Models
Department of Statistics, Penn State Eberly College of Science, STAT 504 Analysis of Discrete Data, Lesson 9: Poisson Regression
Department of Statistics, Penn State Eberly College of Science, STAT 504 Analysis of Discrete Data, Lesson 11: Advanced Topics I (matched pairs and McNemar’s test)
Department of Statistics, Penn State Eberly College of Science, STAT 504 Analysis of Discrete Data, Lesson 12: Advanced Topics II (generalized estimating equations)
Department of Statistics, Penn State Eberly College of Science, STAT 505 Applied Multivariate Statistical Analysis, Lesson 1: Measures of Central Tendency, Dispersion and Association
Department of Statistics, Penn State Eberly College of Science, STAT 505 Applied Multivariate Statistical Analysis, Lesson 6: Multivariate Conditional Distribution and Partial Correlation
Department of Statistics, Penn State Eberly College of Science, STAT 505 Applied Multivariate Statistical Analysis, Lesson 8: Multivariate Analysis of Variance (MANOVA)
Department of Statistics, Penn State Eberly College of Science, STAT 505 Applied Multivariate Statistical Analysis, Lesson 9: Repeated Measures Analysis
Department of Statistics, Penn State Eberly College of Science, STAT 505 Applied Multivariate Statistical Analysis, Lesson 10: Discriminant Analysis
Department of Statistics, Penn State Eberly College of Science, STAT 505 Applied Multivariate Statistical Analysis, Lesson 11: Principal Components Analysis (PCA)
Department of Statistics, Penn State Eberly College of Science, STAT 505 Applied Multivariate Statistical Analysis, Lesson 12: Factor Analysis
Department of Statistics, Penn State Eberly College of Science, STAT 505 Applied Multivariate Statistical Analysis, Lesson 14: Cluster Analysis
Department of Statistics, Penn State Eberly College of Science, STAT 510 Applied Time Series Analysis, Lesson 1: Time Series Basics
Department of Statistics, Penn State Eberly College of Science, STAT 510 Applied Time Series Analysis, Lesson 3: Identifying and Estimating ARIMA Models and Using Them to Forecast
Department of Statistics, Penn State Eberly College of Science, STAT 415 Introduction to Mathematical Statistics, Lesson 2: Estimation (Part I)
Department of Statistics, Penn State Eberly College of Science, STAT 415 Introduction to Mathematical Statistics, Lesson 4: Maximum Likelihood Estimation (Part I)
Department of Statistics, Penn State Eberly College of Science, STAT 415 Introduction to Mathematical Statistics, Lesson 10: Hypothesis Tests (Part II): power
Department of Statistics, Penn State Eberly College of Science, STAT 415 Introduction to Mathematical Statistics, Lesson 11: Bayesian Methods
Higgins JPT, Li T and Deeks JJ (editors), Cochrane, Cochrane Handbook for Systematic Reviews of Interventions, Chapter 6: Choosing effect measures and computing estimates of effect
Schünemann HJ, Vist GE, Higgins JPT and colleagues, Cochrane, Cochrane Handbook for Systematic Reviews of Interventions, Chapter 15: Interpreting results and drawing conclusions
Ronald L. Wasserstein and Nicole A. Lazar, American Statistical Association, The ASA statement on p-values: context, process, and purpose (The American Statistician, 2016)
Sander Greenland, Stephen J. Senn, Kenneth J. Rothman, John B. Carlin, Charles Poole, Steven N. Goodman and Douglas G. Altman, Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations (European Journal of Epidemiology, 2016)
Daniël Lakens, Equivalence tests: a practical primer for t tests, correlations, and meta-analyses (Social Psychological and Personality Science, 2017)
Marie Delacre, Daniël Lakens and Christophe Leys, Why psychologists should by default use Welch’s t-test instead of Student’s t-test (International Review of Social Psychology, 2017)
Marie Delacre, Christophe Leys, Youri L. Mora and Daniël Lakens, Taking parametric assumptions seriously: arguments for the use of Welch’s F-test instead of the classical F-test in one-way ANOVA (International Review of Social Psychology, 2019)
Douglas G. Altman and J. Martin Bland, Standard deviations and standard errors (BMJ, 2005)
M. M. Mukaka, A guide to appropriate use of correlation coefficient in medical research (Malawi Medical Journal, 2012)
Asghar Ghasemi and Saleh Zahediasl, Normality tests for statistical analysis: a guide for non-statisticians (International Journal of Endocrinology and Metabolism, 2012)
Anna Hart, Mann-Whitney test is not just a test of medians: differences in spread can be important (BMJ, 2001)
Francis Sahngun Nahm, Nonparametric statistical tests for the continuous data: the basic concept and the practical use (Korean Journal of Anesthesiology, 2016)
Salvatore S. Mangiafico, Rutgers Cooperative Extension, Summary and Analysis of Extension Program Evaluation in R: Mann–Whitney U test, Wilcoxon signed-rank test, Kruskal–Wallis test, Friedman test, Cochran’s Q test and permutation tests
Shi-Yi Chen, Zhe Feng and Xiaolian Yi, A general introduction to adjustment for multiple comparisons (Journal of Thoracic Disease, 2017)
GraphPad Software, GraphPad Prism Statistics Guide: How the Dunnett T3, Games and Howell, and Tamhane T2 tests work
María J. Blanca, Jaume Arnau, F. Javier García-Castro, Rafael Alarcón and Roser Bono, Repeated measures ANOVA and adjusted F-tests when sphericity is violated: which procedure is best? (Frontiers in Psychology, 2023)
The SciPy community, scipy.stats.pointbiserialr (SciPy reference documentation)
T. G. Clark, M. J. Bradburn, S. B. Love and D. G. Altman, Survival analysis part I: basic concepts and first analyses (British Journal of Cancer, 2003)
M. J. Bradburn, T. G. Clark, S. B. Love and D. G. Altman, Survival analysis part II: multivariate data analysis, an introduction to concepts and methods (British Journal of Cancer, 2003)
Centre for Multilevel Modelling, University of Bristol, What are multilevel models and why should I use them?
Tanya N. Beran and Claudio Violato, Structural equation modeling in medical research: a primer (BMC Research Notes, 2010)
Anna B. Costello and Jason W. Osborne, Best practices in exploratory factor analysis: four recommendations for getting the most from your analysis (Practical Assessment, Research and Evaluation, 2005)
Angela Sorgente, Rossella Caliciuri, Matteo Robba, Margherita Lanz and Bruno D. Zumbo, A systematic review of latent class analysis in psychology: examining the gap between guidelines and research practice (Behavior Research Methods, 2025)
David A. Kenny, Mediation
David A. Kenny, Moderation
James Lopez Bernal, Steven Cummins and Antonio Gasparrini, Interrupted time series regression for the evaluation of public health interventions: a tutorial (International Journal of Epidemiology, 2017)
Peter C. Austin, An introduction to propensity score methods for reducing the effects of confounding in observational studies (Multivariate Behavioral Research, 2011)
Michael Baiocchi, Jing Cheng and Dylan S. Small, Instrumental variable methods for causal inference (Statistics in Medicine, 2014)
Johnny van Doorn, Don van den Bergh, Udo Böhm, Eric-Jan Wagenmakers and colleagues, The JASP guidelines for conducting and reporting a Bayesian analysis (Psychonomic Bulletin and Review, 2021)
Don van Ravenzwaaij, Pete Cassey and Scott D. Brown, A simple introduction to Markov chain Monte-Carlo sampling (Psychonomic Bulletin and Review, 2018)
Stef van Buuren, Flexible Imputation of Missing Data, 2nd edition, 1.2 Concepts of MCAR, MAR and MNAR
Stef van Buuren, Flexible Imputation of Missing Data, 2nd edition, 1.3 Ad-hoc solutions
Stef van Buuren, Flexible Imputation of Missing Data, 2nd edition, 1.4 Multiple imputation in a nutshell
Yiran Dong and Chao-Ying Joanne Peng, Principled missing data methods for researchers (SpringerPlus, 2013)
Jan Van den Broeck, Solveig Argeseanu Cunningham, Roger Eeckels and Kobus Herbst, Data cleaning: detecting, diagnosing, and editing data abnormalities (PLOS Medicine, 2005)
Inter-university Consortium for Political and Social Research (ICPSR), University of Michigan, Guide to Social Science Data Preparation and Archiving, 5th edition (2012)
Andrew Gelman and Eric Loken, Columbia University, The garden of forking paths: why multiple comparisons can be a problem, even when there is no “fishing expedition” or “p-hacking” and the research hypothesis was posited ahead of time (2013)
Megan L. Head, Luke Holman, Rob Lanfear, Andrew T. Kahn and Michael D. Jennions, The extent and consequences of p-hacking in science (PLOS Biology, 2015)
Framework for Open and Reproducible Research Training (FORRT), FORRT Glossary: HARKing
Framework for Open and Reproducible Research Training (FORRT), FORRT Glossary: p-hacking
Framework for Open and Reproducible Research Training (FORRT), FORRT Glossary: Garden of forking paths
Framework for Open and Reproducible Research Training (FORRT), FORRT Glossary: Researcher degrees of freedom
Framework for Open and Reproducible Research Training (FORRT), FORRT Glossary: Multiverse analysis