Johansen’s Test: Simple Definition

Cointegration > Johansen’s test is a way to determine if three or more time series are cointegrated. More specifically, it assesses the validity of a cointegrating relationship, using a maximum likelihood estimates (MLE) approach. It is also used to find the number of relationships and as a tool to estimating those relationships (Wee & Tan, … Read more


Comments? Need to post a correction? Please Contact Us.

Wording Bias

Bias > Wording bias, also called question-wording bias or “leading on the reader” (Gerver & Sgroi, 2017) happens in a survey when the wording of a question systematically influences the responses (Hinders, 2019). Examples of Wording Bias If a survey question asks if people agree that “Welfare helps people to get back on their feet,” … Read more


Comments? Need to post a correction? Please Contact Us.

Pooled Variance

Statistics Definitions > Pooled variance (also called combined, composite, or overall variance) is a way to estimate common variance when you believe that different populations have the same variances. The pooled sample variance formula is: Where: n = the sample size for the first sample, m = the sample size for the second sample, S2x … Read more


Comments? Need to post a correction? Please Contact Us.

Pruning Statistics

Statistics Definitions > What is Pruning? Pruning removes parts of a model that are non-predictive. The process discards statistical noise, reducing the model’s size and usually improving its accuracy. Pruning is often necessary because the number of potential subtrees grows as a function of the size of the tree. Tree pruning algorithms will repeatedly delete … Read more


Comments? Need to post a correction? Please Contact Us.

Fuzzy Statistics: Simple Definition, Examples

Statistics Definitions > Fuzzy statistics usually refers to a combination of fuzzy set theory—the treatment of ambiguous, imprecise, or subjective data—and traditional stats methods. It’s a very loose term that isn’t very well defined; Technically, it could apply to anything to do with fuzzy sets and statistics. Therefore, you’ll want to pay close attention to … Read more


Comments? Need to post a correction? Please Contact Us.

Plausibility, Plausible Values and Measures

Statistics Definitions > In a broad sense, plausibility is usually used as another name for “reasonable.” Let’s say you wanted to check a normal model, obtained from a sample, to see if it is a reasonable model for the population. An effective way to check the assumption that a normal model is plausible is to … Read more


Comments? Need to post a correction? Please Contact Us.

Reservoir Sampling

Sampling > Reservoir sampling is a quota-based random sampling method, used to get a particular sample size when you don’t know the population size (i.e. when you’re dealing with a data stream of unknown length). It can also be used to create a sample for very large data sets. It’s called reservoir sampling because the … Read more


Comments? Need to post a correction? Please Contact Us.

Oversampling in Statistics

Sampling > Oversampling in Statistics In statistics, oversampling involves taking higher, disproportionate samples than would otherwise be collected with random sampling. Depending on the structure of the survey or poll, oversampling might result in bringing low numbers of an underrepresented minority group to a representative proportion. In some cases, it might result in the minorities … Read more


Comments? Need to post a correction? Please Contact Us.

Collider Variable: Definition

Descriptive Statistics > What is a Collider Variable? Graphically, a node on a causal graph (a type of directed acyclic graph) is a collider variable if the path entering and exiting the node both have arrows pointing into it. Essentially, the paths “collide.” Nodes that don’t meet this definition are called noncolliders. It’s possible for … Read more


Comments? Need to post a correction? Please Contact Us.

Absolute Frequency: Definition, Examples

Descriptive Statistics > Absolute frequency is a simple count of the number of cases, items, or things. An absolute frequency distribution displays those counts, usually in a table. Creating an Absolute Frequency Distribution The set up for an absolute frequency distribution is simple: Create Two Columns. Enter the data you want to track in the … Read more


Comments? Need to post a correction? Please Contact Us.