Neurofactor
Back to the knowledge base
Ethics and research quality

Correlation versus causation: meaning and application

Correlation versus causation: why a link is not yet a cause, which research designs support causal claims and how to phrase your findings carefully.

Martijn den Otter 9 min read10/2/2026
Correlation versus causation: meaning and application

You see in your data that customers who use an app stay longer, or that employees with a mentor are less likely to leave. Such links quickly read as cause and effect. Whether that reading holds depends on how the data came about and which other explanations you have ruled out.

Correlation versus causation is the distinction between two characteristics that occur together and one characteristic that causes the other.

A correlation can arise by chance, because the effect influences the cause, through a shared third factor or through how people ended up in your data. A causal claim therefore needs a suitable design, such as a randomised experiment, or explicit and defensible assumptions. Never phrase your finding more strongly than that design allows.

What does correlation versus causation mean?

There is an association or correlation when two characteristics vary together: where one occurs more often, the other occurs more often too, or less often. In quantitative research you express this as a correlation or a difference between groups. In qualitative research you see it as patterns: participants who mention one thing often mention the other.

There is causation when a change in one characteristic brings about a change in the other. A causal claim is therefore about what happens when you intervene. A correlation does not answer that question. "Correlation is not causation" does not tell you what to do instead; that is what this article is about.

Where does the distinction come from?

Shadish, Cook and Campbell summarise a classic analysis they attribute to John Stuart Mill: a causal relationship exists if the cause preceded the effect, if cause and effect are related and if no other plausible explanation for the effect can be found. A correlation does not prove causation, they write, because you often do not know which characteristic came first or whether alternative explanations exist (Shadish, Cook & Campbell, 2002).

In 1965 the epidemiologist Austin Bradford Hill asked which aspects of an association you should consider before deciding that causation is the most likely interpretation. He named nine viewpoints, including the strength of the association, consistency across studies, the order in time and experimental evidence (Hill, 1965).

Why an association is not yet a cause

Besides a causal effect, there are at least four other explanations for a correlation.

  1. Chance. Particularly in small samples, links appear that vanish on repetition.
  2. Reverse causation. The presumed effect influences the presumed cause. Hill described the order in time as the question of "which is the cart and which is the horse" (Hill, 1965).
  3. Confounding. A third factor, a confounder, influences both characteristics. Rohrer's example is intelligence, which can influence both educational attainment and income and so explain part of their link (Rohrer, 2018).
  4. Selection effects. The link arises from who is in your data. In Rohrer's example, if publication depends on both rigour and innovativeness, the two can be negatively related among published studies, even if they are unrelated in reality. Elwert and Winship call this endogenous selection bias: conditioning on a so-called collider, a variable influenced by two other variables (Elwert & Winship, 2014).

Correlation, causation and related concepts

Concept What it describes Typical research question What it does not say
Correlation Two characteristics occur or vary together Do app users stay longer on average? Which characteristic influences the other
Causation A change in one characteristic brings about a change in the other Do members stay longer if you actively offer the app? Whether the effect applies to every individual
Confounder A third factor that influences both characteristics Does engagement explain both app use and membership length? That the original link is zero
Reverse causation The presumed effect influences the cause Do members who already want to cancel stop using the app first? That there is no effect in the other direction
Selection effect The link arises from who is in the data Do we only look at members who stayed the first months? What the link is in the whole group

Which research designs support causal claims?

Designs differ in the claims they support:

  • Randomised experiment. You assign people to conditions at random. If carried out well, the groups are comparable on average, so differences in outcome are likely due to the treatment (Shadish, Cook & Campbell, 2002). For digital products this is often called an A/B test or online controlled experiment (Kohavi, Tang & Xu, 2020). See also testable messages.
  • Quasi-experiment. There is an intervention but no random assignment: people choose for themselves or are assigned by an organisation. Alternative explanations must then be ruled out separately, for example with before-and-after measurements and a comparison group.
  • Observational research. You measure without intervening. A causal conclusion then rests entirely on assumptions about confounders, order and selection.
  • Descriptive and qualitative research. Suited to discovering links, meanings and processes and to forming hypotheses.

Making causal assumptions explicit

Judea Pearl stresses that causal claims cannot be substantiated from associations alone: behind every causal conclusion lies a causal assumption that cannot be tested in observational data (Pearl, 2009). A common tool is a causal diagram (directed acyclic graph, DAG): a chart with arrows for the causal relationships you assume.

Rohrer shows why this helps. Adjusting for a confounder can remove a spurious link; adjusting for a collider can create one. And adjusting for a mediator, an intermediate step through which the cause works, removes the very process you wanted to study. "Include as many variables as possible" is therefore not a safe strategy. Rohrer advises thinking these paths through before data collection, and notes that careful wording does not stop readers from drawing causal conclusions anyway (Rohrer, 2018).

How do you weigh evidence when an experiment is not possible?

When an experiment is impossible or irresponsible, Hill's nine viewpoints help you weigh the evidence: strength, consistency, specificity, temporality, a dose-response relationship (biological gradient), plausibility, coherence with existing knowledge, experimental evidence and analogy. Hill wrote that none of these viewpoints can bring indisputable evidence for or against a cause-and-effect hypothesis and none is a strict requirement. In his view, tests of significance do not answer those questions either (Hill, 1965).

Use them as an aid to thinking, not a checklist. A strong, consistent link makes a causal explanation more plausible but does not rule out a confounder.

Correlation and causation in qualitative research

Maxwell distinguishes two ways of thinking about causation. Variance theory works with variables and the correlations among them. Process theory looks at events and the processes that connect them. In his view, qualitative methods can often investigate such processes directly but face validity threats of their own; the researcher must actively identify and test alternative explanations (Maxwell, 2004).

Interviews can show how participants experience a choice and which steps they see between trigger and behaviour. They do not produce an effect size for a population. And a cause that participants name is their explanation, not automatically the actual mechanism. Combining methods and sources strengthens an interpretation, but agreement between sources is not itself causal evidence.

What does this mean for association and behavioural research?

Much target group and association research is observational: you measure which meanings people attach to a choice and how that relates to their behaviour or preference. See How do you measure associations? and What are association clusters?. Such patterns help you understand and form hypotheses. They do not show whether an association causes behaviour or behaviour shapes the association (editorial application).

Use a link as the starting point for a testable question, not as a proven lever.

Phrasing findings carefully

Match your wording to your design (editorial advice).

Design Appropriate wording Wording that is too strong
Observational "Members who use the app stay longer on average." "The app makes members stay longer."
Observational with an explicit causal model "Under the assumptions in our diagram, the analysis points to an effect of the app." "It is proven that the app works."
Randomised experiment "In this test, members who received the new introduction stayed longer on average than the control group." "The introduction works for every member."
Qualitative "Participants describe the app as support for keeping going." "The app explains why members stay."

Also state population, period, size of the difference and uncertainty.

Step by step: from association to a responsible claim

  1. Formulate the causal question. Which intervention, which outcome?
  2. Draw a causal diagram. Map confounders, possible reverse paths, mediators and points of selection.
  3. Check order and selection. Did the cause come first, and who is missing from your data?
  4. Choose a fitting design. An experiment where possible, otherwise a quasi-experiment or an explicitly justified observational analysis.
  5. Explore the process qualitatively. How would the effect work, and which alternatives do participants raise?
  6. Phrase it appropriately. State design, assumptions and limits in the same sentence as the finding.

Fictional example: a gym chain and its app

This example is fictional and intended as an illustration.

Situation. A gym chain sees in its membership records that members who use the app stay longer on average than those who do not. Management is considering enrolling all new members via the app.

Decision question. Does the app extend membership, or do members who would have stayed anyway use the app more often?

Available information. Observational data only; members decide for themselves whether to install the app.

Suitable approach. The team first draws a diagram. Engagement and training experience can influence both app use and membership length (confounding). Members already considering cancelling may stop using the app first (reverse causation). An analysis of only members who made it through the first three months may create a selection effect. It then sets up a randomised test: half of new members receive a personal app introduction at sign-up, the other half the usual welcome. Interviews clarify how members use the app.

Possible interpretation. The test measures the effect of the introduction, not directly that of app use itself. No results are known in this example.

Next step. Only after the test does the chain decide on a wider roll-out, phrased within the limits of the design.

Common mistakes and limits

  • Presenting a link as a lever. You then invest in something that may merely move along.
  • Adjusting for everything. Adjusting for a mediator or a collider can actually bias your estimate.
  • Forgetting selection. An analysis of only active customers or respondents can create or hide links.
  • Treating cautious words as the solution. "Is associated with" in the text and causal advice in the conclusion contradict each other.
  • Applying group data to individuals. Even an experimentally demonstrated average effect does not tell you what one person will do, nor does it automatically hold elsewhere.

Conclusion

A correlation is a valuable starting point for research. A causal claim needs more: a design that rules out alternative explanations, or explicit assumptions you can justify. Make your causal model visible, choose the design that fits your question and do not phrase your finding more strongly than that design allows.

Key terms

correlation versus causation
Correlation versus causation is the distinction between two characteristics that occur together and one characteristic that causes the other.
confounder
A third factor that influences both the presumed cause and the outcome, so that a link can appear or seem larger than the actual effect.
selection effect (endogenous selection bias)
A link that arises because you analyse within a selection influenced by both characteristics, in other words by conditioning on a collider.
causal diagram (DAG)
A chart with arrows (directed acyclic graph) for the causal relationships you assume, used to decide which variables to include and which not.

Frequently asked questions

What does correlation versus causation mean?

It is the distinction between two characteristics that occur together and one characteristic that causes the other. A correlation shows that there is a link. A causal claim says what happens when you change the cause, and therefore needs a suitable design or explicit assumptions.

Which explanations exist for a correlation without a cause?

The main ones are chance, reverse causation, confounding by a third factor and selection effects. With a selection effect, the link arises because you look at only some people, for example customers who already stayed.

Which research design supports a causal claim?

A well-conducted randomised experiment supports a causal claim most strongly, because groups start out comparable on average. Quasi-experiments and observational analyses can also provide causal evidence, but only with explicit and defensible assumptions about alternative explanations.

What is a confounder?

A confounder is a third factor that influences both the presumed cause and the outcome. As a result, a link can appear or seem larger than the actual effect. A causal diagram helps you decide which factors to adjust for and which not.

Can qualitative research say anything about causes?

Yes, qualitative research can clarify processes: how and through which steps something may work in a particular context. It does not produce an estimated effect size for a population, and a cause that participants name themselves is their explanation, not automatically the actual mechanism.

How do you phrase a finding from observational research?

Describe what you saw, for whom and when, for example "members who use the app stay longer on average". Avoid "causes" or "leads to" unless your design or explicit assumptions support it, and state the size and uncertainty of the difference.

Sources

  1. 1.Hill (1965). The environment and disease: Association or causation?. - Proceedings of the Royal Society of Medicine, 58(5), 295–300 (1965)
  2. 2.Shadish, Cook & Campbell (2002). Experimental and quasi-experimental designs for generalized causal inference. - Houghton Mifflin, Boston (boek, xxi + 623 pp.) (2002)
  3. 3.Pearl (2009). Causal inference in statistics: An overview. - Statistics Surveys, 3, 96–146 (2009)
  4. 4.Rohrer (2018). Thinking clearly about correlations and causation: Graphical causal models for observational data. - Advances in Methods and Practices in Psychological Science, 1(1), 27–42 (2018)
  5. 5.Elwert & Winship (2014). Endogenous selection bias: The problem of conditioning on a collider variable. - Annual Review of Sociology, 40, 31–53 (2014)
  6. 6.Maxwell (2004). Using qualitative methods for causal explanation. - Field Methods, 16(3), 243–264 (2004)
  7. 7.Kohavi, Tang & Xu (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. - Cambridge University Press (boek) (2020)

Related topics

Reviewed by: Martijn den Otter · Last reviewed: 10/2/2026

Martijn den Otter

Martijn den Otter

Oprichter van Neurofactor. Expert in neuromarketing en consumentenpsychologie.

LinkedIn →