Causality: What Does It Mean for One Thing to Cause Another?

Causality is not a hidden label that can be read directly from data. This article separates correlation, temporal order, common causes, counterfactuals, interventions, mechanisms, and study designs used to build a causal inference.

When we say that one thing caused another, we usually claim more than that the two appeared together. We claim that changing the first could make a difference to the second — that the relation is not merely coincidence, a shared pattern, or the result of some third factor.

This is where the difficulty begins. Data often show that two variables move together, but they do not by themselves tell us why. People taking a particular medicine may be sicker on average precisely because sicker people are more likely to receive it. More firefighters may be associated with more damage because larger fires require more firefighters. Correlation can be an important clue, but the direction of the arrow is not written into the correlation itself.

Modern causal inference therefore uses several different tools. We ask what came first, which common causes could generate a spurious association, what would happen in a comparable world without the proposed cause, what happens under an intervention, and whether we understand a mechanism that makes the link intelligible. A randomized experiment is one of the strongest tools, but it is not the only way we learn about causes.

Causality is therefore not a single test and not a single philosophical formula. It is a disciplined way of distinguishing prediction from change, accompaniment from production, and a story that is merely possible from an inference supported by data, study design, and knowledge of the system.

Correlation starts the question; it does not finish the answer

Correlation tells us that two variables are statistically associated. When X changes, Y often changes as well. But the same pattern can arise from very different structures. X may cause Y, Y may cause X, both may be caused by Z, the association may reflect sampling or measurement, and in smaller datasets chance alone can create a pattern.

The phrase “correlation does not imply causation” is therefore not a reason to ignore correlations. Without patterns in data, we would often not know what to investigate. The more accurate lesson is that correlation alone does not determine the causal explanation. Moving from an observed pattern to a cause requires additional assumptions and evidence.

It is also crucial to distinguish prediction from intervention. A barometer can predict a change in weather, but moving the barometer needle will not cause a storm. A variable may be an excellent predictor without being a useful target for action. Causal inference asks precisely about this difference: what would happen to the outcome if we actually changed the proposed cause?

Temporal order: what came first?

For ordinary empirical causal claims, temporal order is a basic constraint. The proposed cause must occur before the effect, or a change in the cause must precede a change in the outcome. If an early stage of disease changes diet before either is measured, we may wrongly conclude that diet caused the disease even though part of the association runs in the opposite direction.

Temporal precedence alone, however, proves very little. A rooster crows before sunrise but does not cause it. An alarm may sound before firefighters arrive without being the cause of the fire. Temporal order mainly helps us rule out some explanations and detect the possibility of reverse causation.

In fast, feedback-driven, or tightly coupled systems, order can be difficult to establish. Economies, behavior, biology, and social systems often contain loops in which X affects Y and Y later feeds back into X. The question “what came first?” therefore needs adequate temporal resolution and a clear model of the process.

The common cause: how a false arrow can appear

One of the most common sources of mistaken causal inference is confounding. A third variable influences both the proposed cause and the outcome, creating or distorting the association between them. Rain, for example, increases both umbrella use and wet roads; if rain is omitted, umbrellas and wet roads may be strongly associated even though umbrellas do not cause the road to become wet.

Medicine has a classic version called confounding by indication: sicker people are more likely to receive stronger treatment. A simple comparison of treated and untreated patients can therefore make an effective treatment look harmful. Similar problems arise whenever people, institutions, or systems select exposures according to risk, need, or expected outcome.

The solution is not simply to “control for everything.” Some variables are mediators on the pathway from cause to effect, while others are common consequences of multiple causes; inappropriate adjustment can remove part of a real effect or even introduce new bias. Causal diagrams are valuable mainly because they force us to state our assumptions about the direction of relationships before running the analysis.

Causal diagram in which a common factor Z affects X and Y, shown before and after control of the common cause.
A causal diagram shows why an association between X and Y does not by itself determine the direction of causation. A common factor Z can create or distort the observed association; whether adjustment is appropriate depends on the assumed causal structure, not on a rule to “control for everything.” Image: Marcos M. López de Prado / Wikimedia Commons CC0 1.0 / Public Domain Dedication

The counterfactual: what would have happened without the cause?

A powerful intuition about causation is counterfactual: if X had not happened, would Y still have happened? If removing X changes the answer, we have reason to regard X as causally relevant. Modern statistical theory often formulates this using potential outcomes: for the same unit, we imagine the outcome under one exposure and the outcome under another.

The difficulty is immediate. For the same person, country, cell, or planet, we cannot observe both histories at the same moment. We observe only one realized path. Causal effects are therefore tied to the problem of the missing counterfactual — how to estimate the unobserved outcome credibly.

Comparing with another person or group works only if it is a sufficiently good stand-in for that missing counterfactual. This is where study design, randomization, natural experiments, replication, and statistical adjustment enter. They do not literally create a second world; they try to construct a comparison that approximates it well enough.

Intervention: what happens if we deliberately change X?

The interventionist perspective asks a practical question: if we changed X while not simultaneously altering other relevant pathways by fiat, would Y change? This distinguishes observation from action. Seeing that people with a particular characteristic have a different outcome is not the same as knowing what would happen if that characteristic were changed.

A randomized controlled trial is powerful because random assignment, on average, breaks systematic links between the assigned intervention and prior causes of the outcome. The groups should therefore differ mainly in the assigned intervention, allowing a cleaner estimate of its effect.

Randomization is not magic. Small samples can still have chance imbalances, participants may not follow assigned treatment, outcomes can be mismeasured, people can drop out, and a result in a narrow study population may not generalize to another population. A good experiment greatly strengthens causal inference, but it does not eliminate the need for judgment.

Simplified event tree with branching paths from an initial state through assigned options to different outcomes.
A simplified event tree illustrates that study design determines which comparisons among paths and outcomes are meaningful. Branching alone does not prove causation; the assumptions about assignment, comparability, and outcome measurement do the real inferential work. Image: JackStorrorCarter / Wikimedia Commons CC0 1.0 / Public Domain Dedication

Mechanism: how is the cause supposed to produce the effect?

A mechanistic explanation identifies intermediate processes through which a cause contributes to an effect. In biology this may involve receptors, signaling pathways, and changes in cell function; in engineering, transmission of force or energy; in social systems, incentives, information, rules, and behavioral responses. Mechanisms help explain why a relation is more than a statistical pattern.

But a mechanism is not independent proof when it is only a plausible story. Almost any pattern can be given a persuasive narrative after the fact. A strong mechanistic argument therefore needs independent support: measurable intermediate steps, the correct temporal order, response to intervention, or other predictions that could have failed.

The reverse also matters: not knowing the mechanism does not prove that causation is absent. Bradford Hill already emphasized in epidemiological reasoning that laboratory and mechanistic evidence can greatly strengthen a causal hypothesis, yet lack of such evidence should not be turned into an absolute veto. Science can identify an effect reliably before the detailed mechanism is fully understood.

When experiments are impossible: observation, natural experiments, and triangulation

Many important questions cannot ethically or practically be randomized. We cannot randomly assign people to decades of smoking, an earthquake, poverty, or dangerous pollution. In such cases causal inference does not disappear; it becomes more dependent on the quality of design and on how well alternative explanations are controlled.

Researchers use cohort and other observational studies, natural experiments, instrumental variables, discontinuities, policy changes, and other quasi-experimental approaches. Every method carries assumptions. The key question is not whether a study is labeled “observational” or “experimental,” but which biasing pathways its design closes and which remain open.

Triangulation is therefore often especially powerful: different methods, populations, measurements, and data sources converge on a similar conclusion while their main weaknesses are not the same. Agreement across independent evidential routes is more persuasive than ten nearly identical analyses that share the same hidden bias.

Causes need not be necessary, sufficient, or deterministic

Everyday language often treats a cause like a switch: if the cause occurs, the effect follows. In many systems, however, causes only change probabilities. Smoking increases the risk of lung cancer, yet not every smoker develops the disease and lung cancer also occurs in non-smokers. This is not an objection to causation; it is a feature of multicausal and probabilistic systems.

One factor can be part of several different pathways to the same outcome. Infection, genetic vulnerability, environment, and behavior can jointly shape risk; removing one factor may reduce risk without eliminating it. The question “is X a cause?” therefore often needs qualification: in which population, under which conditions, for which outcome, and with what effect size?

The terms “necessary” and “sufficient” are also stricter than the ordinary word cause. Something can be a genuine causal factor without being sufficient on its own or necessary for every instance of the effect. This blocks a common false argument: “Y sometimes happens without X, therefore X cannot cause Y.”

A cause in a population is not the same as the cause of one event

The statement “X increases the risk of Y” concerns a general or population-level causal effect. The statement “X caused this particular Y” concerns an individual event. The second question is often harder because we cannot directly observe the alternative history of a single case.

This distinction matters in medicine, history, law, and everyday life. If a treatment is known to reduce risk in a population, we still cannot always say with certainty that the treatment saved a particular individual. If an exposure raises disease risk, that population effect does not automatically show that it was the sole cause of disease in one person.

General and individual causation are not unrelated. Population data, timing, mechanistic evidence, and case-specific facts can be combined to assess individual causation. But the degree of certainty and the type of evidence required are not always the same as when estimating an average effect in a group.

How to evaluate a causal claim in practice

When someone claims that X causes Y, first ask what evidence is actually available. Is there only a correlation? Does X precede Y? Is there a plausible common cause or reverse direction? How were groups selected? What was measured and what was not? Was an intervention randomized, or is there a natural experiment that approximates a credible counterfactual?

Then ask whether different lines of evidence point in the same direction. Does the effect persist across populations and analyses? Does it respond to intervention? Is the effect size meaningful? Do we understand a mechanism and its intermediate steps? Is there an alternative explanation that accounts for more of the data with fewer extra assumptions?

Bradford Hill presented his famous aspects of association as aids to judgment, not as a set of hard rules that mechanically prove causation. Modern causal models are more formal, but they preserve the same methodological lesson: calculation cannot replace the question of what causal world the model assumes.

The safest conclusion is therefore modest but powerful. We usually do not “see” causality directly in a single number. We build a causal inference from temporal structure, credible counterfactual comparisons, interventions, mechanisms, models, and independent evidence. A good causal conclusion is not perfect certainty; it is an explanation that has survived a serious attempt to replace it with a better one.

Sources and further reading

  1. THY-REALITY — Dejstvo, interpretacija, hipoteza in špekulacija niso isto / Fact, Interpretation, Hypothesis and Speculation Are Not the Same (LOCKED): methodological separation between observation and explanatory claim.
  2. THY-REALITY — Kako oceniti vir / How to Evaluate a Source (LOCKED): source quality, independence, triangulation and evidential fit.
  3. THY-REALITY — Vzrok in posledica: dejanja niso brez posledic / Cause and Effect: Actions Have Consequences (LOCKED): distinction between causal contribution, risk, responsibility and moral evaluation.
  4. THY-REALITY — Kako vemo, da nekaj vemo? Dokaz, prepričanje in gotovost / How Do We Know What We Know? Evidence, Belief, and Certainty (LOCKED): degrees of evidence and calibrated confidence.
  5. THY-REALITY — Kako dokazujemo skrite operacije: od suma do dokumenta / How We Prove Hidden Operations: From Suspicion to Document (LOCKED): convergence of independent evidence and avoidance of single-clue inference.
  6. Woodward, J. — Causation and Manipulability. Stanford Encyclopedia of Philosophy, substantive revision 2023. Interventionist accounts, structural equations, causal graphs and limits of manipulability theories.
  7. Menzies, P.; Beebee, H. — Counterfactual Theories of Causation. Stanford Encyclopedia of Philosophy, substantive revision 2024. Counterfactual dependence, actual causation and structural-equation approaches.
  8. Hitchcock, C. — Causal Models. Stanford Encyclopedia of Philosophy. Structural causal models, interventions, counterfactuals, type and token causation.
  9. Craver, C. F.; Tabery, J. — Mechanisms in Science. Stanford Encyclopedia of Philosophy, substantive revision 2024. Mechanistic explanation, organization and the role of mechanisms in scientific reasoning.
  10. Hernán, M. A.; Robins, J. M. — Causal Inference: What If. Chapman & Hall/CRC, 2020; continuously updated online edition. Counterfactual outcomes, exchangeability, randomization, confounding, causal diagrams and observational causal inference.
  11. Hill, A. B. — The Environment and Disease: Association or Causation? Proceedings of the Royal Society of Medicine 58(5), 295–300 (1965). Temporality, experiment, plausibility and other viewpoints as aids to causal judgment rather than hard proof rules.
  12. Hitchcock, C. — Probabilistic Causation. Stanford Encyclopedia of Philosophy, archived 2022 edition. Causes can alter probabilities without deterministically guaranteeing their effects.
  13. Frisch, M. — Causation in Physics. Stanford Encyclopedia of Philosophy. Causal models, interventionist approaches and limits of applying causal language in fundamental physics.