Correlation & Causation

TipLearning Objectives

By the end of this lesson, you’ll be able to:

  • Distinguish correlation from causation.
  • Interpret the correlation coefficient \(r\) qualitatively.
  • Identify positive, negative, strong, and weak linear correlations.
  • Recognize lurking or confounding variables that may explain an association.
  • Distinguish observational studies from experiments when evaluating causal claims.
  • Recognize misleading conclusions drawn from correlated data.

Key Ideas

Two variables are associated when values of one variable tend to occur with particular values of another.

When that association is linear, we can describe it using correlation.

For example:

  • More hours studied may be associated with higher test scores.
  • Higher temperatures may be associated with lower heating costs.

But an association between two variables does not automatically mean that one causes the other.

Causation means that changing one variable produces a change in another variable.

The key principle is:

Correlation does not imply causation.

Even a very strong correlation does not, by itself, prove that one variable causes the other.


The Correlation Coefficient

The correlation coefficient, written as \(r\), measures the direction and strength of a linear relationship between two numerical variables.

Its value is always between:

\[ \boxed{-1\le r\le1} \]

The sign tells the direction:

  • \(r>0\) → positive linear correlation
  • \(r<0\) → negative linear correlation
  • \(r\approx0\) → little or no linear correlation

The magnitude \(|r|\) tells the strength:

  • \(|r|\) close to 1 → strong linear correlation
  • \(|r|\) close to 0 → weak linear correlation

For example:

\[ r=0.92 \]

indicates a strong positive linear correlation.

Meanwhile:

\[ r=-0.88 \]

indicates a strong negative linear correlation.

And:

\[ r=0.08 \]

indicates little or no linear correlation.

NoteThe Sign Does Not Measure Strength

A negative correlation is not automatically weaker than a positive correlation.

For example:

\[ r=-0.95 \]

represents a stronger linear relationship than:

\[ r=0.40 \]

Strength depends on \(|r|\), not whether \(r\) is positive or negative.


Correlation Measures Linear Relationships

The correlation coefficient \(r\) describes linear association.

This distinction matters.

A scatterplot can show a strong curved relationship even when:

\[ r\approx0 \]

Therefore, always look at the scatterplot when one is provided.

A small value of \(r\) does not necessarily mean that the variables have no relationship at all.

It means that they have little or no linear relationship.


Correlation vs. Causation

Suppose two variables tend to increase together.

There are several possible explanations.

Possibility 1: One Variable Causes the Other

For example:

Increasing the amount of fertilizer applied to a plant may cause the plant to grow more under appropriate conditions.

Here, a causal relationship may be plausible.


Possibility 2: A Third Variable Affects Both

Suppose:

  • ice cream sales increase
  • swimming-related incidents also increase

These variables may be positively correlated.

But buying ice cream does not cause swimming incidents.

A third variable helps explain both:

\[ \boxed{\text{warmer weather}} \]

Warmer weather can lead to:

  • more ice cream purchases
  • more people swimming

This is an example of a lurking variable.


Possibility 3: The Association May Be Coincidental

Sometimes two variables happen to move together even though there is no meaningful relationship between them.

With enough variables and enough data, some correlations can occur simply by chance.

Therefore, correlation alone is not sufficient evidence of causation.


Lurking and Confounding Variables

A lurking variable is a variable that is not included in the analysis but may help explain the observed relationship between the variables being studied.

For example, suppose data show:

Larger fires tend to have more firefighters present.

It would be incorrect to conclude:

More firefighters cause larger fires.

The lurking variable is:

\[ \boxed{\text{size or severity of the fire}} \]

Larger fires tend to:

  • cause more damage
  • require more firefighters

The underlying severity of the fire helps explain the association.

A confounding variable is a related idea: it is a variable whose effect is mixed with the effect of another variable, making it difficult to determine which variable is responsible for an observed difference.

For many test questions, the key idea is simply:

Could another variable be influencing the relationship?


Observational Studies vs. Experiments

Whether a study can support a causal conclusion depends heavily on how the data were collected.

Observational Study

In an observational study, researchers observe or measure variables without assigning treatments.

For example, researchers might record:

  • how many hours students study
  • their test scores

They may find that students who study more tend to have higher scores.

This establishes an association, but other variables could also matter, such as:

  • prior knowledge
  • motivation
  • sleep
  • course difficulty

Therefore, an observational study generally does not establish causation by itself.


Randomized Experiment

In a randomized experiment, researchers randomly assign subjects to different treatments or conditions.

Random assignment helps balance other variables between the groups.

If the groups then show meaningful differences, there is stronger evidence that the treatment itself caused the difference.

TipCausation Questions

When a question asks whether a study supports a causal conclusion, look for random assignment to treatments.

Randomized experiments provide much stronger evidence for causation than observational studies.


Common Problem Types

1. Describing Positive Correlation

Suppose:

\[ r=0.85 \]

Because:

\[ r>0 \]

the direction is positive.

Because \(|r|\) is relatively close to 1, the linear relationship is strong.

Therefore:

\[ \boxed{\text{strong positive linear correlation}} \]


2. Describing Negative Correlation

Suppose:

\[ r=-0.90 \]

The negative sign indicates that as one variable increases, the other tends to decrease.

Because:

\[ |-0.90|=0.90 \]

the relationship is strong.

Therefore:

\[ \boxed{\text{strong negative linear correlation}} \]


3. Comparing Correlation Strength

Which represents a stronger linear relationship?

\[ r=0.45 \]

or:

\[ r=-0.80 \]

Compare absolute values:

\[ |0.45|=0.45 \]

\[ |-0.80|=0.80 \]

Since:

\[ 0.80>0.45 \]

the stronger correlation is:

\[ \boxed{r=-0.80} \]

The negative sign describes direction, not weakness.


4. Recognizing Correlation Without Causation

Suppose ice cream sales and swimming-related incidents both increase during the summer.

The variables may be positively correlated.

However, it would be unreasonable to conclude that:

Buying ice cream causes swimming incidents.

A likely lurking variable is:

\[ \boxed{\text{temperature}} \]

Warmer weather contributes to both.


5. Identifying a Lurking Variable

Suppose cities with more firefighters at a fire also tend to experience more fire damage.

Does this mean firefighters cause additional damage?

No.

Larger and more severe fires generally:

  • cause more damage
  • require more firefighters

The lurking variable is:

\[ \boxed{\text{fire size or severity}} \]


6. Evaluating a Causal Claim

Suppose researchers observe that students who sleep more tend to earn higher test scores.

Can they conclude that additional sleep caused the higher scores?

Not necessarily.

Other variables may differ between the students.

The data show an:

\[ \boxed{\text{association}} \]

but do not automatically establish:

\[ \boxed{\text{causation}} \]


7. Recognizing Experimental Evidence

Suppose researchers randomly assign participants to two study methods.

One group uses Method A, and the other uses Method B.

Everything else is kept as similar as possible.

If the groups later show a meaningful difference in performance, the randomized design provides stronger evidence that the study method caused the difference.

The important feature is:

\[ \boxed{\text{random assignment}} \]


Strategies

Separate Association From Causation

When two variables are related, first say:

The variables are associated.

Do not automatically say:

One variable causes the other.


Look for Another Explanation

Ask:

Could another variable influence both variables?

If yes, that variable may help explain the observed correlation.


Use the Sign of \(r\) for Direction

Remember:

\[ r>0 \]

means positive linear correlation.

And:

\[ r<0 \]

means negative linear correlation.


Use \(|r|\) for Strength

To compare strength, ignore the sign temporarily.

For example:

\[ |-0.9|=0.9 \]

is stronger than:

\[ |0.5|=0.5 \]


Look at the Scatterplot

Do not rely only on \(r\) when a scatterplot is available.

Check for:

  • curved relationships
  • outliers
  • clusters
  • unusual patterns

These features may not be captured well by a single correlation coefficient.


Look for Random Assignment

When evaluating whether a study supports causation, ask:

Were subjects randomly assigned to treatments?

If not, be cautious about making a causal conclusion.


Worked Examples

Example 1 — Interpret a Correlation Coefficient

Suppose the correlation between hours studied and test scores is:

\[ r=0.82 \]

Describe the relationship.

Solution

First consider the sign:

\[ 0.82>0 \]

so the relationship is positive.

Next consider the magnitude:

\[ |0.82|=0.82 \]

which indicates a fairly strong linear relationship.

Therefore:

\[ \boxed{\text{strong positive linear correlation}} \]

In context:

Students who study more hours tend to have higher test scores.

This statement describes association and does not, by itself, establish causation.


Example 2 — Compare Correlations

Which relationship is stronger?

\[ r=-0.91 \]

or:

\[ r=0.63 \]

Solution

Compare absolute values:

\[ |-0.91|=0.91 \]

and:

\[ |0.63|=0.63 \]

Since:

\[ 0.91>0.63 \]

we conclude:

\[ \boxed{r=-0.91\text{ is stronger}} \]

It represents a strong negative linear correlation.


Example 3 — Ice Cream and Swimming

A city records both ice cream sales and swimming-related incidents throughout the year.

The two variables show a strong positive correlation.

Can we conclude that ice cream sales cause swimming incidents?

Solution

No.

Both variables tend to increase during warmer weather.

A likely lurking variable is:

\[ \boxed{\text{temperature}} \]

The correlation does not establish that one measured variable causes the other.


Example 4 — Firefighters and Fire Damage

Data show that fires with more firefighters present tend to have more property damage.

A person concludes:

Sending more firefighters causes more damage.

Is this conclusion justified?

Solution

No.

More severe fires generally require more firefighters and also produce more damage.

The lurking variable is:

\[ \boxed{\text{fire severity}} \]

Therefore, the observed correlation does not show that firefighters cause the damage.


Example 5 — Observational Study

Researchers survey 1,000 students and find that students who regularly eat breakfast tend to have higher test scores.

Can the researchers conclude that eating breakfast causes higher scores?

Solution

Not from this evidence alone.

This is observational data.

Students who eat breakfast may differ in other ways, such as:

  • sleep habits
  • schedules
  • study habits
  • family routines

Therefore, the study supports an:

\[ \boxed{\text{association}} \]

but not necessarily:

\[ \boxed{\text{causal conclusion}} \]


Example 6 — Randomized Experiment

Researchers randomly assign participants to either:

  • a new study program
  • a standard study program

After several weeks, the new-program group has significantly higher scores.

Why does this design provide stronger evidence of causation?

Solution

The key feature is:

\[ \boxed{\text{random assignment}} \]

Random assignment helps make the groups similar with respect to other variables.

Therefore, a systematic difference in outcomes provides stronger evidence that the treatment itself caused the difference.


Common Mistakes

WarningCommon Mistakes
  • Assuming that a strong correlation proves causation.
  • Thinking a negative correlation is weaker simply because \(r\) is negative.
  • Saying \(r\approx0\) means there is absolutely no relationship; there could be a nonlinear relationship.
  • Ignoring lurking or confounding variables.
  • Claiming causation from observational data alone.
  • Ignoring the design of the study when evaluating a causal claim.
  • Looking only at \(r\) and ignoring important patterns or outliers in the scatterplot.
  • Assuming that two variables moving together means one must cause the other.

Practice Problems

  1. A dataset has:

\[ r=0.94 \]

Describe the linear correlation.

  1. A dataset has:

\[ r=-0.20 \]

Describe the linear correlation.

  1. Which represents the stronger linear relationship?

\[ r=-0.87 \]

or:

\[ r=0.52 \]

  1. Ice cream sales and swimming pool attendance both increase during summer. Name a likely lurking variable.

  2. A study finds that people who exercise more tend to report lower stress. The researchers only surveyed participants. Can they conclude that exercise caused the lower stress?

  3. Researchers randomly assign plants to receive either Fertilizer A or Fertilizer B and then compare their growth under otherwise similar conditions. Does this design provide stronger evidence for causation?

  4. A scatterplot shows a very strong U-shaped relationship, but \(r\) is close to 0. Does \(r\approx0\) prove there is no relationship?

1

The correlation coefficient is:

\[ r=0.94 \]

Because it is positive, the direction is positive.

Because:

\[ |0.94| \]

is close to 1, the linear relationship is strong.

Therefore:

\[ \boxed{\text{strong positive linear correlation}} \]


2

The correlation coefficient is:

\[ r=-0.20 \]

The negative sign indicates a negative direction.

Because:

\[ |-0.20|=0.20 \]

is relatively close to 0, the linear relationship is weak.

Therefore:

\[ \boxed{\text{weak negative linear correlation}} \]


3

Compare the absolute values:

\[ |-0.87|=0.87 \]

and:

\[ |0.52|=0.52 \]

Since:

\[ 0.87>0.52 \]

the stronger linear relationship is:

\[ \boxed{r=-0.87} \]


4

Both ice cream sales and swimming pool attendance are likely affected by:

\[ \boxed{\text{temperature or warm weather}} \]

Warm weather encourages both activities.


5

No.

The researchers observed existing behavior rather than randomly assigning people to exercise conditions.

Other variables could influence both exercise habits and stress.

The study supports an association but does not, by itself, establish causation.


6

Yes.

Random assignment helps create comparable groups and reduces the influence of other variables.

Therefore, this experiment can provide stronger evidence that the type of fertilizer caused differences in plant growth.


7

No.

The correlation coefficient \(r\) measures linear association.

A U-shaped pattern is nonlinear, so it is possible to have:

\[ r\approx0 \]

while still having a strong relationship.

Therefore:

\[ \boxed{\text{Always examine the scatterplot when available.}} \]

Summary

Correlation describes the direction and strength of a linear relationship between two numerical variables.

The correlation coefficient satisfies:

\[ \boxed{-1\le r\le1} \]

  • \(r>0\) → positive linear correlation
  • \(r<0\) → negative linear correlation
  • \(|r|\) near 1 → stronger linear correlation
  • \(|r|\) near 0 → weaker linear correlation

However:

\[ \boxed{\text{Correlation does not imply causation.}} \]

An observed association may be explained by:

  • a causal relationship
  • a lurking or confounding variable
  • coincidence or other factors

Observational studies can reveal associations, while well-designed randomized experiments provide much stronger evidence for causal conclusions.

  • Sign of \(r\) → direction.
  • \(|r|\) → strength.
  • \(r\) measures linear association.
  • A strong correlation does not prove causation.
  • Ask whether another variable could explain both.
  • Look at the scatterplot when available.
  • Random assignment is a major clue that a causal conclusion may be justified.
  • Observational data usually support association, not causation.