Correlation & Causation
By the end of this lesson, you’ll be able to:
- Distinguish correlation from causation.
- Interpret the correlation coefficient \(r\) qualitatively.
- Identify positive, negative, strong, and weak linear correlations.
- Recognize lurking or confounding variables that may explain an association.
- Distinguish observational studies from experiments when evaluating causal claims.
- Recognize misleading conclusions drawn from correlated data.
Key Ideas
Two variables are associated when values of one variable tend to occur with particular values of another.
When that association is linear, we can describe it using correlation.
For example:
- More hours studied may be associated with higher test scores.
- Higher temperatures may be associated with lower heating costs.
But an association between two variables does not automatically mean that one causes the other.
Causation means that changing one variable produces a change in another variable.
The key principle is:
Correlation does not imply causation.
Even a very strong correlation does not, by itself, prove that one variable causes the other.
The Correlation Coefficient
The correlation coefficient, written as \(r\), measures the direction and strength of a linear relationship between two numerical variables.
Its value is always between:
\[ \boxed{-1\le r\le1} \]
The sign tells the direction:
- \(r>0\) → positive linear correlation
- \(r<0\) → negative linear correlation
- \(r\approx0\) → little or no linear correlation
The magnitude \(|r|\) tells the strength:
- \(|r|\) close to 1 → strong linear correlation
- \(|r|\) close to 0 → weak linear correlation
For example:
\[ r=0.92 \]
indicates a strong positive linear correlation.
Meanwhile:
\[ r=-0.88 \]
indicates a strong negative linear correlation.
And:
\[ r=0.08 \]
indicates little or no linear correlation.
A negative correlation is not automatically weaker than a positive correlation.
For example:
\[ r=-0.95 \]
represents a stronger linear relationship than:
\[ r=0.40 \]
Strength depends on \(|r|\), not whether \(r\) is positive or negative.
Correlation Measures Linear Relationships
The correlation coefficient \(r\) describes linear association.
This distinction matters.
A scatterplot can show a strong curved relationship even when:
\[ r\approx0 \]
Therefore, always look at the scatterplot when one is provided.
A small value of \(r\) does not necessarily mean that the variables have no relationship at all.
It means that they have little or no linear relationship.
Correlation vs. Causation
Suppose two variables tend to increase together.
There are several possible explanations.
Possibility 1: One Variable Causes the Other
For example:
Increasing the amount of fertilizer applied to a plant may cause the plant to grow more under appropriate conditions.
Here, a causal relationship may be plausible.
Possibility 2: A Third Variable Affects Both
Suppose:
- ice cream sales increase
- swimming-related incidents also increase
These variables may be positively correlated.
But buying ice cream does not cause swimming incidents.
A third variable helps explain both:
\[ \boxed{\text{warmer weather}} \]
Warmer weather can lead to:
- more ice cream purchases
- more people swimming
This is an example of a lurking variable.
Possibility 3: The Association May Be Coincidental
Sometimes two variables happen to move together even though there is no meaningful relationship between them.
With enough variables and enough data, some correlations can occur simply by chance.
Therefore, correlation alone is not sufficient evidence of causation.
Lurking and Confounding Variables
A lurking variable is a variable that is not included in the analysis but may help explain the observed relationship between the variables being studied.
For example, suppose data show:
Larger fires tend to have more firefighters present.
It would be incorrect to conclude:
More firefighters cause larger fires.
The lurking variable is:
\[ \boxed{\text{size or severity of the fire}} \]
Larger fires tend to:
- cause more damage
- require more firefighters
The underlying severity of the fire helps explain the association.
A confounding variable is a related idea: it is a variable whose effect is mixed with the effect of another variable, making it difficult to determine which variable is responsible for an observed difference.
For many test questions, the key idea is simply:
Could another variable be influencing the relationship?
Observational Studies vs. Experiments
Whether a study can support a causal conclusion depends heavily on how the data were collected.
Observational Study
In an observational study, researchers observe or measure variables without assigning treatments.
For example, researchers might record:
- how many hours students study
- their test scores
They may find that students who study more tend to have higher scores.
This establishes an association, but other variables could also matter, such as:
- prior knowledge
- motivation
- sleep
- course difficulty
Therefore, an observational study generally does not establish causation by itself.
Randomized Experiment
In a randomized experiment, researchers randomly assign subjects to different treatments or conditions.
Random assignment helps balance other variables between the groups.
If the groups then show meaningful differences, there is stronger evidence that the treatment itself caused the difference.
When a question asks whether a study supports a causal conclusion, look for random assignment to treatments.
Randomized experiments provide much stronger evidence for causation than observational studies.
Common Problem Types
1. Describing Positive Correlation
Suppose:
\[ r=0.85 \]
Because:
\[ r>0 \]
the direction is positive.
Because \(|r|\) is relatively close to 1, the linear relationship is strong.
Therefore:
\[ \boxed{\text{strong positive linear correlation}} \]
2. Describing Negative Correlation
Suppose:
\[ r=-0.90 \]
The negative sign indicates that as one variable increases, the other tends to decrease.
Because:
\[ |-0.90|=0.90 \]
the relationship is strong.
Therefore:
\[ \boxed{\text{strong negative linear correlation}} \]
3. Comparing Correlation Strength
Which represents a stronger linear relationship?
\[ r=0.45 \]
or:
\[ r=-0.80 \]
Compare absolute values:
\[ |0.45|=0.45 \]
\[ |-0.80|=0.80 \]
Since:
\[ 0.80>0.45 \]
the stronger correlation is:
\[ \boxed{r=-0.80} \]
The negative sign describes direction, not weakness.
4. Recognizing Correlation Without Causation
Suppose ice cream sales and swimming-related incidents both increase during the summer.
The variables may be positively correlated.
However, it would be unreasonable to conclude that:
Buying ice cream causes swimming incidents.
A likely lurking variable is:
\[ \boxed{\text{temperature}} \]
Warmer weather contributes to both.
5. Identifying a Lurking Variable
Suppose cities with more firefighters at a fire also tend to experience more fire damage.
Does this mean firefighters cause additional damage?
No.
Larger and more severe fires generally:
- cause more damage
- require more firefighters
The lurking variable is:
\[ \boxed{\text{fire size or severity}} \]
6. Evaluating a Causal Claim
Suppose researchers observe that students who sleep more tend to earn higher test scores.
Can they conclude that additional sleep caused the higher scores?
Not necessarily.
Other variables may differ between the students.
The data show an:
\[ \boxed{\text{association}} \]
but do not automatically establish:
\[ \boxed{\text{causation}} \]
7. Recognizing Experimental Evidence
Suppose researchers randomly assign participants to two study methods.
One group uses Method A, and the other uses Method B.
Everything else is kept as similar as possible.
If the groups later show a meaningful difference in performance, the randomized design provides stronger evidence that the study method caused the difference.
The important feature is:
\[ \boxed{\text{random assignment}} \]
Strategies
Separate Association From Causation
When two variables are related, first say:
The variables are associated.
Do not automatically say:
One variable causes the other.
Look for Another Explanation
Ask:
Could another variable influence both variables?
If yes, that variable may help explain the observed correlation.
Use the Sign of \(r\) for Direction
Remember:
\[ r>0 \]
means positive linear correlation.
And:
\[ r<0 \]
means negative linear correlation.
Use \(|r|\) for Strength
To compare strength, ignore the sign temporarily.
For example:
\[ |-0.9|=0.9 \]
is stronger than:
\[ |0.5|=0.5 \]
Look at the Scatterplot
Do not rely only on \(r\) when a scatterplot is available.
Check for:
- curved relationships
- outliers
- clusters
- unusual patterns
These features may not be captured well by a single correlation coefficient.
Look for Random Assignment
When evaluating whether a study supports causation, ask:
Were subjects randomly assigned to treatments?
If not, be cautious about making a causal conclusion.
Worked Examples
Example 1 — Interpret a Correlation Coefficient
Suppose the correlation between hours studied and test scores is:
\[ r=0.82 \]
Describe the relationship.
Solution
First consider the sign:
\[ 0.82>0 \]
so the relationship is positive.
Next consider the magnitude:
\[ |0.82|=0.82 \]
which indicates a fairly strong linear relationship.
Therefore:
\[ \boxed{\text{strong positive linear correlation}} \]
In context:
Students who study more hours tend to have higher test scores.
This statement describes association and does not, by itself, establish causation.
Example 2 — Compare Correlations
Which relationship is stronger?
\[ r=-0.91 \]
or:
\[ r=0.63 \]
Solution
Compare absolute values:
\[ |-0.91|=0.91 \]
and:
\[ |0.63|=0.63 \]
Since:
\[ 0.91>0.63 \]
we conclude:
\[ \boxed{r=-0.91\text{ is stronger}} \]
It represents a strong negative linear correlation.
Example 3 — Ice Cream and Swimming
A city records both ice cream sales and swimming-related incidents throughout the year.
The two variables show a strong positive correlation.
Can we conclude that ice cream sales cause swimming incidents?
Solution
No.
Both variables tend to increase during warmer weather.
A likely lurking variable is:
\[ \boxed{\text{temperature}} \]
The correlation does not establish that one measured variable causes the other.
Example 4 — Firefighters and Fire Damage
Data show that fires with more firefighters present tend to have more property damage.
A person concludes:
Sending more firefighters causes more damage.
Is this conclusion justified?
Solution
No.
More severe fires generally require more firefighters and also produce more damage.
The lurking variable is:
\[ \boxed{\text{fire severity}} \]
Therefore, the observed correlation does not show that firefighters cause the damage.
Example 5 — Observational Study
Researchers survey 1,000 students and find that students who regularly eat breakfast tend to have higher test scores.
Can the researchers conclude that eating breakfast causes higher scores?
Solution
Not from this evidence alone.
This is observational data.
Students who eat breakfast may differ in other ways, such as:
- sleep habits
- schedules
- study habits
- family routines
Therefore, the study supports an:
\[ \boxed{\text{association}} \]
but not necessarily:
\[ \boxed{\text{causal conclusion}} \]
Example 6 — Randomized Experiment
Researchers randomly assign participants to either:
- a new study program
- a standard study program
After several weeks, the new-program group has significantly higher scores.
Why does this design provide stronger evidence of causation?
Solution
The key feature is:
\[ \boxed{\text{random assignment}} \]
Random assignment helps make the groups similar with respect to other variables.
Therefore, a systematic difference in outcomes provides stronger evidence that the treatment itself caused the difference.
Common Mistakes
- Assuming that a strong correlation proves causation.
- Thinking a negative correlation is weaker simply because \(r\) is negative.
- Saying \(r\approx0\) means there is absolutely no relationship; there could be a nonlinear relationship.
- Ignoring lurking or confounding variables.
- Claiming causation from observational data alone.
- Ignoring the design of the study when evaluating a causal claim.
- Looking only at \(r\) and ignoring important patterns or outliers in the scatterplot.
- Assuming that two variables moving together means one must cause the other.
Practice Problems
- A dataset has:
\[ r=0.94 \]
Describe the linear correlation.
- A dataset has:
\[ r=-0.20 \]
Describe the linear correlation.
- Which represents the stronger linear relationship?
\[ r=-0.87 \]
or:
\[ r=0.52 \]
Ice cream sales and swimming pool attendance both increase during summer. Name a likely lurking variable.
A study finds that people who exercise more tend to report lower stress. The researchers only surveyed participants. Can they conclude that exercise caused the lower stress?
Researchers randomly assign plants to receive either Fertilizer A or Fertilizer B and then compare their growth under otherwise similar conditions. Does this design provide stronger evidence for causation?
A scatterplot shows a very strong U-shaped relationship, but \(r\) is close to 0. Does \(r\approx0\) prove there is no relationship?
1
The correlation coefficient is:
\[ r=0.94 \]
Because it is positive, the direction is positive.
Because:
\[ |0.94| \]
is close to 1, the linear relationship is strong.
Therefore:
\[ \boxed{\text{strong positive linear correlation}} \]
2
The correlation coefficient is:
\[ r=-0.20 \]
The negative sign indicates a negative direction.
Because:
\[ |-0.20|=0.20 \]
is relatively close to 0, the linear relationship is weak.
Therefore:
\[ \boxed{\text{weak negative linear correlation}} \]
3
Compare the absolute values:
\[ |-0.87|=0.87 \]
and:
\[ |0.52|=0.52 \]
Since:
\[ 0.87>0.52 \]
the stronger linear relationship is:
\[ \boxed{r=-0.87} \]
4
Both ice cream sales and swimming pool attendance are likely affected by:
\[ \boxed{\text{temperature or warm weather}} \]
Warm weather encourages both activities.
5
No.
The researchers observed existing behavior rather than randomly assigning people to exercise conditions.
Other variables could influence both exercise habits and stress.
The study supports an association but does not, by itself, establish causation.
6
Yes.
Random assignment helps create comparable groups and reduces the influence of other variables.
Therefore, this experiment can provide stronger evidence that the type of fertilizer caused differences in plant growth.
7
No.
The correlation coefficient \(r\) measures linear association.
A U-shaped pattern is nonlinear, so it is possible to have:
\[ r\approx0 \]
while still having a strong relationship.
Therefore:
\[ \boxed{\text{Always examine the scatterplot when available.}} \]
Summary
Correlation describes the direction and strength of a linear relationship between two numerical variables.
The correlation coefficient satisfies:
\[ \boxed{-1\le r\le1} \]
- \(r>0\) → positive linear correlation
- \(r<0\) → negative linear correlation
- \(|r|\) near 1 → stronger linear correlation
- \(|r|\) near 0 → weaker linear correlation
However:
\[ \boxed{\text{Correlation does not imply causation.}} \]
An observed association may be explained by:
- a causal relationship
- a lurking or confounding variable
- coincidence or other factors
Observational studies can reveal associations, while well-designed randomized experiments provide much stronger evidence for causal conclusions.
- Sign of \(r\) → direction.
- \(|r|\) → strength.
- \(r\) measures linear association.
- A strong correlation does not prove causation.
- Ask whether another variable could explain both.
- Look at the scatterplot when available.
- Random assignment is a major clue that a causal conclusion may be justified.
- Observational data usually support association, not causation.