Материал Causal, probabilistic and statistical fallacies
0%

Causal, probabilistic and statistical fallacies

Causal, probabilistic and statistical fallacies

Data does not “speak for itself.” It answers the question posed by the model, measurement, and selection procedure.

Scene: Happy Bracelet

A person put on a new bracelet, successfully passed an interview, and decided that the bracelet brought good luck. The events indeed happened one after another, but order does not prove cause. Preparation, experience, the interviewer’s questions, or chance could have helped.

The “after this, therefore because of this” fallacy is called post hoc.

Repair: list other explanations, find a mechanism and compare cases with and without the bracelet. Honest conclusion: “the events coincided in time; causal connection is not established”.

Change Matching

People who buy sunglasses more often buy ice cream more often.

Ice cream is not forced to buy. Both events grow due to a third factor — hot sunny weather. It is called the confounding factor.

Question: what could simultaneously affect both the supposed cause and the result?

Percentages without grounds

The ad says: “risk reduced by 50%”. This may mean a decrease from 2 cases per thousand to 1 case. The relative change is large, the absolute difference — one case per thousand.

Repair: always ask:

  • 50% from which initial level;
  • how many events and participants there were;
  • what the absolute risk is;
  • within what timeframe;
  • how uncertain the outcome is.

Base frequency

Imagine a rare event that occurs in one person out of a thousand. The test usually detects it, but sometimes gives false positives. Even a good test among a thousand people can find one true case and several false ones.

Therefore, “positive result” and “probability of an event given a positive result” are not the same. You need to take into account how rare the event was before the check. This is the base rate.

Base frequency on natural frequencies: how many true and false signals arise in the group

Who made it into the sample

An in-app survey speaks not for all residents, but first and foremost for people who have the app and the desire to respond. Reviews show the voices of those who decided to leave them. Winner stories hide those who did the same and did not succeed.

Breakdowns get different names — self-selection, non-representative sampling, survivorship bias — but the repair is general: restore the path of data entering the sample.

Average can hide the data structure

The average salary, average waiting time, or average rating do not show the spread. Two groups with the same average can be arranged completely differently.

Sometimes aggregated statistics show one trend, while each subgroup — the opposite. This is the Simpson’s paradox. Before drawing conclusions, it is necessary to check important subgroups and the reason for their different sizes.

Short practice

After the renovation, the café’s attendance increased by 30%. This means the new design brought more guests.

Call:

  1. three competing reasons;
  2. data before and after, which are needed;
  3. a suitable comparison group;
  4. a conclusion, which is permissible right now.

Repair: “After the repair, attendance increased by 30%; to assess the design’s contribution, one needs to consider the season, advertising, prices, weather, and changes in foot traffic”.

Chapter Rule: separate observation, probabilistic association, and causal assertion. The stronger the conclusion, the more alternatives it must exclude.

Causality: Observed Association, Confounding Factor and Intervention

Нашли неточность? Выделите фрагмент текста — рядом появится жучок.

Нужен разбор именно вашей ситуации?

Статья описывает общий случай. Если у вас частный — можно разобрать его отдельно, платно. А если не хватает целого материала, предложите тему: её оплачивают вскладчину, и она выходит открытой для всех.

Доска запросов
Дальше