PythonMastery
beginner 20 min read · lesson 6 of 6 in Data Science Fundamentals

Exploratory Data Analysis

1 · The lesson

read

Before you train a model, build a dashboard, or write up a finding, you look at the data. Properly. Methodically. With suspicion. This step is called exploratory data analysis (EDA) — coined by John Tukey in 1977, still the cheapest source of value in any DS project.

The phrase that captures EDA best: "You can observe a lot by watching." (Yogi Berra.) Most modelling disasters trace back to skipped EDA. Most "wow" findings come straight out of it, no ML required.

EDA typically eats 40-80% of project time. Budget for it; don't apologise for it.


1. The Standard EDA Flow

Six steps, in order. Skip none.

text
1. Shape & types     "What am I looking at?"
2. Missing values    "What's not there?"
3. Distributions     "What does each column look like alone?"
4. Categoricals      "What are the levels and counts?"
5. Relationships     "How do columns relate to each other?"
6. Outliers          "What's unusually far from the rest?"

Run them every time. The order matters — you can't sensibly look at relationships before knowing the dtypes, and you can't trust a correlation if half the column is missing.


2. A Toy Dataset

We'll explore a small DataFrame inline so every step runs:

python
import pandas as pd
import numpy as np

df = pd.DataFrame({
    "customer_id": range(1, 13),
    "age":         [25, 32, 47, np.nan, 28, 19, 55, 41, 33, 29, np.nan, 38],
    "city":        ["London", "Paris", "Berlin", "London", "Paris",
                    "London", "Berlin", "Paris", "London", "Berlin",
                    "Paris", "London"],
    "tier":        ["free", "pro", "pro", "free", "free",
                    "free", "enterprise", "pro", "free", "pro",
                    "free", "enterprise"],
    "spend":       [12.5, 89.0, 145.0, 8.0, 22.0, 5.0,
                    980.0, 110.0, 30.0, 75.0, 18.0, 1100.0],
})

print(df)

Twelve rows is unrealistic but it lets every output fit on screen.


3. Step 1 — Shape & Types

python
print(df.shape)                       # (12, 5)
print(df.dtypes)
# customer_id      int64
# age            float64       — float because of the NaNs
# city            object
# tier            object
# spend          float64

df.info()
# Non-Null counts, dtypes, memory usage — all in one call
+ setup added so this can run · defines df
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

df = _AutoMock('df')

Look for:

  • Unexpected dtypes — a column you expected to be numeric showing as object usually means it has stray strings ("N/A", "-") mixed in.
  • float instead of int — usually means missing values forced the upcast (NumPy ints can't hold NaN).
  • Row count — does it match the spec you were given?

4. Step 2 — Missing Values

python
print(df.isnull().sum())
# customer_id    0
# age            2
# city           0
# tier           0
# spend          0
+ setup added so this can run · defines df
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

df = _AutoMock('df')

.isnull() returns a boolean DataFrame; .sum() counts the Trues per column. Add .sort_values(ascending=False) to surface the worst offenders first.

Percentages are more useful than raw counts on real data:

python
print((df.isnull().mean() * 100).round(2))
# age            16.67
+ setup added so this can run · defines df
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

df = _AutoMock('df')

Three response strategies — pick per column:

StrategyWhen to use
Drop rows (df.dropna())Few rows missing, plenty of data
Drop column>50% missing and not critical
Fill / impute (df.fillna(value))Few rows missing but every row matters
Leave & flagAlgorithm can handle NaNs (e.g. xgboost)
python
df["age"] = df["age"].fillna(df["age"].median())     # impute with median
+ setup added so this can run · defines df
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

df = _AutoMock('df')

The fuller treatment is in the data cleaning lesson.


5. Step 3 — Distributions of Numerics

python
print(df.describe())
#        customer_id        age        spend
# count        12.00      10.00        12.00
# mean          6.50      34.70       216.20
# std           3.61      11.05       377.83   <- big! suspicious
# min           1.00      19.00         5.00
# 25%           3.75      28.25        19.00
# 50%           6.50      32.50        52.50
# 75%           9.25      40.25       117.50
# max          12.00      55.00      1100.00   <- way above 75th pct
+ setup added so this can run · defines df
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

df = _AutoMock('df')

The gap between the 75th percentile and the max is the single most useful signal in describe(). When max >> 75%, you have outliers — visible here in spend (1100 vs a median of 52.50).

For shape, plot a histogram of every numeric column:

python
# df.hist(figsize=(10, 6), bins=20)
# plt.tight_layout()
# plt.savefig("dists.png")
# Chart preview: one histogram per numeric column laid out in a grid.

df.hist() is the lazy-but-effective EDA chart — one line, every numeric column.


6. Step 4 — Categoricals

For object / category columns, describe(include="object") and value_counts do the work:

python
print(df.describe(include="object"))
#         city  tier
# count     12    12
# unique     3     3
# top   London  free
# freq       5     5

print(df["tier"].value_counts())
# free          5
# pro           4
# enterprise    2

print(df["tier"].value_counts(normalize=True).round(2))
# free          0.42
# pro           0.33
# enterprise    0.17
+ setup added so this can run · defines df
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

df = _AutoMock('df')

Look for:

  • Cardinality — 3 unique cities is fine; 10,000 unique IDs is not a usable categorical, it's an identifier.
  • Class imbalance — if 95% of your target is "no", model accuracy is a useless metric (a constant "no" predictor scores 95%).
  • Typos / dirty levels — "London", "london ", "LONDON" should be one category, not three. Catch this in cleaning.

7. Step 5 — Relationships

For numeric pairs, .corr() gives a matrix of Pearson correlations between -1 and +1:

python
print(df[["age", "spend"]].corr().round(2))
#         age  spend
# age    1.00   0.57
# spend  0.57   1.00
+ setup added so this can run · defines df
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

df = _AutoMock('df')

Then visualise — a scatter is the truth-teller; correlation is a one-number summary that hides shape:

python
# import matplotlib.pyplot as plt
# plt.scatter(df["age"], df["spend"], alpha=0.7)
# plt.xlabel("Age"); plt.ylabel("Spend")
# plt.savefig("scatter.png")
# Chart preview: a noisy upward cloud — older customers tend to spend more.

For numeric × categorical, groupby plus a boxplot:

python
print(df.groupby("tier")["spend"].agg(["mean", "median", "count"]))
#              mean  median  count
# tier
# enterprise  1040.0  1040.0      2
# free          13.1    15.0      5
# pro          104.75   99.5      4
+ setup added so this can run · defines df
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

df = _AutoMock('df')

For all numerics at once, a heatmap of the correlation matrix shows which columns move together. Seaborn's heatmap does this in one line.


8. Step 6 — Outliers

The boxplot is the diagnostic. The whiskers reach to 1.5×IQR; anything beyond is flagged:

python
# df.boxplot(column="spend")
# plt.savefig("box.png")
# Chart preview: a tight box near 0-120 with two distant flier points near 980 and 1100.

For each suspect, investigate before "fixing":

  • Real but rare — keep them; they're informative. (A whale customer at £1100 is real signal.)
  • Data entry error — fix or drop. (An age of 250 is a typo.)
  • Different population — split the dataset. (Enterprise vs free behave so differently that one model for both is wrong.)

Never silently delete outliers. They're often the most interesting rows.


9. EDA Heuristics — When You See X, Suspect Y

SymptomLikely causeFirst move
Huge skew, long right tailPower-law variableTry np.log1p and replot
One column with many NaNsOptional field or join failureCheck upstream pipeline
dtype=object in a numeric columnStrings mixed in ("N/A", "-")pd.to_numeric(col, errors="coerce")
95/5 class imbalanceImbalanced targetUse a different metric (F1, AUC), not accuracy
Correlation 0.99 between two featuresDuplicate or trivially-derivedDrop one
max >> 99th percentileOutliers or a data capBoxplot the column
Bimodal histogramTwo subpopulations mixedGroup by a categorical, replot

The heuristics aren't rules — they're prompts to look closer.


10. Write a Data Dictionary

Before modelling, write a short markdown table — one row per column — covering name, type, allowed values, unit, source. Five minutes; saves hours later.

markdown
| column      | type    | unit      | notes                          |
|-------------|---------|-----------|--------------------------------|
| customer_id | int     | -         | unique, primary key            |
| age         | float   | years     | 16.7% missing, imputed median  |
| city        | string  | -         | 3 levels: London/Paris/Berlin  |
| tier        | string  | -         | 3 levels: free/pro/enterprise  |
| spend       | float   | GBP       | last 30 days, includes outliers|

Future-you and your reviewers will thank you.


11. Automated EDA — A Note

Tools like ydata-profiling (formerly pandas-profiling) and sweetviz generate a full HTML EDA report from a single call:

python
# from ydata_profiling import ProfileReport
# report = ProfileReport(df, title="Customer EDA")
# report.to_file("eda.html")

Useful as a first scan on a dataset you've never seen. Not a replacement for actually thinking — these reports surface symptoms; you still diagnose causes.


Common Mistakes

  • Skipping EDA and jumping to a model. The model will run; it will also be wrong, and you won't know why.
  • Ignoring missing values. They propagate through pandas operations silently — NaN + 1 = NaN — and quietly corrupt aggregates.
  • Not looking at the target's distribution. Imbalanced classes, a heavy tail, or zero-inflation completely change which model is appropriate.
  • Treating describe() as enough. Mean and std hide bimodality, skew, and outliers. Plot the histogram.
  • Trusting correlation without a scatter. Anscombe's quartet — four datasets, identical correlation, wildly different shapes — is the canonical warning.
  • Cleaning outliers without understanding them. Sometimes the outlier is the finding.

🎯 Your Turn — quick_eda()

Write a reusable function that prints the first-pass EDA summary you'll want on every dataset: shape, dtypes, missing counts, and describe().

Skeleton:

python
import pandas as pd

def quick_eda(df):
    """Print a first-pass EDA summary of a DataFrame."""
    # TODO 1: shape and dtype counts
    # TODO 2: missing counts per column (only if > 0)
    # TODO 3: describe() — numerics and objects
    ...

# Test
import numpy as np
df = pd.DataFrame({
    "id":     range(1, 8),
    "age":    [25, 32, np.nan, 41, 28, 55, np.nan],
    "city":   ["London", "Paris", "London", "Berlin", "London", "Paris", "Berlin"],
    "spend":  [50, 120, 35, 200, 85, 410, 60],
})
quick_eda(df)
Hint 1 — Filtering the missing report Compute missing = df.isnull().sum(), then print only the columns where missing > 0. Use missing[missing > 0].
Hint 2 — Describe for both kinds df.describe() covers numerics. df.describe(include="object") covers strings. Wrap the second in try/except in case the DataFrame has no object columns.
Show full solution
python
import pandas as pd

def quick_eda(df):
    print("=" * 50)
    print(f"Shape: {df.shape[0]:,} rows x {df.shape[1]} columns")
    print("\nDtypes:")
    print(df.dtypes.value_counts())

    missing = df.isnull().sum()
    missing = missing[missing > 0]
    if not missing.empty:
        pct = (missing / len(df) * 100).round(2)
        report = pd.DataFrame({"missing": missing, "pct": pct})
        print("\nMissing values:")
        print(report.sort_values("missing", ascending=False))
    else:
        print("\nMissing values: none.")

    print("\nNumeric summary:")
    print(df.describe().round(2))

    try:
        print("\nCategorical summary:")
        print(df.describe(include="object"))
    except ValueError:
        pass

    print("=" * 50)


import numpy as np
df = pd.DataFrame({
    "id":     range(1, 8),
    "age":    [25, 32, np.nan, 41, 28, 55, np.nan],
    "city":   ["London", "Paris", "London", "Berlin", "London", "Paris", "Berlin"],
    "spend":  [50, 120, 35, 200, 85, 410, 60],
})
quick_eda(df)

You now have a reusable first-pass scanner — drop it into every new notebook and you're 30 seconds into useful EDA before your coffee lands.


What You Learned

  • EDA follows a fixed six-step flow: shape → missing → distributions → categoricals → relationships → outliers.
  • .info(), .describe(), .isnull().sum(), .value_counts() are the four commands you'll run on every dataset.
  • Distributions matter more than means — always plot a histogram before trusting describe().
  • .corr() is the summary, the scatter is the truth — never accept correlation without seeing the shape.
  • Outliers are clues, not noise. Investigate before deleting.
  • A short data dictionary before modelling pays for itself within a week.
  • Auto-EDA tools like ydata-profiling are useful starters but never the finish line.

You've now walked the full DS-fundamentals path — questions, Python idioms, NumPy, pandas, visualisation, EDA. The next paths take you deeper: Statistics for inference, Regression for your first models, and the full Data Science track for production-grade tooling.