Exploratory Data Analysis
1 · The lesson
readBefore you train a model, build a dashboard, or write up a finding, you look at the data. Properly. Methodically. With suspicion. This step is called exploratory data analysis (EDA) — coined by John Tukey in 1977, still the cheapest source of value in any DS project.
The phrase that captures EDA best: "You can observe a lot by watching." (Yogi Berra.) Most modelling disasters trace back to skipped EDA. Most "wow" findings come straight out of it, no ML required.
EDA typically eats 40-80% of project time. Budget for it; don't apologise for it.
1. The Standard EDA Flow
Six steps, in order. Skip none.
1. Shape & types "What am I looking at?" 2. Missing values "What's not there?" 3. Distributions "What does each column look like alone?" 4. Categoricals "What are the levels and counts?" 5. Relationships "How do columns relate to each other?" 6. Outliers "What's unusually far from the rest?"
Run them every time. The order matters — you can't sensibly look at relationships before knowing the dtypes, and you can't trust a correlation if half the column is missing.
2. A Toy Dataset
We'll explore a small DataFrame inline so every step runs:
import pandas as pd import numpy as np df = pd.DataFrame({ "customer_id": range(1, 13), "age": [25, 32, 47, np.nan, 28, 19, 55, 41, 33, 29, np.nan, 38], "city": ["London", "Paris", "Berlin", "London", "Paris", "London", "Berlin", "Paris", "London", "Berlin", "Paris", "London"], "tier": ["free", "pro", "pro", "free", "free", "free", "enterprise", "pro", "free", "pro", "free", "enterprise"], "spend": [12.5, 89.0, 145.0, 8.0, 22.0, 5.0, 980.0, 110.0, 30.0, 75.0, 18.0, 1100.0], }) print(df)
Twelve rows is unrealistic but it lets every output fit on screen.
3. Step 1 — Shape & Types
print(df.shape) # (12, 5) print(df.dtypes) # customer_id int64 # age float64 — float because of the NaNs # city object # tier object # spend float64 df.info() # Non-Null counts, dtypes, memory usage — all in one call
setup added so this can run · defines df
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) df = _AutoMock('df')
Look for:
- Unexpected dtypes — a column you expected to be numeric showing as
objectusually means it has stray strings ("N/A","-") mixed in. floatinstead ofint— usually means missing values forced the upcast (NumPy ints can't holdNaN).- Row count — does it match the spec you were given?
4. Step 2 — Missing Values
print(df.isnull().sum()) # customer_id 0 # age 2 # city 0 # tier 0 # spend 0
setup added so this can run · defines df
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) df = _AutoMock('df')
.isnull() returns a boolean DataFrame; .sum() counts the Trues per column. Add .sort_values(ascending=False) to surface the worst offenders first.
Percentages are more useful than raw counts on real data:
print((df.isnull().mean() * 100).round(2)) # age 16.67
setup added so this can run · defines df
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) df = _AutoMock('df')
Three response strategies — pick per column:
| Strategy | When to use |
|---|---|
Drop rows (df.dropna()) | Few rows missing, plenty of data |
| Drop column | >50% missing and not critical |
Fill / impute (df.fillna(value)) | Few rows missing but every row matters |
| Leave & flag | Algorithm can handle NaNs (e.g. xgboost) |
df["age"] = df["age"].fillna(df["age"].median()) # impute with median
setup added so this can run · defines df
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) df = _AutoMock('df')
The fuller treatment is in the data cleaning lesson.
5. Step 3 — Distributions of Numerics
print(df.describe()) # customer_id age spend # count 12.00 10.00 12.00 # mean 6.50 34.70 216.20 # std 3.61 11.05 377.83 <- big! suspicious # min 1.00 19.00 5.00 # 25% 3.75 28.25 19.00 # 50% 6.50 32.50 52.50 # 75% 9.25 40.25 117.50 # max 12.00 55.00 1100.00 <- way above 75th pct
setup added so this can run · defines df
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) df = _AutoMock('df')
The gap between the 75th percentile and the max is the single most useful signal in describe(). When max >> 75%, you have outliers — visible here in spend (1100 vs a median of 52.50).
For shape, plot a histogram of every numeric column:
# df.hist(figsize=(10, 6), bins=20) # plt.tight_layout() # plt.savefig("dists.png") # Chart preview: one histogram per numeric column laid out in a grid.
df.hist() is the lazy-but-effective EDA chart — one line, every numeric column.
6. Step 4 — Categoricals
For object / category columns, describe(include="object") and value_counts do the work:
print(df.describe(include="object")) # city tier # count 12 12 # unique 3 3 # top London free # freq 5 5 print(df["tier"].value_counts()) # free 5 # pro 4 # enterprise 2 print(df["tier"].value_counts(normalize=True).round(2)) # free 0.42 # pro 0.33 # enterprise 0.17
setup added so this can run · defines df
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) df = _AutoMock('df')
Look for:
- Cardinality — 3 unique cities is fine; 10,000 unique IDs is not a usable categorical, it's an identifier.
- Class imbalance — if 95% of your target is "no", model accuracy is a useless metric (a constant "no" predictor scores 95%).
- Typos / dirty levels —
"London","london ","LONDON"should be one category, not three. Catch this in cleaning.
7. Step 5 — Relationships
For numeric pairs, .corr() gives a matrix of Pearson correlations between -1 and +1:
print(df[["age", "spend"]].corr().round(2)) # age spend # age 1.00 0.57 # spend 0.57 1.00
setup added so this can run · defines df
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) df = _AutoMock('df')
Then visualise — a scatter is the truth-teller; correlation is a one-number summary that hides shape:
# import matplotlib.pyplot as plt # plt.scatter(df["age"], df["spend"], alpha=0.7) # plt.xlabel("Age"); plt.ylabel("Spend") # plt.savefig("scatter.png") # Chart preview: a noisy upward cloud — older customers tend to spend more.
For numeric × categorical, groupby plus a boxplot:
print(df.groupby("tier")["spend"].agg(["mean", "median", "count"])) # mean median count # tier # enterprise 1040.0 1040.0 2 # free 13.1 15.0 5 # pro 104.75 99.5 4
setup added so this can run · defines df
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) df = _AutoMock('df')
For all numerics at once, a heatmap of the correlation matrix shows which columns move together. Seaborn's heatmap does this in one line.
8. Step 6 — Outliers
The boxplot is the diagnostic. The whiskers reach to 1.5×IQR; anything beyond is flagged:
# df.boxplot(column="spend") # plt.savefig("box.png") # Chart preview: a tight box near 0-120 with two distant flier points near 980 and 1100.
For each suspect, investigate before "fixing":
- Real but rare — keep them; they're informative. (A whale customer at £1100 is real signal.)
- Data entry error — fix or drop. (An age of 250 is a typo.)
- Different population — split the dataset. (Enterprise vs free behave so differently that one model for both is wrong.)
Never silently delete outliers. They're often the most interesting rows.
9. EDA Heuristics — When You See X, Suspect Y
| Symptom | Likely cause | First move |
|---|---|---|
| Huge skew, long right tail | Power-law variable | Try np.log1p and replot |
One column with many NaNs | Optional field or join failure | Check upstream pipeline |
dtype=object in a numeric column | Strings mixed in ("N/A", "-") | pd.to_numeric(col, errors="coerce") |
| 95/5 class imbalance | Imbalanced target | Use a different metric (F1, AUC), not accuracy |
| Correlation 0.99 between two features | Duplicate or trivially-derived | Drop one |
max >> 99th percentile | Outliers or a data cap | Boxplot the column |
| Bimodal histogram | Two subpopulations mixed | Group by a categorical, replot |
The heuristics aren't rules — they're prompts to look closer.
10. Write a Data Dictionary
Before modelling, write a short markdown table — one row per column — covering name, type, allowed values, unit, source. Five minutes; saves hours later.
| column | type | unit | notes | |-------------|---------|-----------|--------------------------------| | customer_id | int | - | unique, primary key | | age | float | years | 16.7% missing, imputed median | | city | string | - | 3 levels: London/Paris/Berlin | | tier | string | - | 3 levels: free/pro/enterprise | | spend | float | GBP | last 30 days, includes outliers|
Future-you and your reviewers will thank you.
11. Automated EDA — A Note
Tools like ydata-profiling (formerly pandas-profiling) and sweetviz generate a full HTML EDA report from a single call:
# from ydata_profiling import ProfileReport # report = ProfileReport(df, title="Customer EDA") # report.to_file("eda.html")
Useful as a first scan on a dataset you've never seen. Not a replacement for actually thinking — these reports surface symptoms; you still diagnose causes.
Common Mistakes
- Skipping EDA and jumping to a model. The model will run; it will also be wrong, and you won't know why.
- Ignoring missing values. They propagate through pandas operations silently —
NaN + 1 = NaN— and quietly corrupt aggregates. - Not looking at the target's distribution. Imbalanced classes, a heavy tail, or zero-inflation completely change which model is appropriate.
- Treating
describe()as enough. Mean and std hide bimodality, skew, and outliers. Plot the histogram. - Trusting correlation without a scatter. Anscombe's quartet — four datasets, identical correlation, wildly different shapes — is the canonical warning.
- Cleaning outliers without understanding them. Sometimes the outlier is the finding.
🎯 Your Turn — quick_eda()
Write a reusable function that prints the first-pass EDA summary you'll want on every dataset: shape, dtypes, missing counts, and describe().
Skeleton:
import pandas as pd def quick_eda(df): """Print a first-pass EDA summary of a DataFrame.""" # TODO 1: shape and dtype counts # TODO 2: missing counts per column (only if > 0) # TODO 3: describe() — numerics and objects ... # Test import numpy as np df = pd.DataFrame({ "id": range(1, 8), "age": [25, 32, np.nan, 41, 28, 55, np.nan], "city": ["London", "Paris", "London", "Berlin", "London", "Paris", "Berlin"], "spend": [50, 120, 35, 200, 85, 410, 60], }) quick_eda(df)
Hint 1 — Filtering the missing report
Computemissing = df.isnull().sum(), then print only the columns where missing > 0. Use missing[missing > 0].
Hint 2 — Describe for both kinds
df.describe() covers numerics. df.describe(include="object") covers strings. Wrap the second in try/except in case the DataFrame has no object columns.
Show full solution
import pandas as pd def quick_eda(df): print("=" * 50) print(f"Shape: {df.shape[0]:,} rows x {df.shape[1]} columns") print("\nDtypes:") print(df.dtypes.value_counts()) missing = df.isnull().sum() missing = missing[missing > 0] if not missing.empty: pct = (missing / len(df) * 100).round(2) report = pd.DataFrame({"missing": missing, "pct": pct}) print("\nMissing values:") print(report.sort_values("missing", ascending=False)) else: print("\nMissing values: none.") print("\nNumeric summary:") print(df.describe().round(2)) try: print("\nCategorical summary:") print(df.describe(include="object")) except ValueError: pass print("=" * 50) import numpy as np df = pd.DataFrame({ "id": range(1, 8), "age": [25, 32, np.nan, 41, 28, 55, np.nan], "city": ["London", "Paris", "London", "Berlin", "London", "Paris", "Berlin"], "spend": [50, 120, 35, 200, 85, 410, 60], }) quick_eda(df)
You now have a reusable first-pass scanner — drop it into every new notebook and you're 30 seconds into useful EDA before your coffee lands.
What You Learned
- EDA follows a fixed six-step flow: shape → missing → distributions → categoricals → relationships → outliers.
.info(),.describe(),.isnull().sum(),.value_counts()are the four commands you'll run on every dataset.- Distributions matter more than means — always plot a histogram before trusting
describe(). .corr()is the summary, the scatter is the truth — never accept correlation without seeing the shape.- Outliers are clues, not noise. Investigate before deleting.
- A short data dictionary before modelling pays for itself within a week.
- Auto-EDA tools like
ydata-profilingare useful starters but never the finish line.
You've now walked the full DS-fundamentals path — questions, Python idioms, NumPy, pandas, visualisation, EDA. The next paths take you deeper: Statistics for inference, Regression for your first models, and the full Data Science track for production-grade tooling.