Practice datasets
Seven datasets built for practicing on, not scraped from anywhere.
Each one has real structure in it — price genuinely depends on floor area, churn
genuinely depends on contract type — so a model you fit actually finds something.
One of them is deliberately filthy, because cleaning is the skill nobody teaches.
Load any of them in the notebook with one line.
Each CSV loads from this site straight into the Python running in your tab, and nothing you do with it is sent back.
df = pm.load("housing")
df.head()
Housing prices
600 rows × 9 cols · 22 KB
Six Pune neighbourhoods, 600 listings. Price is driven by floor area, bathrooms, age and parking, with a per-area premium on top.
Practice: Linear regression, feature engineering, one-hot encoding, residual plots.
idareasqftbedroomsbathroomsage_yearsparkingfloorprice_inr
df = pm.load("housing")
open the notebook
·
download the CSV
Daily weather
2,920 rows × 6 cols · 97 KB
Two years of daily readings for four cities with genuinely different climates, including a monsoon season that shows up in the rainfall.
Practice: Datetime parsing, resampling, groupby, rolling averages, seasonality.
datecitytemp_chumidity_pctrain_mmwind_kph
df = pm.load("weather")
open the notebook
·
download the CSV
Retail orders
1,400 rows × 10 cols · 94 KB
1,400 orders across four categories and three channels. Electronics skew to the app, books to the web — a pivot table will find it.
Practice: groupby, pivot_table, revenue aggregation, categorical analysis.
order_iddatecategoryproductregionchannelunitsunit_pricediscount_pctrevenue
df = pm.load("retail")
open the notebook
·
download the CSV
Films
500 rows × 8 cols · 27 KB
500 invented films with budgets, revenue and ratings. Revenue follows budget loosely, with a long tail of flops and outsized hits.
Practice: Filtering, correlation, log transforms, outlier hunting, plotting.
titleyeargenreruntime_minbudget_musdrevenue_musdratingvotes
df = pm.load("movies")
open the notebook
·
download the CSV
Signups (deliberately messy)messy on purpose
414 rows × 9 cols · 26 KB
400 rows carrying six kinds of real-world dirt on purpose: four date formats, inconsistent casing and whitespace, missing values, currency stored as text, impossible ages, and exact duplicate rows.
Practice: Everything cleaning: dropna, fillna, str.strip, str.lower, to_datetime, to_numeric, duplicated, range validation.
user_idemailnamesignup_dateplancountryagespendnewsletter
df = pm.load("signups")
open the notebook
·
download the CSV
Student results
800 rows × 8 cols · 24 KB
800 students. Final score is a real function of study hours, prior score, attendance, sleep and tutoring — plus noise you cannot model.
Practice: Regression, train/test split, classification on the pass/fail column.
student_idstudy_hours_weekattendance_pctprior_scoretutoringsleep_hoursfinal_scoreresult
df = pm.load("students")
open the notebook
·
download the CSV
Subscription churn
900 rows × 8 cols · 32 KB
900 customers. Month-to-month contracts and frequent support calls genuinely predict churn; long tenure genuinely protects against it.
Practice: Binary classification, class imbalance, confusion matrix, feature importance.
customer_idtenure_monthscontractmonthly_chargesupport_callspaperless_billingpayment_methodchurned
df = pm.load("churn")
open the notebook
·
download the CSV
Already built in
scikit-learn ships several classic datasets inside the package itself, so they
work here with no network at all. Those are not duplicated above — reach for
them directly:
from sklearn.datasets import load_iris, load_wine, load_diabetes, load_breast_cancer
X, y = load_iris(return_X_y=True, as_frame=True)
Every file above is generated from a seeded random number
generator. The rows are invented: no real person, listing, order or customer is
represented, and nothing here carries a licence you need to worry about.