PythonMastery

Practice datasets

Seven datasets built for practicing on, not scraped from anywhere. Each one has real structure in it — price genuinely depends on floor area, churn genuinely depends on contract type — so a model you fit actually finds something. One of them is deliberately filthy, because cleaning is the skill nobody teaches.

Load any of them in the notebook with one line. Each CSV loads from this site straight into the Python running in your tab, and nothing you do with it is sent back.

df = pm.load("housing")
df.head()

Housing prices

600 rows × 9 cols · 22 KB

Six Pune neighbourhoods, 600 listings. Price is driven by floor area, bathrooms, age and parking, with a per-area premium on top.

Practice: Linear regression, feature engineering, one-hot encoding, residual plots.

idareasqftbedroomsbathroomsage_yearsparkingfloorprice_inr
df = pm.load("housing")

Daily weather

2,920 rows × 6 cols · 97 KB

Two years of daily readings for four cities with genuinely different climates, including a monsoon season that shows up in the rainfall.

Practice: Datetime parsing, resampling, groupby, rolling averages, seasonality.

datecitytemp_chumidity_pctrain_mmwind_kph
df = pm.load("weather")

Retail orders

1,400 rows × 10 cols · 94 KB

1,400 orders across four categories and three channels. Electronics skew to the app, books to the web — a pivot table will find it.

Practice: groupby, pivot_table, revenue aggregation, categorical analysis.

order_iddatecategoryproductregionchannelunitsunit_pricediscount_pctrevenue
df = pm.load("retail")

Films

500 rows × 8 cols · 27 KB

500 invented films with budgets, revenue and ratings. Revenue follows budget loosely, with a long tail of flops and outsized hits.

Practice: Filtering, correlation, log transforms, outlier hunting, plotting.

titleyeargenreruntime_minbudget_musdrevenue_musdratingvotes
df = pm.load("movies")

Signups (deliberately messy)messy on purpose

414 rows × 9 cols · 26 KB

400 rows carrying six kinds of real-world dirt on purpose: four date formats, inconsistent casing and whitespace, missing values, currency stored as text, impossible ages, and exact duplicate rows.

Practice: Everything cleaning: dropna, fillna, str.strip, str.lower, to_datetime, to_numeric, duplicated, range validation.

user_idemailnamesignup_dateplancountryagespendnewsletter
df = pm.load("signups")

Student results

800 rows × 8 cols · 24 KB

800 students. Final score is a real function of study hours, prior score, attendance, sleep and tutoring — plus noise you cannot model.

Practice: Regression, train/test split, classification on the pass/fail column.

student_idstudy_hours_weekattendance_pctprior_scoretutoringsleep_hoursfinal_scoreresult
df = pm.load("students")

Subscription churn

900 rows × 8 cols · 32 KB

900 customers. Month-to-month contracts and frequent support calls genuinely predict churn; long tenure genuinely protects against it.

Practice: Binary classification, class imbalance, confusion matrix, feature importance.

customer_idtenure_monthscontractmonthly_chargesupport_callspaperless_billingpayment_methodchurned
df = pm.load("churn")

Already built in

scikit-learn ships several classic datasets inside the package itself, so they work here with no network at all. Those are not duplicated above — reach for them directly:

from sklearn.datasets import load_iris, load_wine, load_diabetes, load_breast_cancer
X, y = load_iris(return_X_y=True, as_frame=True)

Every file above is generated from a seeded random number generator. The rows are invented: no real person, listing, order or customer is represented, and nothing here carries a licence you need to worry about.