PythonMastery
beginner 14 min read · lesson 1 of 6 in Data Science Fundamentals

What is Data Science?

1 · The lesson

read

Data science is the discipline of extracting knowledge from data — turning raw rows and columns into decisions, predictions, and arguments people will actually act on. It is not just statistics, not just machine learning, and not just programming. It is the working combination of all three with a question in front of it.

If you can write a Python for loop, you already have the smallest piece. This lesson maps the rest of the territory before you touch a single NumPy array.


1. The Working Definition

A data scientist takes a vague question — "why are users churning?", "what should we price this at?", "will this transaction be fraud?" — and answers it with evidence. The evidence comes from data; the answer takes the form of a number, a chart, a model, or a recommendation.

Three ingredients show up every time:

  • A question that someone cares about.
  • Data that plausibly relates to the question.
  • A method — descriptive, statistical, or predictive — that connects the two.

Drop any one and you don't have data science. A model with no question is a toy. A question with no data is a guess. Data and a question without method is a hunch.


2. The Workflow Loop

Real projects loop, they don't march. The typical cycle:

text
       ┌─────────────┐
       │   1. Ask    │  define the question, the success metric
       └──────┬──────┘
              ▼
       ┌─────────────┐
       │  2. Get     │  load CSVs, query databases, scrape, call APIs
       └──────┬──────┘
              ▼
       ┌─────────────┐
       │  3. Clean   │  fix types, handle missing, dedupe
       └──────┬──────┘
              ▼
       ┌─────────────┐
       │  4. Explore │  plot, summarise, form hypotheses
       └──────┬──────┘
              ▼
       ┌─────────────┐
       │  5. Model   │  statistics, ML, simulation — if needed
       └──────┬──────┘
              ▼
       ┌─────────────┐
       │ 6. Communicate │  chart, deck, dashboard, PR
       └──────┬──────┘
              │
              └──► back to 1 with sharper questions

Junior practitioners spend 80% of their time on 2-4 and underestimate how much of the value lives in 1 and 6. A correct model nobody reads changes nothing.


3. DS vs Data Engineering vs ML vs Analytics

These titles overlap and every company defines them differently. A rough map:

RolePrimary outputTooling leanTime horizon
Data AnalystDashboards, reports, ad-hoc answersSQL, Excel, BI toolsDays
Data ScientistModels, experiments, decisionsPython, stats, MLWeeks
ML EngineerProduction models in servicesPython, Docker, MLOpsMonths
Data EngineerPipelines, warehouses, tablesSQL, Spark, AirflowQuarters

A "full-stack" DS does pieces of all four. At a small company, that's the job. At a big one, you specialise.


4. A Day in the Life

A typical DS day, demystified:

  • Morning: open a Jupyter notebook, run a SQL query against the warehouse, eyeball results, ask the PM a clarifying question on Slack.
  • Mid-morning: clean the data, build three plots, notice something weird in a category, dig in.
  • Afternoon: fit a baseline model — usually something boring like logistic regression — to see if the signal is even there.
  • Late afternoon: write up the finding in a doc with one chart and three bullet points. Push to GitHub.
  • Following day: iterate on whichever question the writeup triggered.

Note what's missing: dramatic ML breakthroughs, six-monitor setups, leetcode. Most of the work is patient looking at data.


5. The Python Data Science Stack

You'll meet these in roughly this order across the next five lessons:

LibraryWhat it doesWhen you reach for it
NumPyFast numeric arraysAnything mathy on numbers
pandasLabelled tables (DataFrames)Loading, cleaning, grouping CSVs
matplotlib / seabornStatic chartsExploring and communicating
scikit-learnClassical ML — regression, trees, clusteringPredictive modelling
statsmodelsStatistical models with full diagnosticsInference, hypothesis testing
PyTorch / TensorFlowDeep learningImages, text, audio

Everything else — Plotly, Polars, PySpark, XGBoost, LangChain — is an alternative or specialisation built on the same mental model. Master the shortlist above and the rest is vocabulary.


6. A Tiny First Example — No NumPy Required

Before any specialty library, the standard library can already do useful descriptive work. The statistics module ships with Python:

python
import statistics

monthly_sales = {
    "Jan": 12_400,
    "Feb": 9_800,
    "Mar": 15_200,
    "Apr": 11_700,
    "May": 18_900,
}

values = list(monthly_sales.values())

print(f"Mean:   {statistics.mean(values):.0f}")        # 13600
print(f"Median: {statistics.median(values):.0f}")      # 12400
print(f"Stdev:  {statistics.stdev(values):.0f}")       # 3601
print(f"Max:    {max(values)}  ({max(monthly_sales, key=monthly_sales.get)})")
print(f"Min:    {min(values)}  ({min(monthly_sales, key=monthly_sales.get)})")

That's data science in miniature — a dataset, a question ("how is each month doing?"), a method (summary stats), and an interpretable answer. You'll graduate to NumPy and pandas in the next lessons; the thinking doesn't change.


Common Mistakes

  • Jumping to ML before exploring the data. A scatter plot can answer 60% of business questions. Train the model only when looking isn't enough.
  • Treating DS as just programming. Clean code with no domain understanding produces wrong answers, prettily. Always learn the business context.
  • Skipping step 1. "Build me a churn model" is not a question. "Which signals predict churn within 30 days for users on the free tier?" is. Force precision early.
  • Drowning in tools. You don't need to know Spark, Airflow, dbt, and three vector databases on day one. NumPy + pandas + matplotlib covers most of the job.
  • Confusing correlation with causation. Two columns moving together does not mean one drives the other. Reach for statistics before claiming "X causes Y".

🎯 Your Turn — Build a summarise() Function

Write a function that takes a list of numbers and returns the five summary stats every DS computes in their sleep. Use only the standard library — no NumPy yet.

Skeleton:

python
import statistics

def summarise(data):
    """Return {mean, median, min, max, std} for a list of numbers."""
    # TODO 1: handle the empty list case — return None or raise
    # TODO 2: build and return the dict
    ...

# Test
print(summarise([4, 8, 15, 16, 23, 42]))
# Expected: {'mean': 18.0, 'median': 15.5, 'min': 4, 'max': 42, 'std': 13.34...}
Hint 1 — Which functions? The statistics module has mean, median, and stdev. Min and max are Python built-ins. Round the float values to keep your output readable.
Hint 2 — Edge cases statistics.stdev needs at least 2 values and raises StatisticsError on shorter lists. Decide what your function does for 0 or 1 elements — returning None or an empty dict is fine for now.
Show full solution
python
import statistics

def summarise(data):
    """Return {mean, median, min, max, std} for a list of numbers."""
    if not data:
        return None
    return {
        "mean":   round(statistics.mean(data), 2),
        "median": statistics.median(data),
        "min":    min(data),
        "max":    max(data),
        "std":    round(statistics.stdev(data), 2) if len(data) > 1 else 0.0,
    }

print(summarise([4, 8, 15, 16, 23, 42]))
# {'mean': 18.0, 'median': 15.5, 'min': 4, 'max': 42, 'std': 13.34}

print(summarise([]))     # None
print(summarise([7]))    # {'mean': 7.0, 'median': 7, 'min': 7, 'max': 7, 'std': 0.0}

You've just written your first reusable EDA helper — one that you'll throw away in two lessons when pandas gives you .describe() for free. That's fine. Knowing what .describe() is doing under the hood is the whole point of writing this version first.


What You Learned

  • Data science is question + data + method — drop one and it isn't DS.
  • The workflow is a loop: ask → get → clean → explore → model → communicate → repeat.
  • DS, DA, MLE, and DE overlap but lean toward different outputs and time horizons.
  • The Python stack centres on NumPy, pandas, matplotlib, scikit-learn with deep learning on top.
  • Even before any library, the stdlib statistics module covers basic descriptive work.

Next: Python for Data Analysis — the idioms and standard-library tools that make Python feel built for this work.