PythonMastery
beginner 24 min read · lesson 7 of 7 in Deep Learning Fundamentals

Deploy Your Neural Network

1 · The lesson

read

A model that lives in a Jupyter notebook is a science experiment. A model that lives behind an HTTPS endpoint and answers requests in 80ms is a product. The gap between the two is mostly plumbing, but it's the plumbing that determines whether your work gets used.

This lesson walks the chain — train → serialise → serve → monitor — and shows the practical decisions at each step.

The Keras snippets need TensorFlow. Run in Colab or pip install tensorflow fastapi uvicorn. Outputs shown inline.


1. The Deployment Chain

python
trained model  →  serialise  →  serve  →  monitor
   (notebook)      (.keras)    (HTTP API)   (logs, drift, alerts)

Each step is its own failure mode:

  • Serialise wrong and the model won't load on the server.
  • Serve without preprocessing and you produce confident garbage.
  • Skip monitoring and you discover six months later that your input distribution shifted in March.

Every production ML team you've heard of spends most of their time on stages 2–4. Training is the easy part.


2. Save a Keras Model — One File

python
# After training
model.save("sentiment.keras")
+ setup added so this can run · defines model
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

model = _AutoMock('model')

A single file, ~50–200 MB typical, containing:

  • Architecture (the layer graph)
  • Weights (the learned parameters)
  • Optimiser state (so training can resume)
  • Compile config (loss, metrics)

Loading is the mirror image, on any machine with the same major TF version:

python
from tensorflow import keras
model = keras.models.load_model("sentiment.keras")
preds = model.predict(some_input)
+ setup added so this can run · defines some_input
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

some_input = _AutoMock('some_input')

That's the minimum viable deploy artefact.

Pin your TensorFlow version in production. A model trained with TF 2.15 may not load cleanly in TF 2.17. Lock the version in requirements.txt / pyproject.toml.


3. The Simplest Serving — A Python Script

Before the web framework, the simplest deploy is a script that reads from stdin and writes to stdout. Useful for batch jobs and as a smoke test.

python
# predict.py
import json, sys
import numpy as np
from tensorflow import keras

model = keras.models.load_model("sentiment.keras")

for line in sys.stdin:
    payload = json.loads(line)
    text = payload["text"]
    score = float(model.predict([text], verbose=0)[0][0])
    print(json.dumps({"score": score, "label": "positive" if score > 0.5 else "negative"}))
bash
$ echo '{"text": "loved it"}' | python predict.py
{"score": 0.978, "label": "positive"}

That's deployable. Not scalable, but deployable.


4. Serve Over HTTP — FastAPI in 25 Lines

For anything user-facing, you want an HTTP API. FastAPI is the modern Python default — async, type-checked, free OpenAPI docs.

python
# server.py
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel
from tensorflow import keras

app = FastAPI()
model = keras.models.load_model("sentiment.keras")

class TextRequest(BaseModel):
    text: str

class TextResponse(BaseModel):
    score: float
    label: str

@app.post("/predict", response_model=TextResponse)
def predict(req: TextRequest):
    if not req.text or len(req.text) > 5000:
        raise HTTPException(400, "text must be 1–5000 chars")
    score = float(model.predict([req.text], verbose=0)[0][0])
    return TextResponse(score=score,
                        label="positive" if score > 0.5 else "negative")

Run it:

bash
$ uvicorn server:app --host 0.0.0.0 --port 8000

Hit it:

bash
$ curl -X POST localhost:8000/predict \
    -H "Content-Type: application/json" \
    -d '{"text": "absolutely brilliant film"}'
# → {"score": 0.987, "label": "positive"}

Pydantic validates the request shape for you — bad JSON gets a 422 before your model is even called. See exceptions for the broader pattern of validating at boundaries.


5. Preprocessing — Ship It With the Model

The single most common deployment bug is preprocessing drift. You normalise pixels by dividing by 255 in training. In production someone forgets, the model receives 0–255 inputs, and accuracy quietly collapses.

Two fixes:

Option A — Preprocessing lives in the model

Use Keras preprocessing layers (TextVectorization, Rescaling, Normalization) inside the Sequential model. They get saved as part of model.save() and run automatically at inference.

python
model = Sequential([
    layers.Rescaling(1./255, input_shape=(28, 28, 1)),
    Conv2D(32, (3, 3), activation='relu'),
    # ...
])
+ setup added so this can run · defines Sequential, Conv2D, layers
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def Sequential(*_a, **_kw):
    print('-> Sequential() called')
    return _AutoMock('Sequential()')
def Conv2D(*_a, **_kw):
    print('-> Conv2D() called')
    return _AutoMock('Conv2D()')
layers = _AutoMock('layers')

This is the modern best practice. The model is now a self-contained text → prediction or raw_pixels → prediction function.

Option B — sklearn-style Pipeline

If you have non-trivial Pythonic preprocessing (custom tokeniser, feature engineering), wrap everything in a sklearn.pipeline.Pipeline and pickle the whole thing, or use joblib.dump. Load both pipeline and model on the server.

python
from joblib import dump, load
dump(preproc_pipeline, "preproc.joblib")

# server-side
pipeline = load("preproc.joblib")
model    = keras.models.load_model("model.keras")

def predict(raw_input):
    features = pipeline.transform(raw_input)
    return model.predict(features)
+ setup added so this can run · defines preproc_pipeline, keras
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

preproc_pipeline = _AutoMock('preproc_pipeline')
keras = _AutoMock('keras')

The cardinal rule either way: preprocessing code is part of the model artefact, not a side script.


6. Format Conversions — When You Outgrow .keras

.keras is fine for Python-based servers. When you need to run elsewhere, convert:

FormatUse case
TF SavedModel (a directory)TensorFlow Serving (high-throughput production)
TFLiteMobile / embedded — quantised, tiny binary
ONNXRun in a different framework (PyTorch runtime, .NET, Rust)
TF.jsRun directly in the browser
Core MLiOS native

Conversion is usually one function call:

python
# To TFLite for mobile
converter = tf.lite.TFLiteConverter.from_keras_model(model)
tflite_model = converter.convert()
open("model.tflite", "wb").write(tflite_model)

# To ONNX (needs tf2onnx)
# python -m tf2onnx.convert --keras model.keras --output model.onnx
+ setup added so this can run · defines model, tf
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

model = _AutoMock('model')
tf = _AutoMock('tf')

Don't convert until you have a concrete reason — every conversion is one more thing to test for numerical equivalence with the original.


7. Performance Levers

The cheapest first wins:

  • Batch requests — model.predict(X_batch) is far faster than N × model.predict(x_single). On busy servers, accumulate incoming requests for a few ms and process as a batch.
  • GPU vs CPU — GPUs win on batched inference and large models, CPUs win on small models + low latency. Profile before paying for a GPU.
  • Quantisation — replace 32-bit floats with 8-bit integers. Often 4× faster, 4× smaller, with sub-1% accuracy hit. Built into TFLite.
  • Pruning — zero out small weights. Less common in deploy than quantisation.
  • Use a smaller model — distill the trained model into a small one. Often the biggest win, and the most overlooked.

8. Hosting Options

PlatformBest forCost shape
AWS SageMakerEnterprise, autoscaling, integration with the rest of AWSPer-instance-hour
GCP Vertex AISame idea, Google's stackPer-instance-hour
HuggingFace SpacesDemos, prototypes, public modelsFree tier, generous
Modal / ReplicateServerless GPU, pay-per-requestPer-second of compute
FastAPI on a VPS (DigitalOcean, Hetzner)Cheap, simple, full controlFlat monthly
Cloudflare Workers + ONNXEdge inference, very low latencyPer-request

For your first deploy: HuggingFace Spaces for a demo, FastAPI on a VPS for a real product with predictable traffic, Modal if traffic is spiky.


9. Containerise — A Minimal Dockerfile

Reproducibility means shipping the model with its exact runtime. Docker is the modern unit.

dockerfile
FROM python:3.11-slim

WORKDIR /app

# Pin everything
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt

# Model artefact and code
COPY sentiment.keras .
COPY server.py .

EXPOSE 8000
CMD ["uvicorn", "server:app", "--host", "0.0.0.0", "--port", "8000"]

requirements.txt:

python
tensorflow==2.17.0
fastapi==0.115.0
uvicorn==0.30.0
pydantic==2.9.0

Build, run, deploy:

bash
$ docker build -t sentiment-api .
$ docker run -p 8000:8000 sentiment-api

Pin every version. "It works on my laptop" is not a deployment strategy.


10. Monitoring — The Part Everyone Skips

Once the model is serving, you need to know when it stops being useful. Three things to log:

  • Inputs — sample a percentage of incoming requests. You'll need them to debug regressions and detect drift.
  • Predictions — the model's output and confidence. Watch the distribution shift over time.
  • Outcomes — when you can — the eventual ground truth (the user clicked / didn't click, the doctor agreed / disagreed). This closes the loop on real accuracy.

Drift detection: compare the distribution of recent inputs to the training distribution. If your model was trained on 2024 product reviews and 2026 reviews are talking about features that didn't exist, the model is operating out of its training distribution. Several open-source tools (Evidently AI, WhyLabs, Arize) automate this.


11. The Production Checklist

Before you call a model "deployed":

  • [ ] Reproducible build — pinned versions, single command to rebuild the image
  • [ ] Preprocessing inside the artefact — no preprocessing scripts that could drift
  • [ ] Input validation at the boundary — clear errors for bad input, not crashes
  • [ ] Versioned model file — sentiment-v3.2.keras, not model.keras
  • [ ] Versioned data — you can re-derive the training set six months from now
  • [ ] Health endpoint — GET /health returns 200 if the model loaded
  • [ ] Logged inputs and predictions — at least sampled
  • [ ] Drift monitor — even a weekly summary email beats nothing
  • [ ] Fallback path — what happens when the model service is down? (Cache last result? Return a default? Fail the request?)
  • [ ] Rollback plan — can you revert to the previous model in one command?

A model without these isn't deployed; it's hoping.


Common Mistakes

  • Forgetting to ship preprocessing. Training divides by 255, production doesn't. Predictions look "almost right" — the worst kind of bug.
  • No input validation. Empty strings, 100 MB images, malformed JSON — your /predict endpoint will see all of them. Validate at the boundary and return clear errors (see exceptions).
  • Unpinned dependencies. A pip install tensorflow six months later picks up a new major version and your model file no longer loads. Pin versions; rebuild rarely; test that rebuilds still pass.
  • No monitoring. Models silently degrade with input drift. You won't know until someone complains, which can mean months of bad predictions in the meantime.
  • Optimising before measuring. Quantising a model that's already serving in 20ms with 99th-percentile latency under 100ms is wasted effort. Measure first.

🎯 Your Turn — A predict.py Script

Write a predict.py that:

1. Loads a saved Keras model from model.keras.
2. Reads a single JSON object from stdin: {"text": "..."}.
3. Validates that text is present, is a string, and is 1–5000 characters.
4. Calls model.predict and writes a JSON response to stdout: {"score": float, "label": "positive" | "negative"}.
5. On validation failure, writes {"error": "<message>"} to stdout and exits with code 1.

Skeleton:

python
# predict.py
import json
import sys
from tensorflow import keras

# TODO 1: load the model

def predict_one(payload):
    # TODO 2: validate payload["text"] exists, is a string, length 1–5000
    # TODO 3: run model.predict and build the response dict
    pass

if __name__ == "__main__":
    raw = sys.stdin.read()
    try:
        payload = json.loads(raw)
        result = predict_one(payload)
        print(json.dumps(result))
    except Exception as e:
        # TODO 4: print error JSON, exit 1
        ...
Hint 1 — Validation Check three things in order: "text" in payload, isinstance(payload["text"], str), 1 <= len(payload["text"]) <= 5000. Raise ValueError with a useful message on each failure — the except block formats it as JSON.
Hint 2 — Loading once keras.models.load_model at module level (not inside predict_one) so the script doesn't reload the model on every invocation. In a real server, you do this once at startup.
Show full solution
python
# predict.py
import json
import sys
from tensorflow import keras

MODEL_PATH   = "model.keras"
MAX_TEXT_LEN = 5000

model = keras.models.load_model(MODEL_PATH)

def predict_one(payload: dict) -> dict:
    if "text" not in payload:
        raise ValueError("missing field: text")
    text = payload["text"]
    if not isinstance(text, str):
        raise ValueError("text must be a string")
    if not (1 <= len(text) <= MAX_TEXT_LEN):
        raise ValueError(f"text length must be 1–{MAX_TEXT_LEN}")

    score = float(model.predict([text], verbose=0)[0][0])
    return {
        "score": round(score, 4),
        "label": "positive" if score > 0.5 else "negative",
    }

if __name__ == "__main__":
    raw = sys.stdin.read()
    try:
        payload = json.loads(raw)
        result = predict_one(payload)
        print(json.dumps(result))
    except (json.JSONDecodeError, ValueError) as e:
        print(json.dumps({"error": str(e)}))
        sys.exit(1)

Use it:

bash
$ echo '{"text": "absolutely loved it"}' | python predict.py
{"score": 0.9881, "label": "positive"}

$ echo '{}' | python predict.py
{"error": "missing field: text"}

$ echo 'not even json' | python predict.py
{"error": "Expecting value: line 1 column 1 (char 0)"}

That's a minimum-viable production interface: deterministic, validated, JSON-in / JSON-out, sensible exit codes. Wrap it in the FastAPI shell from Section 4 and you have an HTTP endpoint ready for a Dockerfile.


What You Learned

  • The chain: train → serialise (.keras) → serve (FastAPI) → monitor.
  • Preprocessing ships with the model — Keras preprocessing layers or a pickled sklearn Pipeline.
  • Validate at the boundary: Pydantic in FastAPI, explicit checks in scripts.
  • Pin versions in requirements.txt and a Dockerfile — unpinned ML stacks rot fast.
  • Monitor inputs and predictions for drift; have a fallback and rollback plan.
  • Pick a host based on traffic shape: VPS for steady, Modal/Replicate for spiky, SageMaker/Vertex for enterprise scale.

That closes the deep learning fundamentals path — neuron, framework, training, vision, language, deploy. Next stops are the specialty tracks: deeper into CNNs, RNNs, transfer learning, transformers, and the ethics (ethics) you should think about before any of these ship to real users.