Deploy Your Neural Network
1 · The lesson
readA model that lives in a Jupyter notebook is a science experiment. A model that lives behind an HTTPS endpoint and answers requests in 80ms is a product. The gap between the two is mostly plumbing, but it's the plumbing that determines whether your work gets used.
This lesson walks the chain — train → serialise → serve → monitor — and shows the practical decisions at each step.
The Keras snippets need TensorFlow. Run in Colab or
pip install tensorflow fastapi uvicorn. Outputs shown inline.
1. The Deployment Chain
trained model → serialise → serve → monitor (notebook) (.keras) (HTTP API) (logs, drift, alerts)
Each step is its own failure mode:
- Serialise wrong and the model won't load on the server.
- Serve without preprocessing and you produce confident garbage.
- Skip monitoring and you discover six months later that your input distribution shifted in March.
Every production ML team you've heard of spends most of their time on stages 2–4. Training is the easy part.
2. Save a Keras Model — One File
# After training model.save("sentiment.keras")
setup added so this can run · defines model
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) model = _AutoMock('model')
A single file, ~50–200 MB typical, containing:
- Architecture (the layer graph)
- Weights (the learned parameters)
- Optimiser state (so training can resume)
- Compile config (loss, metrics)
Loading is the mirror image, on any machine with the same major TF version:
from tensorflow import keras model = keras.models.load_model("sentiment.keras") preds = model.predict(some_input)
setup added so this can run · defines some_input
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) some_input = _AutoMock('some_input')
That's the minimum viable deploy artefact.
Pin your TensorFlow version in production. A model trained with TF 2.15 may not load cleanly in TF 2.17. Lock the version in
requirements.txt/pyproject.toml.
3. The Simplest Serving — A Python Script
Before the web framework, the simplest deploy is a script that reads from stdin and writes to stdout. Useful for batch jobs and as a smoke test.
# predict.py import json, sys import numpy as np from tensorflow import keras model = keras.models.load_model("sentiment.keras") for line in sys.stdin: payload = json.loads(line) text = payload["text"] score = float(model.predict([text], verbose=0)[0][0]) print(json.dumps({"score": score, "label": "positive" if score > 0.5 else "negative"}))
$ echo '{"text": "loved it"}' | python predict.py
{"score": 0.978, "label": "positive"}That's deployable. Not scalable, but deployable.
4. Serve Over HTTP — FastAPI in 25 Lines
For anything user-facing, you want an HTTP API. FastAPI is the modern Python default — async, type-checked, free OpenAPI docs.
# server.py from fastapi import FastAPI, HTTPException from pydantic import BaseModel from tensorflow import keras app = FastAPI() model = keras.models.load_model("sentiment.keras") class TextRequest(BaseModel): text: str class TextResponse(BaseModel): score: float label: str @app.post("/predict", response_model=TextResponse) def predict(req: TextRequest): if not req.text or len(req.text) > 5000: raise HTTPException(400, "text must be 1–5000 chars") score = float(model.predict([req.text], verbose=0)[0][0]) return TextResponse(score=score, label="positive" if score > 0.5 else "negative")
Run it:
$ uvicorn server:app --host 0.0.0.0 --port 8000
Hit it:
$ curl -X POST localhost:8000/predict \
-H "Content-Type: application/json" \
-d '{"text": "absolutely brilliant film"}'
# → {"score": 0.987, "label": "positive"}Pydantic validates the request shape for you — bad JSON gets a 422 before your model is even called. See exceptions for the broader pattern of validating at boundaries.
5. Preprocessing — Ship It With the Model
The single most common deployment bug is preprocessing drift. You normalise pixels by dividing by 255 in training. In production someone forgets, the model receives 0–255 inputs, and accuracy quietly collapses.
Two fixes:
Option A — Preprocessing lives in the model
Use Keras preprocessing layers (TextVectorization, Rescaling, Normalization) inside the Sequential model. They get saved as part of model.save() and run automatically at inference.
model = Sequential([ layers.Rescaling(1./255, input_shape=(28, 28, 1)), Conv2D(32, (3, 3), activation='relu'), # ... ])
setup added so this can run · defines Sequential, Conv2D, layers
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) def Sequential(*_a, **_kw): print('-> Sequential() called') return _AutoMock('Sequential()') def Conv2D(*_a, **_kw): print('-> Conv2D() called') return _AutoMock('Conv2D()') layers = _AutoMock('layers')
This is the modern best practice. The model is now a self-contained text → prediction or raw_pixels → prediction function.
Option B — sklearn-style Pipeline
If you have non-trivial Pythonic preprocessing (custom tokeniser, feature engineering), wrap everything in a sklearn.pipeline.Pipeline and pickle the whole thing, or use joblib.dump. Load both pipeline and model on the server.
from joblib import dump, load dump(preproc_pipeline, "preproc.joblib") # server-side pipeline = load("preproc.joblib") model = keras.models.load_model("model.keras") def predict(raw_input): features = pipeline.transform(raw_input) return model.predict(features)
setup added so this can run · defines preproc_pipeline, keras
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) preproc_pipeline = _AutoMock('preproc_pipeline') keras = _AutoMock('keras')
The cardinal rule either way: preprocessing code is part of the model artefact, not a side script.
6. Format Conversions — When You Outgrow .keras
.keras is fine for Python-based servers. When you need to run elsewhere, convert:
| Format | Use case |
|---|---|
| TF SavedModel (a directory) | TensorFlow Serving (high-throughput production) |
| TFLite | Mobile / embedded — quantised, tiny binary |
| ONNX | Run in a different framework (PyTorch runtime, .NET, Rust) |
| TF.js | Run directly in the browser |
| Core ML | iOS native |
Conversion is usually one function call:
# To TFLite for mobile converter = tf.lite.TFLiteConverter.from_keras_model(model) tflite_model = converter.convert() open("model.tflite", "wb").write(tflite_model) # To ONNX (needs tf2onnx) # python -m tf2onnx.convert --keras model.keras --output model.onnx
setup added so this can run · defines model, tf
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) model = _AutoMock('model') tf = _AutoMock('tf')
Don't convert until you have a concrete reason — every conversion is one more thing to test for numerical equivalence with the original.
7. Performance Levers
The cheapest first wins:
- Batch requests —
model.predict(X_batch)is far faster thanN × model.predict(x_single). On busy servers, accumulate incoming requests for a few ms and process as a batch. - GPU vs CPU — GPUs win on batched inference and large models, CPUs win on small models + low latency. Profile before paying for a GPU.
- Quantisation — replace 32-bit floats with 8-bit integers. Often 4× faster, 4× smaller, with sub-1% accuracy hit. Built into TFLite.
- Pruning — zero out small weights. Less common in deploy than quantisation.
- Use a smaller model — distill the trained model into a small one. Often the biggest win, and the most overlooked.
8. Hosting Options
| Platform | Best for | Cost shape |
|---|---|---|
| AWS SageMaker | Enterprise, autoscaling, integration with the rest of AWS | Per-instance-hour |
| GCP Vertex AI | Same idea, Google's stack | Per-instance-hour |
| HuggingFace Spaces | Demos, prototypes, public models | Free tier, generous |
| Modal / Replicate | Serverless GPU, pay-per-request | Per-second of compute |
| FastAPI on a VPS (DigitalOcean, Hetzner) | Cheap, simple, full control | Flat monthly |
| Cloudflare Workers + ONNX | Edge inference, very low latency | Per-request |
For your first deploy: HuggingFace Spaces for a demo, FastAPI on a VPS for a real product with predictable traffic, Modal if traffic is spiky.
9. Containerise — A Minimal Dockerfile
Reproducibility means shipping the model with its exact runtime. Docker is the modern unit.
FROM python:3.11-slim WORKDIR /app # Pin everything COPY requirements.txt . RUN pip install --no-cache-dir -r requirements.txt # Model artefact and code COPY sentiment.keras . COPY server.py . EXPOSE 8000 CMD ["uvicorn", "server:app", "--host", "0.0.0.0", "--port", "8000"]
requirements.txt:
tensorflow==2.17.0 fastapi==0.115.0 uvicorn==0.30.0 pydantic==2.9.0
Build, run, deploy:
$ docker build -t sentiment-api . $ docker run -p 8000:8000 sentiment-api
Pin every version. "It works on my laptop" is not a deployment strategy.
10. Monitoring — The Part Everyone Skips
Once the model is serving, you need to know when it stops being useful. Three things to log:
- Inputs — sample a percentage of incoming requests. You'll need them to debug regressions and detect drift.
- Predictions — the model's output and confidence. Watch the distribution shift over time.
- Outcomes — when you can — the eventual ground truth (the user clicked / didn't click, the doctor agreed / disagreed). This closes the loop on real accuracy.
Drift detection: compare the distribution of recent inputs to the training distribution. If your model was trained on 2024 product reviews and 2026 reviews are talking about features that didn't exist, the model is operating out of its training distribution. Several open-source tools (Evidently AI, WhyLabs, Arize) automate this.
11. The Production Checklist
Before you call a model "deployed":
- [ ] Reproducible build — pinned versions, single command to rebuild the image
- [ ] Preprocessing inside the artefact — no preprocessing scripts that could drift
- [ ] Input validation at the boundary — clear errors for bad input, not crashes
- [ ] Versioned model file —
sentiment-v3.2.keras, notmodel.keras - [ ] Versioned data — you can re-derive the training set six months from now
- [ ] Health endpoint —
GET /healthreturns 200 if the model loaded - [ ] Logged inputs and predictions — at least sampled
- [ ] Drift monitor — even a weekly summary email beats nothing
- [ ] Fallback path — what happens when the model service is down? (Cache last result? Return a default? Fail the request?)
- [ ] Rollback plan — can you revert to the previous model in one command?
A model without these isn't deployed; it's hoping.
Common Mistakes
- Forgetting to ship preprocessing. Training divides by 255, production doesn't. Predictions look "almost right" — the worst kind of bug.
- No input validation. Empty strings, 100 MB images, malformed JSON — your
/predictendpoint will see all of them. Validate at the boundary and return clear errors (see exceptions). - Unpinned dependencies. A
pip install tensorflowsix months later picks up a new major version and your model file no longer loads. Pin versions; rebuild rarely; test that rebuilds still pass. - No monitoring. Models silently degrade with input drift. You won't know until someone complains, which can mean months of bad predictions in the meantime.
- Optimising before measuring. Quantising a model that's already serving in 20ms with 99th-percentile latency under 100ms is wasted effort. Measure first.
🎯 Your Turn — A predict.py Script
Write a predict.py that:
1. Loads a saved Keras model from model.keras.
2. Reads a single JSON object from stdin: {"text": "..."}.
3. Validates that text is present, is a string, and is 1–5000 characters.
4. Calls model.predict and writes a JSON response to stdout: {"score": float, "label": "positive" | "negative"}.
5. On validation failure, writes {"error": "<message>"} to stdout and exits with code 1.
Skeleton:
# predict.py import json import sys from tensorflow import keras # TODO 1: load the model def predict_one(payload): # TODO 2: validate payload["text"] exists, is a string, length 1–5000 # TODO 3: run model.predict and build the response dict pass if __name__ == "__main__": raw = sys.stdin.read() try: payload = json.loads(raw) result = predict_one(payload) print(json.dumps(result)) except Exception as e: # TODO 4: print error JSON, exit 1 ...
Hint 1 — Validation
Check three things in order:"text" in payload, isinstance(payload["text"], str), 1 <= len(payload["text"]) <= 5000. Raise ValueError with a useful message on each failure — the except block formats it as JSON.
Hint 2 — Loading once
keras.models.load_model at module level (not inside predict_one) so the script doesn't reload the model on every invocation. In a real server, you do this once at startup.
Show full solution
# predict.py import json import sys from tensorflow import keras MODEL_PATH = "model.keras" MAX_TEXT_LEN = 5000 model = keras.models.load_model(MODEL_PATH) def predict_one(payload: dict) -> dict: if "text" not in payload: raise ValueError("missing field: text") text = payload["text"] if not isinstance(text, str): raise ValueError("text must be a string") if not (1 <= len(text) <= MAX_TEXT_LEN): raise ValueError(f"text length must be 1–{MAX_TEXT_LEN}") score = float(model.predict([text], verbose=0)[0][0]) return { "score": round(score, 4), "label": "positive" if score > 0.5 else "negative", } if __name__ == "__main__": raw = sys.stdin.read() try: payload = json.loads(raw) result = predict_one(payload) print(json.dumps(result)) except (json.JSONDecodeError, ValueError) as e: print(json.dumps({"error": str(e)})) sys.exit(1)
Use it:
$ echo '{"text": "absolutely loved it"}' | python predict.py
{"score": 0.9881, "label": "positive"}
$ echo '{}' | python predict.py
{"error": "missing field: text"}
$ echo 'not even json' | python predict.py
{"error": "Expecting value: line 1 column 1 (char 0)"}That's a minimum-viable production interface: deterministic, validated, JSON-in / JSON-out, sensible exit codes. Wrap it in the FastAPI shell from Section 4 and you have an HTTP endpoint ready for a Dockerfile.
What You Learned
- The chain: train → serialise (
.keras) → serve (FastAPI) → monitor. - Preprocessing ships with the model — Keras preprocessing layers or a pickled sklearn
Pipeline. - Validate at the boundary: Pydantic in FastAPI, explicit checks in scripts.
- Pin versions in
requirements.txtand a Dockerfile — unpinned ML stacks rot fast. - Monitor inputs and predictions for drift; have a fallback and rollback plan.
- Pick a host based on traffic shape: VPS for steady, Modal/Replicate for spiky, SageMaker/Vertex for enterprise scale.
That closes the deep learning fundamentals path — neuron, framework, training, vision, language, deploy. Next stops are the specialty tracks: deeper into CNNs, RNNs, transfer learning, transformers, and the ethics (ethics) you should think about before any of these ship to real users.