PythonMastery
advanced 28 min read · lesson 7 of 9 in AI & Deep Learning

GANs & Generative Models: From Adversarial Training to Diffusion

1 · The lesson

read

Runtime note — Keras GAN training and Hugging Face diffusers Stable Diffusion both require a local Python environment with a GPU for serious work. CPU runs the diffusers example but takes minutes per image. Neither runs in the Pyodide sandbox.

Discriminative models answer "what is this?". Generative models answer "what does data from this distribution look like?". Until 2014 the answers were either too blurry to be useful (VAEs of the era) or too constrained (autoregressive pixel models, ~30 frames per minute on the day's hardware). Then Ian Goodfellow proposed pitting two networks against each other in a game, and a decade of remarkable images followed. Then — quietly, around 2021 — diffusion models replaced GANs as state-of-the-art for almost everything image-related.

This lesson covers the GAN game, its messy training dynamics, the line of improvements that followed, the principled-but-blurry VAE alternative, and the diffusion models that ate the field. The honest 2026 verdict is at the end.


1. The GAN Game

A GAN has two networks playing zero-sum:

  • Generator G takes a random vector z (the "latent") and tries to produce a sample that looks like real training data.
  • Discriminator D takes a sample (real or generated) and outputs a probability that it's real.

You train them alternately. D's gradient teaches it to spot the fakes. G's gradient teaches it to produce fakes that fool D. As D improves, G is forced to produce better fakes; as G improves, D is forced to look closer; they spiral upward together until — in the ideal case — G's outputs are indistinguishable from real data and D is stuck guessing 50/50.

That's the entire idea. No likelihood function, no explicit density model, no maximum-likelihood objective. Just two networks improving by mutual exploitation.

The Min-Max Intuition (No Proofs)

The training objective written formally is a min-max game:

python
min_G max_D  E_real[log D(x)] + E_fake[log(1 - D(G(z)))]

D is trying to maximise the expression (correctly label reals as 1 and fakes as 0). G is trying to minimise it (make D label fakes as 1). At equilibrium, G's distribution matches the data distribution and D outputs 0.5 everywhere. In practice you don't reach the equilibrium; you reach a configuration that's good enough and stop before something explodes.


2. A Vanilla GAN in Keras

The structure for a small image dataset (think MNIST-scale):

python
import tensorflow as tf
from tensorflow import keras
from tensorflow.keras import layers

LATENT_DIM = 100

def build_generator():
    return keras.Sequential([
        layers.Dense(7 * 7 * 128, input_shape=(LATENT_DIM,)),
        layers.LeakyReLU(0.2),
        layers.Reshape((7, 7, 128)),
        layers.Conv2DTranspose(128, 4, strides=2, padding="same"),
        layers.LeakyReLU(0.2),
        layers.Conv2DTranspose(128, 4, strides=2, padding="same"),
        layers.LeakyReLU(0.2),
        layers.Conv2D(1, 7, padding="same", activation="tanh"),
    ])

def build_discriminator():
    return keras.Sequential([
        layers.Conv2D(64, 4, strides=2, padding="same", input_shape=(28, 28, 1)),
        layers.LeakyReLU(0.2),
        layers.Conv2D(128, 4, strides=2, padding="same"),
        layers.LeakyReLU(0.2),
        layers.Flatten(),
        layers.Dropout(0.3),
        layers.Dense(1, activation="sigmoid"),
    ])

generator = build_generator()
discriminator = build_discriminator()

# Two separate optimisers — historically the most common bug is sharing them
g_opt = keras.optimizers.Adam(1e-4, beta_1=0.5)
d_opt = keras.optimizers.Adam(1e-4, beta_1=0.5)
bce = keras.losses.BinaryCrossentropy()

@tf.function
def train_step(real_images):
    batch = tf.shape(real_images)[0]
    z = tf.random.normal((batch, LATENT_DIM))

    # 1. Train discriminator on real + fake
    with tf.GradientTape() as tape:
        fake_images = generator(z, training=True)
        d_real = discriminator(real_images, training=True)
        d_fake = discriminator(fake_images, training=True)
        d_loss = bce(tf.ones_like(d_real), d_real) + bce(tf.zeros_like(d_fake), d_fake)
    d_grads = tape.gradient(d_loss, discriminator.trainable_weights)
    d_opt.apply_gradients(zip(d_grads, discriminator.trainable_weights))

    # 2. Train generator — wants D to label fakes as real
    z = tf.random.normal((batch, LATENT_DIM))
    with tf.GradientTape() as tape:
        fake_images = generator(z, training=True)
        d_fake = discriminator(fake_images, training=True)
        g_loss = bce(tf.ones_like(d_fake), d_fake)
    g_grads = tape.gradient(g_loss, generator.trainable_weights)
    g_opt.apply_gradients(zip(g_grads, generator.trainable_weights))

    return d_loss, g_loss

Three things to notice:

  • tanh output on G, with real images scaled to [-1, 1]. Don't mix sigmoid outputs with [-1, 1] inputs — a common silent bug.
  • Two separate optimisers, both with beta_1=0.5. The default Adam beta_1=0.9 is too high for GANs and contributes to mode collapse.
  • Conv2DTranspose for upsampling. The alternative — UpSampling2D + Conv2D — avoids the "checkerboard artifact" you sometimes see in early GAN papers.

That's the minimum viable GAN. It will train on MNIST, sometimes; it will mode-collapse, also sometimes. That instability is the headline problem.


3. Why GAN Training Is Hard

Three failure modes you'll hit:

  • Mode collapse. G discovers one or a few outputs that fool D reliably and stops producing variety. You generate 1,000 samples and they're all the same digit. Symptom: D loss low, G loss low, samples homogeneous.
  • Vanishing gradients on G. D gets too good. It assigns 0.0001 to every fake, the BCE gradient becomes tiny, G has nothing to learn from. Symptom: D loss near zero, G loss flat or rising.
  • Hyperparameter sensitivity. Change the LR by 2×, the result is qualitatively different. Change the batch size, the architecture seems "broken". Reproducing a paper's results is genuinely hard, which is why GAN papers got progressively more obsessed with seed-and-config disclosure.

The literature's response was a sequence of fixes — each one a paper, each one cited 10,000+ times.


4. The GAN Improvement Timeline

YearVariantKey idea
2014GANThe original. Often unstable.
2015DCGANDeep convolutional GAN. Convolutional architectures, BN, LeakyReLU, no fully-connected layers. Made GANs trainable for serious image tasks.
2017WGANReplace the BCE objective with the Wasserstein distance. Drops the sigmoid on D, uses weight clipping. Much more stable; loss becomes a meaningful quality signal.
2017WGAN-GPReplaces weight clipping with a gradient penalty. The "just works" GAN training recipe of the late 2010s.
2017Progressive GANGrow the network during training — start at 4×4 and add layers up to 1024×1024. NVIDIA used this for the famous "this person does not exist" demos.
2018StyleGANNVIDIA. Style injection at each resolution. Best-in-class face generation for years.
2019BigGANDeepMind. Scale matters; very large batches; class-conditional. State-of-the-art on ImageNet for a long stretch.
2019StyleGAN2/3Iterative refinements to StyleGAN. The faces got harder to spot.

The pattern is the same as everywhere else in deep learning: better architectures, better losses, more compute, better results — until something fundamentally different came along.


5. Conditional GANs and Image-to-Image

A conditional GAN (cGAN) takes a label alongside the random latent and produces an output conditioned on it. "Generate a 7" instead of "generate some digit". G and D both see the label.

pix2pix generalises this: input one image, output another. Edge map → photograph, day photo → night photo, sketch → finished art. Trained with paired data — you need both the input image and the target output.

CycleGAN removes the paired-data requirement. Two GANs and a cycle consistency loss — translating A → B → A should recover the original. This unlocked unpaired domain translation: horses to zebras, summer to winter, oil paintings to photographs. Still used in research and specific production tasks (medical image augmentation, style transfer pipelines) even after diffusion took over general image generation.


6. VAEs — The Principled Alternative

A Variational Autoencoder is the principled, likelihood-based generative model that GANs were largely replacing. The architecture is an autoencoder with a probabilistic twist:

  • Encoder maps an input x to a distribution over the latent z — outputs mean and log-variance.
  • Sample z from that distribution.
  • Decoder maps z back to a reconstruction of x.
  • Loss = reconstruction error + KL divergence pulling the encoder distribution toward a standard normal.

The KL term forces the latent space to be smooth and well-behaved — sample any z ~ N(0, I), decode it, and you get a plausible image. That's something a vanilla autoencoder doesn't give you.

Trade-offs vs GANs:

VAEGAN
Training stabilityStable. Single loss, single optimiser.Notoriously unstable.
Sample qualityBlurrier — the Gaussian likelihood encourages mean-seeking.Sharper, but variety can collapse.
Latent spaceSmooth, interpolatable, semantically organised.Less structured by default.
LikelihoodExplicit lower bound (the ELBO).None — you can't ask "how likely is this image?".

VAEs remain useful when you care about the latent space — anomaly detection, smooth interpolation, controllable generation — but they were never going to beat GANs on pixel quality. The interesting twist: modern latent diffusion models (Stable Diffusion) use a VAE encoder to compress images to a small latent, then run diffusion in that latent space. The VAE is a load-bearing component of state-of-the-art image generation in 2026.


7. Diffusion — What Replaced GANs

The intuition is almost embarrassingly simple. Take an image; add Gaussian noise; repeat until it's pure noise. Now learn a network that removes one step of noise. To generate: sample pure noise, apply the denoiser repeatedly, end up with a clean image.

Formally there are two processes:

  • Forward process (no learning) — add a tiny bit of noise at each of T steps until the image is indistinguishable from N(0, I). T is typically 1000 during training.
  • Reverse process (the learned model) — predict the noise that was added at each step, subtract a fraction of it, iterate.

The network is usually a U-Net (or, more recently, a transformer — DiT) that takes (noisy image, current timestep) and outputs the predicted noise. Training is regression with MSE — far easier than the min-max adversarial dance.

Why it beat GANs:

  • Training stability — single loss, single optimiser. No mode collapse.
  • Mode coverage — diffusion naturally captures the full diversity of the training distribution.
  • Steerability — classifier-free guidance lets you trade off prompt-fidelity vs diversity by changing one number at inference.
  • Compositionality — text conditioning works astoundingly well via cross-attention over text embeddings.

The cost is inference speed. A GAN samples in one forward pass; diffusion needs 20–50 iterative denoising steps even with the best samplers. The community fixed most of this with techniques like DPM-Solver and distillation — modern samplers run 4-step diffusion at near-original quality.

By 2026 the major image generators — DALL-E 3, Stable Diffusion 3, Midjourney, Imagen — are all diffusion-based. So are the leading video models (Sora, Veo, Runway) and a growing share of audio synthesis (AudioLDM, Stable Audio).


8. Stable Diffusion in Five Lines

The diffusers library wraps Stable Diffusion behind a Hugging Face-style pipeline:

python
from diffusers import StableDiffusionPipeline
import torch

pipe = StableDiffusionPipeline.from_pretrained(
    "stabilityai/stable-diffusion-2-1",
    torch_dtype=torch.float16,
).to("cuda")

image = pipe(
    "a watercolour painting of a snowy mountain at dawn, soft pastel colours",
    num_inference_steps=25,
    guidance_scale=7.5,
).images[0]

image.save("mountain.png")

That's the full code. The first call downloads ~5 GB of weights; subsequent calls are seconds on a recent GPU. CPU works but takes minutes per image.

The two knobs that matter:

  • num_inference_steps — fewer steps = faster but lower quality. 20–30 is the standard range for SD2/SD3. Modern distilled models (SDXL-Turbo, SD3.5-Turbo) work at 1–4 steps.
  • guidance_scale (CFG) — how strongly to follow the prompt. 7.5 is the canonical default; higher means more prompt-faithful but lower diversity, lower means more creative but more wandering.

For image-to-image, inpainting, ControlNet (conditioning on edges, depth, poses), or LoRA fine-tuning, the diffusers library has a pipeline for each — same from_pretrained pattern, different class.


9. The Honest 2026 Verdict

TaskWhat's state-of-the-art
Unconditional image generationDiffusion (latent diffusion). GANs uncompetitive at high resolution.
Text-to-imageDiffusion. GANs not in the conversation.
Image-to-image translationMixed. ControlNet (diffusion) is most common; CycleGAN still used for unpaired pairs.
Super-resolutionDiffusion variants (StableSR) lead; ESRGAN (a GAN) still strong for fast inference.
Face generation (unconditional)StyleGAN3 is competitive; diffusion wins on diversity.
Real-time / on-device generationDistilled GANs and few-step diffusion. Raw multi-step diffusion is too slow.
Audio / videoDiffusion, often with transformer denoisers.

GANs are not dead — they remain the right tool for low-latency inference, image-to-image with unpaired data, and a number of specialised research niches. But for the average "I want to generate good-looking images" problem in 2026, the answer is "use a pretrained diffusion model from diffusers".


10. Ethics — Generative Power Cuts Both Ways

The same model that produces a watercolour mountain produces a convincing deepfake. The same audio model that dubs your video into ten languages clones a senator's voice. The same image model that fills in your concept art trains on web-scraped art without consent.

The standard mitigations are partial:

  • Watermarking — invisible signals embedded in outputs that detectors can find. Easily stripped by re-encoding. Better than nothing, not a solution.
  • Content provenance (C2PA) — cryptographic signatures attesting how a piece of media was produced. Useful when verifying providence, useless if the platform doesn't display the signature.
  • Model-level safety — refuse to generate certain content. Always circumventable with fine-tuning of the open weights.
  • Opt-out registries — for training data. Voluntary. Late.

This deserves its own treatment. Forward link: ethics.


Common Mistakes

1. Training a GAN without seeding.

GAN training is so sensitive to initialisation that two runs with different seeds can produce visibly different results. Always tf.random.set_seed(42) (or the framework equivalent) and log the seed in your config. Otherwise you can never reproduce the good run.

2. Same learning rate for G and D.

D usually learns faster than G — the discriminative task is easier than the generative one. If they share an LR, D pulls ahead and G's gradients vanish. Use different LRs (or different update frequencies — train G twice for every D step). beta_1=0.5 instead of Adam's default 0.9 also helps.

3. Over-training the discriminator.

If D reaches near-perfect accuracy, BCE gradients vanish and G stops improving. The classic recipe is n_critic updates of D per update of G (n=1 by default, n=5 for WGAN). Watch the loss curves: if D loss is near zero while G loss climbs, D is winning too hard. Either reduce D's LR, add noise to D's inputs, or switch to WGAN-GP whose objective is better-behaved.

4. Mismatched output activation and pixel range.

tanh output → train on images in [-1, 1]. sigmoid output → train on images in [0, 1]. Mismatch the two and you'll see gradient pathology that looks like a mode-collapse bug but is actually a normalisation bug.

5. Trusting a single quality number.

GAN papers report FID, IS, and KID. None of them fully captures sample quality. Always eyeball a grid of samples. Cherry-picking a few good ones for the paper figure has, historically, hidden many a mediocre model.


🎯 Your Turn — Generate an Image with Stable Diffusion

Use the diffusers library to generate an image with Stable Diffusion in under ten lines. Save it to disk. Requirements:

  • Use StableDiffusionPipeline.from_pretrained("stabilityai/stable-diffusion-2-1").
  • Use torch.float16 and .to("cuda") if a GPU is available; otherwise CPU.
  • A prompt of your choice (something visually concrete works best).
  • num_inference_steps=25, guidance_scale=7.5.
  • Save the result to generated.png.

Skeleton:

python
from diffusers import StableDiffusionPipeline
import torch

# TODO 1: choose device — "cuda" if torch.cuda.is_available() else "cpu"
device = ...

# TODO 2: load the pipeline; use float16 on GPU, float32 on CPU; .to(device)
pipe = ...

# TODO 3: write a prompt
prompt = ...

# TODO 4: call the pipe with num_inference_steps=25, guidance_scale=7.5
result = ...

# TODO 5: save result.images[0] as "generated.png"
Hint 1 — Choosing dtype by device torch_dtype = torch.float16 if device == "cuda" else torch.float32. Pass that into from_pretrained. float16 on CPU is allowed but very slow; float32 on GPU works but uses 2× the memory.
Hint 2 — Saving the output The pipeline call returns an object with .images — a list of PIL Image objects. Save with result.images[0].save("generated.png"). No extra imports needed.
Show full solution
python
from diffusers import StableDiffusionPipeline
import torch

device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.float16 if device == "cuda" else torch.float32

pipe = StableDiffusionPipeline.from_pretrained(
    "stabilityai/stable-diffusion-2-1",
    torch_dtype=dtype,
).to(device)

prompt = (
    "a watercolour painting of a snowy mountain at dawn, "
    "soft pastel colours, gentle morning light"
)

result = pipe(prompt, num_inference_steps=25, guidance_scale=7.5)
result.images[0].save("generated.png")
print("Saved generated.png")

On an RTX-class GPU this runs in roughly 5–10 seconds. On CPU, expect 1–3 minutes. The first run downloads about 5 GB of weights to ~/.cache/huggingface/; subsequent runs are fast.

For more control, swap the model — stabilityai/sdxl-turbo runs at 1–4 steps for near-real-time generation; stable-diffusion-3-medium is the highest-quality open SD release; runwayml/stable-diffusion-inpainting is the in-painting variant. Same from_pretrained interface, very different capabilities.


What You Learned

  • A GAN trains a Generator and a Discriminator adversarially — G tries to produce realistic fakes, D tries to spot them.
  • GAN training is unstable: mode collapse, vanishing gradients, hyperparameter sensitivity. The literature is a decade of fixes (DCGAN → WGAN → WGAN-GP → StyleGAN).
  • Conditional GANs and pix2pix/CycleGAN extend the idea to controllable and image-to-image generation.
  • VAEs are the principled, likelihood-based alternative — stable to train, smooth latent space, but blurrier samples.
  • Diffusion models train a noise predictor and sample by iterative denoising. They beat GANs on stability, mode coverage, and prompt-following — at the cost of multi-step inference.
  • Stable Diffusion in five lines via the diffusers library. num_inference_steps and guidance_scale are the main knobs.
  • In 2026, diffusion has won for image, video, and most audio generation. GANs remain useful for real-time inference, unpaired image-to-image, and specific super-resolution tasks.
  • Generative power has serious ethical consequences — deepfakes, copyright, consent. Watermarking and provenance are partial mitigations, not solutions. Continued in ethics.

Next: Model Deployment — taking a trained model from a notebook to a production API that handles real traffic, real failures, and real cost.