GAN vs VAE

Two generative models on CIFAR-10, judged by one protocol that was locked before either could win.

Context
Deep Learning coursework, two-person, with Mohamad Aniq
My part
Led the GAN track, shared EDA and the evaluation protocol; made the final call
Stack
Python, TensorFlow, Keras, InceptionV3 (FID), Colab
Year
2026
Conditional GAN samples across CIFAR-10 classes

The problem

Which model family makes the stronger CIFAR-10 image generator: a GAN or a VAE? The honest answer depends less on the models than on the test. Comparisons like this usually drift: each side is tuned against its own favourite metric, the test set leaks into checkpoint choices, and one lucky seed becomes the result.

Generative models also have no single score. A model can be sharp and repetitive, diverse and unreadable, or look good on a loss curve and bad on screen. So the work was as much about building a fair judge as about building the generators.

The two tracks

  • GAN (my track)Three foundations screened under identical conditions, then controlled experiments on latent size, upsampling, capacity, learning rates and augmentation, then class conditioning. Final: a conditional DCGAN with a 100-d latent and a learned 50-d label embedding, 1,290,487 generator parameters.
  • VAE (Aniq’s track)Decoder, latent, KL, conditioning and objective experiments, frozen at V04: a conditional convolutional VAE with a 64-d latent and 394,307 decoder parameters.

The screen

Every foundation trained 100 epochs from the same seed, then got the same evaluation package: milestone FID at epochs 50, 70, 90 and 100, repeated FID on three fresh latent samples, a 100-image manual score, and feature-space diversity against real CIFAR-10.

  • Dense GAN

    Repeated dev FID 327.3Clear / 100 0Nonsense / 100 80

  • DCGAN

    Repeated dev FID 37.0Clear / 100 9Nonsense / 100 16

  • WGAN-GP

    Repeated dev FID 54.7Clear / 100 7Nonsense / 100 22

FoundationRepeated dev FIDClear / 100Nonsense / 100
Dense GAN327.3080
DCGAN37.0916
WGAN-GP54.7722

Development FID is measured against the 50,000 training images. The official test set stayed sealed until the final comparison.

Decisions

  • Screen three foundations before tuning one

    WGAN-GP is the textbook stable choice, so it would have been easy to start there. Running Dense GAN, DCGAN and WGAN-GP through the same 100-epoch screen showed DCGAN ahead on FID, manual scores and diversity, 17 FID points clear of WGAN-GP, with a better quality-to-compute balance.

  • Lock the adoption rule before seeing any tuning result

    Four runs of the unchanged baseline measured the noise: a spread of 0.43 FID between training seeds. From that I fixed the bar in advance: a change is adopted only if it beats the baseline by at least 2.0 FID, a margin set by that measured noise rather than by the results I hoped for.

  • Reject a win that doesn’t reproduce

    A 128-d latent beat the baseline by 2.57 FID under seed 42, cleared the bar, and looked better by eye. Under seed 123 it lost by 1.32. It was not adopted. The same discipline closed the upsampling family (resize-convolution was 94 FID worse) and kept a stronger generator out when its 1.38 gain sat inside the margin.

  • Never pick a checkpoint by one number

    The best epoch is not assumed to be the last or the lowest FID. Each run kept epochs 50, 70, 90 and 100, and the choice weighed FID with fixed-seed grids, uncurated grids, stability and signs of collapse. A reload-verified fallback generator was frozen early as insurance, so an overnight failure could never cost the deliverable.

  • Judge the finalists four ways

    After both candidates and the rules were frozen, each model generated 5,000 balanced images on three fixed seeds, scored once against the protected test set. A blinded visual review, a requested-class check and a nearest-neighbour memorisation audit sat beside FID, so the result could not rest on a single metric.

Results

  • 36.7

    GAN protected-test FID, ±0.12 over three seeds (lower is better)

  • 147.2

    VAE protected-test FID, ±0.61 on the same protocol

  • 48% vs 1%

    Images showing the requested class, GAN vs VAE, in a blinded review

  • Clear / Marginal / Nonsense (blind, 100 each)

    GAN 50 / 45 / 5VAE 0 / 71 / 29

  • Requested class correct / ambiguous / wrong

    GAN 48 / 30 / 22VAE 1 / 65 / 34

  • Unique nearest training neighbours / 5,000

    GAN 3,595VAE 1,006

  • Duplicate neighbour assignments

    GAN 28.10%VAE 79.88%

MeasureGANVAE
Clear / Marginal / Nonsense (blind, 100 each)50 / 45 / 50 / 71 / 29
Requested class correct / ambiguous / wrong48 / 30 / 221 / 65 / 34
Unique nearest training neighbours / 5,0003,5951,006
Duplicate neighbour assignments28.10%79.88%
Contact sheet of the final GAN outputs
From the final 1,000-image deliverable: 100 requested images per CIFAR-10 class, uncurated.

The conditional GAN was selected for the deliverable. Neither model sat unusually close to its training images; the VAE’s heavy neighbour reuse points to mode concentration, not copying. The visual review had one blinded scorer, so it supports FID rather than standing as ground truth.

Limits

  • Class conditioning stayed imperfect, especially for several animal classes.
  • One scorer for the visual review, so no inter-rater agreement.
  • Seed-level confirmation was limited, particularly for the VAE.
  • A nearest-neighbour audit cannot prove the absence of memorisation.