Step 5 of 6
Classification
Four neural networks learn to sort the photos: AlexNet, VGG19, ResNet50 and EfficientNet-B7.
How it works
Four networks, one question
Each network looks at a photo and produces five scores, one per class, that add up to 100%. The highest score is its answer. Training nudges millions of internal weights, one batch of photos at a time, until those answers match the labels.
VGG19, ResNet50 and EfficientNet-B7 started from weights already learned on ImageNet, a collection of over a million everyday photos, and were then fine-tuned on our eye photos. This is transfer learning: the networks arrive knowing edges, curves and textures, so they only have to learn what makes an eye misaligned.

A photo goes in
- Esotropia
- Exotropia
- Hypertropia
- Hypotropia
- Normal
Five scores come out
The four models
From a baseline to the best performer
AlexNet
59.26%accuracy
The baseline, used to understand how a network extracts features from this data.
See the architecture and results →
VGG19
66.67%accuracy
A deep, uniform stack of 3×3 convolutions, pretrained on ImageNet and fine-tuned in stages.
See the architecture and results →
ResNet50
74.00%accuracy
A residual network whose skip connections let a very deep model train stably.
See the architecture and results →
EfficientNet-B7
84.00%accuracy
The largest EfficientNet, which scales depth, width and resolution together. The best performer.
Best performer · See the architecture and results →