Strabismus

Results · published in the paper

How the four models compared

All numbers come from Table I and Figures 2–5 of our AIMLA 2025 paper. They were measured on the 81 test photos, which the models never saw during training.

  • EfficientNet-B784.00%EfficientNet-B7Accuracy 84.00%Precision 85.00%Recall 83.00%F1 score 84.40%
  • ResNet5074.00%ResNet50Accuracy 74.00%Precision 79.00%Recall 73.00%F1 score 74.00%
  • VGG1966.67%VGG19Accuracy 66.67%Precision 67.45%Recall 66.67%F1 score 67.00%
  • AlexNet59.26%AlexNetAccuracy 59.26%Precision 58.88%Recall 59.26%F1 score 58.71%
Accuracy on the held-out test set, as published in the paper (Table I). Hover or focus a bar for all four metrics.
Every metric, every model
ModelAccuracyPrecisionRecallF1 score
AlexNet59.2658.8859.2658.71
VGG1966.6767.4566.6767.00
ResNet5074.0079.0073.0074.00
EfficientNet-B784.0085.0083.0084.40
Accuracy:
Out of all test photos, the share the model classified correctly.
Precision:
When the model names a class, how often it is right, averaged across classes.
Recall:
Of the photos that truly belong to a class, how many the model found, averaged across classes.
F1 score:
A single score that balances precision and recall; it is high only when both are.

Where each model went wrong

Confusion matrices

Each row is a true class and shows where its test photos ended up. A strong model puts its weight on the diagonal. The weaker models spread theirs: VGG19 and ResNet50, for example, send many of their mistakes to exotropia.

Confusion matrix for AlexNet. Rows are the true class, columns the predicted class; each cell is the share of that true class's test photos.
Predicted class →
True ↓EsoExoHyperHypoNormal
Eso0.730.070.070.000.13
Exo0.000.650.060.180.12
Hyper0.190.190.440.190.00
Hypo0.120.120.120.440.19
Normal0.060.060.060.120.71
01 — share of each true class; the diagonal is correct predictions
Confusion matrix for VGG19. Rows are the true class, columns the predicted class; each cell is the share of that true class's test photos.
Predicted class →
True ↓EsoExoHyperHypoNormal
Eso0.800.130.000.000.07
Exo0.060.470.180.240.06
Hyper0.060.250.620.060.00
Hypo0.060.120.060.690.06
Normal0.000.180.060.000.76
01 — share of each true class; the diagonal is correct predictions
Confusion matrix for ResNet50. Rows are the true class, columns the predicted class; each cell is the share of that true class's test photos.
Predicted class →
True ↓EsoExoHyperHypoNormal
Eso0.930.000.000.000.07
Exo0.240.590.120.000.06
Hyper0.000.060.880.060.00
Hypo0.000.250.120.620.00
Normal0.000.180.060.060.71
01 — share of each true class; the diagonal is correct predictions

EfficientNet-B7

About this model →
Confusion matrix for EfficientNet-B7. Rows are the true class, columns the predicted class; each cell is the share of that true class's test photos.
Predicted class →
True ↓EsoExoHyperHypoNormal
Eso0.800.070.000.070.07
Exo0.060.760.000.060.12
Hyper0.060.000.940.000.00
Hypo0.000.060.060.750.12
Normal0.000.000.000.060.94
01 — share of each true class; the diagonal is correct predictions
Next →See it classify test photos