Step 3 of 6
Splitting
The photos are divided 70/15/15 into training, validation and test sets before anything else touches them.
70 / 15 / 15
Three sets, chosen before anything else
The preprocessed photos were divided into training, validation and test sets in a 70:15:15 ratio. The split was made class by class, so every set keeps the same mix of the five classes.
360
training photos
76
validation photos
81
test photos
Three jobs
Why three sets and not two
A model can look good on photos it has studied and still fail on new ones. Keeping photos aside is how you measure what it actually learned.
Validation photos guide training while it runs. Test photos are never used for any decision, so the final score is a clean measure of how the model handles photos it has never seen.
- 360TrainingThe only photos the models learn from.
- 76ValidationChecked after every round of training, to decide when to slow down or stop.
- 81TestLocked away until the very end. Every result on this site comes from these photos.
The order matters
Why we split before augmenting
The next step makes ten versions of every training photo. If we had made the copies first and split afterwards, a photo's flipped version could land in training while its brighter twin landed in the test set.
The model would then be tested on photos it had effectively already seen, and its score would look better than it really is. This is called data leakage. Splitting first keeps the test set honest: none of its photos, or any version of them, is ever used for training.
Split first, then copy
What we did
Training
Test
Photo B stays unseen. The test is fair.
Copy first, then split
What we avoided
Training
Test
Leak: the test holds a near-copy of a training photo, so the score comes out too high.