As usual with ML, I now wonder how “similar” the test set is to the training set, compared to the examples that are neither in the training set, nor the test set:
TODO: 200 million - 170,000 training - 100 test ~= 199.8 million proteins
They trained on 170k sequences/ structures/ proteins, each sequence has 10s to 100s or even 1000s amino acids. Structure is much more conserved than sequence. Out of the 100 targets, roughly 1/4th have no similarity to known structures, so there shouldn't be an overlap for those with the training set. They did very well on those targets.
TODO: 200 million - 170,000 training - 100 test ~= 199.8 million proteins