Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

As usual with ML, I now wonder how “similar” the test set is to the training set, compared to the examples that are neither in the training set, nor the test set:

TODO: 200 million - 170,000 training - 100 test ~= 199.8 million proteins



They trained on 170k sequences/ structures/ proteins, each sequence has 10s to 100s or even 1000s amino acids. Structure is much more conserved than sequence. Out of the 100 targets, roughly 1/4th have no similarity to known structures, so there shouldn't be an overlap for those with the training set. They did very well on those targets.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: