Post #2790144
2023-04-10 08:53 UTC
@GordanKnott@mastodon.social @ngaylinn@tech.lgbt @DrewKadel@social.coop
Yet another flawed benchmark in which the LLM very likely memorised the answers
https://wandering.shop/@janellecshane/110104164829618120
Without any knowledge of how much the training dataset was contaminated by the medical exam questions/answers (and OpenAI's own whitepaper admits there is contamination)
https://youtu.be/PEjl7-7lZLA?t=4m0s
we cannot really know how it would perform in the real world if say a novel virus were to start spreading
Replies (0)
No replies.