Elektrine lite

← Feed

@ingorohlfing@mastodon.social

Post #2547551

2026-05-04 17:47 UTC

@druedin@sciences.social @philipncohen@mastodon.social Interesting point. It is probably correct because, as the post notes, the model likely has been trained on similar research. Would then be interesting to let an LLM produce results for a more or less completely new phenomenon that is not part of the training data. The challenge then may be that no one has a good benchmark for evaluating the LLMs results, except for plausibility, which maybe good enough as a starting point.

Replies (1)

  • @druedin@sciences.social 2026-05-04 18:07

    @ingorohlfing@mastodon.social @philipncohen@mastodon.social Maybe we should do this when pre-registering research? That'd be a useful benchmark, I think.

    Open ##2547552