Post #673817
2023-10-27 16:56 UTC
Many models produce outputs that are hard to verify for an end-user. Our new #emnlp2023 paper shows that providing a quality estimation model, can make a user better at deciding when to rely on the model.
https://arxiv.org/pdf/2310.16924v1.pdf
Replies (1)
-
@Niloufar@hci.social 2023-10-27 16:59
We ran a controlled experiment in which 65 physicians were shown sentences from a real emergency discharge instruction and their Google translations. The group that saw a QE model calibrated to medical communication was almost two times better at deciding when to rely on the translation than a baseline (measure is confidence weighted accuracy).