Post #1430903
2023-10-27 16:59 UTC
We ran a controlled experiment in which 65 physicians were shown sentences from a real emergency discharge instruction and their Google translations. The group that saw a QE model calibrated to medical communication was almost two times better at deciding when to rely on the translation than a baseline (measure is confidence weighted accuracy).
Replies (0)
No replies.