Elektrine lite

← Feed

@apftwb@lemmy.world

Post #2461113

2026-02-06 01:28 UTC

Are you having as much trouble with OCR as the article author? I would have thought OCR was a solved problem in 2026 even with poor font selection.

Replies (3)

  • @kescusay@lemmy.world 2026-02-06 04:06

    I'm not having trouble with it as such, it's just a slow and painstaking process. The source is crappy enough that an enormous number of characters need to be checked manually, and it's ridiculously time-consuming.

    Open ##2461115

  • @floofloof@lemmy.ca 2026-02-06 03:24

    I wonder if they gave considered crowdsourcing this, having many people type in small chunks of the data by hand, doing their own character recognition? Get enough people in and enough overlap and the process would have some built-in error correction.

    Open ##2461116

  • @Taldan@lemmy.world 2026-02-06 13:26

    OCR is mostly good enough. Problem here is we have 76 pages that we need to be read perfectly, with a low fidelity input We also have very little in the way of error correction, since it's mostly not human readable

    Open ##2461117