Elektrine lite

← Feed

@stuebinm@pleroma.stuebinm.eu

Post #1179805

2026-04-14 15:25 UTC

okay, so, I looked at the state of OCR again (for "i have to many crusty-looking pdfs and maybe i can make them nicer to read" reasons) … and now i am kinda mad about AI because, well, my idea was, if i have a bad blurry scan of a paper but reasonable OCR of it, surely i can just replace the image by text characters which the OCR found? And it turns out, this works surprisingly well! … for prose. As soon as there's even one equation, subscript, or anything else non-trivial (or, god forbid, a table, or a *rotated* table), tesseract kinda gives up. No! Worse, it does not give up, it produces nonsense, so now there's just, symbol salad that you can't even easily detect as likely error! yay! but! this looks, I thought, like the kind of problem where the AI hype might've actually produced reasonable advances as a side effect (the way RNNoise is also very neat). And wouldn't you know it, they did! Throwing unlimited funding (and some machine learning) *can* produce results, it turns out! Except, of course, they are AI-hype-brained results. I can throw crusty theoretical compsci papers at their demonstration website (I assume these were all scraped already, so I have little qualms about it — tho ig there's some risk that their were in the training set and someone had to manually do a nice version, so take the following with some salt) and it deals wonderfully with all sorts of complicated formatting stuff … but nobody in this space is interested in using any of that to make the world a place with more accessible, searchable pdfs. They're purely interested in "linearising" the pdf into a markdown document, so it can be thrown into a machine learning pipeline. You might say, so what? I just wanted to read papers, and while I like beautiful typesetting and layout, surely I can just read it as markdown instead if i'm interested in the content? But well, they also discard anything that does not fit into markdown's text model. Bullet lists, tables? Nicely converted. Diagrams? Images? Figures? Simply gone. And the model does not tell one where the text came from, just the text itself, so if you just look at the generated markdown you might not even know you've missed the most important part and just … i dunno, i'm frustrated. there might be something close in this conceptual space that would genuinely help make tools better. But instead we get things that willfully discard information and whose main metric is "dollars per million pages read", to be put in a paper that reads like a marketing press release :BlobhajSadReach:

Replies (0)

No replies.