Post #1311278
2025-07-21 15:58 UTC
I recently spent a bit of time at work seeing how well @webrecorder 's (excellent) Browsertrix and ArchiveWebPage work at archiving Facebook posts:
https://docs.google.com/document/d/1uYTlUZeqsKyf_GHcIF5YYFHn2E5-NjdTh4FLoziAaVI/
It looks like the automated behavior for scrolling and interacting with the page may need adjusting. But more worrisome is that even when content has been archived, and successfully returned to the browser during replay there is some FB JavaScript that is preventing the archived content from being added to the page?
Replies (1)
-
@edsu@social.coop 2025-07-21 16:08
Because Meta et al. are so actively engaged in preventing content from being extracted from the web, the tools that archivists use inevitably get cut off as well. It's a lot of work keeping the browsertrix behaviors up to date, and ensuring that fuzzy matching works during replay. Webrecorder deserve much credit for doing this work as open source tools, but they need support: Open Collective https://opencollective.com/webrecorder or: Subscribe to their service: https://webrecorder.net/browsertrix/#get-started