#mastodonnz

5 posts · Last used 18d

RE: https://mastodon.au/@xrobau/117300844625946496 Edit: We have figured it out, and are manually importing to a new temp table. --- Any #PostgreSQL exports around to help me out here? After a storage server crash, I've ended up with (at least one) table having corrupt time data in it. There is, apparently, no way to pg_dump and ignore errors, and \COPY also doesn't appear to continue-on-error. What I want to do is just get a list of the IDs of the rows with the corrupt time entries so I can either fix them, or just delete them. This is the 'status_stats' table from #MastodonNZ, and has about 70m rows. My current plan is that I've copied the schema of status_stats to status_stats_new, changing the time rows to just 'text' and I'm doing 'insert into status_stats_new select * from status_stats ON CONFLICT DO NOTHING;' then a left join to find missing ones. IS THERE A BETTER WAY? Halp!
Quoting
Sorry #MastodonNZ but there is a good chance I'm going to have to do a complete database dump and restore, as there's still some bits of corruption left over from the crash the other day. This is probably an hour or so outage. I'm actually not even totally sure that #PostgreSQL will even restart properly when I shut it down. More information will be in this thread (so you can look at it on #MastodonAU directly via the web interface.)
Open quoted post
6
5
10
0
Sorry #MastodonNZ but there is a good chance I'm going to have to do a complete database dump and restore, as there's still some bits of corruption left over from the crash the other day. This is probably an hour or so outage. I'm actually not even totally sure that #PostgreSQL will even restart properly when I shut it down. More information will be in this thread (so you can look at it on #MastodonAU directly via the web interface.)
22
4
4
2
Replying to
So far it looks like ONLY the 'status_stats' table is corrupt. This is not a critical table, and fingers crossed I can just skip the bad bits and re-import it. #MastodonNZ
8
0
0
0
Well I gotta say, moving #MastodonAU 's database to a pure SSD storage backend has made a massive difference. Whooohoo, she's much faster! The storage servers are a pair of #Dell R640s, with 2x Gold 6150 (2.7GHz 18 Cores each) and 256GB of RAM. Each one has 8 x 1TB SSDs, in a #ZFS raidz2 pool, and they automatically replicate between themselves. I am doing my best to avoid any more visits from the #fuckup-fairy Now to finish #MastodonNZ
29
14
5
0
Sorry about that long-ish outage, but we were YET AGAIN getting smashed by AI bots. I have given up fighting them, and I'm just scaling up until #mastodonAU can handle the load. This time the #PostgreSQL database was running at a load of 100+ because it was disk IO bound. I did *try* to live migrate it, but it was just insanely slow so I made the call to shut it down (hence the 'Something went wrong' errors) and move it offline. It's moved. Now it's running on a pure SSD-only backing store, which hopefully will handle everything the #AI scrapers throw at us. I'll be doing the same thing with #MastodonNZ too, shortly (but hopefully without the outage, as that's a much more lightweight instance) I also really need to fix the status page that comes up when mau is down.
23
4
11
0
You've seen all posts