PSA for instance admins in case of slow federation

nutomic@lemmy.ml · edit-2 1 year ago

PSA for instance admins in case of slow federation

ProfessionalHandJob@lemmy.beyondcombustion.net · edit-2 1 year ago

my server is just me currently… but it’s got 100GB of RAM and 30CPUs so i kicked my value up to 200k but this shit (the fediverse) is still slow as hell/doesn’t sync with most other servers because their specs are so low. People need to stop running them on $4 VPS shit boxes.

ono@lemmy.ca · edit-2 1 year ago

I have a bunch of lemmy.ml communities still stuck in “Subscribe Pending” state even after waiting for days. Canceling and re-subscribing does not help.

nutomic@lemmy.ml · 1 year ago

cc @smorks@lemmy.ca

snowe@lemmy.ml · 1 year ago

I am also seeing this.

RoundSparrow@lemmy.ml · edit-2 1 year ago

Going from 512 to 160,000 is a massive parameter change.

Network replication like this presents a ton of issues with servers going up and down, database insert locking the tables, desire for backfill and integrity checks, etc.

Today things are going poorly, this posting has an example: https://lemmy.ml/post/1239920 – comments are not showing up on other instances after hours.

From a denial-of-service perspective, intentional or accidental, I think we need to start discussing the protocol for federation. When servers connect to each others, how frequently, how they behave when getting errors from a remote host.

Is lemmy_server doing all the federation replication in-code? I would propose moving it to a independent service - perhaps a shell application - and start looking at replication (and associated queues) as a core thing to manage, optimize, and operate. It isn’t working seamlessly and having hundreds of servers creates a huge amount of complexity.