Monitoring for ipa backups #12992

Closed
opened 2025-12-18 09:31:41 +00:00 by gwmngilfen · 20 comments
Member

In #12969 we found (and fixed) a failure in the ipa backups. We should add a monitor for this to Zabbix.

In #12969 we found (and fixed) a failure in the ipa backups. We should add a monitor for this to Zabbix.
Member

Metadata Update from @james:

  • Issue priority set to: Waiting on Assignee (was: Needs Review)
  • Issue tagged with: low-trouble, medium-gain
**Metadata Update from @james**: - Issue priority set to: Waiting on Assignee (was: Needs Review) - Issue tagged with: low-trouble, medium-gain
Author
Member

Metadata Update from @gwmngilfen:

  • Issue tagged with: sprint-0
**Metadata Update from @gwmngilfen**: - Issue tagged with: sprint-0
Owner

I'll note that ipa01 seems fixed, but 02/03 aren't. :(

I'll note that ipa01 seems fixed, but 02/03 aren't. :(
Owner

I didn't noticed 02/03 are still having issues.

I didn't noticed 02/03 are still having issues.
Owner

Those failing cron jobs are complaining about space, so it's not an issue with cronjob itself, we just need to clear some space.

Those failing cron jobs are complaining about space, so it's not an issue with cronjob itself, we just need to clear some space.
Owner

So we have 13 G available on ipa01, but the whole db has 16 G. So it can't back it up. I assume it got bloated during the spam :/

For ipa02 17 G vs 20 G free, so it should have enough space to create backup. Same for ipa03. It's strange that all of them are failing then.

I remember @gwmngilfen saying something about out of inode error on ipa machines, so that could be it.

So we have 13 G available on ipa01, but the whole db has 16 G. So it can't back it up. I assume it got bloated during the spam :/ For ipa02 17 G vs 20 G free, so it should have enough space to create backup. Same for ipa03. It's strange that all of them are failing then. I remember @gwmngilfen saying something about out of inode error on ipa machines, so that could be it.
Owner

It really needs more space, I tried to do backup manually and it consumed all available space during that.

It really needs more space, I tried to do backup manually and it consumed all available space during that.
Owner

Hm, the backups are being kept in /var/ipa/backups and taking up space. I thought those are moved to backup01 and then deleted.

And I see only ipa01 on backup01:/fedora_backups/, but all of them are already there.

Hm, the backups are being kept in /var/ipa/backups and taking up space. I thought those are moved to backup01 and then deleted. And I see only ipa01 on backup01:/fedora_backups/, but all of them are already there.
Owner

It's keeping the latest 7 backups in cron https://pagure.io/fedora-infra/ansible/blob/main/f/roles/ipa/server/files/data-only-backup.sh, so we either give it more space or lower the number of kept backups on the machine.

It's keeping the latest 7 backups in cron https://pagure.io/fedora-infra/ansible/blob/main/f/roles/ipa/server/files/data-only-backup.sh, so we either give it more space or lower the number of kept backups on the machine.
Owner

Yeah, we keep the last 7 and we back that up to backup01 (on ipa01 only, because all of the cluster servers should be the same due to replication).

I guess we could keep less there, but it seemed handy to be able to easily restore.

Note that stage users were not being deleted correctly, so perhaps after thats fixed things will work better...

Yeah, we keep the last 7 and we back that up to backup01 (on ipa01 only, because all of the cluster servers should be the same due to replication). I guess we could keep less there, but it seemed handy to be able to easily restore. Note that stage users were not being deleted correctly, so perhaps after thats fixed things will work better...
Owner

They are deleted now, but the db size is still the same. I tried to cleanup the tombstones, but it didn't changed the db size. From what I found we probably have to do re-indexing as well, but that needs to be done on stopped db.

They are deleted now, but the db size is still the same. I tried to cleanup the tombstones, but it didn't changed the db size. From what I found we probably have to do re-indexing as well, but that needs to be done on stopped db.
Owner

TL;DR;: The backups are working again and there should be enough space to keep 7 of them.

With the help from LDAP folks I was able to reduce the size of LDAP db to 8G from 16G.

The problem why the size of the DB was so big was the replication_changelog. We still had some replication entries for ipaXX.aid2, which it couldn't reach and for some reason ipa01.rdu3 was there twice. This caused the replication to fail and the changelog to grow incrementally till the migration, the users created by spammers were just the last drop.

So I cleared the RUVs (Replication Update Vectors) and that helped me to finally clean the db manually.

TL;DR;: The backups are working again and there should be enough space to keep 7 of them. With the help from LDAP folks I was able to reduce the size of LDAP db to 8G from 16G. The problem why the size of the DB was so big was the `replication_changelog`. We still had some replication entries for `ipaXX.aid2`, which it couldn't reach and for some reason `ipa01.rdu3` was there twice. This caused the replication to fail and the changelog to grow incrementally till the migration, the users created by spammers were just the last drop. So I cleared the RUVs (Replication Update Vectors) and that helped me to finally clean the db manually.
Owner

Cool. Great detective work!

Shall we close this?

Cool. Great detective work! Shall we close this?
Owner

This ticket was originally opened for adding these checks to zabbix monitoring, backups were just something that we resolved during that. I don't think the checks are in place, so let's keep this open.

This ticket was originally opened for adding these checks to zabbix monitoring, backups were just something that we resolved during that. I don't think the checks are in place, so let's keep this open.
Owner

Oh, indeed. ok.

Oh, indeed. ok.
Author
Member

So, I have a prototype for this on STG, see this page

The setup returns the timestamp of the newest thing in /var/lib/ipa/backup, with a trigger to alert if that value is more than 26 hours old.

However, the backups filder is currently set 700 root:root meaning Zabbix can't read it. For testing I have set that to 750 root:zabbix - @kevin @zlopez is that ok with you?

I can do the Ansible parts and roll it out if you're good with that change.

So, I have a prototype for this on STG, see [this page](https://zabbix.stg.fedoraproject.org/zabbix.php?name=&evaltype=0&tags%5B0%5D%5Btag%5D=&tags%5B0%5D%5Boperator%5D=0&tags%5B0%5D%5Bvalue%5D=&show_tags=3&tag_name_format=0&tag_priority=&state=-1&filter_name=&filter_show_counter=0&filter_custom_time=0&sort=name&sortorder=ASC&show_details=0&action=latest.view&groupids%5B%5D=62&subfilter_tags%5Btype%5D%5B%5D=backup) The setup returns the timestamp of the newest thing in /var/lib/ipa/backup, with a trigger to alert if that value is more than 26 hours old. However, the backups filder is currently set `700 root:root` meaning Zabbix can't read it. For testing I have set that to `750 root:zabbix` - @kevin @zlopez is that ok with you? I can do the Ansible parts and roll it out if you're good with that change.
Owner

@gwmngilfen wrote in #12992 (comment):

However, the backups filder is currently set 700 root:root meaning Zabbix can't read it. For testing I have set that to 750 root:zabbix - @kevin @zlopez is that ok with you?

I don't think that would be a problem.

@gwmngilfen wrote in https://forge.fedoraproject.org/infra/tickets/issues/12992#issuecomment-342935: > However, the backups filder is currently set `700 root:root` meaning Zabbix can't read it. For testing I have set that to `750 root:zabbix` - @kevin @zlopez is that ok with you? I don't think that would be a problem.
Owner

I think thats fine, as long as the backup tool doesn't complain.

I think thats fine, as long as the backup tool doesn't complain.
Author
Member

Given this morning's backup seems to have run, I'd say we're golden. I'll roll it out.

Host Item Value
ipa01.stg.rdu3.fedoraproject.org IPA Backup Age 2026-01-23 04:03:06 AM
ipa02.stg.rdu3.fedoraproject.org IPA Backup Age 2026-01-23 04:03:07 AM
ipa03.stg.rdu3.fedoraproject.org IPA Backup Age 2026-01-23 04:03:05 AM
Given this morning's backup seems to have run, I'd say we're golden. I'll roll it out. Host | Item | Value ---|---|--- ipa01.stg.rdu3.fedoraproject.org | IPA Backup Age | 2026-01-23 04:03:06 AM ipa02.stg.rdu3.fedoraproject.org | IPA Backup Age | 2026-01-23 04:03:07 AM ipa03.stg.rdu3.fedoraproject.org | IPA Backup Age | 2026-01-23 04:03:05 AM
Author
Member

Closed via 01ab355

Closed via [01ab355](https://pagure.io/fedora-infra/ansible/c/01ab35531a532cf441bd0ccf8359a08a573660b2?branch=main)
Sign in to join this conversation.
No milestone
No assignees
4 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
infra/tickets#12992
No description provided.