Filter out some Copr related Zabbix warnings #13450

Closed
opened 2026-07-02 13:17:31 +00:00 by frostyx · 4 comments

Description of request

@gwmngilfen configured a Matrix room for the Copr team where we receive Zabbix warnings. This is much appreciated. We have one problem which is that we receive too many notifications and the important ones get lost in the noise. So I went through all the unique messages that we got and categorized them as either relevant or irrelevant for our team.

We definitely want to get these:

❌ copr-ping check older than 50mins (3754929)
Severity: Average started at 2026.06.10 13:39:23 on copr-be.aws.fedoraproject.org

⚠ Copr: copr-be.aws.fedoraproject.org return code is not 200 (5 checks) (3755533)
Severity: Warning started at 2026.06.10 13:58:54 on copr-fe.aws.fedoraproject.org

⚠ Copr: copr.fedorainfracloud.org does not have correct text (3756438)
Severity: Warning started at 2026.06.10 14:18:52 on copr-be.aws.fedoraproject.org

⚠ Linux: FS [/var/lib/dist-git]: Space is low (used > 80%, total 8857.7GB) (3801957)
Severity: Warning started at 2026.06.12 08:41:49 on copr-dist-git.aws.fedoraproject.org

⚠ COPR CDN Check failed (3869307)
Severity: Warning started at 2026.06.15 09:58:24 on copr-fe.aws.fedoraproject.org

⚠ Apache: Service is down (4136727)
Severity: Warning started at 2026.06.28 07:01:56 on copr-fe.aws.fedoraproject.org

ℹ MD Raid array md2 is in recovery mode on vmhost-p09-copr03.rdu3.fedoraproject.org (3870258)
Severity: Information started at 2026.06.15 11:19:50 on vmhost-p09-copr03.rdu3.fedoraproject.org

🔥 MD Raid array md0 is degraded on vmhost-p09-copr03.rdu3.fedoraproject.org (3818561)
Severity: High started at 2026.06.13 02:13:23 on vmhost-p09-copr03.rdu3.fedoraproject.org

❌ Linux: High memory utilization (>90% for 5m) (3756263)
Severity: Average started at 2026.06.10 14:17:22 on copr-be.aws.fedoraproject.org

These are not helpful for us, and should probably go to #fedora-noc. I don't know how to fix either of these:

🔥 MGMT: Chassis unreachable by HTTP (3754791)
Severity: High started at 2026.06.10 13:35:59 on vmhost-x86-copr01.rdu3.fedoraproject.org

✅ Linux: Interface eth5: Link down (4d 20h 56m 54s)
(kevin) acknowledged, commented and closed problem at 2026.06.10 16:52:17
expected.

⚠ Linux: Number of installed packages has been changed (3931379)
Severity: Warning started at 2026.06.18 13:24:32 on vmhost-p09-copr03.rdu3.fedoraproject.org

ℹ Linux: /etc/passwd has been changed (3931380)
Severity: Information started at 2026.06.18 13:24:35 on vmhost-p09-copr03.rdu3.fedoraproject.org

ℹ Linux: Operating system description has changed (4087316)
Severity: Information started at 2026.06.25 23:01:54 on vmhost-x86-copr03.rdu3.fedoraproject.org

🔥 Linux: Active checks are not available (3930081)
Severity: High started at 2026.06.18 11:28:40 on vmhost-p09-copr03.rdu3.fedoraproject.org

❌ Linux: Zabbix agent is not available (or nodata for 30m) (3930308)
Severity: Average started at 2026.06.18 11:51:30 on vmhost-p09-copr03.rdu3.fedoraproject.org

🔥 Host DNS: Unavailable by ICMP ping (3930508)
Severity: High started at 2026.06.18 12:03:27 on vmhost-p09-copr03.rdu3.fedoraproject.org

The load and high CPU utilization ones, we don't know

⚠ Linux: Load average is too high (per CPU load over 6 for 5m) (3755230)
Severity: Warning started at 2026.06.10 13:46:34 on copr-be.aws.fedoraproject.org

⚠ Linux: High CPU utilization (over 90% for 5m) (3772263)
Severity: Warning started at 2026.06.11 04:44:29 on copr-dist-git.aws.fedoraproject.org

Is there some severity above warning? Like critical? If yes, we would like to only get that one. If not, please continue sending the warnings.

### Description of request @gwmngilfen configured a Matrix room for the Copr team where we receive Zabbix warnings. This is much appreciated. We have one problem which is that we receive too many notifications and the important ones get lost in the noise. So I went through all the unique messages that we got and categorized them as either relevant or irrelevant for our team. We definitely want to get these: ``` ❌ copr-ping check older than 50mins (3754929) Severity: Average started at 2026.06.10 13:39:23 on copr-be.aws.fedoraproject.org ⚠ Copr: copr-be.aws.fedoraproject.org return code is not 200 (5 checks) (3755533) Severity: Warning started at 2026.06.10 13:58:54 on copr-fe.aws.fedoraproject.org ⚠ Copr: copr.fedorainfracloud.org does not have correct text (3756438) Severity: Warning started at 2026.06.10 14:18:52 on copr-be.aws.fedoraproject.org ⚠ Linux: FS [/var/lib/dist-git]: Space is low (used > 80%, total 8857.7GB) (3801957) Severity: Warning started at 2026.06.12 08:41:49 on copr-dist-git.aws.fedoraproject.org ⚠ COPR CDN Check failed (3869307) Severity: Warning started at 2026.06.15 09:58:24 on copr-fe.aws.fedoraproject.org ⚠ Apache: Service is down (4136727) Severity: Warning started at 2026.06.28 07:01:56 on copr-fe.aws.fedoraproject.org ℹ MD Raid array md2 is in recovery mode on vmhost-p09-copr03.rdu3.fedoraproject.org (3870258) Severity: Information started at 2026.06.15 11:19:50 on vmhost-p09-copr03.rdu3.fedoraproject.org 🔥 MD Raid array md0 is degraded on vmhost-p09-copr03.rdu3.fedoraproject.org (3818561) Severity: High started at 2026.06.13 02:13:23 on vmhost-p09-copr03.rdu3.fedoraproject.org ❌ Linux: High memory utilization (>90% for 5m) (3756263) Severity: Average started at 2026.06.10 14:17:22 on copr-be.aws.fedoraproject.org ``` These are not helpful for us, and should probably go to #fedora-noc. I don't know how to fix either of these: ``` 🔥 MGMT: Chassis unreachable by HTTP (3754791) Severity: High started at 2026.06.10 13:35:59 on vmhost-x86-copr01.rdu3.fedoraproject.org ✅ Linux: Interface eth5: Link down (4d 20h 56m 54s) (kevin) acknowledged, commented and closed problem at 2026.06.10 16:52:17 expected. ⚠ Linux: Number of installed packages has been changed (3931379) Severity: Warning started at 2026.06.18 13:24:32 on vmhost-p09-copr03.rdu3.fedoraproject.org ℹ Linux: /etc/passwd has been changed (3931380) Severity: Information started at 2026.06.18 13:24:35 on vmhost-p09-copr03.rdu3.fedoraproject.org ℹ Linux: Operating system description has changed (4087316) Severity: Information started at 2026.06.25 23:01:54 on vmhost-x86-copr03.rdu3.fedoraproject.org 🔥 Linux: Active checks are not available (3930081) Severity: High started at 2026.06.18 11:28:40 on vmhost-p09-copr03.rdu3.fedoraproject.org ❌ Linux: Zabbix agent is not available (or nodata for 30m) (3930308) Severity: Average started at 2026.06.18 11:51:30 on vmhost-p09-copr03.rdu3.fedoraproject.org 🔥 Host DNS: Unavailable by ICMP ping (3930508) Severity: High started at 2026.06.18 12:03:27 on vmhost-p09-copr03.rdu3.fedoraproject.org ``` The load and high CPU utilization ones, we don't know ``` ⚠ Linux: Load average is too high (per CPU load over 6 for 5m) (3755230) Severity: Warning started at 2026.06.10 13:46:34 on copr-be.aws.fedoraproject.org ⚠ Linux: High CPU utilization (over 90% for 5m) (3772263) Severity: Warning started at 2026.06.11 04:44:29 on copr-dist-git.aws.fedoraproject.org ``` Is there some severity above warning? Like critical? If yes, we would like to only get that one. If not, please continue sending the warnings.
Member

Thanks for this. In #noc we use a filter of "above Warning" which still leaves us "Average", "High", and "Disaster" levels, while "Warning" and "Informational" stay in the Zabbix UI. So for COPR checks like the CDN, copr-ping, etc, we can simply set the status to "Average" or higher, and apply the same filter.

The issue is with the generic checks like MD raid, Apache, etc. These come from more generic templates which apply to many hosts, so simply raising the alert level isn't what we want to do. We can match on specific trigger names, but that might mean you miss out on other things you care about later.

My gut feeling is that we need some negative-tagging like "not-copr" or such that we can add to infra-wide templates, but I'm not super happy with that. I'll think on it.

Thanks for this. In #noc we use a filter of "above Warning" which still leaves us "Average", "High", and "Disaster" levels, while "Warning" and "Informational" stay in the Zabbix UI. So for COPR checks like the CDN, copr-ping, etc, we can simply set the status to "Average" or higher, and apply the same filter. The issue is with the generic checks like MD raid, Apache, etc. These come from more generic templates which apply to many hosts, so simply raising the alert level isn't what we want to do. We can match on *specific* trigger names, but that might mean you miss out on other things you care about later. My gut feeling is that we need some negative-tagging like "not-copr" or such that we can add to infra-wide templates, but I'm not super happy with that. I'll think on it.
Author

So for COPR checks like the CDN, copr-ping, etc, we can simply set the status to "Average" or higher, and apply the same filter.

Sounds good to me

> So for COPR checks like the CDN, copr-ping, etc, we can simply set the status to "Average" or higher, and apply the same filter. Sounds good to me
Member

So I don't think a simple severity-level check works, you have some items you want that are low-prio, and some you don't want that are high-prio.

Here's what I've set up on prod just now, which I think covers what you need:

image

Specifically this is either a check marked as copr-team, or one of the named checks above. I used a match on tags, so other problems with those services will be reported, rather than using explicit trigger names.

Let's see how that goes - I haven't made the change yet on staging, so we can perhaps compare.

So I don't think a simple severity-level check works, you have some items you want that are low-prio, and some you don't want that are high-prio. Here's what I've set up on prod just now, which I think covers what you need: ![image](/attachments/d4b5531c-bb7e-41b5-9315-769970e1980a) Specifically this is *either* a check marked as copr-team, *or* one of the named checks above. I used a match on tags, so other problems with those services will be reported, rather than using explicit trigger names. Let's see how that goes - I haven't made the change yet on staging, so we can perhaps compare.
Member

I've set up the above on both prod and stg, so I'll close this for now. We can always tweak it more.

I've set up the above on both prod and stg, so I'll close this for now. We can always tweak it more.
Sign in to join this conversation.
No milestone
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
infra/tickets#13450
No description provided.