request for more s390x koji builder resources #13141

Open
opened 2026-02-12 21:02:38 +00:00 by decathorpe · 17 comments

Description of request

It looks like the current capacity of s390x koji builders is barely enough to keep up with the "base load", but if there is anything "larger" happening, like OpenJDK, Python, or LLVM builds, these saturate all available koji builders and other builds usually wait 30-45 min, sometimes 60 min or even longer, to get an s390x builder assigned.

(This is even worse for CI builds, which usually just time out after 2 hours (?) and never report back success even if they finish after having waited in line for 42 hours.)

Especially when doing multi-package updates where individual builds finish ~quickly (i.e. chain-build), the wait times multiply quickly: For example, assuming two packages that have to be built in order, and take ~5 minutes to build individually; without resource starvation, this takes ~15 minutes to finish (5 min build, 5 min repo regen, 5 min build) - but with resource starvation, this often takes upwards of an hour (or two) to finish (other architectures finish quickly, but the builds are stuck waiting for s390x builder resources for longer than the build itself takes).

I assume builds that have to be done "in order" (i.e. koji chain-builds) are not that common in the RHEL world so the issue might not be "known" (or felt as badly) there, but it is very common in Fedora to hit this, and it is a massive time-suck on days like today (like, things taking ~3 hours that should take 15 minutes).

### Description of request It looks like the current capacity of s390x koji builders is *barely* enough to keep up with the "base load", but if there is anything "larger" happening, like OpenJDK, Python, or LLVM builds, these saturate all available koji builders and other builds usually wait 30-45 min, sometimes 60 min or even longer, to get an s390x builder assigned. (This is even worse for CI builds, which usually just time out after 2 hours (?) and never report back success even if they finish after having waited in line for 42 hours.) Especially when doing multi-package updates where individual builds finish ~quickly (i.e. chain-build), the wait times multiply *quickly*: For example, assuming two packages that have to be built in order, and take ~5 minutes to build individually; without resource starvation, this takes ~15 minutes to finish (5 min build, 5 min repo regen, 5 min build) - but with resource starvation, this often takes upwards of an hour (or two) to finish (other architectures finish quickly, but the builds are stuck waiting for s390x builder resources for longer than the build itself takes). I assume builds that have to be done "in order" (i.e. koji chain-builds) are not that common in the RHEL world so the issue might not be "known" (or felt as badly) there, but it is *very* common in Fedora to hit this, and it is a massive time-suck on days like today (like, things taking ~3 hours that should take 15 minutes).
Owner

Some more data:

We have 2 lpars. One for staging and one for production.
The production one has 560GB of memory and 48 cpus

There's 20 builders on it.
We reserve 3 of them for composes (because they have a rw koji mount and we never want non compose jobs to use those)
We reserve 1 of them for a varnish cache that allows the rest of them to get cached packages much faster.

CI jobs are only allowed to use 4 of them

We could spread the resources out more (make more vms), but then everything will be slower, but at least smaller jobs will finish sooner.

We could ask for more resources, but how much?

We could move those 'big' builds into another channel and restrict them to only a subset of builders, leaving more for 'small' jobs at the cost of all those taking longer.

Some more data: We have 2 lpars. One for staging and one for production. The production one has 560GB of memory and 48 cpus There's 20 builders on it. We reserve 3 of them for composes (because they have a rw koji mount and we never want non compose jobs to use those) We reserve 1 of them for a varnish cache that allows the rest of them to get cached packages much faster. CI jobs are only allowed to use 4 of them We could spread the resources out more (make more vms), but then everything will be slower, but at least smaller jobs will finish sooner. We could ask for more resources, but how much? We could move those 'big' builds into another channel and restrict them to only a subset of builders, leaving more for 'small' jobs at the cost of all those taking longer.
Author

Speaking from personal experience: Builds being a bit slower (even 50% slower) would not be a problem, the big time-suck is when builds do not start to build at all for an hour because they're waiting for resources. So maybe allocating more but smaller VMs could help.

Speaking from personal experience: Builds being a bit slower (even 50% slower) would not be a problem, the big time-suck is when builds do not start to build *at all* for an hour because they're waiting for resources. So maybe allocating more but smaller VMs could help.

I'm one of the package maintainers of LLVM.

We could ask for more resources, but how much?

It would be great if newer builders had at least 200GiB of storage.
Some s390x builders have only ~130GiB of storage, which is not enough to build LLVM with all the features enabled.
"Luckily", LLVM upstream supports only a subset of features on s390x. If the missing features are enabled upstream, we wouldn't be able to enable them on Fedora s390x due to the storage limit.

I don't have strong opinions on memory and CPU, but if you need help, I'd gladly come up with numbers.

I'm one of the package maintainers of LLVM. > We could ask for more resources, but how much? It would be great if newer builders had at least 200GiB of storage. Some s390x builders have only ~130GiB of storage, which is not enough to build LLVM with all the features enabled. "Luckily", LLVM upstream supports only a subset of features on s390x. If the missing features are enabled upstream, we wouldn't be able to enable them on Fedora s390x due to the storage limit. I don't have strong opinions on memory and CPU, but if you need help, I'd gladly come up with numbers.
Owner

Someone noted that things were delayed this morning.

This seems related to:

142736225 19   packit               OPEN     s390x      buildArch (llvm-21.1.8-26.eln155.src.rpm, s390x)
142736248 19   packit               OPEN     s390x      buildArch (llvm-21.1.8-26.fc45.src.rpm, s390x)
142736300 19   packit               OPEN     s390x      buildArch (llvm-21.1.8-26.fc45.src.rpm, s390x)
142736304 19   packit               OPEN     s390x      buildArch (llvm-21.1.8-26.eln155.src.rpm, s390x)
142740673 19   tuliom               OPEN     s390x      buildArch (llvm21-21.1.8-1.fc45.src.rpm, s390x)
142740918 19   packit               OPEN     s390x      buildArch (llvm-22.1.0-1.fc45.src.rpm, s390x)
142740952 19   packit               OPEN     s390x      buildArch (llvm-22.1.0-2.fc45.src.rpm, s390x)
142740982 19   packit               OPEN     s390x      buildArch (llvm-22.1.0-1.eln155.src.rpm, s390x)
142740990 19   packit               OPEN     s390x      buildArch (llvm-22.1.0-2.eln155.src.rpm, s390x)
142741032 19   packit               OPEN     s390x      buildArch (llvm-22.1.0-3.eln155.src.rpm, s390x)
142741038 19   packit               OPEN     s390x      buildArch (llvm-22.1.0-3.fc45.src.rpm, s390x)

I'm not sure if we can filter packit over to the restricted subset of builders, because it also does real builds for users.
Perhaps we can restrict it's scratch builds somehow?
Or just disable s390x in scratch builds for pr's.

Someone noted that things were delayed this morning. This seems related to: ``` 142736225 19 packit OPEN s390x buildArch (llvm-21.1.8-26.eln155.src.rpm, s390x) 142736248 19 packit OPEN s390x buildArch (llvm-21.1.8-26.fc45.src.rpm, s390x) 142736300 19 packit OPEN s390x buildArch (llvm-21.1.8-26.fc45.src.rpm, s390x) 142736304 19 packit OPEN s390x buildArch (llvm-21.1.8-26.eln155.src.rpm, s390x) 142740673 19 tuliom OPEN s390x buildArch (llvm21-21.1.8-1.fc45.src.rpm, s390x) 142740918 19 packit OPEN s390x buildArch (llvm-22.1.0-1.fc45.src.rpm, s390x) 142740952 19 packit OPEN s390x buildArch (llvm-22.1.0-2.fc45.src.rpm, s390x) 142740982 19 packit OPEN s390x buildArch (llvm-22.1.0-1.eln155.src.rpm, s390x) 142740990 19 packit OPEN s390x buildArch (llvm-22.1.0-2.eln155.src.rpm, s390x) 142741032 19 packit OPEN s390x buildArch (llvm-22.1.0-3.eln155.src.rpm, s390x) 142741038 19 packit OPEN s390x buildArch (llvm-22.1.0-3.fc45.src.rpm, s390x) ``` I'm not sure if we can filter packit over to the restricted subset of builders, because it also does real builds for users. Perhaps we can restrict it's scratch builds somehow? Or just disable s390x in scratch builds for pr's.

See https://github.com/packit/packit-service/issues/3023 on why half of these builds are running.
The other half is related to https://github.com/packit/packit-service/issues/3021 .

There are alternatives being discussed in order to improve the situation without reducing pre-commit tests.

See https://github.com/packit/packit-service/issues/3023 on why half of these builds are running. The other half is related to https://github.com/packit/packit-service/issues/3021 . There are alternatives being discussed in order to improve the situation without reducing pre-commit tests.
Author

There seems to be another pile-up happening today, involving COSMIC 1.0.9, mariadb, thunderbird, kernel, uv, ruff, etc. Some of my builds have been pending waiting for an s390x builder for 90+ minutes now (example: https://koji.fedoraproject.org/koji/taskinfo?taskID=144225079 ).

There seems to be another pile-up happening today, involving COSMIC 1.0.9, mariadb, thunderbird, kernel, uv, ruff, etc. Some of my builds have been pending waiting for an s390x builder for 90+ minutes now (example: https://koji.fedoraproject.org/koji/taskinfo?taskID=144225079 ).
Owner

Yeah, mostly it's all the cosmic jobs... they are taking a while and there's many of them. ;(

Looking at monitoring... it seems like it might be memory constrained. I wonder if a good first step is to ask for more memory and increase the guests memory accordingly.

Yeah, mostly it's all the cosmic jobs... they are taking a while and there's many of them. ;( Looking at monitoring... it seems like it might be memory constrained. I wonder if a good first step is to ask for more memory and increase the guests memory accordingly.
Author

This is happening again today. Builds that take 1-2 minutes to build take 30-60 minutes instead instead because they wait for 30+ minutes to get an s390x builder assigned to them.

Example tasks:

https://koji.fedoraproject.org/koji/taskinfo?taskID=144417655
https://koji.fedoraproject.org/koji/taskinfo?taskID=144417775
https://koji.fedoraproject.org/koji/taskinfo?taskID=144417831

I know that you can't do much about this but it's a massive and annoying time-and-joy-suck.

This is happening again today. Builds that take 1-2 minutes to build take 30-60 minutes instead instead because they wait for 30+ minutes to get an s390x builder assigned to them. Example tasks: https://koji.fedoraproject.org/koji/taskinfo?taskID=144417655 https://koji.fedoraproject.org/koji/taskinfo?taskID=144417775 https://koji.fedoraproject.org/koji/taskinfo?taskID=144417831 I know that you can't do much about this but it's a massive and annoying time-and-joy-suck.

Kevin asked me to start keeping track of builds slow to start on s390x. I've done exactly one WebKitGTK build since then, https://koji.fedoraproject.org/koji/taskinfo?taskID=145258748, and it took 10 hours to start.

(That said, these delays are greatly preferable to the unreliable builds we used to have before #12377.)

Kevin asked me to start keeping track of builds slow to start on s390x. I've done exactly one WebKitGTK build since then, https://koji.fedoraproject.org/koji/taskinfo?taskID=145258748, and it took 10 hours to start. (That said, these delays are greatly preferable to the unreliable builds we used to have before https://forge.fedoraproject.org/infra/tickets/issues/12377.)

We have a little more older history in: #12131.

We have a little more older history in: https://forge.fedoraproject.org/infra/tickets/issues/12131.

https://koji.fedoraproject.org/koji/taskinfo?taskID=145425348 - PR for a small package that builds on other arches in a few minutes has been queued for over an hour on s390x

https://src.fedoraproject.org/rpms/libucontext/pull-request/1

https://koji.fedoraproject.org/koji/taskinfo?taskID=145425348 - PR for a small package that builds on other arches in a few minutes has been queued for over an hour on s390x https://src.fedoraproject.org/rpms/libucontext/pull-request/1
Owner

Yeah, 9 python3.15 builds were taking 9 builders up... this was the same case of "tons of large/long builds all happening at the same time" again... we just need more numbers of builders so there are some available when a large build chud comes along. :(

Yeah, 9 python3.15 builds were taking 9 builders up... this was the same case of "tons of large/long builds all happening at the same time" again... we just need more numbers of builders so there are some available when a large build chud comes along. :(
Owner

FYI, this backlog hit again today... There were 93 builds in queue on s390x.

I moved one builder that was dedicated to kernel builds over to the general pool and that seemed to help.
I then also pulled a builder from compose channel (hopefully won't make composes slower) and added it also to the general pool.

I also have filed another plea for more resources on the existing mainframe.

FYI, this backlog hit again today... There were 93 builds in queue on s390x. I moved one builder that was dedicated to kernel builds over to the general pool and that seemed to help. I then also pulled a builder from compose channel (hopefully won't make composes slower) and added it also to the general pool. I also have filed another plea for more resources on the existing mainframe.
Author

This is happening again today, apparently due to the mini-Python-extension-mass-rebuild.

I have builds that would take ~2 minutes to complete, but they have been waiting for s390x builder to pick them up for 45 minutes (and counting).

This is happening again today, apparently due to the mini-Python-extension-mass-rebuild. I have builds that would take ~2 minutes to complete, but they have been waiting for s390x builder to pick them up for 45 minutes (and counting).
Owner

I don't actually think this is due to the python mass rebuild... it's the same ol 'a bunch of large builds are occupying the builders' problem. ;(

There's 3 gcc's, 2 llvms, 2 kernels, 2 nodejs, 2 firefoxes... so thats 11 builders occupied for a while. ;(

I don't actually think this is due to the python mass rebuild... it's the same ol 'a bunch of large builds are occupying the builders' problem. ;( There's 3 gcc's, 2 llvms, 2 kernels, 2 nodejs, 2 firefoxes... so thats 11 builders occupied for a while. ;(

I wonder if it would have helped if we had a hw requalification process similar to Debian [1].
Perhaps that could help to get warnings before having such a critical situation as we have now.

[1] https://release.debian.org/testing/arch_qualify.html

I wonder if it would have helped if we had a hw requalification process similar to Debian [1]. Perhaps that could help to get warnings before having such a critical situation as we have now. [1] https://release.debian.org/testing/arch_qualify.html

I've started a tool to analyze Koji build statistics -- this is what it looks like for yesterday (when the Python archful mass rebuild was happening)

Right now it uses the Koji APIs, once this proves useful we can extend it to use a database dump (hopefully a selective one that only includes the data it needs). It can fetch one day's worth of build data at a time and combine the raw data, so we can slowly build the data to analyze over time and re-analyze it later

Instances: fedora
Gated builds: 3502 (critical-path attribution); unattributed tasks: 0

arch queued med-wait p90-wait built med-time p90-time gated med-delay tot-delay
s390x 1298 2.6h 3.6h 1258 3.7m 50.4m 1120 2.0h 2336.9h
ppc64le 3975 1.3m 57.3m 3791 2.8m 9.4m 2195 1.2m 119.3h
aarch64 4038 51s 1.2m 3839 2.1m 8.5m 153 29s 19.6h
x86_64 4673 52s 2.0m 4058 1.7m 7.3m 26 1.6m 9.0h
i386 1022 55s 3.4m 993 1.9m 10.3m 8 24s 10.0m
noarch 12 51s 1.1m 10 1.3m 1.8m 0 - 0s

Official builds:

arch queued med-wait p90-wait built med-time p90-time gated med-delay tot-delay
s390x 920 2.3h 3.5h 901 3.6m 38.4m 812 1.9h 1618.8h
ppc64le 1003 21.0m 1.1h 986 2.9m 16.9m 109 1.6m 28.7h
aarch64 1015 48s 1.6m 994 2.3m 13.3m 11 49s 7.5h
x86_64 1044 53s 3.1m 1018 2.0m 11.3m 12 1.3m 42.0m
i386 723 56s 3.5m 708 2.0m 10.3m 7 24s 8.0m
noarch 6 56s 1.4m 6 55s 3.5m 0 - 0s

Scratch builds:

arch queued med-wait p90-wait built med-time p90-time gated med-delay tot-delay
s390x 378 3.2h 3.6h 357 4.0m 1.1h 308 2.4h 718.1h
ppc64le 2972 1.1m 11.0m 2805 2.8m 8.0m 2086 1.1m 90.6h
aarch64 3023 52s 1.2m 2845 2.1m 7.2m 142 29s 12.1h
x86_64 3629 52s 1.2m 3040 1.6m 5.8m 14 1.7m 8.3h
i386 299 53s 3.1m 285 1.8m 10.4m 1 2.0m 2.0m
noarch 6 24s 1.1m 4 1.3m 1.6m 0 - 0s

Column legend:

  • queued / built — tasks counted in the wait/time stats.
  • med-wait, p90-wait — task creation until a builder picked it up.
  • med-time, p90-time — builder start until completion.
  • gated — builds where this arch finished last, holding up the build.
  • med-delay / tot-delay — how long after the second-slowest arch the
    gating arch finished; the extra wall-clock time it alone cost those
    builds (median per build / summed over the window).
I've [started a tool ](https://github.com/slopfest/sandogasa/tree/main/tools/koji-lag)to analyze Koji build statistics -- this is what it looks like for yesterday (when the Python archful mass rebuild was happening) Right now it uses the Koji APIs, once this proves useful we can extend it to use a database dump (hopefully a selective one that only includes the data it needs). It can fetch one day's worth of build data at a time and combine the raw data, so we can slowly build the data to analyze over time and re-analyze it later Instances: fedora Gated builds: 3502 (critical-path attribution); unattributed tasks: 0 | arch | queued | med-wait | p90-wait | built | med-time | p90-time | gated | med-delay | tot-delay | |:--------|-------:|---------:|---------:|------:|---------:|---------:|------:|----------:|----------:| | s390x | 1298 | 2.6h | 3.6h | 1258 | 3.7m | 50.4m | 1120 | 2.0h | 2336.9h | | ppc64le | 3975 | 1.3m | 57.3m | 3791 | 2.8m | 9.4m | 2195 | 1.2m | 119.3h | | aarch64 | 4038 | 51s | 1.2m | 3839 | 2.1m | 8.5m | 153 | 29s | 19.6h | | x86_64 | 4673 | 52s | 2.0m | 4058 | 1.7m | 7.3m | 26 | 1.6m | 9.0h | | i386 | 1022 | 55s | 3.4m | 993 | 1.9m | 10.3m | 8 | 24s | 10.0m | | noarch | 12 | 51s | 1.1m | 10 | 1.3m | 1.8m | 0 | - | 0s | Official builds: | arch | queued | med-wait | p90-wait | built | med-time | p90-time | gated | med-delay | tot-delay | |:--------|-------:|---------:|---------:|------:|---------:|---------:|------:|----------:|----------:| | s390x | 920 | 2.3h | 3.5h | 901 | 3.6m | 38.4m | 812 | 1.9h | 1618.8h | | ppc64le | 1003 | 21.0m | 1.1h | 986 | 2.9m | 16.9m | 109 | 1.6m | 28.7h | | aarch64 | 1015 | 48s | 1.6m | 994 | 2.3m | 13.3m | 11 | 49s | 7.5h | | x86_64 | 1044 | 53s | 3.1m | 1018 | 2.0m | 11.3m | 12 | 1.3m | 42.0m | | i386 | 723 | 56s | 3.5m | 708 | 2.0m | 10.3m | 7 | 24s | 8.0m | | noarch | 6 | 56s | 1.4m | 6 | 55s | 3.5m | 0 | - | 0s | Scratch builds: | arch | queued | med-wait | p90-wait | built | med-time | p90-time | gated | med-delay | tot-delay | |:--------|-------:|---------:|---------:|------:|---------:|---------:|------:|----------:|----------:| | s390x | 378 | 3.2h | 3.6h | 357 | 4.0m | 1.1h | 308 | 2.4h | 718.1h | | ppc64le | 2972 | 1.1m | 11.0m | 2805 | 2.8m | 8.0m | 2086 | 1.1m | 90.6h | | aarch64 | 3023 | 52s | 1.2m | 2845 | 2.1m | 7.2m | 142 | 29s | 12.1h | | x86_64 | 3629 | 52s | 1.2m | 3040 | 1.6m | 5.8m | 14 | 1.7m | 8.3h | | i386 | 299 | 53s | 3.1m | 285 | 1.8m | 10.4m | 1 | 2.0m | 2.0m | | noarch | 6 | 24s | 1.1m | 4 | 1.3m | 1.6m | 0 | - | 0s | Column legend: - `queued` / `built` — tasks counted in the wait/time stats. - `med-wait`, `p90-wait` — task creation until a builder picked it up. - `med-time`, `p90-time` — builder start until completion. - `gated` — builds where this arch finished last, holding up the build. - `med-delay` / `tot-delay` — how long after the second-slowest arch the gating arch finished; the extra wall-clock time it alone cost those builds (median per build / summed over the window).
Sign in to join this conversation.
No milestone
No project
No assignees
5 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
infra/tickets#13141
No description provided.