Planned Outage - server moves - 2026-06-25 14:00 UTC #13419

Closed
opened 2026-06-18 21:39:47 +00:00 by kevin · 12 comments
Owner

planned outage

Planned Outage - server moves - 2026-06-25 14:00 UTC

There will be an outage starting at 2026-06-25 14:00 UTC,
which will last approximately 5 hours.

To convert UTC to your local time, take a look at
http://fedoraproject.org/wiki/Infrastructure/UTCHowto
or run:

date -d '2026-06-26 14:00 UTC'

Reason for outage:

In order to reduce power use in 2 racks that are higher than others, datacenter folks will be moving 11 servers from the racks that they are in to two different ones. They will be connected to the same switch ports, etc, just physically in another rack.

Affected Services:

The servers to be moved are:

openqa-x86-worker03.rdu3.fedoraproject.org
openqa-x86-worker04.rdu3.fedoraproject.org
openqa-x86-worker05.rdu3.fedoraproject.org
vmhost-x86-copr01.rdu3.fedoraproject.org
vmhost-x86-copr02.rdu3.fedoraproject.org
vmhost-x86-copr03.rdu3.fedoraproject.org
worker04.ocp.rdu33.fedoraproject.org
worker05.ocp.rdu33.fedoraproject.org
worker01.stg.ocp.rdu33.fedoraproject.org
worker02.stg.ocp.rdu33.fedoraproject.org
worker03.stg.ocp.rdu33.fedoraproject.org

Ticket Link:

https://forge.fedoraproject.org/infra/tickets/issue/13419

Please join #admin:fedoraproject.org / #noc:fedoraproject.org on matrix.
or add comments to this ticket for more information or feedback.

Updated status for this outage may be available at
https://www.fedorastatus.org/

### planned outage Planned Outage - server moves - 2026-06-25 14:00 UTC There will be an outage starting at 2026-06-25 14:00 UTC, which will last approximately 5 hours. To convert UTC to your local time, take a look at http://fedoraproject.org/wiki/Infrastructure/UTCHowto or run: date -d '2026-06-26 14:00 UTC' Reason for outage: In order to reduce power use in 2 racks that are higher than others, datacenter folks will be moving 11 servers from the racks that they are in to two different ones. They will be connected to the same switch ports, etc, just physically in another rack. Affected Services: The servers to be moved are: openqa-x86-worker03.rdu3.fedoraproject.org openqa-x86-worker04.rdu3.fedoraproject.org openqa-x86-worker05.rdu3.fedoraproject.org vmhost-x86-copr01.rdu3.fedoraproject.org vmhost-x86-copr02.rdu3.fedoraproject.org vmhost-x86-copr03.rdu3.fedoraproject.org worker04.ocp.rdu33.fedoraproject.org worker05.ocp.rdu33.fedoraproject.org worker01.stg.ocp.rdu33.fedoraproject.org worker02.stg.ocp.rdu33.fedoraproject.org worker03.stg.ocp.rdu33.fedoraproject.org Ticket Link: https://forge.fedoraproject.org/infra/tickets/issue/13419 Please join #admin:fedoraproject.org / #noc:fedoraproject.org on matrix. or add comments to this ticket for more information or feedback. Updated status for this outage may be available at https://www.fedorastatus.org/
kevin self-assigned this 2026-06-18 21:39:47 +00:00
Author
Owner

Note that we may not need to announce this widely, as we may be able to do things with minimal notice.

We should be able to evacuate the prod openshift workers, but we should check that this will not cause a problem with storage for the other ones.

staging will be down since all those workers are all the workers staging has.

CC @adamwill for openqa workers. I hope we can just take them down and run on reduced capacity while they move things?

CC @praiskup on the copr ones. Same question.

Note that we may not need to announce this widely, as we may be able to do things with minimal notice. We should be able to evacuate the prod openshift workers, but we should check that this will not cause a problem with storage for the other ones. staging will be down since all those workers are all the workers staging has. CC @adamwill for openqa workers. I hope we can just take them down and run on reduced capacity while they move things? CC @praiskup on the copr ones. Same question.
Member

I'd have to resurrect an old ansible workflow and shuffle things around a bit on openQA staging to cope, as 04 and 05 are the tap worker hosts for staging. I'd have to bring back the old 'tap12' path where a worker can be both tap1 and tap2, and make 06 be it for staging. Prod should be fine as 03 is its 'boring not-a-tap-job-host' box.

I'd have to resurrect an old ansible workflow and shuffle things around a bit on openQA staging to cope, as 04 and 05 are the tap worker hosts for staging. I'd have to bring back the old 'tap12' path where a worker can be both tap1 and tap2, and make 06 be it for staging. Prod should be fine as 03 is its 'boring not-a-tap-job-host' box.
Member

@kevin no need to let us know with Copr vmhosts, just shut the machines down and Copr will re-direct the traffic to other workers.

@kevin no need to let us know with Copr vmhosts, just shut the machines down and Copr will re-direct the traffic to other workers.
Author
Owner

@gwmngilfen will shut down the copr hosts, @darknao will handle the openshift workers and hopefully @adamwill will handle openqa hosts.

If folks could have them powered off at the start of the window that would be great.

@gwmngilfen will shut down the copr hosts, @darknao will handle the openshift workers and hopefully @adamwill will handle openqa hosts. If folks could have them powered off at the start of the window that would be great.
Member

I'll be on trains that day, but will try to prepare for this tomorrow I guess.

I'll be on trains that day, but will try to prepare for this tomorrow I guess.
Member

Actually, I changed my mind. If the outage will just be five hours it'll be fine to just let it happen and let tap jobs pile up on staging, I think. The system should recover automatically when the hosts are back online (they'll pick up the waiting jobs and run them). It's only staging so it doesn't affect anything prod.

Heads up @lruzicka that x86_64 tap jobs will likely not run for several hours on staging tomorrow.

Actually, I changed my mind. If the outage will just be five hours it'll be fine to just let it happen and let tap jobs pile up on staging, I think. The system should recover automatically when the hosts are back online (they'll pick up the waiting jobs and run them). It's only staging so it doesn't affect anything prod. Heads up @lruzicka that x86_64 tap jobs will likely not run for several hours on staging tomorrow.

@adamwill Noted, thanks for the info.

@adamwill Noted, thanks for the info.
Author
Owner

Yeah, it may be less, thats just a conservative window.

They do want us to power things off before the outage... can you do the openqa machines? Or perhaps @gwmngilfen could do them when he does the copr ones?

Yeah, it may be less, thats just a conservative window. They do want us to power things off before the outage... can you do the openqa machines? Or perhaps @gwmngilfen could do them when he does the copr ones?
Member

I can do the copr ones, no problem there. I'll aim for 13:00 UTC - @adamwill I can also do the openqa ones if you dont get to it by then.

I can do the copr ones, no problem there. I'll aim for 13:00 UTC - @adamwill I can also do the openqa ones if you dont get to it by then.
Member

We hit an issue with timeouts in the ocp Apache loadbalancer from the slower proxies, which meant users were seeing timeouts. With help from @darknao I removed them from the lb and updated the proxies.

After the outage we'll need to revert infra/ansible@!3439 (commit bdf834a3e9) and then re-run the reverseproxy playbook.

We hit an issue with timeouts in the ocp Apache loadbalancer from the slower proxies, which meant users were seeing timeouts. With help from @darknao I removed them from the lb and updated the proxies. After the outage we'll need to revert https://forge.fedoraproject.org/infra/ansible/pulls/3439/commits/bdf834a3e91480bb6c1229ba4fdef4209ebe2361 and then re-run the reverseproxy playbook.
Member

I've shutdown the copr and openqa nodes. @darknao has done the ocp and ocp.stg nodes. Good to go 👍

I've shutdown the copr and openqa nodes. @darknao has done the ocp and ocp.stg nodes. Good to go 👍
Author
Owner

This outage is over. All the nodes should be back up and working.

As part of this we found that the staging openshift workers aren't properly using 802.11ad... we will need to fix that.

This outage is over. All the nodes should be back up and working. As part of this we found that the staging openshift workers aren't properly using 802.11ad... we will need to fix that.
kevin closed this issue 2026-06-25 18:16:45 +00:00
Sign in to join this conversation.
No milestone
No assignees
5 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
infra/tickets#13419
No description provided.