Planned Outage - server moves - 2026-06-25 14:00 UTC #13419
Labels
No labels
announcement
anubis
authentication
aws
backlog
blocked
bodhi
ci
cloud
communishift
copr
database
day-to-day
dc-move
deprecated
dev
discourse
dns
downloads
easyfix
epel
firmitas
forgejo_migration
Gain
High
Gain
Low
Gain
Medium
gitlab
greenwave
hardware
help wanted
high-trouble
koji
koschei
lists
low-trouble
medium-trouble
mirrorlists
monitoring
Needs investigation
odcs
OpenShift
ops
outage
packager_workflow_blocker
pagure
permissions
Priority
Needs Review
Priority
Next Meeting
Priority
🔥 URGENT 🔥
Priority
Waiting on Assignee
Priority
Waiting on External
Priority
Waiting on Reporter
rabbitmq
release-monitoring
releng
request-for-resources
s390x
security
SMTP
sprint-0
sprint-1
src.fp.o
staging
unfreeze
waiverdb
websites-general
wiki
Backlog Status
Needs Review
Backlog Status
Ready
chore
documentation
points
01
points
02
points
03
points
05
points
08
points
13
Priority
High
Priority
Low
Priority
Medium
Sprint Status
Blocked
Sprint Status
Done
Sprint Status
In Progress
Sprint Status
Review
Sprint Status
To Do
Technical Debt
Work Item
Bug
Work Item
Epic
Work Item
Spike
Work Item
Task
Work Item
User Story
No milestone
No project
No assignees
5 participants
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
infra/tickets#13419
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
planned outage
Planned Outage - server moves - 2026-06-25 14:00 UTC
There will be an outage starting at 2026-06-25 14:00 UTC,
which will last approximately 5 hours.
To convert UTC to your local time, take a look at
http://fedoraproject.org/wiki/Infrastructure/UTCHowto
or run:
date -d '2026-06-26 14:00 UTC'
Reason for outage:
In order to reduce power use in 2 racks that are higher than others, datacenter folks will be moving 11 servers from the racks that they are in to two different ones. They will be connected to the same switch ports, etc, just physically in another rack.
Affected Services:
The servers to be moved are:
openqa-x86-worker03.rdu3.fedoraproject.org
openqa-x86-worker04.rdu3.fedoraproject.org
openqa-x86-worker05.rdu3.fedoraproject.org
vmhost-x86-copr01.rdu3.fedoraproject.org
vmhost-x86-copr02.rdu3.fedoraproject.org
vmhost-x86-copr03.rdu3.fedoraproject.org
worker04.ocp.rdu33.fedoraproject.org
worker05.ocp.rdu33.fedoraproject.org
worker01.stg.ocp.rdu33.fedoraproject.org
worker02.stg.ocp.rdu33.fedoraproject.org
worker03.stg.ocp.rdu33.fedoraproject.org
Ticket Link:
https://forge.fedoraproject.org/infra/tickets/issue/13419
Please join #admin:fedoraproject.org / #noc:fedoraproject.org on matrix.
or add comments to this ticket for more information or feedback.
Updated status for this outage may be available at
https://www.fedorastatus.org/
Note that we may not need to announce this widely, as we may be able to do things with minimal notice.
We should be able to evacuate the prod openshift workers, but we should check that this will not cause a problem with storage for the other ones.
staging will be down since all those workers are all the workers staging has.
CC @adamwill for openqa workers. I hope we can just take them down and run on reduced capacity while they move things?
CC @praiskup on the copr ones. Same question.
I'd have to resurrect an old ansible workflow and shuffle things around a bit on openQA staging to cope, as 04 and 05 are the tap worker hosts for staging. I'd have to bring back the old 'tap12' path where a worker can be both tap1 and tap2, and make 06 be it for staging. Prod should be fine as 03 is its 'boring not-a-tap-job-host' box.
@kevin no need to let us know with Copr vmhosts, just shut the machines down and Copr will re-direct the traffic to other workers.
@gwmngilfen will shut down the copr hosts, @darknao will handle the openshift workers and hopefully @adamwill will handle openqa hosts.
If folks could have them powered off at the start of the window that would be great.
I'll be on trains that day, but will try to prepare for this tomorrow I guess.
Actually, I changed my mind. If the outage will just be five hours it'll be fine to just let it happen and let tap jobs pile up on staging, I think. The system should recover automatically when the hosts are back online (they'll pick up the waiting jobs and run them). It's only staging so it doesn't affect anything prod.
Heads up @lruzicka that x86_64 tap jobs will likely not run for several hours on staging tomorrow.
@adamwill Noted, thanks for the info.
Yeah, it may be less, thats just a conservative window.
They do want us to power things off before the outage... can you do the openqa machines? Or perhaps @gwmngilfen could do them when he does the copr ones?
I can do the copr ones, no problem there. I'll aim for 13:00 UTC - @adamwill I can also do the openqa ones if you dont get to it by then.
We hit an issue with timeouts in the ocp Apache loadbalancer from the slower proxies, which meant users were seeing timeouts. With help from @darknao I removed them from the lb and updated the proxies.
After the outage we'll need to revert infra/ansible@!3439 (commit
bdf834a3e9) and then re-run the reverseproxy playbook.I've shutdown the copr and openqa nodes. @darknao has done the ocp and ocp.stg nodes. Good to go 👍
This outage is over. All the nodes should be back up and working.
As part of this we found that the staging openshift workers aren't properly using 802.11ad... we will need to fix that.