Released on July 6, 2026
Hi Andromeda Community
We have returned Andromeda to back to its standard configuration. It was placed in a power saving configuration until on July 2 at the request of BC facilities as part of a region wide effort to relieve the grid during the heat wave late last week.
Thanks for everyone’s cooperation
Matt
Released on July 2, 2026
Hi Andromeda Community
Due to extreme temperatures we have taken the following actions to reduce power consumption on the Andromeda Cluster:
64-core CPUs have been put into a power saving profile
Half of 48-core CPUs are set to drain and remain inactive
10 GPU nodes are set to drain and will remain inactive.
What this means: CPU nodes that are in power saving profile will remain in service but users may see longer times to job completion. On the nodes set to drain all jobs will complete but no new jobs will start.
Early this morning we received information from BC facilities and our data center team that Eversource was expecting record consumption. Taking these actions will reduce pressure and help Boston College increase the reliability of the electrical grid.
Thanks
Matt Gregas, PhD
Director, Research Services
Boston College
Released on May 28, 2026
Dear Andromeda Researchers
We have been notified by Red Hat of critical, out-of-cycle software updates that must be applied immediately across all compute nodes. Because these updates modify fundamental operating system components, a full reboot of all affected nodes is required.
Our priority is to complete these necessary system updates as quickly as possible while minimizing disruption to your ongoing research and active workflows.
When
Start Time: This afternoon, May 28, 2026, at 5:00 PM.
What to Expect During the Maintenance
Login Node Downtime: The login node will be unavailable for approximately 30 minutes.
Compute Node Draining: All compute nodes will be placed into a “drain” status beginning at 5:00 PM.
Sequential Reboots: As individual compute nodes finish their current workloads and successfully drain, they will be restarted.
Return to Service: As soon as a node finishes rebooting, it will be returned to service and become available to accept new job submissions.
What This Means for Your Work and Active Jobs
Running Jobs Will Not Be Terminated: Any job that is currently active and running on the cluster at 5:00 PM will continue running uninterrupted until it naturally completes.
File Access: Except for the brief 30-minute window while the login node itself reboots, you will retain full access to view and manage your files.
Potential Job Queue Delays: Any jobs that are pending/not yet running, or any new jobs submitted after 5:00 PM, will experience longer-than-usual queue wait times to start.
Gradual Recovery: This scheduling lag will progressively shorten throughout the evening as more rebooted nodes are successfully returned to the active cluster pool.
We sincerely appreciate your cooperation and understanding
Thanks
Matt Gregas