Released on August 11, 2026
Dear Andromeda Researchers
I just wanted to send a friendly reminder that large computational jobs or processes should not be run on the login node. The login node is the gateway to the cluster and is not provisioned to run heavy compute jobs. It has only 8 core and 96GB of RAM. For script testing we do have an interactive sessions which will allow you to test scripts on our more high-powered compute nodes. When a job or process consumes too many resources the login node becomes unresponsive and folks are not able to login, additionally other key services may fail.
If you are not sure if you are executing processes that are too large for the login node please let us know and we will be happy to show you how to submit jobs via Slurm or launch interactive sessions.
We have also noticed that many of the processes on login node have been submitted/executed by AI agents. This makes applying remedies problematic because when we kill a process that is using large amounts of CPU or memory the agent often launches a new job to replace the killed process. Please do not allow or instruct an AI agent or coding tool to launch jobs autonomously on Andromeda.
We are always willing to assist if you have questions. Feel free to submit a ticket using bc.edu/researchhelp.
Thanks
Matt
Released on July 27, 2026
Dear Andromeda Community,
Some of you may have noticed that the login node became unresponsive for two distinct periods Wednesday evening. The root cause of the issue is that an autonomous AI agent was executing processes on the login node that consumed available CPU resources. Our efforts to resolve the issue were complicated because the agent would relaunch the process shortly after our system administrators killed it.
In the short term we have taken steps to mitigate saturation of the login node.
In the bigger picture autonomous AI agents pose a unique risk in HPC environments for a variety of reasons, in particular, they are a risk to the contents of the shared file system. At this time we ask all users not to let AI agents launch jobs, processes or other executables. That is, only humans are allowed to submit jobs on Andromeda. If we discover jobs launched by AI agents or similar we will first kill the jobs and then reach out to the user with further guidance.
We understand that AI coding tools and agents can greatly speed up research and we have benefitted from the efficiency gains by these tools. Our goal is to find a solution that allows researchers to safely benefit from the AI tools and products. The Research Service team is working with their ITS colleagues to develop a set of technical guardrails, policy guidelines and educational resources.
This is an active area of development of Research Computing units nationwide. Although many unique solutions exist there is a consensus that humans remain “in the loop” and general prohibitions against allowing AI agents to submit jobs or start processes on HPC systems.
Thank you for your cooperation in maintaining the integrity and stability of the Andromeda Cluster.
Matt Gregas, PhD
Research Services
Boston College
PS: As a reminder any installed software needs to be approved by the Boston College Get Tech committee. This is required by ITS policy.
Released on July 16, 2026
To: All cluster account holders
Subject: Scheduled downtime: July 25
Greetings
You are receiving this notice because you either have a cluster account or utilize a virtual machine//physical server that is housed in Boston College’s data center.
Following up on the general ITS notice sent on July 7:
On July 25 the Boston College Network and Infrastructure teams will be replacing switches in the Data Center. During this period Andromeda (HPC cluster), virtual machines/physical server and dependent services will be unavailable. Power down is not needed meaning job on Andromeda and applications on virtual machines will continue to run, however you will not be able to login to the cluster or access other services.
As stated in the notice we expect services to return to normal by 12:00 pm on July 25
Thanks
Matt Gregas, PhD
Director, Research Services
Boston College
Released on July 8, 2026
Dear Cluster Community
On July 25 the Boston College Network and Infrastructure teams will be replacing switches in the Data Center where the Andromeda Cluster resides. During this you will not be able to login to the Andromeda. We will return login service to Andromeda once the repairs are completed. Our goal is to have the cluster available on the morning of Monday July 27.
This will affect the logins only, any jobs running prior to July 24 at 8 pm will continue to run
We will send out more notices as we approach July 25
Thanks
Matt Gregas, PhD
Director, Research Services
Boston College
P.S. We will be introducing an Open On Demand Tutorial this summer. For more information on this and other tutorials please visit:
https://www.bc.edu/content/bc-web/offices/its/services/research-services/tutorials.html
Released on July 6, 2026
Hi Andromeda Community
We have returned Andromeda to back to its standard configuration. It was placed in a power saving configuration until on July 2 at the request of BC facilities as part of a region wide effort to relieve the grid during the heat wave late last week.
Thanks for everyone’s cooperation
Matt
Released on July 2, 2026
Hi Andromeda Community
Due to extreme temperatures we have taken the following actions to reduce power consumption on the Andromeda Cluster:
64-core CPUs have been put into a power saving profile
Half of 48-core CPUs are set to drain and remain inactive
10 GPU nodes are set to drain and will remain inactive.
What this means: CPU nodes that are in power saving profile will remain in service but users may see longer times to job completion. On the nodes set to drain all jobs will complete but no new jobs will start.
Early this morning we received information from BC facilities and our data center team that Eversource was expecting record consumption. Taking these actions will reduce pressure and help Boston College increase the reliability of the electrical grid.
Thanks
Matt Gregas, PhD
Director, Research Services
Boston College
Released on May 28, 2026
Dear Andromeda Researchers
We have been notified by Red Hat of critical, out-of-cycle software updates that must be applied immediately across all compute nodes. Because these updates modify fundamental operating system components, a full reboot of all affected nodes is required.
Our priority is to complete these necessary system updates as quickly as possible while minimizing disruption to your ongoing research and active workflows.
When
Start Time: This afternoon, May 28, 2026, at 5:00 PM.
What to Expect During the Maintenance
Login Node Downtime: The login node will be unavailable for approximately 30 minutes.
Compute Node Draining: All compute nodes will be placed into a “drain” status beginning at 5:00 PM.
Sequential Reboots: As individual compute nodes finish their current workloads and successfully drain, they will be restarted.
Return to Service: As soon as a node finishes rebooting, it will be returned to service and become available to accept new job submissions.
What This Means for Your Work and Active Jobs
Running Jobs Will Not Be Terminated: Any job that is currently active and running on the cluster at 5:00 PM will continue running uninterrupted until it naturally completes.
File Access: Except for the brief 30-minute window while the login node itself reboots, you will retain full access to view and manage your files.
Potential Job Queue Delays: Any jobs that are pending/not yet running, or any new jobs submitted after 5:00 PM, will experience longer-than-usual queue wait times to start.
Gradual Recovery: This scheduling lag will progressively shorten throughout the evening as more rebooted nodes are successfully returned to the active cluster pool.
We sincerely appreciate your cooperation and understanding
Thanks
Matt Gregas