Troubleshooting¶
Work through the section that matches your symptom. If nothing here fixes it, open a ticket with the details listed at the bottom of that page — cluster, job ID, exact command, and full error message.
I can't log in¶
- Check for maintenance on the cluster status monitor and UNM IT alerts.
- Password or OTP problems — reset via the steps in password reset. A fresh reset can take a little while to propagate (try again shortly, or reset twice), and if one cluster works but the other doesn't, SSH to the broken one from the working cluster's login node while you sort it out.
Permission denied (publickey)— your SSH key setup is incomplete or has wrong permissions; see SSH keys (~/.sshmust be700, private keys600).- Account exists but no access — you may not be on an active project yet; ask your PI to add you in ColdFront (getting started).
"Disk quota exceeded"¶
You've hit a storage limit (what the limits are):
quotas # show your usage against each quota
du -sh ~/* | sort -rh | head # find what's using home space
Clean up, move bulk data to scratch or project space
(storage layout), or talk to us about
purchasing more.
Remember conda environments and pip caches grow quietly — conda clean --all
and pip cache purge often free gigabytes.
Two quota surprises worth knowing:
- Files in a project directory can still count against you — quota is charged by group ownership, not location, so files you copied into a project space may still bill your personal quota; see storage permissions for finding and fixing ownership.
- The warning email covers only one group — run
quotasfor the full per-tier breakdown before deleting anything. If the reported usage doesn't match any files you can actually find, the accounting itself may be stale — open a ticket instead of deleting more data.
My job won't start¶
squeue -u $USER # state and reason code
squeue --start --job <id> # predicted start time (fairshare-aware)
sinfo # partition and node availability
Common reason codes:
| Reason | Meaning | What to do |
|---|---|---|
Priority |
Others are ahead of you (fairshare) | Wait, or request fewer/shorter resources; see fairshare |
Resources |
Not enough free nodes for your request | Reduce cores/memory/GPUs or choose another partition |
QOSMax* / limits |
You've hit a partition or account limit | Check resource limits |
ReqNodeNotAvail |
Nodes down or reserved (often maintenance) | Check the cluster status monitor |
InvalidAccount |
Wrong --account |
List yours: sacctmgr show assoc user=$USER format=account |
Rejected with uid not in group permitted to use this partition — the
partition is group-gated (on Easley that includes the h100 and l40s GPU
partitions). Access is provisioned through a ColdFront allocation for that
specific partition, requested by your PI — support cannot simply add you to
the group. If no allocation exists yet, ask your PI to submit one in
ColdFront; after
approval, allow some time for group membership to propagate to the cluster.
Need to run something right now? The scavenger partition doesn't have this
gate (jobs there are preemptible). If the error persists well after an
approved allocation, open a ticket — that's a
provisioning problem, not a normal delay.
Stuck in CG (completing) — a few minutes in CG after a job finishes
is normal cleanup. If it persists, scancel will not clear it — the job is
stuck in Slurm's own cleanup, which needs admin action — so don't keep
retrying; open a ticket with the job ID.
My job failed or was killed¶
sacct -j <id> --format=JobID,State,ExitCode,Elapsed,MaxRSS,ReqMem
seff <id> # efficiency summary after completion
OUT_OF_MEMORY/oom-kill(often just a bareKilledfrom your program) — Slurm enforces the memory your job requested, not what the node has free, so a job on a shared node can be killed while the node itself shows plenty of RAM. If you never set--mem, the default is proportional to the CPUs you requested (DefMemPerCPU— ≈3.7 GB/CPU on Easleygeneral, ≈2.9 GB/CPU on Hoppergeneral), so a small--cpus-per-tasksilently caps memory. Resubmit with--mem(or--mem-per-cpu) sized to your data's actual working set — well above the raw data size for tools that process in memory. For variable workloads, runseffon a smaller successful run first to calibrate. Ifseffshows the state was notOUT_OF_MEMORY, or memory used was well under what you requested, more memory is not the fix — suspect an application bug and open a ticket.TIMEOUT— raise--timewithin partition limits, or checkpoint and restart.- Immediate crash — check the job's
.out/.errfiles in the submit directory; a missingmodule loador unactivated conda environment is the usual culprit (modules, conda).
Software and environment problems¶
command not found— load the module first (module spider <name>to find it; modules guide).- Python
ModuleNotFoundErroraftermodule load miniconda3— the module provides only conda's base environment, which doesn't include numpy, scipy, or other packages: create and activate your own conda environment. Old scripts thatmodule load anaconda3must switch tominiconda3— that module is retired. Build environments from an interactive job, not a login node. - Conda is slow or conflicts — prefer clean per-project environments and the conda-forge channel; see channels and pip.
- GPU code can't see the GPU — did you request one in the job
(
--gres=gpu:1or the cluster's GPU partition)? Verify withnvidia-smiinside the job; see example Slurm scripts. - My kernel is missing in JupyterHub — register your environment as a kernel: conda in JupyterHub.
Graphics won't display¶
X11 applications need forwarding enabled — ssh -Y and a local X server;
see X11 forwarding. For heavier
visualization, use ParaView client–server or an
Open OnDemand session instead.
Transfers are slow or failing¶
Use rsync with resume (rsync -avP) rather than scp for large trees,
and transfer to the right storage tier — see
transferring data.
A large transfer that repeatedly hangs or times out usually points to
client-side network stability (wireless, VPN, off-campus path) rather than
CARC. Chunk it: loop over subdirectories with separate rsync calls instead
of one massive invocation — reruns resume where they left off. Still stuck?
Open a ticket noting whether it dies at the same file
or at random, wired vs. wireless, and on- vs. off-campus; support can try
reproducing the transfer to rule out a CARC-side issue.
Still stuck?¶
Open a ticket or bring it to office hours — include your cluster, job ID, command, and the complete error text.