es_cott group resources on Euler

loading…

GPU queues

es_cott has three SLURM accounts, one per partition family, each with its own share and queue. Click a card to select a queue; every chart and table below shows that queue only.
Showing queue

Concurrent GPUs

hourly, from job start/end intervals; black line = entitled steady state (NormShares × partition GPUs)

GPU-hours per day

Fairshare over time

sampled every 15 min from sshare since the dashboard started; Euler keeps no history of this

Running GPUs by card model

Pending GPUs by reason

Priority: waiting on fairshare. AssocGrpGRES: over the coordinator cap. QOSMaxGRESPerUser: over the QOS cap. Resources: nothing free at that size.

Queue wait, last 30 days

GPU utilization of finished jobs, weighted by GPU-hours (60 days)

gpuutil is a 30 s time average per job, divided by the job's GPU count. Jobs shorter than 10 min excluded. Bursty data-loading workloads read low legitimately.

Members

members with any use or queue activity in the window; click a row for detail. FairShare and LevelFS from sshare; caps from sacctmgr; hours from sacct intervals.

Live queue

Partition capacity

GPU counts per node state from sinfo; partition totals include drained nodes.
QOS per-user caps in force (recomputed by the site over time)