Runner guides · Troubleshooting

GitHub Actions job stuck queued? Check your self-hosted runners

Follow a job from its requested labels to a ready runner. Find the missing link before restarting machines or increasing capacity.

Baremetal · Updated

1. Read what the job is actually waiting for

Open the individual job in GitHub Actions. A workflow waiting for an environment approval, a previous job, or a concurrency group needs a different fix from a job waiting for a runner. Record the job URL, requested labels and the time it entered the queue.

Check needs, environment and concurrency in the workflow. An intentionally serialized deployment will not start sooner because you add another machine. GitHub documents concurrency and environment protection rules separately from runner scheduling.

2. Match every label, not just the operating system

A job using an array of runs-on labels needs a runner matching all of them. A machine being online is not enough. Compare the labels shown on the job with those on an actual registered runner, including any custom size or toolchain label.

runs-on: [self-hosted, linux, x64, baremetal-linux]

Here, a Linux x64 runner without baremetal-linux will not match. If that label was renamed on the pool but remains in the workflow, the job can wait even when the fleet has spare resources. Add the intended label or update the workflow after checking that the pool has the right image and resources.

For matrix jobs, inspect the resolved labels of the queued matrix entry. One entry may target a pool that no longer exists while the other entries run successfully.

3. Check repository access and runner status

In GitHub, open Settings → Actions → Runners for the repository or organization. Check that the runner is available to this repository. For an organization runner, review runner-group access as well as the labels.

  • Idle: check labels and access if the job still cannot start.
  • Active: the runner is already executing work; inspect other matching runners.
  • Offline: check the runner process and its connection to GitHub.
  • No matching runner: investigate provisioning and registration.

With ephemeral runners, an empty list can be normal while a pool is idle. During a queued job, watch whether a new runner appears and becomes ready. GitHub’s runner troubleshooting guide describes the status view and runner diagnostic logs.

4. Follow provisioning from the host to the guest

In Baremetal, a connected host and a ready GitHub runner are separate things. The agent can be online while an image is downloading, a VM is booting, or the runner inside the guest is trying to register. Check the machine, pool and current activity in the dashboard.

If provisioning stalls, read the agent log on the affected host:

# Linux host
journalctl -u baremetal-agent --since "30 minutes ago"

# macOS host
tail -n 100 ~/.baremetal/agent.log

Look for image-download failures, insufficient disk space, VM startup errors and registration failures around the queued job’s timestamp. Test network access from the environment that failed: a host downloading an image and a guest connecting to GitHub have different network paths. A successful request from your laptop proves neither.

5. Separate missing capacity from a slow cold start

Compare the pool’s guest CPU and memory requirements with the resources left on eligible hosts after reserves and running VMs. Check per-machine limits too. Spare capacity in a Linux pool cannot run a macOS job, and Baremetal caps each Mac at two concurrent VMs.

A warm-runner target can reduce startup waits when capacity permits. It cannot create RAM or physical machines. If every matching runner is busy, consider reducing unnecessary matrix parallelism, moving other workloads, or adding capacity. If each job waits mostly for image preparation, address the image and warm capacity first.

Inspect runner update failures too. GitHub can stop routing jobs to outdated runner software. Keep the runner current; do not work around the problem by turning off update checks. The runner reference explains routing and update requirements.

Confirm the fix with a small job

Run this manually on a Linux pool with the example labels, replacing the custom label with your own. It requires no repository checkout or deployment credentials.

name: Check runner routing
on: workflow_dispatch
permissions:
  contents: read
jobs:
  check:
    runs-on: [self-hosted, linux, x64, baremetal-linux]
    steps:
      - run: uname -a

Once it runs, retry the original workflow and compare its queue time. Keep a note of the root cause: label mismatch, access policy, guest startup or capacity. That makes the next incident easier to diagnose. For initial enrollment, use the setup guide.

All runner guides →