Runner guides · Architecture

Ephemeral self-hosted runners: what survives a CI job?

A single-job registration and a clean machine solve different problems. Plan the full lifecycle, including the data you need after teardown.

Baremetal · Updated

Single-job registration is one part of the lifecycle

An ephemeral GitHub Actions runner accepts one job, then its registration is removed. That does not, by itself, erase the computer it ran on. You also need a way to provision the environment and discard it after execution. GitHub recommends ephemeral runners for autoscaling and describes this distinction in its self-hosted runner reference.

Baremetal combines a single-job registration with a fresh VM on your own host. On Apple Silicon that VM runs through Tart; on Linux it runs through QEMU/KVM. The guest is removed after the job. The physical host and cached base image remain available for future jobs.

Decide what should survive

Data lifetime in a Baremetal runner
DataAfter the jobYour responsibility
Guest workspace and installed packagesDiscarded with the VMPut repeatable setup in the image or workflow.
Base VM imageRetained on the host for reuseVersion and maintain the image; keep secrets out of it.
Uploaded build artifactsRetained by the artifact serviceChoose access controls and retention periods.
External build cacheRetained by the cache serviceChoose keys, trust boundaries and expiry.

Teardown removes the job environment. It does not revoke credentials the job used or delete files uploaded elsewhere. Review the whole data path, including logs and artifacts, when deciding which workflows may access sensitive material.

Keep the environment repeatable without rebuilding every dependency

Separate three kinds of state. The base image contains tools shared by many jobs, such as a compiler or Xcode. A build cache contains reusable outputs keyed to inputs. Artifacts are the outputs you deliberately retain or ship. They need different update and retention policies.

For an Xcode project, pin the toolchain and simulator runtimes in the image. Cache dependencies with keys that account for the lockfile and toolchain. Upload the app or test results before the guest disappears. For a Linux project, apply the same separation to compiler tools, dependency downloads and release binaries.

Baremetal does not automatically preserve a guest’s Docker layers or dependency cache between jobs. Configure an external cache if you need one. Measure restore and save times: a large cache can cost more time to transfer than it saves in compilation.

Publishing a new image is an explicit rollout. Use a versioned image reference and test it in a separate pool before moving production workflows. For KVM images, a new versioned URL also avoids reusing the old base image cached under the previous URL.

Isolation still depends on what the workflow can reach

A fresh VM reduces accidental state carrying over to the next job. It does not make arbitrary code safe to give a deployment token or signing key. Restrict those credentials to trusted workflows and environments, and keep them out of base images and caches.

Baremetal restricts guest networking as described in the security model. Plan workflows around that policy: a job that needs a private internal service should not assume it can reach every address accessible from the host. Review the public repository guidance before accepting outside contributions.

Plan for startup and failure, not only successful builds

Track image download, VM boot, runner registration and execution separately. A warm runner uses resources before a job arrives but can reduce its wait. A pool with no idle runners releases those resources and pays the startup cost on the next job. Choose the target using your arrival pattern and available host capacity.

Keep diagnostic logs outside the disposable guest when you need them after failure. GitHub recommends preserving ephemeral runner logs externally; see its logging guidance. Workflow logs and runner startup logs answer different questions, especially when a job never reaches its first step.

Before moving an important workflow, test a successful job, a failed job and cancellation. Confirm that the expected artifacts survive, the VM is removed and a subsequent job gets a clean workspace. Then check a cold start after the pool has been idle.

Start with one pool and one workflow

Choose a representative build, create a pool from a known image and route that workflow using a dedicated label. Keep an easy way to return to the previous runner while comparing results. Expand after the cleanup, cache and image-update process works reliably.

Follow the macOS runner guide for Apple builds or GitHub Actions on bare metal for host selection. If a new job never starts, use the queued-job checklist.

All runner guides →