How it works¶
How deployments work¶
Existing image: runway resolves the tag to a digest (registry v2
HEAD …/manifests/<tag> with the ADC token for Google registries or the
anonymous token flow for Docker Hub) and deploys repository@sha256:….
Source build: the build strategy is chosen per stage: an explicit
dockerfile, or an explicit builder (buildpacks); otherwise the context's
Dockerfile if it exists, else buildpacks with
gcr.io/buildpacks/builder:latest. Buildpacks run in Cloud Build with
gcr.io/k8s-skaffold/pack (pack build IMAGE --builder … --network
cloudbuild), and Cloud Build pushes the image and reports its digest, as for
Dockerfile builds. The strategy and builder are part of the source hash, so
switching between them produces a new image.
When an image is rebuilt. The image tag src-<hash> is a content address
of everything that determines the image:
- the uploaded files (code, dependency manifests, Dockerfile) and the build strategy/builder name;
- the current digests of the base images: the Dockerfile's
FROMandCOPY --from=<image>references (withARGdefaults substituted), or the buildpacks builder. A new upstream image under the same tag (security patches innode:22-slim, a newgcr.io/buildpacks/builder:latest) changes the tag, so the next deploy rebuilds.
If an image with that tag already exists, the build is skipped (config-only
changes and promotions to another stage reuse it). The base images are
recorded on the release (runway.dev/base-images), and plan explains a
rebuild caused by a base image update. service.rebuild: always (or
--force-build) rebuilds on every deploy. Not detected: dependencies that are
not pinned (no lockfile) and files excluded by .runwayignore.
Release tags. deploy --tag reads the latest version from the changelog
(CHANGELOG.md, CHANGELOG, CHANGES.md, HISTORY.md, in the build context
or next to runway.yaml): the first heading containing vX.Y.Z or X.Y.Z,
skipping "Unreleased", used as written. It tags the image with that version
in Artifact Registry, refusing to move a version that already tags another
image. deploy --tag-rc tags X.Y.Z-RC<n>, n being one more than the
highest existing X.Y.Z-RC* tag; an image that already has an RC tag for that
version keeps it (re-running a deploy does not create RC2, RC3…). The tag is
shown by deploy and info (annotation runway.dev/release).
- Scan the build context and hash a deterministic tar stream (sorted
entries, fixed timestamps and owners), so identical source gives an
identical SHA-256. Nothing is compressed or buffered at this stage; this is
all
plandoes. Ignore rules: - always excluded:
.git/,.hg/,.svn/,.runway/,.env,.env.*(except.env.example|sample|template),*.pem,*.p12,*.pfx,id_rsa*,id_ecdsa*,id_ed25519*,.ssh/,.aws/,.gcloud/,.config/gcloud/,.netrc,application_default_credentials.json,credentials.json,gha-creds-*.json,*.tfstate*,.terraform/, and the runway config file itself (so config-only changes do not rebuild); - then
.runwayignore(gitignore syntax and semantics: an excluded directory cannot be re-entered) if present, otherwise.dockerignore(Docker semantics: patterns are anchored at the context root, the last matching rule wins, and!re-includes files even inside excluded directories, so an allowlist such as*then!src/main.pyworks); - the Dockerfile and
.dockerignoreare always included. - If
<repo>/<app>:src-<hash>already exists, skip to step 6. If a build of the same source is already running, attach to it. - Compress the tar stream with parallel gzip (all CPU cores, a standard
single-member gzip stream) into an anonymous temporary file, re-checking
that the content still matches the hash, and upload it, streamed from
disk, to
gs://<bucket>/runway/<app>/source-<sha256>.tar.gz. Compression overlaps with the in-flight build lookup. - Submit a regional Cloud Build (
docker build, push, build service account,CLOUD_LOGGING_ONLY), pinned to the uploaded object generation. - Monitor it with bounded polling; on failure show the status, failing step, the last log lines and the log URL. Ctrl-C cancels the build.
- Use the pushed digest from the build results.
Service reconciliation (Cloud Run Admin API v2):
- Read the service. If it exists without runway's labels (or belongs to
another app/stage) runway stops (
--adopttakes over an unlabeled one). - Compare desired and live state field by field; if nothing differs, no request is sent.
- Create (
serviceId=<app>-<stage>) or update with an update mask (labels, annotations, client, client_version, ingress, invoker_iam_disabled, template, traffic) and the currentetag. Labels and annotations set by others are preserved; runway owns the revision template. Traffic goes 100% to the latest ready revision. - Wait for the long-running operation and for reconciliation to finish
(
reconciling=false, terminal condition succeeded, latest ready revision = latest created revision), bounded by--timeout. - Reconcile the
allUsersinvoker binding with a read-modify-write on the service IAM policy (etag-protected, conditional bindings untouched).
Full step order of deploy (each step reads the live state, changes only
what is missing, and is retried per the retry policy):
- Enable missing APIs (
enable_apis): nothing else works without them. - Inspect: read the service (ownership check) while hashing the source and resolving the image.
- Bind
provider.tagsto the project and wait until they are effective (organization policy conditions such asresource.matchTag(...)read them; a service cannot carry a tag before it exists, so policies checked at creation need the tag on the project). If such a tag was bound in this run, organization policy refusals of the service are retried while the policy engine catches up; otherwise they fail immediately. - Create or update buckets (
buckets:, and the build source bucket). - Create the Artifact Registry repository (
create_build_resources). - Create the build and runtime service accounts.
- Grant roles: the build service account's (log writer, repository writer,
source reader), then
identity.roles(IAM policies via read-modify-write that preserves other bindings; BigQuery datasets via their access list). - Build the image (skipped when it already exists).
- Create/update the service (including volumes and
iap_enabled) and wait for readiness. When the service already exists,service.tagsare bound and awaited before this update (an organization policy may need them to accept it); this also resumes an interruptedbootstrapfirst deploy. - Bind
service.tagsof a service created in this run and wait until they are effective (before public access, so a tag that an organization policy requires forallUsersis in place first). - IAP: make sure the IAP service agent exists and can invoke the service;
grant
roles/iap.httpsResourceAccessortoiap.members. - Public access (
allUsersinvoker added or removed).
runway describe prints this order for a given configuration.
First deploy with bootstrap. Tags can only be bound to a service that
exists, but an organization policy such as constraints/run.allowedIngress
may refuse to create the service with its real settings until the tag is
there. With service.bootstrap, when the service does not exist yet runway
creates the real service with bootstrap.ingress (default internal), binds
service.tags to that service only, waits until they are effective, then
switches ingress to the configured value (a service-level change: no new
revision). Organization policy refusals are retried during that run while the
policy engine catches up. With bootstrap.image (for example
us-docker.pkg.dev/cloudrun/container/hello), a minimal placeholder is
created instead and the real image and settings are applied after the tags,
useful when the app cannot start before something else exists. Later deploys
skip all this. Use provider.tags instead when the tag
must be on the whole project.
OpenTelemetry Collector sidecar. service.otel_collector adds the
Google-built collector (otelcol-google) next to the app container: the app
container is named app and starts after the collector's health check
(port 13133) passes. Its configuration is passed through the environment
(--config=env:RUNWAY_OTELCOL_CONFIG). The default one receives OTLP on
localhost:4317/4318 and exports traces and logs with the googlecloud
exporter and metrics with googlemanagedprometheus (validated against
otelcol-google 0.160.0; the image's built-in config only prints to stdout).
runway also sets OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4318 on the
app unless you set it (with a warning if it points elsewhere), grants the
runtime service account roles/cloudtrace.agent,
roles/monitoring.metricWriter and roles/logging.logWriter, and requires
the Cloud Trace, Monitoring and Logging APIs. Without version, the newest
released collector is looked up in the registry at plan/deploy time (falling
back to 0.160.0) and pinned in the service, so a new release rolls out with
the next deploy; set version (or image) to pin it. With request-based
billing the collector only gets CPU while requests are served; exports
happen in batches every few seconds.
runway plan runs the read-only half of every step concurrently and lists
each as = done, + pending or ? unknown (for example, when the deployer
cannot read a policy).
Failure handling and recovery¶
There is no cross-service transaction; every step is safe to retry and
runway deploy is the recovery command.
- Per-step retries (
retryblock /--retries): a failed step is retried with exponential backoff. Retried: transient API errors, timeouts, unhealthy revisions, permission and "does not exist" errors (typical right after a service account is created or a role is granted, while IAM propagates). Not retried: invalid configuration, ownership conflicts, failed Docker builds (FAILURE), organization policy violations (constraints/..., reported with the constraint and how to inspect it) and Ctrl-C. Because each attempt re-reads the live state, a retry never repeats work that already succeeded. - Idempotent reads are retried by the SDK on transient errors (bounded).
- When a create/update/build submission/IAM write has an ambiguous outcome (timeout, connection reset, 5xx), runway re-reads the resource or looks for the tagged build before trying again. Stale-etag conflicts are retried after a fresh read (at most 3 attempts).
- Build failed: nothing was deployed; fix and re-run.
- Build succeeded, deploy failed: re-run; the image is reused, not rebuilt.
- Revision unhealthy: Cloud Run keeps traffic on the previous ready revision. runway reports the condition messages, revision log link and hints (port, secrets, image access). A re-run with an unchanged configuration rolls out a fresh revision (useful after fixing a grant).
- Interrupted or timed out while waiting: the rollout continues in Google
Cloud; check with
runway info, then re-run deploy to converge. - A step after the rollout failed (tags, IAP, access): reported as a partial failure with the live URL; re-running skips everything already done.