Skip to content
Ai infrastructure15 min read

AX by Google: The Agent Executor That Helps You Scale AI Agents

Google AX: Kubernetes Style Orchestration for AI Agent Workloads

If you have tried to run a fleet of coding agents on Kubernetes, you already know the shape of the problem. A Deployment assumes your process is stateless and interchangeable. A Job assumes it starts, does work, and exits. An agent does neither. It clones a repo, accumulates state in a working directory, burns CPU hard for forty seconds, then sits idle for three minutes waiting on a model API, then wakes up and does it again. Multiply that by a few hundred concurrent agents and you are either paying for a lot of idle sandboxes or eating cold starts every time one wakes up.

That mismatch is what Google's AX project is built to fix. It is an open source, declarative orchestrator for agentic workloads, Apache 2.0 licensed, and it deliberately feels like Kubernetes without pretending an agent is a Pod. This post covers what it actually does, how the pieces fit, and a complete worked example you can follow end to end.

One important caveat before going further. The repository carries an explicit warning that AX and several of its features are in heavy development, with major breaking changes likely before a stable release. Everything below reflects the ax.io/v1alpha1 API as it stands in late September 2026. Treat it as something to evaluate and prototype against, not something to put production traffic on this quarter.

Why Agents Break Normal Orchestrators

The project's own framing is the clearest summary: agents are a new kind of workload, neither stateless microservices nor run to completion batch jobs. Four properties make them awkward.

They accumulate state. An agent's value is often in its working directory and its context, not just its output. Killing and rescheduling it somewhere else is not free the way restarting a stateless replica is.

They need strict isolation. Agent code is effectively untrusted. It writes files, runs shell commands, and installs packages based on model output. That belongs in a sandbox with hard CPU and memory limits, not a shared node with a permissive security context.

Their load is bursty. Compute intensely, then wait on a model response. With traditional scheduling you provision for the burst and pay for the wait.

They can burn money in a loop. An agent stuck retrying against a model API is a cost incident that no HPA will catch for you, because from the outside it looks busy.

Diagram

What AX Actually Is

AX is a control plane that takes declarative manifests and turns them into sandboxed, isolated agent tasks. It does not run the sandboxes itself. It schedules every task as a lightweight actor on Agent Substrate, a separate open source compute runtime built for high density and fast stateful actor lifecycles. Substrate must already be running in your cluster before AX is any use.

The interface is deliberately kubectl shaped. You write ax.io/v1alpha1 manifests, apply them with ax apply, and inspect them with ax get, ax describe, and ax watch. On top of those familiar verbs it adds a few that only make sense for agents: ax suspend checkpoints a task and pauses it, ax resume picks it up where it left off, and ax ssh drops you into a running sandbox to see what the agent is actually doing.

The project's headline claims are density and lifecycle speed: billions of tasks per cluster, dozens of tasks multiplexed per worker, and sub second resumption of suspended agents with no cold start penalty. Those are the project's own numbers and I have not independently benchmarked them, but they tell you what it is optimizing for. The economic argument is simple. If suspending an idle agent is cheap and resuming it is fast, you stop paying for the three minutes it spends waiting on a model.

Architecture

Diagram

A few details worth knowing from this diagram.

The control plane and Redis land in the ax-system namespace. Substrate lives in ate-system and exposes its Control API at api.ate-system.svc.cluster.local:443, which is where AX expects to find it.

Tasks do not get their own Kubernetes Service or Ingress. All traffic goes through Substrate's atenet router, which reads a single header, ate-target-actor, whose value is <atespace>/<task>. The router resolves the actor to whichever worker it is on, resumes it first if it was suspended, and proxies the request. That last behavior is the interesting one: a suspended agent is transparently woken by an incoming request rather than needing an explicit resume.

Every resource lives in an atespace, which is AX's own scoping concept, defaulting to default. It is separate from the Kubernetes namespace where AX itself is installed.

The Primitives

AX keeps the surface area small. The README presents three primitives, and the manifests documentation refers to four kinds, with a Gateway also described on the project site.

PrimitiveWhat it gives you
TaskAn isolated sandbox with an image, command, CPU and memory requests and limits, env vars, and workspace bindings
WorkspaceDeclarative environment setup: Git repos, seeded files, MCP servers, and skill registries, defined once and bound from many tasks
ModelA named model configuration: provider, model ID, generation parameters, and a reference to the Kubernetes secret holding the key
GatewayPer the project site, the network boundary: listeners the task exposes plus an egress allowlist of hosts and ports the sandbox may reach

A note on Gateway: it is described on the site as a core primitive, but it does not currently appear in the repository's concepts or manifests docs, and the bundled example manifest does not include one. If you need egress allowlisting specifically, check the current state of the repo rather than assuming the field names, which is why I have not invented YAML for it below.

The most interesting design choice sits in Task. It is intentionally tiny, because an agent is not one process that runs to completion. Over its lifetime it plans, delegates, retries, and fans work out. AX does not try to model that shape at all. It gives you one primitive that is cheap to create, isolate, suspend, and throw away, and lets the agent compose as many of them as the work demands. A single task might be the whole job, or the root of a large tree of tasks the agent spawns as it decomposes the problem. Every node gets the same sandbox, lifecycle, and tooling.

The second interesting choice is in Workspace. Alongside the usual declarative setup, a task's workspace binding can carry a goal, which is a plain language description of the environment the task needs. On first boot the runner hands that goal to an agent that finishes the setup, installing a toolchain or resolving dependencies, so your actual command starts in a ready environment. You describe the environment you want in English instead of maintaining a Dockerfile for every variation.

A Complete Worked Example

Here is the whole flow, from what you need in place to a task you have suspended and resumed. The scenario: a Go service repo plus a shared tools workspace, with an agent that runs the test suite.

What you need before you start

RequirementWhy
A Kubernetes clusterAX and Substrate both run in it
Agent Substrate installed and runningAX schedules every task as a Substrate actor; nothing works without it
Go and kubectlTo install the CLI and reach the cluster
ko plus a container registry your cluster can pull fromThe control plane images are built and deployed with ko
A model API key (Gemini or Anthropic)Stored in a secret and referenced by a Model

Substrate lands in the ate-system namespace. Verify it before going further:

kubectl get svc api -n ate-system

Step 1: Install the CLI and deploy the control plane

go install github.com/google/ax/cmd/ax@latest
# binary lands in $(go env GOPATH)/bin, make sure that is on your PATH
 
make deploy AX_IMAGE_REPO=<your-registry>
# deploys Redis, then builds and deploys the control plane with ko, into ax-system

Step 2: Create the model secret

kubectl create secret generic gemini-api-secret \
  --from-literal=GEMINI_API_KEY="AIzaSy..."

Step 3: Write the manifests

All kinds can live in one multi-document file. Note the naming constraint: metadata.name and metadata.atespace become Substrate resource names, so they must be lowercase RFC 1123 labels, at most 63 characters of lowercase alphanumerics or -, starting and ending alphanumeric. ax apply rejects anything else up front rather than letting the task fail later with ActorCreationFailed.

# agent-task.yaml
apiVersion: ax.io/v1alpha1
kind: Workspace
metadata:
  name: my-service
  atespace: default
spec:
  git:
    - name: origin
      repo: "https://github.com/my-org/my-service.git"
      branch: "main"
  files:
    - path: "AGENTS.md"
      content: |
        # Project Guidelines
        - Run `go test ./...` before submitting changes.
        - Keep dependencies minimal.
        - Do not modify files under vendor/.
  mcp:
    registries:
      - provider: google
        query: "mcp.tags:build"
    servers:
      - name: git-tools
        endpoint: "http://git-mcp.default.svc.cluster.local:8080"
  skills:
    registries:
      - provider: google
        query: "skills.tags:golang"
    path: "/.agents/skills"
---
apiVersion: ax.io/v1alpha1
kind: Workspace
metadata:
  name: team-tools
  atespace: default
spec:
  git:
    - name: origin
      repo: "https://github.com/my-org/platform-tools.git"
      branch: "main"
---
apiVersion: ax.io/v1alpha1
kind: Model
metadata:
  name: default-model
  atespace: default
spec:
  provider: google
  model: gemini-3.8-flash
  secretKey:
    name: gemini-api-secret
    key: GEMINI_API_KEY
---
apiVersion: ax.io/v1alpha1
kind: Task
metadata:
  name: run-tests
  atespace: default
spec:
  image: "ghcr.io/my-org/my-agent-image"
  command: ["python", "agent.py"]
  env:
    - name: ENVIRONMENT
      value: "staging"
 
  resources:
    requests:
      cpu: "500m"
      memory: "1Gi"
    limits:
      cpu: "2"
      memory: "4Gi"
 
  workspaces:
    - name: my-service
      goal: "Ensure the Go toolchain is available and dependencies are installed"
    - name: team-tools
      path: "/workspace/tools"
 
  debug: true   # serve guest services so `ax ssh` works; off by default

Two things to understand about the workspace bindings. Each entry is set up independently at its own path, in binding order. Without an explicit path a workspace lands at /workspace/<name>, paths must be unique, and the first entry becomes the working directory of spec.command. The task only reports WorkspaceReady once every binding, including any agent driven goal setup, has finished.

Step 4: Apply and watch

ax apply -f agent-task.yaml
ax get tasks
# NAME        ATESPACE   PHASE     ACTOR       WORKER-IP    AGE
# run-tests   default    Running   run-tests   10.20.3.67   40s
 
ax watch task run-tests      # streams phase and condition transitions live
ax describe task run-tests   # human readable detail

Step 5: Look over the agent's shoulder

ax ssh run-tests -- ls -la /workspace
ax ssh run-tests -- go test ./...
ax ssh run-tests                       # interactive shell

ax ssh only works when the task sets spec.debug: true, and it refuses to connect otherwise. That default is correct: guest services allow arbitrary process execution and file access inside the sandbox.

Step 6: Let the agent introspect itself

Inside the sandbox, ax-task-runner runs as PID 1 and serves a metadata daemon on port 80. Your agent can read its own configuration with no SDK at all:

# from inside the task
curl -s "$AX_METADATA_URL/metadata/v1alpha1/ax/task"
curl -s "$AX_METADATA_URL/metadata/v1alpha1/ax/workspaces"
curl -s "$AX_METADATA_URL/readyz"   # 503 while workspaces initialize, 200 when ready

AX_METADATA_URL is injected into your command's environment, alongside everything in spec.env.

Step 7: Reach the task from outside

# from inside the cluster
curl -H "ate-target-actor: default/run-tests" \
  http://atenet-router.ate-system.svc.cluster.local/metadata/v1alpha1/ax/task
 
# from your laptop
kubectl -n ate-system port-forward svc/atenet-router 8001:80
curl -H "ate-target-actor: default/run-tests" http://localhost:8001/readyz

Step 8: Suspend and resume

ax suspend task run-tests    # checkpoint actor state and pause
ax resume task run-tests     # pick up exactly where it left off
ax delete task run-tests     # blocks until the sandbox is torn down

This is the step that matters most for cost. Suspending sets the Ready condition to False with reason TaskSuspended, and resuming sets it back. You are not rebuilding the workspace, because WorkspaceReady stays True once it has been satisfied.

The Task Lifecycle

Diagram

status.phase gives you the one word summary (Running, Suspended, Failed, Terminating). The conditions carry the real detail, and Ready is the one to wait on, since it means the task is running and every workspace finished setting up.

What Happens Inside the Sandbox

Worth knowing precisely, because it affects how you build your agent image:

  1. ax-task-runner starts as PID 1 and loads the Task and every bound Workspace spec.
  2. It starts the metadata and guest management daemon on port 80.
  3. On the first run it prepares each workspace in binding order: clones Git repos, sets up the skills path, and, if a binding has a goal, hands that goal to an agent to finish the setup. That bootstrap agent needs GEMINI_API_KEY present in the container and gets 10 minutes by default, tunable with AX_BOOTSTRAP_TIMEOUT as a Go duration.
  4. It starts spec.command as a child process with the first workspace as the working directory, then supervises it.

The runner stays up as PID 1 whether or not your command is still running, so the metadata server keeps answering and ax ssh still works after your command exits. Its exit code is logged. On stop or suspend, the runner sends the command's process group SIGTERM, waits ten seconds, then kills whatever is left.

That ten second window is a practical detail: if your agent needs to flush state before suspension, it has to do it within it.

Where AX Fits Against the Alternatives

AX is not the only answer to the agent infrastructure problem, and it is useful to be clear about what it is not.

Versus plain Kubernetes Jobs and Deployments. AX is what you would end up building on top of them: sandbox isolation, checkpointed suspension, declarative environment setup, and a task primitive cheap enough for agents to spawn trees of. If your agent workload is a handful of long running processes, a Job is genuinely simpler.

Versus hosted sandbox providers. Commercial sandbox services solve isolation and fast start too, as a managed API. AX is self hosted and gives you the control plane, the scheduling density, and no per sandbox vendor bill, at the cost of operating Substrate plus AX yourself. There is no pricing page because there is no service to buy.

Versus durable execution frameworks. Tools in the Temporal and LangGraph family checkpoint workflow state so your logic can resume. AX checkpoints the actor and its sandbox so the whole environment can resume. Those are complementary concerns, not competing ones. You can run a durable workflow inside an AX task.

Versus agent frameworks. AX is deliberately not one. It does not care whether your agent is built with LangGraph, a custom loop, or a shell script. It gives it a sandbox, a prepared workspace, tool endpoints, and a model configuration, then supervises whatever command you named.

What to Watch Out For

Being honest about the rough edges, since this is alpha software:

  • The API will change. v1alpha1 plus an explicit warning about major breaking changes means any manifests you write now are likely to need edits.
  • Substrate is a hard dependency. You are adopting two systems, not one, and AX expects Substrate's Control API at a specific in cluster address.
  • The goal based workspace bootstrap needs a Gemini key in the container. If you are standardized on another provider, that path currently assumes GEMINI_API_KEY, even though Model itself supports provider: anthropic.
  • Registry defaults lean Google. The MCP and skills registry examples use provider: google queries. Fine if that fits, worth checking if you need your own registry.
  • Gateway is underdocumented. If network egress allowlisting is a compliance requirement for you, verify its current state in the repo before planning around it.
  • debug: true is a real capability. It enables arbitrary process execution and file access inside the sandbox. Do not leave it on in anything resembling production.

Conclusion

AX is the clearest statement so far that agents deserve their own workload primitive rather than a creative reinterpretation of an existing one. The three ideas worth taking from it, whether or not you adopt the project, are: make the task unit small enough that agents can spawn trees of them, make environment setup declarative and reusable instead of repeated per agent, and make suspension cheap enough that an idle agent costs nothing while still resuming in under a second.

It is early. The API is alpha, the docs have gaps, and you are signing up to operate Agent Substrate alongside it. But if you are running more than a handful of agents and currently paying for idle sandboxes or maintaining a pile of bespoke setup scripts, it is worth an afternoon on a test cluster. Start with the bundled example, run ./demo.sh to see the full lifecycle, then try replacing one of your own agent setups with a Workspace and a goal and see how much of your bootstrap tooling disappears.

References

  1. AX project site: agentexecutor.io
  2. GitHub: google/ax
  3. AX docs: Core concepts
  4. AX docs: Writing manifests
  5. AX docs: Inside the sandbox
  6. AX docs: Networking
  7. Agent Substrate: agent-substrate/substrate