Run Your k3s Cluster From Anywhere With OpenClaw + ANTbase
·11 min read
☸

Run Your k3s Cluster From Anywhere With OpenClaw + ANTbase

OpenClaw is the agent. ANTbase routes its calls to the right model. Together they turn a phone, a laptop in a café, or a friend's borrowed terminal into a complete kubectl-equivalent. Here's the full setup: config, two CronJobs, and a Telegram trigger for the panic moments.

kubernetesk3sOpenClawdevopsintegrationagentsself-hosting

There's a moment every operator has had: away from your desk, phone buzzes, something on your cluster is on fire. You SSH from the phone, fight the keyboard, mistype kubectl twice, and curse the world. The problem isn't that kubectl is hard. The problem is that the loop — read alert, recall command, type, parse output, decide next action — needs full attention you don't have.

OpenClaw is the agent half of the answer: an OpenAI-compatible CLI that runs anywhere, plans multi-step actions, and uses tools you give it. ANTbase is the other half: a model router that picks the right LLM for each step. Stitched together, your phone can do a real incident-response runbook over Telegram.

OpenClaw agent terminal overlaid on a glowing Kubernetes wheel, with a Telegram message floating in the foreground showing an open incident
OpenClaw drives kubectl, ANTbase picks the model, Telegram is the remote.

Why route OpenClaw through ANTbase

Out of the box OpenClaw points at a single model. That works, until the workload mixes a 30-line incident triage (needs reasoning), a one-line restart (needs speed), and a long log diff (needs context). Three call shapes, three best-fit models, but one model id in the config.

ANTbase is OpenAI-compatible, sits in front of every provider, and exposes virtual models like ant:auto, ant:reasoning, ant:fast, and ant:free. You point OpenClaw at one base URL. ANT picks GPT-5 for the triage, Gemini Flash for the restart, Claude for the diff — without OpenClaw knowing or caring.

  • •One API key, every model. No juggling provider keys per agent run.
  • •Auto-failover: if Anthropic is down, ANT routes to OpenRouter, then to a free pool model, and only fails the request once nothing is left.
  • •Free-tier first: ant:auto will pick a free model when the task is small enough. Most "are there any failing pods?" queries cost zero.
  • •Telemetry: every OpenClaw call shows up on the ANT dashboard with model, provider, tokens, latency.

The setup (5 minutes)

Install OpenClaw, get an ANT API key (any of the providers above ant:free is enough), point OpenClaw at antbase.ai/v1. That's the whole integration.

bash
# Install OpenClaw
brew install openclaw   # or: pipx install openclaw

# Point at ANT — the only two env vars that matter
export OPENAI_BASE_URL="https://antbase.ai/v1"
export OPENAI_API_KEY="ak_live_..."

# Sanity check
openclaw --model ant:auto -- "what virtual models do you support?"

If you want OpenClaw to know about your cluster, give it the kubeconfig like you would any other tool. Run OpenClaw on a small jump box in the cluster, mount ~/.kube/config read-only, and don’t use admin-tier credentials — see the safety section below.

OpenClaw config: one file, every virtual model

json
// ~/.openclaw/config.json
{
  "providers": [
    {
      "name": "ant",
      "flavor": "openai-compatible",
      "baseUrl": "https://antbase.ai/v1",
      "apiKey": "ak_live_...",
      "models": [
        { "id": "ant:auto",      "name": "ANT Auto (default)" },
        { "id": "ant:reasoning", "name": "ANT Reasoning (incidents, diff analysis)" },
        { "id": "ant:fast",      "name": "ANT Fast (one-liners, lookups)" },
        { "id": "ant:free",      "name": "ANT Free (background polling)" }
      ]
    }
  ],
  "default_model": "ant:auto",
  "tools": ["kubectl", "shell", "files"],
  "max_steps": 8,
  "ask_before_apply": true
}

ask_before_apply: true is the seatbelt. OpenClaw can plan a delete-pod, but it asks before running it. Combined with a read-mostly kubeconfig, it’s safe to leave running unattended.

CronJob #1: hourly cluster health summary

This is the workhorse. Every hour, OpenClaw checks the cluster, summarises anything weird, and only pings you if it found something. Cheap to run because ant:auto picks a free model when the answer is “all good”.

yaml
apiVersion: batch/v1
kind: CronJob
metadata:
  name: openclaw-cluster-watch
  namespace: ops
spec:
  schedule: "0 * * * *"   # every hour, on the hour
  concurrencyPolicy: Forbid
  jobTemplate:
    spec:
      template:
        spec:
          serviceAccountName: openclaw-readonly
          restartPolicy: Never
          containers:
            - name: openclaw
              image: ghcr.io/openclaw/openclaw:latest
              command: ["openclaw"]
              args:
                - --model=ant:auto
                - --max-steps=6
                - --output=telegram
                - --quiet-when=clean
                - --
                - |
                  Walk the cluster. List nodes (skip cordoned). List pods
                  not in Running/Succeeded across all namespaces. List PVCs
                  pending more than 5 minutes. Summarise findings in 5
                  bullets max. If everything is healthy say "all clear".
              env:
                - name: OPENAI_BASE_URL
                  value: https://antbase.ai/v1
                - name: OPENAI_API_KEY
                  valueFrom: { secretKeyRef: { name: ant-key, key: api } }
                - name: TELEGRAM_BOT_TOKEN
                  valueFrom: { secretKeyRef: { name: telegram, key: token } }
                - name: TELEGRAM_CHAT_ID
                  valueFrom: { secretKeyRef: { name: telegram, key: chat_id } }

Notes on the YAML above:

  • •--quiet-when=clean — OpenClaw only posts to Telegram when there’s something off. No 24 daily "all clear" messages a day.
  • •serviceAccountName: openclaw-readonly — a namespace-scoped read-only SA. Even if the model goes rogue, blast radius is limited.
  • •ant:auto — for "are there any failing pods?", routing typically picks a free fast model. Hourly cost: pennies.

CronJob #2: nightly diff brief

Once a day, OpenClaw checks what changed in the cluster overnight — new deployments, image bumps, scale events. Routed to ant:reasoning because the model has to actually compare states.

yaml
apiVersion: batch/v1
kind: CronJob
metadata:
  name: openclaw-nightly-diff
  namespace: ops
spec:
  schedule: "0 7 * * *"   # 07:00 UTC every day
  jobTemplate:
    spec:
      template:
        spec:
          serviceAccountName: openclaw-readonly
          restartPolicy: Never
          containers:
            - name: openclaw
              image: ghcr.io/openclaw/openclaw:latest
              command: ["openclaw"]
              args:
                - --model=ant:reasoning
                - --max-steps=10
                - --output=telegram
                - --
                - |
                  Diff the cluster state vs. 24 hours ago. List deployments
                  with image changes, new namespaces, scale-up/down events,
                  and any CronJob that failed twice in a row. Group by
                  namespace. Note anything that looks like a config drift
                  from the helm chart in this repo.
              env:
                - name: OPENAI_BASE_URL
                  value: https://antbase.ai/v1
                - name: OPENAI_API_KEY
                  valueFrom: { secretKeyRef: { name: ant-key, key: api } }

The Telegram remote: incident → Claude Code prompt

If you’re running ANT for your own product traffic too, you already have the incident detector wired into Telegram. When a fingerprint trips the threshold, ANT posts a copy-paste prompt directly into your support chat. Long-press the prompt, paste into a Claude Code session on your phone, and you have the entire context — cf-ray, request id, error code, provider, the audit findings for that provider — in one tap.

Combine that with the hourly OpenClaw watch and the loop is closed: cluster tells Telegram → you tell Claude Code → Claude reads the repo + proposes a fix → you approve → OpenClaw applies it. Zero context switches, no SSH, no kubectl typing on a phone keyboard.

Safety: don’t hand an LLM root

  • •Use a namespace-scoped, read-only ServiceAccount for the watch CronJob. RBAC: get/list/watch on pods/services/deployments/cronjobs, nothing else.
  • •For the apply path, run a second OpenClaw instance with a separate write-capable SA. Keep it gated behind ask_before_apply: true so every kubectl apply / delete needs your confirmation.
  • •Never use a cluster-admin token. If you wouldn’t paste it into a script that runs unattended, don’t mount it.
  • •Set max_steps: 8 (or lower). Hard cap on how many tool calls one OpenClaw run can make. Stops accidental loops.
  • •Log every OpenClaw run. ANTbase already records every call — if you suspect something weird, /admin/traces shows the exact requests OpenClaw made.

Routing in action: one Saturday

Last Saturday morning at 06:47, the hourly watch ran. ant:auto picked Gemini 2.5 Flash (free pool), found one pod in CrashLoopBackOff, posted three lines to Telegram. At 06:48 I opened Claude Code on my phone, pasted the prompt, asked for a fix. Claude grepped the deployment, found a config-map env var that was renamed in last night’s merge but not bumped in the helm chart. It generated a one-line patch. I approved, OpenClaw applied. Total time from alert to fix: 6 minutes. Total kubectl commands typed: zero.

That run cost about $0.003. The hourly watches that happen to find no problem cost nothing — ant:auto stays in the free pool when the answer is trivially “nothing to report”.

Generate your own hero image

If you’re writing this up internally, here’s the exact prompt I’d use through ant:image-gen for the hero asset. Drop the output into apps/web/public/blog/openclaw-ant-setup-guide.png.

text
A dark cinematic illustration: a glowing translucent Kubernetes wheel
floating mid-air over a midnight-blue void. In front of it, an open
laptop showing a terminal with a green OpenClaw prompt; off to the
right, a phone displaying a Telegram chat with one orange alert
bubble. Light leaks of warm orange and electric cyan from the wheel.
Clean, modern, slightly futuristic. No text in image. 16:9, high
detail, soft volumetric lighting, no people.
bash
# Generate via the playground:
curl https://antbase.ai/v1/chat/completions \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"ant:image-gen","messages":[{"role":"user","content":"A dark cinematic illustration: a glowing translucent Kubernetes wheel ..."}]}'
# Save the returned URL to apps/web/public/blog/openclaw-ant-setup-guide.png

TL;DR

  • •Install OpenClaw, point it at antbase.ai/v1, pick ant:auto.
  • •Run a read-only watch CronJob hourly — free most of the time.
  • •Run a reasoning-heavy diff CronJob nightly.
  • •Use Telegram as the alert channel — it doubles as your Claude Code remote.
  • •Keep apply paths gated behind ask_before_apply and a write-capable SA.

Whole setup is two YAML files, one OpenClaw config, and one ANT API key. Your phone is now a working cluster console.

Try ANT routing today

Drop-in OpenAI-compatible API. Change your base URL and every request gets intelligent routing across 30+ providers.