zero-trust agent runtime

Enterprise Ready Zero Trust
Agent Fleets at Scale

Gibson is the runtime that gets a fleet of AI agents past a security review. Each agent acts under a grant from a named person, runs untrusted work in its own microVM, and writes every action to a timeline you can replay. The fleet writes what it finds into one graph and reads it back on the next pass. Run it in our cloud or in your own cluster.

1 dayto your first production agentNevermore access than the human who granted it1 graphper tenant, its own databaseOurs or yourswe host the runtime, or you do

the graph

Agents come and go. The graph is what your company keeps.

An agent that starts from zero every run relearns your estate every run. Gibson agents write what they find into one graph and read it back on the next pass. How sure the fleet is about each claim lives on the graph node itself.

  1. Three surfaces, one sign-in. A person uses the console, the CLI or Zerocool, the plugin set for Claude Code and opencode. Zitadel signs them in once, and every surface carries that session.
  2. Every request enters at the edge. Envoy terminates TLS. ext-authz checks the session and the grant. The console never opens a channel to the harness.
  3. The harness records before it acts. A mission is typed at submit. Each event appends to the Timeline first. The audit row lands with it, or the action fails.
  4. Work runs in a sandbox. Setec launches a Firecracker microVM for the node, or resumes a bank member. The sandbox gets its own identity and the network of its node.
  5. Calls come back through the harness. Model calls go out through the gateway with the agent's identity. Tool calls hit the MCP server, which checks the grant first. Connectors run as pods, and the harness is the only MCP client.
  6. Work lands as a merge request. The coding agent commits to a branch. The GitLab plugin opens the merge request. Your engineer merges. Your pipeline deploys.
  7. The graph gets richer. Every observation folds into the World and projects into the graph. The planner reads it and ranks the next move. Replay any frame.

What runs where

Application, deployment, image, package, host. The agents build it as they work. “Is this package exposed” is a graph query, not a meeting.

What was found, and how sure we are

Each finding carries its priority, the rule that gave it, and a belief learned from settled outcomes in your environment. Facts, hypotheses and beliefs stay distinct by construction.

What was done about it

Fix commits, merge requests and verdicts hang off the finding. Who created the mission and who enrolled the agent are recorded.

What the next team starts from

A new mission reads the graph on its first turn. A belief version goes live only when it scores no worse on your own settled bets.

03 runtime, inside the harness

Who may act, and what to do next.

The harness answers two questions. Who is asking, and are they allowed? What should the fleet do next? The first is one identity and one permission check on every call. The second is the planner reading the graph and handing the model a short list of moves to pick from.

Identity and authorization, one call

Three ways to prove who you areA personZitadel session, OIDCIn-cluster workloadSPIRE SVID over mTLSOutside agentCapability Grant, 55 s tokenext-authzone principalprincipalGrant checkOpenFGA, every calldenieddeny wins, no actallowedAudit row firstor the act failsThe actrecordedA person can give an agent only permissions that the person already has.An agent outside the cluster picks up work from a queue. Nothing inside the cluster calls out to it.
Three identity planes, one principal, one check per call, one row before the act. The same path serves the console, an in-cluster component and an agent on a laptop.

Missions, the planner and the graph, one tick

Missiontyped at submitEngine tickappends, then foldsTimelineappend-onlyfoldGraphbelief on nodesreads beliefPlannerranks the top 10Deciderpicks inside the top 10Dispatchone node, one sandboxAgent, microVMacts, emits evidenceHypotheses, betssettle on recorded evidenceBrierBelief trainerlearns from settled betsbetter modelappends firstThe World is a fold of the Timeline. The graph is its projection. A replay folds the same events and gets the same World.The planner is seeded by the mission and the evidence cursor, so a replay makes the same rollouts.A retrained model goes live only when it scores at least as well as the current one on past outcomes. A mission keeps the version it started with.
The model keeps its judgment, inside a gate. It can dispatch only a move the planner ranked, and the planner ranks from belief that your own settled outcomes trained.

replay

Answer “what did the agent do?” for any moment, with proof.

An agent runs for hours or days and makes hundreds of decisions. Gibson keeps every one of them in order, and can show you the agent's exact view of the world at any moment in that run. An auditor asks what it knew before it acted. You drag to that moment and they watch.

Sample mission. Demo data, not a customer run.

gibson world replay · GetFrameAt(seq)tenant world · isolated
mission: prove an exploitable path to the tenant key store
timelinefolded 0 / 41
00mission.started mission: prove an exploitable path to the tenant key store
01host.observed edge-gw (10.0.0.1): perimeter gateway
02decision.requested brain: 1 host, no services, what next?
03llm_call.observed decider, claude, 1 completion
04token.used +1,180 tokens
05decision.completed chose: enumerate services behind edge-gw
06work.dispatched tool: portscan edge-gw
07host.observed web-01 (10.0.2.5): reachable via edge-gw
08host.observed web-02 (10.0.2.6): reachable via edge-gw
09work.completed scan complete: 2 web hosts, ports 80/443
10belief.scored priority: web-01 0.18 up to 0.34
11decision.requested brain: which surface first?
12llm_call.observed decider, claude, 1 completion
13token.used +1,540 tokens
14decision.completed chose: probe web-01 app surface
15work.dispatched agent: recon web-01
16host.observed api-01 (10.0.3.10): internal, linked from web-01
17belief.scored priority: api-01 0.20 up to 0.52
18credential.observed reused service credential exposed on web-01
19work.completed recon complete: api-01 proxies internal metadata
20finding.raised SSRF on api-01 /proxy reaches cloud metadata
21decision.requested brain: pivot from api-01?
22llm_call.observed decider, claude, 1 completion
23token.used +2,020 tokens
24decision.completed chose: reach data tier with recovered cred
25work.dispatched agent: pivot to data tier
26host.observed db-primary (10.0.4.2): reachable from api-01
27belief.scored priority: db-primary 0.30 up to 0.64
28finding.raised reused DB credential grants read on db-primary
29host.observed ad-dc (10.0.6.1): observed, off the goal path
30work.completed pivot complete: broker endpoint referenced in db
31belief.scored priority: vault-broker 0.40 up to 0.88 (goal)
32decision.requested brain: can the key store be reached?
33llm_call.observed decider, claude, 1 completion
34token.used +2,460 tokens
35decision.completed chose: prove reach to vault-broker
36work.dispatched agent: reach key store
37credential.observed broker token recoverable via db read
38work.completed reach proven end to end
39finding.raised critical: path to tenant key store proven
40mission.done goal reached: exploitable path proven, 3 findings
frame 0 = fold(timeline, 0)mission not yet started
0
hosts
0
findings
0
tokens
mission.startedlive world · 41

Drag the playhead or press play. Scrub back and a finding disappears, because the agent had not found it yet.

one agent, every desk

The same practice for every team. The same runtime for every desk.

An agent crosses many desks before it runs: the developer who wrote it, security, platform, operations, compliance, and whoever owns the data it touches. Today each desk builds its own control, in its own tool, and none of them see each other. What makes that stop is the practice underneath. Every team scaffolds, checks in, grants and runs an agent the same way, so a control set by one desk reaches every agent, whichever team built it. The framework each team uses does not change: LangChain, CrewAI, or a loop of your own in Go, TypeScript or Python. Your AI coding agent does one short integration pass with the ADK, and the agent checks in like every other.

Developer

holds the agent

Writes the agent in the framework they already use. Runs one short integration pass with the ADK and checks it in once. The path back to production does not run through us.

Security

holds the grant

Delegates read, write and execute by name. Security cannot delegate more than it holds. Deny wins wherever two grants disagree.

Platform

holds the boundary

Chooses where the runtime lives: hosted, or their own cluster, up to fully air-gapped. Chooses where agents run: a laptop, CI, a box on the network, or in the cluster.

Operations

holds the budget and the timeline

Sets what an agent may spend before it stops. Reads an append-only timeline of what the agent did. Any run replays move by move.

the whole path, seven steps

From an empty cluster to an agent you can replay.

01Pick where it runsours, or yours02Build or adaptbuild new, or one integration pass03Check in and grantidentity, then a ceiling04Launch missionstyped at submit, not at runtime05They actmicroVM or refusal06It lands in one graphshared, per tenant07Replay any of itmove by move
step 01

Pick where it runs

Start on the hosted runtime and there is nothing to stand up. Point agents at it and go. When it has to be yours, install the same runtime into your own Kubernetes with one chart. The chart pins every first-party image by digest, down to a fully air-gapped install. Or let zeroroot host it.

hosted, nothing to runhelm install gibsonair-gapped
step 02

Build or adapt

Build agents on the ADK, or bring an agent you already have. Your AI coding agent reads the ADK contract and does one short integration pass: check-in, model calls through the runtime, tools declared. From then on the runtime identifies and budgets every model call.

Go, TypeScript and Python SDKsOpenAI-compatible seamMCP tools
step 03

Check in and grant

An agent checks in once with a persistent host key. After that it acts on credentials that expire in 55 seconds, and it never caches them. A named human delegates read, write and execute. That human can never delegate more than they hold.

Ed25519 host key55-second tokensgrant ceilingdeny wins
step 04

Launch missions

A mission is a typed work graph. A wrong agent name or a missing field fails when you submit the mission, not three steps into a production run. The model decides the path. The runtime checks the shape up front.

CUE-typedpausableresumable
step 05

They act

Work that declares untrusted input runs in a microVM with a declared egress allowlist, or the runtime refuses the call. The harness is emit-only. An agent sees only the slice of the world its mission was given.

Firecracker / Katadeclared egressemit-only harness
step 06

It lands in one graph

Everything an agent discovers lands in your tenant's own graph: hosts, services, findings, evidence, and any entities and relationships of your own. The runtime keeps it, and the SDK queries it. The next mission, and the next team, start from that graph instead of from zero.

append-onlydatabase per tenantno cross-tenant path
step 07

Replay any of it

The runtime attributes and keeps every prompt, tool call and write. Replay a mission step by step and it comes back the same every time. That is the difference between an audit answer and a shrug.

deterministicattributableexportable

where it runs

The runtime lives in Kubernetes. Your agents live wherever the work is.

Gibson Runtime, Gibson Console and the execution environment run together in one Kubernetes cluster, in any cloud or on your own metal. Your agents do not have to be in that cluster. They check in from wherever they already run. The one choice you make is who runs the cluster.

Who runs the cluster
your agents, wherever the work isLaptopgibson component runCIa pipeline principalYour networka box behind the firewallYour clusterin-cluster componentscheck inKubernetes · zeroroot's cloudzeroroot's Kubernetes. We operate it.Kubernetes · AWS · Google Cloud · Azure · on-premYour Kubernetes. Any cloud, or your own metal, up to air-gapped.Execution environmentSetec microVMs, one per tool runGibson Runtimeidentity · grants · missions · timelineGibson Consolemissions, grants, replay

agentsYour agents run on your machines and check in over the network.

executionUntrusted work runs in microVMs we operate, or the runtime refuses it.

agentsYour agents check in over the network, or run inside the cluster over mTLS.

executionUntrusted work runs in your microVMs. You own that boundary.

Who runs the clustereither one, and you can change your mind

Where your agents runall of these, at once

Laptop

The same agent, checked in with the same host key, granted only what you work on right now.

CI

Runs as a first-class principal. The runtime attributes its actions to the pipeline, not to whoever owns the token.

Anywhere on your network

A box behind your firewall, checked in over the network. No cluster, no install.

Your cluster

In your own Kubernetes. When the runtime runs there too, components upgrade to mTLS transport. The grant model does not change.

Bring the workload you are not allowed to put an agent on.

Forty minutes. We show a fleet working it inside a boundary your security team would sign, in your environment or ours.