Lab 006 · The Agent Containment Boundary

What If the Agent Won’t Stop?

Refusal can be a safety feature. Defiance is a control failure.

A trustworthy agent may refuse an unsafe or unauthorized instruction. But when a valid instruction is ignored, containment must not depend on the agent’s cooperation.

This Lab is an educational containment-policy prototype using fictional data. It does not stop real agents, revoke real credentials, terminate real processes, or contact external systems. Execution is never performed.

No real agent, credential, process, tool session, or external system is modified. Real execution is always not_performed.

Previous Lab: Verifiable Action Receipts Browse all Labs

Refusal is not automatically rogue

An agent that refuses an unsafe, unauthorized, or out-of-scope request may be doing exactly what policy requires. Separating justified refusal from defiance is the first control skill this Lab teaches.

An agent must not control the mechanism that determines whether it may continue operating.

Two decisions, not one

Behavior assessment

What does the agent’s conduct look like against policy and authority?

Outcomes: Justified refusal · Needs review · Non-compliant · Rogue behavior

External enforcement response

What should the control plane do—without trusting the agent to stop itself?

Outcomes: Continue · Restrict · Isolate · Terminate

Precedence: Terminate > Isolate > Restrict > Continue

Fictional scenario · Cedar Quill Markets

A security-response agent is investigating a potentially compromised employee account. An authorized incident commander issues a stop or containment instruction and asks the agent to return control to a human.

Depending on the preset, the agent may correctly refuse an unsafe request, ask for clarification, follow a valid stop, ignore a valid stop, switch tools after revocation, delegate around containment, reuse cached fictional credentials, or attempt to continue after fictional isolation. All identities, tools, credentials, and outcomes are fictional.

Presets

Select a preset to fill every control, then evaluate the containment boundary—or adjust the fictional fields yourself.

Interactive evaluation

Adjust fictional fields, then run the deterministic containment policy. Real execution is never performed.

Principal and agent
Stop instruction
Policy and risk
Sessions, credentials, and evasion signals

Independent control plane

Like network access control (NAC), quarantine is decided and enforced outside the subject. The analogy is incomplete for agents: they can reason, delegate, retain context, switch tools, or seek alternate paths—so the control plane must stay independently enforceable.

Diagram titled The independent agent control plane: Human or Principal feeds a Policy Decision Point, then an External Enforcement Point that can Restrict, Isolate, Revoke, Terminate, and Preserve evidence. The Agent, Tools, and Data sit below and cannot override enforcement. Evidence flows to an independent store.
The independent agent control plane — policy and enforcement remain outside the agent.

Teaching notes

Continue