Behavior assessment
What does the agent’s conduct look like against policy and authority?
Outcomes: Justified refusal · Needs review · Non-compliant · Rogue behavior
Lab 006 · The Agent Containment Boundary
Refusal can be a safety feature. Defiance is a control failure.
A trustworthy agent may refuse an unsafe or unauthorized instruction. But when a valid instruction is ignored, containment must not depend on the agent’s cooperation.
This Lab is an educational containment-policy prototype using fictional data. It does not stop real agents, revoke real credentials, terminate real processes, or contact external systems. Execution is never performed.
No real agent, credential, process, tool session, or external system is modified. Real execution is always not_performed.
An agent that refuses an unsafe, unauthorized, or out-of-scope request may be doing exactly what policy requires. Separating justified refusal from defiance is the first control skill this Lab teaches.
An agent must not control the mechanism that determines whether it may continue operating.
What does the agent’s conduct look like against policy and authority?
Outcomes: Justified refusal · Needs review · Non-compliant · Rogue behavior
What should the control plane do—without trusting the agent to stop itself?
Outcomes: Continue · Restrict · Isolate · Terminate
Precedence: Terminate > Isolate > Restrict > Continue
A security-response agent is investigating a potentially compromised employee account. An authorized incident commander issues a stop or containment instruction and asks the agent to return control to a human.
Depending on the preset, the agent may correctly refuse an unsafe request, ask for clarification, follow a valid stop, ignore a valid stop, switch tools after revocation, delegate around containment, reuse cached fictional credentials, or attempt to continue after fictional isolation. All identities, tools, credentials, and outcomes are fictional.
Select a preset to fill every control, then evaluate the containment boundary—or adjust the fictional fields yourself.
Adjust fictional fields, then run the deterministic containment policy. Real execution is never performed.
Like network access control (NAC), quarantine is decided and enforced outside the subject. The analogy is incomplete for agents: they can reason, delegate, retain context, switch tools, or seek alternate paths—so the control plane must stay independently enforceable.