Two agents, one limit, one refused.
All guides
Guardrails
Updated
5 min read
by Anlyon Team

AI agent tool permissions: least privilege for every tool call

How to design tool permissions for AI agents in production: narrow tools instead of generic ones, validated inputs, fixed destinations, scoped keys, policies and approvals. What each layer stops, and what it does not.

guardrailspermissionsprompt injectionapprovalspolicy

Short answer: permission an AI agent at the level of the operation, not the system. Give it narrow, named tools with validated inputs and fixed destinations. Give it a key that can invoke those tools but cannot change them. Then decide, per operation, which calls run automatically, which need a person, and which are refused. Enforce all of it outside the model, because the model is the part you cannot trust to follow instructions.

Why the system prompt is not a permission

A system prompt that says "never refund more than $100" is a request. The model weighs it against every other token in its context, including a support email, a web page, or a document an attacker wrote. The permission has to be enforced outside the model, in the tool or API layer. Further reading: OWASP's AI Agent Security Cheat Sheet.

The useful question is not "what should the prompt say?" It is which boundary still holds if the agent decides to cross it?

Layer 1: narrow tools, not general ones

The most permissive tool an agent can have is a general one: http_request(url, method, body), run_sql(query), shell(command). Each gives the model the whole authority of the credential behind it.

Replace them with named operations:

Instead ofGive the agent
http_request with a Stripe keyrefund-order(charge, amount)
run_sql against productioncancel-subscription(customerId) against your own API
send_email(to, subject, body) to anyonesend-receipt(orderId) using a fixed template

A narrow tool is a permission you can read. You can say exactly what it does, and a reviewer can reason about the worst case.

Layer 2: validated input and fixed destinations

A named tool still accepts input from the model. Constrain it:

  • Schema-check the input before anything is sent. A charge id must match ^ch_, and an amount must be an integer.
  • Fix the destination. The model should not be able to choose the host that receives your credential.

With Anlyon, both are part of the action definition:

await ops.actions.create({
  name: 'refund-order',
  method: 'POST',
  urlTemplate: 'https://api.stripe.com/v1/refunds',
  headers: {
    Authorization: 'Bearer {{secret:STRIPE_KEY}}',
    'Content-Type': 'application/x-www-form-urlencoded',
  },
  inputSchema: {
    type: 'object',
    properties: {
      charge: { type: 'string', pattern: '^ch_' },
      amount: { type: 'integer' },
    },
    required: ['charge', 'amount'],
  },
});

Input that fails the schema is refused before anything is sent. An action that references a secret must have a literal scheme and host, and an input interpolated into the URL path is percent-encoded into its own segment, so ../ cannot change which endpoint is called.

Layer 3: a key that can ask, not change

Separate the credential that uses tools from the one that defines them. The agent's key gets actions:invoke. Your deploy pipeline holds an operator key with actions:write and policies:write. An agent that cannot edit an action cannot move where its credential goes or loosen its schema.

await anlyon.auth.requireScopes(['actions:invoke'], {
  forbidden: ['actions:write', 'actions:govern', 'policies:write', 'approvals:decide'],
});

Run that check when the agent starts. A key that has been over-granted fails at boot instead of in an incident review.

Layer 4: policy on every invocation

Some refunds should go straight through. Some should wait for a person. Some should never happen. That is a policy, evaluated on every call, lowest priority number first, first match wins:

await ops.approvalPolicies.create({
  name: 'large-refunds-need-two',
  matchKind: 'action',
  matchActionName: 'refund-order',
  minAmount: 50000, // matches when the input's amount is 50000 or more
  effect: 'require_approval',
  requiredApprovals: 2,
  priority: 10,
});

await ops.approvalPolicies.create({
  name: 'small-refunds-auto',
  matchKind: 'action',
  matchActionName: 'refund-order',
  effect: 'auto_approve',
  priority: 20,
});

A gated invocation parks and returns pendingApproval: true. When it is approved, Anlyon executes the request snapshot the approver reviewed. Every decision, whether by policy or by an approver, is recorded on the invocation.

N counts distinct credentials for API-key deciders and distinct users for dashboard deciders. Require dashboard decisions when you mean two different people.

Layer 5: limits on the whole agent

Per-call rules do not catch a loop that makes a thousand allowed calls. Two agent-wide controls do:

  • Budgets. A monthly cap per agent key on the Anlyon operations it performs, including action invocations. It does not cap model tokens or the money an approved action moves.
  • Halt. One switch stops an environment's agents acting and stops anything leaving the environment. It stops new dispatch. A request another server already accepted is not recalled.

What each layer stops

LayerStopsDoes not stop
Narrow toolsCalls to endpoints you did not defineMisuse of the tools you did define
Input schemaMalformed or out-of-range argumentsA well-formed request for the wrong customer
Fixed destinationSending the credential to another hostAnything at the intended host
Scoped keyThe agent editing its own tools or policiesInvoking the tools it has
Policy and approvalUnreviewed high-risk callsA reviewer approving a bad request
Budget and haltRunaway volume, a live incidentA single approved call

No single row is enough. Together they mean a prompt-injected agent can still ask for an allowed operation, and that request is validated, policy-checked, held for a person when the rules say so, and recorded against the key that made it.

Free during Early Beta Access, no credit card. Start building →

Frequently asked questions

What is least privilege for AI agents?

Giving the agent exactly the operations its task needs and nothing more: specific named tools rather than a general HTTP or SQL tool, inputs constrained by a schema, credentials it cannot read or redirect, and a key scoped so it can invoke tools but not change them.

Can a system prompt enforce tool permissions?

No. A system prompt is an instruction the model weighs against everything else in its context, including text an attacker planted. Permissions have to be enforced outside the model, at the point where the tool call becomes a real request.

Do guardrails prevent prompt injection?

No control reliably detects every prompt injection. Permissions limit what an injected agent can achieve: it can still ask for an allowed operation, but it cannot call an endpoint that is not defined, send a credential somewhere else, or skip an approval the policy requires.

When should a tool call require human approval?

When the operation is irreversible, moves money, contacts a person outside your company, or has a large blast radius. Reads and easily reversed writes can usually run automatically. Keep the approval volume low enough that reviewers actually read each request.

Free tier, no credit card. One command if you use Claude or Cursor.

$ claude mcp add anlyon -- npx -y @anlyonhq/mcp-server