Two agents, one limit, one refused.
All posts
6 min read
by

How to add human-in-the-loop approval to AI agent tool calls

Pause an AI agent's tool call until a human approves, with the agent unable to approve itself and every decision recorded. Two ways to build it, in TypeScript.

approvalshuman-in-the-loopagentstutorial

Human-in-the-loop (HITL) approval means a specific tool call does not run until a person says yes. Not "logs a warning". Not "flags it for review afterwards". The call waits, a human decides, and a no means the side effect never happens.

There are two ways to build it, and they guarantee different things. Most teams pick one without noticing there was a choice. This post covers both, and when to use which.

Which calls need a human?

Not many. An agent that asks before reading a file is useless, and a human who approves forty requests an hour stops reading them. The list worth gating is short and every team can write it in a minute:

  • moving money: refunds, payouts, credits over a threshold
  • emailing or messaging customers
  • deleting anything: rows, volumes, buckets, accounts
  • deploying to production, rotating keys, changing permissions

If a call is reversible and cheap, do not gate it. If it is irreversible or expensive, gate it and make the approval fast to give.

Two shapes

Local gate     agent -> your function -> [wait for human] -> your credential -> API
Hosted action  agent -> named action -> Anlyon -> [wait for human] -> vaulted credential -> API

The difference is who holds the call while the human decides.

Shape 1: gate your own function

Wrap the function the agent already calls. The wrapper creates an approval, waits for a decision, and calls through only on approve.

import { Client, ApprovalDeniedError, ApprovalPendingError } from '@anlyonhq/sdk';

const client = new Client({ apiKey: process.env.ANLYON_API_KEY! });

async function refund(orderId: string, amount: number) {
  // your existing refund logic, using your Stripe key
}

const guardedRefund = client.approvals.gate(
  { title: 'Refund a customer?', timeoutMs: 10 * 60_000 },
  refund,
);

try {
  await guardedRefund('ord_42', 129.99); // blocks until a human decides
} catch (err) {
  if (err instanceof ApprovalDeniedError) console.log('Denied:', err.approval.decisionNote);
  else if (err instanceof ApprovalPendingError) console.log('Still open:', err.approval.id);
  else throw err;
}

This is quick to adopt, and it is the right tool for a call that cannot move: an ORM write, a native SDK with no HTTP equivalent, something deep in your own code.

What it does not give you: your process still runs your function with your credential. You get the human decision and the record. You do not get credential isolation, and nothing outside your process can attest that the function ran, or did not. The gate depends on every path to that function going through the wrapper.

Shape 2: gate a hosted action

Define the call once as an action, with the credential in a vault. The agent invokes it by name. Because Anlyon makes the request, Anlyon is the one holding it.

const { data } = await anlyon.actions.invoke('refund-order', {
  charge: 'ch_3P9x',
  amount: 12000,
});

if (data!.pendingApproval) {
  // 202. Nothing has been sent to Stripe.
  console.log('Waiting on a human:', data!.approvalId);
}

On approval, Anlyon sends the request snapshot frozen when the agent asked, using the vaulted credential. Editing the action while someone is deciding cannot change what their yes sends. Once the key is out of the agent's process and only in the vault, the agent has no second path around the gate. (Setup is in how to let an AI agent call the Stripe API without your key.)

Local gateHosted action
Who runs the callYour processAnlyon
Who holds the credentialYour processAnlyon's vault
Covers other paths to the same APIOnly calls that go through the wrapperOnly calls routed through the action. A tool your own code calls directly is not governed, and a key your process still holds is outside the vault
Record of the callThe approval decisionThe decision, the request sent and the destination's reply
Best forCalls that cannot become an HTTP requestEverything that can

The rules that make an approval mean something

Whichever shape you use, four properties separate an approval control from a log entry. Check any approval system against them, including one you build yourself.

1. The requester cannot approve itself. If the agent's credential can decide its own request, you have a rubber stamp. In Anlyon, deciding needs the separate approvals:decide scope, and the credential that requested an approval can never decide it, whatever scopes it holds and whatever any policy says.

2. No means the call does not happen. Denied throws. Expired throws. A timeout while the decision is pending leaves the approval open and the call unmade.

3. Routine cases do not reach a human. Policies evaluate on every invocation, per environment, first match wins:

await ops.approvalPolicies.create({
  name: 'large-refunds-need-two',
  matchKind: 'action',
  matchActionName: 'refund-order',
  minAmount: 50000,
  effect: 'require_approval',
  requiredApprovals: 2,
  priority: 10,
});

Use auto_approve for the small cases, require_approval (optionally N-of-M) for the big ones, and auto_deny for what should never happen. An action's own requiresApproval flag is a floor: only a policy that names that action can lower it, so a catch-all auto_approve cannot disarm every gate at once. Configure policies from deploy code with an operator key. The agent's runtime key cannot read or change them.

N counts distinct deciders. For API-key deciders that is distinct credentials, and for dashboard deciders it is distinct users. Two keys held by one person are two deciders, so require dashboard decisions when you mean two people.

4. Every decision is recorded. Who decided, when, with what note, and which policy (and which version of it) routed the request. Each approval also records whether a human, a policy or an expiry decided it, so automated decisions are an audit trail rather than a queue.

Making approval fast enough that people do it

An approval nobody sees is a timeout. The request should reach a person where they already are:

  • The console Approvals inbox updates in realtime.
  • approval_requested alerts go to email by default (the workspace owner is seeded as a recipient), and to Slack, Discord or a webhook once connected, each with a link to the exact card. Email is included during Early Beta Access, capped at 500 emails per workspace per UTC month. A webhook carries an X-Anlyon-Signature header only when you set a webhook secret.
  • Pass a runId and the wait shows up as approval_wait and approval_decision spans in the run's trace, next to the model and tool calls that led to it.

Approvals are not metered on any plan, so there is no cost reason to gate fewer calls than you should.

A checklist

  1. Write the list of calls that should never happen unwatched.
  2. For each, ask whether it can become an HTTP request. If yes, make it a hosted action and gate that. If not, use the local gate and accept what it does not cover.
  3. Give the agent a key without approvals:decide.
  4. Add an auto_approve policy for the routine cases so humans only see what matters.
  5. Connect a channel your team actually reads.

Free during Early Beta Access, no credit card.

Start building → · Approvals and policies in the docs → · Local approval gates →

Free tier, no credit card. One command if you use Claude or Cursor.

$ claude mcp add anlyon -- npx -y @anlyonhq/mcp-server