The Direct Answer

Enterprises should treat an AI agent as a non-human identity with narrowly scoped credentials, not as a trusted extension of the employee who started a task. A defensible authorization architecture connects each agent to a named human or workload principal, assigns a short-lived identity, grants permission only for particular tools, data, actions, and environments, and evaluates every consequential request at runtime. The practical pattern combines standard identity and access management, policy decision and enforcement points, tool gateways, complete audit records, and independent evaluation of whether the agent completed the task correctly. This matters because an agent can plan, call multiple tools, and operate across systems faster than a person can inspect each action. Traditional login controls remain necessary, but they are insufficient once one authenticated user can indirectly trigger thousands of API calls.

Also worth reading: How Should Indonesian Enterprises Structure AI Knowledge Architecture in 2026? · How Should Enterprises Govern AI Agents in Indonesia in 2026? · How do Indonesian enterprises design cloud financial governance models amid localization laws and generative AI deployment costs?

The governing rule should be: authenticate the principal, authorize the requested operation in its current context, constrain the agent’s available capabilities, and record the result. Authentication only establishes identity; it does not prove that the action is appropriate. For example, an analyst may be allowed to view a market report but not export customer-level data, change a forecast, or approve a vendor payment. A 2 a.m. action from a production agent should not receive the same permission merely because the same human can perform it during working hours. Authorization decisions should therefore consider user delegation, agent purpose, resource sensitivity, risk, environment, time, data classification, and sometimes transaction size. The objective is not to make agents inflexible. It is to place measured boundaries around autonomy so that mistakes, prompt injection, compromised dependencies, and unexpected reasoning do not become enterprise-wide incidents.

Identity, Delegation, and Least Privilege

Every agent needs a machine identity that is distinguishable from the user interface, application, API key, or human account that initiated it. Enterprises commonly represent that identity through a workload identity, OAuth client credential, certificate, or signed workload token rather than a permanently stored username and password. A useful naming convention identifies the business owner, environment, agent version, and purpose, such as finance-forecast-agent-production-v7. The identity must not be shared by agents with different responsibilities, even if both run in the same vendor account. If two agents receive identical access because they share one integration credential, the organization cannot determine which agent performed an action or revoke one without disrupting the other.

Delegation is where many designs become unclear. When a person asks an agent to prepare a report, the system must state whether the person can delegate only the necessary read operations or all actions available to that person. The safer default is purpose-bound delegation: report preparation permits reading approved datasets and generating a draft, while publishing, notifying external parties, or modifying records requires a separate permission. High-impact actions can be conditioned on explicit human approval, a lower transaction threshold, or a second agent that checks a policy. A request to send one internal Slack message is different from sending the same content to 5,000 customers, even when both use the same tool. Scope must reflect the requested operation and its context, not merely the agent’s broad job description.

Least privilege should apply to connected tools as carefully as to business data. An agent authorized to query a database should receive a read-only role, limited schemas, row filters, column masking, query duration limits, and perhaps a maximum result count. An email agent should initially be restricted to drafting, with sending disabled until a person approves the recipient list and content. Service accounts should rotate automatically, and dormant or decommissioned agents should lose access within minutes rather than waiting for a quarterly access review. In a 30-agent pilot, 10 or 20 well-tested permissions per agent may be more realistic than 200 generic platform privileges; the exact number depends on the task, but a low ceiling makes anomalies easier to identify. The key metric is not the number of policies. It is the number of distinct actions that can cause material harm if the identity is abused.

Policy Decisions, Enforcement, and Agent Capabilities

A sound architecture separates deciding whether an action is allowed from performing the action. The policy decision point evaluates identity, action, resource, context, and current risk, while the policy enforcement point sits beside each tool or data source and rejects unauthorized calls. This separation allows different tools to use one policy model without giving every tool unrestricted access to the identity provider. For MCP-based agents, the relevant control point is the server that exposes a tool, because that server knows the full tool name, arguments, caller, and target resource. A model’s statement that it “only needs read access” is not an enforcement mechanism. Enforcement must occur outside the model, in code or infrastructure that the model cannot bypass.

Policies can be static, contextual, or risk-based. Static rules are predictable and inexpensive, such as denying production database writes to all research agents. Contextual rules add conditions such as environment, delegation chain, data classification, ticket ID, or approval status. Risk-based rules are useful for ambiguous requests, but they need defined inputs and conservative fallbacks; otherwise, a probabilistic score can grant access unpredictably. A practical threshold might deny all external publishing by default, allow internal drafts for every request, and require review when more than 50 records, 10,000 currency units, or 100 recipients are involved. Those numbers are examples rather than universal standards, and regulated companies may set them at zero.

Agent capabilities should be designed as a layered control rather than as another policy engine. Tool allowlists limit which interfaces can be called, argument schemas restrict the shape of requests, sandboxing limits network and filesystem access, and data-loss controls filter outputs. Model Context Protocol is relevant because authorization and tool governance sit at the server boundary, but adopting MCP does not itself create enterprise-grade access control. The June 2025 MCP authorization revision emphasized stateless request-oriented behavior and removed protocol-level session tracking, changing how authorization metadata is carried. Consequently, a gateway or authorization server must tolerate retries without assuming that a protocol session exists, preserve security boundaries, and avoid storing unnecessary state. Stateless operation can improve scaling, but it also requires careful token validation and replay protection.

A Reference Architecture for Enterprise Teams

A reference design begins with an enterprise identity plane that manages employees, workload identities, service accounts, federation, and revocation. An agent registry then records each agent’s owner, purpose, version, model, tools, datasets, identity, risk tier, and retirement date. A policy layer receives contextual claims from the orchestration service and returns allow, deny, or approval-required decisions. Tool gateways enforce those decisions for APIs, databases, browsers, code runners, messaging systems, and MCP servers. Between them sit a planning or orchestration service, a memory store, a secrets system, and an approval service. An independent audit pipeline captures requests, decisions, tool arguments, outputs, approvals, and errors without recording prohibited data or exposing credentials to the agent.

The orchestration service should request a token scoped to one session or task, not mint permanent broad credentials for the model. When the agent calls a tool, the tool gateway validates the audience, issuer, expiration, nonce where relevant, agent identity, and user delegation. It then constructs an enforceable policy request and returns a short-lived result. External APIs may require a second token exchange because the enterprise identity and the downstream provider use different trust domains. Approval requests should show the intended action, affected resource, expected cost or impact, and any changed fields, rather than merely saying “Approve agent action.” Approval tokens should be single-use and bound to exact parameters so that approval for a $500 transfer cannot be reused for $50,000.

The architecture should distinguish three classes of control. Preventive controls include least-privilege roles, deny-by-default policies, network isolation, and secretless tool access. Detective controls include anomaly detection, unusual tool sequences, repeated authorization failures, and deviations from an agent’s normal operating profile. Responsive controls include token revocation, agent quarantine, tool shutdown, session termination, and notification to the owner. A useful pilot target is to revoke an agent’s access in under 5 minutes and contain a compromised tool connection in under 15 minutes; organizations should set their own objectives based on incident severity. No single component provides all three classes. Identity stops an unknown caller, policy limits a known one, telemetry helps investigate behavior, and rapid revocation limits persistence.

Comparing Authorization Approaches

There is no single product category called an “agent authorization layer.” Enterprises usually combine existing controls with newer agent-specific products, and each option has a different balance of flexibility, assurance, and operating cost. The comparison below concerns architectural approaches, not endorsements of particular vendors.

FeatureConventional IAM and role-based accessAgent gateway or authorization layerHuman approval for every action
Authorization speedMilliseconds to secondsMilliseconds to secondsMinutes to hours
Best control pointUser, service account, API roleAgent-to-tool request and downstream resourcePerson reviewing intent and impact
GranularityRole, group, resource, and API scopeAgent identity, task, tool, argument, context, and riskExact proposed action and transaction
Typical operating costLowest incremental cost; existing staff and licensesPlatform, gateway, policy-engine, and integration costStaff time, workflow tooling, and delay
Main weaknessStatic roles may become broad or sharedCan become another pass-through unless independently enforcedBottleneck and rubber-stamp risk
Appropriate autonomyLow to moderateLow, moderate, or high within enforced boundariesNarrow or high-risk workflows
These options are alternatives only at the margin. A production architecture normally uses all three: conventional IAM for foundational identities, an authorization layer for agent-specific context, and human approval for exceptional actions. A gateway cannot safely mediate a database operation if the agent still holds unrestricted database credentials that bypass it. Human review also fails as a control when reviewers approve hundreds of routine actions without inspecting them. The preferred design makes ordinary low-impact actions automatic, rejects clearly prohibited actions, and reserves human attention for decisions where mistakes are difficult to detect or reverse.

Policy-as-code engines are another relevant option. Cedar, used by AWS for multi-agent authorization, illustrates how explicit policy can express relationships and least-privilege decisions without embedding authorization logic in prompts. Explicit policies improve testability and review, especially when the same chain contains several agents. They can also become difficult to maintain if every attribute becomes a bespoke rule. Regulatory or standards-based engines may fit regulated enterprises, while application-level checks remain necessary for domain controls such as transaction limits, segregation of duties, and approved business calendars. The decision should be driven by existing architecture, audit requirements, and the team’s ability to test policies. Adding a new engine is not automatically better than extending well-governed IAM and a thin tool gateway.

Implementation Steps for B2B Teams

Start with a bounded workflow rather than an “autonomous enterprise agent.” Select one process with identifiable value, limited data, reversible actions, and a clear owner; invoice triage, weekly market monitoring, or draft research briefs are more suitable starting points than treasury movement or customer termination. Inventory every human, service account, model provider, connector, API, dataset, and downstream action involved. Then create a threat model that includes prompt injection, stolen credentials, excessive tool permissions, delegated authority, unsafe outputs, privacy leakage, and actions performed by one agent on instructions from another. The inventory is not bureaucratic overhead. It often reveals that the same cloud integration key is shared by research, reporting, and publishing components, creating a wider blast radius than the initial agent requires.

Next, establish a minimum permission contract for the pilot. Replace shared credentials with workload identities, issue task-scoped tokens, separate read and write tools, restrict network destinations, and add argument validation. Write allow, deny, and approval-required policies before connecting the model, and test them for direct calls, tool substitution, chained agents, replay, token expiry, and attempts to bypass the gateway. A reasonable policy test suite contains at least 100 cases for a production-critical pilot, including 20 positive cases, 50 abuse or boundary cases, 20 revocation or concurrency cases, and 10 failure-mode cases. These are engineering targets, not industry benchmarks. Track authorization latency at the 50th, 95th, and 99th percentiles, because a control that adds several seconds to every step can make an agent impractical even if it is secure.

Run the pilot in shadow mode or with a read-only tool set before granting write access. Compare intended actions with actual calls, inspect false denials, and measure how often users repeat requests because approvals or recoveries fail. For Indonesian and Southeast Asian teams, the operating model must also account for data residency, sector rules, cross-border transfers, local-language content, and provider availability; security architecture cannot ignore where prompts, logs, and embeddings are stored. After 4 to 8 weeks, review the evidence with the business owner, security team, legal or privacy function, and operations team. Promote the agent only when the organization can answer who owns it, revoke it, explain each sensitive action, and produce records sufficient for an incident investigation.

Common Mistakes and Cost Trade-Offs

The most common design error is confusing identity with authorization. A valid workload token proves which workload called, while a policy still determines what that workload may do. Another error is giving the model broad access to a connector so it can discover features at runtime. The model may then invoke an undocumented operation, use a destructive variant, or combine tools in a sequence that no individual test covered. Tool names such as files.update should expose narrower operations, and the platform should remove unused methods. A prompt that asks the model to “stay within policy” is useful for behavior but is not a replacement for server-side enforcement because prompt injection can alter model behavior.

Teams also err by granting permanent secrets, giving all agents one identity, and treating approval as a universal brake. Shared keys defeat attribution; permanent credentials increase persistence after a breach; and approval fatigue can make reviewers approve dangerous actions mechanically. Excessive logging creates a different problem by capturing customer records, prompts, or secrets in central logs. Audit systems should preserve decisions and evidence while applying retention, redaction, and access controls. For example, storing a hash of a sensitive query plus resource, actor, policy, and decision may be sufficient for investigation, whereas copying the full customer dataset into logs is not.

Pricing is usually based on a combination of identity volume, policy decisions, gateway requests, tool calls, log volume, evaluation runs, and enterprise support. Existing IAM, API management, security information and event management, and data-platform licenses may reduce incremental cost, but integration and policy maintenance are rarely free. A pilot with 30 agents making 20,000 tool calls per day could remain affordable on existing infrastructure, yet a regulated deployment with 300 agents, 2 million monthly decisions, long-term evidence storage, and 24/7 support can become a material platform expense. Buyers should price not only the authorization engine but also token brokering, secrets management, evaluation, incident response, and the engineers who maintain policy. The cheapest product is not necessarily the least expensive architecture if it produces manual reviews, emergency credential rotation, or repeated integration work.

When to Act and How to Measure Success

Organizations should act before deploying agents that can change business records, communicate externally, execute code, handle personal data, or spend money. Waiting for a fully autonomous enterprise program is unnecessary because agents already create risk when they can read internal systems and call APIs. Immediate action is justified when a single user can trigger many actions, when credentials can be delegated through tools, or when one agent can influence another agent’s permissions. Lower-risk read-only assistants still need inventory, data classification, monitoring, and a revocation path, but they may not require approval for every query. The appropriate response follows the action’s consequence, not the marketing label “autonomous.”

Useful measures include percentage of agents with named owners and machine identities, percentage of tool calls passing through a policy enforcement point, mean time to revoke access, number of shared credentials, percentage of privileged actions with attributable audit evidence, and authorization latency at the 95th and 99th percentiles. Security outcomes include blocked cross-tenant access, detected token replay, reduced unauthorized tool use, and speed of containment. Business outcomes include completion rate, human review time, false-denial rate, recovery success, and cost per completed task. A target such as 99.9% availability for the authorization service may be reasonable for a critical production path, but 99.9% still represents roughly 43 minutes of unavailability per month, so downstream services need safe behavior when policy checks are unavailable. Security-critical writes should fail closed; low-risk reads may sometimes use a tightly bounded cached decision, depending on risk tolerance.

As of 28 September 2026, the practical direction is toward workload identity, explicit delegation, short-lived credentials, server-side policy enforcement, authorization-aware gateways, and continuous evaluation of both permission and behavior. The market contains experimental agent-specific access-control projects, MCP security controls, runtime enforcement layers, and established IAM vendors all addressing parts of this problem. That variety means buyers should demand interoperability, testability, revocation speed, and evidence quality rather than assume a new category has replaced identity fundamentals. The correct architecture does not promise perfect autonomy. It makes autonomy conditional, observable, and stoppable at a cost the business can understand.

The Governance Decision

The main decision is not whether an enterprise should allow AI agents to act, but how much independent action each agent can safely hold. Start with read access and reversible operations, then expand permissions only when telemetry shows that the agent, identity, tool contract, and recovery process work as designed. Reserve irreversible, regulated, financial, or externally visible actions for explicit policy and human approval. Revoke temporary delegation when the task ends, and review persistent permissions at least quarterly for high-impact agents and monthly for privileged ones; higher-risk deployments may need continuous review.

The architecture should be owned jointly by security, identity, platform engineering, data governance, and the business process owner. A model vendor can secure its own service, but it cannot decide which records an Indonesian finance team may export or which customer communications require local approval. Likewise, a gateway can enforce a policy but cannot repair a weak identity model or an inaccurate business rule. A mature program treats authorization as a living control informed by incidents, tool changes, model updates, new data sources, and actual behavior. It also measures whether controls preserve speed. If every step requires a meeting, agents will be routed around the system; if almost every step is allowed, the architecture is mostly documentation. Effective authorization combines a small number of clear restrictions with rapid, attributable decisions for ordinary work.