agentgateway in Production: One Endpoint for MCP Servers, Agents, and Models

An agent that calls tools usually holds a list of MCP server addresses and a credential for each one. A gateway replaces that list with one address. Every request passes through it, and it checks the caller, decides which tools that caller can use, and records the call. That gives you one place to answer the question after an incident: who called which tool, and why was the call allowed? I ran agentgateway 1.6.0 on my laptop and measured what each policy does to a real request. This guide builds one setup step by step, from the first call to production.


Step 1: What a gateway does

The gateway in this guide is agentgateway, an Apache-2.0 proxy written in Rust. It speaks MCP, A2A, and the common LLM APIs on the same listener. Solo.io donated it to the Linux Foundation on August 25, 2025. On June 4, 2026, it joined the Agentic AI Foundation. Version 1.6.0 shipped on October 2, 2026, and every number in this guide comes from that release.

One protocol change shapes how you run a gateway. Before July 28, 2026, MCP opened with a handshake. The client sent initialize, the server picked a protocol version, and the client confirmed with notifications/initialized. The server then issued an Mcp-Session-Id, and every later request carried that header. The version lived in the session.

The 2026-07-28 revision moves the version into each request. A request declares it in _meta under io.modelcontextprotocol/protocolVersion. On Streamable HTTP, the same value travels in the MCP-Protocol-Version header. The server accepts or rejects each request on its own. A client can call server/discover first, but it does not have to.

For a gateway, this is the change that matters most. A request that carries its own version needs no memory of earlier requests. So any gateway replica can forward it to any healthy server. Old-protocol sessions still need care, and step 9 shows why. Version 1.4.0 added full support for 2026-07-28, one day before that revision became current.

Application state does not go away. A pending refund approval or an idempotency key still needs an owner. Kai's analysis on mubibai.com makes this point well. The gateway can send a request to any server replica, so that state has to live in the server or in a store behind it.

This guide uses one example from start to end. Two Python MCP servers sit behind the gateway. The notes server keeps no sessions and has search_notes and get_note. The billing server keeps sessions and has get_invoice and refund_invoice, with fake data. Users carry an ES256 JWT with a role claim. Alice has the support role, and Bob is an admin.

The rule is simple: support can read everything and refund only small amounts. Later steps add a billing agent that other agents call, and a model behind the same gateway. Each step adds one thing and shows what a real request does.

Step 2: Run it and send the first call

Start with the two servers behind one listener and no policies. One URL replaces the list of server addresses in each client. Imagine Learning grew from 5 to 12 servers behind one endpoint this way, though they published no numbers. This is the whole config. Every field comes from the agentgateway documentation and its schema files.

config:
  adminAddr: 127.0.0.1:15000
  statsAddr: 127.0.0.1:15020
  readinessAddr: 127.0.0.1:15021
binds:
- port: 3000
  listeners:
  - routes:
    - backends:
      - mcp:
          targets:
          - name: notes
            mcp:
              host: http://127.0.0.1:8101/mcp
          - name: billing
            mcp:
              host: http://127.0.0.1:8102/mcp

Check the file with agentgateway -f gateway.yaml --validate-only, then start it without the flag. The gateway was ready in about 4 ms. A client on either protocol version saw four tools: notes_search_notes, notes_get_note, billing_get_invoice, and billing_refund_invoice.

The config sets the admin, stats, and readiness addresses on purpose, and step 7 uses all three. The schema now marks binds as deprecated in favor of gateways. The examples in the docs still use binds, so this guide uses it too. Pin the gateway version, and read the breaking changes in each release note before an upgrade.

Two client details can stop the first call. In 2026-07-28 mode, the Python MCP SDK 2.3.0 wants an mcp-name header on tools/call that matches params.name. Without it, the server returns HTTP 400 with code -32020. Also send the Accept header as one line with both media types. Separate Accept lines can produce an unexpected HTTP 406.

Now follow one tools/call through the gateway. The listener on port 3000 accepts the POST to /mcp. The gateway splits billing_get_invoice at the first underscore into the target billing and the tool get_invoice. It forwards the call, relays the answer, and writes a log line, two spans, and metric samples.

Policies sit between the listener and the upstream call. Each request passes through them in order, and the first one that says no wins. The order decides what the client sees, which telemetry exists, and whether the server ever learns about the call. Steps 4 to 6 add those policies one at a time.

The first figure shows the whole setup this guide builds. A person, two agents, two MCP servers, and a model provider all talk to the gateway. None of them talk to each other directly. Pick a scenario to send its packets one hop at a time.

Fig 1 · Gateway traffic replay

Each scenario replays measured requests through one agentgateway. A stage lights when the packet passes it, the log records its verdict, and the readout counts hops and policies.

personchat app user
support-agentcarries the alice JWT
agentgatewayone listener, five stages
jwtAuth
localRateLimit
mcpAuthorization
route
trace
billing-agentA2A :9999
notesMCP :8101
billingMCP :8102
model provider:11999 dead, :11434 Ollama
Log
Readout

Measured on agentgateway 1.6.0. MCP calls use the configs from steps 2 to 6, A2A uses the route from step 11, and failover uses the llm block from step 10. A dashed stage is not on that path.

The figure also shows two shorter paths. A request to /v1/chat/completions names a model, and the gateway resolves the model to a provider. It applies prompt guards and token limits, and its log line carries gen_ai.operation.name, gen_ai.provider.name, and token counts. A route with the a2a: {} policy forwards agent-to-agent traffic and logs a2a.method. Steps 10 and 11 cover both paths.

Step 3: Name tools across two servers

Each tool name gets a <target>_ prefix, so two servers can both offer a tool called search. The prefixMode field accepts always, conditional, or never. Lin Sun pins always before a migration, so tool names stay the same when someone adds a target later. I agree with that default.

The gateway finds the target by splitting the name at the first underscore. That rule decides what happens when an agent sends an old, unprefixed name. A call to get_invoice did not return an unknown tool error. It returned HTTP 500 with -32603 and unknown service get, because the gateway read get as a target name.

The error also depends on the client. A 2026-07-28 client that sent the same name got HTTP 400 with -32602 invalid request parameters: unknown service get. So an agent on the old protocol sees a server fault, and an agent on the new one sees a bad request.

Give each target a short name with no underscore, such as billing or notes. Then search agent prompts, skills, and tests for old tool names before agents depend on the gateway. Handle both errors while the names change.

Step 4: Add identity and tool rules

Authentication and authorization are separate policies. The jwtAuth policy proves who is calling. The mcpAuthorization policy decides which tools that caller can see and use. This guide uses local keys in place of an identity provider. Christian Posta's MCP authorization series and his Entra SSO walkthrough cover the identity provider side. Add this block to the route that holds the two backends:

    - policies:
        jwtAuth:
          mode: strict
          issuer: lab.local
          audiences: [agw-lab]
          jwks:
            file: ./keys/jwks.json
        mcpAuthorization:
          rules:
          - allow: 'jwt.role == "admin"'
          - allow: 'jwt.role == "support" && mcp.tool.target == "notes"'
          - allow: 'jwt.role == "support" && mcp.tool.target == "billing" && mcp.tool.name == "get_invoice"'

The jwks.file field needs a JWKS JSON file, not a PEM key. Its path is relative to the working directory of the process. In CEL, mcp.tool.name is the unprefixed upstream name, and mcp.tool.target is the target. Since version 1.5.0, the gateway checks that iss matches. It also checks aud when audiences are set, so tokens from an older setup can return 401.

Now send the same calls again. A request with no token got HTTP 401 with the text authentication failure: no bearer token found. Alice's tools/list returned three tools, without the refund tool. Bob's list had all four. When Alice called billing_refund_invoice anyway, she got HTTP 400 with -32602 Unknown tool: billing_refund_invoice on both protocol versions.

That denied call never reached the billing server, and the trace had a server span with no client span. So a denied tool looks exactly like a tool that does not exist. That hides your catalog from a curious caller, and Hoon Jo traces it to a deliberate choice against tool enumeration. It also means a client cannot tell "forbidden" from "missing." Write agent prompts so that an unknown tool error ends the attempt, instead of starting a search for the right name.

The docs show two shapes for rules. The shipped example uses bare CEL strings. The schema describes objects with one key: allow, deny, or require. Both shapes pass --validate-only, and bare strings act as allow at runtime. A list that mixes the two shapes is also valid.

An item with two keys, such as allow and require, fails with data did not match any variant of untagged enum RuleSerde. An unknown key such as permit fails the same way. Invalid CEL fails with a parse error that points at the column. I use the object form because it states intent. I avoid deny, because the schema warns that an expression that fails to evaluate also fails to deny.

The second figure replays six real requests through the policy stack. It lights the stages each request passed and marks the stage that decided. Three boxes show what the client received, what the log recorded, and which spans exist.

Fig 2 · Policy stack replay

Pick a request. The stack lights the stages it passed, marks the stage that decided, and shows the response and telemetry captured for it.

Measured on agentgateway 1.6.0 with the identity config from step 4 and the limit from step 6. Response bodies are trimmed. Where telemetry was not captured, the box says so.

Step 5: Check tool arguments in the right place

Role rules are not enough for billing. Support should refund small amounts, so a rule has to read the amount_usd argument. Hoon Jo (sysnet4admin) found the problem with that in his enforcement study on kuberneteslab.dev. The CEL reference marks mcp.tool.arguments as available only after the request. So inside mcpAuthorization, the field is absent.

I gave refund_invoice an optional amount_usd argument and put the same intent in three places. With the check inside mcpAuthorization, the rule never matches, so it fails closed. The refund tool disappears from support's list, and every refund call returns Unknown tool.

The second placement adds a !has(mcp.tool.arguments) || guard before the same check. The guard is always true at that point, so the rule fails open. An 80 dollar refund went through.

The third placement works. The mcpAuthorization rules allow the billing target, and a route-level authorization rule with require checks the amount. A 20 dollar refund succeeded, and an 80 dollar refund returned HTTP 403. The validator accepted all three configs without a warning, and the log flags none of them. The guarded version is the dangerous one, because it looks like a working policy.

Fig 3 · Argument rule bench

One intent, three placements. Support can refund only when amount_usd is under 50. Each row compares the intent with what the gateway did.

Measured on agentgateway 1.6.0 as user alice with role support. Old-protocol and 2026-07-28 clients gave the same results.

This is the working placement. It replaces the mcpAuthorization block from step 4:

        mcpAuthorization:
          rules:
          - allow: 'jwt.role == "admin"'
          - allow: 'jwt.role == "support" && mcp.tool.target == "notes"'
          - allow: 'jwt.role == "support" && mcp.tool.target == "billing"'
        authorization:
          rules:
          - require: '!has(mcp.tool) || mcp.tool.name != "refund_invoice" || jwt.role == "admin" || (has(mcp.tool.arguments.amount_usd) && mcp.tool.arguments.amount_usd < 50)'

It has two costs. The denial is a plain-text HTTP 403 authorization failed, not a JSON-RPC error. Some clients treat a 403 as a reason to sign in again. The refund tool also stays in support's tool list, because the list filter only knows the mcpAuthorization rules.

The project closed the upstream report on argument rules on September 8, 2026, as a documentation change. The route-level rule is the documented way to check arguments. For every CEL variable in a policy, check when the gateway fills it in. Then send one call that the rule must deny, and confirm the denial. A policy test without a denial case proves nothing.

Step 6: Limit calls per user

Give each user a localRateLimit bucket beside jwtAuth, with type: requests and key: jwt.sub. The bucket holds maxTokens: 5 and adds tokensPerFill: 1 every fillInterval: 1s. Alice's sixth call in a burst was limited, and Bob had his own bucket. After 2.1 seconds of rest, Alice got two more calls.

The bucket behaved exactly as configured. The surprise was the response shape, which depends on the request type. A limited tools/call returns HTTP 200 with result.isError: true. Its text starts with rate limit exceeded and gives the limit of 5 with 0 remaining. Other MCP methods, such as tools/list, return HTTP 200 with JSON-RPC error -32003. Its data holds limit, remaining, and retryAfterSeconds.

Only a non-MCP body on the same route gets HTTP 429. That response carries x-ratelimit-limit, x-ratelimit-remaining, x-ratelimit-reset, and retry-after: 1.

My first test script printed only status codes and saw no limiting at all. A limited tool call becomes a tool error that the model reads in its own context. The model can then wait, try another tool, or tell the user.

It does change your retry logic. HTTP retry middleware never sees a 429, so it never backs off. The agent loop has to read the isError text or the -32003 code. Dashboards that count 429s will show zero, so count reason=RateLimit in the access log instead.

Local buckets live in one process, so three replicas behind a load balancer give each user three buckets. For one limit across the fleet, use remoteRateLimit. It calls an external rate-limit service with a domain and a list of descriptors. Its default failureMode is failOpen, so traffic keeps flowing when that service is down. Choose that mode on purpose. Token limits for LLM routes use the same policy with type: tokens.

Step 7: Watch every call

The gateway sends OpenTelemetry traces over OTLP. In v1.6.0, config.tracing still works but logs a deprecation warning, and the replacement is frontendPolicies.tracing. I used it with host: 127.0.0.1:4318, protocol: http, and randomSampling: "true".

The documentation did not list span names, so I captured them. Each tool call produces a server span named tools/call and a child client span named tools/call billing_get_invoice. The server span carries mcp.method.name, mcp.target, mcp.resource.type, gen_ai.tool.name, jwt.sub, and http.response.status_code. The client span carries agentgateway.outbound.kind=Primary and agentgateway.outbound.subtype=Mcp, plus the target and the tool name. In one call, the server span took 2,153 microseconds and the client span took 1,348.

The server span covers the client's request, and the client span covers the upstream call, so the gap between them is gateway time. A federated tools/list produces one client span per target, such as tools/list billing. A denied call produces only the server span, with status ERROR and error.type=MCP. That turns "blocked by policy" into a simple trace query: a server span with no child. Put an alert on that query.

The access log has the same fields, plus trace.id and span.id. Error records log at the error level and all others at info. During an incident, a search for level=error protocol=mcp is a fast first look.

Metrics come from :15020/metrics, in OpenMetrics or Prometheus protobuf format. The agentgateway_mcp_requests_total counter carries method, resource_type, server, and resource, so it names the tool. The denied refund still counts there under its real tool name. Join it with agentgateway_requests_total, which carries protocol, status, and reason, to separate attempts from successes. LLM routes add the agentgateway_gen_ai_* family, including time to first token and agentgateway_gen_ai_client_cost_usd_total.

The admin UI lives at 127.0.0.1:15000/ui, and the same port serves a full config dump at /config_dump. Readiness answers ready at :15021/healthz/ready. Because the config sets every address, probes and scrape configs always match. Keep the admin port on localhost. If you must expose the UI, the schema recommends OIDC sign-in for it.

Step 8: Measure what the gateway costs

Before you put every call through one more hop, measure the hop. I sent one 2026-07-28 tool call, search_notes, 500 times after 50 warmup calls. Each condition ran three rounds, mixed with direct calls, and the table reports the median round. Each gateway run has its own direct baseline from the same batch. I ran the benchmark twice, hours apart.

RunGateway configDirect p50Gateway p50Added
1No policies1.64 ms2.28 ms+0.64 ms
1JWT, CEL rules, full tracing1.33 ms1.87 ms+0.54 ms
2No policies1.04 ms1.27 ms+0.23 ms
2JWT, CEL rules, full tracing1.04 ms1.37 ms+0.33 ms

With one client, the gateway added about 0.2 to 0.6 ms at p50. At p99 it added about 3 ms in the first run and about 2 ms in the second. The laptop moved the number more than the gateway did, and the direct baseline itself fell from 1.64 ms to 1.04 ms between runs. JWT, three CEL rules, and 100% trace sampling added at most about 0.1 ms. In the first run, one client got 587 requests per second direct and 406 through the gateway.

With 16 clients, the single-process Python server was the limit, at about 450 to 690 requests per second. Gateway throughput stayed close to direct, and the gateway's p99 was lower, around 180 to 200 ms against 245 to 290 ms. These numbers measure the Python server, not the gateway's own ceiling. I did not run a faster upstream to find that ceiling. Gateway memory was about 20 MB at idle and peaked near 31 MB, or 35 to 36 MB with auth and tracing on.

Other measurements agree on the shape. Hoon Jo measured under 1 ms at p50 per tool call on Kubernetes, and 1 to 3 ms per tool list. In his study, p99 never improved through the gateway. Tool-list filtering cost about 0.7 ms at p50 whatever the tool count, because the gateway fetches the full upstream list first. A slow upstream list stays slow.

The project's own benchmarks test other limits. Against LiteLLM with a mock backend, they report 36,933 requests per second at a 1.97 ms p99 and 22 MB of memory. On real H100 GPUs, peak output rose from 6,910 to 16,178 tokens per second, a 134% gain from smarter routing. John Howard's Gateway API benchmark found 4 MB to 40 MB of memory at 5,000 routes. All three come from people close to the project, so treat them as upper bounds and measure your own path.

Step 9: Run more than one copy

Before you add a second replica, look at the session token. Old-protocol clients get a gateway session ID, and by default anyone can read it. Mine was base64 JSON. Decoded, it held "t":"mcp", a list with one entry per target, and a request ID. The notes entry had only its name. The billing entry also held the real upstream session ID, a 32-character hex value.

So any client that holds the token can read the session ID of every stateful server behind the gateway. Set config.session.key to 64 hex characters from openssl rand -hex 32. The gateway then encrypts the token with AES-256-GCM, and the token becomes opaque. The same key also decides whether replicas can share sessions.

To test that, I ran two identical gateway replicas on ports 3001 and 3002. A small proxy sent requests to them in strict turns. I made nine calls in an old-protocol session, then repeated the test on 2026-07-28.

The old session survived the alternation with no key set. The token carries every upstream session ID, so any replica that can decode it can resume the session. No shared store is needed. The same was true when both replicas had the same session.key.

With different keys, every request that reached the other replica failed with HTTP 400 mcp: invalid session ID header. That included notifications/initialized. In a rolling deploy that changes the key, half of all old-protocol calls would fail. Keep the key in a secret and give every replica the same one. The 2026-07-28 calls succeeded in every variant.

The gateway also has statefulMode: stateless, where it issues no session ID and opens a fresh upstream session for each request. Everything succeeded, but the upstream paid for it. One old-protocol client call caused 10 HTTP requests and 3 short-lived sessions at the billing server. The same call on 2026-07-28 caused one upstream request.

Where affinity still matters

The token approach works because both replicas reach the same billing process. If billing runs as three pods, each session lives in one of them. The gateway then has to send each session to the same pod. Version 1.6.0 adds a backend sessionAffinity policy for this, with consistent hashing over a CEL expression such as request.headers["x-session-id"]. The schema calls it best-effort, so requests can move when the set of healthy endpoints changes. I did not test it.

Lin Sun's rollout guide on the AAIF blog shows what happens when sessions move. At a 10% canary weight, 100 old-protocol session calls produced 20 "Session not found" errors. A separate stateful backend stayed at 100 of 100. The lasting fix is the protocol, not the router: move clients to 2026-07-28.

Draining

Version 1.6.0 also drains connections on shutdown. The config.connectionTerminationDeadline field sets how long the gateway waits for open connections. The config.connectionMinTerminationDeadline field sets a minimum wait, and it defaults to zero. On SIGTERM, my logs showed "drain started, waiting 0ns-5s," and then the readiness listener closed. Set the Kubernetes terminationGracePeriodSeconds longer than the deadline. Otherwise the kubelet kills streams that the gateway was still draining.

Step 10: When a server or a model dies

I killed the billing server with SIGKILL during a test, then started it again. Calls to notes kept working, and calls to billing returned HTTP 500 with -32603 and Connection refused. But tools/list also failed for the whole federation with HTTP 500, and so did a new old-protocol initialize.

So by default, one dead server blocks every new session and every tool list. New clients cannot reach its healthy neighbors. The default failureMode for an mcp backend is failClosed. With failureMode: failOpen, tools/list returned HTTP 200 with only the two notes tools, and new sessions opened. When billing came back, the list returned to four tools on the next request, with no gateway restart. If every server fails, the list call returns an error, not an empty success.

After the restart, old sessions got HTTP 404 Session not found from billing, because the upstream session died with the process. That error comes from the server, not the gateway. The client has to open a new session. The 2026-07-28 calls recovered on the first request.

Errors from a live server pass through unchanged, so the agent sees the code the server sent. The gateway's own -32603 with Connection refused means something else: the gateway could not reach the server at all.

Failover for models is health-based

The model side works differently. I set up an llm: block on port 4000 with two models, both with provider: openAI and the model qwen3-4b-instruct-8k:latest. The primary's baseUrl pointed at http://127.0.0.1:11999/v1, where nothing listens. The second pointed at a local Ollama on http://127.0.0.1:11434/v1. A health.eviction rule on the primary set consecutiveFailures: 1 and duration: 30s. A virtualModels entry named chat listed both under routing.failover.targets, with the primary at priority 0 and Ollama at priority 1.

The first request to chat returned HTTP 503 in 3 ms, and that failure evicted the primary. The next two requests went to Ollama and returned 200. So failover protects the next request, not the current one. The llm: shorthand has no retry field. A route retry policy exists in the binds form, but I did not test it. Expect one failed request per eviction, and let the client retry.

Step 11: Send agent-to-agent traffic through it

Now add the billing agent. For agent-to-agent traffic, the route needs one policy line, a2a: {}, and a backend host, which was 127.0.0.1:9999 here. A message/send through the gateway returned the agent's answer. The access log carried a2a.method=message/send, a2a.response.outcome=success, a2a.result.kind=message, and a stable a2a.context.id.

The gateway also rewrites the agent card at /.well-known/agent-card.json and /.well-known/agent.json. The card then points clients at the gateway, so later calls do not go around it. A card can publish its endpoint in two places, though. A2A v0.3 clients read a top-level url, and v1.0 clients read supportedInterfaces[].url.

A card with only the v0.3 field had its url rewritten. A card with both fields had only the v1.0 field rewritten, and its top-level url still held the backend address. A v0.3 client that reads such a card calls the agent directly, outside every gateway policy. Hoon Jo documented this in his reproducible cluster study, and it reproduced on v1.6.0. Publish one endpoint field per card, or test which field your clients read.

There is also no A2A authorization. The schema offers only the a2a marker, and the CEL reference has no a2a.* variables. Route policies such as jwtAuth, authorization, and rate limits still apply to the HTTP request, though I did not test them on A2A. You cannot write "agent X can send tasks of skill Y to agent Z" today.

Until that changes, use route JWT and authorization rules for coarse caller checks. Block direct network access to agents with a Kubernetes NetworkPolicy or a firewall, so the gateway is the only way in.

Step 12: Run it on Kubernetes

On Kubernetes, agentgateway has its own Gateway API controller, conformant for HTTPRoute, GRPCRoute, TCPRoute, and TLSRoute. It turns resources into configuration and streams it to the proxies over xDS. Standalone proxies can use the same channel with config.xdsAddress. The agentgateway_xds_* metrics show the health of that stream.

I did not run Kubernetes for this guide, so this step comes from the Kubernetes MCP quickstart and the AgentgatewayPolicy reference. A Service marks MCP with appProtocol: agentgateway.dev/mcp. An AgentgatewayBackend from agentgateway.dev/v1alpha1 lists the MCP targets, each with a backendRef to a Service. A gateway.networking.k8s.io/v1 HTTPRoute names the agentgateway-proxy Gateway in parentRefs. Its backendRefs send the /mcp path prefix to that backend. This is the quickstart's own example, with its SSE fetch server:

apiVersion: agentgateway.dev/v1alpha1
kind: AgentgatewayBackend
metadata:
  name: mcp-backend
spec:
  mcp:
    targets:
    - name: mcp-target
      static:
        backendRef:
          name: mcp-website-fetcher
        port: 80
        protocol: SSE
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: mcp
spec:
  parentRefs:
  - name: agentgateway-proxy
    namespace: agentgateway-system
  rules:
  - matches:
    - path:
        type: PathPrefix
        value: /mcp
    backendRefs:
    - name: mcp-backend
      group: agentgateway.dev
      kind: AgentgatewayBackend

Policies attach through AgentgatewayPolicy, and its authorization fields map to the standalone ones. The field spec.backend.mcp.authorization covers MCP servers and tools. The field spec.traffic.authorization covers routes, and spec.backend.authorization runs after the gateway picks a backend. So the argument check from step 5 goes in the route-level field here too.

Two v1.6.0 changes matter on Kubernetes. When two policies of equal specificity set the same field, the oldest one now wins. A backend with one invalid inline policy is now accepted as PartiallyValid instead of dropped, so check resource status after every apply. The docs also warn against running the HPA and the VPA together on the controller, because they can fight each other. The data plane is lighter, and the replica rules from step 9 apply to it.

Step 13: Put the whole config together

This is the full config for the support role. It puts the two servers behind one endpoint, requires a JWT, limits support to safe tools, and allows small refunds only. It gave the results in Fig 3.

config:
  adminAddr: 127.0.0.1:15000
  statsAddr: 127.0.0.1:15020
  readinessAddr: 127.0.0.1:15021
binds:
- port: 3000
  listeners:
  - routes:
    - policies:
        jwtAuth:
          mode: strict
          issuer: lab.local
          audiences: [agw-lab]
          jwks:
            file: ./keys/jwks.json
        mcpAuthorization:
          rules:
          - allow: 'jwt.role == "admin"'
          - allow: 'jwt.role == "support" && mcp.tool.target == "notes"'
          - allow: 'jwt.role == "support" && mcp.tool.target == "billing"'
        authorization:
          rules:
          - require: '!has(mcp.tool) || mcp.tool.name != "refund_invoice" || jwt.role == "admin" || (has(mcp.tool.arguments.amount_usd) && mcp.tool.arguments.amount_usd < 50)'
      backends:
      - mcp:
          targets:
          - name: notes
            mcp:
              host: http://127.0.0.1:8101/mcp
          - name: billing
            mcp:
              host: http://127.0.0.1:8102/mcp

Add the remaining pieces one at a time. The localRateLimit block from step 6 goes beside jwtAuth. The frontendPolicies.tracing block from step 7 goes at the top level. The config.session.key from step 9 goes on every replica, with the same value. If a partial tool list suits your agents, add failureMode: failOpen from step 10 under mcp.

After each addition, run the same eight calls with any MCP client. The full set takes about two minutes:

  1. No token: expect 401.
  2. Support lists tools: expect four, because the billing target is allowed.
  3. Support gets an invoice: expect 200.
  4. Support refunds 20 dollars: expect 200.
  5. Support refunds 80 dollars: expect 403.
  6. Support refunds with no amount: expect 403.
  7. Admin refunds 80 dollars: expect 200.
  8. Six calls in one second: expect the sixth to carry isError: true.

Four of the eight calls test a denial or a limit. Keep that share as the rules grow, because a test set with no denial cases cannot catch a rule that fails open.

Step 14: Move your existing setup behind the gateway

Most teams start with servers listed straight in client configs. The move can happen in small steps, and you can undo each one until the last. First, list every MCP server in every client config, with its URL, transport, and embedded credentials. Each credential becomes a gateway concern later.

Next, pick short target names with no underscore and pin prefixMode: always. Tool names change once, so change them before agents depend on them. Run the gateway beside the servers with no policies, and point one test client at it. Compare its tool list with the direct lists, tool by tool. Then search agent prompts, skills, and evals for old tool names, because an unprefixed name now returns a 500 or a 400.

Then add identity. Turn on jwtAuth with tokens from your identity provider, and match the iss and aud claims that the gateway checks. Write mcpAuthorization rules per role, and diff the tool list for each role. Put argument checks in route authorization, never in mcpAuthorization. Run --validate-only in CI on every config change, and run the eight-call test set after every deploy.

Then add rate limits per jwt.sub, and teach the agent loop to read isError. Send traces to your collector and save the query for denied calls. Run two replicas with one shared session.key, point readiness probes at /healthz/ready, and set the drain deadline and grace period.

Move servers and clients to 2026-07-28, where Lin Sun's dark launch, canary, and promote sequence fits. Last, remove the direct network paths from clients to servers. Until you do, anyone who knows an old URL can go around the gateway. Keep the old client configs until this step, so you can point any client back at its servers.

Step 15: When things go wrong

Most problems I hit came back as a status code that pointed somewhere else. This step maps each symptom to its cause and its fix, roughly in the order you are likely to meet them.

A 500 on the old protocol or a 400 on the new one, with unknown service get, means an unprefixed tool name. The gateway split it at the first underscore. Send prefixed names, and keep the underscore out of target names.

When a tool vanishes from one role's list and every call returns Unknown tool, an mcpAuthorization rule probably reads mcp.tool.arguments. That field is absent there, so the rule never matches. The opposite symptom, a rule that allows calls it should deny, points at a !has(mcp.tool.arguments) guard. In both cases, move the argument check to route authorization with require, and add a denial test that must fail.

If agents report rate limit text but your dashboards show no 429s, the limits work as designed. MCP calls get isError or -32003 inside an HTTP 200. Read tool errors in the agent loop, and count reason=RateLimit in the logs.

HTTP 400 invalid session ID header on half of all calls means the replicas have different session.key values. Give every replica one key from one secret. HTTP 404 Session not found after a deploy means the stateful upstream restarted or moved, and its session died with it. The client has to open a new session, and 2026-07-28 removes the problem. If the upstream sees about 10 requests per old-protocol call, statefulMode: stateless is on. Use session keys instead, or move clients to 2026-07-28.

When every tool list and every new session fails while one server is down, the federation runs with the default failClosed. Set failureMode: failOpen where a partial list is acceptable. When the first model request after an outage returns 503, failover evicted the primary on that failure but did not retry the request. Retry in the client, or test a route retry policy.

A2A traffic that reaches an agent with no gateway log line usually means a dual agent card left the v0.3 url unrewritten. Publish one endpoint field per card, and block direct network paths. A startup warning about config.tracing means you use the deprecated field, so move to frontendPolicies.tracing. JWTs that worked before an upgrade and now return 401 point at the issuer and audience checks from step 4.

Send the call, read the answer

A gateway turns twenty client configs into one place that decides who can call what. On the paths I measured, agentgateway does that job well, with well under a millisecond added at the median and 36 MB of memory. The problems sit at the edges: a denied tool looks missing, and a rate limit looks like success. One dead server can hide the healthy ones, and one argument rule can do nothing while it looks correct.

None of these show up in a feature table, but all of them show up when you send the call and read the answer. To compare gateways, start with Christian Posta's "Can You Use an API Gateway as an MCP Gateway?". It explains why MCP needs a gateway that understands the protocol. Then run the steps of this guide against each candidate, agentgateway included. Do that for every policy before you trust it.

rg
Rohit Ghumare

CNCF Ambassador, AAIF Ambassador, and Google Developer Expert. agentgateway is an AAIF project, so read this guide with that in mind. I have no role in the agentgateway project and no relationship with Solo.io. Every number comes from agentgateway 1.6.0 on my laptop, run on October 5, 2026.

Related: Stateless MCP · All About the A2A Protocol · More posts · X