Skip to content

Securing grpc.aio with TLS and Auth Metadata

A gRPC service needs two things to be secure: an encrypted, authenticated connection (TLS, so clients know they are talking to the real server and nobody can read the traffic), and per-call credentials (a token in the call's metadata, so the server knows who is calling). Tested with grpcio 1.84 and a private CA: a call over TLS without a token was rejected by a server interceptor with UNAUTHENTICATED; with a bearer token attached through call credentials it succeeded; a client that did not trust the CA, or that expected a different host name, failed during the TLS handshake with UNAVAILABLE. The cost was small: the first call on a new TLS channel took a median 2.46 ms against 0.93 ms in plaintext, and steady-state calls 430 µs against 269 µs (the TLS server also ran the auth interceptor). A plaintext channel with the same token in metadata also "worked" — with the token readable by anyone on the path. This guide sets up both layers.

Prerequisites

1. Serve over TLS

Load the server's private key and certificate chain and bind a secure port:

import grpc


def read(path: str) -> bytes:
    with open(path, "rb") as f:
        return f.read()


async def serve() -> None:
    server = grpc.aio.server(interceptors=[AuthInterceptor()])
    greet_pb2_grpc.add_GreeterServicer_to_server(Greeter(), server)
    credentials = grpc.ssl_server_credentials(
        [(read("server.key"), read("server.pem"))],     # key, then certificate (chain)
    )
    server.add_secure_port("[::]:50051", credentials)
    await server.start()
    await server.wait_for_termination()

The certificate's subject alternative names must include the name clients use — localhost, greeter.internal, or the IP — or clients will reject it. In production the certificate usually comes from your platform: cert-manager in Kubernetes, a cloud certificate service, or a service mesh that terminates TLS in a sidecar, in which case the application listens in plaintext on localhost only. Reload certificates before they expire; grpcio supports dynamic server credentials through a certificate configuration fetcher for rotation without restarts.

Verify: openssl s_client -connect host:50051 -alpn h2 shows the certificate chain and negotiates h2.

2. Connect with TLS and verify the server

Clients trust a CA and check the server's name against the certificate:

channel_credentials = grpc.ssl_channel_credentials(root_certificates=read("ca.pem"))

async with grpc.aio.secure_channel("localhost:50051", channel_credentials) as channel:
    stub = greet_pb2_grpc.GreeterStub(channel)
    ...

Tested failures: with the system trust store instead of the private CA, the call failed with UNAVAILABLE and a TLS handshake error; with a host name override that did not match the certificate, it failed with UNAVAILABLE and a verification error. Both are what you want — a client that silently accepted either would be open to impersonation. Omit root_certificates only for servers with publicly trusted certificates. For mutual TLS, where the server also verifies the client's certificate, pass private_key and certificate_chain here and root_certificates plus require_client_auth=True on the server.

Verify: a client configured with the wrong CA cannot connect, and the error mentions the handshake.

What each misconfiguration produced A grid of 5 rows by 3 columns. What each misconfiguration produced client setup result meaning TLS, no token UNAUTHENTICATED interceptor rejected the call TLS + bearer token OK authenticated and encrypted TLS, untrusted CA UNAVAILABLE (handshake) server not verified TLS, wrong host name UNAVAILABLE (handshake) name not in certificate plaintext + token in metadata OK works, token sent in clear grpcio 1.84 with a private CA; the last row is the one to forbid.

3. Attach tokens with call credentials

Per-call credentials put a token in the authorization metadata of every call. Combine them with the channel's TLS credentials so they are sent only over encrypted channels:

call_credentials = grpc.access_token_call_credentials(token)        # "authorization: Bearer <token>"
credentials = grpc.composite_channel_credentials(channel_credentials, call_credentials)

async with grpc.aio.secure_channel("localhost:50051", credentials) as channel:
    reply = await greet_pb2_grpc.GreeterStub(channel).Hello(greet_pb2.HelloRequest(name="x"))

For tokens that expire, use grpc.metadata_call_credentials with a plugin that fetches or refreshes the token and calls back with the metadata; it runs before every call. Passing metadata=[("authorization", ...)] on each call works too, and is what the plaintext test did — but nothing then stops it from being sent on an insecure channel. Composite credentials make "token only over TLS" a property of the channel rather than of every call site.

Verify: inspect a call on the server: context.invocation_metadata() contains the bearer token, and the channel is TLS (context.auth_context() shows the transport security type).

4. Check tokens in a server interceptor

Authentication belongs in one place, before any servicer runs. A server interceptor reads the metadata and rejects unauthenticated calls:

class AuthInterceptor(grpc.aio.ServerInterceptor):
    def __init__(self, verify) -> None:
        self.verify = verify                                   # token -> principal, or None

    async def intercept_service(self, continuation, handler_call_details):
        metadata = dict(handler_call_details.invocation_metadata or ())
        header = metadata.get("authorization", "")
        principal = self.verify(header[7:]) if header.startswith("Bearer ") else None
        if principal is None:
            async def deny(request, context):
                await context.abort(grpc.StatusCode.UNAUTHENTICATED, "missing or invalid token")
            return grpc.unary_unary_rpc_method_handler(deny)
        CURRENT_PRINCIPAL.set(principal)                       # contextvar for the servicer
        return await continuation(handler_call_details)

Tested: a call without a token received UNAUTHENTICATED with the interceptor's message, and the servicer never ran. The deny handler shown is for unary-unary methods; for streaming methods return the matching handler type, or check the method name in handler_call_details.method to pick one. Exempt the health service, so probes work without credentials, and distinguish UNAUTHENTICATED (no valid identity) from PERMISSION_DENIED (valid identity, not allowed) for authorization checks inside servicers.

Verify: every RPC except health checks returns UNAUTHENTICATED without a token, and the servicer logs show the principal for authenticated calls.

Two layers of gRPC security on one call A flow of 4 stages. Two layers of gRPC security on one call TLS handshake verify CA + host name call credentials authorization metadata server interceptor validate token servicer runs with a principal TLS proves the server to the client; the token proves the client to the server.

5. Keep the cost low

TLS adds a handshake per connection and encryption per byte. Both are small when connections are reused:

# One long-lived channel per target, created at startup and shared
@asynccontextmanager
async def lifespan(app):
    app.state.greeter = grpc.aio.secure_channel(
        "greeter.internal:50051",
        grpc.composite_channel_credentials(channel_credentials, call_credentials),
        options=[("grpc.keepalive_time_ms", 30_000)],
    )
    yield
    await app.state.greeter.close()

Measured over 20 fresh channels each: the first call took 2.46 ms with TLS against 0.93 ms in plaintext — the handshake — and steady-state calls 430 µs against 269 µs, a difference that includes the auth interceptor's work on the TLS server. A channel created per request would pay the handshake every time; a shared channel pays it once. Token validation can be the larger cost if it calls an identity provider on every request — cache verified tokens until shortly before their expiry.

Verify: the number of TLS handshakes per second on the server (from metrics or connection counts) is close to the rate of new client processes, not the rate of calls.

Latency with and without TLS 4 horizontal bars comparing first call, TLS with the others. Latency with and without TLS first call, TLS 2.46 ms first call, plaintext 0.93 ms steady call, TLS + auth 0.43 ms steady call, plaintext 0.27 ms grpcio 1.84 on localhost; medians over 20 channels, 200 calls each. Reuse channels and the handshake is paid once.

Verification

A grpc.aio service is secured when:

  • It serves only over TLS, with certificates whose names match how clients connect.
  • Clients verify the server against a specific CA and host name.
  • Tokens travel as composite call credentials, never on insecure channels.
  • A server interceptor authenticates every call, with health checks exempt.

Diagnostic Hook: count UNAUTHENTICATED and PERMISSION_DENIED responses per calling service and handshake failures per client. A sudden rise in handshake failures across clients points at certificate rotation or expiry; a rise in UNAUTHENTICATED from one caller points at its token refresh.

Pitfalls & edge cases

  • Tokens on plaintext channels. Tested: it works, and the token is exposed.
  • Disabling verification to fix handshake errors. Fix the CA or the certificate's names instead.
  • Authentication in each servicer. One missed method is an open endpoint; use an interceptor.
  • A channel per request. Every call pays the TLS handshake.

Frequently Asked Questions

How do I enable TLS on a grpc.aio server?

Create grpc.ssl_server_credentials with the private key and certificate chain and bind with server.add_secure_port. Clients use grpc.ssl_channel_credentials with the CA certificate and grpc.aio.secure_channel.

How do I send a bearer token with grpc.aio calls?

Combine the channel's TLS credentials with grpc.access_token_call_credentials(token) using grpc.composite_channel_credentials, so every call carries the authorization metadata over TLS.

How do I check authentication on a gRPC server?

Add a grpc.aio.ServerInterceptor that reads the authorization metadata and returns a handler that aborts with UNAUTHENTICATED when the token is missing or invalid.

How much does TLS slow down gRPC?

In testing, the first call on a new channel took 2.46 ms with TLS versus 0.93 ms in plaintext; on a reused channel the difference was under 0.2 ms per call.