One gateway for model lookup, spend, and inference calls from the terminal
Connect Vercel AI Gateway and Luumen can read your model catalog, per-provider endpoint pricing and latency, credit balance, and spend reports on demand. When you want a call made — a chat completion, an embedding batch, an image generation — Luumen shows the model, the routing, and the request first, then waits for your approval.
22 tools: 9 read, 13 write. Reads answer instantly. Writes require approval by default. Everything is logged.
What a governed Vercel AI Gateway run looks like inside Luumen.
Authorize once with API token. Luumen lists the scopes each action needs before you approve the connection, and credentials never appear in the chat.
Read actions answer immediately. Anything that writes — compact response, create anthropic message, create chat completion, create embeddings, and more — is shown as a plan and requires approval by default. Administrators configure that per tool, so you decide exactly which actions can ever run unattended.
You decide. Actions are granted per agent, skill, and team, and per environment — production is not staging. Read access can be broad while writes stay narrow.
Every call to Vercel AI Gateway — read or write, approved or declined — is recorded with the actor, the input, and the result, and can be linked to the ticket or change record.
Connect in minutes. Every action scoped, approved, and audited from day one.