Members share the owner's AI budget, so each now gets a Switchboard key minted on first use with a daily cap and a per-request ceiling. That gives real spend limits and per-user attribution: the app's own rate limiter lives in memory and resets on every deploy, so it could never be a spend control. The gateway enforces the caps and answers 402 guardrail, which the error mapper already turns into a budget message. Resolution order is the user's own key, then mint, then borrow the owner's key for a single request if the gateway is unreachable - a served request beats a hard failure, and the fallback logs loudly because no per-user cap applies to it. The owner key is used only for minting, never for inference. API key management is now owner-only, enforced on GET, POST and DELETE. GET matters as much as POST because it returns the masked key and the gateway URL. Settings became a server component so the role is known before first render: members never see the card and never issue the request, rather than having it flash and disappear. AiCall records one row per request from the gateway's own response metadata, so member spend is queryable instead of a journald grep. The write is fire-and-forget - tracking must never fail a working request. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
13 KiB
13 KiB