Commit Graph

2 Commits

Author SHA1 Message Date
JP
34afc497f4 Give members their own budget-capped gateway key
Members share the owner's AI budget, so each now gets a Switchboard key
minted on first use with a daily cap and a per-request ceiling. That
gives real spend limits and per-user attribution: the app's own rate
limiter lives in memory and resets on every deploy, so it could never be
a spend control. The gateway enforces the caps and answers 402
guardrail, which the error mapper already turns into a budget message.

Resolution order is the user's own key, then mint, then borrow the
owner's key for a single request if the gateway is unreachable - a
served request beats a hard failure, and the fallback logs loudly
because no per-user cap applies to it. The owner key is used only for
minting, never for inference.

API key management is now owner-only, enforced on GET, POST and DELETE.
GET matters as much as POST because it returns the masked key and the
gateway URL. Settings became a server component so the role is known
before first render: members never see the card and never issue the
request, rather than having it flash and disappear.

AiCall records one row per request from the gateway's own response
metadata, so member spend is queryable instead of a journald grep. The
write is fire-and-forget - tracking must never fail a working request.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 21:16:14 +00:00
JP
a0e1619072 Route all AI features through the Switchboard gateway
Replace the direct Anthropic and OpenAI integrations with a single
provider that talks to Switchboard, an OpenAI-compatible gateway that
routes each request to the best available model. The app no longer pins
a model id anywhere: it sends switchboard/auto and lets the gateway
choose, then logs which model answered and what it cost.

Routing levers are set per feature in src/lib/ai/routing.ts. Three of
those choices came from measuring against the live gateway:

- category and prefer_free are set explicitly on every request. An API
  key carries its own routing defaults, and anything left unset inherits
  them - drink prompts were being sent to a free coding model.
- Token budgets are generous because the router may pick a reasoning
  model, and reasoning tokens come out of the same max_tokens budget as
  the answer. At 512 tokens a request returned null content; at 4096 the
  same request returned correct JSON.
- No tier lever on text features. tier "cheap" pinned a slow reasoning
  model (42-180s, two timeouts and one truncated response in five
  trials) and tier "frontier" escalated as far as Opus at $0.02 a call,
  while unconstrained routing answered in about a second. Vision keeps
  "frontier", where the accuracy is worth a few tenths of a cent.

Gateway failures are mapped to actionable messages rather than passed
through: a 401 relayed as 401 would read as an expired session and
bounce the user to login, and a 429 would collide with the app's own
rate limiter.

Also collapses the key lookup that was duplicated across ten call sites
into getUserProvider(), which fixes a latent bug where a bare findFirst
with no ordering let different features pick different providers.

Existing claude/openai key rows are ignored at runtime and offered for
removal in Settings, so no migration is needed before deploying.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 16:41:00 +00:00