Route all AI features through the Switchboard gateway

Replace the direct Anthropic and OpenAI integrations with a single
provider that talks to Switchboard, an OpenAI-compatible gateway that
routes each request to the best available model. The app no longer pins
a model id anywhere: it sends switchboard/auto and lets the gateway
choose, then logs which model answered and what it cost.

Routing levers are set per feature in src/lib/ai/routing.ts. Three of
those choices came from measuring against the live gateway:

- category and prefer_free are set explicitly on every request. An API
  key carries its own routing defaults, and anything left unset inherits
  them - drink prompts were being sent to a free coding model.
- Token budgets are generous because the router may pick a reasoning
  model, and reasoning tokens come out of the same max_tokens budget as
  the answer. At 512 tokens a request returned null content; at 4096 the
  same request returned correct JSON.
- No tier lever on text features. tier "cheap" pinned a slow reasoning
  model (42-180s, two timeouts and one truncated response in five
  trials) and tier "frontier" escalated as far as Opus at $0.02 a call,
  while unconstrained routing answered in about a second. Vision keeps
  "frontier", where the accuracy is worth a few tenths of a cent.

Gateway failures are mapped to actionable messages rather than passed
through: a 401 relayed as 401 would read as an expired session and
bounce the user to login, and a 429 would collide with the app's own
rate limiter.

Also collapses the key lookup that was duplicated across ten call sites
into getUserProvider(), which fixes a latent bug where a bare findFirst
with no ordering let different features pick different providers.

Existing claude/openai key rows are ignored at runtime and offered for
removal in Settings, so no migration is needed before deploying.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
JP
2026-08-08 16:41:00 +00:00
parent 7c41b15ecc
commit a0e1619072
31 changed files with 830 additions and 429 deletions

View File

@@ -3,6 +3,7 @@ import { auth } from "@/lib/auth"
import { prisma } from "@/lib/prisma"
import { encrypt, decrypt, maskApiKey } from "@/lib/encryption"
import { apiKeySchema } from "@/lib/validators"
import { switchboardBaseUrl } from "@/lib/ai/switchboard-provider"
export async function GET() {
const session = await auth()
@@ -44,7 +45,9 @@ export async function GET() {
}
})
return NextResponse.json(maskedKeys)
// The gateway endpoint is server config, so surface it here rather than making the
// user guess which Switchboard instance this deployment points at.
return NextResponse.json({ keys: maskedKeys, gatewayUrl: switchboardBaseUrl() })
}
export async function POST(request: Request) {
@@ -90,6 +93,15 @@ export async function POST(request: Request) {
},
})
// Keys from the old direct Claude/OpenAI integration are already ignored when
// selecting a provider. Clear them now that a working replacement exists.
await prisma.userApiKey.deleteMany({
where: {
userId: session.user.id,
provider: { in: ["claude", "openai"] },
},
})
return NextResponse.json({
id: key.id,
provider: key.provider,