Route all AI features through the Switchboard gateway
Replace the direct Anthropic and OpenAI integrations with a single provider that talks to Switchboard, an OpenAI-compatible gateway that routes each request to the best available model. The app no longer pins a model id anywhere: it sends switchboard/auto and lets the gateway choose, then logs which model answered and what it cost. Routing levers are set per feature in src/lib/ai/routing.ts. Three of those choices came from measuring against the live gateway: - category and prefer_free are set explicitly on every request. An API key carries its own routing defaults, and anything left unset inherits them - drink prompts were being sent to a free coding model. - Token budgets are generous because the router may pick a reasoning model, and reasoning tokens come out of the same max_tokens budget as the answer. At 512 tokens a request returned null content; at 4096 the same request returned correct JSON. - No tier lever on text features. tier "cheap" pinned a slow reasoning model (42-180s, two timeouts and one truncated response in five trials) and tier "frontier" escalated as far as Opus at $0.02 a call, while unconstrained routing answered in about a second. Vision keeps "frontier", where the accuracy is worth a few tenths of a cent. Gateway failures are mapped to actionable messages rather than passed through: a 401 relayed as 401 would read as an expired session and bounce the user to login, and a 429 would collide with the app's own rate limiter. Also collapses the key lookup that was duplicated across ten call sites into getUserProvider(), which fixes a latent bug where a bare findFirst with no ordering let different features pick different providers. Existing claude/openai key rows are ignored at runtime and offered for removal in Settings, so no migration is needed before deploying. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
35
src/lib/ai/switchboard-log.ts
Normal file
35
src/lib/ai/switchboard-log.ts
Normal file
@@ -0,0 +1,35 @@
|
||||
import type { SwitchboardMeta } from "./switchboard-types"
|
||||
|
||||
/**
|
||||
* One line per gateway call so the cost and the model actually used are visible in
|
||||
* the server log. Called from inside the provider, so every feature gets it for free.
|
||||
*/
|
||||
export function logSwitchboardMeta(
|
||||
feature: string,
|
||||
meta: SwitchboardMeta | null
|
||||
): void {
|
||||
if (!meta) {
|
||||
console.warn(`[switchboard] feature=${feature} no meta block in response`)
|
||||
return
|
||||
}
|
||||
|
||||
console.log(
|
||||
`[switchboard] feature=${feature} model=${meta.model_id} ` +
|
||||
`provider=${meta.provider} locality=${meta.locality} category=${meta.category} ` +
|
||||
`cost=${meta.cost_usd ?? "?"} latency_ms=${meta.latency_ms} request_id=${meta.request_id}`
|
||||
)
|
||||
|
||||
// Both of these silently degrade output quality, so they warn rather than log.
|
||||
if (meta.failover) {
|
||||
console.warn(
|
||||
`[switchboard] feature=${feature} FAILOVER intended=${meta.intended_model} ` +
|
||||
`actual=${meta.model_id} reason=${meta.reason}`
|
||||
)
|
||||
}
|
||||
if (meta.context_overflow) {
|
||||
console.warn(
|
||||
`[switchboard] feature=${feature} CONTEXT OVERFLOW model=${meta.model_id} ` +
|
||||
`- the provider may have truncated this request`
|
||||
)
|
||||
}
|
||||
}
|
||||
Reference in New Issue
Block a user