Route all AI features through the Switchboard gateway
Replace the direct Anthropic and OpenAI integrations with a single provider that talks to Switchboard, an OpenAI-compatible gateway that routes each request to the best available model. The app no longer pins a model id anywhere: it sends switchboard/auto and lets the gateway choose, then logs which model answered and what it cost. Routing levers are set per feature in src/lib/ai/routing.ts. Three of those choices came from measuring against the live gateway: - category and prefer_free are set explicitly on every request. An API key carries its own routing defaults, and anything left unset inherits them - drink prompts were being sent to a free coding model. - Token budgets are generous because the router may pick a reasoning model, and reasoning tokens come out of the same max_tokens budget as the answer. At 512 tokens a request returned null content; at 4096 the same request returned correct JSON. - No tier lever on text features. tier "cheap" pinned a slow reasoning model (42-180s, two timeouts and one truncated response in five trials) and tier "frontier" escalated as far as Opus at $0.02 a call, while unconstrained routing answered in about a second. Vision keeps "frontier", where the accuracy is worth a few tenths of a cent. Gateway failures are mapped to actionable messages rather than passed through: a 401 relayed as 401 would read as an expired session and bounce the user to login, and a 429 would collide with the app's own rate limiter. Also collapses the key lookup that was duplicated across ten call sites into getUserProvider(), which fixes a latent bug where a bare findFirst with no ordering let different features pick different providers. Existing claude/openai key rows are ignored at runtime and offered for removal in Settings, so no migration is needed before deploying. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
31
README.md
31
README.md
@@ -20,6 +20,37 @@ You can start editing the page by modifying `app/page.tsx`. The page auto-update
|
||||
|
||||
This project uses [`next/font`](https://nextjs.org/docs/app/building-your-application/optimizing/fonts) to automatically optimize and load [Geist](https://vercel.com/font), a new font family for Vercel.
|
||||
|
||||
## AI Gateway (Switchboard)
|
||||
|
||||
All AI features — menu scanning, label identification, drink search, the bartender
|
||||
and the recommendation engine — go through [Switchboard](http://192.168.2.11:8787/v1/guide),
|
||||
an OpenAI-compatible gateway that routes each request to the best available model.
|
||||
The app never pins a model id; it always sends `switchboard/auto` and lets the gateway
|
||||
choose, then logs which model answered and what it cost.
|
||||
|
||||
Setup:
|
||||
|
||||
1. Set `SWITCHBOARD_BASE_URL` in your env file (defaults to `http://192.168.2.11:8787/v1`).
|
||||
2. Mint an API key in the Switchboard UI under Settings → API keys.
|
||||
3. Add that key in the app under Settings → AI Gateway.
|
||||
|
||||
Per-feature routing (cost/quality levers, token budgets, timeouts) lives in
|
||||
`src/lib/ai/routing.ts`. Note that a Switchboard key carries its own routing defaults,
|
||||
so the app sets `category` and `prefer_free` explicitly on every request rather than
|
||||
inheriting whatever the key was minted for.
|
||||
|
||||
### Migrating from the old Claude/OpenAI integration
|
||||
|
||||
Earlier versions stored a per-user Anthropic or OpenAI key. Those rows are ignored at
|
||||
runtime and the Settings page offers to remove them, so no migration is required. To
|
||||
clear them in bulk instead:
|
||||
|
||||
```sql
|
||||
DELETE FROM "UserApiKey" WHERE provider IN ('claude','openai');
|
||||
DELETE FROM "SearchCache" WHERE provider IN ('claude','openai');
|
||||
UPDATE "UserPreference" SET "defaultProvider" = NULL;
|
||||
```
|
||||
|
||||
## Learn More
|
||||
|
||||
To learn more about Next.js, take a look at the following resources:
|
||||
|
||||
Reference in New Issue
Block a user