Commit Graph

41 Commits

Author SHA1 Message Date
JP
7ff3f68b2c Add a backlog, consolidating deferred work that was only in commit messages
Several decisions had been consciously postponed - the Docker purge date, the
unencrypted dumps, the unestablished Docker-in-LXC cause - but they lived in
commit messages and chat, which is the same as nowhere.

Each entry records what would trigger picking it up, so none of it has to be
re-derived. Starts with the MCP tool-consolidation question: the answer is that
merging CRUD into one tool would cost the annotations, the scope boundary, the
schema clarity and the audit trail, and the useful rule is to merge only when
it is the same gesture with a different parameter.

Nothing here is urgent; that is the point of writing it down rather than doing it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1Ee4Mc1X1SX8HgYa52zu7
2026-08-10 01:22:05 +00:00
JP
a8c959bfab Add a saved recipe to the collection or the wishlist so it can be rated
A recipe was previously a dead end: you could save one and see whether your bar
could make it, but there was no way to record that you had actually drunk it.

Promoting links rather than converts. A drink is created and Recipe.sourceDrinkId
is set to point at it, so the recipe survives and then renders on that drink's
page through the relation the drink detail view already reads. The wishlist path
creates a "try later" entry tagged source=recipe and leaves the recipe alone.

Both are idempotent. Promoting an already-linked recipe returns the existing
drink with a 200 rather than creating a second one, because pressing the button
twice is a double click and not a request for a duplicate. Creation and linking
happen in one transaction - a drink that existed but was not linked back would
show no recipe and would be promoted again on the next attempt.

The logic lives in src/lib/recipe-promotion.ts and is shared by the REST route
and the MCP tool, so the ownership check, the idempotency rule and the generated
description cannot drift apart.

Adding to drinks navigates straight to the rating page: being able to rate it is
the entire point, and an unrated cocktail is easy to lose in a list of 88. There
is no Toaster mounted in this app, so the wishlist path reports inline instead.

One MCP tool with a target parameter rather than two, since the tool list is
already at the point where selection quality suffers - 23 now. Gated on
drinks:write, not bar:write: it changes the collection, not the bar.

Verified: promote to drink then rate it, idempotency on both targets, the
recipe stays linked and undeleted, invalid target, unknown recipe, unauthenticated
access, read-only tokens, and that another user's recipe is refused identically
on both the REST route and the tool without confirming the row exists.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1Ee4Mc1X1SX8HgYa52zu7
2026-08-09 21:30:40 +00:00
JP
23c4e63a68 Fix all five findings from the independent OAuth security audit
An independent model reviewed the MCP and OAuth surfaces. All five findings
were verified against the code before changing anything; none were false
positives. audit.md is kept as the record of what was reviewed.

F1 (High) - refresh reuse detection forgot the token family. Only the current
hash and one predecessor lived on the grant, and each rotation overwrote the
predecessor. A thief who rotated a stolen token twice made the victim's
original unrecognisable: replaying it returned "unknown token" instead of
revoking the family, and the thief kept working. Refresh tokens are now rows in
OAuthRefreshToken, one per generation, retained for the life of the family and
consumed by a guarded update on usedAt - which also means two concurrent uses
of the same token can no longer both succeed. This corrects a claim I made when
the OAuth server shipped: reuse detection covered one generation, not the family.

F2 (Medium) - issueAccessToken wrote scopes back onto the grant, so redeeming a
stale authorization code redefined standing consent. Token issuance is not
consent; the consent endpoint is now the only writer. Codes are additionally
bound to a grant id and epoch, with a coversScopes check behind that.

F3 (Medium) - revocation was reversible. Reconnecting a disconnected app cleared
revokedAt and left credentials that had raced the revoke usable again. Every
approval now starts a clean epoch: the counter advances and prior access tokens,
refresh tokens and unconsumed codes are destroyed. Token writes are conditional
on the epoch they validated, so a revoke that wins a race aborts them. Refresh
also now requires offline_access to still be granted.

F4 (Medium) - loopback redirect matching compared only scheme, host and path,
silently accepting a differing query, fragment or userinfo. RFC 9700 2.1 wants
exact matching apart from the RFC 8252 port exception; that is what it does now.

F5 (Low) - get_collection_stats returned bar and recipe counts under
drinks:read. Gated on the caller actually holding bar:read.

Verified with regression tests for each: the two-rotation attack now revokes the
family, concurrent refresh yields exactly one winner, a pre-narrowing code is
refused, disconnect-reconnect leaves old credentials dead, and Claude Code's
ephemeral-port callback still works while query/userinfo/fragment variants are
rejected. Existing protections re-checked - code replay, PKCE mismatch, deny,
confidential-client rejection, and the MCP tools themselves.

Note for deploy: OAuthAuthCode gains a required grantId, so existing rows must
be cleared first. They are 60-second ephemeral codes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1Ee4Mc1X1SX8HgYa52zu7
2026-08-09 21:05:08 +00:00
JP
5c05464269 Retract the unprivileged-LXC claim in the Docker post-mortem
/proc/self/uid_map reads 0 0 4294967295 - an identity mapping - so this
container is privileged. The AppArmor and sysctl failures were real and
reproduced, but "because unprivileged" was never the diagnosis, and leaving it
in the doc invites someone to build on a false premise.

Says what is actually known and marks the root cause as unestablished rather
than substituting a fresh guess.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1Ee4Mc1X1SX8HgYa52zu7
2026-08-09 19:24:03 +00:00
JP
1710e6147b Document the off-site backup leg now that it exists
Nightly backups now copy to TrueNAS over SSH rather than living only on the
LXC they are protecting. Chose scp over the existing NFS share deliberately:
a mount that goes stale writes into the local mountpoint instead, silently
filling the LXC disk while every run reports success. No mount, no such class
of failure - and the backup script already spoke scp, so nothing changed but
two Environment lines.

Set as a systemd drop-in so re-installing the service cannot quietly return to
local-only backups.

Recorded the failure mode that cost time here: sshd refuses authorized_keys
when the user's home directory is group- or world-writable, which TrueNAS
datasets often are. The key is offered and silently rejected, which looks
exactly like a key that was never added.

Verified end to end rather than assumed: the copy lands, the dump passes
gzip -t, contains 25 CREATE TABLE statements including the new MCP and OAuth
tables, and the MinIO tarball reads back with 194 entries.

Still open: the dumps are unencrypted, and TrueNAS is on the same site.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1Ee4Mc1X1SX8HgYa52zu7
2026-08-09 19:20:16 +00:00
JP
82680d9430 Correct stale deploy docs and document the new env knobs
deploy/README.md described a rewrites() block that no longer exists, and told
the reader to verify the image bucket's anonymous-download policy was set -
which is precisely the exposure that was closed. Following it would have
re-opened every stored image to the public internet.

Replaced with what is actually true: /minio-images is a session-gated route
handler with per-user key-prefix ownership, the bucket must have no anonymous
policy, and MinIO must stay bound to loopback. Also noted that the suggested
`mc anonymous get` check does not work on this host - the alias has no
credentials and returns Access Denied either way - and gave an external curl
that actually tells you something.

The same stale rationale sat in deploy.sh's build comment. Building without the
env file is still right, just for a different reason: nothing needs baking in.

Added the MCP and OAuth journal greps, an unauthenticated discovery check, and
a note that a 307 there means /.well-known fell out of PUBLIC_ROUTES. Documented
in .env.example that NEXTAUTH_URL must be the exact public origin with no
trailing slash, since it is now the OAuth issuer and MCP resource identifier,
plus the MCP_CHATGPT_TOOLS opt-out.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1Ee4Mc1X1SX8HgYa52zu7
2026-08-09 19:04:42 +00:00
JP
e77f605c9c Add OAuth 2.1 server so claude.ai and ChatGPT can connect natively
Phase A needed a token pasted into a header, which Anthropic documents as an
org-admin-scoped beta on claude.ai and is undocumented on ChatGPT. This makes
the app its own authorization server so both connect through their normal
"add a connector" flow, with a per-user consent screen.

The whole thing funnels into the existing verifyMcpToken: an issued access
token is an ordinary McpAccessToken row with source "oauth", so the resource
server gained no OAuth awareness and Phase A's verification path is unchanged.

Notes on the parts that fail quietly if got wrong:

- scopes_supported is published in the protected resource metadata. When a 401
  challenge carries no explicit scope Claude requests exactly what is
  advertised there, so omitting it silently makes every connection read-only.
- Loopback redirect URIs match with the port ignored. Claude Code registers
  http://localhost/callback and then redirects to an ephemeral port; exact
  matching would reject every native client. RFC 8252 7.3 requires this.
- CIMD is deliberately not advertised. Claude only selects it when the metadata
  carries client_id_metadata_document_supported, and supporting it would mean
  fetching a client-supplied URL server-side from a host that also reaches
  MinIO, Postgres and the gateway on the LAN. DCR costs nothing by comparison.
- The consent page validates client_id and redirect_uri before it will redirect
  anywhere, because an unvalidated redirect_uri is an open redirect.
- Authorization codes are consumed by one guarded updateMany, so a replay under
  concurrency cannot mint a second token. A PKCE mismatch burns the whole grant
  rather than just the code.
- Refresh tokens rotate and keep the previous hash; presenting it revokes the
  grant, since that is either theft or a client that cannot be trusted to hold
  state.

Also fixes the login page, which hardcoded callbackUrl and so dropped anyone
sent to sign in for the consent screen onto /dashboard instead. Same-origin
destinations only, absolute or relative - the middleware writes absolute.

"/.well-known" joins PUBLIC_ROUTES. It is a prefix match, so nothing else
should be served from there.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1Ee4Mc1X1SX8HgYa52zu7
2026-08-09 05:33:48 +00:00
JP
4fd339dabc Add MCP server so Claude and ChatGPT can read and write drink data
Exposes the collection over the Model Context Protocol at /api/mcp, with 22
tools covering drinks, ratings, bar inventory, recipes, wishlist and taste
preferences, plus search/fetch aliases for ChatGPT's deep-research mode.

Authentication is a bearer token, the app's first header-borne credential -
every other route derives identity from the NextAuth cookie, which a machine
client cannot present. /api/mcp sits under the middleware's /api exclusion so
it can answer a JSON 401 with an RFC 9728 WWW-Authenticate challenge instead
of an HTML redirect to /login.

Tokens are stored as a SHA-256 hash rather than plaintext like Invite.token
and PasswordReset.token. Those are single-use and short-lived; this one is
long-lived and grants read/write over a whole collection, and the nightly
pg_dump keeps 14 days of history. Not encrypt(), which is reversible AES and
right only for outbound keys we must replay; not bcrypt, which cannot be
indexed and would turn verification into a table scan per request.

verifyMcpToken joins User.status on every call, mirroring the jwt callback, so
suspending a member kills their MCP access immediately rather than leaving the
token as a documented way to outlive suspension. It fails closed on a database
error, deliberately unlike the jwt callback, which keeps the session because a
throw there would sign out every user at once.

No tool reaches the Switchboard gateway. Claude and ChatGPT are language
models already, so they can reason over a bar inventory without the app paying
to do it a second time, and a remote client looping a vision call is not a
failure mode worth having. Account deletion, restore, gateway keys, admin
routes and shared-list creation are excluded too.

The OAuth models ship now but are unused; the token endpoint will write the
same McpAccessToken rows, so adding it later touches no verification code.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1Ee4Mc1X1SX8HgYa52zu7
2026-08-09 04:29:54 +00:00
JP
13793d43ca Make MenuItem.userId required
Backfilled from the parent scan (176/176 rows), so the column can now be
NOT NULL rather than relying on every writer remembering to set it. The
constraint was applied on the server alongside this, so the schema and
the database stay in step for the next db push.

Also rewrites two MenuScan rows that stored absolute
http://localhost:9000 URLs from before the app used the /minio-images
proxy. Those resolved against the viewer's own machine, so they have
always been broken images for anyone not running MinIO locally, and are
unreachable now that MinIO is bound to loopback.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 22:03:11 +00:00
JP
70db6314e3 Phase 5 hygiene: image ownership, real deletes, scan resilience
Images were readable by any signed-in user, not just their owner. Keys
are now namespaced per user in both shapes the app writes, the three
bar/ objects that predated that were migrated under the owner's prefix,
and the image route enforces the prefix - answering 404 rather than 403
so it does not confirm someone else's key exists. Barcode-mirrored
images now write under the user prefix too.

deleteImage existed but was never called, so deleting a drink, bar item
or scan removed the row and left the file in storage forever, still
fetchable. All three now clean up through one shared helper, after the
database work and best-effort, so a storage hiccup cannot block a delete
the user asked for or leave the row behind.

Menu scans are processed detached from the request, so every deploy
abandoned anything in flight and left the row PROCESSING forever with a
spinner that never resolved. They are now reaped lazily when a user
lists their scans - no scheduler needed. Concurrent extractions are
capped at two per process: each is a vision call with a four minute
timeout that costs real money, and nothing else bounded them.

MenuItem gains userId. Ownership was only ever transitive via scanId,
which held because nothing queries MenuItem directly, but left any
future direct query an IDOR with nothing to stop it.

Restore was correctly scoped but unbounded, so a crafted file could
create unlimited rows in one transaction against shared Postgres - self
harm with one user, denial of service with several. Capped, and imageUrl
from the CSV now goes through the same validation the API enforces
instead of reaching the column unchecked.

Members can delete their own account and data. Once the app holds other
people's history, including Rating.location, that is the minimum.

Registration rate limiting took the first x-forwarded-for hop, which the
client controls; it now takes the last, which our proxy appends.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 21:57:57 +00:00
JP
0ead1385f7 Add the owner admin area: invites, people, reset links
Creating an invite previously meant writing SQL by hand, which made the
whole feature unusable in practice. The owner can now create invites
with a use count and expiry, copy the link, show a QR code, and revoke.

QR codes are generated locally by node-qrcode as inline SVG. The CSP
forbids loading from any other origin, so an external generator was not
an option, and img-src 'self' already covers a same-origin SVG.

People lists everyone with what they have added and what their AI use
has cost over 30 days, with suspend, reactivate, per-user AI toggle and
delete. Deleting collects image keys and the minted gateway key id
before the cascade removes the rows, then cleans both up best-effort -
storage or the gateway being unavailable must not leave an account
half-deleted. Guards refuse to suspend or delete yourself or the last
owner.

Password reset links close the gap that came with keeping email and
password sign-in: with no email infrastructure, a member who forgets
their password had no way back in and the owner had no way to help. The
owner generates a single-use 24 hour link and delivers it the same way
as an invite. Issuing one invalidates any earlier unused reset, and the
reset endpoint answers identically for unknown, used and expired tokens
so it cannot be used to probe which exist.

The admin area 404s for members rather than 403ing, so its existence is
not advertised, and every /api/admin handler independently requires the
owner role.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 21:26:33 +00:00
JP
34afc497f4 Give members their own budget-capped gateway key
Members share the owner's AI budget, so each now gets a Switchboard key
minted on first use with a daily cap and a per-request ceiling. That
gives real spend limits and per-user attribution: the app's own rate
limiter lives in memory and resets on every deploy, so it could never be
a spend control. The gateway enforces the caps and answers 402
guardrail, which the error mapper already turns into a budget message.

Resolution order is the user's own key, then mint, then borrow the
owner's key for a single request if the gateway is unreachable - a
served request beats a hard failure, and the fallback logs loudly
because no per-user cap applies to it. The owner key is used only for
minting, never for inference.

API key management is now owner-only, enforced on GET, POST and DELETE.
GET matters as much as POST because it returns the masked key and the
gateway URL. Settings became a server component so the role is known
before first render: members never see the card and never issue the
request, rather than having it flash and disappear.

AiCall records one row per request from the gateway's own response
metadata, so member spend is queryable instead of a journald grep. The
write is fire-and-forget - tracking must never fail a working request.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 21:16:14 +00:00
JP
c81d1eefb3 Build invite redirects from the public origin
request.url carries the app's internal bind address (0.0.0.0:3000)
because it runs behind a reverse proxy, so every invite link - valid or
not - redirected to an unreachable https://0.0.0.0:3000/join.

Uses NEXTAUTH_URL, which is the configured public origin and cannot be
influenced by a request header, falling back to forwarded headers.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 21:03:06 +00:00
JP
a6cabc5178 Gate signup behind an invitation and revoke sessions on suspend
Registration was open to anyone who could reach the app. Signup now
requires an invite: /invite/<token> validates the link and parks the
token in an HttpOnly cookie, /join re-checks it and renders the form,
and the register route consumes a use inside the same transaction that
creates the user - so a duplicate email rolls the use back instead of
burning it. The token never reaches the client, so the form cannot forge
or replay one.

The jwt callback now revalidates the user on every call and returns null
when the account is missing or suspended, which clears the session
cookie. Every API route and server component already branches on
session?.user?.id, so this revokes access everywhere without editing any
of them. The try/catch around that lookup is load-bearing: Auth.js
treats a throw in this callback the same as a null return, so an
unguarded transient database error would sign out every user at once.

authorize() rejects non-ACTIVE users too, so a suspended account cannot
sign in again to mint a fresh token.

/register redirects to /join; it is linked from elsewhere and may be
bookmarked.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 20:59:52 +00:00
JP
f0c745c50c Remove Google and GitHub sign-in
Both providers were registered but never usable: all four client
env vars are empty in production and no Account row has ever been
created. Email and password is the only path that has ever worked.

This also simplifies the invite gating that follows. With OAuth there
were two redemption paths, an invite token that had to survive the
provider round trip in a SameSite=Lax cookie, and an
OAuthAccountNotLinked dead end for anyone who signed up with a password
and later clicked a provider button. Now there is one path.

Account, Session and VerificationToken stay in the schema. They are
Auth.js's tables and dropping them would be a destructive migration for
no benefit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 20:42:40 +00:00
JP
b3a0cc75eb Add roles and invite models
Additive schema only; no behaviour changes yet. Verified with prisma
migrate diff against production: two enums, three new tables, new User
columns and foreign keys, and no destructive statements.

Role lives on the User row rather than an OWNER_EMAIL env check, because
email is nullable and user-editable and so a poor thing to authorize
against. The jwt callback will need a per-request lookup for status
anyway, so reading role in the same query is free.

Invites are bearer tokens in a URL, shared out of band as a link or QR
code, because the app has no email capability. claimInvite consumes a
use with a single UPDATE guarded on usedCount < maxUses: Prisma cannot
compare two columns in a where clause, and one statement means one row
lock, so two people redeeming the last use cannot both succeed.

SearchCache finally gets its user relation. Every other user-owned model
cascades; without it, deleting a user left orphaned rows holding their
raw search queries. Verified zero orphans before adding the constraint.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 20:35:48 +00:00
JP
181ef13c7e Make the image bucket private
The bucket carried an anonymous download policy, which is what made the
old unauthenticated proxy work. Now that images are served through a
session-gated route using credentials, anonymous access is unnecessary
and was the second half of the public exposure.

Applied on the running host; both compose files updated so bringing the
stack up elsewhere does not re-apply 'download'.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 20:31:01 +00:00
JP
996b1b5360 Serve image bytes via the SDK's documented transform
The SDK returns a Node IncomingMessage, not a web ReadableStream. undici
accepts it (verified against the live bucket, 21882 bytes), but the cast
claimed a type it never had. transformToByteArray is the documented API;
buffering is fine with uploads capped at 10MB.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 20:26:58 +00:00
JP
e0c1814cf1 Put images behind the session and make routes private by default
The /minio-images rewrite proxied straight onto the MinIO bucket with no
authentication, and the bucket carries an anonymous read policy. Verified
from the public internet: GET /minio-images?list-type=2 returned a full
ListBucketResult naming all 33 objects, and any key then fetched 200. Every
menu scan, drink photo and bar photo was enumerable and downloadable by
anyone. Replaced with a route handler that requires a session and streams
via the existing getImage(). The URL shape is unchanged, so stored
imageUrls, uploadImage/getImageUrl and every <img> keep working.

Nothing uses next/image and it must stay that way here: the optimizer
fetches server-side without the session cookie.

Route protection was an allowlist that had to be updated by hand for each
new page, and had already been missed for /bar, /bartender and /recommend,
which rendered to logged-out visitors. /recipes was in the middleware
matcher but not the authorized() list, so it fell through too. Inverted
both to a denylist so new pages are private by default. /api stays out of
the matcher because authorized() answers with an HTML redirect and API
clients need the JSON 401 those routes already return.

Also fixes a cross-user write: POST /api/recipes took sourceDrinkId from
the client with no ownership check, and the drink detail page loaded its
recipes relation unfiltered, so one user could attach an arbitrary recipe
to another user's drink where it rendered permanently and the owner could
not delete it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 20:20:39 +00:00
JP
593f68138c Add a camera option wherever a photo can be attached
The bar add-item form could only pick an existing file, so adding a
bottle meant taking a photo first and then hunting for it. The scan flow
already had a working camera; this reuses that CameraCapture component
rather than adding a second implementation.

The change is in DrinkImageUpload, which the bar form, the drink form
and the drink detail view all share, so all three gain the camera.

Falls back to a file input with capture="environment" when getUserMedia
is unavailable - it needs a secure context, so it is absent when the app
is reached over plain http on the LAN.

Also sets type="button" on CameraCapture's controls. They previously had
no type, which defaults to submit, so capturing a photo inside the bar
item form would have submitted the form.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 19:16:20 +00:00
JP
058735b1a0 Mirror external images into our own storage so they render
Bar items added by barcode stored the Open Food Facts image URL
directly. The CSP in next.config.mjs restricts img-src to our own
origin, so the browser blocked those and showed a broken image - the
picture was fine, we just could not display it.

Copy externally-hosted images into MinIO at lookup time and hand back a
/minio-images path instead. That fixes the class rather than the
instance: no CSP entry is needed per image source, the picture survives
the source deleting or reorganising it, and the user's browser never
has to talk to a third party to render their own bar.

Also fixes a latent bug this uncovered: imageUrl was validated with
z.string().url(), which rejects the relative /minio-images/... paths
that uploadImage returns, so saving an uploaded drink image would fail
validation. That matches production having zero drinks with an image.
Both schemas now accept either form.

CSP keeps two third-party entries for OAuth avatars, which the provider
hosts and we only ever receive as a URL at sign-in.

Existing rows were backfilled separately.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 19:06:45 +00:00
JP
80fd99bc40 Fix deploy: sudo to root once and su down, not sudo -u
sudoers grants drinkadmin NOPASSWD only as root, so every 'sudo -u
drinktracker' in the deploy prompted for a password. Restructured to
pipe one script per phase to 'sudo -n bash -s' and use su inside, which
needs no privilege change and cuts the SSH round-trips.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 17:59:48 +00:00
JP
87f951bf94 Run natively under systemd; Docker was unworkable in this LXC
Docker in this unprivileged Proxmox LXC was broken two independent ways:
AppArmor could not load the docker-default profile, so the daemon could
not start any new container, and runc could not set the
net.ipv4.ip_unprivileged_port_start sysctl, which is why every service
needed network_mode: host. `docker build` was impossible outright since
Docker 29 removed the classic builder. The stack only survived because
the containers predated the breakage - a reboot would have left the app
down, and nothing could be redeployed.

Postgres, MinIO and the app now run natively under systemd. Deploys
build out-of-place into releases/<sha> and swap a symlink, so the build
happens while the old release keeps serving and downtime is the ~3s
restart rather than the ~3min build. Rollback is the same swap in
reverse with no rebuild.

Dockerfile and docker-compose.prod.yml are unchanged and still work on
a normal host; the compose app service gains `build: .` so it can come
up on a VPS that cannot reach the LAN-only Gitea registry.

Documents three traps found during the migration: rewrites() in
next.config.mjs is evaluated at build time so MINIO_ENDPOINT changes
silently do nothing, prisma migrate deploy would fail because
production has no _prisma_migrations table, and the image proxy relies
on the bucket's anonymous-download policy.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 17:58:34 +00:00
JP
f3865c7901 Fix deploy: avoid cd into /root as the deploy user
The remote shell runs as drinkadmin, which cannot enter /root even
though sudo can operate there, so 'cd $REMOTE_DIR && sudo docker ...'
failed at the cd. Pass the build context and compose file as paths
instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 17:06:01 +00:00
JP
2beb2d5d43 Version the deploy skill with the repo
.claude/ was ignored wholesale, which meant the deploy skill lived only
on one machine. Narrow the rule so .claude/skills/ is tracked while
settings.local.json stays local.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 17:02:51 +00:00
JP
45d9bd12d7 Add deploy script and skill for the production LXC
There is no CI, so a git push deploys nothing - the image has to be built
and the stack restarted separately. This captures that as one command
rather than a sequence to remember.

The image is built on the LXC and tagged with the registry name the
compose file already references, so compose finds it locally and never
pulls. That means no registry credentials are needed on either machine.

Also records the two things that cost the most time to work out: the
public hostname resolves to the reverse proxy rather than the container
(192.168.2.169), and the live stack runs from /root/drinktracker while
the checkouts under /home/drinkadmin are stale.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 17:02:20 +00:00
JP
a0e1619072 Route all AI features through the Switchboard gateway
Replace the direct Anthropic and OpenAI integrations with a single
provider that talks to Switchboard, an OpenAI-compatible gateway that
routes each request to the best available model. The app no longer pins
a model id anywhere: it sends switchboard/auto and lets the gateway
choose, then logs which model answered and what it cost.

Routing levers are set per feature in src/lib/ai/routing.ts. Three of
those choices came from measuring against the live gateway:

- category and prefer_free are set explicitly on every request. An API
  key carries its own routing defaults, and anything left unset inherits
  them - drink prompts were being sent to a free coding model.
- Token budgets are generous because the router may pick a reasoning
  model, and reasoning tokens come out of the same max_tokens budget as
  the answer. At 512 tokens a request returned null content; at 4096 the
  same request returned correct JSON.
- No tier lever on text features. tier "cheap" pinned a slow reasoning
  model (42-180s, two timeouts and one truncated response in five
  trials) and tier "frontier" escalated as far as Opus at $0.02 a call,
  while unconstrained routing answered in about a second. Vision keeps
  "frontier", where the accuracy is worth a few tenths of a cent.

Gateway failures are mapped to actionable messages rather than passed
through: a 401 relayed as 401 would read as an expired session and
bounce the user to login, and a 429 would collide with the app's own
rate limiter.

Also collapses the key lookup that was duplicated across ten call sites
into getUserProvider(), which fixes a latent bug where a bare findFirst
with no ordering let different features pick different providers.

Existing claude/openai key rows are ignored at runtime and offered for
removal in Settings, so no migration is needed before deploying.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 16:41:00 +00:00
JP Scott
7c41b15ecc Update prod compose to pull from Gitea container registry
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-04 23:01:47 -07:00
JP Scott
dc1ad4d0c0 Add recipes, images, AI photo ID, barcode scanning & ingredient matching
- Fuzzy ingredient matching for bar inventory against recipes
- AI photo identification API for bottles/labels (drink + bar context)
- Barcode scanner with photo toggle for My Bar
- Barcode scan + photo ID buttons on Add Drink form
- Auto-pull product images from Open Food Facts barcode lookup
- Recipes section on drink detail pages with bar availability
- Dedicated Recipes page in sidebar navigation
- Bar item image support (schema, upload, display)
- Drink detail image upload component
- MinIO image proxy through Next.js rewrites (fixes broken image links)
- Improved category mapping (energy drinks → Mixers, not Spirits)
- Re-process saved recipe ingredients against current bar inventory

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-04 22:26:17 -07:00
JP Scott
2ac2c4b2d4 Add My Bar, Bartender, Recommend features + drink images
- Drink Images: upload/display photos of bottles/cans on drink cards and detail pages
- My Bar: inventory tracker for spirits, liqueurs, mixers, bitters, garnishes, tools
- Bartender: AI-powered cocktail recipe generation, "what can I make" suggestions,
  saved recipes. Cross-references bar inventory for ingredient availability.
- Recommend: AI flavor profile analysis, personalized drink recommendations,
  "find similar" drinks based on highly-rated favorites
- Navigation: desktop sidebar with all 8 routes, mobile bottom nav with
  4 primary items + "More" popup menu
- New Prisma models: BarItem, Recipe, FlavorProfile
- Backup/restore updated to include bar items

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 18:28:02 -07:00
JP Scott
d8f069cce4 Fix CSV restore: strip BOM and extract embedded ratings from drink rows
- Strip UTF-8 BOM that Excel/editors add to CSV files
- When drink rows contain score/notes/wouldReorder fields, automatically
  create rating entries (supports manually edited CSVs)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 16:45:04 -07:00
JP Scott
247a21b6a7 Add AUTH_TRUST_HOST for reverse proxy deployments
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 16:16:00 -07:00
JP Scott
4ed53d0fd7 Switch to host networking for Proxmox LXC compatibility
network_mode: host avoids Docker creating separate network namespaces
which trigger sysctl writes blocked in LXC containers. All service
references updated from container names to localhost.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 15:37:14 -07:00
JP Scott
c401c8f2e0 Add privileged: true to all services for LXC Docker compatibility
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 15:12:57 -07:00
JP Scott
959cf57a46 Add apparmor:unconfined to all services for Proxmox LXC compatibility
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 14:45:35 -07:00
JP Scott
491f8f2c7b Switch to pre-built Docker Hub image for production
- Push app image to jpscott84/drinktracker on Docker Hub
- docker-compose.prod.yml uses image instead of build
- install.sh pulls image instead of building from source
- Much faster deploys (no npm ci/build on target server)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 14:38:44 -07:00
JP Scott
9212fd4acd Fix compose env var interpolation with --env-file flag
Docker Compose reads ${VAR} interpolation from .env by default,
not from the env_file directive (which only sets container vars).
Added --env-file .env.production to all docker compose commands
so POSTGRES_USER, POSTGRES_PASSWORD, etc. are available for
compose file interpolation.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 14:32:25 -07:00
JP Scott
44c70e7825 Add auto-install for Docker and dependencies in install script
- Automatically installs Docker via get.docker.com if not found
- Installs Docker Compose plugin if missing
- Installs OpenSSL and curl if missing
- Detects package manager (apt, dnf, yum, apk)
- Handles docker group permissions for current user
- Falls back to sudo for docker commands when needed

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 14:28:01 -07:00
JP Scott
1d454d84b2 Add production install script and migrate service
- install.sh: Interactive setup script for Linux VPS/LXC deployment
  - Checks prerequisites (Docker, Docker Compose, OpenSSL)
  - Auto-generates all secrets (Postgres, MinIO, NextAuth, encryption)
  - Creates .env.production with proper Docker service hostnames
  - Builds and starts all services via docker-compose.prod.yml
  - Health check loop with status reporting
  - Idempotent (safe to re-run)

- docker-compose.prod.yml: Add migrate service
  - One-shot container that runs prisma db push before app starts
  - App depends on migrate completing successfully
  - Override DATABASE_URL and MINIO_ENDPOINT for Docker networking

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 13:40:48 -07:00
JP Scott
8a582bfa7f Security hardening for production readiness
- Add security headers (CSP, HSTS, X-Frame-Options, X-Content-Type-Options, etc.)
- Strengthen password requirements (10+ chars, mixed case, numbers)
- Increase shared list slug entropy from 4 to 16 bytes
- Add rate limiting to login, registration, upload, and restore endpoints
- Add file magic number validation for image uploads (JPEG, PNG, WebP, HEIC)
- Add CSV row limit (50k) to restore endpoint
- Update client-side registration form to match new password policy

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 12:55:16 -07:00
JP Scott
969bc9347a Initial commit: DrinkTracker full-stack app
Next.js 14 drink collection tracker with AI-powered search,
menu scanning, ratings, wishlist, sharing, and CSV backup/restore.

Features:
- Auth (credentials + OAuth ready)
- Drink collection with ratings and reviews
- AI search via Claude/OpenAI with search history
- Menu photo scanning with AI extraction
- Wishlist / Try Later system
- Public sharing via slug URLs
- CSV backup and restore (merge/replace modes)
- Docker Compose for Postgres + MinIO + dev server

Security: docker-compose files use env var interpolation
instead of hardcoded secrets.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 12:42:11 -07:00