Files
drinktracker/deploy/README.md
JP 82680d9430 Correct stale deploy docs and document the new env knobs
deploy/README.md described a rewrites() block that no longer exists, and told
the reader to verify the image bucket's anonymous-download policy was set -
which is precisely the exposure that was closed. Following it would have
re-opened every stored image to the public internet.

Replaced with what is actually true: /minio-images is a session-gated route
handler with per-user key-prefix ownership, the bucket must have no anonymous
policy, and MinIO must stay bound to loopback. Also noted that the suggested
`mc anonymous get` check does not work on this host - the alias has no
credentials and returns Access Denied either way - and gave an external curl
that actually tells you something.

The same stale rationale sat in deploy.sh's build comment. Building without the
env file is still right, just for a different reason: nothing needs baking in.

Added the MCP and OAuth journal greps, an unauthenticated discovery check, and
a note that a 307 there means /.well-known fell out of PUBLIC_ROUTES. Documented
in .env.example that NEXTAUTH_URL must be the exact public origin with no
trailing slash, since it is now the OAuth issuer and MCP resource identifier,
plus the MCP_CHATGPT_TOOLS opt-out.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1Ee4Mc1X1SX8HgYa52zu7
2026-08-09 19:04:42 +00:00

112 lines
8.6 KiB
Markdown

# Deploying drinktracker
```bash
./deploy/deploy.sh # typecheck, push, build, swap, restart, verify
./deploy/deploy.sh --yes # skip the confirmation prompt
./deploy/deploy.sh --rollback # repoint at the previous release (~3s, no rebuild)
```
The app runs **natively under systemd — there is no Docker on this host.**
## Where things live
| | |
|---|---|
| Production LXC | `192.168.2.169` (hostname `drinktracker`, Proxmox vmid 100) |
| Deploy access | `ssh -i ~/.ssh/drinktracker_ed25519 drinkadmin@192.168.2.169` |
| App | `/opt/drinktracker/{repo,releases/<sha>,current}`, service user `drinktracker` |
| Secrets | `/etc/drinktracker/drinktracker.env` (`root:drinktracker` `0640`) |
| Services | `drinktracker`, `minio`, `postgresql@16-main` |
| Database | `drinkman`, user `drinkdb`, `127.0.0.1:5432` |
| Object storage | MinIO at `127.0.0.1:9000`, bucket `drink-images` |
| App port | 3000 on `0.0.0.0`; unauthenticated requests return 307 → `/login` |
| Backups | `/var/backups/drinktracker`, nightly 03:25, 14-day retention |
## Traps
**`drinktracker.tenseconddelay.net` resolves to the reverse proxy (192.168.2.172), not the container.** `NEXTAUTH_URL` points at the public name, so resolving it and SSHing there fails in a way that looks like a rejected key but is simply a different machine.
**Images are served by an authenticated route, not a rewrite.** `next.config.mjs` has no `rewrites()` block at all. `/minio-images/:path*` is a route handler (`src/app/minio-images/[...key]/route.ts`) that requires a session, rejects `..` in the key, and enforces per-user ownership by key prefix (`<userId>/` or `scans/<userId>/`). It reads `MINIO_ENDPOINT` at request time, so changing it takes effect on restart with no rebuild.
> This replaced an unauthenticated rewrite that, combined with an anonymous-download bucket policy, exposed every stored image to the public internet. **The bucket must not have an anonymous policy, and MinIO must stay bound to loopback** (`127.0.0.1:9000`, verify with `ss -lntp | grep 9000`). `mc anonymous get local/drink-images` is not a usable check here — the `local` alias on this host has no working credentials and returns `Access Denied` regardless. Check from outside instead: `curl -si https://drinktracker.tenseconddelay.net/minio-images?list-type=2` must **not** return a `200` with an XML key listing.
**Never use `next/image` for stored images.** The optimizer fetches them server-side without the session cookie, so every image would 401. All renders are plain `<img>`, and there is deliberately no `images.remotePatterns` config.
**`NEXTAUTH_URL` must exactly equal the public origin, with no trailing slash.** Beyond login redirects it now drives the OAuth `issuer` and the MCP `resource` identifier, both of which clients compare byte-for-byte. A mismatch does not error — connectors just fail to complete with an unhelpful "couldn't reach the server".
**`prisma db push`, not `migrate deploy`.** Production has no `_prisma_migrations` table, so `migrate deploy` would try to apply the init migration against populated tables and fail. `--accept-data-loss` means removing a field from `schema.prisma` **drops the column and its data** — the deploy script dumps first for this reason.
## How a deploy works
Build happens in `/opt/drinktracker/repo` and is assembled **out-of-place** into `releases/<sha>`, so the ~3 minute build runs while the current release keeps serving. Only the symlink swap and `systemctl restart` cause downtime, about 3 seconds. Rollback is the same swap in reverse and needs no rebuild.
The build deliberately does **not** source the env file, same as the Dockerfile. Nothing in this app needs configuration baked into the bundle: `NEXTAUTH_URL`, `MINIO_ENDPOINT` and the rest are read at request time by dynamic routes, so the build stays identical regardless of which host it runs on. (This originally guarded against a `rewrites()` block that baked a MinIO URL into `routes-manifest.json`; that block is gone, but building without secrets is still the right default.)
## Docker still works, just not here
`Dockerfile` and `docker-compose.prod.yml` are maintained for running this app elsewhere (a VPS). `docker compose build && docker compose up -d` works on any normal Docker host.
It does **not** work on this LXC, which is why the app runs natively. Two independent problems, both from unprivileged-LXC nesting:
1. AppArmor — the container can't load the `docker-default` profile (`apparmor_parser: Access denied`), so any container without `--security-opt apparmor=unconfined` fails.
2. runc can't set `net.ipv4.ip_unprivileged_port_start`, so containers only work with `--network host`.
`docker build` is impossible regardless: Docker 29 removed the classic builder and BuildKit spawns a builder container without the unconfined option. Docker is installed but disabled; the old containers and volumes are retained until 2026-09-07 as a rollback path.
## Verifying by hand
```bash
ssh -i ~/.ssh/drinktracker_ed25519 drinkadmin@192.168.2.169
systemctl is-active postgresql@16-main minio drinktracker
sudo journalctl -u drinktracker -n 50
sudo journalctl -u drinktracker | grep '\[switchboard\]' # AI routing + cost per call
sudo journalctl -u drinktracker | grep '\[mcp\]' # one line per MCP tool call
sudo journalctl -u drinktracker | grep '\[oauth\]' # connector grants and token issuance
readlink /opt/drinktracker/current # which release is live
```
Each AI call logs one `[switchboard] feature=… model=… cost=… latency_ms=…` line. `FAILOVER` warnings are expected — the gateway's `:batch` model variants currently fail on every request and fall back. `CONTEXT OVERFLOW` is worth investigating.
MCP tool calls log `[mcp] user=… tool=… ok=… ms=…`, and the same events are rows in `McpAuditLog` (arguments are deliberately never recorded). `[oauth]` lines cover consent grants and token issuance; a `PKCE mismatch` or `refresh token replay` warning there means a grant was revoked defensively and is worth looking at.
Discovery can be checked without credentials:
```bash
curl -s https://drinktracker.tenseconddelay.net/.well-known/oauth-protected-resource | jq
curl -si https://drinktracker.tenseconddelay.net/api/mcp -X POST \
-H 'content-type: application/json' -H 'accept: application/json, text/event-stream' \
-H 'MCP-Protocol-Version: 2026-07-28' -H 'Mcp-Method: tools/list' \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{"_meta":{"io.modelcontextprotocol/protocolVersion":"2026-07-28","io.modelcontextprotocol/clientCapabilities":{}}}}'
```
The second must return **401** with a `WWW-Authenticate: Bearer … resource_metadata="…"` header. A `307` there means `/.well-known` fell out of `PUBLIC_ROUTES` in `src/lib/auth.ts` and discovery is broken.
## Backups
`drinktracker-backup.timer` runs nightly at 03:25: `pg_dump` + a tar of `/var/lib/minio/data` + a copy of the env file, into `/var/backups/drinktracker`, retained 14 days (~1 GB).
**TODO — the chain has no off-site leg.** `/usr/local/bin/drinktracker-backup` supports `OFFSITE_DEST` (an scp target) and `OFFSITE_KEY`; set them via `Environment=` in `drinktracker-backup.service`. Until then everything lives on one LXC on one Proxmox host, and the nightly job prints a warning saying so.
Restore:
```bash
sudo systemctl stop drinktracker
zcat /var/backups/drinktracker/pg-<stamp>.sql.gz | sudo -u postgres psql -d drinkman
sudo systemctl stop minio
sudo tar -C /var/lib/minio -xzf /var/backups/drinktracker/minio-<stamp>.tgz
sudo chown -R minio:minio /var/lib/minio/data
sudo systemctl start minio drinktracker
```
## Notes
- The deploy needs unattended sudo on the LXC (`/etc/sudoers.d/drinkadmin-deploy`, currently
`NOPASSWD: ALL`). A narrowly-scoped allowlist was considered and rejected: the script calls
root for `install`, `tee`, `systemctl`, `journalctl` and `sudo -u postgres pg_dump`, so an
allowlist would need updating every time the script changes and would fail in confusing ways
when it drifted. Be aware this buys little security either way — `drinkadmin` is in the `sudo`
group, so full root is a password away regardless. What it grants is *unattended* root.
- `nodejs` is `apt-mark hold`-ed at v20.20.0 so the native runtime can't drift from the Dockerfile's pinned `node:20-alpine`.
- Postgres was created with `en_US.UTF-8`; the Ubuntu default of `C`/`SQL_ASCII` would re-sort every drink list and mangle names like `Grüner Veltliner`. Preserve the locale if the cluster is ever rebuilt.
- `unattended-upgrades` is not installed, so nothing will surprise-restart these services — but postgres will never auto-patch either. A deliberate open question.