A health check that reads environment variables is not a health check
Decided
/api/health now answers two different questions and says which one it answered. The default GET reports credential presence and labels itself checks: "configuration". GET /api/health?probe=deep sends a real request to each of the four dependencies and labels itself checks: "reachability". The deep arm is cached in module scope for 30 seconds and rate limited on a cache miss. scripts/check-health-honesty.mjs holds the shape in CI.
What was actually wrong. The old endpoint had a careful, plausible argument in its header for why checking environment variables was the right thing to do, and on 4 Aug 2026 it returned
{"status":"ok","dependencies":{"database":"configured","email":"configured",
"helpdesk":"configured","scheduling":"configured"}}while a single POST /support was failing three separate ways at once: the Supabase URL named a project that did not contain the intake tables (D-030), the service-role variable held a browser-level publishable key (D-029), and the helpdesk token had been revoked (D-031). Two external uptime monitors polled this URL throughout and recorded 100% availability. support_tickets contained zero rows. Every consultation request, support ticket and job application the site had ever received had been lost.
None of those three faults changes the shape of an environment variable. All four variables were set and all four values were plausible. That is the whole of it: a credential check can prove that somebody filled in a form, and "somebody filled in the form" is not what anyone reads a health endpoint to learn.
Why the old argument was kept rather than deleted. It was not wrong, only incomplete. An unauthenticated endpoint that makes four outbound calls per request is a free amplifier pointed at four vendors and a free way to hold a database connection open. So the round trip is opt-in rather than default, the result is cached for PROBE_CACHE_MS, and only a cache miss consults the limiter — serving the cache costs nothing, and a monitor that receives a 429 records an outage that did not happen, which is precisely the class of lie this endpoint exists to stop telling.
Why each probe is the one it is. Each reproduces one of the three failures that were live that day, cheaply and without side effects. probeDatabase() selects from support_tickets with head: true — a table present in both Supabase projects would have stayed green through the incident, and head means no row data crosses the wire. The count is discarded, so the number of support tickets the company has received is not published to anyone who appends a query string. probeEmail() does not stop at "the key works": it reads GET /domains and requires the domain in the deployment's own From line to be *verified*, because Resend refuses a send from an unverified domain and a valid key on an unverified domain is a perfectly healthy-looking transport that delivers nothing. probeHelpdesk() calls users/me, which a revoked token fails. probeScheduling() uses GET rather than HEAD, because a vendor's 404 for a deleted event type is produced by the application while a HEAD is answered by the edge.
What the status code still excludes, and why that did not change. Without a database a submission is lost; without a transport nobody is told it arrived. Without a helpdesk the submission is stored, emailed, and answered by a human reading an inbox; without a scheduler the visitor asks for a time in the site's own modal and we confirm it by hand. Worse in each case, but not a failure — rolling them into degraded would put a permanent 503 on a deployment that is doing its job, which is how a health check stops being read.
Why a script guards this and not the type system. The failure mode is additive. A developer adding a fifth dependency to the dependencies object gets no compiler error for forgetting its probe; the deep arm goes on reporting ok for four things while the fifth is unmeasured, which is this bug again with a smaller blast radius. So check:health-honesty asserts what a compiler cannot: both arms labelled, the default arm not probing, dependencies and reachability naming the same set, and every probe imported, exported and actually awaited in the fan-out. It was verified by mutation — nine deliberate regressions, including the swap of the two labels and a fifth dependency with no probe, and each one fails the script.
No detail string is ever a vendor's own words. The endpoint is unauthenticated, and vendor error bodies quote back request URLs, account identifiers and occasionally the tail of a key. Every phrase a caller can see comes from the fixed vocabulary in lib/health/probe.ts.
What would make this wrong. If the endpoint ever needs authentication — if the deep arm grows a check that cannot be made side-effect-free, or the detail strings need to carry vendor text to be useful — then the right move is a second authenticated endpoint, not loosening this one. And if a monitor is ever pointed at the deep URL, PROBE_CACHE_MS must stay comfortably under its interval or the monitor measures the cache instead of the deployment.