Trust Center / Security
Continuity Plan
Our CEO leads recovery, as in our Incident Response Plan. We aim to lose at most 1 day of data (RPO) and to recover within 7 days (RTO).
Failures
| Failure | Effect | Recovery |
|---|---|---|
| Cloudflare | API and dashboard down | Wait for Cloudflare, or move our Workers elsewhere if the outage runs long |
| Customer data lost or corrupted | Sign-in and token checks fail | Restore the latest daily backup, held within Cloudflare |
| One analysis region | The other regions keep answering | Rebuild its servers from our deploy scripts |
| Every analysis region | Cached verdicts are still served; new analysis waits | Rebuild servers, most-used region first |
| Verdict master | Scan servers fail over to its replica | Rebuild the master from the replica |
| Our LLM | Grading falls back to OpenRouter, or to no second opinion | Restart or rebuild it |
| Our on-prem dataset | None; serving customers doesn't depend on it | Restore from R2 or ZFS snapshots |
| A sign-in provider | Its users can't reach the dashboard; API tokens still work | Wait for the provider |
| Stripe | New purchases fail; existing access continues | Wait for Stripe, which retries its webhooks |
Steps
- Declare: start a dated log and tell affected customers what is down
- Restore: the API first, then customer data, then analysis, then the dataset
- Verify: confirm sign-in, token checks, and fresh analysis all work
- Review: within two weeks, write up what happened and what we changed
Testing
Each year we restore customer data from backup and rehearse losing a region, and keep the notes.