gabriel / musehub public
musehub-production-infrastructure-launch-todo.md markdown
286 lines 41.5 KB
Raw
sha256:cd5c2fcb91a44ac9e38e9c36176c27ce079a074586d9852296da628a29fb01ff Merge 'docs/aws-identity-and-deploy-fixes' into 'dev' — pro… Human 3 days ago

MuseHub Production Infrastructure Launch TODO

Goal

Reproduce the proven staging architecture in the dedicated MuseHub Production AWS account, deploy and validate the application, configure production DNS and TLS, and establish a safe, repeatable production deployment and rollback workflow.

This is an infrastructure reproduction and promotion project—not a blind copy. Do not copy credentials, secrets, mutable data, IAM identities, or obsolete resources from Nonproduction without an explicit decision.

Account map

Purpose AWS account Account ID Root mailbox
Governance, billing, IAM Identity Center MuseHub Management Record in password manager/docs [email protected]
Staging and existing workloads MuseHub Nonproduction 992382692655 [email protected]
New production environment MuseHub Production 672469410277 [email protected]

Targets (decided 2026-08-25, closing out readiness ticket #149)

Framed against a "become the new GitHub" north star, phased against what the infra can actually support today — not aspirational numbers presented as already achievable.

Now (bootstrap) Next ~1 month GitHub-scale (aspirational)
RPO (max data loss) ≤24h — daily automated backup, once #157 lands ≤1h — more frequent snapshots or WAL-based PITR Near-zero — continuous replication
RTO (max downtime to recover) ≤4h — untested manual restore ≤1h — once restore is actually rehearsed at least once ≤15min — requires multi-AZ/automated failover, not being built yet
Traffic Low hundreds of users, sporadic pushes Hundreds–low thousands of users Thousands of users, hundreds of daily pushes (12-month marker)
Deploy downtime Hard zero for routine deploys — already true today via the existing blue/green pattern Same Same; brief (<5min) maintenance windows remain acceptable only for rare infra-level changes

Practical trigger to revisit t3.medium instance sizing: crossing into "hundreds of daily active pushers," not a calendar date. These RPO/RTO numbers directly size the backup cadence for #157.

Non-negotiable safety rules

  • [ ] Perform production work through IAM Identity Center using [email protected] and the AdministratorAccess permission set—not root credentials.
  • [ ] Use an explicit AWS CLI profile for every command; never rely on an implicit default profile.
  • [ ] Before any mutation, verify the target account with aws sts get-caller-identity.
  • [ ] Production mutations must target account 672469410277.
  • [ ] Nonproduction discovery must target account 992382692655 and remain read-only until a change is separately approved.
  • [ ] Add account-ID guards to deployment and provisioning scripts so they fail closed in the wrong account.
  • [ ] Do not reuse the old long-lived musehub-infra credentials for Production.
  • [ ] Do not expose secrets in terminal output, screenshots, commits, agent prompts, or logs.
  • [ ] Do not modify public DNS until Production passes pre-cutover validation.
  • [ ] Define and test rollback before routing production traffic.

Phase 0 — Access and account preflight

  • [x] Confirm [email protected] has the required IAM Identity Center access to:

    • [x] MuseHub Nonproduction for read-only discovery initially. — assigned AdministratorAccess 2026-08-23
    • [x] MuseHub Production with AdministratorAccess for initial provisioning. — already assigned
  • [x] Configure named AWS CLI SSO profiles, preferably:

    • [x] musehub-nonproduction
    • [x] musehub-production
  • [x] Use the working dual-stack AWS access portal URL for SSO configuration. — https://ssoins-7223bcb94598e297.portal.us-east-1.app.aws; the Regional IPv4 URL (*.us-east-1.portal.amazonaws.com) fails SAML sign-in the same way the legacy awsapps.com/start URL does — dual-stack is the only one that works for this SAML app config. Configured directly in ~/.aws/config as [sso-session musehub] rather than via the aws configure sso wizard, whose start-URL format validator rejects this URL even though it's valid (client-side bug, not a URL problem).

  • [x] Set the SSO region and default working region to us-east-1 unless discovery proves a resource belongs elsewhere.

  • [x] Authenticate and verify both profiles:

    aws sso login --sso-session musehub
    aws sts get-caller-identity --profile musehub-nonproduction
    aws sts get-caller-identity --profile musehub-production
    
  • [x] Confirm the first command reports 992382692655 and the second reports 672469410277. — verified 2026-08-23

  • [ ] Confirm production billing visibility and create an initial AWS Budget with email alerts. — needs Gabriel, not yet done

  • [x] Record the chosen AWS region, SSO profile names, and account IDs in the repository documentation without recording credentials. — this section

Phase 1 — Read-only staging and repository discovery

Repository inspection

  • [x] Locate all infrastructure and deployment assets using repository search:
    • [x] Deployment scripts and commands. — deploy/push.sh, deploy/deploy.sh, deploy/set-active-slot.sh, deploy/aws-provision*.sh, deploy/setup-ec2*.sh, deploy/bootstrap-instance.sh, deploy/backup.sh, deploy/cloudwatch-alerts.sh
    • [x] Dockerfiles and Compose files. — Dockerfile, docker-compose.yml, docker-compose.override.yml
    • [x] Infrastructure-as-code templates. — none exist; confirmed by docs/musehub-production-readiness-checklist.md Section 4 (IaC tool not yet chosen, Terraform/OpenTofu recommended but not built)
    • [x] Environment templates. — .env.example; live .env is hand-provisioned per instance, not templated
    • [x] Reverse-proxy configuration. — deploy/nginx-cf.conf (Cloudflare Origin Certificate model, not Let's Encrypt/ACM)
    • [x] Database migration and backup commands. — Alembic via deploy/deploy.sh; deploy/backup.sh exists but is not wired in (no cron, no rclone configured — see checklist Accepted Risks)
    • [x] Existing staging/production documentation. — docs/infrastructure.md, docs/deploy.md, docs/musehub-production-readiness-checklist.md and its production-readiness/ companion docs
  • [x] Document the exact current staging deployment command and its inputs. — bash deploy/push.sh staging; builds via Docker, pushes to ECR with crane, triggers deploy.sh over SSM
  • [x] Determine how the existing blue/green deployment works, including ports, health checks, active-slot selection, and rollback. — blue=1337/green=1338, /healthz gate, musehub-set-slot script writes the nginx upstream file, rollback via IMAGE_TAG=<prev> bash deploy/push.sh <env>
  • [x] Identify every place where staging-specific account IDs, hostnames, IP addresses, paths, ports, repository names, or secrets are hard-coded. — ECR registry 992382692655... is hard-coded in deploy/push.sh, deploy/deploy.sh, deploy/bootstrap-instance.sh; instance IDs hard-coded in deploy/push.sh, deploy/bootstrap-instance.sh, .museagent.md; see discovery report for full list

Key discovery note (2026-08-23): docs/musehub-production-readiness-checklist.md (dated 2026-08-04/08-08) already contains a much deeper, earlier production-readiness sweep than this TODO. It documents that the "prod" EC2 instance (i-0855d6efe7fa1a49d) referenced throughout deploy/ scripts lives in the old shared Nonproduction/Stori account (992382692655), is currently broken (no IAM instance profile, push.sh prod fails with InvalidInstanceId), and that DNS/TLS actually run through Cloudflare (DNS hosting + Origin Certificate), not Namecheap/Let's Encrypt/ACM as this TODO's Phase 5 assumes. See the full discovery report delivered separately for details before Phase 2 proceeds.

AWS inventory in MuseHub Nonproduction

  • [x] Inventory relevant resources in every used region, starting with us-east-1:
    • [x] VPCs, subnets, route tables, internet gateways, NAT gateways, and network ACLs. — single default VPC (vpc-0552bdd377de8a9a8, 172.31.0.0/16); no custom networking
    • [x] EC2 instances, AMIs, launch templates, volumes, snapshots, and Elastic IPs. — musehub-staging (i-07547cd20bee2dea5, t3.medium, running, EIP 23.22.27.39) and musehub-prod (i-0855d6efe7fa1a49d, t3.small, running, EIP 98.89.99.211); no launch templates/ASGs
    • [x] Security groups and inbound/outbound rules. — musehub-sg (sg-05815872537fcfe76): 443 open to Cloudflare IP ranges only, as documented. Finding, resolved 2026-08-23: an additional, undocumented rule opened TCP 1337 directly to a single IP (98.51.144.179/32), bypassing nginx/Cloudflare entirely. Confirmed dead (both docker-compose.yml and deploy/deploy.sh bind the app container to 127.0.0.1 only, so nothing was listening on the public interface for it) and revoked (sgr-03be7da55eece078f). Only port 443 remains open.
    • [x] IAM roles, instance profiles, policies, and deployment identities. — Users: musehub-infra, musehub-app, stori-assets-app (all expected), plus s3-migration-temp — created 2026-02-10 with AmazonS3FullAccess, used exactly once the day it was created (R2 migration, per deploy/migrate_r2_keys.py) and never since. Resolved 2026-08-23: access key deactivated (reversible — reactivate with aws iam update-access-key --status Active if ever needed). Roles: musehub-ec2-ssm, rds-monitoring-role, tus-image-processing-1-role-7dl2zvb1 (Stori-owned, not ours), plus the new SSO role.
    • [x] ECR repositories, image tags/digests, scan settings, and lifecycle policies. — one repo musehub/musehub, scan-on-push enabled, AES256 encryption, mutable tags, ~420 images accumulated with no pruning. Resolved 2026-08-23: lifecycle policy applied — expire untagged images after 14 days, keep only the most recent 60 tagged images.
    • [x] S3 buckets, policies, encryption, lifecycle rules, and CORS settings. — musehub-releases (ours) and stori-assets-992382692655-useast2 (Stori's, confirms shared-account risk from the checklist); no bucket found for DB backups.
    • [x] Databases, caches, queues, and other stateful services. — Finding: a managed RDS Postgres instance exists (musehub-staging-db, db.t3.micro, 20GB, single-AZ, not publicly accessible, 7-day automated backups) that isn't mentioned in any doc. The app's actual DATABASE_URL points at the containerized postgres:5432 service per docs/infrastructure.md, so this RDS instance is very likely orphaned/unused — needs Gabriel to confirm before considering decommission (it's actively costing money and, confusingly, is backed up, unlike the real database). No Lambda, DynamoDB, SQS, or SNS found.
    • [x] Systems Manager parameters and Secrets Manager secret names only—not their values. — Secrets Manager: empty. SSM: /musehub/staging/PACK_WORKER_URL, /musehub/staging/WEBHOOK_SECRET_KEY, /musehub/staging/WORKER_INTERNAL_KEY — a partial, abandoned attempt at the SSM secrets pipeline (missing DB_PASSWORD/RUNNER_TOKEN), consistent with docs/infrastructure.md noting deploy/secrets.sh "is not the live path."
    • [x] CloudWatch log groups, alarms, dashboards, and retention settings. — one log group /musehub/staging, retention set to Never Expire, already 1.9GB stored; zero alarms configured anywhere (cloudwatch-alerts.sh has evidently never been run for real — matches checklist Section 12 status).
    • [x] Load balancers, target groups, listeners, and health checks, if present. — none; direct EC2 + Elastic IP + nginx, as documented.
    • [x] DNS records hosted in AWS, if any. — none; DNS is Cloudflare-hosted (Route 53 not used).
    • [x] Backup plans, retention policies, and restore procedures. — none for the real (containerized) database; confirms the checklist's highest-priority open risk.
  • [x] Determine whether the existing musehub-prod EC2 resource inside Nonproduction is active, obsolete, or safe to retire later. Do not alter it during discovery. — running but non-functional: it's up and billing, but per docs/infrastructure.md has no IAM instance profile attached, so SSM can't reach it and push.sh prod fails outright. Not touched during this discovery. Disposition (fix in place vs. retire in favor of the new Production account) is a Phase 2/7 decision for Gabriel.

AWS starting state in MuseHub Production

  • [x] Confirmed clean: default VPC only (vpc-038dc9e464d82a05d), zero EC2 instances, zero S3 buckets, zero ECR repositories, zero IAM users, zero Elastic IPs, zero RDS instances. Only roles present are the SSO AdministratorAccess role and the standard AWS Organizations OrganizationAccountAccessRole. Ready for a from-scratch build.

Staging host inspection

  • [x] Record the staging host's:
    • [x] OS and architecture. — Ubuntu 22.04.5 LTS, x86_64
    • [x] Docker and Compose versions. — Docker 29.3.1, Compose v5.1.1
    • [x] Reverse proxy and configuration paths. — nginx, /etc/nginx/sites-available/musehub-staging (only site enabled); Cloudflare Origin Cert at /etc/ssl/cloudflare/origin.{pem,key} with correct permissions (644/640)
    • [x] Listening ports and firewall rules. — publicly: 443 (nginx) and 22 (sshd) on 0.0.0.0/[::]. Note: the security group has no port-22 ingress rule, so SSH isn't reachable from the internet today — but sshd is running and bound to all interfaces as a live daemon, meaning the SG is the only thing preventing exposure. Worth disabling sshd outright on the new Production host (SSM-only, per the existing "no SSH" design intent) rather than relying solely on the SG. App ports 1337/1338 confirmed bound to 127.0.0.1 only (docker-proxy), matching docker-compose.yml/deploy.sh — this is what made the now-removed SG rule for port 1337 a dead rule.
    • [x] systemd units, scheduled jobs, and deployment hooks. — no user crontab (no crontab for ubuntu); /etc/cron.d/ has only default OS entries (certbot, e2scrub_all) — deploy/backup.sh is confirmed never installed. A leftover certbot.timer fires every ~14h from the default Ubuntu certbot package install but has nothing to renew (Cloudflare handles TLS) — harmless noise, not a real risk.
    • [x] Environment-file and secret paths without displaying values. — /opt/musehub/.env, root-owned, 644 permissions (world-readable) — should be tightened to 600/640 on the new Production host.
    • [x] Certbot/Let's Encrypt configuration and renewal timers, if used. — not used for TLS (Cloudflare Origin Cert instead); default certbot package/timer present but inert.
    • [x] Log locations and rotation policy. — no dedicated log files; app logs go to stdout, captured by the awslogs Docker log driver into CloudWatch (/musehub/staging, never-expire retention — see Phase 1 AWS inventory finding above).
  • [x] Map all external dependencies: databases, object storage, email, OAuth, webhooks, APIs, payment services, analytics, and DNS. — containerized Postgres (not the orphaned RDS instance), Cloudflare R2 for object storage, Cloudflare for DNS/edge/TLS; no email/payment/analytics dependencies identified in this pass.
  • [x] Produce a staging inventory and dependency map before provisioning Production. — see this section plus the Phase 1 AWS inventory above.

Additional findings from host inspection (2026-08-24), not yet acted on:

  • /opt/musehub on the staging host contains the entire repositorynode_modules/, tests/, tools/, .mypy_cache/, .cursor/, and large standalone design docs (MUSEHUB_V2.md at ~100KB, several *_REDESIGN.md files) — not just the running app. This is a leftover from the pre-image-based deploy method (the old rsync-based setup scripts), and contradicts docs/deploy.md's current claim that deploys are "image-based via ECR. No SSH, no rsync, no code on the instance after provisioning." It's dead weight (likely a meaningful share of the 14GB/20GB, 70% disk usage observed) but not a safety issue — recommend a separate, deliberate cleanup task rather than doing it mid-discovery.
  • The one existing manual backup is empty: /opt/backups/musehub/musehub_staging_manual_20260613_003724.sql.gz is 20 bytes — independently reproduces the checklist's flagged finding that even the sole manual dump ever taken is unusable. Reinforces that this is the single highest-priority risk to close before or immediately after production launch.
  • Disk usage on staging is already at 70% (14G/20G) with no automated cleanup beyond the on-host image prune (last 3) — worth sizing the Production EBS volume larger and/or cleaning up the stray repo copy before it becomes a real constraint.

Phase 2 — Decide the production design

  • [x] Decide whether the current repository already contains usable infrastructure as code. — no, confirmed none exists (04-iac.md). Decided 2026-08-24: provision Production manually now (matches the proven staging pattern, fastest path to live), then codify into Terraform/OpenTofu as a follow-up task via import (the tool's own recommended adoption path for hand-created infra) — not before initial launch. Priority is getting Production live and handing Aaron real operating ownership; IaC's main benefit (reviewable, shared change process) matters most once Aaron starts operating independently, not before.
  • [ ] If it does, adapt it with isolated production variables and state. — N/A, no existing IaC.
  • [ ] Capture the production infrastructure as code — deferred to the post-launch follow-up above.
  • [x] Choose the production ingress/TLS topology: Direct EC2 + Elastic IP + nginx + Cloudflare Origin Certificate — matches staging's proven, working setup exactly. Neither Let's Encrypt nor ACM apply (no ALB/CloudFront in this architecture; TLS terminates at Cloudflare's edge).
  • [x] Confirm the production hostname(s). — Decided 2026-08-24: musehub.ai is canonical; www.musehub.ai will 301-redirect to it (matching the github.com/www.github.com convention). Not yet implemented — planned for Phase 5.
  • [x] Decide how application images are promoted: Decided 2026-08-24: dedicated Production ECR repo (see below), promoted via crane copy of the exact staging-validated image digest rather than rebuilding — implements the "build once, promote the tested artifact" pattern the CI/CD research flagged as missing.
  • [x] Decide whether Production uses a dedicated ECR repository or controlled cross-account access. — Dedicated repository, created in Phase 3 below. Rationale: Nonproduction is shared with the unrelated Stori project, and a cross-account trust relationship would directly violate this project's own rule ("production access, credentials, secrets, DNS, certificates, and deployment targets must be independent from nonproduction").
  • [ ] Decide how state and data will be handled: fresh production database (not a copy of staging), fresh musehub-prod R2 bucket — still to be executed in Phase 4.
  • [x] Establish resource naming and tagging conventions. — Applied starting 2026-08-24: every new resource tagged Environment=production, Application=musehub, Owner=gabriel, Managed-by=manual-pending-terraform, plus a Name tag. This is also the first consistently-tagged environment in the account (staging/legacy-prod have inconsistent/absent tags per 15-cost-controls.md).
  • [ ] Explicitly identify which staging settings must differ in Production — still to be executed in Phase 4 (env vars, secrets, R2 bucket name, CORS origins).

Phase 3 — Provision the production baseline ✅ complete (2026-08-24)

  • [ ] Create or confirm an AWS Budget and cost alerts for MuseHub Production. — needs Gabriel/Management-account access, not yet done.
  • [x] Confirm CloudTrail/account audit coverage and appropriate log retention. — Done 2026-08-25. The prior sweep's "permissions gap" was re-verified with full admin access, and the honest finding is: none of CloudTrail, GuardDuty, Security Hub, or AWS Config existed in either account — not just unconfirmed, genuinely never enabled. Decided (per Gabriel): per-account now, org-wide from Management deferred to later. Set up in both Nonproduction and Production:
    • Dedicated CloudTrail S3 bucket per account (musehub-<env>-cloudtrail-<account-id>), public access fully blocked, AES256 encryption, 365-day lifecycle expiration, consistent tags.
    • Multi-region trail with log file validation enabled, logging confirmed active in both accounts.
    • GuardDuty detector enabled in both accounts (default 6-hour finding frequency).
    • Also confirmed while checking: no Route53 zones (DNS is 100% Cloudflare, as already known), no CloudFront/WAF (Cloudflare's edge is the substitute layer, as already documented), Budgets API blocked because these are linked accounts in an Organization (must be enabled from Management — separate, deferred task).
  • [x] Provision the production network or intentionally use the account default VPC only if documented and accepted. — using the default VPC (vpc-038dc9e464d82a05d), matching staging's approach; documented here as the accepted choice.
  • [x] Provision production security groups:
    • [x] Public access only on required ingress ports — musehub-production-sg (sg-02afc8844c9bf135a): TCP 443 only, restricted to the 15 Cloudflare IPv4 + 7 IPv6 ranges (identical list to Nonproduction's musehub-sg). No port 80, no port 22.
    • [x] Restrict SSH to known sources or prefer AWS Systems Manager Session Manager. — no SSH rule exists at all; SSM is the only access path, by design.
    • [x] Restrict database and internal service access to required security groups. — N/A at this stage (DB is containerized, no separate SG needed yet).
  • [x] Create the compute layer using the chosen architecture. — EC2 instance i-043aaed71bef11903, t3.medium (matches staging's proven size, not the undersized t3.small the old broken "prod" used), Ubuntu 22.04 LTS (current AMI ami-06e78a71af43ef21a, resolved live via the AWS public SSM parameter rather than reusing the older hardcoded AMI ID), IMDSv2 enforced (HttpTokens=required) — a hardening default the older instances don't have.
  • [x] Encrypt storage and configure snapshots/backups. — 30GB gp3 root volume (larger than staging's 20GB, to avoid repeating the disk-pressure issue found there); snapshot/backup automation is Phase 4/8 work, not yet configured.
  • [x] Allocate an Elastic IP. — 100.60.14.220 (eipalloc-0797b2cd35c017a5e), associated with the instance.
  • [x] Create least-privilege production IAM roles and instance profiles. — musehub-production-ec2-ssm role + instance profile: AmazonSSMManagedInstanceCore managed policy, an inline policy scoping ECR pull actions to exactly the new Production repository ARN (not *), and a second inline policy (added 2026-08-24 after the first container launch failed with AccessDeniedException on logs:CreateLogStream) scoping logs:CreateLogStream/PutLogEvents/DescribeLogStreams to exactly the /musehub/production log group ARN. SSM registration verified — the instance showed Online on the very first poll, confirming this is correctly wired (unlike the old musehub-prod instance, which has never had a working instance profile).
  • [x] Create production ECR repositories/access and lifecycle policies. — 672469410277.dkr.ecr.us-east-1.amazonaws.com/musehub/musehub, scan-on-push enabled, AES256 encryption, and the same lifecycle policy applied to Nonproduction (expire untagged >14 days, keep last 60 tagged).
  • [x] Create production parameters and secrets separately; never copy staging secrets wholesale. — Done 2026-08-25, upgraded from hand-provisioned .env to the SSM Parameter Store pipeline (deploy/secrets.sh), resolving readiness-ticket #156's "Contradiction 2" for good rather than repeating staging's partial (3-of-7 parameters) setup:
    • All 6 secrets (DB_PASSWORD, WEBHOOK_SECRET_KEY, RUNNER_TOKEN, BLOB_STORAGE_ACCESS_KEY_ID, BLOB_STORAGE_SECRET_ACCESS_KEY, WORKER_INTERNAL_KEY) migrated into /musehub/production/* as KMS-encrypted SecureString parameters (default alias/aws/ssm key, matching staging's existing 3 parameters) — same values already generated for production, just relocated to a durable, audited source of truth instead of only living in a file on disk.
    • Migration itself never exposed a secret value to the assistant: a temporary, narrowly-scoped ssm:PutParameter policy was granted to the instance role, the migration script read .env and wrote to SSM entirely server-side, then the policy was removed and replaced with the actual least-privilege runtime policy (ssm:GetParameter/GetParametersByPath + kms:Decrypt, scoped to exactly /musehub/production/*).
    • deploy/secrets.sh run for real against production — regenerated .env from SSM, verified identical structure/non-secret values, old .env auto-backed-up by the script itself. Containers were not restarted (unnecessary — values are unchanged, this only affects the next deploy or restart).
    • UVICORN_WORKERS is intentionally not set by secrets.sh; confirmed the entrypoint's own default (4) meets the test-enforced minimum and exceeds staging's manually-set 2 — not a regression.
    • Staging still only has 3 of 7 parameters in SSM — completing that migration is a reasonable follow-up but wasn't in scope for "day one done right" on the new production environment specifically.
  • [x] Configure CloudWatch logs, retention, alarms, and basic health monitoring. — /musehub/production log group created with 30-day retention set from day one (fixing the never-expire mistake found on staging's /musehub/staging group). Alarms/SNS still pending — planned alongside Phase 4/6.
  • [ ] Verify the instance can pull the approved image and access only its required production services. — Phase 4 work (first deploy).

Phase 4 — Deploy and validate before public DNS

  • [x] Bootstrap the instance: Docker 29.7.2 installed, nginx installed and configured (site musehub with hostnames musehub.ai www.musehub.ai), blue-green active-port file initialized to blue (1337), /opt/musehub app directory created. Completed 2026-08-24.
  • [x] Cloudflare Origin Certificate provisioned and installed. Decided 2026-08-24: rather than hunt for the original staging-era private key (unknown if it was ever saved anywhere retrievable — Cloudflare never re-displays a private key after creation), generated a new Origin Certificate (RSA 2048, 15-year validity, musehub.ai + *.musehub.ai) dedicated to the production host. This coexists fine with the existing cert staging uses; no conflict. Private key handling was done entirely by Gabriel via an interactive SSM session (aws ssm start-session) and the Cloudflare dashboard — never passed through the assistant's tool calls, consistent with not routing secrets through an AI agent's context. Verified live via SSM: nginx serves the correct Cloudflare Origin cert (issuer CloudFlare, Inc. / CloudFlare Origin SSL Certificate Authority, valid through 2041); HTTP 502 on /healthz is expected at this stage since no app container is running yet.
  • [x] Cloudflare account migrated to [email protected]. Decided/completed 2026-08-24: rather than move the musehub.ai zone/DNS/R2/Workers to a new Cloudflare account (high-risk, could cause live DNS downtime), added [email protected] as a Super Administrator member of the existing account (renamed from "[email protected]'s Account" to "MuseHub" — a cosmetic label change only, same underlying account ID bed873d46de5273abf843468a7833f09, zero resource migration). [email protected]'s membership is intentionally still present — removing it is deferred as a "nice to have," not required for the production launch. An unrelated, separate, empty Cloudflare account auto-created for the [email protected] identity (5760f45bb74895367e52da0e1ee2a2b2) was identified as harmless noise and left alone (no self-serve "delete account" button was found in the dashboard; not worth pursuing further).
  • [x] Selected the immutable application image digest known to work in staging: sha256:b57d94e44b5a785607d16d27c4641a51f2623584838848c6f28e5dfa1a23f743 (ECR tag 1705df12-20260808183652, confirmed still the live tag on both musehub-green and musehub-worker on staging at copy time). Promoted via crane copy directly from the Nonproduction ECR repo into the new dedicated Production ECR repo, preserving the exact tag plus an additional :latest tag — no rebuild, digest verified identical on both sides. This implements the "build once, promote the tested artifact" pattern the CI/CD research had flagged as missing.
  • [x] Production .env created at /opt/musehub/.env (root:root, 640 — tighter than staging's 644). Variable set mirrors staging's actual live .env (confirmed by listing variable names only via SSM, never values): DEBUG, DATABASE_URL, DB_PASSWORD, WEBHOOK_SECRET_KEY, RUNNER_TOKEN, WORKER_INTERNAL_KEY, UVICORN_WORKERS, BLOB_STORAGE_ENDPOINT, BLOB_STORAGE_ACCESS_KEY_ID, BLOB_STORAGE_SECRET_ACCESS_KEY, BLOB_STORAGE_BUCKET, BLOB_STORAGE_REGION, PUBLIC_URL, MUSE_ENV, CORS_ORIGINS. Secret-handling approach:
    • DB_PASSWORD, WEBHOOK_SECRET_KEY, RUNNER_TOKEN, WORKER_INTERNAL_KEY — pure-random values, generated directly on the instance via SSM (openssl rand / stdlib-only Fernet-equivalent) and written straight to .env server-side. Never printed to or captured by the assistant at any point.
    • BLOB_STORAGE_ACCESS_KEY_ID / BLOB_STORAGE_SECRET_ACCESS_KEY — a new, dedicated R2 API token (musehub-production-r2) was created in the Cloudflare dashboard scoped to Object Read & Write on the musehub-prod bucket only (least-privilege, matches this project's existing per-environment R2 bucket convention). Bucket musehub-prod did not exist before this session and was created fresh (Standard storage class, Automatic location) — production has never shared staging's musehub-staging bucket. Gabriel pasted these two values directly into the instance himself via an interactive SSM session (nano), same secret-handling principle as the TLS key — never passed through the assistant.
    • BLOB_STORAGE_ENDPOINT, BLOB_STORAGE_REGION, UVICORN_WORKERS — non-secret, mirrored from staging's actual values for consistency.
    • PUBLIC_URL=https://musehub.ai, MUSE_ENV=production, CORS_ORIGINS=["https://musehub.ai","https://www.musehub.ai"] — production-specific, per the canonical-hostname decision.
    • Known gap, deliberately deferred: WORKER_INTERNAL_KEY is a shared secret with a Cloudflare Worker (musehub-pack-receiver-staging is the only Worker route that currently exists, per the DNS records inventoried earlier — there is no production equivalent yet). A value was generated so the app starts cleanly, but the Worker→MuseHub callback path won't actually round-trip correctly until a production Worker is deployed and configured with this same key. This is separate follow-up work, not a blocker for getting /healthz green or validating the rest of the app.
  • [x] Deploy it to Production using the same blue/green pattern where practical. Completed 2026-08-24. Brought up the full stack via SSM (no repo checkout on the instance — purely image + generated config, cleaner than staging's legacy full-repo-copy pattern):
    • Docker network musehub_musehub-internal created; Postgres (postgres:16-alpine, network-alias postgres) started and healthy.
    • AWS CLI v2 installed on the instance (this was missing — deploy/bootstrap-instance.sh normally does this but wasn't run since a custom bootstrap script was used instead; documented here so it isn't missed on any future fresh instance).
    • 74 Alembic migrations applied cleanly to a brand-new database; schema gate passed (73 tables in sync).
    • musehub-set-slot installed at /usr/local/bin/; app container (musehub-blue, network-alias musehub) started, health-checked internally, then promoted live via musehub-set-slot blue (writes .active-slot + nginx upstream + reload — same script staging uses).
    • Fixed a real gap found live: the instance role initially had no CloudWatch Logs write permission, so the first container start failed with AccessDeniedException on logs:CreateLogStream. Added a scoped inline policy (logs:CreateLogStream/PutLogEvents/DescribeLogStreams on exactly the /musehub/production log group ARN) and retried successfully.
    • musehub-worker started (background job processor).
    • musehub-runner was deliberately not started. Per the prior CI-pipeline research, this container is scaffolding for a job-execution client that was never actually built (RUNNER_TOKEN/MUSEHUB_URL are wired but nothing polls MuseHub for jobs) — running it on production today would add a privileged docker:24-dind container for zero functional benefit. Revisit once Section 9's CI runner work is real.
    • Verified end-to-end: curl https://localhost/healthz (through nginx + the new Cloudflare Origin Cert + the app container) returns HTTP 200. All three containers (musehub-blue, musehub-worker, musehub_postgres) report healthy/running.
  • [x] Automated backups — done 2026-08-25, closing readiness ticket #157 (the single highest-priority item across the entire readiness effort). Real, verified, end-to-end — not just "a cron job exists":
    • Fixed a real bug found while wiring this up: deploy/backup.sh used docker compose exec postgres, but production has no docker-compose.yml at all — fixed to docker exec musehub_postgres directly.
    • Extended deploy/secrets.sh to also source BACKUP_R2_BUCKET from SSM (non-secret, plain String parameter) — needed since secrets.sh overwrites .env on every run, so a manually-appended line would've been silently wiped on the next refresh.
    • Dedicated backup bucket (musehub-production-backups) and a separate, narrowly-scoped R2 API token (musehub-production-backups-r2, Object Read & Write on that bucket only) — deliberately not reusing the app's own musehub-production-r2 credential, so a compromise of one doesn't grant access to the other.
    • rclone (apt version 1.53.3) configured with the r2 remote by Gabriel via an SSM session, matching the private-key handling pattern used throughout — never passed through the assistant. Hit and fixed a real gotcha: writes failed with 403 AccessDenied even though reads worked, because the object-scoped token lacks bucket-level permissions (HeadBucket/CreateBucket) that this old rclone version checks for by default before uploading — fixed with no_check_bucket = true in the remote config. Also had to copy the rclone config from the SSM session user's home directory to /root/ since backup.sh runs via sudo.
    • Ran backup.sh for real: produced a genuine 16K compressed dump (not the historical 20-byte stale/empty dump this whole effort kept finding) and successfully synced it to R2.
    • Installed the daily 3 AM cron job.
    • Performed an actual restore test (the ticket's explicit ask, not just "backups run"): restored the fresh dump into a completely separate, throwaway postgres:16-alpine container (zero risk to the live database), confirmed all 74 tables came back correctly (matching the schema gate's own count), then tore the scratch container down. Row counts on musehub_repos/musehub_identities were 0 — expected and correct, since production has no real user data yet, only hours after cutover. The pipeline itself (dump → compress → upload → download → restore → working schema) is what's being proven, and it's fully verified.
    • Staging still has zero backups — this work was production-only; replicating it for staging is a reasonable, lower-urgency follow-up.
  • [ ] Run database migrations only after:
    • [ ] Taking a backup/snapshot.
    • [ ] Reviewing forward and rollback behavior.
    • [ ] Confirming the target is the production database.
  • [ ] Test Production without changing public DNS using an Elastic IP, load-balancer hostname, temporary hostname, or local hosts-file override.
  • [ ] Validate:
    • [ ] Application health endpoint.
    • [ ] Home page and critical routes.
    • [ ] Authentication and authorization.
    • [ ] Database reads/writes.
    • [ ] Uploads and media delivery.
    • [ ] Email and webhook integrations.
    • [ ] Background jobs/queues.
    • [ ] Logs and alarms.
    • [ ] Restart and reboot recovery.
  • [ ] Run a rollback rehearsal to the previous production slot/image.

Phase 5 — Namecheap DNS and TLS

  • [ ] Confirm the final production hostname(s) and current DNS records in Namecheap.
  • [ ] Lower DNS TTL in advance of cutover if appropriate.
  • [ ] Create the appropriate record:
    • [ ] A record to an Elastic IP for direct EC2 ingress; or
    • [ ] CNAME/alias-compatible record to a load balancer hostname.
  • [ ] Verify propagation with dig/nslookup from more than one resolver.
  • [ ] Complete TLS according to the chosen topology:
    • [ ] Direct EC2: issue a Let's Encrypt certificate with Certbot only after DNS resolves correctly and ports 80/443 are reachable.
    • [ ] Load balancer: request and validate an ACM certificate, then attach it to the HTTPS listener.
  • [ ] Configure HTTP-to-HTTPS redirection.
  • [ ] Validate the full certificate chain, hostname coverage, and expiration.
  • [ ] For Let's Encrypt, confirm the renewal timer and run certbot renew --dry-run.
  • [ ] Re-run the complete production smoke test over the final HTTPS hostname.

Phase 6 — Make production deployment repeatable

  • [ ] Update the deployment tooling to require an explicit environment, such as staging or production.
  • [ ] Add expected-account guards:
    • [ ] Staging/nonproduction requires 992382692655.
    • [ ] Production requires 672469410277.
  • [ ] Require an explicit production release version/tag or immutable image digest.
  • [ ] Make production deployment use the production SSO profile and production-specific configuration.
  • [ ] Preserve the working blue/green deployment behavior and health gate.
  • [ ] Add a documented one-command or short-command rollback.
  • [ ] Prevent accidental use of staging secrets, hostnames, databases, or buckets in Production.
  • [ ] Ensure deploy logs clearly state the target account, region, environment, host, image digest, and resulting active slot without printing secrets.
  • [ ] Add CI/CD later only after the manual scripted flow is deterministic and verified.

Phase 7 — Cutover and post-launch

  • [ ] Capture the pre-cutover production backup/snapshot and rollback point.
  • [ ] Confirm owners, decision-maker, and rollback trigger.
  • [ ] Route production DNS only after all acceptance checks pass.
  • [ ] Monitor application health, error logs, CPU/memory/disk, database health, and external integrations during the launch window.
  • [ ] Verify AWS billing/cost telemetry after resources have run long enough to report.
  • [ ] Raise DNS TTL after the environment is stable, if it was lowered.
  • [ ] Document the deployed architecture, resource identifiers, operational commands, backup/restore procedure, and incident contacts.
  • [ ] Decide whether and when to retire the obsolete musehub-prod resource in Nonproduction. Treat retirement as a separate, approved task.

Acceptance criteria

  • [ ] Production resources reside only in AWS account 672469410277.
  • [ ] Staging remains functional and unchanged throughout the launch.
  • [ ] Engineers authenticate through Google Workspace and IAM Identity Center; no root access keys are used.
  • [ ] Production is accessible on its intended HTTPS hostname with a valid, renewable certificate.
  • [ ] Production secrets and state are isolated from staging.
  • [ ] A known-good image can be deployed to Production through a documented, guarded command.
  • [ ] Rollback is documented and successfully rehearsed.
  • [ ] Monitoring, logs, backups, budgets, and basic alerts are active.
  • [ ] Repository documentation accurately describes the final architecture and operating procedure.

Decisions that must be resolved during discovery

  • [x] Exact production hostname(s). — musehub.ai canonical, www.musehub.ai redirects to it (decided 2026-08-24).
  • [x] Direct EC2/Let's Encrypt versus load balancer/ACM. — Direct EC2 + Cloudflare Origin Certificate, matching staging (decided 2026-08-24).
  • [x] Existing infrastructure-as-code strategy and state location. — None yet; provision manually now, codify in Terraform/OpenTofu via import as a post-launch follow-up (decided 2026-08-24).
  • [ ] Production database and data-migration plan. — still open; fresh database confirmed (not copied from staging), execution is Phase 4.
  • [x] ECR promotion/cross-account strategy. — Dedicated Production ECR repository, promoted via crane copy of the staging-validated digest (decided and provisioned 2026-08-24).
  • [x] Production instance size and scaling requirements. — t3.medium, matching staging's proven size rather than the old undersized t3.small (decided and provisioned 2026-08-24).
  • [ ] Required OAuth, email, webhook, storage, and third-party production configuration. — still open, Phase 4 work.
  • [ ] Fate of the old musehub-prod resource in Nonproduction. — still open, deferred to Phase 7 per plan.

Agent execution checkpoints

An implementation agent should stop and report at these gates:

  1. After access verification: report both caller identities and profiles.
  2. After read-only discovery: deliver the inventory, dependency map, proposed design, cost-impacting resources, and unresolved questions.
  3. Before provisioning: obtain approval for the exact production plan.
  4. Before any state/data migration: provide backup and rollback procedures.
  5. Before DNS changes: provide validation results and the DNS/TLS cutover plan.
  6. After launch: provide smoke-test evidence, deployed image digest, resource summary, and rollback command.
File History 1 commit
sha256:cd5c2fcb91a44ac9e38e9c36176c27ce079a074586d9852296da628a29fb01ff Merge 'docs/aws-identity-and-deploy-fixes' into 'dev' — pro… Human 3 days ago