Files
marketplaces/docs/DEPLOYMENT.md
sdarbinyan 55634b3b57
Some checks failed
Architecture Governance / architecture (push) Has been cancelled
Deploy Frontend / deploy (push) Has been cancelled
Merge improvements/fork-harvest into main
Fork-harvest brings: the ip-api.com geo fix, credential bundle scan,
mock gateways out of production, JIT compiler dropped (1.55->1.04 MB),
host hardening, provider-agnostic identity + VK/Yandex + account linking,
and the backend contracts consolidated into one BACKEND-INTEGRATION.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

# Conflicts:
#	docs/backend/BACKEND-HANDOFF.md
#	docs/backend/TRACK-S-SECURITY-RBAC-CONTRACT.md
2026-08-22 16:19:55 +04:00

12 KiB

Deployment — server provisioning, CD, TLS

Frontend deployment plus API-domain edge configuration. The backend service is a separate developer's responsibility. API hostnames are separate reverse proxies and return 502 until their configured upstream exists (production currently defaults to https://127.0.0.1:445).

Multi-tenant, one bundle. Every customer domain is served by the same build. The SPA resolves its tenant from the Host header (BACKEND-HANDOFF §1a). One deploy updates every domain simultaneously — there is no per-tenant build and no per-tenant deploy.

The SPA derives one API origin from the storefront's base domain: example.com, store1.example.com, and www.example.com all use api.example.com. Tenant identity still comes from the complete storefront host; tenant subdomains do not create additional API DNS names.


1. Files

Path Purpose
scripts/deploy/server-setup.sh One-time server provisioning. Idempotent. Run as root.
scripts/deploy/add-domain.sh Attach one domain + issue TLS. Run per domain, as root, after DNS resolves.
scripts/deploy/configure-api-domain.sh Configure shared api.<base domain> TLS, storefront-origin CORS, backend proxy, and JSON bootstrap verification.
.github/workflows/deploy.yml CD: build → upload → atomic swap → verify. Triggers on push to main.

2. Layout on the server

/srv/marketplaces/
├── releases/
│   ├── a1b2c3d4e5f6/frontend/    <- one directory per deployed commit
│   └── ...                        (last 5 kept)
└── current -> releases/a1b2c3d4e5f6

nginx root is /srv/marketplaces/current/frontend. Activation is a symlink swap, so no request is ever served from a half-written directory, and a rollback is a symlink change rather than a rebuild.

On the current production host there is one extra hop. That server predates server-setup.sh and was provisioned by hand, so instead of the catch-all vhost it has per-domain configs (gorbushka.conf, dexarmarket.conf, gorbushka-admin.conf, gorbushka-landing.conf) whose root is /var/www/dexarmarket/browser. That path is itself a symlink:

/var/www/dexarmarket/browser -> /srv/marketplaces/current/frontend

so the release/current model above still holds and the workflow needs no per-host special-casing. Until 2026-08-22 browser pointed straight at one pinned release directory with no current in between, which is why two successfully-uploaded releases sat unserved.


3. First-time setup

3.1 Generate a CI deploy key

On your machine, not on the server:

ssh-keygen -t ed25519 -C "ci@marketplaces" -f ./marketplaces_deploy -N ""

Two files result. marketplaces_deploy.pub goes to the server; marketplaces_deploy (private) goes into CI secrets and nowhere else.

3.2 Provision the server

Copy scripts/deploy/ to the server and run:

sudo bash server-setup.sh --pubkey "$(cat marketplaces_deploy.pub)"

This installs nginx + certbot, creates a key-only deploy user with no password, writes the catch-all nginx config, opens 80/443/OpenSSH in ufw, and installs a root-owned, argument-validating API-domain helper. The deploy user may run that helper and reload nginx, but cannot replace the helper.

It also applies host hardening (added 2026-08-21, FH-D.3) — three drop-in files, so a re-run replaces its own config and never edits a distro file in place:

File Effect
/etc/ssh/sshd_config.d/10-marketplaces-hardening.conf Password and keyboard-interactive auth off, root key-only, no agent/X11 forwarding, MaxAuthTries 3, 30 s login grace
/etc/fail2ban/jail.d/marketplaces.local sshd, nginx-http-auth, nginx-bad-request jails — 5 failures in 10 min, 1 h ban
/etc/sysctl.d/99-marketplaces-hardening.conf No redirects or source routing, reverse-path filtering, SYN cookies, forwarding off, restricted kernel pointers and dmesg

Both accounts on this host are key-only by construction, so disabling password auth cannot lock anyone out — it only closes unlimited guessing against a credential nobody intended to exist. The script runs sshd -t before reloading and removes its own drop-in if the test fails, because a bad sshd config taking effect on a remote box is how people lock themselves out permanently.

Confirm after provisioning:

sudo fail2ban-client status sshd
sudo sshd -T | grep -E 'passwordauthentication|permitrootlogin|maxauthtries'

Verify before continuing:

curl -I http://<server-ip>/health

Expect 200. A placeholder page is served until the first real deploy.

3.3 Capture the host key

ssh-keyscan -H <server-ip>

The output is the DEPLOY_KNOWN_HOSTS secret. Pinning it means a rebuilt or impersonated server fails the deploy instead of being trusted silently.

3.4 Add CI secrets

Required for every deploy:

Secret Value
DEPLOY_HOST server IP or hostname
DEPLOY_USER deploy
DEPLOY_SSH_KEY contents of the private key file
DEPLOY_KNOWN_HOSTS output of ssh-keyscan -H <server-ip>

Required only when running the workflow with reconcile_api_domains on (§4.6) — a normal release deploy never reads these:

Secret Value
STOREFRONT_DOMAINS space-separated full hosts, e.g. gorbushka.market store1.example.com
CERTBOT_EMAIL operations email used for Let's Encrypt
BACKEND_UPSTREAM optional; defaults to https://127.0.0.1:445

When that step does run, point each base domain's shared API hostname at the server first. For gorbushka.market and store1.gorbushka.market, only api.gorbushka.market is required. The workflow deduplicates STOREFRONT_DOMAINS by base domain and deliberately stops before release activation if DNS, certificate issuance, nginx validation, or the JSON /bootstrap check fails.

3.5 Deploy

Push to main, or run the workflow manually with a ref. The workflow refuses to swap the symlink unless the uploaded release contains an index.html, so a failed upload leaves the previous release serving.


4. Domains and TLS — dynamic by default

Domains arrive continuously: one today, five tomorrow. Nothing here requires a person per domain.

HTTP already needs zero configuration. The nginx catch-all serves any Host, and the SPA resolves its tenant from that header. Point a domain's A record at the server and it works over port 80 immediately. Only TLS needs a certificate per name — that is the whole problem this section solves.

Two mechanisms, used together:

4.1 Wildcard — tenants on our own apex

One certificate covers every <slug>.<apex>. A new tenant subdomain is then live over HTTPS the moment DNS resolves, with no certificate work at all.

sudo bash setup-wildcard-tls.sh \
  --apex marketplaces.example.com \
  --email ops@example.com \
  --dns cloudflare --creds /root/cloudflare.ini

Wildcards require DNS-01 validation, so certbot must write a _acme-challenge TXT record. With a provider plugin (cloudflare, route53) renewal is unattended. --dns manual works but prompts for a TXT record at every renewal — fine to prove the setup out, not acceptable as a steady state.

Hostinger has no certbot plugin. If DNS lives there: either move DNS to a provider that has one (Cloudflare is free, minutes of work), or drive issuance from the Phase 9 domain-automation API once it exists.

4.2 Reconciler — tenants on their own domains

A wildcard cannot cover a customer's own domain. sync-domains.sh runs on a 10-minute timer and reconciles the live set against a desired list:

  • issues certificates for domains that lack one
  • skips domains whose certificate has more than 30 days left
  • skips subdomains already covered by WILDCARD_APEX
  • leaves domains alone while their DNS has not propagated yet, and retries next tick
  • disables server blocks for domains removed from the source — without deleting the certificate, so re-adding one later is instant
  • caps issuance per run, so a misconfigured source cannot burn the weekly ACME budget in a single pass

Configure /etc/marketplaces/domains.env:

DOMAINS_SOURCE=file:/etc/marketplaces/domains.txt
CERTBOT_EMAIL=ops@example.com
MAX_ISSUE_PER_RUN=10

Then:

sudo systemctl enable --now marketplaces-domains.timer
sudo /srv/marketplaces/bin/sync-domains.sh --dry-run   # see the plan, change nothing

Adding a domain becomes: append a line to /etc/marketplaces/domains.txt (or add the row in the backend registry), point DNS, wait one tick.

4.3 Backend-driven, once Phase 9 ships

Point the reconciler at the registry instead of a file and the loop closes — MarketplaceDomain already carries exactly the statuses this needs (planned → dns_pending → ssl_pending → active → failed):

DOMAINS_SOURCE=https://api.example.com/api/admin/v2/domains
DOMAINS_API_TOKEN=...

The script accepts a bare JSON array of hostnames, or objects with domain + status, in which case it acts only on active rows. A fetch failure aborts the run rather than reading as "remove every domain."

4.4 One-off

For a single domain, outside the reconciler:

sudo bash add-domain.sh shop.example.com --email ops@example.com --with-www

4.5 API domains in CD are opt-in

The deploy workflow's Reconcile tenant API domains step is gated behind the reconcile_api_domains input and is off for push-triggered deploys.

configure-api-domain.sh writes /etc/nginx/sites-available/api.<domain> and enables it. The API vhosts on the current production host were created by hand under different filenames (gorbushka-api.conf), so running the helper there produces a second server block for a server_name that already has one, and re-runs certbot against a live API — on every deploy. Shipping frontend files needs none of that.

Turn it on from the workflow-dispatch form only when standing up a new base domain. Before the first such run, reconcile the naming: either delete the hand-made vhost and let the helper own the name, or leave the step off and keep managing API domains manually.

4.6 Verify

curl -I https://shop.example.com/health
sudo certbot certificates
journalctl -u marketplaces-domains.service --since "1 hour ago"

5. Rollback

ssh deploy@<server-ip>
ls -1dt /srv/marketplaces/releases/*/     # newest first
ln -sfnT /srv/marketplaces/releases/<sha> /srv/marketplaces/current.new
mv -Tf /srv/marketplaces/current.new /srv/marketplaces/current
sudo systemctl reload nginx

Only the last 5 releases are retained. Older ones need a rebuild from the tag.

The production host reaches releases through /var/www/dexarmarket/browser -> /srv/marketplaces/current/frontend (§2), so moving current is all a rollback needs there too — do not repoint browser at a release directly, or the next deploy's swap will silently stop taking effect.


6. Operational checks

curl -I http://<host>/health                    # 200 from nginx
readlink -f /srv/marketplaces/current           # which commit is live
sudo nginx -t                                   # config valid
systemctl status nginx certbot.timer            # both active
sudo tail -f /var/log/nginx/marketplaces.error.log

7. Known limits

  • /api/ 502s until the backend runs. Expected. nginx proxies to 127.0.0.1:8080; nothing listens there yet.
  • No staging environment. main goes straight to production. Adding one means a second server plus a staging branch trigger.
  • No smoke test beyond HTTP 200. The verify step confirms nginx serves the shell, not that the app boots. A real check needs the E2E harness from Track Q.
  • Caching. index.html is no-store; hashed assets are immutable for a year. A deploy therefore takes effect on the next page load, with no cache purge.