Scaling

Where the load goes

Path Cost Scales by
Camera/selection/preview relay 1 Redis PUBLISH per message + fan-out on each subscribed node (O(participants) sends) More realtime nodes; Redis pub/sub throughput
Mutations 1 Lua script (a few hash/stream ops) + PUBLISH Redis CPU; sessions are independent (shardable by hash tag)
Presence state HSET on heartbeat (every 5 s per client) and on presence changes Redis
Persistence Batched (≤500) INSERT … SELECT FROM jsonb_to_recordset per batch Postgres write throughput; consumer group spreads across nodes
REST Indexed single-tenant queries API instances + Postgres read replicas later

Per-session ceilings (defaults)

  • Participants per session: plan-limited (2 – 32).
  • Camera traffic per moving camera ≤ 15 msg/s; per-connection hard caps in protocol.md.
  • Replay log: last 10 000 events per session (SESSION_LOG_MAX); older reconnects get a full state sync instead.

Growing

  1. Hundreds of sessions / thousands of users: add realtime and API replicas behind the LB. No code changes — nodes share nothing but Redis/Postgres.
  2. Redis CPU-bound: move to Redis Cluster. All per-session keys share the {sess:<id>} hash tag so every Lua script touches one slot; sessions spread across shards automatically. Pub/sub channels can move to sharded pub/sub (SSUBSCRIBE) — a contained change in node.ts.
  3. Postgres write-bound: the changes table is append-only and indexed by (project_id, created_at); partition it by month and/or move history to a columnar store. Retention purges run in batches.
  4. Global teams: run realtime clusters per region with session-to-region affinity (choose the region when the session is created; the ticket's URL points there).

What to watch

  • skein_rt_connections, skein_rt_sessions — capacity per node.
  • skein_rt_commit_seconds — Redis latency on the critical path (p99 should stay low single-digit ms).
  • skein_rt_dropped_total{reason} — backpressure/slow_consumer mean clients or the network can't keep up; rate_limit means abusive or buggy clients.
  • skein_rt_persist_errors_total and consumer-group lag (XPENDING changes:stream persister) — sync failures.
  • skein_http_request_duration_seconds, skein_db_pool_connections{state="waiting"}, skein_db_ping_seconds.

Load testing

npm run test:load -- --sessions 20 --clients 6 --seconds 60 (see tests/load) creates real users, projects and sessions, connects simulated editors that move cameras and edit objects at realistic rates, and reports end-to-end delivery latency percentiles and drop counts.

Measured results

Run on a 2-vCPU / 2 GB VM that was also running unrelated production services, with the load generator, the API and one realtime node in the same Node.js process (so every number includes the generator's own JSON parsing). Real Postgres and Redis (Docker). Every editor sends camera updates at the protocol maximum (15/s, i.e. constantly flying) — far above typical use, where cameras are mostly still.

Editors Sessions × size Camera frames in / out per s Delivered Camera latency p50 / p95 / p99 Edit latency p50 / p99 Drops
40 10 × 4 600 / 1 800 100% 4 ms / 15 ms / 33 ms 4 ms / 20 ms 0
120 20 × 6 1 640 / 8 200 100% 18 ms / 122 ms / 292 ms 25 ms / 150 ms 0

At 120 constantly-moving editors the single process is CPU-bound; latency grows but nothing is dropped. In production the generator isn't co-located and realtime nodes scale horizontally (see the two-node test).