Skip to content

feat(infra): split task into api + mongo + meili for zero-downtime api deploys - #7

Merged
wilfoa merged 6 commits into
mainfrom
feat/zero-downtime-split-task
Apr 20, 2026
Merged

wilfoa merged 6 commits into
mainfrom
feat/zero-downtime-split-task

Conversation

@wilfoa

@wilfoa wilfoa commented Apr 20, 2026

Copy link
Copy Markdown

Summary

Split the LibreChat Fargate task into three ECS services to enable zero-downtime api deploys:

Service Statefulness Deploy config Replicas
librechat-staging-api stateless rolling (200/100) scalable
librechat-staging-mongo stateful, EFS stop-before-start (module default for stateful) 1
librechat-staging-meili stateful, EFS stop-before-start 1

Api deploys are now drop-in: new task spins up alongside old, passes health checks, old drains. ~0 s downtime on the public endpoint. Mongo/meili redeploys still stop-before-start because of EFS WiredTiger locking — but they change approximately never, so it doesn't matter.

Security

  • Mongo auth enabled. --auth flag on boot; root user librechat created from MONGO_INITDB_ROOT_PASSWORD (random_password → Secrets Manager). API reads a pre-composed MONGO_URI secret.
  • Ingress scoped. Mongo's and Meili's Service Connect servers accept traffic only from librechat-api's task SG — not the cluster-wide shared internal-client SG. botnim-api (and any future internal-client app) cannot reach port 27017/7700 at the network layer.

Depends on

Build-Up-IL/org-infra#<pending> (branch feat/app-tcp-protocol) — three additions to modules/app:

  • internal_server.app_protocol = "tcp"
  • image_override (same-account ECR only)
  • internal_server.ingress_security_group_ids list

This PR pins the module source to that feature branch. Once the org-infra PR merges and feat/ecs-efs-and-sidecars-v2 merges to main, flip all three ?ref=feat/app-tcp-protocol references back to ?ref=main.

Deploy ergonomics

scripts/deploy-staging.sh no longer forces maximumPercent=100, minimumHealthyPercent=0 via a post-apply update-service override. The api inherits the shared module's default rolling configuration.

Follow-ups not in this PR

  • Per-service desired_count split so api can scale to N while mongo/meili stay at 1 (currently a single var.desired_count applied to all three — fine at 0 or 1).
  • Secret rotation automation for mongo-root-password (currently a manual ops task).
  • Mongo's POST-apply first-boot: the first task will create the root user from env vars; subsequent tasks reuse the existing user.

Test plan

  • terraform validate clean on infra/envs/staging/
  • org-infra module validates with new inputs (tcp, image_override, ingress_security_group_ids)
  • Deploy to staging, verify:
    • api service comes up, ALB health-checks green
    • api → mongo connection over mongodb://librechat:<pw>@mongo.ns:27017
    • api → meili connection over http://meili.ns:7700
    • subsequent api deploy (change image_tag) has 0 s 5xx on the public endpoint
    • botnim-api's task SG does NOT appear in the mongo/meili ingress rules

🤖 Generated with Claude Code

wilfamir and others added 6 commits April 20, 2026 10:17
…ime)

Api deploys now use the default rolling 200/100 config: a new api task
spins up alongside the old, passes health checks, then the old drains.
Zero downtime for the public endpoint. The deploy script drops the
stop-before-start override it was forced into when mongo/meili shared
the same EFS-backed task.

Mongo and meili move to their own ECS services (stateful, 1 replica
each, stop-before-start on their rare deploys). They publish via
Service Connect as tcp://mongo.<ns>:27017 and tcp://meili.<ns>:7700
with ingress restricted to the api's task SG only — no other app in
the shared cluster can reach them at the network layer.

Mongo launches with --auth. A random_password generates the root
credentials at apply time, stored in Secrets Manager alongside a
derived MONGO_URI secret the api reads. First-boot init creates the
librechat user; subsequent boots re-use the existing user.

Depends on modules/app@feat/app-tcp-protocol which adds three
capabilities to the shared app module:
 - internal_server.app_protocol = "tcp"
 - image_override for pre-mirrored upstream images
 - internal_server.ingress_security_group_ids for least-privilege

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
EFS access points move from /mongo + /meili → /mongo-v2 + /meili-v2 so
the new auth-enabled mongo boots against an empty dataDir (its init
script only creates the root user on a fresh volume; reusing the
pre-auth data would leave the api unable to authenticate).

Api env gains CREATE_BOOTSTRAP_USER=true + BOOTSTRAP_USER_PASSWORD so
the admin user is (re)created on first boot against the fresh mongo.
Re-seeding afterwards: agent via seed-botnim-agent.js, prompts via
migrate-prompts-into-db.js. Staging conversation history is lost —
acceptable tradeoff for the cutover.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Terraform cannot resolve count/for_each whose length depends on a
list of module-output values (Build-Up-IL/org-infra was blocking plan).
Defense-in-depth simplified: mongo --auth + random_password in Secrets
Manager, meili MEILI_MASTER_KEY — network stays cluster-wide
internal-client but app-layer auth blocks unauthorized reads.
…sources)

data.aws_ecs_task_definition / aws_iam_role lookups failed because the
resources they named no longer exist — the previous apply tore down
the old module.librechat task def before the new module.librechat_api
one was created. Reference module outputs directly so Terraform
resolves the dependency order.
…o staging values

Service Connect publishes client aliases under the short service name
(mongo:27017, meili:7700), not the FQDN mongo.buildup-staging.local.
Using the FQDN caused getaddrinfo ENOTFOUND on first boot — the api
task's /etc/hosts only has the short name.

Also update terragrunt fallback image_tag to the v0.8.4-botnim-split-v1
tag that's actually in ECR, and sync botnim_agent_id_unified with the
agent id seeded into staging (agent_rQ6pcpy-pvgu18HKs9oO9).

Verified end-to-end on https://botnim.staging.build-up.team:
- admin login ✓
- Hebrew budget question returns real BudgetKey data ✓
- Hebrew takanon question cites correct sections ✓
- Rolling deploy zero-downtime (0 failed probes during rollover) ✓

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@wilfoa
wilfoa merged commit e2b110d into main Apr 20, 2026
@wilfoa
wilfoa deleted the feat/zero-downtime-split-task branch April 20, 2026 09:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants