Task: Get all 18 services running end-to-end on Newton
Table of Contents
This page documents a task in the Offload service and DB runtime to a WSL host over SSH story. It captures the goal, current status, acceptance, and any notes or results.
1. Goal
Redeploy the containerized service runtime to Newton with everything
learned since PR #1747 closed with a partial-success caveat: the
ores.nats/service/client.cpp cert-load-error fix, and
--userns=keep-id (already used by docker/run-pod.sh locally, but
not used in the manual podman run --network=host commands used on
Newton in that earlier session). Determine whether this resolves the
remaining service failures there, or whether a genuinely distinct
process_supervisor corruption bug still needs chasing – important
caveat carried over from that session: on Newton the controller itself
demonstrably succeeded at NATS connect and launched children, unlike
the uid-remap failure reproduced locally, so this fix may not fully
explain Newton's failures and this needs verifying empirically, not
assumed.
2. Status
| Field | Value |
|---|---|
| State | DONE |
| Parent story | Offload service and DB runtime to a WSL host over SSH |
| Now | Nothing. |
| Waiting on | Nothing. |
| Next | Nothing. |
| Last touched | 2026-07-30 |
3. Acceptance
- All 18 supervised services reach phase
runninginpublic.ores_controller_service_instances_tblon Newton's Postgres, reproducibly across at least 3 consecutive restarts (the earlier session's results were unstable across runs). - The local Qt client connects to the services/DB running on Newton over the network and is usable end-to-end.
- If a distinct corruption bug remains after the uid-remap fix, documented with a precise repro (not just "unstable"), and filed as its own task rather than left as a vague caveat.
4. Plan
(Implementation strategy. Written when work starts; key decisions
are distilled into the parent story's * Decisions at close, but the
plan itself stays — it is the historical record of what we did.)
5. Notes
5.1. Everything the uid-remap fix alone didn't cover
--userns=keep-id (already used by docker/run-pod.sh) plus the
ores.nats/service/client.cpp fix got the fleet from 2-6/18 to close
to full, but several more real, distinct bugs surfaced getting all the
way to 18/18 reproducibly and then to an actual working end-to-end
scenario. None of these were the corruption this task set out to
chase down as a separate concern – there was no such distinct
corruption bug once the following were fixed:
- Newton-specific storage EIO: writing to
/app/log~/~/app/run(baked into the image's own overlay layer) hit a genuineEIOon Newton specifically, not reproducible locally – fixed by host-bind- mounting those paths (plus/app/storage, added for the same reason onceores.http.serverneeded it) instead of relying on the image's own copy. - DB connection exhaustion: repeated redeploys during this session
leaked stale connections faster than Newton's default
max_connections=100could absorb; bumped to 800 and stale connections terminated once. - Stale schema: Newton's DB predated some recent
mainschema changes (marketdata.feed_bindings.asset_classmissing); fixed with a fullcompass db recreateagainst Newton. - Missing
ores.http.server~/~ores.compute.wrapperin the staged image:docker/stage-runtime.sh's default glob (ores.*.service) never matched these two binaries at all (naming convention gap) – fixed, they're now explicitly included. ores.http.serveradvertising the wrong URL: defaults--http-base-urltohttp://localhost:{port}when not set, and it was never set in this service's DB-drivenargs_template– remote clients got told to reach the HTTP server on their ownlocalhost. Fixed by setting--http-base-url http://192.168.1.22:20200explicitly in thatargs_templaterow (deployment-specific DB data, not a code fix).- podman
--env-filedoesn't strip quotes: the JWT private key's.envline (KEY="...\n...", quoted to keep an embedded-newline PEM on one line for bash) got the literal surrounding quote characters passed through as part of the value under podman's env-file parser (unlike bash'ssource, which strips them) – corrupted the PEM, breaking IAM's own JWT signing. Fixeddocker/generate-env.shto strip one matching pair of quotes generically, and to pass this one variable via--env "$(cat build/keys/iam-rsa-private.pem)"instead of via the env-file at all, since podman's env-file format can't represent a multi-line value regardless of quoting (mirrors whatcompass_services.py'schild_env()already does for the native launch path). - NATS server cert SAN wrong type for an IP hostname:
compass nats certs --hostname 192.168.1.22added it as aDNS:SAN entry, notIP Address:– fails strict TLS hostname verification for any client connecting by IP (the Qt client,ores.shell). Fixednats_certs.pyto detect an IP-shaped hostname and emit the correct SAN type. - Compute dispatch tenant mismatch:
process_supervisor.cppsubstituted{tenant_id}with the NATSsubject_prefixstring (not a UUID) when buildingores.compute.wrapper's launch args, so its NATS subscription for dispatched work (compute.v1.work.assignments.{tenant_id}.{platform}) never matched what the server actually publishes to. Root-caused and fixed at the source: the whole compute domain (hosts, batches, workunits) is consistently created under the system tenant by design, so this now passestenant_id::system()explicitly.
5.2. Deliberately not fixed here – filed as captures instead
- Compute repositories should query the system tenant, not the caller's:
even with the dispatch fix above,
host_repository.cpp(and likelybatch_repository.cpp~/~workunit_repository.cpp) filter reads/writes by the caller's own session tenant rather than the system tenant the data actually lives under – a real business-tenant login (tenant_admin@acme) sees 0 compute hosts/tasks; onlysuper_adminhappens to work, since its session resolves to the system tenant already. Bigger fix (read+write paths, 3 repositories), explicitly scoped out of this task. - Vendored ORE engine package missing its own shared libraries:
external/ore/packages/ore-1.8.15.0-x64-linux.tar.gznever bundledlibOREAnalytics.so~/~libQuantExt.so~/~libQuantLib.so.1, so a dispatched compute job's engine process fails at startup. Confirmed reproducible natively (not container/Newton-specific) – a packaging gap upstream of this repo, not something to fix blind here.
6. Test Scenarios
Manual QA scenarios (scaffolded via compass add test_scenario, run
through the QA Validation Runner panel) that verify this task. Link
new ones here as they're created; the scenario doc itself links back
via its "Verifies task" field.
| Scenario | State | Notes |
|---|---|---|
| Acme end-to-end smoke test against the Newton deployment | PASSED | Currency CRUD + eventing + history, plus a compute job submitted and run to completion, all against the remote Newton deployment. Surfaced two further real bugs along the way, fixed/captured separately (see Notes). |
7. PRs
| PR | Title |
|---|---|
| #1779 | [docker] Get all 18 services running end-to-end on Newton |
8. Review
| # | Comment summary | File | Decision | Notes |
|---|---|---|---|---|
| 1 | No blocking issues on the initial diff | (none) | N/A | First review pass, no fixable findings. |
| 2 | run-pod.sh still passes JWT key via –env-file only, contradicting PR description | docker/run-pod.sh | Accepted | Added an explicit –env passthrough reading the real PEM file, mirroring compass_services.py's child_env(). |
| 3 | Same JWT key gap, independently confirmed with root-cause detail (escaped \n can't parse as PEM) | docker/run-pod.sh | Accepted | Same fix as #2. |
| 4 | stage-runtime.sh comment overstates the fix as DB-driven; silent skip on missing extra binary | docker/stage-runtime.sh | Accepted | Reworded comment to describe the actual glob+allow-list implementation; missing extra binary now logs a warning instead of silently skipping. |
9. Result
All 18 supervised services (22 rows counting ores.compute.wrapper's
5 replicas) reach phase running on Newton, confirmed reproducibly
across 3+ consecutive container restarts. No distinct
process_supervisor corruption bug remained once the fixes in Notes
were applied – the original PR #1747 partial-success caveat is fully
resolved.
The local Qt client connects to Newton's remote NATS/DB end-to-end and
is genuinely usable: the Acme end-to-end smoke test scenario passed in
full, including currency CRUD/eventing/history and a compute job
submitted and run to completion. Getting the scenario to actually pass
(not just the fleet reaching running) surfaced several further real
bugs beyond the acceptance criteria's original scope – all fixed at
the source except two genuinely separate, larger issues, filed as
captures rather than fixed blind in this session (see Notes).
Also set up, as a side effect of testing: two new connections.db
environment entries (swift_curie (local) and swift_curie (Newton))
with the full persona set mirrored from prime_origin, so this
checkout's Qt client can switch between local and Newton targets going
forward without re-deriving connection profiles each time.