Live cross-host filesystem-check contract (diagnostics.stat_paths)#
ops-db-api #111, epic data-transfer #139. Worker side: data-transfer #141.
Context#
The DB-only diagnosis (#110) reports what the database believes about an
operation’s files. The next tier asks the per-site hosts whether the bytes are
actually on disk right now. ops-db-api owns the admin trigger/poll API; the
data-transfer workers own the actual stat on each host (they alone are
deployed at the sites, with the mounts and credentials). The two repos are
separate codebases — ops-db-api must not import ccat_data_transfer — so
the only thing binding producer and worker is a wire contract, which this ADR
pins.
Two further constraints shaped the design:
Broker vs. buffering Redis. In observatory deployments
ccat_ops_db_api_settings.redis_hostisredis-ccat, the read-only transaction-buffering Redis — not the Celery broker. The existingdependencies.get_redis()points at that buffering Redis. Dispatching the stat task there would put it on a Redis no worker consumes from.No new DB table. The correlation between a poll request and its dispatched tasks is short-lived operational state, not a durable record.
Decision#
The contract between the ops-db-api producer and the data-transfer worker is:
Task name —
diagnostics.stat_paths.Args — a single positional argument: the list of path strings to stat. The producer calls
send_task("diagnostics.stat_paths", args=[[p1, p2, …]]).Per-site queue routing — one task per relevant location, routed to the owning
Site.short_namequeue (cologne/us/fyst). These are the site-level static queues that data-transfer’sdisk_monitordeclares and the stat worker (#141) consumes from — NOT the compound{short_name}_{name}_{operation}queues thatroute_task_by_locationbuilds for the transfer/monitor tasks.Result shape — the worker returns a list of per-path dicts, each
{path, exists, size, readable, error}(size/readablenull when absent,errora string when the stat itself failed).
The producer (ccat_ops_db_api/live_check/producer.py) is a producer-only
Celery app built against a new, separate broker connection:
New settings
celery_broker_host/celery_broker_port/celery_broker_db/celery_broker_ssl_*, kept explicitly distinct fromredis_host. They point at the MAIN broker (the broker data-transfer’s workers consume from) and reuse the existing/etc/redis/certs/SSL material (redis_ca_cert/redis_certfile/redis_keyfile). They default to localhost so local dev and CI need no broker.The producer’s result backend uses the same
celery_broker_dbindex as the broker. This index MUST match the data-transfer worker’sREDIS_DB: the worker writes its stat result to itsREDIS_DB, and the poll endpoint readsAsyncResultfromcelery_broker_db. A mismatch silently lands results in a different db and the poll stayspendingforever (no error surfaces).It dispatches purely by task name via
send_task; it registers no worker code.get_redis()is not reused for dispatch.
Only DISK locations are filesystem-checked. A site can own both a disk
buffer and an S3 long_term_archive under the same short_name (the production
Cologne topology). Since the worker stats with os.stat, an S3/tape object key
would always read as absent and falsely report a present archive as “missing”.
The producer skips non-disk copies and surfaces them in the trigger response’s
skipped list so the caller knows they were not live-checked (the #110 DB
diagnosis still covers them).
The correlation mapping {correlation_id → [(task_id, host_label, queue), …]}
is a TTL’d Redis key on the MAIN broker Redis (the producer’s own sync
client, producer.get_broker_redis()), never on the buffering redis-ccat. The
poll endpoint reads the mapping and each task’s AsyncResult, then reduces each
host’s per-path dicts to one status: error if any path errored, else missing
if any path is absent, else present; an unready task is pending; a vanished
correlation key is state="expired". A finished task with an empty/falsy
result is error, not present — every dispatched host has ≥1 path, so an
empty result means the payload expired/was evicted from the result backend, and
reporting present there would be a false all-OK on a diagnostic tool.
Contract test and its limitation#
A TRUE cross-repo round-trip (API → real data-transfer worker) is not
feasible in ops-db-api CI: the worker isn’t installed here and can’t be
imported. tests/test_stat_paths_contract.py is the strongest self-contained
equivalent and exercises the real producer send_task path:
The producer dispatches over an in-memory broker (
memory://). The test drains the per-site queue and inspects the published message, pinning the task name, the args ([[paths]]), and the queue routing on the wire. (task_always_eageris deliberately not used: Celery’ssend_taskignores it —AlwaysEagerIgnored— so the memory transport is what removes the live broker.)A local task registered under the exact name
diagnostics.stat_pathsproduces the{path, exists, …}result, which is fed through the real poll-assembly logic to prove the result shape drives the per-host status.
Limitation: the worker body in the test is a stand-in, not data-transfer’s
real implementation. The contract is pinned by shape and routing, not by the
worker’s actual stat behaviour. When data-transfer #141 lands, the real
worker must register a task under this exact name, accept args=[[paths]],
consume the cologne/us/fyst queues, and return the per-path dict shape.
Any change to name, args, routing, or result keys is a breaking change to this
contract and must update both repos together.
Consequences#
Two Redis connections from ops-db-api: the buffering
redis-ccat(get_redis(), unchanged) and the new MAIN broker (producer). Operators must configurecelery_broker_*to the real broker per environment; the localhost default is a dev/CI convenience, not a production value.The live tier is best-effort and ephemeral: if the broker is unreachable the trigger returns 503; if the correlation key expires the poll returns
expired. No durable record is kept (no new table), matching the operational, throwaway nature of a “is it on disk right now?” check.celerybecomes a direct dependency of ops-db-api (it was previously only transitive).