How to decide whether a new job needs another process, and how to run that process with Apple Container or Docker Compose. For implementing a platform sync job itself, see jobs.md.
There are three different roles in this application:
| Role | Command | Responsibility | How many? |
|---|---|---|---|
| HTTP server | Insights serve ... |
Handles API and dashboard requests | One or more |
| Scheduler | Insights queues --scheduled |
Runs every registered clock-based schedule | Exactly one |
| Named worker | Insights queues --queue <name> |
Drains one named queue | At least one per queue name in use |
Adding another AsyncScheduledJob does not require another scheduler container. Register it
with app.queues.schedule(...), and the existing --scheduled process runs it alongside all the
other schedules.
Adding another AsyncJob usually does not require another worker either. If it is dispatched
onto the existing metrics queue, the existing --queue metrics worker processes it.
Add a new worker service only when the job is intentionally dispatched to a new queue name,
such as emails, exports, or maintenance.
%%{init: {"theme":"base","themeVariables":{"primaryColor":"#DBEAFE","primaryTextColor":"#172554","primaryBorderColor":"#2563EB","secondaryColor":"#DCFCE7","secondaryTextColor":"#14532D","secondaryBorderColor":"#16A34A","tertiaryColor":"#F3E8FF","tertiaryTextColor":"#581C87","tertiaryBorderColor":"#9333EA","lineColor":"#64748B","noteBkgColor":"#FEF3C7","noteTextColor":"#78350F","actorBkg":"#E0E7FF","actorBorder":"#4F46E5","actorTextColor":"#1E1B4B","signalColor":"#475569","signalTextColor":"#334155"}}}%%
flowchart LR
S[One scheduler container]
R[CollectDueResources<br/>hourly due scan]
A[CollectAccountStats<br/>monthly]
E[WarnExpiringServiceTokens<br/>daily, alerts inline]
Q[(Valkey: metrics)]
W1[metrics worker 1]
W2[metrics worker N]
J[Registered sync jobs]
N[FailureNotifier]
S --> R
S --> A
S --> E
R --> Q
A --> Q
E --> N
Q --> W1
Q --> W2
W1 --> J
W2 --> J
A job whose resource or account no longer exists logs at notice and returns, rather than
throwing. QueueWorker decides whether to retry from the remaining attempt count alone — there is
no per-error hook, and the retry budget is fixed at dispatch — so a thrown error would be retried
four times across roughly ten minutes to rediscover a row that is gone.
A deleted subject is not a failure the job can recover from; it is work that no longer needs doing. The usual cause is a resource deleted between dispatch and execution, occasionally a queued payload outliving the database state that produced it.
Scheduled jobs stay small: they query for work and enqueue typed jobs. Workers own remote API requests, so a slow synchronization cannot delay the scheduler's next tick.
Valkey stores the queued payloads. PostgreSQL stores application data and collection due dates. The HTTP server handles requests independently from the queue workers.
All clock expressions are registered in Sources/Insights/configure.swift. The single
insights-scheduled container evaluates them; insights-queues executes everything dispatched
to the metrics queue.
| Platform / scope | Dispatcher | Clock and eligibility | Executed job | Queue / worker | Metric state behavior |
|---|---|---|---|---|---|
| GitHub repository resources | CollectDueResources |
Sweep hourly at minute 00; each resource runs only when nextCollectionAt <= now. Its configured interval defaults to 7 days and is capped at 14. |
SyncGitHubRepoStats |
metrics / queues --queue metrics |
Stars, forks, and subscribers are snapshots. Clone/view daily values fold into all-time rows through MetricWatermark, excluding the incomplete UTC day. |
| Hugging Face resources | CollectDueResources |
Same hourly due-resource sweep; interval defaults to 7 days and is capped at 30. | SyncHuggingFaceHubStats |
metrics / queues --queue metrics |
Likes and rolling downloads are snapshots; the provider's lifetime downloads value replaces the stored all-time value with setAllTime. No daily watermark is needed. |
| GitHub accounts | CollectAccountStats |
First day of every month at 03:00; every GitHub account is eligible because accounts have no nextCollectionAt. |
SyncGitHubOrgStats |
metrics / queues --queue metrics |
Followers are a current-value snapshot. No watermark or resource interval applies. |
| GHCR resources | Planned | Awaiting schedule and dispatch routing. | SyncGHCRStats prototype |
Queue pending | Activation requires registration, routing, tests, and a queue decision for the HTML-scraping workload. |
| npm resources | None | Not scheduled; dispatchSync logs and skips them. |
None | None | No collection implementation yet. |
| Webhook token expiry | WarnExpiringServiceTokens |
Daily at 07:00; warns at 14, 7, 3, and 1 days remaining. | None — alerts inline | Scheduler only | Writes nothing. Queries live, unexpired tokens and calls FailureNotifier; critical at 3 days or fewer, warning above that. |
| PyPI resources | None | Not scheduled; dispatchSync logs and skips them. |
None | None | No collection implementation yet. |
The hourly resource schedule is a due-work scanner. After a successful dispatch,
scheduleNextCollection(from:) advances that
resource by collectionIntervalDays. The scheduler can therefore support different per-resource
cadences without registering a separate clock for each platform.
The ResourceController.create handler dispatches a resource sync immediately and then books
its next collection date. The POST route currently awaits authentication and authorization;
enabling the route activates creation-time dispatch.
Schedule times use the scheduler process's calendar/time zone. Container deployments should set
and document an explicit TZ value if 03:00 must mean a particular local zone; otherwise treat
the deployment's configured zone as authoritative.
Use an AsyncScheduledJob when work begins on a clock rather than from an HTTP request or
another job:
struct RemoveExpiredData: AsyncScheduledJob {
func run(context: QueueContext) async throws {
// Query for work and preferably dispatch jobs to a named queue.
}
}Register its cadence in Sources/Insights/configure.swift:
app.queues.schedule(RemoveExpiredData()).daily().at(2, 30)That is the only container-related step. The existing services already run:
Insights queues --scheduledDo not create one scheduler per scheduled job. Multiple scheduler replicas run the same set of schedules and can dispatch duplicate work. Keep the scheduler at one replica unless the jobs have explicit distributed locking and are designed for concurrent scheduling.
An AsyncJob needs to be registered so the worker can decode and execute its payload:
app.queues.add(SyncNewPlatformStats())Dispatch it through the existing queue:
try await context.queues(.metrics).dispatch(
SyncNewPlatformStats.self,
.init(id: id)
)The existing worker already covers this job type:
Insights queues --queue metricsOne worker can execute every registered job type placed on metrics; workers are selected by
queue name, not by Swift job type.
A separate queue is useful when one workload needs independent concurrency, scaling, resource limits, or failure isolation. Examples include:
- slow exports that should not delay metric collection;
- email work that needs different retry or scaling behavior;
- CPU-heavy jobs that deserve their own container limits;
- high-priority work that must not sit behind a large metrics backlog.
A new job type can share an existing queue. Introduce another named queue when its independent scaling, isolation, or resource policy justifies another worker service.
If the queue-name extension does not already contain the new name, add it using the Queues API convention used by the project, then dispatch explicitly:
try await context.queues(.exports).dispatch(
BuildExport.self,
.init(id: id)
)Register BuildExport with app.queues.add(...) as usual.
Copy the structure of the queues recipe in justfiles/apple-container.just, but give the
container and queue distinct names:
exports: build
container stop exports >/dev/null 2>&1 || true
container delete exports >/dev/null 2>&1 || true
container run --detach --name exports \
--network '{{container_network}}' \
{{container_app_env}} \
'{{container_image}}' queues --queue exportsThen add exports to the stack-scheduled dependency list and to the loop in stop.
container_app_env holds the environment every application container shares, layered in
precedence order: --env-file .env first, --env-file .env.container second, then the --env
flags for the in-network DNS names. Repeated --env-file flags merge, with the later file
winning on conflicts, and an explicit --env outranks both regardless of position. That is what
lets .env hold the credentials — it is gitignored, and the only place they belong — while the
tracked .env.container overrides the deployment database values with local ones. A missing
--env-file is a hard error rather than a skipped file, so .env must exist; copy .env.example.
Add a service using the shared environment and the new queue name:
exports:
image: insights:latest
build:
context: .
environment:
<<: *shared_environment
depends_on:
- db
- valkey
command: ["queues", "--queue", "exports"]Compose starts this service with docker compose up. No new Valkey instance or database is
needed; all named queues can share the existing Valkey service.
It is safe to run multiple workers for the same named queue when more throughput is needed. Valkey atomically moves a queued payload to one worker's processing list, so two workers cannot claim the same available payload at the same time. No application-level coordination or queue configuration is required: both processes consume the same named list. Scale the worker, not the scheduler:
docker compose up --scale queues=3For deployments that honor the Compose deploy section, a fixed replica count can also be
recorded on the worker service:
queues:
# image, environment, dependencies, etc.
command: ["queues", "--queue", "metrics"]
deploy:
replicas: 2For Apple Container, create multiple containers with unique names that all execute
queues --queue metrics:
container run --detach --name insights-queues-1 ... \
icicle-insights queues --queue metrics
container run --detach --name insights-queues-2 ... \
icicle-insights queues --queue metricsEach Apple container is a lightweight VM with its own configured CPU and memory allocation. On an orchestrated deployment, scale worker replicas with the same image, command, and queue name.
Atomic claiming does not mean exactly-once execution. A worker can fail after an external side effect but before acknowledging the payload, and retry handling can execute it again. Jobs must remain safe to retry.
Horizontal workers also introduce real concurrency. This project's
Metric.foldDailyIntoAllTime takes a PostgreSQL advisory transaction lock keyed by resource and
metric type, so two folds cannot corrupt the same all-time value while unrelated resources can
still proceed concurrently. New read-modify-write jobs need equivalent concurrency protection.
Worker scaling has two practical limits:
- They consume a provider's API allowance concurrently. GitHub 403 responses or an exhausted
x-ratelimit-remainingvalue call for throttling or delayed retries, not more workers. - A single shared queue still has head-of-line pressure: a backlog of slow HTML-scraping jobs can delay API jobs. Split by provider/workload when independent scaling or isolation becomes valuable; job type alone is not a reason to create another queue.
The resulting rule is: drainers scale horizontally; the scheduler stays single-replica.
Start the complete local Apple Container stack:
just stackInspect each role independently:
container logs -f insights-scheduled
container logs -f insights-queues
container logs -f insights-app
container exec insights-valkey valkey-cli pingThe scheduler only runs a scheduled job when its registered clock fires. For resource metrics,
CollectDueResources additionally selects only rows whose nextCollectionAt is in the past.
Changing a resource's collectionIntervalDays does not make it immediately due; it controls the
next date assigned after a successful collection.
Use the one-shot recipes rather than editing and restoring production schedules:
just collect-now # resources whose nextCollectionAt is already due
just collect-all-now # force every active resource due, then dispatch
just collect-accounts-now # run the monthly account dispatcher onceImmediate runs still use metric watermarks. Triggering controls when a sync is enqueued;
Metric.foldDailyIntoAllTime still decides which completed daily values are new. The complete
test procedure, database effects, log commands, and native equivalents are in
testing-collection.md.
For automated tests, call a job's dequeue or scheduled job's run directly with test
contexts, as the existing queue suites do. For local manual testing, prefer a dedicated one-shot
application command that invokes the same scheduled-job code over temporarily changing the
production cadence in configure.swift.
When adding queue work, ask these in order:
- Is it an
AsyncScheduledJobor anAsyncJob? - Is every
AsyncJobregistered withapp.queues.add(...)? - Is every scheduled job registered once with
app.queues.schedule(...)? - Which named queue receives the dispatched job?
- Does that queue already have a worker service?
- If the queue is new, was a worker added to both Apple Container and Compose?
- If the queue is new, was it added to Apple Container
stackandstop? - Is there still exactly one scheduler replica?
- Can the job safely retry without corrupting or duplicating data?
- Are scheduler and worker logs visible in the deployment?
#icicle-insights# #queues# #scheduling# #containers# #operations# #developer-documentation#