| Branch | Status |
|---|---|
| develop | |
| main |
- Table of Contents
- Summary Of The Project
- Technologies
- Setup and Run
- Deployment configuration
- Debugging
service-print-renderer is a background worker service responsible for consuming print jobs from an SQS queue and rendering them into PDF documents. The state of the print job is being updated in a dynamodb.
When service-print-api receives a print request from a client, it enqueues a job to SQS and returns a job ID. The renderer continuously polls that queue, picks up pending jobs one at a time, launches a headless Chrome browser via Playwright, renders the web-portal page as a PDF, uploads it to S3 under the deterministic key <S3_PDF_PREFIX>/<job_id>.pdf, and updates the job status in DynamoDB to finished. The renderer does not store the PDF URL: because the S3 key is deterministic, service-print-api derives the PDF URL from the job_id once the status is finished. Clients can then query service-print-api with the job ID to check the status and retrieve the resulting document once it is ready.
Malformed SQS messages (unparseable body or missing job_id) are deleted directly from the main queue. Failed rendering jobs are not deleted — the worker lets the visibility timeout (SQS_VISIBILITY_TIMEOUT) expire so SQS redelivers the message and retries up to SQS_MAX_RECEIVE_COUNT times. Only on the final attempt is the job marked as error in DynamoDB; SQS then routes the message to the DLQ automatically via the redrive policy.
- AWS SQS - job queue
- AWS DynamoDB - job status tracking
- AWS S3 - PDF storage
- Playwright (Python) - browser automation
- Chrome headless - PDF rendering
- Python 3.14+
- uv
dockeranddocker compose
Copy the default env file and install dependencies:
make setupmake setup creates the virtual environment, installs all dependencies.
Start the local AWS stack (DynamoDB, SQS, S3) and create the required resources:
Note
Maybe you want to start the local stack from the project service-print-api. It starts exactly the same stack as in this project. Doing so, you have the possibility to test the entire print procedure.
make start-motoThis starts a moto server container and runs the following init containers:
| Container | Action |
|---|---|
init-dynamo |
Creates the DynamoDB table (DYNAMODB_TABLE_NAME) |
init-sqs |
Creates the DLQ (SQS_DL_QUEUE_NAME) and the main SQS queue (SQS_QUEUE_NAME) with a redrive policy pointing to the DLQ |
init-s3 |
Creates the S3 bucket (S3_BUCKET_NAME) with a public-read policy |
If a moto server is already running (e.g. started from service-print-api), make start-moto reuses it and only reruns the init containers.
To verify the S3 bucket was created:
AWS_ACCESS_KEY_ID=123 AWS_SECRET_ACCESS_KEY=123 aws s3 ls --endpoint-url http://localhost:5000make runThe worker polls the SQS queue continuously, processes incoming jobs, updates the dynamodb with the state of the process and uploads the resulting PDFs to S3.
make test # run tests with HTML coverage report
make test-ci # run tests with XML coverage report (used in CI)
make lint # run ruff linter + ty type checker
make ci-check-format # run ruff format and checks if any files changed (used in CI)The service is configured entirely via environment variables:
| Env | Default | Description |
|---|---|---|
AWS_LOCAL |
false |
Set to true to point AWS clients at the moto server instead of real AWS |
MOTO_HOST |
localhost |
Hostname of the moto server (local development only) |
MOTO_PORT |
5000 |
Port of the moto server (local development only) |
AWS_REGION |
eu-central-1 |
AWS region |
AWS_CONNECT_TIMEOUT |
5 |
Timeout in seconds for establishing a connection to AWS services |
AWS_READ_TIMEOUT |
30 |
Timeout in seconds for reading a response from AWS services |
DYNAMODB_TABLE_NAME |
service-print-jobs-local |
DynamoDB table storing print job status |
SQS_QUEUE_NAME |
service-print-jobs-queue-local |
SQS queue name |
SQS_DL_QUEUE_NAME |
service-print-jobs-dlq-local |
SQS dead-letter queue name |
SQS_MAX_RECEIVE_COUNT |
3 |
Number of times a message can be received before SQS routes it to the DLQ automatically |
SQS_VISIBILITY_TIMEOUT |
90 |
How long (in seconds) a received message is hidden from other consumers; after expiry SQS redelivers it (or routes to DLQ if maxReceiveCount is reached). Must exceed the worst-case render time so a slow job is not redelivered while still processing |
SQS_WAIT_TIME_SECONDS |
20 |
Long-polling wait time in seconds when reading from SQS |
SQS_DLQ_WAIT_TIME_SECONDS |
2 |
Polling wait time in seconds when reading from SQS DLQ |
SQS_MAX_MESSAGES |
1 |
Maximum number of messages to retrieve per SQS poll |
S3_BUCKET_NAME |
service-print-jobs-local |
S3 bucket where rendered PDFs are stored |
S3_PDF_CACHE_CONTROL_MAX_AGE |
3600 |
TTL in seconds for the Cache-Control: max-age header set on PDFs uploaded to S3; controls how long CloudFront and browsers cache the PDF |
PORTAL_URL |
- | web-portal endpoint (required). The per-job URL is built as <PORTAL_URL without trailing /?>/<print_lang>/print?<query>, e.g. https://www.dev.sgdi.tech/en/print?state=…&z=4&print_format=a3&… |
TIMEOUT_LOADING_WEB_PAGE |
30000 |
Browser page-load timeout in milliseconds |
USE_GPU |
false |
Set to true to use the local machine's GPU (native OpenGL) instead of the SwiftShader software rasterizer. For local development only |
BROWSER_RECYCLE_AFTER_JOBS |
10 |
Restart Chrome after this many jobs to prevent memory accumulation; set to 0 to disable |
BROWSER_NAVIGATION_RETRIES |
3 |
Number of times to retry page navigation on ERR_NETWORK_CHANGED before failing the job |
TMP_DIR |
/tmp |
Writable scratch directory used for the probe files, the temporary PDF of the job being rendered and the scratch data of Playwright and Chrome (temp profile, artifacts, caches). See read-only root filesystem |
The container runs as a non-root user and needs no write access to its root
filesystem, but Chrome, Playwright and the worker itself all need some writable
space. Everything they write goes underneath TMP_DIR, so a deployment with
readOnlyRootFilesystem: true only has to mount one writable volume and point
TMP_DIR at it:
securityContext:
readOnlyRootFilesystem: true
env:
- name: TMP_DIR
value: /scratch
volumeMounts:
- name: scratch
mountPath: /scratch
volumes:
- name: scratch
emptyDir: {}Mounting the volume at /tmp works just as well and needs no TMP_DIR. Note
that HOME stays on the read-only filesystem: XDG_CONFIG_HOME and
XDG_CACHE_HOME are pointed at TMP_DIR so Chrome does not try to write there.
The worker checks TMP_DIR on startup and exits immediately with an explicit
error if it is not writable.
The worker writes probe files to signal its state to Kubernetes:
| Probe | File (default) | Env var | Behaviour |
|---|---|---|---|
| Startup | $TMP_DIR/startup_probe |
STARTUP_PROBE_FILE |
Created once when the worker starts; never removed |
| Liveness | $TMP_DIR/liveness_probe |
LIVENESS_PROBE_FILE |
Touched after every polling and every printing cycle |
Configure the Kubernetes probes as exec checks:
startupProbe:
exec:
command: ["test", "-f", "/tmp/startup_probe"]
livenessProbe:
exec:
command: ["sh", "-c", "test $(( $(date +%s) - $(date +%s -r /tmp/liveness_probe) )) -lt 60"]The liveness check passes only if the file was touched within the last 60 seconds, catching a stalled worker even if the file still exists from a previous cycle.
You can verify the probe files manually while the worker is running:
# startup probe: exists once the worker loop has started
test -f /tmp/startup_probe && echo "started" || echo "not started"
# liveness probe: touched every polling cycle, must be < 60s old
test $(( $(date +%s) - $(date +%s -r /tmp/liveness_probe) )) -lt 60 && echo "alive" || echo "not alive"The worker exports traces and metrics via OpenTelemetry (OTLP) by default, and can also
export logs via OTLP when the OTEL logging config is used (see
Logging implementation). With OTEL_ENABLE_BOTOCORE=true, every
DynamoDB, SQS, and S3 call is also captured as a span. The service's custom metrics are
described under Metrics.
| Env | Default | Description |
|---|---|---|
OTEL_SDK_DISABLED |
false |
Set to true to disable all OTEL instrumentation |
OTEL_ENABLE_METRICS |
true |
Set to false to disable OTLP metrics export |
OTEL_METRIC_EXPORT_INTERVAL |
60000 |
Metric export interval in ms (read straight from the env by the OTEL SDK; only relevant when metrics are enabled) |
OTEL_METRIC_EXPORT_TIMEOUT |
30000 |
Metric export timeout in ms (read straight from the env by the OTEL SDK; only relevant when metrics are enabled) |
OTEL_EXPORTER_OTLP_ENDPOINT |
http://localhost:4317 |
OTLP gRPC endpoint of the collector |
OTEL_EXPORTER_OTLP_INSECURE |
false |
Set to true for an insecure (non-TLS) connection. Required for a plaintext local collector |
OTEL_EXPORTER_OTLP_HEADERS |
- | Optional headers for the OTLP collector (e.g. for authentication) |
OTEL_RESOURCE_ATTRIBUTES |
- | Extra resource attributes attached to all telemetry. |
When the OTEL logging config (app/config/logging-cfg-otel.yaml) is used, logs are exported
through the OpenTelemetry LoggerProvider (the otel handler).
The meter provider is set up in app/helpers/otel.py and exports via OTLP
alongside traces and logs. Instruments are defined in
app/helpers/metrics.py (scope.name = app.helpers.metrics,
scope.version = 1.0.0); bump METRICS_SCHEMA_VERSION on any schema change. Two follow the
OpenTelemetry semantic conventions for messaging
metrics; the third is
custom because the conventions model has no queue-wait instrument. All three reuse the convention's
attribute keys.
| Metric | Type | Unit | Attributes | Description |
|---|---|---|---|---|
messaging.client.consumed.messages |
Counter | {message} |
messaging.operation.name = print,messaging.system = aws_sqs,error.type = processing-retries-exceeded (permanent failures only) |
Print jobs the renderer finished with. One message is one print job |
messaging.process.duration |
Histogram | s |
messaging.operation.name = print,messaging.system = aws_sqs,error.type = processing-retried (failed, will be retried) or processing-failed (failed, last attempt) |
Time the renderer spent processing one message (render + upload). Excludes the queue wait |
swissgeo.messaging.queue.duration |
Histogram | s |
messaging.operation.name = print,messaging.system = aws_sqs,error.type = processing-retried (redeliveries only) |
Time from a message being sent to SQS to it being received by the renderer |
For the two semantic-convention instruments the name, unit and description are not written out as
literals. They come from opentelemetry-semantic-conventions, so a spec update propagates on the
next dependency bump. messaging.operation.name is the domain operation (print), not the SQS
API call.
messaging.client.consumed.messages is recorded once per job, at its terminal outcome: a
successful render, or a permanent failure once the SQS redrive policy is exhausted
(ApproximateReceiveCount reaches SQS_MAX_RECEIVE_COUNT), the latter carrying
error.type = processing-retries-exceeded. Redeliveries in between are not counted, so the
series without error.type is exactly the jobs that rendered successfully. Jobs that only ever
hit infrastructure errors crash the worker and are redriven to the DLQ without being counted here.
This pairs with the API's messaging.client.sent.messages: sent counts enqueue attempts,
consumed counts jobs picked up and finished, so the two can be compared as rates.
messaging.process.duration is measured with time.perf_counter() around the processing in
handle_message and recorded once per processing attempt - a redelivered job adds a sample
per attempt, so its _count is attempts (not distinct jobs) and _sum accumulates a job's total
processing time across retries. A failed attempt carries error.type = processing-retried
while it still has retries left, then processing-failed on the final attempt - so
processing-retried is the time burned on work that gets redone.
swissgeo.messaging.queue.duration is custom: the messaging conventions have no instrument for
the time a message sat in the queue (messaging.client.operation.duration times the receive
call, not the wait). It is now - SentTimestamp (the SQS system attribute, epoch ms), clamped
at 0 for clock skew, recorded once per delivery:
- first delivery (no
error.type): The clean queue wait, sent as first pickup. - redelivery (
error.type = processing-retried): SQS does not resetSentTimestamp, so this is the message's total age: it spans the failed attempt(s) and theirSQS_VISIBILITY_TIMEOUTwaits. Kept as its own series so it never pollutes the first-pickup percentiles; it quantifies how far behind a job that needed retries has fallen.
A CloudWatch ApproximateAgeOfOldestMessage alarm still has its place as this metric goes quiet
exactly when nothing is being consumed.
Locally, make start-otel also brings up Prometheus (http://localhost:9090) so metrics can
be queried and graphed — the OTLP collector only has a debug (log-dump) exporter otherwise.
Prometheus does not scrape the worker: the collector pushes metrics into Prometheus' native
OTLP receiver, so instrument names, units and attributes arrive unchanged.
OTEL names are rewritten on the way in: . becomes _, counters gain _total, histograms are
split into _bucket / _count / _sum, and annotation units ({message}) are dropped while
s becomes a _seconds suffix. So the metrics above are
messaging_client_consumed_messages_total, messaging_process_duration_seconds_* and
swissgeo_messaging_queue_duration_seconds_*. Both service-print processes share the
label job="service-print" (from service.name); otel_scope_name="app.helpers.metrics"
isolates the renderer's instruments from the API's.
# Raw count since the process started. Use this to check a job was counted at all
messaging_client_consumed_messages_total{otel_scope_name="app.helpers.metrics"}
# Print jobs rendered successfully, per second
sum(rate(messaging_client_consumed_messages_total{
otel_scope_name="app.helpers.metrics", error_type=""}[5m]))
# Permanent-failure ratio (jobs that exhausted their SQS retries)
sum(rate(messaging_client_consumed_messages_total{otel_scope_name="app.helpers.metrics", error_type!=""}[5m]))
/ sum(rate(messaging_client_consumed_messages_total{otel_scope_name="app.helpers.metrics"}[5m]))
# p95 processing time over 1m (render + upload), successful attempts
histogram_quantile(0.95, sum by (le) (rate(
messaging_process_duration_seconds_bucket{
otel_scope_name="app.helpers.metrics", error_type=""}[1m])))
# p95 queue wait over 1m (first pickup — exclude the retry-cycle series)
histogram_quantile(0.95, sum by (le) (rate(
swissgeo_messaging_queue_duration_seconds_bucket{
otel_scope_name="app.helpers.metrics", error_type=""}[1m])))
-
Start the local OTEL collector, Jaeger and Prometheus:
make start-otel
-
Run the worker. Traces and metrics export by default; to also export logs via OTLP, point
LOGGING_CFGat the OTEL logging config:LOGGING_CFG=app/config/logging-cfg-otel.yaml make run
View the full traces in the Jaeger UI at http://localhost:16686 and metrics in the
Prometheus UI at http://localhost:9090. Stop the stack with make stop-otel.
The OTEL stack (collector, Jaeger, Prometheus) is shared with
service-print-apivia theservice-print-local-otelcompose project — the compose file is identical in both repos, somake start-otelfrom either service brings up (or reuses) the same containers.
To verify that headless Chrome can access WebGL and report the expected renderer, run the worker with the -i / --renderer-info flag:
make renderer-info
# or directly: uv run python -m app.worker --renderer-infoThis launches a headless Chrome instance, evaluates a WebGL probe, logs the hardware-acceleration status and the renderer name, then exits. No queue polling or AWS calls are made.
nginx, one config file: Every request (any path, query, method) returns the same 200 HTML that fires gaMapReady:
# dummy-portal.conf
server {
listen 80 default_server;
location / {
default_type text/html;
return 200 '<!doctype html><html><head><meta charset="utf-8"><title>dummy print</title></head><body><h1>dummy print page</h1><script>window.postMessage({type:"gaMapReady"},"*");</script></body></html>';
}
}edit .env:
PORTAL_URL=http://localhost:8090/?Run it:
docker run -d --name dummy-portal -p 8090:80 \
-v "$PWD/dummy-portal.conf:/etc/nginx/conf.d/default.conf:ro" \
nginx:alpine