Background jobs

Learn how the Reaction async Command, async logging, and analytics queues work, how they retry, and where to troubleshoot them.

Saavo uses Cloudflare Queues for work that should not block the current request. The project includes three dedicated queues for asynchronous Commands generated by business events, asynchronous logging, and first-party analytics events.

Reaction is not a general-purpose asynchronous queue. An Event represents a business fact that has already occurred, and a Command represents an action the system should take next. Each Command determines whether it runs synchronously or asynchronously.

Warning

Do not insert custom events into these predefined queues. Each has its own purpose and processing logic. For asynchronous business operations, you can use Reaction's Event / Command mechanism or create a separate business queue.

Built-in queues

QueueBindingDefault batch size and concurrencyPurpose
async-policy-taskASYNC_POLICY_TASK_QUEUE1 message per batch, concurrency of 1Run asynchronous Reaction Commands
async-loggerASYNC_LOGGER_QUEUEUp to 50 messages per batch, concurrency of 1Write logs to KV
analytics-eventsANALYTICS_QUEUEUp to 8 messages per batch, concurrency of 1Write first-party analytics events to ANALYTICS_DB

All three consumers currently set max_retries to 3. This platform configuration allows up to three redeliveries after the initial delivery, for a total of four consumption attempts. Consumers can acknowledge a message earlier to stop retries. The Reaction consumer handles up to four deliveries, while the analytics and logging consumers acknowledge messages on the third processing failure and record an error or alert.

wrangler.jsonc configures the queue names, bindings, and *_QUEUE_NAME values. The entry point selects a consumer by an exact match on the actual queue name. If the names do not match, messages are not automatically routed to another queue.

How Reaction Commands run

Business code emits an Event through reaction.processor.emit(). The processor first persists the Event and Command execution records, then runs the Commands in order:

  • mode: 'sync': Attempt execution immediately in the current request, webhook, or scheduled entry point.
  • mode: 'async': Send the Event execution ID to async-policy-task for the consumer to continue execution.

Synchronous Commands are currently used primarily to grant or revoke roles and entitlements. Asynchronous Commands include sending emails, sending notifications, creating in-app notifications, synchronizing payment customer email addresses, processing entitlement cycles that are due, and cleaning up expired data.

Synchronous execution only means that the Command runs within the current call chain. It does not guarantee that a failure throws an exception. A normal failure returned by a Command is written to its execution record, and emit() may still return accepted. If the caller must ensure that a Command succeeded, it must check the returned Command status or read the execution record in the admin dashboard, in addition to checking whether the Event was accepted.

The admin dashboard provides two places to investigate:

/dashboard/reaction/events
/dashboard/reaction/commands

Retries, deduplication, and failure handling

Events use a stable idempotencyKey for deduplication. Emitting the same Event type with the same idempotency key again returns duplicate without creating a second set of execution records.

Each Command also has its own maxAttempts, usually defaulting to 3. This controls business-level retries of the Command handler. Cloudflare's max_retries controls queue message redelivery. These are separate retry layers and need to be distinguished when troubleshooting.

Queues use at-least-once delivery, so messages can arrive more than once. Reaction uses persistent state, stable execution IDs, and leases to reduce duplicate execution. Commands with external side effects, such as sending emails or calling third-party APIs, should still use stable business keys to make those operations idempotent.

The async-policy-task queue currently has no dead-letter queue. If a message still fails on its final delivery, the consumer acknowledges it and sends a high-priority alert. The execution records then need manual investigation. If an external side effect has completed but its result cannot be persisted, the consumer also stops redelivery and sends an alert to avoid repeating the side effect.

Enqueuing a job does not mean it has completed

An accepted result from emit() means that the Event and Commands have been accepted and dispatch has started. Asynchronous Commands may still be queued. An abandoned result means that the Event could not be reliably accepted after multiple attempts to emit it. The caller must not treat this as success.

Emitting events from business code

Business code should use Reaction instead of sending custom JSON directly to ASYNC_POLICY_TASK_QUEUE:

import { reaction } from '@/core/reaction';

const result = await reaction.processor.emit(
    workerCtx,
    SomeEvent.create({
        idempotencyKey: 'some-event:stable-business-id',
        payload: { /* Business data */ },
    }),
);

emit() can return the following results:

ResultMeaning
acceptedA new Event has been persisted, and execution or enqueuing has started
duplicateAn Event with the same idempotency key already exists, and its original execution record is returned
ignoredThe Event produced no Commands
abandonedThe emit operation ultimately failed, and logs and alerts have been recorded

Whether callers need to check these results depends on what the business operation requires before it can be considered complete. Critical flows such as granting access at signup or granting payment entitlements must check the corresponding synchronous Command status, rather than just confirming that the function returned. For asynchronous work that only sends notifications, the current request can end once the work has been reliably enqueued.

HTTP routes obtain their context through resolveFetchWorkerCtx(c). Cron and queue entry points use the WorkerCtx already created by their respective runtimes. Application code should not call createWorkerCtx directly.

Choosing how to process work

  • Business facts and their subsequent actions: Define Reaction Events / Commands.
  • Page visits and product tracking: Use the analytics API, which sends events to analytics-events.
  • Application logs: Use the existing logger, which sends logs to async-logger.
  • Long-running computations or general-purpose jobs unrelated to existing business operations: Evaluate the runtime platform separately instead of putting them directly into async-policy-task.

The logger.enableAsyncLoggerQueue setting in config/deploy.ts is enabled by default. Disabling it stops new logs from entering the asynchronous logging queue. The consumer acknowledges and discards any remaining messages it receives. Disabling analytics as a whole also causes the analytics consumer to acknowledge and discard events already in the queue. Before changing these switches, assess whether discarding the backlog is acceptable.

Do not increase batch sizes or concurrency without measurements. The async-policy-task queue uses one message per batch and a concurrency of one so that it can safely advance business Commands according to execution record order and leases.

Pre-launch checklist

  • All three queue names, bindings, and *_QUEUE_NAME values match exactly.
  • Producers and consumers are deployed with the same environment configuration.
  • Reaction execution records match expectations after signup, a test purchase, and the end of a subscription.
  • New Events use stable business idempotency keys.
  • External side effects can be retried safely without assuming that a queue delivers each message only once.
  • Critical callers recognize abandoned and check synchronous Command status as required by the business operation.
  • Before disabling logging or analytics, you have confirmed that remaining queue messages can be discarded.
  • Logging and alert channels receive notifications for final consumption failures and failures to persist results.

FAQ

Next steps