The worker finishes a document, saves the result, and crashes before it deletes the queue message. A second worker receives that message later. Should it process the document again? The answer depends on where the result was saved and how the second worker can recognize it. A box labeled "queue" settles neither point.
Consider a fictional document-processing service. The client submits a document reference to an API and, if the request succeeds, gets a job ID. A worker extracts text from the document. The client asks the API for job status until it can read the result. This example uses an Amazon SQS standard queue and one database for job status and extracted text. It leaves authentication, document storage, capacity, and monitoring outside the view.
Trace one job before it fails
The useful first view has six elements: client, API, status and result database, SQS queue, worker, and dead-letter queue. Its arrows need verbs. The client submits a reference to the API; the API records the job and arranges for its ID to reach SQS. SQS delivers that ID to a worker. The worker reads the referenced document and writes extracted text and status to the database. Separately, the client polls the API, which reads that database. The dead-letter queue receives messages under an SQS redrive policy after repeated unsuccessful receives. It is not the worker's ordinary output path.
Suppose job doc-42 is pending. Worker A receives its message, extracts the text, and commits the text and a succeeded status in one database transaction. Then A crashes before deleting the SQS message. The client can already see success. When the message becomes visible again, worker B may receive it. SQS standard queues can also deliver a duplicate before that timeout expires, so the timeout cannot serve as a lock or guarantee one delivery. AWS on standard queue delivery; AWS on visibility timeouts.
The diagram should show the worker checking durable job state before accepting work and before committing a result. For this example, the result row has a unique key on job ID. The worker commits a result and changes pending to succeeded in the same database transaction, only while the job is still pending. A duplicate delivery of doc-42 finds the committed result, skips the extraction, and deletes its own queue receipt. If two workers race, the database constraint and conditional status change allow at most one committed result for that job. One worker may still waste computation before losing the race. "Idempotent" here means repeat deliveries cannot create a second committed result for the same ID, not that the worker runs exactly once.
That boundary matters. If extraction also charges an external provider or sends a notification, the database transaction cannot roll back that outside effect. If the provider supports an idempotency key, the job ID may help guard its operation too; otherwise the design needs another way to reconcile uncertain outcomes. A stable job ID by itself prevents nothing. The effect being protected, the durable check, and the system that enforces uniqueness belong in the explanation.
Separate a failed attempt from a failed job
Now move the crash earlier. If A fails before the database transaction commits, doc-42 remains pending. The message can be delivered again after its visibility timeout, and B can try the work. A temporary read failure may clear on that attempt. A malformed document may fail every time. The queue cannot tell those cases apart from the diagram alone.
With an SQS redrive policy, maxReceiveCount sets the receive limit before SQS moves an undeleted message to a dead-letter queue. It counts receives, not business failures. That route is from the source queue under queue policy; the worker does not send every exception there. A dead-letter queue isolates a message for diagnosis. Its presence does not prove extraction failed or automatically update the job's status row. AWS on dead-letter queues and redrive policy.
For this service, an operator-facing process must reconcile dead-letter messages with their job IDs. A job whose result committed before repeated deletion failures is already successful. A pending row alone does not show whether a worker is still running. The design needs a rule for when an active attempt can no longer commit; only then should the process conditionally mark that job failed. The worker's own conditional commit must respect that terminal status. Otherwise the reconciler can call a job failed while its result is being saved. Preserve the error context for investigation. Until reconciliation runs, a client may still see pending even though normal retries have stopped. The team has to choose a visibility timeout that fits the extraction work, decide when to extend it, and choose a receive limit that gives temporary faults a chance to recover without hiding persistent faults indefinitely. Those are workload decisions, not numbers a generic queue diagram can supply.
Account for the gap before the queue
There is another failure before any worker runs. The API could save a pending status and crash before publishing the message. If it has already returned the job ID, the client can poll a job that no worker will see. If it crashes before replying, the client may time out while a pending row remains, without knowing its job ID. Retrying the submission could create a second job unless the API recognizes the same client request. Reversing the order creates the opposite gap: a worker could receive a message before the status row exists. Two independent writes do not become atomic because the diagram draws them close together.
One plausible choice is a transactional outbox. The API saves the pending job and an "enqueue this job" record in the same database transaction. A separate publisher reads that record, sends the job ID to SQS, and marks the record delivered only after SQS accepts it. If the publisher crashes after sending but before marking delivery, it may send twice; the worker's job-ID guard still matters. The outbox gives the system a durable record to retry, provided the publisher resumes, keeps retrying, and someone notices records that remain unsent. It does not make the database and SQS one transaction or tell a client the job ID after a lost API response. AWS on transactional outboxes.
This is where the diagram becomes useful in a review. Label the database transaction, the outbox publisher, the SQS redrive route, the worker's conditional result write, and the client's status read. A reviewer can then point to the exact gap behind a duplicate result or a job stuck at pending. In Lycana, you can speak or type a change to the components and relationships as the failure path becomes clearer, then inspect the revised drawing with the team. If the processing effect moves outside the database, revisit the duplicate guard before trusting the same picture.


