Background queues are one of the quiet workhorses of a Laravel application. Emails, payment captures, report exports and webhook calls all run there, out of sight of the user. That also means queue bugs tend to stay out of sight. A detailed post on DEV Community by the team behind laravel-o11y, a Laravel observability project, digs into one of the most common and least visible of them: the relationship between a job's timeout and the queue connection's retry_after setting. Get the two in the wrong order and Laravel can run the same job twice, or mark a successful job as failed, without any error pointing at the real cause.

Two settings, two owners

The confusion starts with naming and placement. The timeout belongs to the worker. It is set on the queue:work command, in a Horizon supervisor in config/horizon.php, or on the job class itself through a timeout property or attribute. It limits how long one job may occupy a worker process, and when it is exceeded the worker kills itself.

retry_after belongs to the queue connection in config/queue.php. Despite the name, it is not a delay before retrying. It behaves like a lease. When a worker pops a job, the job is reserved until now plus retry_after seconds. If that reservation expires, the job becomes available to any worker again, whether or not the first worker is still running it.

As the article explains with reference to the framework source, on Redis the reservation is stored in a sorted set scored by its expiry time. Nothing watches that score actively. Instead, every time a worker pops from the queue, Laravel first sweeps expired reservations back onto the ready list. It cannot tell a job whose worker crashed during a deploy from a job that is simply still running. The database driver reaches the same result through its query for the next available job.

The defaults are closer than they look

Laravel's own Horizon documentation says the timeout should always be at least a few seconds shorter than retry_after, otherwise jobs may be processed twice. The stock configuration respects that: the default worker timeout is 60 seconds, and the Redis connection in the shipped config/queue.php uses 90 seconds for retry_after.

The article points out a subtle trap. That 90 lives in your configuration file. If you add a connection by hand and leave out retry_after, the framework falls back to 60, the same as the default worker timeout. And equal values are not safe. The lease clock starts the moment the job is popped, while the timeout alarm is armed a few steps later, after the reservation is written and reservation events have fired. With identical values, the lease always expires first.

Queue drivers differ too. Amazon SQS ignores retry_after entirely, because the lease is the queue's visibility timeout configured on the AWS side, which defaults to 30 seconds. That is shorter than Laravel's default worker timeout, so a queue created with console defaults and processed by a stock worker is misconfigured from day one.

Two symptoms, chosen by the tries setting

What happens after a lease expires depends on how many attempts the job is allowed.

  1. With a single try, which is common in Horizon setups, the second worker sees an attempt count above the limit and immediately fails the job with MaxAttemptsExceededException. Meanwhile the first worker finishes the work successfully. You end up with a failed job in the dashboard for work that actually completed, and a stack trace that blames a timeout that never happened.

  2. With two or more tries, the second worker simply runs the job again while the first is still running it. Both copies complete, nothing is logged as an error, and the only evidence is in your data: a customer charged twice, two identical emails, duplicate rows.

The article also notes why this is hard to reproduce locally. Expired reservations are only swept when someone pops the queue, so with a single worker busy on the long job, nobody notices. The bug appears when a second worker exists, typically in production or after autoscaling.

How correct configurations drift

The most practical part of the article lists the everyday changes that break a once-correct setup. A slow export gets its own longer timeout on the job class, which takes precedence over the supervisor value while the connection's lease is unchanged. A supervisor timeout is raised to stop jobs being killed, but retry_after lives in a different file and is forgotten. A new connection is added without retry_after and inherits the 60-second fallback. Or the queue moves to SQS, where the setting no longer applies at all. Setting a timeout of zero is also dangerous, because it disables the alarm rather than triggering it.

Going the other way and disabling retry_after on Redis avoids duplicates but means jobs whose workers die are never recovered.

The ordering rule

The rule the article proposes is a simple chain: job timeout must be lower than the supervisor timeout, which must be lower than retry_after, with real margin at every step. Its worked example starts from a measured worst-case runtime of 180 seconds and sets the job timeout to 240, the supervisor timeout to 300 and retry_after to 390. Erring on the long side for retry_after has a small cost, because a job whose worker genuinely died waits a little longer before recovery. Erring on the short side causes duplicate execution.

Because Laravel does not enforce the rule, the article suggests writing a test that iterates over Horizon supervisors and asserts that each timeout is below its connection's retry_after. Such a test cannot see per-job timeout overrides, so runtime logging that records when a reservation expired and a job was handed to another worker is the complementary safeguard. The post also describes how the authors' own package surfaces these events.

Why it matters

Duplicate execution is one of the most expensive kinds of bug because it is silent and lands directly on customers and finances. It also undermines trust in your monitoring: a dashboard that shows a successful run while money moved twice is worse than no dashboard.

Practical takeaways for Laravel teams

Treat retry_after as a lease and set it explicitly on every connection rather than relying on framework fallbacks.

Base timeouts on measured worst-case runtimes, then keep the chain of job timeout, supervisor timeout and retry_after strictly ordered with generous gaps.

On SQS, set the visibility timeout on the queue itself to exceed your longest job timeout.

Review any change that touches a timeout as a change to the whole chain, and add an automated test for supervisor values.

Finally, make jobs idempotent wherever money or messages are involved, using unique keys or checks before side effects, so that if a job does run twice, the second run does no harm.


Source: laravel-o11y, DEV Community, original article linked below.

Cover photo: Ciara Ní Riain, CC BY-SA 4.0, via Wikimedia Commons.