Add ccJobAttempts, retry strategy, and metrics for attempts for failed jobs - #48
jonathanjouty wants to merge 7 commits into
Conversation
Rebase Jonathan's 2021 proof of concept (github.com//pull/5) onto current master. ConsumerConfig's job type is now always Job idx info, exposing jobIndex, jobRunAt, jobFinishedAt and jobAttempts (the queue bookkeeping columns already maintained by reserveJobs) alongside the caller-supplied payload as jobInfo, instead of leaving attempts opaque to everything outside the reservation query. Update ccJobFetcher/ccProcessJob/ccOnException/ccJobLogData signatures accordingly, and add exponentialBackoff, a ready-made ccOnException handler built on jobAttempts. Update the example and test consumers to select the new columns and construct Job values; fix a crash in the ported diffTimeToInterval (the original used `^` with a negative Integer exponent). This only touches the consumers package; consumers-metrics-prometheus is deliberately left as-is for now, since it doesn't need any changes to keep compiling (it only threads job values through opaquely). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016NeDfW7w3TwyS1UM82uvkL
…empt
Revert the previous Job idx info restructuring (63e226f) in favour of a
much smaller approach worked out with a colleague: add a single new
selector ccJobAttempts :: job -> Int to ConsumerConfig, mirroring
ccJobIndex. Existing job types don't change shape at all, only need to
also select the attempts column and expose it via this new field.
Add Database.PostgreSQL.Consumers.RetryStrategy as a separate module with
ready-made ccOnException handlers built on ccJobAttempts:
- constantBackoff: fixed delay every retry
- linearBackoff: delay grows linearly with attempt count
- exponentialBackoff: delay doubles each attempt, capped at a max delay
- exponentialBackoffWithJitter: like exponentialBackoff, but with random
jitter to avoid multiple consumer instances retrying in lockstep after
a shared dependency recovers
Update the example and test consumers to select attempts, wire up
ccJobAttempts, and use exponentialBackoff/constantBackoff respectively.
consumers-metrics-prometheus is still untouched; ccJobAttempts is what a
future failed-attempts histogram there would read from.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016NeDfW7w3TwyS1UM82uvkL
Add a new histogram, consumers_job_failed_attempts, labelled by job_name, observing ccJobAttempts for jobs whose result wasn't Ok (i.e. Failed, an exception, or an abort). This is the metric Jan asked about: it lets you tell one-off failures (attempt 1) apart from jobs stuck in a persistent retry loop, to denoise alerting/dashboards for e.g. the document sealing queue. Reuses the existing reportJob release action from generalBracket alongside consumers_job_execution_seconds, now also passed the job value so it can read ccJobAttempts. Bucket boundaries are configurable via the new jobFailedAttemptsBuckets field on ConsumerMetricsConfig, same pattern as jobExecutionBuckets. Requires consumers >= 2.4.0.0 for ccJobAttempts. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016NeDfW7w3TwyS1UM82uvkL
jonathanjouty
left a comment
There was a problem hiding this comment.
Self-review
| @@ -1,2 +1,10 @@ | |||
| # consumers-metrics-prometheus-1.1.0.0 (2026-??-??) | |||
There was a problem hiding this comment.
Note: TODO release date
| @@ -1,4 +1,11 @@ | |||
| # consumers-2.3.5.0 (2026-??-??) | |||
| # consumers-2.4.0.0 (2026-??-??) | |||
There was a problem hiding this comment.
Note: TODO release date
| -- consecutive failed attempts, including the current one: it's reset to 1 | ||
| -- once a job has succeeded, so it's a streak of failures rather than a | ||
| -- lifetime total. See "Database.PostgreSQL.Consumers.RetryStrategy" for | ||
| -- ready-made 'ccOnException' handlers built on it. |
There was a problem hiding this comment.
Nit: seems like dear Claude made this a bit longer than it needs to be
There was a problem hiding this comment.
Try to use my rules wrt. code comments etc. from https://github.com/arybczak/claude/blob/master/CLAUDE.md and tell it to rewrite the text in this PR using current rules, that should help.
There was a problem hiding this comment.
For the record, haskell-ci shenanigans are solved by #49 (merged soon).
This definitely needs a downstream kontrakcja, journey and eid service (AFAIR they have a lot of consumers) PRs, ideally not only adding the new field, but also using these new backoff functions where they fit and to potentially see if they need any tweaking. Would be a shame to merge this and then these utils staying unused.
| -- consecutive failed attempts, including the current one: it's reset to 1 | ||
| -- once a job has succeeded, so it's a streak of failures rather than a | ||
| -- lifetime total. See "Database.PostgreSQL.Consumers.RetryStrategy" for | ||
| -- ready-made 'ccOnException' handlers built on it. |
There was a problem hiding this comment.
Try to use my rules wrt. code comments etc. from https://github.com/arybczak/claude/blob/master/CLAUDE.md and tell it to rewrite the text in this PR using current rules, that should help.
| -- ^ Function that transforms the list of fields into a job. | ||
| , ccJobIndex :: !(job -> idx) | ||
| -- ^ Selector for taking out job ID from the job object. | ||
| , ccJobAttempts :: !(job -> Int) |
There was a problem hiding this comment.
I guess this should be Int32? so that downstream doesn't have to convert everything.
There was a problem hiding this comment.
Should it be? The library just says
attempts [...]
[...] Needs to be not nullable, of type INTEGER.
I don't remember what the PG type maps, or can map to.
There was a problem hiding this comment.
integer maps to Int32 only.
Started from a discussion with @Siprj on having
consumers_job_failed_attempts, with us realising theattemptsinformation is not cleanly available toconsumers-metrics-prometheus, which led me to revive #5 but with a simplified approach suggested by @arybczak of having a new fieldccJobAttemptsonConsumerConfiginstead of changing thejobsparameter.This PR adds three things to the
consumerslibrary:ccJobAttempts, a new field onConsumerConfig. It returns the attempt count for a job.Database.PostgreSQL.Consumers.RetryStrategy, with four ready-made retry functions:constantBackoff: the same delay before every retry.linearBackoff: the delay grows by a fixed amount on each attempt.exponentialBackoff: the delay doubles on each attempt, up to a set maximum.exponentialBackoffWithJitter: the same asexponentialBackoff, plus a random adjustment. This spreads out retries from different consumer instances, so they do not all retry at the same moment.consumers-metrics-prometheus:consumers_job_failed_attempts. This histogram records the attempt count for jobs that do not finish with success. It usesccJobAttempts.