Skip to content

Add ccJobAttempts, retry strategy, and metrics for attempts for failed jobs - #48

Open
jonathanjouty wants to merge 7 commits into
masterfrom
dev-jj-add-job-structure-2
Open

jonathanjouty wants to merge 7 commits into
masterfrom
dev-jj-add-job-structure-2

Conversation

@jonathanjouty

@jonathanjouty jonathanjouty commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Started from a discussion with @Siprj on having consumers_job_failed_attempts, with us realising the attempts information is not cleanly available to consumers-metrics-prometheus, which led me to revive #5 but with a simplified approach suggested by @arybczak of having a new field ccJobAttempts on ConsumerConfig instead of changing the jobs parameter.


This PR adds three things to the consumers library:

  • ccJobAttempts, a new field on ConsumerConfig. It returns the attempt count for a job.
  • A new module, Database.PostgreSQL.Consumers.RetryStrategy, with four ready-made retry functions:
    • constantBackoff: the same delay before every retry.
    • linearBackoff: the delay grows by a fixed amount on each attempt.
    • exponentialBackoff: the delay doubles on each attempt, up to a set maximum.
    • exponentialBackoffWithJitter: the same as exponentialBackoff, plus a random adjustment. This spreads out retries from different consumer instances, so they do not all retry at the same moment.
  • A new metric in consumers-metrics-prometheus: consumers_job_failed_attempts. This histogram records the attempt count for jobs that do not finish with success. It uses ccJobAttempts.

Rebase Jonathan's 2021 proof of concept (github.com//pull/5)
onto current master. ConsumerConfig's job type is now always Job idx info,
exposing jobIndex, jobRunAt, jobFinishedAt and jobAttempts (the queue
bookkeeping columns already maintained by reserveJobs) alongside the
caller-supplied payload as jobInfo, instead of leaving attempts opaque to
everything outside the reservation query.

Update ccJobFetcher/ccProcessJob/ccOnException/ccJobLogData signatures
accordingly, and add exponentialBackoff, a ready-made ccOnException handler
built on jobAttempts. Update the example and test consumers to select the
new columns and construct Job values; fix a crash in the ported
diffTimeToInterval (the original used `^` with a negative Integer exponent).

This only touches the consumers package; consumers-metrics-prometheus is
deliberately left as-is for now, since it doesn't need any changes to keep
compiling (it only threads job values through opaquely).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016NeDfW7w3TwyS1UM82uvkL
…empt

Revert the previous Job idx info restructuring (63e226f) in favour of a
much smaller approach worked out with a colleague: add a single new
selector ccJobAttempts :: job -> Int to ConsumerConfig, mirroring
ccJobIndex. Existing job types don't change shape at all, only need to
also select the attempts column and expose it via this new field.

Add Database.PostgreSQL.Consumers.RetryStrategy as a separate module with
ready-made ccOnException handlers built on ccJobAttempts:
  - constantBackoff: fixed delay every retry
  - linearBackoff: delay grows linearly with attempt count
  - exponentialBackoff: delay doubles each attempt, capped at a max delay
  - exponentialBackoffWithJitter: like exponentialBackoff, but with random
    jitter to avoid multiple consumer instances retrying in lockstep after
    a shared dependency recovers

Update the example and test consumers to select attempts, wire up
ccJobAttempts, and use exponentialBackoff/constantBackoff respectively.

consumers-metrics-prometheus is still untouched; ccJobAttempts is what a
future failed-attempts histogram there would read from.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016NeDfW7w3TwyS1UM82uvkL
@jonathanjouty jonathanjouty self-assigned this Sep 22, 2026
claude and others added 3 commits September 22, 2026 22:03
Add a new histogram, consumers_job_failed_attempts, labelled by job_name,
observing ccJobAttempts for jobs whose result wasn't Ok (i.e. Failed, an
exception, or an abort). This is the metric Jan asked about: it lets you
tell one-off failures (attempt 1) apart from jobs stuck in a persistent
retry loop, to denoise alerting/dashboards for e.g. the document sealing
queue.

Reuses the existing reportJob release action from generalBracket
alongside consumers_job_execution_seconds, now also passed the job value
so it can read ccJobAttempts. Bucket boundaries are configurable via the
new jobFailedAttemptsBuckets field on ConsumerMetricsConfig, same pattern
as jobExecutionBuckets.

Requires consumers >= 2.4.0.0 for ccJobAttempts.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016NeDfW7w3TwyS1UM82uvkL

@jonathanjouty jonathanjouty left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-review

@@ -1,2 +1,10 @@
# consumers-metrics-prometheus-1.1.0.0 (2026-??-??)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note: TODO release date

Comment thread consumers/CHANGELOG.md
@@ -1,4 +1,11 @@
# consumers-2.3.5.0 (2026-??-??)
# consumers-2.4.0.0 (2026-??-??)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note: TODO release date

-- consecutive failed attempts, including the current one: it's reset to 1
-- once a job has succeeded, so it's a streak of failures rather than a
-- lifetime total. See "Database.PostgreSQL.Consumers.RetryStrategy" for
-- ready-made 'ccOnException' handlers built on it.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: seems like dear Claude made this a bit longer than it needs to be

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Try to use my rules wrt. code comments etc. from https://github.com/arybczak/claude/blob/master/CLAUDE.md and tell it to rewrite the text in this PR using current rules, that should help.

@jonathanjouty
jonathanjouty marked this pull request as ready for review September 23, 2026 14:11
@jonathanjouty jonathanjouty changed the title WIP: Add ccJobAttempts, retry strategy, and metrics for attempts for failed jobs Add ccJobAttempts, retry strategy, and metrics for attempts for failed jobs Sep 23, 2026

@arybczak arybczak left a comment •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For the record, haskell-ci shenanigans are solved by #49 (merged soon).

This definitely needs a downstream kontrakcja, journey and eid service (AFAIR they have a lot of consumers) PRs, ideally not only adding the new field, but also using these new backoff functions where they fit and to potentially see if they need any tweaking. Would be a shame to merge this and then these utils staying unused.

-- consecutive failed attempts, including the current one: it's reset to 1
-- once a job has succeeded, so it's a streak of failures rather than a
-- lifetime total. See "Database.PostgreSQL.Consumers.RetryStrategy" for
-- ready-made 'ccOnException' handlers built on it.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Try to use my rules wrt. code comments etc. from https://github.com/arybczak/claude/blob/master/CLAUDE.md and tell it to rewrite the text in this PR using current rules, that should help.

-- ^ Function that transforms the list of fields into a job.
, ccJobIndex :: !(job -> idx)
-- ^ Selector for taking out job ID from the job object.
, ccJobAttempts :: !(job -> Int)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I guess this should be Int32? so that downstream doesn't have to convert everything.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should it be? The library just says

attempts [...]
[...] Needs to be not nullable, of type INTEGER.

I don't remember what the PG type maps, or can map to.

@arybczak arybczak Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

integer maps to Int32 only.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants