Delivery Tracking & Acknowledgment

Reliable push notification infrastructure requires deterministic tracking of every dispatch event. Delivery tracking establishes the boundary between message submission and service ingestion, enabling precise state management, auditability, and scalable queue orchestration. This guide details production-grade implementation patterns, secure logging architectures, and validation strategies for high-throughput web push systems.

Prerequisites

Acknowledgment to ledger state flow A dispatch carries a correlation ID to the push service, which returns an HTTP status that the acknowledgment handler maps to a ledger state: accepted leads to pending delivery, 4xx to terminal states, 5xx and 429 to retryable. Dispatch correlation_id Push Service HTTP status Ack Handler map to state PENDING 201 / 200 RETRYABLE 429 / 5xx TERMINAL 410 / 404 / 4xx
The acknowledgment handler is a pure function from HTTP status to ledger state — accepted, retryable, or terminal.

1. The Role of Acknowledgments in Push Delivery

Web push delivery tracking relies on synchronous HTTP acknowledgments returned by browser push services immediately after payload submission. Unlike email or SMS gateways, push services only confirm queue ingestion, not end-user receipt or rendering. Understanding this architectural boundary is foundational to building a resilient Backend Delivery Architecture & Queue Management that scales without over-provisioning compute resources or misinterpreting delivery states.

Core Principles:

  • Ingestion vs. Rendering: A 201 response guarantees the push service accepted the payload into its internal queue. It does not guarantee device wake, network delivery, or user interaction. (Some older push services return 200; RFC 8030 specifies 201 Created.)
  • Ledger State Mapping: Map push service HTTP responses to internal delivery ledger states immediately upon receipt. Treat acknowledgments as immutable facts.
  • Idempotent Correlation: Attach cryptographically secure UUIDv4 correlation IDs to every outbound request. These IDs must survive retries, network partitions, and service restarts to maintain end-to-end audit trails.

Compliance Directive: Log only subscription endpoint hashes, correlation IDs, and delivery status codes. Never store raw notification payloads or user-identifiable content in delivery logs to maintain GDPR/CCPA data minimization standards.

When receipts go missing — a dispatch logs sent but the ledger never advances, or a correlation ID has no matching acknowledgment row — the failure is almost always a buffering or idempotency-guard bug rather than a push service problem. The systematic triage for that class of issue is covered in Debugging missing push delivery receipts.

The same ingestion-versus-rendering boundary produces a reporting anomaly that reliably confuses stakeholders: the ledger’s delivered total runs ahead of the number of notifications users actually saw on screen. Why push delivered counts exceed displayed counts accounts for where each unit of that difference goes.

2. Implementing Acknowledgment Handlers

Wrap push API calls in a structured response parser that extracts HTTP status codes, service headers, and correlation IDs. When integrating with high-volume dispatch systems, ensure your tracking layer decouples from Message Batching & Throughput Optimization by routing acknowledgment events through asynchronous streams rather than blocking synchronous waits.

Production-Ready Handler Implementation

/**
 * Tracks push service acknowledgments and routes events to the delivery ledger.
 * Implements idempotency guards, header extraction, and structured error handling.
 *
 * @param {Response} response - Fetch API response from push service
 * @param {string} correlationId - Unique request identifier (UUIDv4)
 * @param {Object} logger - Structured logging interface
 * @returns {Promise<{messageId: string, status: string, state: string}>}
 */
export async function trackPushAcknowledgment(response, correlationId, logger) {
  const status = response.status;
  const headers = response.headers;
  const messageId = headers.get('x-push-message-id')
    || headers.get('x-message-id')
    || correlationId;

  const eventPayload = {
    messageId,
    correlationId,
    timestamp: Date.now(),
    statusCode: status,
    retryAfter: headers.get('retry-after')
      ? parseInt(headers.get('retry-after'), 10)
      : null
  };

  try {
    if (status === 201 || status === 200) {
      await logger.info('PUSH_ACCEPTED', eventPayload);
      return { messageId, status: 'ACCEPTED', state: 'PENDING_DELIVERY' };
    }

    if (status >= 400 && status < 500) {
      await logger.warn('PUSH_CLIENT_ERROR', eventPayload);
      return { messageId, status: 'REJECTED', state: 'CLIENT_ERROR' };
    }

    if (status >= 500) {
      await logger.error('PUSH_SERVER_ERROR', eventPayload);
      return { messageId, status: 'SERVER_ERROR', state: 'RETRYABLE' };
    }

    await logger.warn('PUSH_UNKNOWN_STATUS', eventPayload);
    return { messageId, status: 'UNKNOWN', state: 'UNRESOLVED' };
  } catch (err) {
    await logger.fatal('LEDGER_WRITE_FAILURE', { correlationId, error: err.message });
    throw new Error(`Acknowledgment tracking failed: ${err.message}`);
  }
}

Implementation Notes:

  • Extract x-push-message-id or equivalent proprietary headers for cross-system traceability.
  • Implement a circuit breaker around the ledger write operation to prevent cascading failures during push service outages.
  • Route successful acknowledgments to a message broker (e.g., Kafka, RabbitMQ) for downstream analytics and CRM sync.

Every field the handler reads comes from the response head — there is no body to parse. The annotated response below marks the four things worth extracting and the one thing whose absence is structural rather than accidental.

Anatomy of a push acknowledgment response A rendered HTTP 201 response from a push service showing the status line, location, x-push-message-id, ttl and retry-after headers, plus an empty body. Four callouts explain how each element maps into the delivery ledger and why the empty body means display can never be confirmed server-side. What the handler can actually read from an acknowledgment HTTP/2 201 Created location: /message/0:1699…a41 x-push-message-id: 0:1699…a41 ttl: 3600 retry-after: (absent) content-length: 0 — no response body — Status line → ledger state 201 means PENDING_DELIVERY, never DELIVERED Provider message id prefer it over the local UUID when raising vendor tickets Retry-After present only on 429 and 503 — clamp it before obeying it Empty body, by design RFC 8030 defines no receipt — display is client-side only Anything the handler needs beyond these headers has to come from the client, not the push service.
Four extractable signals and one structural absence: the acknowledgment carries no body, so it can never confirm display.

3. Mapping Status Codes to Delivery States

Push services return standardized HTTP codes that dictate downstream routing and retry logic. Time-sensitive routing must align with TTL & Expiration Handling to prevent stale retries. For permanent subscription invalidations such as 410 Gone, implement automated cleanup workflows detailed in Handling 410 Gone responses at scale to maintain list hygiene.

HTTP Status Delivery State Routing Action Retry Policy
200 / 201 PENDING_DELIVERY Log & advance state machine None (await client event)
400 BAD_REQUEST Drop payload, alert engineering None
401 / 403 AUTH_FAILURE Rotate VAPID keys, quarantine endpoint None
404 NOT_FOUND Flag for pruning, log for audit None
410 GONE Trigger immediate subscription deletion None
413 PAYLOAD_TOO_LARGE Reject, enforce size limits (< 4 KB plaintext) None
429 RATE_LIMITED Defer to the retry/backoff layer Exponential backoff, respect Retry-After
500 / 503 SERVICE_DEGRADED Log for infra alerting, queue for retry Linear backoff, max 3 attempts

Compliance Directive: Maintain immutable audit logs for 30–90 days per regional data retention laws. Implement automated purging jobs post-retention window to prevent compliance drift.

4. Secure Logging & Analytics Integration

Delivery tracking data must be sanitized before ingestion into analytics, BI, or CRM platforms. Raw subscription endpoints constitute persistent identifiers and must be hashed using HMAC-SHA256 with a rotating key to preserve user anonymity while enabling cohort analysis.

Security Architecture Checklist:

  1. Endpoint Hashing: hash = HMAC-SHA256(endpoint, rotating_secret_key). Store only the hash for deduplication.
  2. Structured Logging: Emit JSON-formatted logs with standardized severity levels (INFO, WARN, ERROR, FATAL). Include trace_id, span_id, and correlation_id for distributed tracing.
  3. Data Pipeline Decoupling: Use a write-ahead log (WAL) or message queue to buffer acknowledgment events, preventing dispatch latency spikes from blocking tracking writes.
  4. Access Controls: Restrict dashboard access to engineering and compliance teams. Mask hashed endpoints in UI views and enforce IP-allowlisted API access for raw log exports.

Read the pipeline below as a one-way valve: each stage removes something, and nothing downstream of the buffer can reconstruct what an earlier stage stripped. That irreversibility is the property auditors care about.

Acknowledgment sanitization pipeline A raw acknowledgment event is hashed with a rotating HMAC key, stripped of payload and identifiers, serialised as a structured JSON log, buffered through a write-ahead log, and only then fanned out to the delivery ledger and to analytics. A footer lists the fields that never cross the dispatch boundary. What leaves the dispatch boundary — and what never does Raw ack event status, endpoint, payload, user id HMAC the endpoint SHA-256, rotating key 30-day cadence Strip and redact payload and PII dropped here Structured JSON trace_id, span_id, correlation_id WAL / broker buffer decouples dispatch latency Delivery ledger idempotent by correlation_id Analytics and CRM joins on the hash only Never crosses the boundary: raw endpoints, notification payloads, user identifiers, plaintext p256dh and auth keys.
Each stage removes information irreversibly, so a downstream breach or a careless dashboard query cannot recover a raw endpoint.

5. Validation & Testing Strategy

Validate acknowledgment handlers using mock push service endpoints that simulate 200, 404, 410, and 429 responses. Implement chaos testing to verify system resilience under partial push service outages, network partitions, and malformed header responses. Monitor acknowledgment latency percentiles (p50, p95, p99) to detect upstream degradation before it impacts user engagement metrics.

Percentiles matter more than the mean here, because the failure mode is a stretched upper tail rather than a shifted average. A healthy handler looks like the profile below: p50 and p95 comfortably inside the SLO, with p99 the only series that touches it.

Acknowledgment processing latency percentiles Three bars show acknowledgment handler latency: p50 at 12 milliseconds, p95 at 41 milliseconds, and p99 at 96 milliseconds. A dashed line marks the 50 millisecond p95 service level objective, which only the p99 bar exceeds. Acknowledgment processing latency by percentile 120 ms 80 ms 40 ms 0 ms 12 ms p50 41 ms p95 96 ms p99 p95 SLO — 50 ms A p99 above the line is almost always the ledger write, not the push service — check the circuit-breaker timeout first.
Only the tail crosses the SLO line. Alert on sustained `p95` movement, not on the mean, which hides this shape entirely.

Testing & Debugging Protocol:

  • Contract Testing: Validate handler behavior against Web Push Protocol RFC specifications. Ensure all required headers are parsed and unexpected fields are safely ignored.
  • Concurrency Stress Testing: Simulate high-volume acknowledgment floods (10k+ RPS) to test queue backpressure, connection pooling limits, and circuit breaker thresholds.
  • Latency SLOs: Establish strict SLOs for acknowledgment processing (<50ms p95). Alert on sustained degradation to prevent ledger write bottlenecks.
  • Debugging Checklist:
    • Verify VAPID key rotation isn’t causing 401/403 spikes.
    • Check Retry-After header parsing logic for 429 responses.
    • Audit HMAC key rotation schedules to prevent hash mismatches during cohort joins.
    • Confirm ledger idempotency guards prevent duplicate state transitions on network retries.

If the symptom is that receipts simply never arrive — rather than arriving wrong — follow the dedicated playbook in Debugging missing push delivery receipts.

6. Acknowledgment Handler Configuration Reference

Tune these parameters per provider and workload. Defaults assume a Redis-buffered ledger writing asynchronously to Postgres.

Parameter Type Default Notes
ackBufferSize integer 10000 Max in-memory acknowledgment events before back-pressuring dispatch
ledgerWriteTimeoutMs integer 500 Circuit-breaker threshold for a single ledger write
correlationIdHeader string x-push-message-id Provider header preferred over the locally generated UUID
retryAfterCapSec integer 120 Upper clamp on a Retry-After value before treating it as a service incident
hashKeyRotationDays integer 30 HMAC key rotation cadence for endpoint hashing
auditRetentionDays integer 90 Ledger retention window before automated purge
dedupeWindowSec integer 604800 TTL on the idempotency key that blocks duplicate state transitions

Verification

After wiring the handler, confirm the full path with a mock push service and a real subscription:

# Simulate the four canonical responses against your handler endpoint
for code in 201 410 429 503; do
  curl -s -o /dev/null -w "ack(%{http_code}) -> ledger state\n" \
    -X POST "http://localhost:8080/internal/mock-ack?status=$code" \
    -H "x-correlation-id: $(uuidgen)"
done

Then inspect the ledger to confirm each code mapped to exactly one row and one state transition. A 410 must also have flipped the corresponding subscription to gone. Acknowledgment latency should sit under your p95 SLO (target <50ms).

Mocks prove the mapping; only a live endpoint proves the header extraction. Send one real request with an empty body and a zero TTL, which every push service accepts and immediately discards, and read the response head directly:

curl -i -s -X POST "$TEST_ENDPOINT" \
  -H "TTL: 0" \
  -H "Content-Length: 0" \
  -H "Authorization: vapid t=$VAPID_JWT, k=$VAPID_PUBLIC_KEY" | head -12

A healthy response looks like this — note that there is no body at all, which is exactly the point the handler has to be built around:

HTTP/2 201
location: https://fcm.googleapis.com/fcm/send/0:1699...a41
ttl: 0
content-length: 0

If location is present but your ledger row still shows the locally generated UUID as its provider id, the header lookup is failing — usually because the fetch implementation lowercases header names and the lookup does not. If the status is 401, the JWT audience or the worker clock is wrong rather than the handler; if it is 403, the key pair no longer matches the subscription.

Next confirm idempotency at the database level rather than by reading logs. Replay the same correlation ID twice and assert the ledger did not grow:

-- Expect exactly one row and exactly one terminal state per correlation_id
SELECT correlation_id, count(*) AS rows, count(DISTINCT state) AS states
FROM delivery_log
WHERE correlation_id = '3f2b9c14-7d51-4a0e-9d2c-8e6b1f0a4471'
GROUP BY 1;
            correlation_id            | rows | states
--------------------------------------+------+--------
 3f2b9c14-7d51-4a0e-9d2c-8e6b1f0a4471 |    1 |      1

Two rows means the idempotency guard is keyed on something that varies between attempts — commonly the attempt counter or a timestamp got folded into the key. Then walk the client half of the path in DevTools, because everything above this line only proves the server side:

  1. Open Application → Service Workers on the subscribed origin and confirm the worker is activated and is running, not waiting.
  2. Tick Update on reload and use the Push input on the same panel to fire a synthetic push. The worker’s push listener should run without the network being involved at all — if it does not, the failure is in the worker, not in your acknowledgment path.
  3. Switch to the Network panel, filter to your beacon endpoint, and confirm a request appears when the notification renders. That request is the only evidence of display your ledger will ever get.
  4. In Application → Storage → IndexedDB (or whatever store the worker uses), verify the worker recorded the correlation ID it received, so a missing beacon can be told apart from a missing push.
  5. Finally, kill the network in the Network panel’s throttling dropdown and fire the push again. The beacon should be buffered and replayed, not lost — if it vanishes, your display numbers will silently undercount on mobile.

A useful invariant to assert continuously in production: the count of ledger rows in PENDING_DELIVERY older than the message TTL should trend to zero. Rows that stay pending past their own expiry are the population that will never be reconciled, and their size is the honest measure of how much of your delivery data is unknowable.

Error & Edge-Case Matrix

Most acknowledgment bugs are not wrong status handling — they are timing, key rotation, and clock problems that surface as impossible-looking ledger states.

Condition Cause Fix
Ledger row stuck in PENDING_DELIVERY forever The push was accepted but the device never woke, or the beacon never fired Age pending rows out at TTL expiry into an UNCONFIRMED state; never leave them pending
Two rows for one dispatch Idempotency key includes the attempt number or a timestamp Key strictly on correlation_id; the attempt belongs in a column, not the key
Provider message id missing on some rows Header lookup is case-sensitive, or the provider omits it entirely Normalise header case and fall back to the local UUID rather than writing null
Retry-After of several hours Provider signalling a real incident, not a per-request pause Clamp with retryAfterCapSec and raise an incident instead of sleeping the worker
Acknowledgment latency p99 spikes, p50 flat Ledger write contention, not push-service slowness Check the circuit-breaker timeout and the write buffer, not the provider status page
401 burst confined to one worker That worker’s clock drifted past the JWT exp window Sync NTP; shorten JWT lifetime and refresh with a margin
403 burst across all workers The subscription was created under a superseded application server key Roll keys forward gradually and re-subscribe the affected cohort
Beacon count exceeds dispatch count Service worker replayed a buffered beacon after reconnecting Deduplicate beacons on correlation_id at the intake endpoint
Hash joins to analytics return nothing HMAC key rotated without dual-writing Store the key generation with each hash; back-fill before retiring a generation
410 volume spikes with no churn Subscription table restored from a stale backup Compare 410 rate to the 24-hour baseline before letting the pruner run

An incident worth internalising: a team once saw delivered collapse by 40% overnight with no change to dispatch volume or error rates. The cause was a service-worker deployment that changed the beacon URL; every push still arrived and rendered, but nothing reported back. The ledger was correct and the reporting was wrong, which is the opposite of the usual assumption. The lesson is that the beacon endpoint deserves the same deployment discipline as the dispatch path — version it, keep the old route alive for a release cycle, and alert on the dispatch-to-beacon ratio rather than on the beacon count alone.

At ten times volume the acknowledgment path fails before the dispatch path does, because every dispatch produces at least one ledger write and successful displays produce a second. Buffer both, batch ledger writes into multi-row inserts, and make the intake endpoint independently scalable. The one thing not to batch is the 410 handling: a subscription flip should stay synchronous, because the cost of one extra send to a dead endpoint is small but the cost of a backlog of un-pruned endpoints is a rising error rate that will eventually trip your breakers.

Cross-Browser & Cross-Provider Notes

The acknowledgment contract is the same everywhere — RFC 8030 gives you a status line and headers, never a body — but what appears in those headers varies enough to break a handler that was only ever tested against one provider.

Acknowledgment header divergence across push services A five-row matrix comparing FCM, Mozilla Autopush and Apple web push on the message id header they return, whether Retry-After accompanies a 429, which status signals a dead endpoint, whether an error body is present, and what the acknowledgment actually proves. The final row is identical across all three: queued, not displayed. Same protocol, different acknowledgment surface FCM Autopush Apple web push message id header location + message id location only apns-id Retry-After on 429 reliably present sometimes absent on 429 and 503 dead endpoint code 404 or 410 410 only 410 only body on error usually empty JSON with errno JSON reason string what a 201 proves queued for delivery — identical everywhere, and never proof of display Treat the top four rows as adapter concerns and the bottom row as an architectural fact you cannot engineer around. A handler that only ever saw FCM in testing will write null provider ids the first time it meets Autopush. Quotas and error vocabularies shift without notice — assert on shape, not on exact strings.
Only the last row is guaranteed by the specification. Everything above it is provider behaviour that an adapter has to normalise.

Two consequences follow for the handler. First, never treat a missing provider message id as an error; Autopush returns only a location URL, so extract the trailing path segment as the id and fall back to your own UUID rather than writing null into the ledger. Second, do not build the retry decision on the presence of Retry-After, because Autopush frequently omits it on 429; the backoff curve has to work without it, which is why the Retry Logic & Backoff Strategies guide treats the header as an upper bound rather than an instruction. Where you need the byte-level detail of how two authorities diverge, that comparison lives in FCM vs Mozilla Autopush endpoint differences.

The client side diverges too, and it changes what your ledger can observe. Chrome and Edge wake the service worker reliably and fire your beacon within a second on desktop. Firefox behaves similarly but is stricter about how long the push handler may run before the browser considers the event abandoned. Safari on iOS only has a subscription at all once the site has been installed to the Home Screen, and it batches wake-ups aggressively when the device is in a low-power state, so beacons can arrive minutes after the acknowledgment — which shows up in the ledger as an unusually wide dispatch-to-beacon distribution rather than as an error. Set the window in which you expect a beacon per platform, not globally, or iOS traffic will look permanently broken.

Back to Backend Delivery Architecture & Queue Management

FAQ

Does a push acknowledgment confirm the user saw the notification?

No. The HTTP acknowledgment (201/200) only confirms the push service queued the encrypted payload. Display is observable solely through client-side events forwarded from the service worker. Treat the acknowledgment as proof of ingestion, not receipt.

How do I make ledger writes idempotent across retries?

Use the correlation ID as the primary key and guard every state transition with a short-lived idempotency key (a Redis SET NX EX). A retried dispatch reuses the same correlation ID, so the second write is a no-op instead of a duplicate row.

What should I log without violating GDPR?

Log the endpoint hash (HMAC-SHA256 with a rotating key), the correlation ID, the HTTP status, and timestamps. Never log raw payloads, PII, or full subscription endpoints. Mask hashed endpoints in UI views and IP-allowlist raw exports.

Why is my ledger missing acknowledgment rows?

The most common causes are a dropped event in the buffering layer, an idempotency guard rejecting a legitimate first write because of a reused key, or a ledger write timing out under the circuit breaker. Walk the missing receipts playbook to isolate which stage drops the event.

How long should I retain delivery logs?

Typically 30–90 days, governed by regional data-retention law and your audit needs. Run an automated purge job at the boundary to prevent compliance drift, retaining only anonymized aggregate metrics beyond the window.