Webhooks are the connective tissue for modern SaaS: they trigger automations, sync data with customers’ systems, and drive event-driven integrations. But unreliable webhook delivery is a top cause of support tickets, lost revenue, and churn. This guide walks engineering and product teams through a concrete, production-ready approach to designing a scalable, secure, and observable webhook delivery system in 2026 — with implementation patterns you can apply today.
Why build a dedicated webhook delivery system?
Many teams start by firing HTTP requests directly from application code. That works until a spike of events, a slow downstream, or a misbehaving integration creates backpressure. A dedicated delivery system gives you:
- Durability: persistent queues and retries until success or a dead-letter policy
- Visibility: delivery metrics, logs, and dashboards for SLAs
- Security: signing, secrets, and key rotation separate from app code
- Scale and isolation: rate limiting and per-customer controls to protect downstream systems
- Operational controls: replay, pause, and per-subscription configuration
Core principles
- At-least-once delivery by default, with clear deduplication/ idempotency to prevent duplicate processing.
- Durable persistence of events until acknowledged or dead-lettered; don’t rely on in-memory buffers.
- Observable behavior: track status transitions, latency, and error categories.
- Configurable retry policy per-customer/integration to handle slow or rate-limited endpoints.
- Security by design: signature verification, secret rotation, optional TLS pinning for high-risk environments.
High-level architecture
Use a modular pipeline that decouples event production from delivery:
- Event producer (your app) writes webhook events to an event store or message broker (Kafka, Kinesis, or durable DB).
- Delivery service reads events, enqueues per-subscription work items, and orchestrates retries.
- Worker fleet performs HTTP deliveries, handles responses, computes idempotency keys, and writes delivery outcomes.
- Control plane provides per-subscription config, secret management, pause/replay, and dashboards.
Keep the control plane (configs and secrets) separated from the high-throughput delivery path to reduce blast radius during schema changes.
Typical data model
Below is an example SQL-ish schema for a durable delivery table. You can adapt to a broker+DB or a fully event-sourced model.
webhook_events - id (UUID PK) - subscription_id (FK to customer subscription) - event_type - payload (JSONB) - created_at - delivered_at (nullable) deliveries - id (UUID PK) - webhook_event_id (FK) - attempt (integer) - status (queued, in_flight, success, failed, dlq) - next_attempt_at (timestamp) - response_code (int nullable) - response_body (text nullable) - error_type (timeout, network, 4xx, 5xx) - idempotency_key (varchar) - created_at - updated_at
Idempotency and deduplication
Design for at-least-once delivery but enable safe deduplication at the receiver and in your own system:
- Create a stable idempotency key for each logical event: e.g., sha256(subscription_id + event_type + event_id). Store it with deliveries.
- Include that idempotency key in a request header (e.g., X-Webhook-Idempotency-Key) and in the payload for receivers.
- Document your deduplication semantics to customers: how long keys are honored (e.g., 7 days) and whether keys are globally unique per subscription.
- Use a short-lived cache (Redis) on the delivery side to avoid reprocessing near-duplicate work when retries overlap.
Retry strategy and backoff
Retries must balance promptness with not overwhelming a struggling receiver. A practical policy:
- Immediate retry for transient network errors with small exponential backoff (e.g., base 1s, factor 2, jitter 0.1–0.5).
- Respect 429/503 with exponential backoff and consider honoring Retry-After header if present.
- Stop retrying after a maximum number of attempts (e.g., 10) or a time window (e.g., 48 hours), then mark as dead-letter.
- For 4xx permanent errors (400/401/403/404), treat as customer configuration error and move to dead-letter immediately after a couple attempts.
Sample schedule (example): attempt intervals — 0s, 2s, 8s, 32s, 2m, 8m, 30m, 2h, 8h, 24h (10 attempts).
Handling HTTP semantics
- Consider any 2xx response as success. Log and surface non-2xx codes with error categories.
- Treat 3xx redirects cautiously: follow up to a safe limit or mark as failed (configurable per-customer).
- 400–499 are likely client errors; after 1–2 attempts, surface to the customer as configuration issue.
- For 429/503, back off and retry; use Retry-After if provided.
Security: signing, secrets, rotation
Security is non-negotiable for webhooks used in financial, HR, or healthcare apps.
- Sign the payload with HMAC SHA-256 using a per-subscription secret; include signature header (e.g., X-SRHub-Signature).
- Store secrets in a secrets manager (AWS Secrets Manager, HashiCorp Vault) and avoid plaintext in DB.
- Support secret rotation: issue a new secret, accept old secrets for a configurable grace period, and allow manual rotation by customers.
- Use TLS for all deliveries and optionally support mutual TLS for high-security integrations.
- Log only metadata in plaintext; strip or encrypt sensitive payload fields from logs and dashboards to satisfy compliance (e.g., GDPR).
Observability and SLOs
Define measurable SLOs and instrument the delivery pipeline:
- Key metrics: delivery success rate (1h/24h/30d), p99 latency from event creation to delivery success, number of DLQs, retry counts per subscription.
- Traces: instrument path through producer → delivery queue → worker → HTTP call to get latency breakdowns.
- Dashboards and per-customer views: show top failing endpoints, recent errors, and replays available.
- Alerting: fail-safe alerts for system-wide degradations and per-customer spikes (e.g., >1% of deliveries to a tenant are in DLQ in 30m).
Example SLO: 99.9% of webhooks delivered successfully (2xx) within 6 hours of event creation, excluding customer-configured paused subscriptions.
Operational playbook: incidents and support
Prepare clear runbooks for common scenarios so support and engineering can act quickly:
- Delivery spike: throttle worker concurrency per subscription, apply queue backpressure, notify customers with rate-limited emails.
- High DLQ rate: surface top failing subscriptions, mark as "requires customer action" where error is 4xx, and offer replay once fixed.
- Delayed deliveries: check broker lag, worker errors, and DB write performance; scale workers or temporarily offload to a cheap burst queue.
- Security compromise: rotate secrets immediately, notify affected customers, and provide signed audit logs of actions.
Provide a self-serve control panel for customers to pause deliveries, view failure reasons, rotate secrets, and manually replay events with canned responses to educate them about common misconfigurations.
Scaling and cost controls
Design to scale horizontally. Practical notes:
- Partition work by subscription_id to maintain per-customer ordering when needed, or use per-subscription queues for strict ordering guarantees.
- Use a combination of broker (Kafka, Kinesis) for high-throughput durable ingestion and a delivery DB for per-attempt bookkeeping.
- Implement per-subscription concurrency limits and rate limits to protect both your workers and customers’ endpoints.
- Monitor egress costs: HTTP traffic can be expensive at scale—compress payloads, support batched webhooks, and provide optional webhooks via message queue (SQS, Pub/Sub) for heavy customers.
Managed vs. self-hosted vs. third-party broker
Options to accelerate development:
- Build in-house — maximum control; best if webhooks are core to your product or require custom policies.
- Use a managed integration broker (Hookdeck, Pipedream-style services) — faster time to market, built-in replay/observability, but vendor lock-in and recurring costs.
- Hybrid — use a managed ingestion/queue and implement delivery layer yourself for control over security, idempotency, and SLAs.
Choose based on required SLAs, regulatory constraints, and engineering bandwidth. For high-risk verticals (finance, healthcare), lean toward owning the delivery path and secrets.
Developer ergonomics and customer-facing docs
Good developer experience reduces support load:
- Provide clear webhook contract docs, sample SDKs, and language-specific verification helpers for your signature scheme.
- Offer a webhook tester UI that sends sample events and shows signature headers and delivery logs.
- Document retry semantics, idempotency expectations, and error categories so integrators can handle duplicates and failures properly.
Checklist: launch-ready webhook delivery
- Persistent event storage and delivery table with DLQ
- Idempotency keys surfaced and documented
- Default retry/backoff policy and per-subscription override
- Signing and secrets stored in a secrets manager with rotation support
- Per-subscription rate limiting and concurrency controls
- Dashboards for success rate, p99 latency, DLQ counts, and replay controls
- Runbooks for common incidents and customer-facing self-serve tools
Example: quick implementation notes
Minimal working flow using widely available components:
- App publishes events to Kafka topic partitioned by subscription_id.
- Delivery service consumes, writes a deliveries row to Postgres with status=queued and computes idempotency_key.
- Workers poll deliveries where status=queued and next_attempt_at < now(); set status=in_flight (CAS update).
- Worker performs HTTP POST with X-SRHub-Signature and X-Idempotency-Key headers, logs response, updates delivery row to success or schedules next_attempt_at.
- Expose APIs for pause/replay/rotate-secret and admin dashboard for metrics.
Conclusion
Webhooks are deceptively complex: reliability, security, and observability require deliberate architecture and operational playbooks. By adopting durable persistence, idempotency, sensible retry semantics, and clear customer-facing tools, SaaS teams can reduce support load, increase integration success, and build trust with customers that rely on real-time automation.
Start small: implement a durable queue and a delivery worker with HMAC signing and a basic retry schedule. Iterate toward per-subscription controls, rich observability, and incident runbooks. The result: a webhook system that scales with your product and keeps customer integrations healthy in production.