I rolled out a new “order‑matching” worker last quarter, thinking the extra Redis cache would make the pipeline bullet‑proof. Six minutes later the ops dashboard lit up with 12 k duplicate matches, and my team was scrambling to delete phantom orders. The root cause? Two instances thought they owned the same job because the lock they were using silently expired while the process was still crunching data.

⚡ TL;DR — Key takeaways
  • Use a lease‑based lock (TTL + heartbeat) instead of a fire‑and‑forget mutex.
  • Pair every lock acquisition with a monotonically increasing fencing token.
  • Pick the right coordination service: Redis (AP) for speed, ZooKeeper/etcd (CP) for strong guarantees.
  • Implement exponential‑backoff retries and idempotent workflow handlers.
  • Measure latency, watch clock skew, and size your lock granularity to avoid contention.

Before you start: Go 1.24, Redis 7.x (or Redis 7.2+), ZooKeeper 3.8, etcd 3.5, PostgreSQL 16, AWS DynamoDB, Azure Blob Storage SDK v12, Apache Curator 5.5, and a basic understanding of lease‑based coordination.

How do you safeguard concurrent workflows with distributed locks?

Implement distributed locking for concurrent workflows using a dedicated coordination service (like Redis, ZooKeeper, or etcd) to manage a mutual exclusion lock. The lock must use a time‑bound lease, include heartbeat renewal, and pair with idempotent workflow logic and fencing tokens to ensure safety during process failures or network partitions.

The Problem: Why Distributed Workflows Demand Coordinated Locks

Race Conditions in Concurrent Executions

When multiple workers poll the same queue, they can fetch the identical message before any of them have a chance to lock it. The resulting race condition often manifests as duplicate database writes, double‑charged payments, or, as in my case, phantom order matches.

Idempotency Failure and Duplicate Processing Risks

Even if your downstream API is “idempotent” on paper, the surrounding state (counters, timestamps, side‑effects) can still diverge. Without a reliable lock, two workers may each think they are the first, leading to inconsistent aggregates and costly rollbacks.

**My take:** Most teams treat idempotency as a blanket safety net, but in practice you still need *exclusive* access to the critical section. A lock isn’t a silver bullet; it’s the guardrail that lets idempotency shine.

—

Fundamentals of Distributed Locking for Backend Engineers

Understanding Lease Times and Heartbeats

A lease is a lock with an explicit TTL. The holder must periodically *renew* the lease—think of it as a heartbeat that tells the coordination service, “I’m still alive.” If the heartbeat stops, the lease expires automatically, freeing the resource for others.

// go 1.24
// redis-py 5.0 equivalent in Go: go-redis v9.3
package main

import (
	"context"
	"time"

	"github.com/redis/go-redis/v9"
)

func acquireLease(ctx context.Context, rdb *redis.Client, key string, ttl time.Duration) (string, error) {
	token := fmt.Sprintf("%d-%d", time.Now().UnixNano(), rand.Int())
	ok, err := rdb.SetNX(ctx, key, token, ttl).Result()
	if err != nil {
		return "", fmt.Errorf("redis setnx failure: %w", err)
	}
	if !ok {
		return "", fmt.Errorf("lock already held")
	}
	return token, nil
}

// Heartbeat runs in a goroutine; stops on context cancel.
func startHeartbeat(ctx context.Context, rdb *redis.Client, key, token string, ttl time.Duration) {
	ticker := time.NewTicker(ttl / 2)
	defer ticker.Stop()
	for {
		select {
		case <-ctx.Done():
			return
		case <-ticker.C:
			// Use Lua script to extend only if token matches (prevent steal)
			script := redis.NewScript(`
				if redis.call("GET", KEYS[1]) == ARGV[1] then
					return redis.call("PEXPIRE", KEYS[1], ARGV[2])
				else
					return 0
				end`)
			_, err := script.Run(ctx, rdb, []string{key}, token, int64(ttl/time.Millisecond)).Result()
			if err != nil {
				// Log and continue; lease may have been lost
				log.Printf("heartbeat error: %v", err)
			}
		}
	}
}

Notice the Lua script that only refreshes the TTL if the stored token still matches ours. That’s the first line of defense against a stale lock.

The Critical Role of Fencing Tokens (and Why Most Tutes Skip It)

A fencing token is a monotonically increasing integer (or UUID) issued **by the lock store** each time a lock is successfully acquired. When you later write to a shared resource, you include the token; the downstream system rejects any write with a token lower than the current high‑water mark. This eliminates the “old client still thinks it owns the lock” problem.

Redis’ Redlock algorithm doesn’t bake fencing tokens in the spec, but you can add them yourself by storing a separate counter key:

func getFencingToken(ctx context.Context, rdb *redis.Client) (int64, error) {
	// INCR is atomic, guarantees ever‑increasing token
	val, err := rdb.Incr(ctx, "global:fencing:counter").Result()
	if err != nil {
		return 0, fmt.Errorf("failed to get fencing token: %w", err)
	}
	return val, nil
}

When you later write:

func writeIfFresh(ctx context.Context, db *sql.DB, token int64, payload string) error {
	_, err := db.ExecContext(ctx,
		`INSERT INTO jobs (id, token, payload)
		 VALUES ($1, $2, $3)
		 ON CONFLICT (id) DO UPDATE
		 SET token = EXCLUDED.token,
		     payload = EXCLUDED.payload
		 WHERE jobs.token < EXCLUDED.token`,
		jobID, token, payload)
	if err != nil {
		return fmt.Errorf("conditional write failed: %w", err)
	}
	return nil
}

If a stale worker tries to write with an old token, the `WHERE jobs.token < EXCLUDED.token` guard stops it dead in its tracks.

—

2026 Implementation Deep Dive: Code & Architectural Trade‑offs

Below is a quick‑fire comparison of five patterns you’ll encounter in the wild. The table summarises latency, consistency guarantees, and operational overhead.

PatternConsistency ModelAvg. Lock Latency (us)Ops OverheadTypical Use‑case
Redis + RedlockAP (eventual)150‑300Low (managed cluster)High‑throughput caches, short‑lived jobs
ZooKeeper/etcdCP (strong)400‑800Medium (requires quorum)Financial transactions, strong ordering
DB Advisory LocksCP (via ACID)200‑500Low (existing DB)Legacy monoliths, low volume
Cloud‑Native (DynamoDB, Azure Blob Lease)AP/CP hybrids250‑600Low‑to‑Medium (managed)Multi‑region microservices
Leader Election (Curator)CP500‑900Medium‑High (ZK ensemble)Scheduler election, global coordination

Pattern 1: Redis + Redlock (Production Ready, With Caveats)

**Why you might pick it:** You already run Redis for caching; adding a lock adds negligible latency.

**Caveats:** Redlock relies on synchronized clocks across nodes (NTP jitter > 150 ms can break safety). You must also implement fencing tokens manually, as shown earlier.

**Full implementation (Go 1.24):**

// redis_redlock.go
// Requires go-redis v9.3
package lock

import (
	"context"
	"crypto/rand"
	"encoding/hex"
	"time"

	"github.com/redis/go-redis/v9"
)

// Redlock client holds multiple Redis endpoints
type Redlock struct {
	clients []*redis.Client
	ttl     time.Duration
}

// NewRedlock builds a client from a list of addresses.
func NewRedlock(addresses []string, ttl time.Duration) *Redlock {
	clients := make([]*redis.Client, len(addresses))
	for i, addr := range addresses {
		clients[i] = redis.NewClient(&redis.Options{
			Addr:         addr,
			DialTimeout:  2 * time.Second,
			ReadTimeout:  2 * time.Second,
			WriteTimeout: 2 * time.Second,
		})
	}
	return &Redlock{clients: clients, ttl: ttl}
}

// generateToken creates a 16‑byte random string.
func generateToken() (string, error) {
	b := make([]byte, 16)
	if _, err := rand.Read(b); err != nil {
		return "", err
	}
	return hex.EncodeToString(b), nil
}

// Acquire tries to lock the key on a majority of nodes.
func (r *Redlock) Acquire(ctx context.Context, key string) (string, error) {
	token, err := generateToken()
	if err != nil {
		return "", err
	}
	success := 0
	for _, c := range r.clients {
		ok, err := c.SetNX(ctx, key, token, r.ttl).Result()
		if err != nil {
			// Log and continue; a single failure is tolerable
			continue
		}
		if ok {
			success++
		}
	}
	// Need > N/2 successful locks
	if success <= len(r.clients)/2 {
		// Cleanup partial locks
		r.Release(ctx, key, token)
		return "", fmt.Errorf("failed to acquire majority lock")
	}
	return token, nil
}

// Release removes the lock only if the token matches.
func (r *Redlock) Release(ctx context.Context, key, token string) error {
	script := redis.NewScript(`
		if redis.call("GET", KEYS[1]) == ARGV[1] then
			return redis.call("DEL", KEYS[1])
		else
			return 0
		end`)
	for _, c := range r.clients {
		_, _ = script.Run(ctx, c, []string{key}, token).Result()
	}
	return nil
}

**Operational tip:** Run the Redis nodes in separate AZs and enable *Redis‑Cluster* with at least 3 master shards. This keeps the majority requirement safe even if an AZ goes offline.

Warning: Do not rely on Redis’ `EXPIRE` alone; if the client crashes right after `SETNX` but before setting the TTL, the lock becomes permanent.

Pattern 2: ZooKeeper/etcd for CP‑System Guarantees

Both ZooKeeper 3.8 and etcd 3.5 implement the *Paxos/Raft* consensus algorithm, guaranteeing linearizable reads/writes. Their APIs expose *leases* (ZooKeeper’s `EphemeralNode`, etcd’s `LeaseGrant`).

**ZooKeeper example (Java 21 + Curator 5.5):**

// ZKLock.java
import org.apache.curator.framework.CuratorFramework;
import org.apache.curator.retry.ExponentialBackoffRetry;
import org.apache.curator.framework.recipes.locks.InterProcessMutex;

public class ZKLock {
    private final CuratorFramework client;
    private final InterProcessMutex lock;

    public ZKLock(String zkConn, String lockPath) {
        client = CuratorFrameworkFactory.newClient(
            zkConn,
            new ExponentialBackoffRetry(1000, 3));
        client.start();
        lock = new InterProcessMutex(client, lockPath);
    }

    public boolean acquire(long timeoutMs) throws Exception {
        return lock.acquire(timeoutMs, TimeUnit.MILLISECONDS);
    }

    public void release() throws Exception {
        lock.release();
    }
}

Curator handles session expiration and auto‑re‑acquisition for you. The trade‑off is higher latency (≈ 600 µs) and the need to maintain a ZK ensemble (minimum 3 nodes).

**etcd example (Go 1.24 + clientv3 v3.5):**

// etcd_lock.go
package lock

import (
	"context"
	"time"

	clientv3 "go.etcd.io/etcd/client/v3"
)

type EtcdLock struct {
	cli    *clientv3.Client
	lease  clientv3.LeaseID
	key    string
	cancel context.CancelFunc
}

// NewEtcdLock creates a lease‑based lock.
func NewEtcdLock(endpoints []string, key string, ttl int64) (*EtcdLock, error) {
	cli, err := clientv3.New(clientv3.Config{
		Endpoints:   endpoints,
		DialTimeout: 5 * time.Second,
	})
	if err != nil {
		return nil, err
	}
	lease, err := cli.Grant(context.Background(), ttl)
	if err != nil {
		return nil, err
	}
	// Acquire lock by putting key with lease
	_, err = cli.Put(context.Background(), key, "owner", clientv3.WithLease(lease.ID))
	if err != nil {
		return nil, err
	}
	// Keepalive goroutine
	ctx, cancel := context.WithCancel(context.Background())
	ch, kaerr := cli.KeepAlive(ctx, lease.ID)
	if kaerr != nil {
		cancel()
		return nil, kaerr
	}
	// Drain keepalive channel in background
	go func() {
		for range ch {
		}
	}()
	return &EtcdLock{cli: cli, lease: lease.ID, key: key, cancel: cancel}, nil
}

// Release removes the key and revokes the lease.
func (l *EtcdLock) Release() error {
	l.cancel()
	_, err := l.cli.Delete(context.Background(), l.key)
	if err != nil {
		return err
	}
	_, err = l.cli.Revoke(context.Background(), l.lease)
	return err
}

**When to choose CP:** Your domain cannot tolerate even a single duplicate (financial ledgers, inventory decrements). The added latency is a price you pay for safety.

Pattern 3: Database‑Based Advisory Locks (PostgreSQL 16, MySQL 8)

If you already run a relational DB, use its built‑in advisory lock primitives. They’re cheap, ACID‑safe, and require no extra service.

**PostgreSQL advisory lock (Go 1.24 + pgx v5):**

// pg_advisory.go
package lock

import (
	"context"
	"github.com/jackc/pgx/v5"
)

func AcquirePGAdvisory(ctx context.Context, conn *pgx.Conn, lockID int64) error {
	_, err := conn.Exec(ctx, "SELECT pg_advisory_lock
Written by

’m Nilesh, a Software Development Engineer with 2+ years of experience, specializing in Go, JavaScript, Python, Docker, Kubernetes, Git, Jenkins, microservices, and system design (LLD/HLD), backed by a strong foundation in data structures and algorithms. Alongside my engineering journey, I bring 4+ years of hands-on experience in SEO, where I’ve worked extensively on content strategy, keyword research, technical SEO, and organic growth, helping products and businesses scale efficiently by aligning solid technology with search-driven performance.