Skip to content

Preventing Overlapping Runs of Async Cron Scripts

A cron job scheduled every five minutes that sometimes takes seven will eventually run twice at once — two copies of the same sync job writing the same rows, sending the same emails, or fighting over the same API quota. The fix is a lock that the second copy checks on startup, and the choice of lock decides what happens after a crash. Measured on Linux with Python 3.14, using an asyncio job and a second copy started 0.3 s into the first run: both a pidfile and an fcntl.flock lock made the second copy exit immediately with status 75 (EX_TEMPFAIL) in about 40 ms. The difference showed after the first copy was killed with SIGKILL: the pidfile stayed behind and every later run refused to start, while the kernel released the flock and the next run completed normally. Waiting for the lock with await asyncio.to_thread(fcntl.flock, ...) under a timeout had its own trap: the timeout fired after 0.30 s, but the abandoned thread kept waiting and acquired the lock anyway once the holder exited, leaving the process holding a lock it believed it had given up. This guide locks async jobs so they neither overlap nor wedge.

Prerequisites

1. Take an flock before starting the loop

flock locks belong to an open file description, and the kernel releases them when the last descriptor closes — including when the process dies for any reason. Take the lock non-blockingly and exit if someone else holds it:

import asyncio
import fcntl
import os
from contextlib import contextmanager

EX_TEMPFAIL = 75


@contextmanager
def single_instance(path: str):
    fd = os.open(path, os.O_RDWR | os.O_CREAT, 0o644)
    try:
        fcntl.flock(fd, fcntl.LOCK_EX | fcntl.LOCK_NB)
    except BlockingIOError:
        os.close(fd)
        raise SystemExit(EX_TEMPFAIL)          # another run is in progress
    try:
        os.ftruncate(fd, 0)
        os.write(fd, str(os.getpid()).encode())  # informational only
        yield
    finally:
        os.close(fd)                           # releases the lock


if __name__ == "__main__":
    with single_instance("/run/lock/nightly-sync.lock"):
        asyncio.run(main())

Measured: the overlapping copy exited with status 75 about 38 ms after it started, while the first copy finished normally with status 0. Taking the lock outside asyncio.run keeps the check synchronous and cheap — no loop is created for a run that will not happen — and raising SystemExit there is safe because no tasks exist yet. The PID written into the file is for humans; the lock is the flock, not the file's contents or existence.

Verify: start the job twice in quick succession; the second exits 75 within milliseconds and the first is unaffected.

Pidfile versus flock, measured A grid of 2 rows by 4 columns. Pidfile versus flock, measured lock overlapping run after holder SIGKILLed needs cleanup code pidfile (exists = locked) exit 75 in 40 ms next run refused: stale file yes, and it cannot run on SIGKILL fcntl.flock, LOCK_NB exit 75 in 38 ms next run completed no, the kernel releases it Only the lock that the kernel releases survives the crash that makes locking necessary.

2. Do not use pidfile existence as the lock

The pidfile approach — create a file, delete it on exit — is common because it is easy to read:

@contextmanager
def pidfile(path: str):
    if os.path.exists(path):
        raise SystemExit(EX_TEMPFAIL)
    with open(path, "w") as f:
        f.write(str(os.getpid()))
    try:
        yield
    finally:
        os.unlink(path)            # never runs after SIGKILL or a power cut

Measured: after the running copy was killed with SIGKILL, the file remained and the next scheduled run exited 75 — and so would every run after it, until someone deleted the file by hand. SIGKILL is not exotic: it is what the out-of-memory killer sends, and what Kubernetes and systemd send when a stop exceeds its grace period. The check-then-create sequence is also racy: two copies started at the same instant can both see "no file" before either creates it. Checking whether the recorded PID is still alive narrows the stale-file problem but not the race, and PIDs are reused.

Verify: kill -9 a running job, then start the next one; it runs.

3. Wait for the lock by polling, not in a thread

Some jobs should wait for the previous run instead of skipping — a short backlog-drain job, say. Blocking flock in a worker thread looks like the async way to wait, but it cannot be cancelled:

# Wrong: the thread keeps waiting after the timeout, and may acquire the lock later
async with asyncio.timeout(30):
    await asyncio.to_thread(fcntl.flock, fd, fcntl.LOCK_EX)


# Right: non-blocking attempts with an async sleep between them
async def acquire(fd: int, timeout: float, interval: float = 0.05) -> None:
    async with asyncio.timeout(timeout):
        while True:
            try:
                fcntl.flock(fd, fcntl.LOCK_EX | fcntl.LOCK_NB)
                return
            except BlockingIOError:
                await asyncio.sleep(interval)

Measured with a holder that kept the lock for one second and a 0.3 s timeout: the thread version raised TimeoutError at 0.30 s as intended — and 0.3 s after the holder exited, a check showed the lock held by this process, because the abandoned thread had finally completed its flock call. Cancelling to_thread cancels the await, not the thread, as explained in cancelling asyncio.to_thread calls. The polling version timed out cleanly and left the lock free. Its cost is latency of up to one interval: it acquired a lock released after 0.5 s at 0.552 s, against 0.502 s for the blocking call.

Verify: after a timed-out wait, a second process can take the lock immediately.

What the abandoned thread did after the timeout 3 lanes over time. What the abandoned thread did after the timeout holder process holds the lock (1.0 s) waiting coroutine awaits to_thread(flock) TimeoutError: moves on worker thread still blocked in flock() acquires lock time, holder releases at 1.0 s → Measured: the lock was held by the timed-out process 0.3 s after the holder exited.

4. Bound the job's run time below the schedule interval

A lock prevents overlap but turns a hung job into a permanently skipped one: the lock is held, and every later run exits 75. Put an overall deadline on the async work, shorter than the schedule interval:

async def main() -> int:
    try:
        async with asyncio.timeout(240):          # job runs every 300 s
            await sync_everything()
    except TimeoutError:
        log.error("run exceeded 240 s; aborting so the next run can start")
        return 1
    return 0


if __name__ == "__main__":
    with single_instance(LOCK_PATH):
        sys.exit(asyncio.run(main()))

The deadline cancels the job's tasks and lets their cleanup run, then the process exits and the flock is released by the descriptor closing. The next run starts fresh. Choose the deadline from the job's normal duration distribution rather than guessing, the way deriving timeouts from latency percentiles does for requests. Monitoring the count of 75 exits catches the other failure: runs being skipped because the previous one is legitimately slow.

Verify: a job that hangs exits at its deadline with a non-zero status, and the following run starts.

5. Prefer the scheduler's own guarantee where it exists

Several schedulers already refuse to start a run while the previous one is active, which makes the lock a second line of defence rather than the only one:

# systemd: a .timer never starts its .service while that service is still active
[Timer]
OnCalendar=*:0/5
Persistent=true

[Service]
Type=oneshot
ExecStart=/usr/bin/python3 /opt/jobs/nightly_sync.py
TimeoutStartSec=240
# Kubernetes CronJob: skip a run while the previous Job is still running
spec:
  schedule: "*/5 * * * *"
  concurrencyPolicy: Forbid

Plain cron has no such guarantee, which is why the flock matters most there; flock(1) from util-linux gives a shell-level equivalent (flock -n /run/lock/job.lock python job.py). Kubernetes runs each Job in a fresh pod, so a file lock inside the container protects nothing across runs — concurrencyPolicy: Forbid, or a distributed lock as in Distributed Locks & Coordination, is the tool there.

Verify: the scheduler's configuration states its overlap policy explicitly, and file locks are used only where runs share a filesystem.

How should overlapping runs be prevented? A decision on Where does the job run with 4 outcomes. How should overlapping runs be prevented? Where does the job run? cron on one host flock LOCK_NB, exit 75 plus a run deadline systemd timer timer never overlaps its service TimeoutStartSec Kubernetes CronJob concurrencyPolicy: Forbid file locks do not span pods several hosts distributed lock with a lease Redis or Postgres The kernel-released flock is the right default whenever runs share a host.

Verification

Async cron jobs never overlap or wedge when:

  • An flock with LOCK_NB is taken before asyncio.run, and an overlapping run exits 75.
  • No pidfile's existence is treated as a lock.
  • Waiting for a lock polls with LOCK_NB, never a cancelled thread blocked in flock.
  • Every run has a deadline shorter than its interval, so a hang cannot hold the lock forever.

Diagnostic Hook: count exits with status 75 per job per day. A steady non-zero count means runs regularly outlast their interval; a run of consecutive 75s with no successful run between them means a holder is stuck — or, with pidfiles, that a stale file is blocking everything.

Pitfalls & edge cases

  • Stale pidfiles. Measured: after SIGKILL, every later run refused to start.
  • to_thread(flock) with a timeout. Measured: the thread acquired the lock after the timeout fired.
  • No run deadline. A hung job holds the lock and every later run is skipped.
  • File locks across pods or hosts. They only coordinate processes that share the file.

Frequently Asked Questions

How do I stop a Python cron job from running twice at the same time?

Open a lock file and take fcntl.flock(fd, LOCK_EX | LOCK_NB) before starting the event loop; if it raises BlockingIOError, exit with status 75. The overlapping run exited in about 40 ms in testing, and the kernel released the lock when the holder was killed.

Why is a pidfile not enough to prevent overlapping runs?

It is not removed when the process is killed with SIGKILL. In testing the stale file blocked every later run, and the check-then-create sequence is racy when two runs start together.

How do I wait for a file lock in asyncio?

Poll with non-blocking flock and await asyncio.sleep between attempts, inside asyncio.timeout. Blocking flock in asyncio.to_thread cannot be cancelled: in testing the abandoned thread acquired the lock after the timeout had fired.

Do file locks work for Kubernetes CronJobs?

No — each run is a new pod with its own filesystem. Use concurrencyPolicy: Forbid, or a distributed lock if several workloads must coordinate.