Preventing Overlapping Runs of Async Cron Scripts¶
A cron job scheduled every five minutes that sometimes takes seven will eventually run twice at once — two copies of the same sync job writing the same rows, sending the same emails, or fighting over the same API quota. The fix is a lock that the second copy checks on startup, and the choice of lock decides what happens after a crash. Measured on Linux with Python 3.14, using an asyncio job and a second copy started 0.3 s into the first run: both a pidfile and an fcntl.flock lock made the second copy exit immediately with status 75 (EX_TEMPFAIL) in about 40 ms. The difference showed after the first copy was killed with SIGKILL: the pidfile stayed behind and every later run refused to start, while the kernel released the flock and the next run completed normally. Waiting for the lock with await asyncio.to_thread(fcntl.flock, ...) under a timeout had its own trap: the timeout fired after 0.30 s, but the abandoned thread kept waiting and acquired the lock anyway once the holder exited, leaving the process holding a lock it believed it had given up. This guide locks async jobs so they neither overlap nor wedge.
Prerequisites¶
- Linux or macOS (
fcntlis POSIX-only); Python 3.11+. - Exit statuses for scripts, from exiting async scripts with the right status code.
- Cross-host coordination, for jobs that run on several machines: running singleton jobs across replicas.
1. Take an flock before starting the loop¶
flock locks belong to an open file description, and the kernel releases them when the last descriptor closes — including when the process dies for any reason. Take the lock non-blockingly and exit if someone else holds it:
import asyncio
import fcntl
import os
from contextlib import contextmanager
EX_TEMPFAIL = 75
@contextmanager
def single_instance(path: str):
fd = os.open(path, os.O_RDWR | os.O_CREAT, 0o644)
try:
fcntl.flock(fd, fcntl.LOCK_EX | fcntl.LOCK_NB)
except BlockingIOError:
os.close(fd)
raise SystemExit(EX_TEMPFAIL) # another run is in progress
try:
os.ftruncate(fd, 0)
os.write(fd, str(os.getpid()).encode()) # informational only
yield
finally:
os.close(fd) # releases the lock
if __name__ == "__main__":
with single_instance("/run/lock/nightly-sync.lock"):
asyncio.run(main())
Measured: the overlapping copy exited with status 75 about 38 ms after it started, while the first copy finished normally with status 0. Taking the lock outside asyncio.run keeps the check synchronous and cheap — no loop is created for a run that will not happen — and raising SystemExit there is safe because no tasks exist yet. The PID written into the file is for humans; the lock is the flock, not the file's contents or existence.
Verify: start the job twice in quick succession; the second exits 75 within milliseconds and the first is unaffected.
2. Do not use pidfile existence as the lock¶
The pidfile approach — create a file, delete it on exit — is common because it is easy to read:
@contextmanager
def pidfile(path: str):
if os.path.exists(path):
raise SystemExit(EX_TEMPFAIL)
with open(path, "w") as f:
f.write(str(os.getpid()))
try:
yield
finally:
os.unlink(path) # never runs after SIGKILL or a power cut
Measured: after the running copy was killed with SIGKILL, the file remained and the next scheduled run exited 75 — and so would every run after it, until someone deleted the file by hand. SIGKILL is not exotic: it is what the out-of-memory killer sends, and what Kubernetes and systemd send when a stop exceeds its grace period. The check-then-create sequence is also racy: two copies started at the same instant can both see "no file" before either creates it. Checking whether the recorded PID is still alive narrows the stale-file problem but not the race, and PIDs are reused.
Verify: kill -9 a running job, then start the next one; it runs.
3. Wait for the lock by polling, not in a thread¶
Some jobs should wait for the previous run instead of skipping — a short backlog-drain job, say. Blocking flock in a worker thread looks like the async way to wait, but it cannot be cancelled:
# Wrong: the thread keeps waiting after the timeout, and may acquire the lock later
async with asyncio.timeout(30):
await asyncio.to_thread(fcntl.flock, fd, fcntl.LOCK_EX)
# Right: non-blocking attempts with an async sleep between them
async def acquire(fd: int, timeout: float, interval: float = 0.05) -> None:
async with asyncio.timeout(timeout):
while True:
try:
fcntl.flock(fd, fcntl.LOCK_EX | fcntl.LOCK_NB)
return
except BlockingIOError:
await asyncio.sleep(interval)
Measured with a holder that kept the lock for one second and a 0.3 s timeout: the thread version raised TimeoutError at 0.30 s as intended — and 0.3 s after the holder exited, a check showed the lock held by this process, because the abandoned thread had finally completed its flock call. Cancelling to_thread cancels the await, not the thread, as explained in cancelling asyncio.to_thread calls. The polling version timed out cleanly and left the lock free. Its cost is latency of up to one interval: it acquired a lock released after 0.5 s at 0.552 s, against 0.502 s for the blocking call.
Verify: after a timed-out wait, a second process can take the lock immediately.
4. Bound the job's run time below the schedule interval¶
A lock prevents overlap but turns a hung job into a permanently skipped one: the lock is held, and every later run exits 75. Put an overall deadline on the async work, shorter than the schedule interval:
async def main() -> int:
try:
async with asyncio.timeout(240): # job runs every 300 s
await sync_everything()
except TimeoutError:
log.error("run exceeded 240 s; aborting so the next run can start")
return 1
return 0
if __name__ == "__main__":
with single_instance(LOCK_PATH):
sys.exit(asyncio.run(main()))
The deadline cancels the job's tasks and lets their cleanup run, then the process exits and the flock is released by the descriptor closing. The next run starts fresh. Choose the deadline from the job's normal duration distribution rather than guessing, the way deriving timeouts from latency percentiles does for requests. Monitoring the count of 75 exits catches the other failure: runs being skipped because the previous one is legitimately slow.
Verify: a job that hangs exits at its deadline with a non-zero status, and the following run starts.
5. Prefer the scheduler's own guarantee where it exists¶
Several schedulers already refuse to start a run while the previous one is active, which makes the lock a second line of defence rather than the only one:
# systemd: a .timer never starts its .service while that service is still active
[Timer]
OnCalendar=*:0/5
Persistent=true
[Service]
Type=oneshot
ExecStart=/usr/bin/python3 /opt/jobs/nightly_sync.py
TimeoutStartSec=240
# Kubernetes CronJob: skip a run while the previous Job is still running
spec:
schedule: "*/5 * * * *"
concurrencyPolicy: Forbid
Plain cron has no such guarantee, which is why the flock matters most there; flock(1) from util-linux gives a shell-level equivalent (flock -n /run/lock/job.lock python job.py). Kubernetes runs each Job in a fresh pod, so a file lock inside the container protects nothing across runs — concurrencyPolicy: Forbid, or a distributed lock as in Distributed Locks & Coordination, is the tool there.
Verify: the scheduler's configuration states its overlap policy explicitly, and file locks are used only where runs share a filesystem.
Verification¶
Async cron jobs never overlap or wedge when:
- An
flockwithLOCK_NBis taken beforeasyncio.run, and an overlapping run exits 75. - No pidfile's existence is treated as a lock.
- Waiting for a lock polls with
LOCK_NB, never a cancelled thread blocked inflock. - Every run has a deadline shorter than its interval, so a hang cannot hold the lock forever.
Diagnostic Hook: count exits with status 75 per job per day. A steady non-zero count means runs regularly outlast their interval; a run of consecutive 75s with no successful run between them means a holder is stuck — or, with pidfiles, that a stale file is blocking everything.
Pitfalls & edge cases¶
- Stale pidfiles. Measured: after
SIGKILL, every later run refused to start. to_thread(flock)with a timeout. Measured: the thread acquired the lock after the timeout fired.- No run deadline. A hung job holds the lock and every later run is skipped.
- File locks across pods or hosts. They only coordinate processes that share the file.
Frequently Asked Questions¶
How do I stop a Python cron job from running twice at the same time?
Open a lock file and take fcntl.flock(fd, LOCK_EX | LOCK_NB) before starting the event loop; if it raises BlockingIOError, exit with status 75. The overlapping run exited in about 40 ms in testing, and the kernel released the lock when the holder was killed.
Why is a pidfile not enough to prevent overlapping runs?
It is not removed when the process is killed with SIGKILL. In testing the stale file blocked every later run, and the check-then-create sequence is racy when two runs start together.
How do I wait for a file lock in asyncio?
Poll with non-blocking flock and await asyncio.sleep between attempts, inside asyncio.timeout. Blocking flock in asyncio.to_thread cannot be cancelled: in testing the abandoned thread acquired the lock after the timeout had fired.
Do file locks work for Kubernetes CronJobs?
No — each run is a new pod with its own filesystem. Use concurrencyPolicy: Forbid, or a distributed lock if several workloads must coordinate.
Related¶
- Async Scripts & CLIs — up to the topic overview.
- Exiting async scripts with the right status code — where status 75 fits.
- Asyncio Fundamentals & Event Loop Architecture — the section overview.