Tailing Log Files Asynchronously¶
Following a log file — like tail -F — from an asyncio service sounds like a loop around read(), and the loop is easy. What makes it correct is handling rotation, partial lines and the polling interval. Measured on Python 3.14 on Linux, with a writer appending a line every 5 ms: a polling tail found new lines with a median latency of 5.2 ms when it polled every 10 ms, 50.8 ms at 100 ms and 508.6 ms at 1 second, using 0.01–0.02 s of CPU over each run of about 2.5 s. When the log was rotated by renaming it and creating a new file halfway through 400 lines, a tail that kept reading its open file received 200 and missed the other 200; checking the path's inode on each idle poll and reopening received all 400. When the log was truncated in place instead, the naive tail missed 188 lines and produced a malformed line; checking the size against the read position received 388, missing only the 12 written between its last read and the truncation. This guide builds the rotation-aware version.
Prerequisites¶
- Python 3.11+; Linux or macOS for inode semantics.
- File I/O from async code, from async file I/O with aiofiles vs asyncio.to_thread.
- The topic overview, Subprocesses & File I/O.
1. Read new data and keep partial lines¶
Open the file, seek to its end, and repeatedly read whatever has been appended. A write can end mid-line, so keep the unfinished tail of each read in a buffer:
async def follow(path: str, interval: float = 0.1):
f = open(path, "rb")
f.seek(0, os.SEEK_END)
buf = b""
while True:
chunk = f.read()
if chunk:
buf += chunk
*lines, buf = buf.split(b"\n") # last element is an incomplete line
for line in lines:
yield line
continue
await asyncio.sleep(interval)
f.read() on a regular local file returns quickly — it reads what is in the page cache — so calling it on the event loop is acceptable for local logs. On network filesystems a read can block for much longer; there, run the read in asyncio.to_thread. Open in binary mode and decode complete lines, so a multi-byte character split across two writes is never decoded in halves.
Verify: a writer that emits half a line, pauses, and writes the rest produces exactly one complete line in the tail.
2. Choose the polling interval¶
The poll interval bounds the delay between a line being written and the tail seeing it. Measured with a line every 5 ms for 400 lines: polling every 10 ms gave a median latency of 5.2 ms and a p99 of 10.3 ms; every 100 ms, 50.8 ms and 100.5 ms; every second, 508.6 ms and 991.1 ms. The median is about half the interval, as expected. The CPU cost was small at all three — 0.01 to 0.02 s over runs of about 2.5 s — because an idle poll is one read() returning nothing and one stat().
INTERVAL = 0.1 # 100 ms: ~50 ms median delay, negligible CPU
For alerting or dashboards, 100 ms to 1 s is usually plenty. When lower latency matters, use filesystem notifications instead of polling, as in watching files for changes in asyncio, and keep a slower poll as a fallback, since notifications are not delivered on every filesystem.
Verify: the measured delay between write and tail fits the consumer's needs at the chosen interval.
3. Follow renames by watching the inode¶
Most log rotation renames the current file — app.log becomes app.log.1 — and the application starts writing a new app.log. A tail holding the old file descriptor keeps reading the renamed file and never sees the new one. Measured with rotation after 200 of 400 lines: the naive tail received 200 lines and missed all 200 written after rotation. On each idle poll, compare the inode at the path with the one you have open:
def rotated(f, path) -> bool:
try:
return os.stat(path).st_ino != os.fstat(f.fileno()).st_ino
except FileNotFoundError:
return False # between rename and re-create: try again later
When it differs, read whatever remains in the old file — lines written just before the rename — then open the new one from its beginning. With this check, the tail received all 400 lines with no change in latency. A FileNotFoundError in the short gap before the new file exists is normal and should just wait for the next poll.
Verify: a test that renames the log mid-stream and starts a new file delivers every line exactly once.
4. Handle truncation, and prefer rename¶
Some setups rotate by copying the file and truncating it in place — logrotate's copytruncate. The inode does not change, but the file becomes shorter than the tail's read position. Measured: the naive tail missed 188 of 400 lines and produced one malformed line, because it resumed reading at its old offset in the middle of new data. Detect the shrink and start again from the beginning:
def truncated(f, path) -> bool:
try:
return os.stat(path).st_size < f.tell()
except FileNotFoundError:
return False
# in the idle branch of the loop:
if rotated(f, path):
buf += f.read() # drain the old file
f.close()
f = open(path, "rb")
elif truncated(f, path):
f.seek(0)
With the size check, the tail received 388 of 400 lines. The 12 missing lines had been written after its last read and before the truncation; the truncation destroyed them in the live file, and only the rotated copy has them. That window is inherent to copy-and-truncate, so prefer rename-based rotation where the application can reopen its log, and use copy-and-truncate only when it cannot.
Verify: a truncation test loses at most the lines written within one poll interval before the truncation, and produces no malformed lines.
5. Run the tail as a managed task¶
Wrap the generator in a task that the service starts and stops with everything else, and hand lines to consumers through a bounded queue so a slow consumer cannot make the tail buffer unbounded data:
async def tail_into(path: str, queue: asyncio.Queue, interval: float = 0.1):
async for line in follow_rotating(path, interval): # steps 1-4 combined
await queue.put(line.decode("utf-8", errors="replace"))
async def main():
lines: asyncio.Queue[str] = asyncio.Queue(maxsize=10_000)
async with asyncio.TaskGroup() as tg:
tg.create_task(tail_into("/var/log/app/app.log", lines))
tg.create_task(ship(lines)) # e.g. batch and send
Persist the read offset and inode periodically if the tail must resume after a restart without re-sending or skipping lines. Close the file in a finally block, so cancellation of the task releases the descriptor; a held descriptor on a rotated-away file keeps its disk space allocated until it is closed.
Verify: stopping the service closes the tailed file — no (deleted) log files remain in /proc/<pid>/fd — and a slow consumer causes back-pressure rather than memory growth.
Verification¶
The tail is correct when:
- Partial lines are buffered until their newline arrives.
- Rename rotation is followed by comparing inodes, with the old file drained first.
- Truncation is detected by comparing size with position.
- The task closes its file on cancellation and feeds a bounded queue.
Diagnostic Hook: when a log shipper stops forwarding lines after midnight while the application keeps logging, compare the inode the shipper has open, from /proc/<pid>/fd, with stat on the log path. A tail that never reopened received only the 200 lines written before rotation in this test.
Pitfalls & edge cases¶
- Holding the old descriptor after rotation. Measured: 200 of 400 lines missed.
- Ignoring truncation. Measured: 188 lines missed and a malformed line.
- Copy-and-truncate rotation. Measured: 12 lines lost even when handled.
- Decoding partial reads. Split multi-byte characters become replacement characters.
Frequently Asked Questions¶
How do I tail a file in asyncio?
Seek to the end, read appended bytes in a loop, emit complete lines and keep the partial one, and sleep between empty reads. Polling every 100 ms gave a 51 ms median delay.
Why does my Python tail stop after log rotation?
It still reads the renamed file. Compare os.stat(path).st_ino with the open file's inode and reopen when they differ; that recovered all 200 lines the naive tail missed.
How do I handle copytruncate rotation when tailing?
Seek to the start when the file's size drops below your read position. Lines written between your last read and the truncation are lost; 12 were in this test.
Is it safe to read files on the asyncio event loop?
For local log files the read returns from the page cache quickly. On network filesystems it can block, so use asyncio.to_thread there.
Related¶
- Subprocesses & File I/O — up to the topic overview.
- Sending signals to subprocesses — reopening logs on SIGHUP is the other half of rotation.
- Network I/O & Protocol Handling — the section overview.