Python - Python mmap Module for Memory-Mapped File Processing

The Python mmap module provides a way to access the contents of a file through a memory-mapped region. Instead of repeatedly using normal file operations such as read() and write(), a program can map a file into memory and work with it almost like a byte sequence. This can be especially useful when working with large files because the operating system manages the mapping and loads portions of the file into memory as needed.

What Is Memory Mapping?

Normally, when a Python program reads a file, the data is transferred from the file into a Python object in memory. For example:

with open("data.txt", "rb") as file:
    data = file.read()

If the file is very large, read() may require a large amount of memory because the entire requested content is loaded into the program's memory.

Memory mapping provides another approach. A file can be mapped to a region of virtual memory, allowing the program to access portions of the file directly through that mapping.

import mmap

with open("data.txt", "r+b") as file:
    mm = mmap.mmap(file.fileno(), 0)

    print(mm[:20])

    mm.close()

Here, mmap.mmap() creates a memory-mapped object associated with the file. The 0 means that the entire file is mapped.

Why Use mmap?

The main advantage of mmap is that it can make working with large files more efficient.

Suppose a file contains several gigabytes of data, but the program only needs to locate a particular piece of information. Loading the entire file into memory may be unnecessary. With memory mapping, the program can access the required portions without explicitly reading the entire file into a Python variable.

It can be useful for:

  • Searching large files

  • Processing large binary files

  • Random access to file contents

  • Modifying portions of a file

  • Working with structured binary data

  • Applications that repeatedly access different locations within a file

Creating a Memory-Mapped File

A typical example is:

import mmap

with open("example.bin", "r+b") as file:
    mm = mmap.mmap(file.fileno(), 0)

    print(mm[0:10])

    mm.close()

The file must generally be opened in an appropriate mode depending on what you intend to do. For read-only access, "rb" can be used. For modification, "r+b" is commonly used.

The fileno() method returns the operating-system file descriptor associated with the file.

The resulting mm object behaves in many ways like a mutable byte sequence.

Reading Data Using mmap

A memory-mapped object can be indexed and sliced.

import mmap

with open("data.txt", "rb") as file:
    mm = mmap.mmap(file.fileno(), 0, access=mmap.ACCESS_READ)

    print(mm[0])
    print(mm[0:10])

    mm.close()

When the file contains text, the result is still represented as bytes. Therefore, it may need to be decoded:

print(mm[0:10].decode("utf-8"))

This distinction is important because mmap works primarily with the file's underlying byte representation.

Searching a Large File

One of the useful features of mmap is searching.

import mmap

with open("large_file.txt", "rb") as file:
    mm = mmap.mmap(file.fileno(), 0, access=mmap.ACCESS_READ)

    position = mm.find(b"Python")

    if position != -1:
        print("Found at position:", position)

    mm.close()

The find() method searches for a sequence of bytes. Notice that b"Python" is used rather than "Python" because the memory-mapped data is byte-oriented.

This approach can be particularly convenient when searching large files because the program does not have to explicitly load the entire file into a Python string.

Modifying a File Through mmap

A memory-mapped file can also be modified when it is opened with appropriate permissions.

import mmap

with open("data.txt", "r+b") as file:
    mm = mmap.mmap(file.fileno(), 0)

    mm[0:5] = b"Hello"

    mm.flush()
    mm.close()

Here, the first five bytes of the mapped region are replaced with Hello.

The flush() method requests that modifications made through the mapping be written back to the underlying file.

The replacement data must have the same length as the portion being replaced. For example:

mm[0:5] = b"Hello"

works because both sides contain five bytes.

File Position and seek()

An mmap object also supports operations similar to file objects.

import mmap

with open("data.txt", "rb") as file:
    mm = mmap.mmap(file.fileno(), 0, access=mmap.ACCESS_READ)

    mm.seek(10)
    data = mm.read(20)

    print(data)

    mm.close()

seek() changes the current position within the mapped region, while read() retrieves data starting from that position.

The tell() method can be used to determine the current position:

print(mm.tell())

This makes mmap convenient for applications that need to move around different portions of a large file.

Access Modes

Python provides different access modes for controlling how a memory-mapped region can be used.

mmap.ACCESS_READ creates a read-only mapping.

mm = mmap.mmap(file.fileno(), 0, access=mmap.ACCESS_READ)

mmap.ACCESS_WRITE allows changes to be written back to the underlying file.

mm = mmap.mmap(file.fileno(), 0, access=mmap.ACCESS_WRITE)

mmap.ACCESS_COPY creates a copy-on-write mapping. Changes made to the mapping are not written back to the original file.

Choosing the correct access mode is important because it determines whether the program can modify the underlying data.

Working With Parts of a File

It is not always necessary to map an entire file. A specific portion can be mapped by providing an offset and length where supported.

Conceptually, this allows a program to work with a particular region of a large file rather than treating the entire file as one mapped area.

For example, applications processing very large binary datasets may divide their work into sections and access only the relevant regions.

mmap With Binary Data

The module is particularly useful when working with binary files.

For example:

import mmap

with open("data.bin", "rb") as file:
    mm = mmap.mmap(file.fileno(), 0, access=mmap.ACCESS_READ)

    first_bytes = mm[0:16]

    print(first_bytes)

    mm.close()

The returned value is a sequence of bytes. This makes it possible to combine mmap with other Python modules such as struct when interpreting binary structures.

For example:

import mmap
import struct

with open("data.bin", "rb") as file:
    mm = mmap.mmap(file.fileno(), 0, access=mmap.ACCESS_READ)

    number = struct.unpack("I", mm[0:4])[0]

    print(number)

    mm.close()

Here, four bytes from the mapped region are interpreted as an integer.

Advantages of mmap

Memory mapping provides several important benefits.

First, it can be useful for large-file processing because the application does not necessarily need to load the entire file into a Python object.

Second, it provides efficient random access. A program can jump directly to a particular location rather than sequentially reading everything before it.

Third, it can simplify binary-file processing, because the mapped region can be accessed using indexes and slices.

Fourth, the operating system handles much of the underlying memory management, including bringing required portions of the mapped file into physical memory.

Limitations and Precautions

mmap is not automatically faster for every file-processing task. For small files, ordinary open(), read(), and write() operations may be simpler and sufficiently efficient.

Memory mapping also depends on operating-system facilities, and some behavior can differ between platforms.

Another important consideration is file size and system address space. Mapping a very large file does not necessarily mean that the entire file is physically loaded into RAM, but the mapping itself still consumes virtual address-space resources.

When modifying mapped data, programmers must also ensure that the mapping is writable and that changes are properly flushed when necessary.

mmap vs Normal File Reading

Consider normal reading:

with open("large.txt", "rb") as file:
    data = file.read()

This explicitly creates a Python bytes object containing the requested data.

With mmap:

with open("large.txt", "rb") as file:
    mm = mmap.mmap(file.fileno(), 0, access=mmap.ACCESS_READ)

the file is represented through a memory-mapped object, allowing the program to access portions of it as needed.

Therefore, normal file reading is often preferable for simple and relatively small files, while mmap can be valuable when applications need efficient access to large files or repeated random access.

Conclusion

Python's mmap module provides an alternative way to work with files by mapping file contents into virtual memory. It allows programmers to access, search, and sometimes modify file data using operations similar to those used with byte sequences. Its biggest practical value comes when dealing with large files, random-access workloads, and binary data. Understanding mmap gives Python programmers another powerful technique for designing efficient file-processing applications without unnecessarily loading entire files into Python memory.