Unix - UNIX mmap()-Based Memory Sharing Between Processes
1. Introduction
In UNIX operating systems, multiple processes sometimes need to exchange data or access the same information simultaneously. One way to achieve this is through memory mapping, which allows a file or a region of memory to be mapped into a process's virtual address space. The mmap() system call is used to establish this mapping.
Normally, each process has its own virtual address space, meaning that one process cannot directly access another process's private memory. However, UNIX provides mechanisms that allow processes to share specific memory regions. When two or more processes map the same underlying memory object using suitable sharing options, they can access common data without repeatedly transferring it through pipes or other communication mechanisms.
Memory sharing through mmap() is useful in applications that require efficient data exchange, such as database systems, multimedia applications, large-data processing programs, and applications that run multiple cooperating processes.
2. What Is the mmap() System Call?
The mmap() system call creates a mapping between a file or another supported memory object and a process's virtual address space. Once the mapping is established, the process can access the mapped region using ordinary memory operations, such as reading and writing through pointers.
The general syntax is:
C
#include <sys/mman.h>
void *mmap(
void *addr,
size_t length,
int prot,
int flags,
int fd,
off_t offset
);
The parameters have the following meanings:
-
addr: Specifies a preferred starting address for the mapping. In most applications,NULLis supplied so the operating system can select a suitable address. -
length: Specifies the number of bytes to map. -
prot: Defines the permitted memory operations, such as reading, writing, or executing. -
flags: Specifies the mapping type, including whether changes are shared or private. -
fd: Identifies the open file to be mapped. For anonymous mappings, this parameter is generally unused. -
offset: Specifies the offset within the file where the mapping begins. It must satisfy the system's alignment requirements.
If the call succeeds, mmap() returns the starting address of the mapped region. If it fails, it returns MAP_FAILED.
The mapping can later be removed using munmap(), which releases the specified mapped address range from the process's virtual address space.
3. How Memory Sharing Works Between Processes
Memory sharing occurs when multiple processes map the same underlying memory object with the MAP_SHARED flag. The mapped regions do not have to appear at the same virtual address in each process. The operating system manages the relationship between the mappings and the underlying object.
For example, suppose Process A and Process B need to share a counter. Instead of sending every update through a pipe, both processes can map the same shared-memory region. Process A can update the counter, and Process B can read the updated value through its own mapping.
The basic sequence is as follows:
-
A shared memory object or suitable backing file is created or opened.
-
The object is prepared to hold the required data, if necessary.
-
Each process maps the object into its virtual address space using
mmap()withMAP_SHARED. -
The processes read and write the mapped region according to an agreed data layout.
-
Synchronization mechanisms are used when processes may access or modify the same data concurrently.
-
Each process removes its mapping with
munmap()when it is no longer needed.
A common backing object for this purpose is a POSIX shared-memory object created using shm_open(). It can be sized with ftruncate() and then mapped using mmap(). A suitable file can also be used when file-backed shared memory is appropriate.
It is important to understand that mapping the same file with MAP_PRIVATE does not provide the same shared-write behavior. With MAP_PRIVATE, modifications are private to each process and are not intended to become visible to other mappings.
4. Example of Sharing Memory Between Processes
The following C program demonstrates the main idea using a POSIX shared-memory object. The parent process creates and maps the shared object, writes a value, and then starts a child process. The child inherits the mapping through fork() and reads the value from the shared region.
C
#include <stdio.h>
#include <stdlib.h>
#include <unistd.h>
#include <fcntl.h>
#include <sys/mman.h>
#include <sys/stat.h>
#include <sys/wait.h>
int main(void) {
const char *name = "/unix_shared_example";
int fd = shm_open(name, O_CREAT | O_RDWR, 0600);
if (fd == -1) {
perror("shm_open");
return 1;
}
if (ftruncate(fd, sizeof(int)) == -1) {
perror("ftruncate");
close(fd);
shm_unlink(name);
return 1;
}
int *shared_value = mmap(
NULL,
sizeof(int),
PROT_READ | PROT_WRITE,
MAP_SHARED,
fd,
0
);
if (shared_value == MAP_FAILED) {
perror("mmap");
close(fd);
shm_unlink(name);
return 1;
}
*shared_value = 100;
pid_t pid = fork();
if (pid == -1) {
perror("fork");
munmap(shared_value, sizeof(int));
close(fd);
shm_unlink(name);
return 1;
}
if (pid == 0) {
printf("Child reads: %d\n", *shared_value);
munmap(shared_value, sizeof(int));
close(fd);
_exit(0);
}
waitpid(pid, NULL, 0);
munmap(shared_value, sizeof(int));
close(fd);
shm_unlink(name);
return 0;
}
A typical output is:
Child reads: 100
This example demonstrates that the child can access the shared value through the inherited MAP_SHARED mapping. The example uses fork() to create the child; unrelated processes can also share a memory object if they open and map the same object appropriately.
To compile this program on a system that supports POSIX shared memory:
Bash
gcc shared_memory.c -o shared_memory -lrt
On some modern systems, the real-time library option -lrt is unnecessary. The program can then be executed with:
Bash
./shared_memory
The example shows shared access, but it does not demonstrate simultaneous updates. If several processes modify a shared variable at the same time, synchronization is required to avoid race conditions.
5. Synchronization and Data Consistency
Sharing memory does not automatically make concurrent access safe. If two processes attempt to update the same variable simultaneously, one update may overwrite another. This is known as a race condition.
For example, if two processes both increment a shared counter, each might read the value 10 before either writes its result. Both could then write 11, even though the expected result is 12.
To prevent this, processes need a synchronization mechanism. Depending on the application and platform, suitable mechanisms include process-shared POSIX semaphores, process-shared mutexes configured with the appropriate attributes, or atomic operations supported for interprocess use by the platform.
Synchronization is particularly important when shared memory contains complex data structures, such as queues, tables, indexes, or records. Processes must agree on how data is organized, when it is valid to read, and how updates are coordinated.
Programmers must also consider memory visibility and ordering. A process should not assume that simply writing a value guarantees that every other process can safely observe a consistent multi-step update. Appropriate synchronization establishes the required coordination between readers and writers.
6. Advantages of mmap()-Based Memory Sharing
Memory mapping provides several advantages for applications that exchange large amounts of data.
Efficient data access: Processes can access mapped data using normal memory operations rather than repeatedly copying data through an explicit communication channel.
Reduced copying: Depending on the implementation and workload, mapping can reduce the number of data copies required to exchange information. This can improve performance when handling large files or shared datasets.
Convenient file access: File-backed mappings allow programs to work with file contents through memory-like access instead of repeatedly using file-read operations.
Support for cooperating processes: Multiple processes can access a common data structure, making memory mapping useful for applications that divide work among several processes.
Flexible memory management: Mappings can be established for selected regions and removed when no longer required. The operating system manages the virtual-memory mappings and the associated physical-memory pages.
However, memory mapping is not always faster than every alternative. Performance depends on the size of the data, access patterns, synchronization overhead, memory pressure, and the operating system's implementation.
7. Limitations and Important Considerations
Although mmap()-based memory sharing is powerful, it requires careful design.
First, the processes must agree on the shared data structure and its layout. If one process interprets a region differently from another, the data may become corrupted or unusable.
Second, pointers stored inside shared memory require special attention. Different processes can map the same object at different virtual addresses. Therefore, an ordinary pointer created by one process may not refer to the intended object when interpreted by another. Relative offsets or carefully designed position-independent structures are often more suitable.
Third, the application must manage synchronization correctly. Shared memory can improve communication efficiency, but incorrect coordination may introduce race conditions, inconsistent data, or deadlocks.
Fourth, file-backed mappings have specific behavior. Changes made through MAP_SHARED are shared with other mappings of the same object, but persistence to storage and the exact visibility of changes depend on the backing object and operating-system semantics. Applications that require data to be written to storage may need mechanisms such as msync() or appropriate file-synchronization operations.
Finally, the application must handle errors and resource cleanup. Mapping failures, insufficient permissions, invalid offsets, truncated backing files, and resource limits can affect program behavior.
8. Practical Applications
UNIX mmap()-based memory sharing is used in a variety of real-world applications.
-
Database systems: Shared buffers and memory-mapped data structures can help database processes access common information efficiently.
-
Large-file processing: Applications can map portions of large files and access their contents without loading the entire file into a conventional application buffer.
-
Interprocess communication: Cooperating processes can exchange data through a common mapped memory region.
-
Multimedia applications: Programs may use mapped buffers to coordinate access to large image, audio, or video data.
-
Scientific computing: Multiple processes can work with shared datasets while reducing unnecessary data copying.
-
Caching systems: Shared data structures can provide efficient access to frequently used information when synchronization and consistency are handled correctly.
The precise implementation varies across UNIX systems and UNIX-like operating systems, so developers should consult the documentation for their target platform.
9. Difference Between mmap() and Traditional Interprocess Communication
Traditional interprocess communication methods include pipes, message queues, and sockets. These methods allow processes to exchange data through defined communication interfaces. With pipes and message queues, data is generally transferred through the communication mechanism. With mmap() and MAP_SHARED, processes access a common mapped memory object directly.
The key difference is that memory mapping provides a shared data area, while traditional communication mechanisms often provide an explicit means of sending and receiving messages. Shared memory can be advantageous for large or frequently accessed datasets, but it places greater responsibility on the application to coordinate access and maintain data consistency.
Neither approach is universally better. Pipes and message queues may be simpler for structured communication, while shared memory may be more suitable when several processes need frequent access to the same large data structure.
10. Conclusion
UNIX mmap()-based memory sharing is an important technique for enabling efficient communication between processes. By mapping the same underlying memory object into multiple virtual address spaces with MAP_SHARED, applications can let processes read and modify common data without relying exclusively on explicit message transfers.
Its effectiveness depends on proper memory mapping, suitable permissions, consistent data structures, synchronization, and resource management. When these aspects are handled correctly, shared memory can improve efficiency and provide a practical foundation for high-performance, multi-process applications.