Unix - UNIX System Call Tracing with strace and truss?

1. Introduction to System Call Tracing

System call tracing is a debugging and diagnostic technique used to observe how a program interacts with the operating system. In UNIX, applications cannot directly access many protected resources, such as files, devices, process information, or network connections. Instead, they request these services from the operating system kernel through system calls.

System call tracing helps developers and system administrators understand which requests a program makes, what results it receives, and where errors occur. Tools such as strace and truss record these interactions, making it easier to investigate application failures, performance problems, and unexpected system behavior.

For example, suppose a program fails to open a configuration file. The program might display only a general error message. A system call tracer can reveal whether the program attempted to open the correct file, whether the file existed, and whether the operating system returned a permission error. This information helps identify the actual cause of the problem.

2. Understanding strace and truss

strace and truss are command-line utilities used to trace system calls and, depending on the platform and configuration, related signals or process events.

strace is widely used on Linux systems. It records system calls made by a program, along with their arguments, return values, and associated error information. It can also follow child processes created by the traced program, which is useful when an application launches other programs.

truss is available on certain UNIX systems, including Solaris and some other UNIX variants. It provides similar diagnostic capabilities by showing system calls and their results. Its supported options and output format differ according to the operating system.

Although both tools serve a similar purpose, they are not interchangeable in every environment. Linux generally uses strace, while systems that provide truss may use it for equivalent tracing tasks. Before using either tool, users should consult the manual page available on their particular system.

3. How System Call Tracing Works

When a program runs, it frequently requests services from the operating system. These requests may involve opening files, reading data, writing output, allocating resources, creating processes, or communicating through sockets.

A system call tracer observes these interactions and displays useful information about them. The general process involves the following steps.

First, the user starts a program under the control of a tracing utility. The utility establishes a tracing relationship with the program and monitors the system calls it makes.

Second, the program begins executing normally. Whenever it makes a system call, the tracing utility records relevant information, such as the system call name and its arguments.

Third, after the operating system processes the request, the tracer records the return value. If the operation fails, the output may include an error code and a descriptive message.

Finally, the user examines the trace output to understand the program's behavior and identify possible problems.

For example, consider the following command:

Bash

strace ls

This command runs the ls program under strace on Linux. The output may contain system calls such as openat(), getdents64(), newfstatat(), and write(). These calls help explain how the program accesses directory information and displays filenames.

The exact system calls and their order depend on the program, operating system version, environment, and available resources.

4. Basic Commands for System Call Tracing

Several commands can be used to perform common tracing tasks.

A. Trace a program from startup

Bash

strace ls

This command traces the execution of ls, including system calls made during program startup and directory listing. It is useful for understanding the overall interaction between a program and the operating system.

B. Trace an existing process

Bash

strace -p 1234

This command attaches strace to a running process whose process ID is 1234. Replace this number with the actual process ID.

This method is useful when a program is already running and experiencing a problem. For example, if an application appears to be stuck, tracing it may reveal whether it is waiting for input, repeatedly checking a resource, or blocked on a system call. Attaching to a process may require appropriate permissions.

C. Trace file-related system calls

Bash

strace -e trace=%file ls

This command filters the trace to file-related system calls supported by the installed version of strace. It helps investigate file access, directory lookups, and path-related errors without displaying every system call.

For example, it can help determine whether an application is searching for a configuration file in an unexpected directory.

D. Save trace output to a file

Bash

strace -o trace.log ls

This command saves the trace output to trace.log rather than displaying it directly in the terminal. Saving output is useful when a trace is lengthy or needs to be examined later.

E. Follow child processes

Bash

strace -f ./program

The -f option follows child processes created by the traced program. This is particularly useful when a program launches additional processes or executes other commands. Without following child processes, important operations performed by those children might not appear in the trace.

On systems that use truss, similar tasks can be performed with the options supported by that implementation. Users should check man truss for the correct syntax.

5. Understanding System Call Output

A system call trace typically includes the system call name, the arguments passed to it, and the result returned by the operating system.

Consider this simplified example:

openat(AT_FDCWD, "/tmp/example.txt", O_RDONLY) = 3
read(3, "Hello UNIX\n", 11) = 11
close(3) = 0

The first line indicates that the program requested to open /tmp/example.txt in read-only mode. The operation succeeded and returned the file descriptor 3, which the program can use to refer to the open file.

The second line shows that the program read 11 bytes from that file descriptor. The returned value 11 indicates the number of bytes successfully read.

The third line shows that the program closed the file descriptor successfully. A return value of 0 indicates success for this operation.

Now consider a failed file operation:

openat(AT_FDCWD, "/tmp/missing.txt", O_RDONLY) = -1 ENOENT (No such file or directory)

This output indicates that the program attempted to open a file that could not be found at the specified path. The ENOENT error explains why the operation failed.

Common errors that may appear in traces include:

  • ENOENT: The specified file or directory does not exist.

  • EACCES: The operation was denied because of insufficient permissions.

  • EINTR: A system call was interrupted by a signal.

  • EBADF: The program used an invalid file descriptor.

Understanding these return values helps users distinguish between application logic errors, missing resources, permission problems, and other operating system errors.

6. Practical Applications of System Call Tracing

System call tracing is useful in several areas of UNIX system administration and software development.

Debugging file access problems: A trace can reveal which files an application attempts to open and whether those operations succeed. This helps diagnose missing configuration files, incorrect paths, and permission issues.

Investigating application failures: When an application exits unexpectedly or produces an error, the trace may identify the system call that failed immediately before the problem occurred.

Diagnosing performance issues: A trace can reveal repeated system calls, excessive file access, unnecessary polling, or long waits on certain operations. However, tracing adds overhead, so timing measurements should be interpreted carefully.

Understanding process behavior: By following child processes and observing process-related system calls, developers can investigate how an application starts other programs and manages its execution.

Examining network activity: Network-related system calls, such as socket(), connect(), sendto(), and recvfrom(), may reveal how a program establishes connections and exchanges data. A system call trace does not necessarily show the full contents of encrypted network traffic.

Learning operating system concepts: Students can use tracing to observe how high-level operations in a program translate into requests to the kernel. This provides a practical connection between application code and UNIX operating system concepts.

7. Limitations and Precautions

Although system call tracing is powerful, it has limitations. A trace records system-level interactions rather than explaining the entire internal logic of an application. A successful system call does not necessarily mean the overall program operation succeeded, and a failed call does not always indicate a serious problem because some applications intentionally test for the existence of resources.

Tracing can also slow down a program, especially when it generates a large number of system calls. This may affect the behavior being investigated and distort performance measurements.

Permissions are another consideration. Attaching to another process may be restricted by system security settings, user privileges, or operating system policies. System call traces may also contain sensitive information, including file paths, usernames, command-line arguments, and data passed to system calls. Trace files should therefore be handled carefully.

Finally, strace and truss are platform-specific tools with different supported features. Commands written for Linux strace may not work with truss on another UNIX system. Users should verify the appropriate options and consult the relevant manual pages.

8. Conclusion

UNIX system call tracing with strace and truss is an important technique for understanding how applications interact with the operating system kernel. By recording system calls, arguments, return values, and errors, these utilities help developers and administrators diagnose file access failures, investigate process behavior, identify potential performance problems, and understand application execution.

Learning to read trace output is a valuable skill for anyone studying UNIX internals, software debugging, or system administration. It enables users to move beyond general error messages and examine the underlying operating system operations responsible for a program's behavior.