Why are my builds slow? Using CMake’s Instrumentation Feature

October 11, 2026
CMake Logo

Summary

Resources

Overview

CMake 4.2 dropped the experimental gate on a major new feature: the CMake Instrumentation API. This enables detailed tracking of the entire CMake workflow, including configure time, build execution, testing, and installation, and provides developers and teams with actionable insights into build performance. Support for automated, user-written scripts to handle instrumentation data make it easy to track performance data over time across different machines and environments.

This post covers an explanation of the feature, a summary of what’s new since the last post about CMake Instrumentation, and some practical guidance for how to use CMake instrumentation to diagnose build time inefficiencies in your own projects.

What is CMake Instrumentation?

At a high level, CMake Instrumentation allows you to collect detailed metrics about every step of your CMake workflow. This turns opaque build or test times into actionable results, enabling users to find bottlenecks and inefficiencies, develop real performance improvements, or make data-driven decisions about how and when to run their build or tests.

Without CMake instrumentation, your build output might look something like this:

myproject/build $ cmake ..
-- Configuring done (30.2s)
-- Generatingdone (5.2s)
myproject/build $ ninja
[3041/3041] Linking CXX static library project/libproject.a
myproject/build $

Configure and generate give us some high level timing information, but there is very little to work with. Even if we timed the build command, we’d have no visibility into where that time was going. Using CMake instrumentation, you could write a callback to print out additional details after each build, customized to your liking:

=== CMake build instrumentation summary ===
---------------------------------------------
Build elapsed: 184.732 s
Role          Commands     Time (s)   Peak RSS (MiB)
----------------------------------------------------
compile           2860     1248.615           1842.7
link               121       38.924            612.3
custom              60       26.481            128.6
Total command time: 1314.020 s; failed commands: 0

Top 3 targets by time:
---------------------------------------------
Target                    Commands     Time (s)
-----------------------------------------------
project                        842      386.214
geometry                       516      241.873
rendering                      438      207.562

Top 3 commands by memory:
---------------------------------------------
Peak RSS (MiB)     Time (s)  Command
--------------------------------------
1842.7             8.614     compile: project
                             source: project/src/Registry.cpp
1624.3             7.283     compile: geometry
                             source: geometry/src/MeshOperations.cpp
1496.8             6.941     compile: rendering
                             source: rendering/src/Render.cpp

Trace: /path/to/myproject/build/instrumentation/traces/postBuild-3044-2026-10-09T14-20-00-0123.json

Or, generate Google Trace Event format files to visualize your build at a much more granular level:

A build profiling trace visualization showing multi-threaded CPU compilation activity across 19 threads over an 18-second execution timeline.

Furthermore, because user callbacks run automatically, you could easily submit the data from every build to a database, and track changes in performance over time across different machines or environments.

CMake instrumentation works by wrapping every compile, link, custom command, and test invocation in a launcher that collects various metrics about runtime and resource usage. These metrics get written to a JSON file called a “snippet”. A single snippet file is generated for every compile, link, and test invocation, as well as for the overall configure, generate, build and test phases. Over the course of a CMake workflow, these snippet files will accumulate in the project build tree until indexing occurs. Indexing can be run manually with  ctest --collect-instrumentation, or automatically at configured intervals, such as after each build. This process gathers up all these snippet files and generates a top-level “index” JSON file. The index is then passed off as an argument to any user-written callbacks as an entry-point to the data, after which CMake cleans up the snippet files, preventing infinite accumulation. Some advice on writing instrumentation callbacks is included later in this post.

A flowchart titled "How Instrumentation Works" illustrating CMake generating JSON snippet files, indexing them into an index file, triggering user callbacks, and deleting the temporary instrumentation data.

Enabling CMake instrumentation can be done at the “user-level” by placing a JSON query in the user’s CMAKE_CONFIG_DIR to run on all CMake projects. Alternatively, project-specific queries like the one below can be configured by using the cmake_instrumentation() command in the project’s CMakeList.txt.

cmake_instrumentation(
  API_VERSION 1
  DATA_VERSION 1
  OPTIONS
    staticSystemInformation  # collect static system data
    dynamicSystemInformation # collect dynamic system data 
    trace                    # generate a Google Trace Event file
    cdashSubmit              # submit data to CDash
    captureOutput            # include stdOut and stdErr in data
    compiileTrace            # collate compiler generated traces
    processMetrics           # collect maxRSS, userTime and systemTime
  HOOKS postGenerate postBuild postCTest
     # when to index data and run CALLBACK
  CALLBACK "${CMAKE_COMMAND}" -P /path/to/handle_data.cmake
     # index file will be appended to args
)


Here’s a summary of what commands are instrumented, and what are our available hooks for running a custom instrumentation callback.

Instrumented Commands

  • configure: the CMake configure step
  • generate: the CMake generate step
  • compile: an individual compile step invoked during the build
  • link: an individual link step invoked during the build
  • custom: an individual custom command invoked during the build
  • build: a complete make or ninja invocation (not through cmake --build).
  • cmakeBuild: a complete cmake --build invocation
  • cmakeInstall: a complete cmake --install invocation
  • install: an individual cmake -P cmake_install.cmake invocation
  • ctest: a complete ctest command invocation
  • test: a single test executed by ctest

Instrumentation Hooks

  • postGenerate
  • preBuild (called when ninja or make is invoked)
  • postBuild (called when ninja or make completes)
  • preCMakeBuild (called when cmake --build is invoked)
  • postCMakeBuild (called when cmake --build completes)
  • postCMakeInstall
  • postCMakeWorkflow
  • postCTest

See the documentation for the full use case and available options.

What’s new with CMake Instrumentation?

CMake 4.4 and the upcoming CMake 4.5 release introduce several big features for instrumentation. See the Data Version documentation for details about which features are available in which versions of CMake. You can also parse the output of cmake -E capabilities to see which CMake instrumentation data versions your CMake supports.

The newest features include the following options:

  • trace: automatically generate a Google Trace Event format file when indexing occurs, for easy visualization
  • captureOutput: enable reporting stdOut and stdErr of each command
  • compileTrace: capture any compiler trace files generated by clang’s -ftime-trace flag, and link them to the associated compile data in the output
  • processMetrics: capture maxRSS, userTime and systemTime for each command

These additional metrics provide further insight into a complete CMake workflow. Metrics such as the complete compiler trace and memory usage per-command allow users to identify why certain segments take as long as they do with greater precision, and to track more complex performance regressions beyond the total duration.

There’s also been updates to how instrumentation data is viewable in CDash, including the ability to visualize instrumentation data for tests, and to see a heat map of memory usage for each instrumented command.

A CDash test dashboard interface titled "Linux-c++" displaying a color-coded test execution timeline over three minutes, a hovered test tooltip, and a detailed test results table highlighting failed build tests.
A build instrumentation analysis view in CDash showing a parallel execution timeline filtered by MaxRSS, with a tooltip detailing compilation metrics for a C++ source file.

CMake’s nightly CI includes a build with instrumentation. You can find the latest runs here for an interactive example. After selecting a recent build, find Instrumentation in the sidebar to visualize it.

A CDash build summary page for "nightly-cmake-fedora44_ninja_instrumentation" with a red arrow pointing to the "Instrumentation" navigation item, alongside build error tables and a total execution time history chart.

You can also check out the original blog post about CDash integration with CMake instrumentation for more detail.

Writing Callbacks for CMake Instrumentation

Writing custom callback scripts is required because the consumption of instrumentation data will be specific to the needs of the project or user. These scripts are called whenever indexing occurs, passing an index JSON file as an argument to serve as an entry-point to the collected data. Once all user callbacks have been executed, CMake removes the generated snippet files, so it’s up to the callback to preserve any data, by storing it elsewhere or submitting to a database.

When writing an instrumentation callback, it’s best to consult the documentation for the index file contents. Be sure to reference the documentation for your current CMake version. Here’s a small, example callback that simply copies the instrumentation data to a separate directory:

import json
import os
import shutil
import sys

if name == "main":
    # The instrumentation feature will pass the path to an index file to this callback
    index = sys.argv[1]

    # Get the buildDir and dataDir from the index file
    with open(index) as f:
        data = json.load(f)
        buildDir = data["buildDir"]
        dataDir = data["dataDir"]

    # Get a unique output directory name based on the index filename
    indexName = os.path.basename(index).split(".")[0]

    # Create an output directory that CMake won't clear, to copy the files into
    outputDir = os.path.join(buildDir, "instrumentation")
    indexDir = os.path.join(outputDir, indexName)
    for d in [outputDir, indexDir]:
        if not os.path.exists(d):
            os.makedirs(d)

    # Copy all the instrumentation data into the newly created output directory
    for snippet in data["snippets"]:
        shutil.copyfile(
            os.path.join(dataDir, snippet),
            os.path.join(indexDir, snippet)
        )
    shutil.copyfile(
        index,
        os.path.join(indexDir, os.path.basename(index))
    )

You can find a more complicated example callback here. The example prints out various instrumentation metrics after each build, install, and ctest invocation.

The following code-snippet could be inserted into a CMakeLists.txt to enable this callback.

cmake_instrumentation(
  API_VERSION 1
  DATA_VERSION 1.2
  OPTIONS processMetrics trace
  HOOKS postCMakeBuild postBuild postCMakeInstall postCTest
  CALLBACK "${Python3_EXECUTABLE}" /path/to/summary.py
)

If you want to enable this callback for every CMake project run by your user, you could instead place a JSON query under $CMAKE_CONFIG_DIR/instrumentation/v1/query

{
  "version": { "major": 1, "minor": 2 },
  "options": [ "processMetrics", "trace" ],
  "hooks":  "postCMakeBuild", "postBuild", "postCMakeInstall", "postCTest" ],
  "callbacks": ["/path/to/python3 /path/to/summary.py"]
}

Earlier in this blog, we showed the output of this example callback running in a post CMake build context. We can also expect some output summarizing an invocation of ctest.

=== CTest memory diagnostics ===
--------------------------------------------------------
  Peak RSS (MiB)     Time (s)     Result  Test
----------------------------------------------
           206.6        0.107     passed  memory.192MiB
            78.5        0.053     passed  memory.64MiB
            30.6        0.035     passed  memory.16MiB
            19.5        0.236     passed  cpu.500000
            19.2        0.018     passed  app.smoke
            19.2        1.755     passed  cpu.4000000

Trace: /path/to/instrumentation/traces/postCTest-7-2026-10-09T13-45-59-0657.json

In general, instrumentation callbacks can be configured to run after each cmake --build invocation (postCMakeBuild), or after each invocation of the build tool: make, ninja or fbuild (postBuild). However, note that the postBuild hook runs as a background process, and the output from our callback will not be printed to the command-line in this case.

A Practical Exercise: Why are my builds slow?

Thanks for reading this far! To finish, here’s a walkthrough example to demonstrate how CMake instrumentation can help you make informed decisions about your build performance, and point out some common ways CMake projects might suffer from reduced parallelism.

For this exercise, an application links two groups of libraries:

  • A series of “stage” libraries, that each link to a “config” library. The config library has a private header that gets generated at build time.
  • A chain of “compute” and “schema” libraries with generated source files, that depend on one another.

1. Establish a baseline

To establish a baseline, I’ll run a clean CMake build with a fixed parallelism. The instrumentation callback from earlier is enabled, and I’m generating a Google Trace Event file that will be used to visualize the build.

cmake -S . -B build -G Ninja
cmake --build build-original --clean-first --parallel 6

Because I enabled the example instrumentation callback linked above, this will immediately print the output of that report.

=== CMake build instrumentation summary ===
-------------------------------------------
Build elapsed: 12.217 s

Role         Commands     Time (s)   Peak RSS (MiB)
----------------------------------------------------
compile      22           28.610     234.4
link         64           1.970      194.0
custom       9            7.137      20.2
Total command time: 37.717 s; failed commands: 0

Top 4 targets by time:
----------------------------------------------------
Target                    Commands     Time (s)
-----------------------------------------------
stage_1                          4        4.941
stage_3                          4        4.854
stage_2                          4        4.815
stage_0                          4        4.812

Top 4 commands by memory:
---------------------------------------------------
  Peak RSS (MiB)     Time (s)  Command
--------------------------------------
           234.4        4.718  compile: stage_2
           234.2        4.761  compile: stage_3
           233.7        4.720  compile: stage_0
           233.6        4.850  compile: stage_1

Trace: /path/to/build/instrumentation/traces/postCMakeBuild...json

This tells us the build took just over 12 seconds, and gives an overview of the slowest commands, but it doesn’t give the full picture. Opening the generated Google Trace Event file will show us more. You can open these files in Perfetto UI, or any other compatible viewer.

A build profiler Gantt chart across 7 threads showing a build bottleneck caused by a single long-running header generation task (config.hpp) on Thread 1 prior to parallel compilation.

It’s clear from this visualization that the 6 threads we gave our build are not being fully utilized.

2. Define the visibility of our generated files

One thing that might stick out is the generation of config.hpp, a private header file, which is slow, and bottlenecking the rest of our build.

The generated header is initially attached like this:

add_library(config STATIC config.cpp)
target_sources(config PRIVATE "${generated_dir}/config.hpp")
target_include_directories(config PRIVATE "${generated_dir}")

add_library(stage_0 STATIC stage_0.cpp)
target_link_libraries(stage_0 PRIVATE config)

add_custom_command(
  OUTPUT "${generated_dir}/config.hpp"
  COMMAND …
)

Although we specified the generated source file as PRIVATE, CMake creates an extra dependency here. In case any of our consuming targets end up needing access to this source file, we wait for its generation before beginning to compile them. This is why we can see the compilation of “stage_0”, “stage_1”, “stage_2”, and “stage_3” waiting for the generation to finish before starting.

In this case, however, the generated file is a private implementation detail. If we define it as such, using FILE_SETs, our dependent targets will be able to start compiling much sooner.

target_sources(config PRIVATE
  FILE_SET implementation_headers TYPE HEADERS
  BASE_DIRS "${generated_dir}"
  FILES "${generated_dir}/config.hpp"
)

After running a clean build and opening the generated trace file, we can see the impact this had on our build parallelism:

A build profiler Gantt chart showing an optimized build pipeline (build: private) where compilation tasks (stage_0 through stage_2) run immediately on parallel threads alongside background header generation (config.hpp).

Now we can compile these sources in parallel with the code generation, and our overall build time is reduced.

3. Optimize Static Library Dependencies

There’s still another bottleneck in our trace. Looking at the “compute_*” targets, we can see a repeated sequence of code generation, compilation, and linking. This leaves much of the build running one stage at a time.

Here’s a simplified piece of that dependency chain:

add_custom_command(
  OUTPUT "${generated_dir}/schema_1.cpp"
  COMMAND …
  DEPENDS "${schema_input}"
  VERBATIM
)

add_library(schema_1 STATIC "${generated_dir}/schema_1.cpp")
target_link_libraries(schema_1 PRIVATE compute_0)

add_library(compute_1 STATIC compute_1.cpp)
target_link_libraries(compute_1 PRIVATE schema_1 compute_0)

By default, when a static library links to another target through target_link_libraries(), CMake adds build-order dependencies on that target, even though creating the static archive does not require the linked library to be built. Those ordering constraints can also cause custom commands that generate sources for the library to wait for upstream targets. CMake does this conservatively because an upstream target might produce headers or other files needed by downstream work. Enabling the OPTIMIZE_DEPENDENCIES option examines the dependencies of static and object libraries and removes unnecessary build-order constraints, while retaining explicit ordering and dependencies associated with code generation or other relevant side effects. The final executable still waits for the libraries it needs to link, but independent generation and compilation can start sooner.

set(CMAKE_OPTIMIZE_DEPENDENCIES ON)

This initializes the OPTIMIZE_DEPENDENCIES property on new targets. If we wanted to apply it to just one existing target, we could instead use:

set_property(TARGET schema_1 PROPERTY OPTIMIZE_DEPENDENCIES ON)

After another clean build, we can see the difference:

The schema generators can now run without waiting for the preceding compute archives. We can see generation advancing through the chain while the compute sources compile in parallel. The final application still waits for all of its required libraries before linking.

A build profiler Gantt chart showing a fully optimized build pipeline (build: optimized) with parallel compilation tasks executing immediately across threads while config.hpp generates in the background, achieving a reduced 11.187-second runtime

4. Lessons Learned

A slow build isn’t always caused by long running compilations or links. Sometimes the problem is the time spent waiting for work that isn’t actually needed yet. Visualizing the parallelism of our build makes it easy to see those waits, and validate our performance improvements. Specifying the visibility of our generated files, and using less conservative build edges with OPTIMIZE_DEPENDENCIES, can be a few easy ways to gain significant speed up to our build.

In order to dig deeper into the performance of individual commands, instrumentation supports the collation of generated compiler trace files. You can read more about understanding those files here.

If you wanted to take a closer look at where CMake configure time is going, you can generate a Google Trace event file for the configure and generate stages as well, without needing to use CMake instrumentation. See the –profiling-output and –profiling-format options when running CMake.

Tags:

Leave a Reply