docs: note the call overhead of a non-LTO precompiled library Release build, trivial bound functions, against an LTO header-only module: AppleClang arm64 adds 2-4 ns per call (7-11%), GCC 15 Linux aarch64 adds 1-2 ns (1-7%). With LTO on the library, both are within a few percent. Assisted-by: ClaudeCode:claude-opus-5-5
diff --git a/docs/compiling.rst b/docs/compiling.rst index 42cd320..a81b16f 100644 --- a/docs/compiling.rst +++ b/docs/compiling.rst
@@ -450,7 +450,10 @@ status message reports the directory that created the library. * The library is not compiled with link-time optimization, and the per-target ``THIN_LTO`` and ``OPT_SIZE`` options of ``pybind11_add_module`` do not - apply to it. To change this, call ``pybind11_precompile()`` yourself and + apply to it. In a Release build this adds a few nanoseconds to each call of + a bound function (up to about 10% for a function that does nothing, + depending on the compiler; a few percent with LTO on the library). To change this, call ``pybind11_precompile()`` + yourself and set the properties on the created target, ``pybind11_precompiled`` (the real target behind the ``pybind11::precompiled`` alias; CMake does not let you set properties through an alias):