docs: note the call overhead of a non-LTO precompiled library

Release build, trivial bound functions, against an LTO header-only
module: AppleClang arm64 adds 2-4 ns per call (7-11%), GCC 15 Linux
aarch64 adds 1-2 ns (1-7%). With LTO on the library, both are within a
few percent.

Assisted-by: ClaudeCode:claude-opus-5-5
diff --git a/docs/compiling.rst b/docs/compiling.rst
index 42cd320..a81b16f 100644
--- a/docs/compiling.rst
+++ b/docs/compiling.rst
@@ -450,7 +450,10 @@
   status message reports the directory that created the library.
 * The library is not compiled with link-time optimization, and the per-target
   ``THIN_LTO`` and ``OPT_SIZE`` options of ``pybind11_add_module`` do not
-  apply to it. To change this, call ``pybind11_precompile()`` yourself and
+  apply to it. In a Release build this adds a few nanoseconds to each call of
+  a bound function (up to about 10% for a function that does nothing,
+  depending on the compiler; a few percent with LTO on the library). To change this, call ``pybind11_precompile()``
+  yourself and
   set the properties on the created target, ``pybind11_precompiled`` (the
   real target behind the ``pybind11::precompiled`` alias; CMake does not let
   you set properties through an alias):