Modern C++ Python Bindings Reflection C++26 Performance Cloud

Compute has a turkey problem

Reported compute price increases in 2026: Hetzner +40%, Intel CPUs +10%, TSMC wafers +10%, OVHcloud +8%

We are not used to seeing charts like this one. Most of us have spent our careers assuming that CPUs are abundant and will cost less next year. In 2026 that stopped being true, at least for a while. Hetzner raised its prices, citing component costs, and OVHcloud told customers to expect increases too. Further down the supply chain, Intel raised prices on CPUs it launched three years ago, and TSMC is charging more for all of its advanced nodes.

Things were not always this way. The computer that landed people on the Moon had about 4 KB of RAM and ran roughly 85,000 instructions per second. Printed out, its source code made a stack of paper as tall as Margaret Hamilton, who led the team that wrote it.

During Apollo 11’s descent, a radar left in the wrong switch position flooded the computer with counter interrupts that ate about 15% of its cycles. Minutes before touchdown, it ran out of headroom and raised the famous 1201 and 1202 alarms. Those alarms were part of the design. The software restarted, dropped the low-priority jobs and kept the landing-critical ones running on the cycles that were left. Mission control said “go”, and Eagle landed. Don Eyles, who wrote the landing software, tells the story in Tales from the Lunar Module Guidance Computer.

Left: Margaret Hamilton beside the printed Apollo Guidance Computer listings, 1969. Right: the LUMINARY source lines that raise the 1201 and 1202 alarms, and the Apollo 11 transcript of the moment they fired Left: Hamilton with the printed listings of the software her team wrote. Right: the lines in the lunar module software that raise the 1201 and 1202 alarms, and the transcript of the moment they fired.

Later, cycles became plentiful and we stopped counting them. Niklaus Wirth was already complaining about bloated software in 1995, in A Plea for Lean Software, and “software is getting slower more rapidly than hardware becomes faster” is now known as Wirth’s law. Casey Muratori’s “Clean Code, Horrible Performance” is a recent take on the same problem. To be fair, putting an interpreted language in the hot path was sometimes the rational choice, because engineer time cost more than CPU time.

CPU Thanksgiving

In The Black Swan, Taleb tells the story of a turkey that is fed every day for a thousand days. Plot its well-being and you get a steadily rising line, and every new day of data makes extrapolating that line look safer. Then, a few days before Thanksgiving, the process behind the data changes, and nothing in the previous thousand points warned about it. That is the turkey problem. A long, consistent trend describes the past of a process, but it tells you nothing about whether the process will keep running.

Decades of cheaper compute is a trend like that, and it was never a law of nature. It came from fab economics and from demand staying below supply, and the demand side has changed.

The chart from The Black Swan: 1000 and 1 Days in the Life of a Thanksgiving Turkey, well-being rising for a thousand days then collapsing at day 1,001 The turkey’s data, from Nassim Taleb’s The Black Swan.

On August 19, 2026, the same argument played out on X. Chamath Palihapitiya posted the cost-of-computation chart, 18 orders of magnitude in 80 years, and said it sits upstream of everything our species has accomplished, as long as we keep it going. Elon Musk replied with a cartoon of a turkey presenting its own weight chart, captioned “I see no reason why excellent growth shouldn’t continue”, and three words: “Nothing is guaranteed”.

The exchange: Chamath's cost-of-computation chart next to the turkey-weight cartoon Musk replied with, captioned "Nothing is guaranteed"

We might be on day 1,001. The big cloud providers are spending hundreds of billions of dollars a year on new datacenters, and every one of those servers competes with your next CPU instance for the same wafers, power and rack space. North American datacenter vacancy has been around 1% for three years (CBRE), grid transformers have 128-week lead times, and Microsoft’s CFO has said “We have been short of power and space.” The capacity being built through 2028 is already spoken for, so I don’t expect this to be fixed in a quarter or two.

The Python tax

For compute-heavy loops, C++ is often one to two orders of magnitude faster than Python. Even I/O-heavy services, where the interpreter spends most of its time waiting, tend to get 2x to 5x faster. At scale, saving CPU time means running fewer servers. An engineer at Meta used Strobelight, the company’s fleet profiler, to find a one-character C++ fix (a vector copy that should have been a reference) worth an estimated 15,000 servers of capacity per year (see “The Biggest Ampersand”).

So when did you last profile your heaviest workloads? An afternoon with a profiler usually turns up a few Python loops that account for a surprising share of the bill, and those are natural candidates for a C++ port. Finding them is the easy part. Moving them to C++ meant building a bridge back to Python: hand-written pybind11 glue, a second build system, and type stubs that drift out of date. Often, paying for more CPU was cheaper.

C++26 reflection changes that, because a binding generator can now read a class and derive the glue from it. mirror_bridge turns C++ classes into a Python module with one command and no hand-written binding code, so the bridge is now almost free.

Try it

Here is a real run on a header mirror_bridge has never seen:

Animated terminal: cat a plain C++ header, run one mirror_bridge command, then import the module from Python and call it

The command, if you want to copy it:

mirror_bridge generate include/ --module vec3 --lang python --release-gil --stubs

generate include/ finds every class in the directory. --module vec3 names the module you import, and --lang python picks the target (Lua and JavaScript also work). --release-gil releases Python’s global interpreter lock while the C++ code runs, so a long call doesn’t block your other Python threads. --stubs writes a vec3.pyi with the signatures from the C++ class, which gives you autocomplete and type checking in your editor.

There is not much to maintain. The C++ stays an ordinary header without macros, annotations or wrapper classes, and when the class changes you run the same command again to regenerate the module and its stubs.

How much you save depends on what your hot path does. I have written about two larger examples: porting Open3D (25,262 hand-written binding lines down to 71) and benchmarking the generated bindings against V8, PyPy and LuaJIT.

The point

I am not saying you should rewrite your service in C++. The old rule of thumb, leave it in Python, the rewrite isn’t worth the bridge, depended on cheap CPUs and on glue that took weeks to write. Both of those changed. Reflection is standard C++26, and with stock GCC 16.1 (-std=c++26 -freflection) the bridge is one command. Move the hundred-odd lines that burn 80% of your cycles to C++, and leave everything else where your team is productive.


Written with LLM assistance. The opinions and the mistakes are mine.

2026-08-25