Bend the GPU
I made bend-bench which I introduce in this tweet thread. Taelin has been upfront that Bend's compiler has blind spots, and bend-bench is meant to help find them.
I find the ability to write the same code and have it run on CPU and GPU to be quite magical. Others do it for regular data parallelism (JAX, Futhark, Kokkos, OpenMP offload, Taichi) but Bend does it for plain recursive, pattern-matching programs over trees, with no kernels. There is one downside to the beauty of running the same code on CPU and GPU: it leaves lots of performance on the table. The GPU is just fundamentally different and opens up a lot of avenues for optimization that don't exist on the CPU. How can we close this gap in Bend going forward, in the most elegant way, that stays true to its ethos?
What Bend does now
Taelin pitches Bend as "as secure as Lean, as fast as Rust on CPU, as fast as CUDA on GPU", and his optimization target is competent C running the same algorithm, within about 2x. Bend does have some GPU-specific code currently, but it only modifies the scheduler. There is no specific GPU optimization in terms of data layout, tiling, warp cooperation, memory coalescing, or shared memory. This can be seen in the GPU optimized programs that beat it: HotSpot uses shared-memory tiling, the CUDA lexer scans byte arrays where Bend walks lists, and Mandelbrot sums its histogram in shared memory and stops each pixel's loop as soon as it escapes, where Bend's port runs every iteration branch-free.
One fix so far
Using bend-bench to compare results between Bend 2.0.3 vs 2.0.26 flagged GPU ray tracing being 38% slower. This led to a more principled PR to decide inlining by whether a function loops, instead of only its line count. My suspicion is that similar elegant improvements are going to be hard to find, if we limit ourselves to changes that don't grow the codebase. Beyond fixes like this one, most improvements to GPU compilation require adding cases that don't affect the CPU, and therefore must be net new. How could this be done in the most elegant way going forward?
Path 1: New IR
What if instead of running the exact same code, we open the door to more GPU-specific optimization. My first thought was to target a lower-level intermediate representation than C, which then gets compiled to GPU code anyway. This could be LLVM IR, MLIR, or even tinygrad's in-flight UOp. LLVM IR and MLIR are general compiler IRs with GPU backends, and tinygrad's UOps are built for GPU tensor work. Using them would allow running on any GPU they support as a target.
Why not this path? Taelin has already weighed in: compiling to C is surprisingly effective, and C over LLVM was the realistic way to target C, Metal and CUDA. And UOps are built for tensor and array computation: loops over buffers, element-wise operations and reductions. Bend's programs are recursive calls over trees, run by a task scheduler, and don't map onto that model naturally. MLIR has the same mismatch: its GPU optimizations live in dialects for loops and tensors, and a Bend dialect plus its lowerings would be a large C++ dependency next to a compiler that's one TypeScript file. LLVM IR is the closest fit, and it can even carry facts C loses, like which functions don't touch memory. But it doesn't reach Metal, Bend's first GPU target, so Bend would keep C for Apple GPUs anyway. And the hard parts, scheduling recursive tasks, memory placement and synchronization, sit above any of these IRs in Bend's runtime.
Path 2: Lean into CUDA
CUDA already has lots of code optimized for it, so why not just re-use that? Bend could use its strengths to find and verify parts of its own code that are equivalent to existing CUDA primitives, then swap them in during the compilation step as needed. The CUDA primitives we benchmarked against were up to 280x more efficient per unit of work (CUB summation), and at least 89x on sorting.
Path 3: Skip the IR
Where Path 1 would use tinygrad's IR, this copies its approach: generate GPU code and drive the hardware directly, bypassing CUDA. That means taking on all the work of a compiler like tinygrad, but having full control of how your specific workload needs to be structured. A much bigger task, that doesn't have a clear payoff, while creeping scope that Taelin would probably hate.
Path 4: C is the IR
Opus keeps telling me to keep it simple, and that C is effectively an IR already. This could be built on by adding small pieces of code for each new target, like AMD's HIP, SPIR-V for Vulkan, or WebGPU. Taelin wants WebGPU ASAP and thinks AMD could be added in an evening because the code is that well structured.
The catch is what C forgets. Taelin points out that GCC can't use purity and linearity to optimize, and those are exactly the facts Bend knows about every program. Once they're lowered to C, nothing downstream can use them. So C works as the IR only if Bend spends that knowledge itself before emitting C, and every optimization like that is compiler code that has to fit under the cap.
What Taelin wants
Taelin currently has a cap on comp.ts at 64000 ttok, which he'll raise case by case for the compiler while the trusted kernel stays fixed. The code is basically at the cap already, at 62,584. Why does he do that? He warns that Bend only has a small team, and is worried about the additional maintenance work. But at the same time he's super happy leaning on AI for development, so I'm not sure why that lesson doesn't transfer to maintenance. He's also worried about AI being unable to grok a larger codebase: the cap keeps the whole kernel and compiler inside a model's context, with room left to work. That seems a bit overblown to me. Coding agents don't read a codebase whole. They search it and load the parts a change touches, so what has to fit in context is one change's working set, not the repo. And comp.ts at 62,584 ttok is already a small fraction of the long-context windows current models offer. The problem shrinks further as models get trained on Bend itself, which Taelin has asked the labs to do, and which he suspects Anthropic already did with his earlier work.
His sharper argument is about refactoring. A codebase is easy to fix when it fits in half a model's context, because you can fix it all in a one-shot full refactor: under 100k tokens is the sweet spot, and above that you and your AI lose control. Bigger projects are fine as long as they're broken into isolated modules, but he treats Bend's core as one irreducibly atomic module. The working-set argument above doesn't answer that, since a whole-codebase refactor does need the whole codebase in context. So the question is whether GPU-specific compilation has to live inside that atomic core. The trusted kernel does. But GPU lowering, or swapping in CUDA primitives, looks like exactly the kind of isolated component his own rule says to split out, like the extensions he's already floated.
Of the four paths, a new IR runs against his choice of C, detached extensions suit swapping in CUDA primitives, C with a small header per target suits the portability he's asking for, and skipping the IR would blow through his size cap.