I'm glad to see the donations from big companies to open source maintainers are making measurable difference to the Rust experience. Telling these companies that their employees spend 5% less time waiting for compilation might motivate future investment in people like Nick and the others mentioned.
I've moved from Rust to Go for most things because in the era of agents, being able to iterate quickly on a project is a huge advantage and Rust is way, way slower than Go for compilation. There are times when Rust is more appropriate, but for the vast majority of things Go is perfectly fine.
OpenAI is a Rust Foundation platinum level donor. They gave straight cash which can go toward paying maintainers, which is better IMO. This actually made a lot of people angry because they saw it as a risk of OpenAI gaining influence or something. Can't make everyone happy.
OpenAI also hands out free subscriptions to open source maintainers: https://developers.openai.com/community/codex-for-oss They're time limited for now, but they've been pretty generous about who gets them. Anyone who can show Rust contributions would have been able to get one before.
This is one of the most important objectives in software today.
Rust is the best language to serialize agentic LLM output to. It's native, well constructed, low-defect due to design. It's also easy for humans to read and debug if necessary.
The biggest problem with Rust is the compile times. The cycle has to get faster. And we need to start thinking about making the artifact cache non-blocking so multiple agents can work simultaneously - that'll be a big task, but essential if we want to speed up work on one machine rather than spinning up clusters of agent sandboxes (the alternative, perhaps superior solution).
Doing stuff isn't free. For instance, Go compiles relatively quickly for a modern language, but the biggest reason for that is that it does less stuff than most compilers... less optimization, less checking, and some stuff built into the language to avoid some of the problems with having to read lots of headers just to compile a file and other ways of doing less stuff, but mostly the key is it does less stuff, in both the good and bad senses of that.
If you want something like Rust that offers guarantees and checks and cross-checks by the boatload, it adds up. Macros, monomorphization, implicit code generation with traits and all those other things add up too. And you can't always get O(n) or O(n log n) code to implement those checks. Maybe it can be sped up and maybe there's tricks here or there, but at the Pareto frontier, a language that has more checks will be slower to compile than one that has fewer.
And that's not a bad thing or a deficit in Rust, it's just the nature of the beast.
Rust wants badly to have its cake and eat it too. I get it, but it's not the sensibilities I have. The rustc compiler, if it really can achieve everything all at once, will be a beast like none other, making C++ compilers look simple by comparison.
I'd love something halfway. Go is maybe a bit radical in some regards, but also, with the news of the new SIMD package for Go, it has occured to me just how little I missed having things like, say, autovectorization.
(I know also that some people have tried halfway, but the big thing is figuring out how to keep a relatively simple type system that can still support a borrow checker. Even if there is some middleground, is it truly worth it? As nice as it sounds, I've been more skeptical. Go seems to exist in a very narrow space where its simplifications barely can be made to work.)
And since you bring up Go and contrast the compile time with Rust, it's been experienced at Google (and also Volvo and other places) that Rust and Go teams are as productive whereas C++ is less than half as productive: https://www.youtube.com/watch?t=27012&v=6mZRWFQRvmw&feature=...
So the focus on Rust compile time is misplaced. It's not a big deal in terms of overall productivity.
I don't think that follows - just that compile times aren't the only thing that matters. Maybe rust could be twice as productive as go if compile times went to zero for instance (or maybe not - just disputing that the evidence proves the claim here).
But it can't go to zero, that was the point -- at some level you have to pay for what you get. I'm not saying the compile times couldn't be better, but Rust can't be Go in terms of compile times if it wants to offer the guarantees and options it does. Meanwhile if Go wants to offer more options and stronger guarantees, it won't be able to preserve the short compile times everyone points to.
The Cranelift backend is extremely fast. Most of the slowness is all the crazy optimizations that LLVM does + other things (generating debug info etc).
For symbol heavy projects, linking is a surprising bottleneck.
Some ideas to speed up compilation by not evaluating items that are not used might pan out significantly for big crates in your dep tree (that behavior might never be stable because that would allow items with compile errors in a crate that would still let your application compile, which is against the Rust approach). The same work to do that would also allow overlapping of crate evaluation between different rustc instances called by cargo, as it would require partial evaluation of crates (to do name res only and gather the symbols needed from its deps).
Another thing is that stable rust doesn't treat macros as idempotent (because that wasn't a requirement from the start, there are crates that do dynamic IO to generate types), but if they are then incr comp can be faster by not evaluating them unnecessarily.
I know people are working on a bunch of different strategies to improve both first and incremental compile times, and I'm looking forward to the fruit of their labor.
Compile times are rarely the bottleneck for me, but that doesn't mean I won't welcome any improvements on that front.
Zero-cost abstractions aren't zero cost in compilation time. High-level abstractions translate to a lot of boilerplate that the compiler has to optimize out.
In unoptimized builds often the linker is the bottleneck. Rust/Cargo can parallelize most of the build, generating tons of code and debug info, but then the poor linker has to consume all of it at once. The object/exe formats were designed in ancient times, so they're hard to build incrementally or in parallel (some linkers are trying).
Is there any experimentation with new ABIs? I mean certain linker flags like --f-lto literally hijack the linker protocol to dump an AST into the backend.
At least for the fully static binary part of rust, there should be some optimizations there w.r.t. compilation. Sure you're not going to interface with shared libraries well but maybe a small experimental feature for fully owned projects? Idk.
I once heard (and don't know if it's still true) that the biggest compile-time sinks are macros and codegen.
Both are kind of outside the Rust compiler's influence. Macros can be almost arbitrarily complex: you pay for what you order. Codegen is LLVM, and that's a fixed choice. You can use Cranelift to get around it, but then you pay elsewhere.
Also, generics and monomorphization regularly come up in these discussion, while common wisdom seems to be that cost for the additional static analysis over other languages like C++ is no a major contributor.
Regardless, it's nice to see performance improvements in the compiler, even if you have to cooperate to benefit from them (e.g. by keeping your macros light and use less generics).
First of all, the fact that the article talks about "speeding up" the rust compiler doesn't automatically mean that the compiler is "slow"[0].
Now, is rustc slower than e.g. clang? by how much? why?
Those are different (and complicated) questions. It really depends on what you're compiling, but I'd say rustc can be 1-5x slower (maybe more at times?).
The reasons are many and varied, but in general rust compilation is slower because the compiler is doing way more things compared to C (monomorphization, complex trait resolution + type inference, borrow checker..)
[0]: Also I'd argue that "slow" without a concrete point of reference is a meaningless term in this context.
The go compiler is considered fast, yet Google spends effort speeding up the go compiler. Not to the same extent, but Google has a lot of go code, so it does make a difference.
Its similar to C++ compilers, there is no one reason. It is a language that tries to optimize a lot, its a big language, it does safety checks, it uses llvm which is a bit slow, its a language that makes use of generics which generate extra code etc etc.
This is not actually the main reason, most of the time.
Generics/monomorphization and how iterators work results in a lot of compiler bytecode that has to be churned through. More bytecode = longer compilation. It increases the size of the (debug) binaries, the debuginfo in general, causes performance issues with debug binaries in some situations unless you bump the optimization level, causes more IO, etc.
It's unlikely that many people are using Rust without using Option<T> or Result<T, U> a fair bit. Idiomatic Rust fundamentally uses a lot of generics. And a basic for loop expands into quite a lot of intermediate representation due to Iterator.
Other languages support generics, monomorphization, and iterators (e.g. Zig or D), but they're not as slow as Rust to compile, what's the reason for that?
It's likely easier to answer for specific languages
All three of those words can mean "just like Rust" but equally "Not at all like Rust" for different languages.
Two examples to contrast: In C++ the iterators are basically a pointer analog (in some cases they're just literally pointers) and that's a very difference "feature" but it's still definitely iterators. In Ginger Bill's Odin, the iterators are a function, possibly generic, which returns a pair, the next item and a boolean telling you whether the iterator was exhausted.
Trait solving isn't a bottleneck for Rust compilation unless you're doing some extremely cursed things like trying to implement Doom in the type system. At the end of the day it's still mostly that Rust just generates a ton of IR for LLVM to chew on (because of monomorphization), then LLVM takes a while to process all of it (because LLVM is designed primarily to produce high-performance code, often resulting in reasonable tradeoffs against compilation speed), and then the linker takes a while to connect it all together. With an alternative backend to LLVM (e.g. Cranelift) you could choose to design it with a greater emphasis on compilation speed (likely trading off generated code quality in the process). And with a more radical vertically-integrated model where Rust controls the linker you could do in-place linking and nearly completely skip the final step (at least for debug builds), but that's a big change compared to the classic C-style build model.
Incremental seems to keep winning the easy wins, so the interesting part is whether the remaining compile-time still lives in the same places as last year.
The EverInitializedPlaces example really stands out. Going from ~1.5M to ~90K apply_effects_in_block calls by changing the CFG traversal is a reminder that the biggest compiler optimizations often come from changing the algorithm, not optimizing the hot loop itself.
It also seems like the new Polonius/trait-solver work is pushing compiler performance toward a more interesting problem: doing expensive analysis only when it is actually needed.
4.57% mean wall-time reduction across 629 benchmarks in two months is pretty remarkable. Great progress.
> reminder that the biggest compiler optimizations often come from changing the algorithm, not optimizing the hot loop itself.
Isn't this true for most optimizations, not just in compilers? My usual goto process for optimizing is "Find stuff we're doing that we don't have to do, re-evaluate what data structures we use and then re-evaluate what algorithms we use" basically, with minor changes depending on the results. Served me well so far, and haven't (intentionally) written any compilers.
I've stopped using Tauri and gone with 100% egui. It's cross platform and excellent, and if you give it design constraints it will look beautiful.
Check out my 100% adobe clean room reimplementations:
https://github.com/storytold/filmcraft
https://github.com/storytold/photocraft
https://github.com/storytold/drawcraft (going to rename this vectorcraft)
The #1 thing for the Rust project to do is make Rust faster to compile.
Rust is the agentic AI language. It just needs to lean in and go faster.
#0 get rid of the orphan rule
Sometimes we really can have our cake and eat it too.
OpenAI also hands out free subscriptions to open source maintainers: https://developers.openai.com/community/codex-for-oss They're time limited for now, but they've been pretty generous about who gets them. Anyone who can show Rust contributions would have been able to get one before.
https://forge.rust-lang.org/policies/llm-usage.html
https://forge.rust-lang.org/policies/llm-usage.html#experime...
Rust is the best language to serialize agentic LLM output to. It's native, well constructed, low-defect due to design. It's also easy for humans to read and debug if necessary.
The biggest problem with Rust is the compile times. The cycle has to get faster. And we need to start thinking about making the artifact cache non-blocking so multiple agents can work simultaneously - that'll be a big task, but essential if we want to speed up work on one machine rather than spinning up clusters of agent sandboxes (the alternative, perhaps superior solution).
What is the performance killer?
If you want something like Rust that offers guarantees and checks and cross-checks by the boatload, it adds up. Macros, monomorphization, implicit code generation with traits and all those other things add up too. And you can't always get O(n) or O(n log n) code to implement those checks. Maybe it can be sped up and maybe there's tricks here or there, but at the Pareto frontier, a language that has more checks will be slower to compile than one that has fewer.
And that's not a bad thing or a deficit in Rust, it's just the nature of the beast.
I'd love something halfway. Go is maybe a bit radical in some regards, but also, with the news of the new SIMD package for Go, it has occured to me just how little I missed having things like, say, autovectorization.
(I know also that some people have tried halfway, but the big thing is figuring out how to keep a relatively simple type system that can still support a borrow checker. Even if there is some middleground, is it truly worth it? As nice as it sounds, I've been more skeptical. Go seems to exist in a very narrow space where its simplifications barely can be made to work.)
So the focus on Rust compile time is misplaced. It's not a big deal in terms of overall productivity.
Some ideas to speed up compilation by not evaluating items that are not used might pan out significantly for big crates in your dep tree (that behavior might never be stable because that would allow items with compile errors in a crate that would still let your application compile, which is against the Rust approach). The same work to do that would also allow overlapping of crate evaluation between different rustc instances called by cargo, as it would require partial evaluation of crates (to do name res only and gather the symbols needed from its deps).
Another thing is that stable rust doesn't treat macros as idempotent (because that wasn't a requirement from the start, there are crates that do dynamic IO to generate types), but if they are then incr comp can be faster by not evaluating them unnecessarily.
I know people are working on a bunch of different strategies to improve both first and incremental compile times, and I'm looking forward to the fruit of their labor.
Compile times are rarely the bottleneck for me, but that doesn't mean I won't welcome any improvements on that front.
In unoptimized builds often the linker is the bottleneck. Rust/Cargo can parallelize most of the build, generating tons of code and debug info, but then the poor linker has to consume all of it at once. The object/exe formats were designed in ancient times, so they're hard to build incrementally or in parallel (some linkers are trying).
At least for the fully static binary part of rust, there should be some optimizations there w.r.t. compilation. Sure you're not going to interface with shared libraries well but maybe a small experimental feature for fully owned projects? Idk.
Both are kind of outside the Rust compiler's influence. Macros can be almost arbitrarily complex: you pay for what you order. Codegen is LLVM, and that's a fixed choice. You can use Cranelift to get around it, but then you pay elsewhere.
Also, generics and monomorphization regularly come up in these discussion, while common wisdom seems to be that cost for the additional static analysis over other languages like C++ is no a major contributor.
Regardless, it's nice to see performance improvements in the compiler, even if you have to cooperate to benefit from them (e.g. by keeping your macros light and use less generics).
Now, is rustc slower than e.g. clang? by how much? why?
Those are different (and complicated) questions. It really depends on what you're compiling, but I'd say rustc can be 1-5x slower (maybe more at times?).
The reasons are many and varied, but in general rust compilation is slower because the compiler is doing way more things compared to C (monomorphization, complex trait resolution + type inference, borrow checker..)
[0]: Also I'd argue that "slow" without a concrete point of reference is a meaningless term in this context.
Well, if it weren't "slow" for some definition of "slow", nobody would bother speeding it up, would they?
Generics/monomorphization and how iterators work results in a lot of compiler bytecode that has to be churned through. More bytecode = longer compilation. It increases the size of the (debug) binaries, the debuginfo in general, causes performance issues with debug binaries in some situations unless you bump the optimization level, causes more IO, etc.
All three of those words can mean "just like Rust" but equally "Not at all like Rust" for different languages.
Two examples to contrast: In C++ the iterators are basically a pointer analog (in some cases they're just literally pointers) and that's a very difference "feature" but it's still definitely iterators. In Ginger Bill's Odin, the iterators are a function, possibly generic, which returns a pair, the next item and a boolean telling you whether the iterator was exhausted.
It also seems like the new Polonius/trait-solver work is pushing compiler performance toward a more interesting problem: doing expensive analysis only when it is actually needed.
4.57% mean wall-time reduction across 629 benchmarks in two months is pretty remarkable. Great progress.
Isn't this true for most optimizations, not just in compilers? My usual goto process for optimizing is "Find stuff we're doing that we don't have to do, re-evaluate what data structures we use and then re-evaluate what algorithms we use" basically, with minor changes depending on the results. Served me well so far, and haven't (intentionally) written any compilers.