i've seen static analyzers catch many ub patterns, but guaranteeing zero ub needs whole‑program analysis that blows up compile time and still produces false positives that drown developers
Yes, it's pretty much impossible to prove anything about arbitrary programs.
However if you are willing to restrict what programs you allow, you can make guarantees possible.
Silly example: if you compile valid (safe) Rust programs to C, you know that the resulting code will not invalidate Rust's borrowing rules by construction; and in principle you could try to establish this guarantee just from the C code alone, never having seen the Rust original.
However, you still wouldn't be able to have an algorithm that tells you for any arbitrary C code whether it has these problems or not.
The concept of undefined behaviour specific to C/C++ has always seemed batshit insane to me, and I'm yet to read anything about it that has made it seem any less so.
It doesn't make sense to define what happens when you e.g. read from NULL because it's hardware specific - if you have virtual memory of some sort, you'd probably get a page mapping error. If you don't (embedded, WASM), you'd read back whatever value is at that address.
Does Rust define what I get when I dereference NULL in unsafe code? I doubt it, since it would require a NULL check before every pointer dereference.
The only insane thing about UB is that compiler writers took what everyone understand meant "the compiler emits what it emits and you get what you get" and turned it into "since it's undefined it means it can never happen so we can delete your null check".
> It doesn't make sense to define what happens when you e.g. read from NULL because it's hardware specific - if you have virtual memory of some sort, you'd probably get a page mapping error. If you don't (embedded, WASM), you'd read back whatever value is at that address.
That's an argument in favour of 'implementation defined behaviour'. Not 'undefined behaviour'.
I don't disagree with your general thesis, but I don't think it's right to say that defining behaviors like dereferencing a null would require a null check before every dereference. For example, Java defines the behavior of dereferencing null—it throws a `NullPointerException`. My understanding is that JVMs implement this by (a) representing Java `null` references as the zero pointer, (b) mapping the zero page with write permission disabled, so that accesses are guaranteed to segfault, and (c) trapping `SIGSEGV` to translate that back into a Java exception.
So, while it's accurate to say that inserting null checks before every dereference is one way that you could implement this to make it well-defined, that is not the only way. We have lots of clever tricks to solve problems more efficiently than may seem possible at first glance—Fil-C is a bit of a modern marvel in that regard!
Undefined behavior is just what happens when you violate a mandatory precondition.
if (x > 0) {…}. But what if you entered the body when x <= 0?
if (false) {…}. But what if you execute the body?
These are “impossible”. What happens when the impossible occurs is “undefined”.
When the older standards said signed integer overflow for addition is undefined what they are actually saying is that the real definition of + is:
int +(int x, int y) {
assert(in_range(actual_math_add(x, y), signed_int_min, signed_int_max));
return machine_add(x, y);
}
So of course what happens when you get signed overflow is undefined; you should hit that assert and your program should explode and die. You should “never” get to the next instruction.
But, in the interest of performance, “release mode” (which in this case is just any compilation) elides asserts since as a programmer you should not write code that asserts in much the same way that you should not write assert(false) in a normal code path that is supposed to run. Assertions are intended for “impossible” code paths and usually get compiled out in “release mode” though maybe your code is buggy and can actually hit them and then your program goes off the rails because it had a bug.
Put another way, if you did write assert(false) in a regular code path, would you find it unreasonable for the compiler to just delete the code after it? That is what is happening.
I disagree, though I wouldn't mind adding a mechanism to say that you want signed overflow to be well defined.
C23 already requires 2's-complement representation for signed integer types, but signed overflow still has undefined behavior. I think that mandating 2's-complement wraparound would be a mistake.
Some instances of undefined behavior can be detected at compile time. For example, if I write
int too_big = INT_MAX + 1;
a reasonably clever compiler can warn about it (and in fact both gcc and clang do so). If the result of INT_MAX + 1 were defined by the language to be INT_MIN, there would be no basis for such a warning.
If you evaluate n + 1 and it's possible for n to be equal to INT_MAX before the addition what do you want the result to be? Would quietly yielding INT_MIN really be useful?
Ideally, if I (accidentally) evaluate INT_MAX + 1, I'd like to be told that I've made a mistake. C doesn't have a good mechanism for doing so.
gcc has a non-standard option "-fsanitize=signed-integer-overflow" that can be used to catch signed overflow at runtime. If signed overflow yielded a well defined result, that option would be non-conforming.
> a reasonably clever compiler can warn about it (and in fact both gcc and clang do so). If the result of INT_MAX + 1 were defined by the language to be INT_MIN, there would be no basis for such a warning.
Compilers warn about perfectly well defined behaviour all the time. That's why these are warnings, not errors.
Why would this help? Unsigned integer overflow is defined behavior, that causes basically the same set of bugs in practice. In a way it is worse, because at least runtime UB checkers have a reason to complain about signed integer overflow, but they won't complain about unsigned integer overflow.
OS 2200 has 36-bit words. It is still a supported platform.
https://en.wikipedia.org/wiki/UNIVAC_1100/2200_series
This platform was the first SMP UNIX implementation:
"Any configuration supplied by Sperry, including multiprocessor ones, can run the UNIX system."
https://www.nokia.com/bell-labs/about/dennis-m-ritchie/other...
Yes, but that's perhaps an argument for 'implementation defined behaviour', not in favour of 'undefined behaviour'.
1790673092 | Reducing undefined behavior in the C language | https://lwn.net/SubscriberLink/1095811/efcdbcf080cfa4c6/ | https://news.ycombinator.com/item?id=49890290 | 0 comments
However if you are willing to restrict what programs you allow, you can make guarantees possible.
Silly example: if you compile valid (safe) Rust programs to C, you know that the resulting code will not invalidate Rust's borrowing rules by construction; and in principle you could try to establish this guarantee just from the C code alone, never having seen the Rust original.
However, you still wouldn't be able to have an algorithm that tells you for any arbitrary C code whether it has these problems or not.
Does Rust define what I get when I dereference NULL in unsafe code? I doubt it, since it would require a NULL check before every pointer dereference.
The only insane thing about UB is that compiler writers took what everyone understand meant "the compiler emits what it emits and you get what you get" and turned it into "since it's undefined it means it can never happen so we can delete your null check".
That's an argument in favour of 'implementation defined behaviour'. Not 'undefined behaviour'.
(Here's a link with a bit more detail: https://courses.cs.vt.edu/cs3214/spring2026/questions/catchi...)
So, while it's accurate to say that inserting null checks before every dereference is one way that you could implement this to make it well-defined, that is not the only way. We have lots of clever tricks to solve problems more efficiently than may seem possible at first glance—Fil-C is a bit of a modern marvel in that regard!
Rust's unsafe mode, for instance, has undefined behavior. A whole list of them, in fact.
if (x > 0) {…}. But what if you entered the body when x <= 0?
if (false) {…}. But what if you execute the body?
These are “impossible”. What happens when the impossible occurs is “undefined”.
When the older standards said signed integer overflow for addition is undefined what they are actually saying is that the real definition of + is:
So of course what happens when you get signed overflow is undefined; you should hit that assert and your program should explode and die. You should “never” get to the next instruction.But, in the interest of performance, “release mode” (which in this case is just any compilation) elides asserts since as a programmer you should not write code that asserts in much the same way that you should not write assert(false) in a normal code path that is supposed to run. Assertions are intended for “impossible” code paths and usually get compiled out in “release mode” though maybe your code is buggy and can actually hit them and then your program goes off the rails because it had a bug.
Put another way, if you did write assert(false) in a regular code path, would you find it unreasonable for the compiler to just delete the code after it? That is what is happening.
There wasn't room for safe programming practices, and direct manipulation of the hardware was a design requirement.
It assumes that you know what you are doing.
There are also ports to the Zilog Z80, an architecture with similar limitations (UZI, FUZIX).
C23 already requires 2's-complement representation for signed integer types, but signed overflow still has undefined behavior. I think that mandating 2's-complement wraparound would be a mistake.
Some instances of undefined behavior can be detected at compile time. For example, if I write
a reasonably clever compiler can warn about it (and in fact both gcc and clang do so). If the result of INT_MAX + 1 were defined by the language to be INT_MIN, there would be no basis for such a warning.If you evaluate n + 1 and it's possible for n to be equal to INT_MAX before the addition what do you want the result to be? Would quietly yielding INT_MIN really be useful?
Ideally, if I (accidentally) evaluate INT_MAX + 1, I'd like to be told that I've made a mistake. C doesn't have a good mechanism for doing so.
gcc has a non-standard option "-fsanitize=signed-integer-overflow" that can be used to catch signed overflow at runtime. If signed overflow yielded a well defined result, that option would be non-conforming.
Compilers warn about perfectly well defined behaviour all the time. That's why these are warnings, not errors.
However if you wan, you can already get that via a flag in pretty much any C compiler you care about.