Honestly, this just looks like one of those lingo-heavy-but-surface-level blog posts that used to make functional programming spaces so insufferable to everyone on the outside
You want to avoid branches in hot paths. If you branch inside the loop, lots of branches. If you branch outside the loop (into different specialized loops), few branches. Big fucking deal.
These things are so divorced from the reality of programming, even when they involve actual code instead of fancy lingo. Like in Scala, not a pure functional language, tutorials used to find the most convoluted higher-order functional way to do simple things.
Yeah, sounds human-written to me, too. Out of habit I checked with Pangram - which identifies some parts as AI-written. (I think that might be false positives but am not 100% sure.)
Erm, no? You write f(w: Walrus) -> Walrus and then let the caller handle Walrus|None and Iterable[Walrus] however they wish!
And if someone decides the codebase needs an abstraction over (and therefore specific functions to handle) Iterable[Walrus|None] then you check the weather and suggest they take a break and go for a stroll. (You check the weather to see if you should lend them your brolly.)
At least in JVM land, it's pretty easy to thwart that optimization. Particularly if the condition is on a mutable yet unchanged in the loop value.
For example:
var map = new HashMap<String, String>();
map.put("foo", "bar");
for (var i : items) {
if ("bar".equals(map.get("foo")) {
doStuff(i);
}
}
Even though `map` isn't mutated, it's hard enough for the JVM to detect and the underlying `get` functions are complex enough that it'll run the `get("foo")` every time, which can be quite expensive.
Can't speak for C# but in C/C++ the optimization can rarely be applied safely due to aliasing. If any part of the data you're working with involves a char* then C/C++ optimizers refrain from doing these kinds of optimizations because of how difficult it is to guarantee the absence of mutability.
Didn’t see it mentioned in the article but isn’t leading with if-statement called a “guard clause”.
I like that pattern but it’s just general best practice I thought.
They're one of those good practices that look like bad practice to everyone who just got a CS degree. Seems ex-students are unsettled by asymmetry or want to minimize the number of return statements.
I think it's sort of obvious that the limit to this general rule is when data dependencies between fors and ifs forbid you from pushing things further up/down.
"the loop runs without a branch, and is a candidate for vectorization".
That's it, that's the article. This matters a lot in huge-scale / scientific computing / HPF, where if you can express something as an operation on vectors on matrices, you win big (those ops parallelize well, can be run on GPUs, clusters, what have you).
I am continually impressed by the ability of LLMs to take trivial ideas and turn them into lengthy and obtuse blog posts with unnecessary analogies.
And also generate a shorter version.
Semantic compressor and decompressor
Honestly, this just looks like one of those lingo-heavy-but-surface-level blog posts that used to make functional programming spaces so insufferable to everyone on the outside
You want to avoid branches in hot paths. If you branch inside the loop, lots of branches. If you branch outside the loop (into different specialized loops), few branches. Big fucking deal.
These things are so divorced from the reality of programming, even when they involve actual code instead of fancy lingo. Like in Scala, not a pure functional language, tutorials used to find the most convoluted higher-order functional way to do simple things.
Yet another encroachment on traditionally human activity.
I am the
I know we’re not supposed to comment just for that, but this might be my single favorite joke comment I’ve ever read here. Good job.
Except that TFA is a bog standard example of traditional human activity and the GP's comment is nonsensical trolling.
I'm continually impressed by the ability of humans to spam low-effort whining about suspected LLM writing in almost every HN thread.
It's an LLM-generated article.
How do you know? Doesn't read particularly LLM written to me, and even if it is, it's quite well written.
The author has been blogging about this kind of stuff for over 20 years. I'd be surprised if they suddenly let bots autonomously spam their blog.
Yeah, sounds human-written to me, too. Out of habit I checked with Pangram - which identifies some parts as AI-written. (I think that might be false positives but am not 100% sure.)
I've done this for years. Not every time of course but where it makes the code easier to understand and maintain.
Speed was almost never the reason.
I take it you never rewrote a Matlab for loop as a vector/matrix op for insane speedups then :)
Erm, no? You write f(w: Walrus) -> Walrus and then let the caller handle Walrus|None and Iterable[Walrus] however they wish!
And if someone decides the codebase needs an abstraction over (and therefore specific functions to handle) Iterable[Walrus|None] then you check the weather and suggest they take a break and go for a stroll. (You check the weather to see if you should lend them your brolly.)
What am I missing?
What is missing here is any benchmarks backing up this argument for code structure.
Of note, as of C#9 (and maybe prior), the dotnet runtime does this automatically whenever it is deemed safe. https://devblogs.microsoft.com/dotnet/performance-improvemen...
The same technique is applied as an optimization, when deemed safe, in all current gen c compilers (gcc, llvm, etc).
I'm very confused why neither measurements nor references to when this is done automatically in most modern languages is included in the article.
At least in JVM land, it's pretty easy to thwart that optimization. Particularly if the condition is on a mutable yet unchanged in the loop value.
For example:
Even though `map` isn't mutated, it's hard enough for the JVM to detect and the underlying `get` functions are complex enough that it'll run the `get("foo")` every time, which can be quite expensive.Can't speak for C# but in C/C++ the optimization can rarely be applied safely due to aliasing. If any part of the data you're working with involves a char* then C/C++ optimizers refrain from doing these kinds of optimizations because of how difficult it is to guarantee the absence of mutability.
Just the branch predictor gains are probably worth it.
Didn’t see it mentioned in the article but isn’t leading with if-statement called a “guard clause”. I like that pattern but it’s just general best practice I thought.
Guard clauses are things that return early for trivial or problematic cases. Like https://en.wikipedia.org/wiki/Guard_(computer_science)#Flatt...
They're one of those good practices that look like bad practice to everyone who just got a CS degree. Seems ex-students are unsettled by asymmetry or want to minimize the number of return statements.
Particularly for high performance code when branch prediction is taken to account
Swift explicitly has a guard statement for this. Rust's let .. else { ... } is also very similar.
https://docs.swift.org/latest/documentation/the-swift-progra...
I have always phrased this as "Never do one of something".
I like it, but to do fizzbuzz in this way, you'd have to separate what's inside the loop into a reused function.
I think it's sort of obvious that the limit to this general rule is when data dependencies between fors and ifs forbid you from pushing things further up/down.
Or just use lazy list operations with a single if test at the end:
https://github.com/taolson/Admiran/blob/main/examples/fizzBu...
/s
TL;DR in one sentence:
"the loop runs without a branch, and is a candidate for vectorization".
That's it, that's the article. This matters a lot in huge-scale / scientific computing / HPF, where if you can express something as an operation on vectors on matrices, you win big (those ops parallelize well, can be run on GPUs, clusters, what have you).