I keep seeing software engineers spending copious amounts of time arguing about performance, using performance as a metric when comparing the relative merits of various code constructs, or using performance as an argument about doing something in a certain way vs. a different way. (Heck, even I sometimes make this mistake.) This post explains why this line of thinking is, more often than not, misguided.
- Modern computers are quite fast, and modern strongly typed, compiled programming languages are fairly efficient; so, in most scenarios, our software tends to run quite fast already. This gives us the luxury of usually not having to worry about performance.
- Desktop applications spend 99% of their time doing absolutely nothing but waiting for the user to press a key on the keyboard or make a click with the mouse. There is a famous saying about this: "All software waits at the same speed". So, in 99% of the time, it does not matter how fast your software is.
- When the user does eventually press a key or make a click, most operations tend to involve fairly small amounts of computation, so the vast majority of those operations complete fairly quickly.
- Any operation that completes within 100 milliseconds or less appears instantaneous to the user, so it does not matter whether that operation completes in 1 millisecond or in 99 milliseconds; there is no perceptible difference as far as the user can tell.
- Of the small percentage of operations that take longer than 100 milliseconds, the vast majority complete within a second or two at most; so, although they are not perceived as instantaneous, they are not prohibitive: they merely degrade user experience.
- User experience is only important in consumer markets, which are highly competitive, and a user may at any moment choose a competitor's product simply because it is less sluggish. In other markets, such as Business-to-Business (B2B), and especially DeepTech B2B, where products are chosen exclusively based on their unique capabilities, sluggishness is a non-issue. It is not usually a non-issue; it is not almost always a non-issue; is always a non-issue.
- Of the exceedingly small percentage of operations that take so long to complete as to be classified as genuinely slow instead of just sluggish, many are still acceptable by the user; for example, when a user editing a very large document decides to save that document, they are likely to find it perfectly acceptable if saving the document takes several seconds; it makes sense that a big thing takes longer time to complete.
- Therefore, only an infinitesimally small percentage of operations can be considered problematic in terms of performance, which means that a correspondingly small percentage of code is actually worthy of performance considerations.
- Fun fact: sometimes certain tweaks are advocated in the name of performance, which would take a certain length of time to implement, but the total number of clock cycles that these tweaks would save, through constant use of the software, by the entire customer base, throughout the lifetime of the software, added up, would amount to less than the development time of the tweaks.
- In light of the above, it becomes obvious that we need to be very highly selective about the situations in which performance should be given any consideration at all, because in the vast majority of situations, being concerned about performance is a complete waste of time.
- In order to be methodical, we have to first of all establish (i.e. decide, and document,) the Performance Requirement for each operation carried out by our software. In other words, "when the user clicks such and such, then so-and-so should happen within no more than this-many milliseconds" and so on. Then, we have to measure the actual performance of our product, and compare it against the performance requirements, in order to prove that a certain operation needs to be optimized. Then, and only then, are we justified in optimizing anything.
- So, no software engineer should ever be talking about performance unless they have actually used the CPU profiler on the almost completed product or feature, and have hard facts to talk about. This, by the way, implies that there can be no discussion about the performance of anything unless there is an almost completed product or feature at hand first. That is the light under which we should understand Donald Knuth's famous quote: "Premature optimization is the root of all evil". Any optimization before you have an almost completed product is premature.
- When it has been decided that the performance of a certain piece of code must be improved, you can almost always find some nice and formal algorithmic optimizations to make, (for example an index here, a cache there, etc.) which will make your software meet its performance requirements, instead of tweaking and hacking things to squeeze clock cycles here and there. Algorithmic optimizations tend to yield non-linear performance benefits, and tend to be self-contained, whereas squeezing clock cycles only ever yields linear performance benefits, and tends to have a detrimental effect on the readability, maintainability, and sometimes even the reliability, of anything it touches.
A little autobiographical story
Back in the mid-nineties I went for a few job interviews with gaming companies in Southern California, as a C-and-Assembly programmer. In one of them, I was interviewed by the lead developer. He told me that the company was using only one programming language, and that was 8086 Assembly. He had tried C, and he had decided it was awful. So, if I was hired, he expected me to set C aside, and code in nothing but Assembly.
Now, I knew 8086 assembly; the demos that I had circulated, which had landed me the job interviews, owed their impressiveness to my ability to do some fairly fancy stuff in 8086 assembly. However, I found Assembly very counter-productive, so I preferred to work as much as possible in C, and to only write in Assembly the individual, isolated small routines where performance really mattered.
I tried to argue with the guy. I tried to explain to him that there was no need to write the entire game in Assembly; certainly not things like the game menus, the scripted sequences, the cutscenes, etc. I argued that even in the main loop of the game, if some code contributing to 1% of the frame time became twice as slow, it would result in a 1% slowdown, but that might be a whole 50% of all the code involved in the main loop. He was adamant. He said that the entire game had to be written in Assembly language for maximum performance, and for minimum size, and C was just never going to be good enough.
Other than that, the interview went fine. When it ended, and as I was greeting him on my way out, we were still talking about stuff, including the relative merits of C versus Assembly. One of his last words to me were: "And what are these pointers in C anyway?" I would have explained to him if I had been hired, but they were not paying as much as I would earn doing a less cool job, so it never happened.
So, imagine: this guy is the lead developer in a published gaming house in Southern California. He is definitely a very smart guy. And yet he has failed to understand that pointers in C correspond exactly to address registers in Assembly; essentially, he has refused to understand C; and as a result, he is forfeiting some huge productivity improvement by sticking to Assembly, under the false assumption that every part of the game is equally critical performance-wise. Well, sorry, but it is not.