Skip to main content

Niall's naive list of unsolved problems in software reliability

·3 mins

This post is a part of Project Cinnte (KIN-TEH), an effort to look at some foundational questions in software reliability.

I participate in the Resilience in Software Foundation, the Slack for which contains a large number of people interested in questions that also obsess me - software, systems thinking, safety science, and so on.

A recent thread about metrics provoked this response from me, which one of the thread participants thought was helpful, so I reproduce it here. This is basically my Hilbert problem list, but I think it’s more appropriate to call it “Niall’s naive unsolved problems of the field” list:

  1. We have a lot of qualitative models and not many (perhaps no?) quantitative models. Is this necessarily true?
  2. We have a lot of impossibility theorems and tradeoff models, but not many constructive models. This is in the sense of “you can achieve X if you build it in way Y”) models. Is this a necessary outcome? What’s stopping us getting more positive constructive models?
  3. We have no real theory of software failure, other than Rice’s theorem. Can we do better?

We have well-founded intuitions about software failures, and a ton of experience about how they happen, but we can neither show they are infinite nor finite, not characterise their behaviour. At the other end of the spectrum, evidence-based software engineering papers are fairly hopeless when it comes to long-lasting fundamental results, and we basically have Brook’s Law and Conway’s Law.

I accept that socio-technical observations about wider organisational behaviour are of course not going to have six-sigma output, will suffer from reproducibility problems, et cetera. I know all that. I should be clear that in many cases we have valuable observations and occasionally even working models that can inform behaviour (as above, e.g. “do not add more humans to an already late software project”, “you should think about how your org structure is going to get reflected in your software, because it will” and so on). But the occasions, if I can put it that way, are few and far between, and I see, particularly in the LFI end of things, a lot of epistemic objections entirely at odds with management frameworks, both of them unfortunately rooted in the reality of their lived experience.

The current state of play is that a practitioner working in software - a sharp-end person speaking to a blunt-end person in the learning from incidents/software reliability context cannot explain:

  • On what basis we should believe incidents will happen
  • Predict what might happen with an individual future incident, or indeed any aggregate metric of all future incidents
  • Articulate suitable replacement metrics for understanding behaviour or improvement in the space
  • Understand or explain what is likely to happen to organisations turning to AI-inflected software development

Am I wrong about any of the above? Where am I missing nuances?

When I emphasise “necessary” above, what I mean is that “I don’t know if all of the above is a necessary result of the space, i.e. we have a bunch of limits theorems mostly because it’s impossible to do anything better, or is it merely an accidental result of us not looking in the right direction with the right tools”.

[Continued…]

Software Reliability: what we believe

·5 mins

Software Reliability: what we believe #

This is part one of Project Cinnte (KIN-TEH), an effort to look at some foundational questions in software reliability.

Introduction #

Being professionally interested in questions of software reliability, we have often wondered whether or not we can find good models to help explain, characterise, or predict failure. The industry doesn’t have much in the way of such models today. What we have is practitioner’s intuition, and it’s often the best thing we do have, but it’s hard to know if it’s the best thing we could have.

In the act of reflecting on this question, we realised that there isn’t a good source that documents precisely what that intuition actually says. The rest of this document is an attempt to do that.

The standard doctrine #

Here is our take on the “standard doctrine” on bugs, outages, and how software development typically happens - happened? - in the pre-AI SWE era.

We label the following propositions for ease of reference, together with their commentary.

P0: Bugs happen as a function of the sustained intellectual effort underpinning writing software. The longer the effort or the more intense the effort required, the larger the chance of incorrect software being written.

Software engineers believe this because they model the act of writing software as having a problem to deal with, thinking of a way to solve that problem, and writing down the solution. The tougher the problem, the more chance you get it wrong; the larger the problem, the more chance you get it wrong.

P1: Other things being equal, the bigger or more complex the software, the more bugs there are.

P2: Other things being equal, the newer the software, the more bugs there are.

These look like simple corollaries of the above, though there is more complexity there when you tease it apart. When we say “bigger” above, that follows clearly from P0, but “more complex” also brings in scenarios where the problem might not require a large mass of software to solve, but might be algorithmically complex, have a very large number of business domain exceptions, have uncommon constraints (e.g. latency), or similar. P2 in particular introduces a new idea: newly made software has not encountered reality yet, and will have more bugs than software which has encountered reality.

P3: Bugs can occur for a variety of reasons, and on a number of different levels. The category of “bug” itself is quite wide, and includes at least typos, inaccurately writing what you want, accurately writing what you want but it turns out it’s not the right thing to want, incorrectly capturing your own requirements, incorrectly understanding other people’s requirements, working with third-party software or services with different models of the problem domain, compiler errors, microcode errors, firmware errors, or cosmic rays.

The current authors have had all of these happen to them at one stage or another. There is probably some kind of long-tail distribution of the frequency of occurrence, but the structure doesn’t matter for most of these observations.

[Continued…]

Appearing productive at work

·1 min

A rather forceful description of the current experience of using GenAI at work. Prisoners are not taken.

There’s a lot that resonates with my narrow area of interest: the continuing necessity for human-in-the-loop (despite much pressure for that to be eliminated); senior-versus-junior productivity dichotomies; and the idea that “accomplishing the work itself used to teach the judgment” now going away means that there is no clear way to acquire judgement when you don’t have the feedback loop to acquire it. Which, I suppose, increases the pressure to remove the humans and use LLM judgement instead.

It’s hard to reconcile in the one model of the world both this piece and the like of the pieces that I scroll past every day, which are insanely enthused about the sheer possibilities of the tools and the huge rate at which they’re improving. Erdős problems, passable literary works, code (of course). And yet I don’t feel that either of those kinds of experience are wrong, or lies.

Reclaim the SREets

·1 min

Came across this site the other day. Though the Changelog seems to suggest it was last updated in May 2016, which makes it a bit more likely to be slop, I think the overall message is not misplaced.

Career counselling

·1 min

A short while ago, an ex-colleague from an old team asked for some career advice. I did what I could (though I didn’t think it was very useful). He was kind enough to send me a card and a (very large) collection of chocolates in return. It was a great kindness that reminded me of a great team that was always under pressure.

As the tech industry is collectively doing a lot to erase every aspect of being human from working in it, and especially any question of the vulnerability that naturally attaches to it, it does some good to remind ourselves that there is virtue in being human.

A nice thank-you card