Guest: Sam Newman, bestselling author of Building Microservices and of the forthcoming Building Resilient Distributed Systems
Watch on YouTube
Listen on Spotify
Listen on Apple Podcasts
Read shownotes & transcript below
In this episode of Asynchronous and Unreliable, I talk to Sam Newman, author of *Building Microservices*, about his new O’Reilly book, *Building Resilient Distributed Systems*.
Distributed systems are hard — and, as they scale, they become hard in distinctly inhuman ways. Rare events stop being rare, components inevitably fail, resources run out, and the million-to-one chance starts happening nine times out of ten.
Sam’s new book is intended as a survival guide for developers who find themselves building distributed systems, deliberately or otherwise. We talk about why microservices don’t automatically make systems more resilient; why resilience is a socio-technical property involving people, organisations and incentives as well as software; and why expecting failure can actually be the foundation of extraordinarily robust systems.
We also discuss what software can learn from aviation, railways and the internet; incident response and game days; Google’s wonderfully pragmatic early hardware; the dangers of adding complexity in pursuit of resilience; and why sometimes accepting that your system will go down is exactly the right engineering decision.
Plus rabbits, Godzilla, smart litter boxes, Terry Pratchett, and why Sam thinks we need to have a serious conversation about what the word “asynchronous” actually means.
Anne (00:00)
Hello and welcome to episode twenty six of Asynchronous and Unreliable, a weekly podcast where we discuss the latest ideas and concepts in tech. I'm your host, Anne Currie, Co-author of Building Green Software, the Cloud Native Attitude, and author of the science fiction Panopsicon series. And today I have the great pleasure to welcome Sam Newman, author of the best-selling Building Microservices from O'Reilly. And today he's here to tell us all about his new book, Building Resilient Distributed Systems, also from O'Reilly, which I've read a pre-preview copy of and it's great. So, Sam, welcome to the show.
Sam Newman (00:36)
Thank you so much for having me.
Anne (00:40)
so it's an interesting one because this particular episode is going live while both one of your previous books and my current O'Reilly book, are in the Humble Bundle for software architecture. that's a really cheap way of buying a really good set of books. and it will be available for about a week after this podcast goes live. So viewers might want to go and grab a copy of that. I like the fact that we're both in it. building microservices is part of software architecture, and building green software is part of software architecture. One of the other books in it is a book about distributed systems but it's quite a different book to your book. So I'm very interested to hear about your book, who it's aimed at, why you wrote it, and I have some thoughts about why you wrote it. But tell me all about it.
Sam Newman (01:42)
I mean, we should do a very brief digression is that we're originally when we talked about me being on this podcast, I was going to take you to task over using the word asynchronous because I don't think anyone has a good definition of what that means. But that's a separate conversation and maybe a separate episode., I mean, I wrote the book primarily because I feel that the topic of distributed systems is very daunting. And, partly I think it's rightfully daunting because there are some complex aspects to it.
But as a result, a lot of the discourse around distributed systems become quite intimidating. And so I wanted to take a look at one specific aspect of distributed systems. And that was really looking at it through the lens of resiliency. Cause there were two kind of in my life over the last 20 years, there were two quite intimidating topics I've been trying to deal with: distributed systems on the one hand and resilience engineering on the other. And I found that those topics can be quite hard to get into because you often hit a wall of complexity and inside discussions and acronyms and everything else. And that was happening at the same time that more and more people are building distributed systems, whether it's a good idea or not.
Despite me spending a lot of time in building microservices, second edition and my subsequent books and subsequent work saying, you probably shouldn't do microservices. Some people still ignore me. And as a result, people are building not just distributed systems, but increasingly distributed systems without really understanding the trouble they're going to get themselves into.
And so I wanted to write a book that would almost be like a survival guide, like a rubber ring that you throw someone into, like a developer has been thrown into the waters of building resilient systems, like what would be the thing you'd want to give them to help them get to safety. And so it was a book that kind of was very much from almost like a persona point of view was I want to meet the junior developer who finds himself on a distributed system, but take them on a journey to more and more complex topics.
So like the second chapter is about the server, the third chapter is about timeouts, which is quite a basic concept, but you go more complicated through that chapter. But like you have to have that stuff sorted before you move on to the next piece. And the hope is that for experts, they can jump into the topics that they're most interested in. And they can get real deep on the timeout stuff. I mean, I spent 7,000 words talking about timeouts. I enjoyed myself writing that chapter.
But if you're maybe new in your software delivery career, you can just kind of go at your own pace. And so that was kind of the idea. And so it was about trying to make distributed systems seem more approachable, not in terms of meaning you should do them, but in terms of explaining why they're difficult. And then the other angle was coming at this one, the resilience engineering space. And that's been brewing in my head for a much longer time because it was my first chat with John Allspaw at Velocity when O'Reilly was doing that conference years and years and years ago. And I think he mentioned a couple of papers and I started looking into the world of Resilience Engineering.
And this is kind of the body of work that comes out of safety critical systems. And these ideas are often explained in quite academic contexts from much more formal domains than the ones that most software developers work in. And so I found that space quite unapproachable.
And so it was after kind of multiple years of hitting my head against that topic and talking to people who explained these things well, that I could try and unpack some of those concepts and make them seem approachable. So I'm trying to do two things really, which is introduce you to resilience engineering, introduce you to distributed systems and to see both as kind of fundamentally socio-technical problems.
Anne (05:30)
When I read your building resilient systems, I thought it's an interesting step on from building microservices. Because oddly enough, I started the other way around. I started with building distributed systems because I'm old, I'm all I'm not that much older, but I'm a I'm a fair chunk older than you, which meant that when I started writing systems, the hardware was terrible. And so if we wanted to write any kind of scalable, high performance system, you had to use something which was effectively the same as a distributed system, but on a smaller scale on a couple of machines. because it's much more efficient.
But it's incredibly hard. and I know when I first heard about microservices, I thought of them as an operational tool. You know, well, microservices, you can scale up, you can scale out, you can. But it was it was only a few years later that I realized that actually microservices are a development tool. They're a developer tool, they're a way of making CI C D & fast flow work by dividing up work so that teams can work better in parallel.
They're really about making developers' lives easier. And to start with, they don't really hit the problems of high scale DS systems. But I'm guessing that eventually they do, which is the kind of where your second book comes in.
Sam Newman (07:13)
Well, I got into microservices purely because of working with people doing kind of last mile optimization. So, going into organizations and they'd written some code, now it's getting to production. And so I do things like test automation, build pipeline modeling, know, path production type stuff, all theory of constraints, good old lean manufacturing stuff, right? And increasingly I found it was the architecture that was my main impediment.
So my first talk on this topic was in 2011, I think. It was called designing for rapid release. I was literally looking at architectural patterns about how you get software out of the door more easily. And that was the lens through which I kind of came at microservices.
But ultimately semantic diffusion being a thing for many people, they've just become a replacement for the name service oriented architecture. Although I think when done right, they're a bit, they're kind of more opinionated than general SOA.
I also think there's a lot of, you know, I come back to a quote that James Lewis said to me once years ago, so James Lewis, if people that don't know was the person who kind of helped popularize microservices. Initially, he wrote the first paper on microservices with Martin Fowler, it actually on his website.
We spoke a lot about the topic at the time when he was working through these ideas. And I was working much more on the infrastructure space and he was working on the architecture space. And he said to me once that microservices buy you options. And that the idea is that you don't get these things by using the architecture. You get the option to buy these things.
And so I do see a lot of people that assume that if they use a Microservice architecture, that they'll have a more autonomous team or that their system will be scalable or that their system will be more resilient. No, absolutely not. You have to do work to get them, but you get the option to do that work. And actually many cases I work with teams that have suffered significant degradation in resiliency by going to a Microservice architectures because they don't understand things like network calls can fail.
So actually, I find that for a lot of teams, their resiliency can take a step back and not a step forward by using these architectures. And so for me, i was trying to make some of these ideas approachable. Because I think when we kind of talk about the endless list of things I cover in this book, I could have done the same thing about security, for example. Again, it just feels too daunting. And it's like, you got to try and meet people where they are a little bit. People are building these architectures, whether they should or not.
So what can I do to try and help defang that a little bit?
Anne (10:06)
Distributive systems are really hard. And if you get to it through microservices, it's quite easy to just get slowly kind of boiling a frog. It gets harder and harder and harder. It started quite straightforwardly. And then the irony is the more successful you become, the more scale, the more users you've got, then suddenly it becomes much harder.
At an almost exponential rate.
Sam Newman (10:38)
I think it's partly that the problems of distributed systems are largely about the problems of failure. But that when you don't have much load and you've got a small number of calls, failure is less likely to occur.
You're less likely to run out of resources. You're less likely to see partitions in the live.
And so you're not as exposed to those things. So it's only really once you're already to the point where you're successful, where you're basically saturating resources or where you're starting, you've got a sheer enough volume of calls, you're starting to see component failure more regularly, you start to understand where you've got yourself.
And that's kind of why I try to distill down why are digital systems complicated. Like, because I think I remember the eight fallacies of distributed computing but I can never remember the things.
As I've got older, I don't know if it's, you know, myself diagnosed ADHD brain or whatever else it might be, or just, you know, cognitive decline caused by age. I couldn't remember all eight.
And so I kind of worked out that actually, if you explain to people, there are three reasons why distributed systems are difficult. One, you can't make information move between two points instantly. Two, sometimes the thing you want to talk to isn't there. And three, resources run out.
That fundamentally underpins every single challenge that we experience at a technical level with distributed systems. It's just that as you do more of it, these things are much more likely to bite you on your bum. I mean, technically speaking, if you build a single, if you build a, say a monolithic web application, you have a distributed system, right?
The clients are looking at your website in a browser. So that's a program running on your device, looking at a server on the backend. The database is probably on a third machine, let alone the other pieces in the middle, like, know, CDNs, any in-house proxies or WAFs you might have and things. It's just, it's less likely going to expose you to the challenges just because you've not gone too far, right?
So I think it was like about trying to recognize, help people recognize that's kind of true. I think if you go microservices, you almost accept this premise that more is better, which is incorrect by the way.
And so, you're storing this pain up. so it's like, okay, well, if you're doing this stuff anyway, I don't want you to do it anyway. Let's try and hand hold you a little bit, but also do so in a way that tries to bridge a little bit of that academic resilient engineering world in and also try and make concepts like socio-technical systems a little bit more understandable.
Anne (13:33)
Well talk me through socio technical systems. What do you mean by that and how does that work with the with the problems of distributed systems?
Sam Newman (13:43)
I mean, think you'd argue that any digital system is by definition a socio-technical system. So I'll start by that. we talked about the term is so socio and technical, socio-techni. It's basically people working together with technology. So any system where that occurs is described as a socio-technical system. It's based on some theory done by the Tavistock Institute during World War II, believe it or not, and looked into coal mining and stuff like that. But it's basically a study, the space of socio-technical systems is a study of how people and technology interact.
When we think of technology, you know, that would be the hardware, that would be the software, that would be the infrastructure. But more broadly, you throw in the other things, it's like your process, it's your organizational culture, it's the goals of the organization as well. And the reason this becomes really important when you're looking at this real resiliency mindset is that you've got to look at the components of the system. Who are the components of the system that can help this system be resilient or not?
Well, it's not just the software. It's not just the computers. It's not just the network equipment. It's also the people that operate the system. It's the goals those people are set, right? How that organizational system reacts to things like failure, how it reacts to failure. How well can it do instant analysis? Can it do a surprise? Well, and so you can absolutely look at resiliencythrough the lens of just the tech, the technology part of a socio-technical system, if you want to. But to do so means that you're looking probably at, it's a really, I think, inappropriate lens if you really care about resiliency. Because so many aspects of resiliency are actually about the operators and the users and the stakeholders in the system. And so I think you have to kind of understand the system as a whole.
And so I like to share this like hexagonal model that I came across through some research around sort of how to break down that socio and the technical side of things. And so you can kind of think about like, not on the technology side of it, you've got the hardware, you've kind of got the infrastructural piece, you've got the software, you've got the processes involved with dealing with the human side of it. You've got the culture, the organization lives in, you've got the goals that organization sets for people.
So you will quite often see failures or disasters caused by goal setting that encourages maintenance activity to not be dealt with. One of the things I look at, for example, is a train derailment that occurred. one of the main underpinning causes there was that the train company was trying to focus on modernization. That was an overt goal.
And so they were deprioritizing, there was less incentive for anybody to focus on maintenance activity. So less maintenance was being done. And people were being penalized if they raised maintenance issues, because it was seen that they weren't doing their job. So these issues weren't getting raised. And so like you all start unpicking all that and you said, the root cause is that they should have fixed a bit of track. Well, you can't look at it like that, right? It's way more stuff that goes on. So if you care about fixing those things or like improving the resiliency of your own organisation, you kind of have to take all those bits together.
Anne (17:07)
transport is always an amazing example of where they had a lot of failures and then learned from them. So, you know, crew resource management, the amazingly small number of aircraft accidents there are now compared to how many there used to be, because they deal with reporting error, they unearth issues much better than they used to.
But we're very good at learning from failure, we're not always so good at learning from success. And I always think that there's a really great example of success in the tech industry that we don't talk anywhere near enough about, which is the internet, an amazing example of socio technical success because everybody, literally all the way down to the people sitting in front of a TV at home, are part of the system.
That keeps it alive because it runs on a massively variable resource. We think that, you know, solar and wind, we think of those as variably available resources. But bandwidth has been a variably available resource forever. And it's totally unpredictable, and yet we manage.
Sam Newman (18:25)
But it's got to be the ancestry of the internet was based on the goals of developer and interconnected computer network that could withstand arbitrary loss of nodes in a military setting. I mean, it's sort of unsurprising in a way that it's worked and operated under those variable operating conditions when that was very much at its core. you start looking at these patterns that these application developers put into place to kind of deal with robustness issues.
So a good example is like a token bucket for like doing rate limiting. There's a client side, you can limit how many calls you make. I talk about the leaky button pattern in the book. And like, that's actually the TCP specification, right? Because it's there to do flow limiting and stuff like that. these things are built all the way kind of through the stack. And that variable condition thing is a really important part of that.
There's this fundamental shift that goes from, say, traditional safety management to resilience engineering. Traditional safety management was all about saying, let's make sure that nothing bad happens. Right? Nothing bad's going to happen, and we can make sure nothing bad happens. The problem, of course, is if something bad does happen, you're not prepared for it. So resilience engineering is about making sure that as many things go right as possible. And I share quote in the book from Alec Kohlnagel, who's kind of quite well known in space. And he talks about the idea that a resilient system is one that should be able to vary its operations when things go wrong to keep working.
So we do also have to have this idea of what do we give up on? What do we accept can fail? Like in the internet, for example, when you're watching streaming at home, if you start suffering from less bandwidth, the streaming quality starts to degrade. We have graceful degradation happening at quite low levels in the protocols.
And so taking those ideas into our own software, it's like, okay, well, I need to make sure as many things go right as possible, even when something wrong happens. So what are the things that are most important and what things can I afford to give up on? And what does good look like and what does good enough look like? And I think it lifts the debate a little bit to a more sensible level.
Anne (20:41)
absolutely.
It's the absolute foundation of resilience. I mean you could argue that the internet is, as you say, the internet is the most successful machine that humankind has ever developed and it's incredibly resilient. and as you say, it was built on a culture of not expecting everything to work. It's like expecting failure. If you expect failure, if your culture is expecting failure, you get an unbelievably resilient system.
Sam Newman (21:11)
Well, it comes down to Postel's law, right? John Postel, who wrote part of TCP spec, his law was, he said, be conservative in what you do and liberal in what you expect. He was talking about the need at the TCP spec to just expect nuts packets to arrive sometimes some kind of, you know, bit of networking equipment goes crazy on you just going to like, that's going to happen. How do I keep operating? And I think that's it moves us beyond that kind of naive mindset a little bit.
And think it's also partly a facet that more of us are being exposed to large scale systems. I think it does come back to how frequently do you see an issue? Like if you've got a system that runs entirely on a single computer, if that single computer fails, really are, you're screwed, right? So therefore it's probably beholden upon you to really make sure that one computer has got a very low failure rate, right? You're going to run it with like redundant power supplies, with redundant rack local networking, probably a rack local UPS and all this other stuff going on, right. So try and reduce the chance that one component failing. When you've got a thousand computers operating, where's that tipping point where you stop caring about the health of the individual machine?
Well, an example I use in the book is the early Google servers, which I got to see when I worked there. But there actually, I think it was a rack in a science museum and there's one I think in the computer, the computing museum in Silicon Valley, it was elsewhere as well. But their server racks, the old ones were like bare motherboards in racks, no cases on top. And that was to make it easy to get hold of components to replace them when they failed. The components that are most likely to fail were things with moving parts, which were the spinning plate hard drives and the power supplies,
The power supplies were just bolted on the back and held in place with Velcro. The hard drives were held in place with Velcro. They weren't screwed down because they'd switched their mindset, which is we've got so many computers that are going to fail, we'll build our systems to handle failure. And then what we're going to do instead is optimize for getting those computers up and running again. This is what David Woods calls rebound effectively, how quickly can we get up and running again? So literally, pull the rack out, rip out the Velcro, throw the hard drive in the bin, stick it back in again, right?
And so that kind of, but that shift doesn't happen overnight. And it's, and I think comes back to that boiling frog problem we talked about a little bit as well. People start with a very, simple system. start with a topologically quite simple system, like say, you know, a monolithic single process, monolithic application, your network topology is simple, your infrastructure is simple. You can't see these issues as much.
But you start increasing the number of deployed units you've got and you'll start seeing these issues soon enough.
Anne (24:02)
And it becomes... it's almost inhuman the scale, when you reach a certain point, then human ideas of probability just cease to mean anything anymore. It always reminds me of a Terry Pratchett joke where, in fiction, the million to one chance happens nine times out of ten. And a sufficiently large system behaves more like a novel than reality. coincidences happen all the time.
Sam Newman (24:47)
it comes back to Murphy's law, aka, sods law as its called the UK, right, which is anything bad that can happen, it will. And I just think you increase that certainty level to like close to one, right.
Anne (25:00)
absolutely.
Sam Newman (25:02)
And I don't know if that feels real unless you've experienced it, but I don't think it's healthy or practical that we expect everybody to experience that failure. that's kind of why I talk in the book quite a bit about... like the reason I talk about air traffic crashes and the rail crashes is because the stakes are so high in those domains that people do proper incident reports. And so we can learn from that. And that's kind of, you don't often get the same level of detail from computer related disasters.
Companies aren't interested in sharing that stuff as publicly. But like I talk a lot about reading other people's incident reports, going and looking at the void database, for example, [ed: VOID is the Verica Open Incident Database, a public collection of software/technology incident reports. It was created to collect and curate real-world outages and failures so people doing resilience/SRE work could learn from incidents beyond their own organisation.] which is where there's a whole load of curated case study, know, incidents that happened, or doing your own game day exercises and things like that to at least try and put yourself in other people's shoes a little bit.
I saw a great talk from incident.io at LDX earlier this year, and they said they did a game day exercise on what happens if their cloud region went down. That's in the wake of US East one crashing, like going down last year. They weren't actually directly themselves impacted by it. But some of their, although some of their sort of third parties were, but they thought, well, actually let's run through that scenario that happened to other people for ourselves. What would that do to us? How would that operate? And I think that's a big part of being resilient.
So, Woods outlined this, that's Eric Woods' four concepts of Resilience Engineering, which are: robustness, rebound, graceful extensibility, which is to do with surprise, and sustained adaptability, which is kind of your ability to keep changing and adapting. And so you have to keep learning. You have to learn from what happened, what went wrong. You learn what went wrong for other people to apply that, to change the system. The system being the people and the technology altogether, right?
And if you don't continually learn and adapt, then you don't have a resilient system, you have a brittle system, right? So that's kind of interesting challenge. I think absolutely, you know, never miss the opportunity to learn from a crisis and all that. But it doesn't mean you have to wait for your own crisis to learn from. So I do try and kind of weave through like ways of testing yourself.
testing at a low level, things like game day exercises and things like that can be really useful to kind of make sure that you're happy with what you're doing or change something or whatever else.
Anne (27:47)
maximizing your exposure to failure doesn't have to be your failure.
Sam Newman (27:55)
No, absolutely not. I came across a fantastic company in my research called Uptime Labs and they do a very cool thing. They do incident handling training. So if the first time you've done incident response is during an incident, it's not a go well, right? You're going to be stressed, adrenaline is running high. You not gonna do a great job. There's a lot to it. So they do like these online exercises, which is kind of deliberate practice because they've got like experts that give you guidance and feedback, but it's like set you up in the slack room with like monitor, you know, so you get observability data, give you information about the incident and you've got to work out how to respond to that incident. And so that's a fantastic way of you doing drills that don't have to be the real thing.
I'd spoken to some people and they're like, it's not the same. It's like, well, no, I it's not the same, but I remember England football team used to not practice penalties. Because the theory was that you couldn't reproduce what it was like to take a penalty in practice, in training, which of course, by logical extension means we shouldn't do any training. They should just go and play the game.
It turned out that when England started actually practicing penalties, they won more penalty shootouts. And for me, like incident training, like something like Uptime Labs' ability to do that and give you good feedback on what you're learning from. And that was chatting to someone yesterday, he's got lots of juniors, they haven't had many incidents.
So he says, that's great, but I'm worried because I haven't got that many people that know how to do incident analysis. So I just pointed at uptime and said, go run a few of those, just run people through them, get them to do like an hour a week or something. It's better than nothing. But you're also, you know, you're going to expose more people to that side of things. And again, that's directly a people element part of it, right? I mean, that's...
There's obviously things we can do technically to make something like incident analysis, response easier, like giving people good data to work with and tools that they know how to use to navigate that. But it's as much about people and behavior as anything else.
Anne (30:04)
And you do have to practice it because a lot of the behaviours when things go wrong are not natural. You learn that things that you naturally want to do are not the right things to do.
Sam Newman (30:17)
and they're not linear. They're not the same. Like I think you see all these, I've seen all these enterprise playbooks for how you do incident analysis and stuff. And it's always like, well, when you have an incident, this will happen. And then you will do this and then you will do this and then you will do this and then profit. And it's like, no, I think, you know, incidents are messy and chaotic. And sometimes you don't realize there's an incident until halfway through.
And some incidents that come up that the right person was in the right place has seen it before and goes, you'd do that sorted. Right. And that may never even get recorded as an incident. And so it's a whole messy set of different cognitive activities that happen during incident response that don't happen in a nice, easy order.
And all of the different aspects of it, you can break down and you can look at, and you can decide, you know, how do I get better at this or that or the, know, whatever other part of that might be.
And so that I found that really fascinating because I'd done incident response and done quite a lot of posts and reviews as well. But without ever really having any proper sort of formal training on that, if that makes sense. And so I've never had the fortune to be able to sort of step outside myself and look at what smart people say about how you should do these things and structure these things. So, you know, I think it was selfishly the act of writing the book has been me having a lot of fun reading lots of research and stuff like that, which has been great.
Anne (31:57)
But there's no problem with having fun. Oddly enough, even the coping with practice critical runs, it's stressful and it's scary, but it's also fun. But it's also team bonding, isn't it? There's nothing that bonds a team like a bit of fear. And it's better if that fear isn't really like real fear, it's just like going to a theme park fear.
Sam Newman (32:23)
but you will. I mean, you swim faster if you drop a shark in the swimming pool. Right. and that, was actually the nice thing. I thought the talk I saw from, for LDX from Luda was, she talks about how when they do their game day exercise, they're two parts, right? The first half of the day is the incident. And they have like people who are villains who are like people in the company that are going to run the incident. They, are the baddies who are pretending to be either a really bad internal stakeholder that won't answer their email about the issue or customers who are complaining or whatever else.
And they've got their own back channel and they're sort of running it. And then you've got the team trying to deal with the incident. then, but the second half of the day is them actually having a bit of a debrief and then they have pizzas coming in and then they actually have limited edition stickers that you get for coming to that game day. And so you try and collect as many stickers as you can, right?
And so they've kind of made it fun and bonding exercise from that point of view. There's absolutely nothing wrong with that whatsoever. And, you know, I think it's...
I think it also lifts that approach to things like that. Take those things away from what could otherwise be like a corporate box ticking exercise. People have to do some training and then now elevate it to being, this is a fun exercise. It's also for the company, incredibly valuable.
Anne (33:54)
I mean, people do escape rooms for fun. The corporate equivalent.
Sam Newman (34:02)
Security people gets to capture the flag type stuff, right? I mean, this is kind of no different. I remember the Google SREs. don't know if they still do it. They'd come around and use a thing called the wheel of misfortune because part of it is like, don't, know, as the people who are building and operating the system, almost by definition, you've tried to deal with the things that you think you need to deal with. It's very difficult to step outside yourself and sort of think about those surprising scenarios.
And so like the wheel of misfortune, what's happening today, it is a catastrophic fire in the data center. What's gonna happen to your system? I mean, I had a big network system problem, one of my first jobs, it was caused by rabbits moving into the network, ducting between buildings. You can't prepare for rabbits, right? So
Anne (34:48)
No.
Sam Newman (34:50)
it's having that kind of external element to those things can be really useful as well.
because it takes you outside yourself a little bit. And you can have a bit of fun with that
Anne (35:02)
there's so much that can be done if people aren't too pofaced about it.
Sam Newman (35:06
Yes, you don't have an outage caused by regional conflict. You have an outage caused because Godzilla has woken up and he's angry. Right.
So it's the same circumstances, but a bit of levity is all right.
I think also coming off the back of that though, as well as, is an acknowledgement about what you're not going to try and do. And this is the point I kind of try and make in the book a lot. It's like... When you say is a system resilient or not, it's not like an absolute property on which everyone agrees. Resiliency is a quality of a system and you sort of have to decide how much of it you want. We can measure that, you know, in terms of SLOs and things like that, But that is a conversation that you have. And I think being explicit about what you aren't going to try and deal with is as important.
I saw lots of... You know, in the wake of the US East one outage, lots of companies being lambasted for only being in one region. Like they were idiots and didn't know this. And it's like, well, those big companies that were impacted, they absolutely knew that this was an issue. They knew this was a risk. The people in a place of... like zoom that was severely impacted, for example, you're telling me they don't have people there that were aware that they were beholden to a single AWS region? Of course they did, but they would have made the calculation that going multi-region would not only have been much more expensive and added complexity, but there's a nasty paradox when it comes to resiliency.
it is something that David Woods kind of outlines in the Four Concepts paper. He kind of says that as we try and change our systems to deal with new sources of issue, new sources of errors, to make it more robust in some way, we actually can end up increasing the source of new risks to our system.
So basically, as we make our system more complex, that new complexity itself can open us up to new kinds of failure modes. So, you know, there are a bunch of challenges around, for example, doing, say, multi-region setups that beyond the cost can actually open you up to new kinds of outage models and new kinds of contagious bugs and things. And so more resiliency isn't always the answer. I think it should be a trade-off. So it comes back to seeing it as a socio-technical system, right? That includes the users of your software. What do they need? What's appropriate? And being very explicit about that.
Anne (37:46)
it's interesting. I mean things like US East has gone down is such a good excuse that maybe actually you don't really need anything more than that?
Sam Newman (38:00)
there is a kind of thing where if you're the only company impacted, you stand out when lots of companies are impacted less so.
I think maybe five or 10 years ago, that would have given you more play. I don't think that happens now because I think we're too dependent on the software. And I think there was some quite egregious issues. Like I don't know why so many UK banks were impacted by us East one going down.. Perfectly good cloud regions here, but they were, right?
So those are interesting discussions points for me. So I think there's I think you've got a little bit of that window But it kind of comes back to nobody ever got fired for buying IBM, right?
When you bought IBM I'm old enough that right you bought IBM software if it turned out to be terrible It can't be your fault because everyone buys IBM if you don't buy IBM and the software you bought goes wrong There's definitely your fault right? So it's safer. You sort of had an air cover when something went wrong
It was interesting to see the fallout from US East one where the blame publicly, there was a lot to go around and a lot of it went to those companies that were impacted. And I think the really interesting thing about that was seeing the contagion. know, seeing, for example, people not be able to unlock smart locks to get access to Airbnb's because Alexa is completely single home in US East one.
People's alarms didn't go off. so they were late for work. I mean, I would definitely have used that excuse too. Sorry, I'm late. US East one went down again. you had a smart litter box that stopped working. I didn't know what a smart litter box was. The eight sleep mattresses, which connect to the cloud entirely. I don't know if it should have been changed by now. So if it was heating, it would keep heating. The schedules were all off the cloud. It would get stuck in positions.
And I think that exposed a bunch of questionable decisions made by those vendors. But that, and, it's kind of going through a lot of the discourse around that. I think we've moved beyond that now where it's acceptable for people to... and people are just more generally angry with software. You know, there was a great tweet I saw the other day, which was trying to explain why the younger generations are so anti AI.
And you and I techies we've been around a long time and they sort of pointed out that we've been around our generation and the older generation have been around to see some significant step changes in the tech industry. All right. So the web, right. was around at 95, I started doing software development, right? So you're just starting to the web being a thing. Public cloud agile with DevOps movement. We've had every five to 10 years, this interesting shift.
that has made software delivery better, has made the software quality and what we can do with it better. And then the younger generation, what have they seen? I'm thinking about my son who's developer now. He's seeing web 3.0 and blockchain and AI, right? Which by and large are both giant steaming piles of doodoo in terms of what they're able to do for us. mean, the benefits of AI is really only on software delivery and isn't financially viable yet, yet software has never been more important and we're never more impacted when it goes wrong.
And so we're in this interesting point where people are really not happy with the continual enshitification of their services. They can't not use them. The appetite for these things failing is going down at the same time as companies are putting functionality into our systems that create new sources of instability. And I think you're going to see a lot more societal pushback all over the world to these sorts of things. governments are catching up very slowly. You've got things like the Dora Act in the EU, for example, which for financial services institutions, which is really quite limp, but it's better than nothing., I mean, that's not why I wrote the book, but I'll definitely use that in the marketing push for it,
Anne (42:36)
So when can we get our hands on the book? When does the book come out? Or is it too soon to ask?
Sam Newman (42:43)
No, you can ask. So if you're desperate, you can get about half the book in early access form. If you have an O'Reilly subscription, can find out something for a free trial and you can read it, but that's all kind of the rough version. The files we going to the printer on the 18th of September. You can pre-order it now from all good bookshops and also Amazon, I guess, if that's your thing. If you're in the UK, use something like Hive where you can order books through local bookshops.
You should be getting print copies probably arriving in your hands second week of October. E-books should be out by the end of September. For those of you weren't sure, authors make more money from e-books than we do print books. They're also better for the environment. But, that's kind of the plans we're looking at right now. There will be as normal like a local version going out in India, which is a better price point as well but I'm not sure when that's happening yet so you can pre-order now if you want to but really looking forward to getting this one out.
Anne (43:44)
as it gets closer and closer, you're just desperate for the book to appear, aren't you really?
Sam Newman (43:51)
I mean, I was really lucky. It took a year longer than it was supposed to. It took three years, which is supposed to be a two and a half. It was supposed to be a one 18 month book. It took me three years. but I was still really enjoying the process of writing it. But I had to stop because I'd keep changing it. Right. It's a problem with books. So I've got to move on to the next thing. We've already got another book planned. So I'm really lucky that the thing I enjoy most is actually writing.
And so my goal is to make enough money to be able to write. So that's what, so for me, writing is the happy place. And so it's kind of, I'll be sad, I'm sad it's over in a way, but I hope it just opens me up to be able to have a lot more conversations around that topic as well. Cause that's one of the great ways, you there was a reason I wrote a second edition of my first book was cause I got to have loads of great conversations with people.
And if I save just a few people from their miserable distributed systems that they've accidentally built without realizing what they were doing, then that'll be enough for me.
Anne (44:59)
That sounds fantastic. So basically, your community of people who read Building Microservices, if they if you haven't read it yet, read that first probably. And then don't build a distributed system. But if you find that you have accidentally built distributed systems, then definitely read Building Resilient Distributed Systems, because it will make your life much less painful.
Sam Newman (45:19)
I mean, and it's called resilient distributed systems because it is really for any distributed system. It's not specific to people build with microservice architecture. So I'm talking generically about the challenges of that. So whatever type of distributed architecture you might have, it's a good, I think, introduction to you. And if you're an expert, I build it into topics so you can jump right into the more detailed stuff. I'll start talking about pack, hell can harvest and yield, or you can start with easy stuff. So I'm hoping it's a book that kind of tries to meet people where they are.
Anne (45:53)
It's a really fun read, as was Building Microservices. it's enjoyable and it's informative at the same time. So
Sam Newman (46:00)
Thank you.
Anne (46:01)
thank you very much for being on the podcast.
Sam Newman (46:04)
You're very welcome. Thank you for inviting me. a nice opportunity to a chat. I really enjoyed it.
Anne (46:10)
Very nice opportunity to have a chat. And thank you very much to my viewers and listeners to for asynchronous... and we didn't talk about how I called this podcast asynchronous and unreliable because you said don't call it asynchronous and unreliable. I thought, do you know actually, I think that might be a good name for a podcast.
Sam Newman (46:28)
Hahaha
Anne (46:32)
But we should have a separate conversation.
Once your book comes out about asynchronous and unreliable, why that's a terrible name for a podcast.
Sam Newman (46:41)
don't think it's a terrible name for a podcast. It's more,
we should talk about the word asynchronous. Because
Anne (46:46)
we do need to talk about the word.
Sam Newman (46:48)
I spent a lot of time looking into that word and that's a discussion for another day.
Anne (46:52)
that is a whole chapter.
like, what does asynchronous really mean? But in the meantime, thank you very much for to my listeners and viewers. I hope you enjoyed the episode and I will catch you again on next week's episode of Asynchronous and Unreliable. Thank you very much.