Guest: Sarah Wells
Watch on YouTube
Listen on Spotify
Listen on Apple Podcasts
Read shownotes & transcript below
Navigating the Evolution of Enterprise Tech: Cloud Native, AI, and Beyond
In this episode, Anne Currie and Sarah Wells, ex Financial Times engineering leader and O'Reilly author, explore how enterprises survive rapid technological shifts from cloud native transformations to the rise of AI. They discuss lessons learned from the Financial Times’ cloud journey, the comparative impact of AI on software development, security concerns, and the importance of fostering resilient, adaptable teams in the ever-changing tech landscape.
Key Topics:
Lessons from the Financial Times’ migration from on-premise to cloud, emphasizing gradual transition and cultural change.
Insights into cloud native adoption and what enterprises learn through trial and error in cloud journeys.
How AI is transforming development practices, including benefits, pitfalls, and the critical need for asking the right questions.
The widening gaps in AI governance, security risks, and the importance of automating security in code reviews and deployment pipelines.
Comparing the industry’s early days of internet boom with current AI evolution and understanding the long arc of technological adoption.
The enduring necessity of software design, architecture, and human oversight in an AI-enhanced environment.
Challenges around junior engineers working with AI tools, the risk of over-reliance, and the importance of effective collaboration.
The legal, security, and ethical implications of AI, including managing vulnerabilities and maintaining trust.
Reflecting on how past technological disruptions, like the decline of physical newspapers and the rise of digital, inform current industry shifts.
Timestamps:
00:00 - The impact of weather and garden work on daily routines
00:29 - Solar and heat pump innovations at home for sustainable living
01:50 - Patterns in enterprise digital transformation and cloud native adoption
02:23 - Introducing episode 21 of the podcast, focusing on documentation failures
03:19 - Sarah Wells’ background and her role in transitioning the Financial Times to cloud native
04:22 - The cultural and technical challenges faced during enterprise cloud migration
05:34 - The gradual shift from on-premise to cloud and the importance of cultural change
06:11 - Early technology choices and innovations at the Financial Times
07:10 - The journeys of building internal tools before mature cloud services existed
08:06 - Lessons learned from cloud bills and cost visibility initiatives
09:22 - The importance of cost transparency in cloud environments
10:36 - Metrics and practices to improve system operability and documentation
11:39 - Continuous improvement driven by clear policies and team collaboration
12:13 - Drawing parallels between cloud native transformation and the current rise of AI
12:44 - The industry’s response to AI, including fears and opportunities
13:47 - The speed of AI adoption versus organizational capability for governance
14:04 - Addressing governance gaps and the use of source code in AI training
15:27 - The security challenges brought by AI and the increasing sophistication of vulnerabilities
16:53 - The risks posed by AI in security patch management and vulnerability detection
17:22 - Current research on security vulnerabilities and AI’s limitations in security testing
18:35 - The implications of insecure Java code in AI training data
19:52 - The importance of patching and the challenges faced by enterprise Java systems
20:20 - AI's strengths in modern languages versus older or more complex codebases
21:17 - The plateau of AI improvement in security vulnerability detection and the concerns around it
22:22 - The gap in AI’s understanding of broader security contexts like cross-site scripting
23:04 - Practical governance tools, automated validation, and good practices for secure AI integration
24:07 - The value of clear policies, checklists, and guidance in security and compliance
25:33 - Supporting non-technical staff in secure AI usage and code review processes
26:00 - The role of threat modeling and security reviews in organizational safety
26:57 - The importance of framing questions for AI to avoid false assurances
27:42 - How AI might mislead or escalate security risks without proper guidance
28:14 - Challenges in junior engineers’ reliance on AI and fostering critical thinking
29:48 - The emotional investment in code and the risk of over-attachment to AI-generated work
30:28 - The importance of collaborative feedback and continuous learning in AI-augmented development
31:31 - The risks of siloed knowledge and the importance of teamwork in AI-enabled environments
32:01 - Research on developer productivity with AI and the emerging dependency risk
33:12 - Historical perspectives on tool reliance, from books to Stack Overflow and now AI agents
34:42 - The next iteration of coding assistance and the enduring need for core engineering skills
35:25 - The ongoing relevance of senior engineers amidst rapid AI-driven change
36:31 - The long-term view of AI’s influence and the patience required for meaningful industry shifts
37:37 - Reflections on technological disruptions like COBOL, Blockbuster, and the internet boom
38:52 - Lessons from past overestimations and early tech failures like Urban Fetch
39:34 - The comparison between past internet hype and current AI enthusiasm
40:27 - The realization that the true value of new technology unfolds over decades
41:02 - The challenge of prioritizing meaningful features in rapid development cycles
42:16 - The importance of curated product features and simplicity in user experience
43:02 - Using AI for language learning and personal growth applications
44:21 - The limitations of traditional language
**Anne (00:00)**
Hello and welcome to episode 21 of Asynchronous and Unreliable, a weekly podcast where we discuss the latest ideas and concepts in tech.
I'm your host, Anne Currie, co-author of *Building Green Software*, *The Cloud Native Attitude*, and author of the science fiction *Panopticon* series. Today, I have the great pleasure to welcome Sarah Wells, independent tech consultant, author of O'Reilly's *Enabling Microservices Success*, and former technical director at the Financial Times newspaper, which is how we met. In fact, I have a whole chapter in *Cloud Native Attitude* devoted to what Sarah was doing at the Financial Times and their work going cloud-native. Basically, Sarah, your background is all about getting enterprises to adopt and benefit from modern practices. So, welcome.
**Sarah (00:50)**
Hi, thank you for inviting me. I'm looking forward to this.
**Anne (00:57)**
Yeah, you've interviewed me before because you do—was it GOTO or O'Reilly?
**Sarah (01:03)**
I think it was GOTO, yeah. They like to get two people together who are going to chat, and we've done that before.
**Anne (01:10)**
Yeah, so we have been on a podcast before when Sarah was hosting and I was the guest, so it's a reversal. Although to be honest, it doesn't really make that much difference, does it?
**Sarah (01:19)**
No.
**Anne (01:21)**
One of the reasons why I've been quite keen to talk to you on the podcast—well, for many reasons, because it's a lot of fun to talk to you—is that we've run into one another quite a few times in the past month or so at AI conferences in London.
**Sarah (01:38)**
Mm-hmm.
**Anne (01:39)**
You've actually been out there as a tech consultant talking to a lot of companies.... the Financial Times is probably the most successful online newspaper, isn't it?
**Sarah (01:59)**
It certainly is very successful. I think it has a very good understanding of the audience it's trying to address. The other thing the FT did that was very good was a successful transfer from thinking about the newspaper first to thinking online first. That was a challenge, I think, for a lot of newspapers to start thinking in terms of what it is going to look like on the website. As soon as journalists got used to that, it became, "What is it going to look like on mobile?" It's very different from the layout on paper.
**Anne (02:32)**
Right.
**Sarah (02:33)**
Different from paper layout.
**Anne (02:36)**
It is one of my favourite chapters in *Cloud Native Attitude* because I talked to people who were really doing the work to find out what their problems and issues were. I followed back over and over again over the course of years and the FT and your experience there was a classic experience in that it took a long time. You didn't expect to be in the cloud from on-prem overnight, and you weren't.
**Sarah (03:04)**
No. I mean, there are two sides to that, I think, which is the actual time it takes to get everything out of a data center—it's not an instant thing. But there's also the cultural change that you need to have to really benefit from the cloud, to really be cloud-native. So it wasn't just, "We are decommissioning our data center," it was, "Well, we're going to build microservice-based systems and use containers or Lambdas or cloud components rather than just deploying our applications in exactly the same way on AWS as we did in the data center."
I think that cultural change is also big. We were also doing it at a time where many of these things people hadn't done. It wasn't easy to go and look at what other people were doing and say, "This is how you do testing in microservices and containerized applications," because it was very early days for a lot of these technologies. So you're working it out.
There are a load of things where I talk to people and say, "Well, it would have been great to use this particular technology, it just didn't exist when we were doing it." Kubernetes did not exist when we were first running containers, so we built our own orchestrator. Obviously, you don't do that now. Backstage and portals like that didn't exist when we were first starting to struggle with the idea of keeping track of our software estate, so we built our own. It would never be my first choice—if someone else has built it, I would rather use it—but we were so early we had to do a lot of that work.
**Anne (04:45)**
You told me a lot of stories about the early days of the FT. When I spoke to you, I think you were a couple of years in. It wasn't the earliest days; you were actually successfully partially in the cloud. By the time the last edition of *Cloud Native Attitude* came out, you were totally in the cloud. But you didn't attempt to do it in one go, and there were lots of interesting things that you learned through trial and error on the way.
**Sarah (05:17)**
Yeah, you don't get it right first time all the time. Actually, as a consultant, you quite often go in somewhere and you're not there for that long. But I was at the FT for 11 years, so I was there the whole way through this transformation. When you've tried one thing, discovered it doesn't work, and moved to something else, you start to really hit the challenges of understanding whether a sensible approach was taken.
**Anne (05:43)**
Indeed, yes. One realization that you came to—that I repeated to other people more often than any other—was when you first went live in the cloud, the AWS bills were crazily expensive. You learned that you had to give the information about the bills to the people who were running up those bills, along with the responsibility and agency to drive those bills down. Then they did it, and that was completely transformational.
**Sarah (06:21)**
I think visibility helps in lots of aspects. Certainly in terms of your costs for things like cloud, if you show people what they're using, they'll turn around and say, "Well, we're not actually using that; we could get rid of it." Visibility in terms of what services you have, who owns them, and how secure they are—anything that gives people something to look at is really useful.
One example we had was creating the System Operability Score when trying to improve our runbooks. We had a couple of thousand services across everything at the FT, and some of them had very limited documentation. We had this runbook format and created a score where we basically said, "Here are the most important fields, here are the next most important fields." We aggregated it all together to calculate the score for a service, across a team, and across the group. That visibility really encouraged people to think, "Okay, I know what I need to do to improve the runbooks." We had teams that went from nothing to 100% runbook coverage very quickly.
**Anne (07:40)**
Sharing best practice and sharing what good looks like?
**Sarah (07:46)**
Yeah, sharing what good looks like. But also, the very first thing like this at the FT was way before I was a director of technology. Our operations team would send out an email once a week with the five noisiest alerts. The moment I actually concentrated on this was when my team was in that top five noisiest alerts. I was like, "That's embarrassing, let's fix that." Anything that shows people, "Hey, this is actually really loud, can you fix it?" is good.
**Anne (08:19)**
That's very good. One of the reasons I'm very interested to talk to you at the moment is that you were right at the beginning of Cloud Native in a proper major enterprise, you successfully did it, it took a long time, was really difficult, and you learned loads of things through trial and error. Now you're out and about at what is almost the equivalent moment in the next giant change in the industry: AI. How is it the same and how is it different?
**Sarah (08:56)**
I think there's more worry about where this is going with AI. It feels like a bigger transformation than moving to the cloud. I don't know where I stand on that; at the moment, my view is that it's an extremely useful tool. It is going to change the way that we approach coding, but it looks more to me like a next generation of programming.
You're still thinking about software engineering; you're just using a different tool to do it. I think it affects everyone. There's always been an ability for lots of software engineers to just get on with what they're doing and not worry about moving out of the data center or changes in architecture. But this is going to impact pretty much everybody. It's right there, being used by engineering teams, and it's changing the way those teams work quite significantly. Far more people are sitting there thinking, "What does it mean for me and my career?"
**Anne (10:04)**
Speed-wise, how do you think it compares to the beginning of Cloud Native?
**Sarah (10:12)**
In terms of adoption, it feels like it's going very quick. It really does. It's interesting because it's outpacing organizations' ability to understand what they're doing.
There's a big gap between the number of organizations that have people using AI and organizations that have really thought about how they want to govern that. Recent research showed almost no one has AI-specific governance. About 25% of companies don't have an AI policy, 38% have a comprehensive policy, and yet 90% of organizations have people using AI. So there's a gap.
The other gap is that in lots of those organizations, AI use is being done through personal accounts, which means you really don't know what's happening. Interestingly, research says the number one thing pasted into these personal chats by developers is source code. Your source code is being loaded into ChatGPT by your employees, and you're not really on top of that.
There's real pressure to say that you're using AI to get on with it, but there's also a gap. It exacerbates the perception that compliance and security slow you down. People have invested in the last five to ten years to make it less like that, but now you really have to work super hard to keep up with what's happening and help people be more secure and compliant without being seen as a gatekeeper in front of all of this.
**Anne (12:05)**
The security risks at the moment are crazy. I saw one of your fellow people who were featured heavily in *Cloud Native Attitude*, Greg Hawkins of Starling Bank. I was talking to him about this last week, and he pointed out Microsoft's Patch Tuesday, the patch release for their enterprise software, was the biggest ever: 600 patches needed to be applied that week. That was second only to the previous week, which had been 200 patches. These have to be applied by enterprises to everyone's desktops and systems. It's astonishing, isn't it?
**Sarah (12:59)**
It's terrifying. AI makes a lot of things easy to do with less investment. Within an engineering team, you might say, "I'll just build this one-off tool because I can do it really quickly using my AI agent," in a way that would never have been worth it doing it yourself.
The same is happening with security: it's much easier to check for vulnerabilities. It's going to improve the security of everything eventually, but in the meantime, it's going to be extremely painful.
The other aspect—I'm writing a workshop at the moment about governance for AI-enabled organizations, particularly looking at the security side—is research showing that while AI agents are getting really good at writing code of acceptable quality, they're not getting any better at catching security issues.
If you are overwhelmed in an organization because there are so many PRs from people using agents, and code reviews were always problematic for catching stuff like that, you can't trust that AI agents are going to be good at catching security vulnerabilities. Interestingly, they are particularly bad at Java, which looks like it's due to the training data—there's a lot of insecure Java code used to train these agents. Security is being left behind.
**Anne (15:14)**
It's interesting you say that about Java, because Java has come up several times in previous podcast episodes. I talked to Holly Cummins from Quarkus. For many years, Java hasn't been great at being turned off to apply patches, so they invented Quarkus so it could be turned off and on to apply patches. But that's still not what the majority of people are running.
Most enterprise Java systems are still not fully in the cloud, so they can't necessarily patch these systems easily. Justin Cormack, one of the founders of Unikernels, was on a couple of weeks ago talking about how good AI was at writing in particular languages. He pointed out that it was very good at writing Rust, mostly because Rust is modern and follows good best practices. With C, C++, or languages with a huge historical mass of code of variable quality, the AI gets very confused.
**Sarah (17:01)**
Yeah, I imagine there is a large amount of legacy Java code. If you're not telling the model to only look at modern code, you're going to end up predicting average code. The average of all code is probably pretty bad.
What's interesting is that it's not improving either. A lot of improvements for AI agents are around the LLM harness to make coding agents usable and less frustrating, but it's not happening for security. Research suggests agents are better at localized issues like SQL injection, but struggle with vulnerabilities requiring broader context, like cross-site scripting. I've been doing work with security engineering teams recently, and it's tough to ensure everyone stays secure when code is being produced in high volumes. How do you implement automated governance in your CI/CD pipeline to flag issues?
**Anne (19:34)**
Those are the right things to do. What kind of uptake is there on good practice?
**Sarah (19:43)**
You need to make expectations really clear. Software engineers generally want to do the right thing, but they don't always know what that is unless you're clear. If you automate validation and provide feedback on what to fix, that's appreciated.
At the FT, we created an engineering checklist for building code and getting it to production—covering procurement, security, accessibility, and documentation. A principal engineer told me, "I really like this because I know if I tick all the boxes, I've done my job."
Clear policies help—for instance, specifying approved AI tools while routing new tool requests through security, legal, and compliance. You also have non-software engineers writing code now; you have to guide them on security. You hear stories like someone visiting their GP and discovering the GP "vibe coded" an app that inadvertently exposed patient data. Organizations must provide clear guidance.
One organization I work with requires a security review and threat modeling before production release. A customer success manager built a Chrome extension and proactively arranged a security review. That's fantastic because it shows the security message landed without being an intimidating process.
**Anne (22:41)**
You mentioned last time we spoke a failure mode you observed: people asking the AI, "Is this the right thing to do?" and the AI reassuringly saying, "Yes, that's fine," which escalates risk because the AI is so encouraging.
**Sarah (23:17)**
You have to ask the right questions and provide the right context. You can't just say "make this secure"; you need to instruct it to run a review against specific benchmarks like the OWASP Top 10.
Junior engineers don't necessarily have that context. The danger is a junior engineer working with a coding agent goes down the wrong path. Working with a senior engineer, they would get early course correction. With an agent, they might present finished code after days of work, only for a senior engineer to realize it's completely the wrong, overly complicated approach. It's hard to fix once someone has invested that much time. Where is the collaboration and learning that helps junior engineers grow?
**Anne (25:05)**
They learn the hard way, but discarding everything at that stage doesn't teach them as effectively as learning from mistakes one at a time.
**Sarah (25:15)**
It's much easier to adjust someone's approach before they've finished. When I code with an agent, I have 25 years of experience to anticipate edge cases—like checking for unencoded commas in CSV files. We need ways to let people benefit from agents while continuing to learn from colleagues.
**Anne (26:18)**
My husband Jon mentioned that managers have always needed to touch base regularly with team members to review their progress. Now, people can write a vast amount of code between touchpoints.
**Sarah (26:48)**
People often want to figure things out independently before showing they don't understand something. That's why pair programming or ensemble programming was great—you could learn together in a high-trust environment. Agents make it easy to privately ask questions, but it stops you from admitting what you don't know to your team. We have to adapt collaboration to account for agent usage.
METR did research into developer productivity in 2025 and found that people's perception of their productivity with AI agents wasn't very accurate. When trying to repeat the research in 2026, many developers refused to complete tasks without an AI agent, even when paid. They had to change the format of the research because developers could not be separated from their agents. It becomes a crutch very quickly.
**Anne (28:33)**
Decades ago, if you didn't have a physical technical book, you couldn't do your job because you had to look everything up. Then developers couldn't program without Stack Overflow or IDE auto-complete. Agents are the next iteration of assisting with the basics, except the definition of "the basics" is much larger now.
**Sarah (30:18)**
It's the next iteration.
**Anne (30:20)**
It's been a long time since we wrote code without any external references. How are we going to train future senior engineers, or will the existing senior engineers just age out?
**Sarah (30:57)**
A year ago, people claimed vibe coding meant you didn't need to be a software engineer to build apps. People build personal tools that way, but turning something into a maintainable product at scale still requires software engineering.
Some think that in five years agents will just rewrite code from specifications every time, but I still see value in the software engineering mindset, design, and architecture. We will still need those skills.
**Anne (32:06)**
We're in a transition that could last 20 or 30 years. COBOL engineers are still around and earning good money because the systems remain.
**Sarah (32:55)**
We're in the very early days. In five years, we'll look back and realize how little we understood about where this was heading. It reminds me of starting in IT during the dot-com boom. Many early companies went bust because of wrong predictions, but the internet ultimately transformed everything.
In my first IT job, there was a company called Urban Fetch where you could order food, CDs, or goods delivered to your office for just the cost of the item. It was hugely popular, but they went bust because they had no path to profitability. Years later, instant delivery is everywhere. The early AI winners might not be the ones around in 10 years.
**Anne (34:44)**
When we met 10 years ago around cloud-native, it felt like anyone not adopting it would go out of business due to the speed and resilience advantages. Yet ten years later, most organizations still haven't fully transitioned.
**Sarah (35:25)**
You realize that for many organizations, software delivery speed isn't the primary bottleneck. Writing code faster with AI doesn't solve the problem of knowing *what* to build.
When I managed development teams, I stopped worrying about backlog items below the top three slots because priorities shifted by the time we reached them. My worry with AI is that we'll ship tons of features fast that nobody actually wants.
**Anne (36:49)**
Like the Simpsons episode where Homer designs a ridiculously over-engineered car that flops. People want curated products, not every possible feature.
**Sarah (37:08)**
Maybe we'll see a return to simple products that just do one thing well, like original Google search.
**Anne (37:25)**
I do find AI useful for research—it points me to things I might not have unearthed through traditional search.
**Sarah (37:57)**
It's great for conversational practice. I've been using Claude to brush up on my Italian before a holiday. It adapts to my level much better than apps like Duolingo, which force you to practice obscure phrases instead of practical scenarios like ordering in a restaurant.
**Anne (39:02)**
Like learning how to say "fill it up" at a petrol station in school French when you're 12 years old and can't drive.
**Sarah (39:50)**
AI technology is mind-blowing given that it's essentially predicting the next token.
**Anne (40:13)**
You once told me a story about Kodak. They didn't fail because they failed to invent digital photography—they built the first digital camera—but because they couldn't pivot away from their core film business.
**Sarah (41:22)**
The FT board wanted to avoid that constrained thinking. They wanted the flexibility to see where things were going and change quickly. Cloud, microservices, and DevOps allowed us to build and test new ideas rapidly. Blockbuster is another example—they failed to adapt when physical rentals became obsolete.
**Anne (42:56)**
We're both speaking at the Fast Flow conference later this year, which focuses on team topologies and rapid iteration. True organizational fast flow isn't just about speed of delivery; it's about the resilience to change direction fast.
**Sarah (43:39)**
It's about fast feedback loops. Releasing code frequently allows you to test a hypothesis, see if it works, and roll it back if it doesn't.
At the FT, the web team added star ratings to film review listings, thinking it would drive readership. Instead, traffic dropped because readers just looked at the rating and skipped the article. Because it didn't take months to build, they easily removed the feature without ego.
**Anne (45:12)**
In old waterfall development, nobody wanted to throw away two years of work. That brings us full circle: if a junior developer builds an entire feature with AI without check-ins, they become emotionally attached to it, making it hard to roll back.
**Sarah (46:06)**
Once you invest time, it's hard to accept feedback. When writing my book, I had fantastic technical reviewers like Sam Newman, Daniel Bryant, Anna Shipman. Their feedback made the book so much better, but it required extensive rewriting.
There's a saying: "There are no finished books, only exhausted authors." You need early feedback loops before you're set in your ways.
**Anne (47:45)**
Having co-authors on *Building Green Software* helped because every chapter was reviewed internally as we wrote it.
**Sarah (49:40)**
You have to get input from others before you get set in a direction. With AI, it's easy to make a lot of progress quickly in the wrong direction.
**Anne (50:43)**
It's frustrating reading content that's clearly AI-generated. Writing is a thinking tool—you can't offload it.
**Sarah (51:16)**
When using AI for coding, my core responsibility is software engineering: defining requirements, security, and quality standards, even if I'm not typing every line of syntax. Designing software is the actual thinking tool.
**Anne (52:39)**
Designers, architects, and ops engineers will be needed for a long time to shape systems to specific business contexts.
We're approaching the hour mark—this has been a great, wide-ranging discussion. Thank you very much, Sarah.
**Sarah (53:39)**
Me too, thank you!
**Anne (53:43)**
Thank you to our listeners and viewers on YouTube. Catch you next time on *Asynchronous and Unreliable*.