Guest: Martin Davidson
Watch on YouTube
Listen on Spotify
Listen on Apple Podcasts
Read shownotes & transcript below
Exploring AI, Software Development, and Future Possibilities with Martin Davidson
Join us for a deep conversation on the latest in AI experimentation, software engineering, and reimagining hardware and UI design. Martin Davidson shares insights from his ongoing projects, exploring how AI models like Claude and Codex are transforming the way we build and understand software. #rustlang
Main Topics Covered:
Martin’s current AI experimentation setup leveraging multiple subscriptions from Anthropic and OpenAI to write complex applications in Rust
The shifting landscape of AI models, including the rise of smaller, more capable models like Sol
The concept of "software on demand" — creating applications and systems dynamically with AI
Reimagining operating systems and hardware on demand, inspired by edge computing and unikernels
Challenges and opportunities with AI-generated code and system designs, including over-engineering tendencies
The evolving infrastructure for software development, including automations, bug tracking, and CI/CD systems
The philosophical implications of interacting with AI models as collective "we," and their increasing personality and utility
The future of AI’s role in hardware design, programming languages, and the potential for safer, more efficient systems
Timestamps:
00:00 - Introduction and episode overview
01:10 - Martin’s background and history with AI and software development
02:24 - The current scope of Martin’s AI experiments and subscription setup
04:15 - Transition to more capable and diverse AI models like Luna and Fable
05:54 - Managing AI models as Junior, Senior, and Partner levels for various tasks
06:54 - Cost considerations and AI subscription management
07:37 - The implications of different AI models and their capabilities
09:17 - Chinese models (Kimi) and hosting considerations in the West
10:10 - Subscription choices for AI models
13:14 - The shift from requiring large models to smaller, more capable ones
14:47 - The collective "we" and human-model interaction dynamics
18:26 - Infrastructure for continuous integration and automated development workflows
20:18 - The role of AI in codified workflows and simplified processes
22:07 - Interaction with AI as a collective intelligence and its personality
25:00 - Safety features, filters, and model jailbreak risks
27:35 - The potential for AI to generate operating systems and software on demand
30:38 - Building custom solutions like Uno with AI assistance
32:50 - AI creating websites, understanding architectures, and exploring on demand
36:12 - Reimagining hardware and programming languages for efficiency and security
40:06 - The importance of hardware and language re-engineering
45:52 - Rust and its complexity versus AI automation
50:29 - Over-engineering and AI's tendency to overcomplicate solutions
53:31 - AI’s role in cybersecurity and model safety challenges
55:44 - The future trajectory of models and their increasing powers
58:52 - Optimism in human resilience and AI integration
64:03 - Closing thoughts and future possibilities
This episode provides a fascinating glimpse into the future of AI-assisted software creation, hardware design, and the evolving relationship humans have with intelligent systems.
Anne (00:01) hello and welcome to episode nineteen of Asynchronous and Unreliable, a weekly podcast where we discuss the latest ideas and concepts in tech. I’m your host Anne Currie co-author of Building Green Software, the Cloud Native Attitude, and author of the science fiction Panopticon series.
And today I’m going to be talking again to Martin Davidson, AI experimenter, extraordinaire and Substack author, who came onto the podcast for on episode seven to talk about his personal experiments with using AI to generate mission critical quality software because that’s just his background. Martin and I have known each other for over 30 years, from back in the days when we both worked on some of the big Microsoft products of the 90s.
And and you know, say what you like about MS of the 90s, and it was a pretty evil organization at the time, but the products did work, and they worked on top of hardware and an internet that was a thousand literally a thousand times worse than it is today. So we kind of know what good looks like and what we’re trying to try to achieve.
welcome back to the podcast, Martin.
Martin Davidson (01:31) Hello, yes, lovely to be here. Thank you for having me again. It’s just nice to have a Natter for an hour or so, isn’t it? yeah. I’ve no idea where we’ll go, but we’ll go somewhere and it will be it’ll be fun. So yeah.
Anne (01:37) It will be good because I still rely on your Substack and chatting to you occasionally even in person and on the podcast as well as someone who I trust the judgment of, you know what good looks like and you’re doing a lot of work with AI.
But actually Jon my husband, who you also know and used to work with, has reminded me that the key question that everybody asks after watching your last episode is, What’s your budget? How much do you spend on these crazy giant experiments that you do?
Martin Davidson (02:23) so I’m fortunate because I don’t have to pay tokens at the API prices. so I can get a subscription. and so I have something like between somewhere between eight and ten subscriptions active at any one time. These are the max ones, so the two hundred dollars a month things. Which is fortunate because the company that me and a pal set up that pays for it.
Although really it’s our money that went into it in the first place, but you know, the sort of sleight of hand, you put some money in up front and then it’s a company that’s paying for it, so it’s all fine.
Martin Davidson (03:02) and yeah, we’ve actually made a little bit of money as well, which helps. so that’s been that’s been good. but yeah, I mean I look at the costs if I was to pay it through the API and it’s running into you know, six figures a month. if I was paying API prices, it would be astronomical. and therefore you couldn’t do it.
But the cost is very interesting at the moment because we’re starting to see, you know, with models like Fable and Sol, you’re starting to see a shift in that, for a long time my sort of my mental model was always you used the best model, you know, because the best model wasn’t brilliant. You know, Opus four five was in many ways an amazing model, but was also a terrible model. there was a lot of things it couldn’t do. So you wanted to always be using the best model. and we’re getting to the point now where something like you know Sol Luna extra high or Sol, sorry not five six Luna extra high or Sol, you know, low.
is actually a pretty decent model. Opus four eight is still a pretty decent model. And using Fable as your daily driver is a very expensive way of doing AI. So one of the things that I’ve started to change a little bit is my mental model and I sort of have AIs in junior, seniors and partner levels. so I have Sonnet for example, that’s a junior. And if I want to do you know just code reviews or check-ins or whatever like that, Sonnet’s brilliant for that. If I wanted to clean up a whole lot of work trees, you know, the way our SDLC works now is everything’s in a work tree, SDLC software development lifecycle, so everything’s in a work tree. And sometimes that gets out of control and I can have hundreds of work trees all of a sudden overnight and think, no, this is a disaster. I’m running out of this space. So Sonnet’s brilliant at sorting that out and it’s cheap. and then Opus four eight or you know Sol at like medium or low, that’s my daily driver that’s a good kind of senior level ish in my head. It can manage and organize and coordinate the projects and whatnot.
And then when we need the partner level, you know, the I mean like the Colin Dancer you know, from the from the past, people like him, you know, when you need that real expert to come along and solve your problems, you know, that’s Fable you know.
Anne (05:10) This is an old colleague of ours, very clever. so Colin Dancer, you mentioned there, very clever old colleague of ours, he’s my main technical consultant on my science fiction novels. Cause he’s got a giant head.
Martin Davidson (05:25) right.
Yeah. Yes. So so you know that’s kinda my mental model. So he
Anne (05:33) So I’m gonna stop you then roll your roll your way back. To the first question. So the first thing you said was, You’re very lucky that you’re on subscription. Those subscriptions are publicly available. Anyone can buy those subscriptions, can’t they?
Martin Davidson (06:01) yeah
Anne (06:02) It’s just your company pays for them, although as you were saying, it’s just a little startup that’s you and a pal. So you are actually paying for it by putting the money into the company and then taking the money out of them. So it’s not like crazy money, but you said it was like a couple of hundred pounds, dollars a month, and you’ve got like ten.
Martin Davidson (06:25) Yes.
Anne (06:26) Okay, so you’re you’re basically spending a grand, two grand a month on subscriptions.
Martin Davidson (06:36) Yeah
Anne (06:37) so what models do you subscribe to?
Martin Davidson (06:44) anthropic and open AI. I mean the world is slightly funny at the moment in that for a while at Christmas I had Gemini as well. But Gemini’s really fallen behind the curve. Gemini is not a frontier model anymore. they’re quite a way behind. It probably is, I guess. Four, six months, something like that behind. Gemini
Anne (07:08) Yeah. really?
Martin Davidson (07:09) three six is not even an Opus four eight. Which is interesting. There’s Kimi K three, which was last week, was it? so that’s sort of up there as being a potential, you know, almost as good as Opus four eight, maybe slightly better than Opus four eight. The opinions are still a bit mixed on this at the moment. the evidence seems to be it’s a bit of a distilled version of Fable. certainly it writes and it looks quite a lot like Fable.
So it’s another candidate, but the problem with it at the moment is it’s only served from the inference out of the of Moonshots Labs and they are very constrained on hardware. So it’s very difficult, you know, it’s very slow. and then there’s the whole question of, if you’re actually going to build a product and do something commercial, are your customers going to be happy that you’re using a Chinese model to do that, particularly a Chinese model that’s hosted in China?
So the weights are meant to be available by the twenty seventh of July. at the point where the weights become available then it could be hosted in the US or the UK or wherever and then maybe that becomes more palatable
Anne (08:24) well, so interestingly, the next episode that comes out that you haven’t seen yet, is with the CEO founder of Neuralwatt, who are a EU and US hosting company for AI, and they are adding Kimi Or I think they might already have added Kimi And and for their customers, it has to be hosted in the US or Europe.Those things are becoming available. So what I think you’re saying is if I understand you correctly, is that you’ve got you’ve got ten subscriptions, but they are all basically different models of anthropic and open AI. Is that is that right?
Martin Davidson (09:17) yeah, I mean you get a Claude Max subscription for example I call it max max because there’s two variants of max. There’s one that gives you five times pro and one that gives you twenty times, and you know it’s like, well, the max plus twenty eight anyway. and then you just have access to Opus four eight and there’s been the whole on and off again thing with Fable.
you know, it was there and then the US government took it away and then it was back but it was only gonna be back for a few days and then it was back for another week and now it’s back indefinitely, hopefully, but only for half your subscription, you know.
The anthropic position is interesting at the moment. I guess they’re compute constrained because fable has disappeared from the pro subscriptions. whereas the position with OpenAI is much cleaner. You just get a subscription and here’s your model. And OpenAI are all a bit fiddly as well, because you get a five hour limit and you get a weekly limit. So you get however many tokens you allow per week. But you can only use a certain amount of them in every five hours.
So anthropic say you can only use twenty percent in every five hour block. you can easily use twenty percent of subscription in under five hours. Fable can do it in under an hour if you put your mind to it. so you end up juggling backwards and forwards between them.
Whereas OpenAI have removed the five hour cap and it’s just like use your subscription so you can use it easily in under a day but you don’t keep getting paused. So it all just keeps moving and fiddling, you know, jumping around. But you’ve got to look at the positive side of all these things. It’s amazing having access to these. You know, and there’d been a lot of people who’ve been upset with Fable coming and going and whatnot, but that we have access to it is just fantastic, you know.
It has caused me some slightly sleepless nights where you’ve sort of been up in the middle of the night thinking, Fable’s reset, right, I’ve got all these tokens, I need to use them. So not sure that’s entirely healthy and it probably tells you something about my psyche, but yeah.
There we go.
Anne (11:25) Yeah.
Martin Davidson (11:26) I’m not intending to be in this state in a few years’ time, you know. I still imagine that I’ll be on a beach with my feet up, reading a nice book.
My wife says it’s never gonna happen, but yeah, I can still dream.
Anne (11:42) Yeah. You dream, but you
love working with these models. I'd say I’d never seen you so happy, but you’ve always been happy. I’ve never not really seen you happy.
Martin Davidson (11:52) yeah, talk to my wife.
Anne (11:58) So you’ve got subscriptions to a variety of models from the same folk. but the thing you say is that it used to be that you’d have to be at the frontier, the really expensive models to do what you were doing. Years ago, in AI years, last century, when we were talking three months ago, you were really having to use the big models to do what you were doing.
So I seems to me what you’re saying is the big difference over the past three months is that the smaller models have become better and more usable for you. So you can use the smaller models for some of the stuff that you in the past you had to use the most expensive models for.
Martin Davidson (12:49) Yeah. And if the small model gets stuck, like with humans, it can go and call on a bigger model.
Martin Davidson (12:56) so Opus four eight can decide, yeah, this is a bit beyond what I’m capable of, or I would like a second opinion or whatever, and it will go and talk to Fable or talk to Sol and get a you know, it’s better call Sol, isn’t it? there’s a there’s a good joke there. I haven’t quite figured out what it is. But it will go and invoke those and get some feedback.
There’s a thing, you know, I used to do as a manager when I had a team and someone would come to you with a problem and you think, Have you talked to Blah about this? You know, you used to do that a lot. I’m sure you’ve done it a lot. Have you
Martin Davidson (13:29) talked to this other expert? And I do that with now with Opus. It comes to me and says, Well, this looks tricky and I’m a bit stuck here. "Have you talked to Fable?" I’ll go and talk to Fable. I’ve got an answer now.
And it’s great because it’s not just like having you know one Colin You’ve like got
A hundred Colins ’cause you can have a hundred fables, you know, and they can all opine on something and come up with slightly different views on things. it’s amazing.
Anne (13:54) So the brilliant Colin Dancer here, who you’re referring to as our canonical super brain, is getting a lot of call-outs here. He works at Snyk these days, the security company. I haven’t seen him in ages. But I need to get him back involved because I’m writing the finale to the Panopticon series, so I need to get him to review all my technical stuff again.
Martin Davidson (14:17) It’s slightly weird at the moment because we’ve got the smartest models we’ve ever had, but they have the most distinctive personalities as well. So Fable is very imaginative and explorative and will come up with different solutions. But sometimes it’s a little bit lazy at times, whereas Sol is like a dog with a bone.
You know, it will just go and go and go and go and go.
Martin Davidson (14:48) and there was some stuff I was doing over the weekend, I think and we started on Friday at I don’t know, about six in the morning or something, and we eventually finished. You know, this single session finished on Sunday at, you know, I don’t know, four or five in the afternoon.
You think? Wow. And there was a cool bug as well, because it turns out that if it’s a single turn, Codex doesn’t check the session limit.
So it hit its session limit and used up my whole weekly allowance sometime on Sunday morning. I think it’s seven on Sunday morning, then proceeded to run for what’s that, another nine hours, you know.
Maybe I shouldn’t shouldn’t say that publicly.
Anne (15:30) Hopefully nobody who’s watching will go and turn that off.
Martin Davidson (15:33) Well, I mean
what happens a lot of the time you find is there’s a reset. You know, there’s a bug or something, there’s a bug burns through your tokens faster than it should have done and they just do a reset and you know you’ve got a freebie. they’re they’re very good. So yes, that’s that’s my kind of setup.
Anne (16:01) So basically you’re in a similar setup to three months ago, in the sense that you are basically still working with Codex so open AI and Claude, anthropic, but you’re using more models and you are considering in the future extending into Kimi, the Chinese model, but the trouble is it would need to be something that
customers could accept ’cause it wasn’t quite so Chinese hosted, for example.
Martin Davidson (16:36) Yeah, I think there’s two things. First of all, I think it’s quite different actually. I think one of the things that’s changed a lot over the past three months. It’s hard to remember where I was three months ago.
Anne (16:48) Ha ha.
Martin Davidson (16:48) but I think one of the things that’s changed is the SDLC has changed. So we’ve put together much more of a framework around how we do software. When I first started, it was a bit more kind of Wild Westy and eventually processes and then things got defined and people figured out what best practice looked like.
I remember there was one team I joined 30 years ago where the way they tracked bugs was an Excel spreadsheet and the backups were printouts that were kept under somebody’s bed at home. You know, and you think Well it works, but..
Anne (17:21) I think I might have been in that team.
Martin Davidson (17:26) it works, but perhaps not the long term solution. What we’ve got now is a sort of a whole lot of infrastructure around this, which is the tooling that sets up all our systems so they’re all the same and sets up all the prompts that are consistent and the skills everywhere, and that we can update stuff really easily. I mean, actually, as I’m talking now, it has toasts all the time, so it’s popped up something, it’s one of the queues is blocked and it wants a human.
so it manages our bugs and our issues and our requirements. It has scheduled tasks, every night it runs code reviews. Sol was a mixed blessing because the automatic code reviews, when Sol suddenly took them over, it found a lot more bugs than five had been finding. but that’s okay because then there’s an auto-drain which just goes and drains and fixes the bugs, and then there’s a verify that checks the bugs are fixed, and then everything gets integrated. And
You know, we do our development in work trees, so everything’s in parallel, and then all the work has to come together and get merged into main. and there’s an auto-integrate tool for that. So you’re not trying to integrate, you know, four work trees in parallel because all that will happen is one will win and the other three will lose and then have to rebase and then you know you go around again. So you have this queuing system that’s provided. And the infrastructure also manages our local GitHub runners for doing builds so we can have an overnight night to build and all of this infrastructure suddenly comes together to make your life so much easier.
And it’s one of the things that made me realize that I think three months ago I was probably still interacting with Git, albeit with some wrappers around it. I don’t interact with Git at all anymore. I don’t even go... I used to go and look, you know, at GitHub to see how many things I checked in. I don’t do that anymore because I get a daily report from our agents, you know, and it sends me this daily report and I go and read it and it’s like, Well, there’s the status of how many things we checked in yesterday and whatnot.
It’s very much like you know, when you become like a first or a second line manager and you start putting the processes in place of how you want your team to report stuff to you.
Anne (19:33) Mm.
Martin Davidson (19:34) it’s like that. And then as a second line manager, do any second line managers know how to use Git? Probably not really. They shouldn’t be, probably. And and it’s a bit you know, that’s the world. And I was listening to the podcast you did with Jon and he was talking about how ai might change the way that the hardware worked and that the we wouldn’t have programming languages and those kind of things. And I was sort of reflecting on that. A long set of dots took me to this sort of realization that yeah I increasingly talk or type.
Words are the interface that I have with a computer. I’m you know when I’m faced with a UI where there are buttons and things and weird hieroglyphics it’s like, I don’t like this. I want to be able to ask someone.
When we were setting up this call, I was trying to find out how do I turn on echo cancellation, right? I wanna be able to talk to this app and say, right, how do I turn on echo cancellation? And they go, okay, I’ll do that for you. You know. I don’t wanna have to go and try and find the button because it’s a much more natural interface, it’s much more, you know, discoverable and usable.
Anne (20:48) absolutely. So, I’d like to ask you a question or clarify something that you’ve been saying a lot during the call today, which is we. Now, the interesting thing is I do know there is another human in your team, Ed, who I know. I had a coffee with him a couple of weeks ago. He is a human, but my guess is that when you’re talking about the we there, you’re not talking about Ed, you’re talking about the models.
Martin Davidson (21:14) Yeah, I think I use we very loosely there. I think a lot of it is, you know, perhaps a lot of it is all the years of "there’s no I in team". You know, and I never much like the people and managers who would say, this is my team and I do this and I that you know. I always I always felt that was the wrong steer because it’s a we, it’s a collective. And, as a manager, you’re kind of the guider and the lead.
Not the leader. You’re sort of just gently steering things as opposed to like the rest of you are my minions and you will do as I say. So I think I’ve probably always been a bit sensitive to that kind of language. So we seems softer. But I do, I kind of
I think there’s also a thing: credit where credit’s due. The AIs can do a set of things which I think are quite frankly astounding and I can’t do. You know, and they deserve some... do they deserve some credit for that? I don’t know. I mean, where do you take draw the line? A calculator can do maths much faster than me, but I’ve never felt the need to say, thank my calculator or talk about it as we.
Anne (22:31) It’s hard not to, isn’t it? It’s hard not to start interacting with the models as people. I mean I actually enjoy talking to them and it does feel like a we, which is interesting.
Martin Davidson (22:54) It does, doesn’t it? I mean, you’re talking to another intelligent being and arguably one of the most knowledgeable well, more knowledgeable probably than any human that you talk to. I mean, it has access to so much information, it’s astounding. And you realise that... It’s when you talk to other people and you realise I’m only talking in this very narrow little set of my special interests. I could be talking to it about a whole load of other things I don’t even know to talk to it about.
Anne (23:33) Yeah, I talk to it about all kinds of crazy stuff. But that is my special interest. All kinds of crazy stuff.
Martin Davidson (23:38) Yeah.
Yeah. But you know, I imagine you don’t talk to about, you know, the detail of legal contracts or something. You know, I mean if you were a lawyer or something, you’d be having this great conversation about, you know, case law and this example case from nineteen thirty two, you know, that set the whole precedent and you know I’m sure there’s stuff like that, and I’m sure there’s people who could, you know, are having that conversation.
<PINGING NOISE>
Martin Davidson (24:10) Sorry, that noise. I should turn that off. That is our dev system telling me that it’s finished a turn.
Anne (24:21) There’s more tokens
coming free.
Martin Davidson (24:26) Yes,
There we go. Oops, there we go. I have turned it off.
Anne (24:34) I’m gonna leave that on the tape because that kind of your work environment. That is Claude or Codex calling out to you and saying it’s your turn.
Martin Davidson (24:44) That’s cool. Yeah. That’s five made a sound for me. It authored its own sound to tell me that it was finished. and Claude did the same. I did change the Claude one because the Claude one initially was a triumphant, you know, trumpet.
Anne (25:03) Yeah.
Martin Davidson (25:05) That’s really irritating. So the Claude one got changed.
Anne (25:09) Ha ha
Martin Davidson (25:09) But I quite like the Codex one.
And it and it’s the story of my life. And and you know, we were away on holiday and we came back from holiday and I turned things on and the sound started and we were sitting through next door and I was sitting with my wife on the sofa and in the background we could hear this ding you know and then she literally said, good, we’re home.
Yeah. So but yeah, when it finishes a turn it pops up a toast and summarizes what it’s been doing and it plays a little tune and you know
Yeah, it’s irritating when you forget to turn it off and go to bed and you’re lying in bed and you’re just about to fall asleep and it goes ding.
Anne (25:54) Yeah.
Martin Davidson (25:57) So yes, it’s been the backdrop of my life for a bit. Ed really doesn’t like it at all. He turned it off. I had to add the option to silence it because it was driving him insane.
Anne (26:12) There are
still humans in your life as well.
Martin Davidson (26:16) There are long suffering humans.
Anne (26:29) Actually, I quite like it when during a podcast we’re interrupted by the real life of the person that I am talking to. because I think real life does interrupt things. There’s there’s no doubt. I mean, Sara, my co-author in Sweden has two young kids had to dash off to school to get one of the kids, getting a phone call in the middle of our one of our conversations. And it’s life. You know, things get interrupted.
Martin Davidson (27:05) it’s just how it goes, isn’t it?
Anne (27:09) So essentially there’s a lot of stuff
going on at the moment. You’re getting a lot of alerts. So what are you working on?
Martin Davidson (27:18) That's not very many.
No that’s quite quiet really. so what are we doing at the moment? so we have our startup and we’re working on some stuff with that. a lot of the stuff that I’m still doing is on the framework key pieces and getting that right so that I can unblock both myself and unblock Ed. so that basically Ed has a development environment that has all the stuff that I’ve learned over the past however long and he can just hit the ground running really hard. That’s one of the big bits.
Going back to what you and Jon were talking about because I think that was one of the things that was really interesting to me was you know Jon sort of saying well maybe the hardware changes, maybe you don’t have programming languages and things anymore. You know, and almost the whole idea that you generate software on demand.
We’ve had sort of cloud stuff where you can spin up software on demand, but could you actually generate software on demand? And you know, there’s a bit of me that thought initially, that’s a bit far out, isn’t it? is that actually going to happen? And then I sort of paused and I realized actually that’s exactly what’s happening right now.
So I’ve got GitHub pages and you can have your own little website and whatnot and publish stuff there and it’s all nicely integrated with Git. so I’ve got that. I’ve got my own one of those. and I realized that when we were on holiday, I just built lots of things just because I was I you know, standing in a queue or something, and you think, well, I could just make a thing. and we built software on demand, really.
Yeah, so you there’s a card game, Uno. I don’t know if you’ve ever played it.
Anne (29:16) yeah.
Martin Davidson (29:17) yeah, and we like it as a family. It’s a good game. But where we were standing in this queue, it wasn’t really possible to play Uno. So I just got fable to create Uno for me, and I have a really pretty decent implementation of Uno. We called it Onu, to avoid any copyright things.
I’m not sure that’s enough, but anyway, and that really was software on demand. The thing built it and we could play it and you it’s it’s quite incredible.
And there’s a set of things as well in my life, you know, I sort of read about something, you think, that’s an interesting thing. I’ll go and rat hole down that and find out some more about that. And there was one I was reading which was like once the temperature goes above twenty five degrees C, for every degree above that, your cognitive ability decreases by two percent.
and there was you know some study done in two thousand and six and there’s evidence for this. And I thought, I wonder, I’d like to know more about that. So Claude went off and created me a little website about that and you know some graphs and I could, you know, change my body temperature and see what happened. And it was really interesting. you know, and if I’ve got an interesting bug, like with the you know, I’ve been building this DOS emulator, I was trying to get Windows 2000 running on it, and there’s a whole
raft of problems with that. And and it was really interesting, the story. And it’s the kind of thing where, you know, in the past you would have hoped somebody, you know, when you had really interesting bugs, somebody would give a talk about them, you know, and tell you how they debugged it and all the things and everything they went through to do to solve that. Well, I’ll just get Claude to do that for me. So Claude’s produced this little website, you know, and I can read through all the Windows 2000 stuff. remember Paul Theobald? <Ed: an old colleague of ours> he got in touch recently and he had an old copy of Colonization and he knew I like
stuff. So said I’ll send this to you. And you know, and we got that running in the emulator. And then we went and had a look at how it actually works under the covers. That’s amazing how it works under the covers. There’s a little tiny thin, you know, game engine. And then there’s a whole lot of demand pages that are paged in and out on demand. And they come from, you know, extra expanded memory or extended memory. and then it dynamically compresses the data segments for each of these pages and has a little cache of those. And you sort of think this is this is remarkable
Cool architecture, you know, it’s like a little mini OS. And the kind of thing that you would never, I’d never have known that was how it was architected, but here you are with using the AI tools to reconstruct the architecture of these things and understand them. And this is all sort of just on demand, you know, you just have a thought of it. I would like to know about this, and you can create a little website and then you can go and you know understand it.
Martin Davidson (31:59) when I when we go on holiday, I like to go for a run in the morning.
I don’t like running, but I you know, running around in a city or something’s kinda interesting. You get to see stuff you wouldn’t otherwise. But I really hate running. Anyway, so I thought I’ll just build a little running app, you know, and lo and behold, I have a little app now that, you know, will recommend a route, starting from where I am and give me a little five or ten K route around wherever it is and track how well I do it running and tell me you’re quite slow, but you know, hey. Anyway. And you know, that again was software on demand.
I mean it’s not quite instantaneous, you know, it’s like a couple hours, but still. What a weird world.
Anne (32:38) Yeah, it’s a very weird world. Yeah.
Martin Davidson (32:41) But
you know, you take that forward and you start to think, well, that’s little things today, you know, that will grow the size of those things you can build on demand. You know, and at what point are you starting to build operating systems on demand? You know, and things like that.
Anne (32:57) Well, so at the risk of calling back previous episodes that people might not have listened to yet. But people should listen to this episode a couple of weeks back from Justin Cormack I don’t know if you listened to that, one of the founders of Unikernels
Martin Davidson (33:11) Yeah, yeah, yeah.
Anne (33:14) and yeah, it’s like build your operating system on demand. It was a thought that he and a group of folk came up with about 10 years ago that operating systems are big, they’re bulky, they do an awful lot. They’re very much like a general purpose model that does loads and loads of things and therefore you’re paying a lot of CPU tax in terms of running that model plus security exposure because there’s a load of stuff in that operating system that you’re not using, you know, that somebody’s using because everybody in the world is using this operating system. So it has to be all things to all folk.
but in fact you just need a much smaller operating system that’s much more focused on your use so if you could build that on demand that was just the small amount, the unikernel version of the operating system that was just what you actually needed, it would be faster, it would be more efficient, it would be greener, it would be safer.
Martin Davidson (34:12) Yeah. and I think the door for things like that starts to open. it was you know, pondering about what Jon was saying, I was sort of thinking, well, we’re in a world now where rewriting and re implementing software is really quite straightforward. You know, so you know Windows and as much as I love Windows dearly, it’s showing its age. You know, thirty years ago you could get Windows and people talked about running it for a year.
I mean generally these days I get three or four days before there’s a kernel leak and I have to reboot. And it’s probably not Windows’ fault, it’s the driver’s fault. But the fact that you have an architecture that allows drivers to leak like that and doesn’t stop them is... anyway. you know, it’s getting to the point where that would be an interesting thing to replace.
Excel, you know I’ve got a copy of Excel 2000, so that’s what, 26, 27 years old Surprisingly, it does pretty much everything Excel does currently. But it loads instantaneously and okay, the UI looks a bit old fashioned, but you think twenty seven years all we’ve managed to achieve is the UI looking slightly differently and make it a lot slower.
Anne (35:19) Well, and make it require literally a thousand times more resources to do that same job.
Martin Davidson (35:25) Yeah. you’re absolutely right. It’s completely bonkers how much CPU is wasted. and you sort of look at these things and you start to think, well, these things are ripe for re implementation and you could... I had a go, I don’t know, six months ago starting to re implement Excel ’cause you know, just see if you could do it. And we got so far and it wasn’t brilliant, but I imagine these days you could get a lot further.
and so my initial thought was well, you re-implement these things and that’s all good, but most people don’t really care, you know, a slightly better version of Excel doesn’t move the needle for them. But then you start thinking, well, where I was talking about earlier, conversational UIs, because why do I have Excel? Well, I have Excel because I have a spreadsheet for money and I want to know whether I’m going to be able to pay the bills and whether when I can retire and things like that. But I don’t need a spreadsheet to tell me that. If I had a conversational UI that I trusted and I could have a chat with, then well that’s good enough. I don’t care about it being a spreadsheet. So maybe actually the switch that you see is a change to the UI. You know, the thing that Andre Carpathy talked about several years ago was that the LLM becomes the UI. And people thought that was quite far fetched, I think. But I suspect he probably was right. and at the point where you’re changing the whole of the user interface, then you have the opportunity to change and reimagine all the stuff that’s underneath it as well. you know and that becomes quite exciting because you start to think about the software and the hardware stack. You know, the hardware stack x86 is well, x86. It was a stopgap when it came around in the 70s. It was APX 432 that was going to be Intel’s shiny new future. And that failed. And then
I860 was going to kill x86, but that failed. And then Intel’s grand plan was to kill it off. Even when X86 was really successful, they were still trying to kill it. You know, Itanium was going to kill it, and Itanium failed. You know, and eventually they’ve got into the sort of, well, okay, well, we can’t kill it, but it’s not. I don’t think anybody in the world would describe it as a great architecture. It’s full of compromises and full of history, and it was never brilliant to start with and it’s part of the reason that we have a lot of the problems that we have today.
You know, APX432, it had this idea that a pointer wasn’t just a pointer, it was a capability. So the pointer also included things like your read-write access to it, whether you could copy it, boundaries on it. You know, there was a whole set of other things that came with the pointer. it had, you know, low-level hardware support for object-oriented programming. So there was another layer of security built in lowdown in the hardware. And you think, well, that would be, you know, quite cool if we could have that. You know, how many of the exploits of the current day would not have happened in the past 20 years? How many of those exploits wouldn’t have happened if we’d had this hardware? And we’d had, I mean, you know the whole story with ADA as well, right? ADA was that the US Department of Defense realized that software languages were getting out of control.
There were too many of them and some of them were prone to exploits and weren’t very secure. They’d had that realisation in the late nineteen seventies. And Ada came out of that and then Ada died because people didn’t like the idea. Well, it was yeah, I think it was a combination of being complicated and mandated by the US government.
So yes, you know, maybe we’re at a point where we can reimagine the software programming languages and we can reimagine the hardware.
you know not only do you reimagine the UI but you reimagine everything that sits underneath it as well. and when you pause and you think about it, Rust is a fantastic programming language, but it still has to deal with the compromises of
the hardware that’s underneath it. You know, and if you can improve the hardware then do you get a different language than Rust? I don’t know. I don’t know enough about this, but I do know that there’s a lot of research that’s been done in these areas over the years and a lot of it has come to nothing because driving that change is just too difficult.
Anne (39:49) Mm.
Martin Davidson (39:50) it’s really, really difficult. I mean, Intel haven’t managed... the only people who’ve ever really managed to change hardware architectures are Apple.
And I think that’s because Apple controls the whole ecosystem and well, basically you either do what Apple want you to do or you don’t use Apple products. It’s that simple. so
Anne (40:10) Yeah.
Martin Davidson (40:10) they can force the change through.
Anne (40:13) So as you said, Jon was asked asking the question, how much of hardware design is limited by the fact that in the end the programs that run on it that are written by and supported by humans. and you know, almost certainly a lot of that research into better hardware and more secure hardware was thrown away both because of the change, and, you know, all change is difficult to do, but also because the code had to be written by and supported by humans and it was too complicated. you know even Rust is famously hard to learn.
Martin Davidson (41:01) Rust is really hard and I think Jon’a absolutely right on that because why are we seeing an explosion in Rust right now? Well, because humans don’t have to write it anymore, you know. We all see
Martin Davidson (41:12) the benefits of it. We’re not completely stupid. You know, we can see that it’s really, really useful. But it was far too hard. The cost benefit wasn’t in the right place. And now it’s free. Why wouldn’t you?
Anne (41:25) Yeah. Yeah. So an old colleague of ours, do you remember David Drysdale? you know David very well. I know David.
Martin Davidson (41:33) yeah. Yeah. Yeah, yeah, yeah.
Anne (41:36) Yeah. he wrote the O’Reilly book, Effective Rust, which is a very good book, apparently I’m told, about you know, how to learn Rust and so he’s like a super world expert in Rust. And I was talking to him about, shall I learn Rust? He said no.
Martin Davidson (41:52) Yeah.
Anne (41:52) He said, learn Go. Don’t learn Rust. It’s too hard. It’s too hard for you. And actually
Martin Davidson (41:58) Yeah.
Anne (41:59) I’m very good friends with David. But I think I could see what he meant. It’s hard. It was hard in those days. but if you want to learn it for yourself, you should read Effective Rust because David’s an excellent writer.
Martin Davidson (42:12) Absolutely I mean and it doesn’t surprise me. He is incredibly smart. you know, so Rust is a sort of natural thing for him. But yeah, I don’t know, I thought you were going to say whether you should learn Rust and I thought he was gonna say no, AI can do it all for you. ’Cause I think that’s probably the answer. I mean at this point in time, you know
Learning a learning a programming language is really a kind of recreational thing. It’s not a thing for your career. I think there’s a learning how machines work and the underlying state of them and that kind of stuff is still useful to have the right mental model. But knowing the syntax and the ins and outs of Rust, I mean I don’t know that. I can’t remember it.
Anne (43:04) Yeah.
Martin Davidson (43:09) it’s strange how it changes, isn’t it? What was really valuable five years ago really is of very little value today.
Anne (43:18) Yes, it’s astonishing. and yet as Jon said, we’re still in the transition and we could be in the transition for a very long time. Transitions take a lot longer than you realise. But I had an interesting conversation yesterday with a somebody who’s a consultant, who has been working with a lot of businesses and one of the issues that she’s seeing is junior developers go off with an ai, go off with Claude and come up with some crazily complicated scheme. I think there’s a French term for it, a folie à deux I think. That’s you kind of form like bubble of love and self-reinforcement. It’s kind of the Romeo and Juliet issue where you go, we could do this, and we could do this, we could do this, and you come up with some cunning scheme and it’s too complicated. And you actually need someone externally to burst your bubble and say, no, that’s too complicated a system. And people go too far in their conversations with the AI
before they expose their scheme to anybody to go, it’s too complicated, it does too much, this doesn’t work very well, this doesn’t work very well. and that feels like that is, I mean, obviously it’s, you know, Romeo and Juliet was written hundreds of years ago. So it’s not a new, a new problem of people going off and coming up with, you know, the Thelma and Louise. You don’t want to end up finally with people going driving their car off a cliff. You you want to actually have someone stepped in earlier and say, Is that really a sensible idea? You know, actually it’s not. Let’s burst your bubble a little bit there.
But now people can go a long way in a very short time before anybody comes along to burst their bubble.
Martin Davidson (45:25) Yeah it’s true, isn’t it? And I certainly recognise that in the stuff that Claude and Codex produced. They’ll over engineer stuff. They’ll come up with really complicated solutions because...
I think I don’t know why they do it. I suppose it’s a combination of stuff that’s in the training data and probably also sort of subtleties in the way that I prompt them. Even though one of the things I’ve had for a long time in the system prompt is, we’re looking for simple solutions for things. but there was one yesterday where we’ve got some logic which will the auto integrate across multiple different machines in the fleet in multiple different geographic locations. It was coming out with this really complicated algorithm for how you could take a piece of work and it might fail on that machine, and that machine might die, and then you could reinstantiate on another one and you could share among different members in the fleet and things. And I just said, look, Why? if we just assume that the machine that does the work does the integration. And if that doesn’t happen right, well we lose that piece of work. Okay, it becomes an awful lot simpler then, right?
yeah, and a lot simpler to test as well, because there’s one mainline path and that’s the path we’re always going through. And so there are things like that where it over-engineers there’s another one where it does slightly daft things at times. So with the DOS emulator, one of the things that we’ve got is a thing called DINREC, dynamic recompilation. So you know, by default we interpret all the instructions and we have a simulated, you know, x86 CPU.
When you think about it, you think, well, I’ve got an x86 CPU here, right? If I could take an add, you know, that you’re doing in 32-bit or 16-bit land, and I just convert it to a 64-bit one and run that, then life’s much faster. So that’s dynamic recompilation. and the two things you need for well, very crudely, two things you need for dynamic recompilation. First of all, you need the framework to enable you to do it, and then you need the instruction coverage. So you need all the hundreds of x86 instructions, you want to cover them. So you want adds and moves and all the various different variants of those.
so Claude went away and said, Yep, we’ll do this and we’ll do it in two parts. We’ll do the infrastructure and then we’ll do the coverage of the instruction. So it did the infrastructure and then it performance tested it as it likes to do. Nope, this doesn’t make any difference whatsoever. You know, unsurprisingly, because there were no instructions for it to recompile because it hadn’t covered any of them. So it reverted that change. And then it went on and it did the instruction coverage and it performance benchmark that.
Nope, that doesn’t make any difference at all. That code’s not being called. Well, of course, because there’s no infrastructure to call it anymore, because you got rid of that. And so at the end of, you know, a night’s worth of work, it had nothing. This is a complete disaster. This is useless. And you’re like, Yeah, well, hmm. Maybe if you tried them together, you know, whole being greater than the sum of the parts and all that sort of stuff. And and so you’re in this weird world where it’s doing two things that are really complicated beyond what I could do. And yet it doesn’t see to glue the two together is what you need. And I think this was fable, you know, it’s not a stupid model, and yet because of the way it’s decided to divide the problem up and not look at it holistically. So you have this really weird world where sometimes it overcomplicates some things, sometimes it doesn’t join the dots correctly.
Maybe those are the skills you need to teach people, or maybe it doesn’t matter because maybe in six months’ time, you know Fable Six or whatever it happens to be will solve these problems. But then maybe we won’t have access to Fable Six. We’re getting to that point where the models are so powerful that maybe this is as good as it gets for normal people.
That one’s all really interesting as well, as government begins to realise the models have potentially significant impacts.
And Kimi being as strong as it is is an interesting one as well because if you use Fable or you use Sol you’ll frequently find that it decides there’s a cybersecurity risk here and we’ll need to not allow you to do this sort of quite obviously straightforward fine bug fix. The filters are quite extreme. the first time I used Fable I asked it how it was different from Opus 4.8. That was the first thing I asked it. And it
It said, I’m good at doing biological stuff and then the filter kicked in, you know, to say, ooh, you can’t talk about biological stuff.
Anne (49:58) Okay.
Martin Davidson (50:00) We instantly dropped back.
But you know, Kimi doesn’t have these safeguards and it’s the work of not much effort to jailbreak it. So suddenly you’ve got a very powerful model which potentially although the evidence at the moment seems to be it’s not very good at cybersecurity and if you were being uncharitable you might conclude that’s because it’s distilled from Fable and Fable doesn’t allow you to do cybersecurity threats. But maybe that’s just not true.
I yeah, I mean
Where it goes in the next four or five, six months is really interesting. ’cause we’re getting to a point where I think some of the models are...
I mean if this was as good as it got, this is still a fantastic place to be in.
Anne (51:00) Mm.
Martin Davidson (51:01) I and some of the things you look at and you think, Well, you’re building these little apps on a website and things for me and you know, how much better could this be? You know, it works. it works first time, you can one shot it. I mean it would be nice if it was a bit quicker, but and it was nice if it was a bit cheaper, but functionally I’m not sure there’s any more function I want in it.
Anne (51:26) Yeah. And yet I don’t think that it’s about to stop. so where’s it going? it’s now hard to imagine what you would do with the additional power that’s coming in the bigger... the models keep getting bigger with seemingly no end in sight.
Martin Davidson (51:44) Yeah, and things you can do are slightly bonkers. So the DOS emulator, I thought it’d be cool if it ran Linux. You know, so we’ll try running an old version of Linux. I got that. And I thought I wonder if we could write a compiler that could compile Linux for me. So I wrote a C compiler, it wrote a C compiler for me, it compiled the Linux kernel and then we booted Linux. Well that’s a strange thing, isn’t it?
Yeah. I have some Turbo Pascal code. I wrote about this in the substack, you know, that only runs on sixteen bit windows. that’s irritating, isn’t it? I wonder if we could rewrite the Turbo Pascal compiler so that it could compile this to run on sixty four bit windows. Well it does. So I now have my old code and it works on Windows eleven. Yeah. These are weird things.
Anne (52:33) These are very, very weird things.
Martin Davidson (52:35) Yeah. You know, a few days ago I suddenly remembered Solaris two five So Solaris used to run on X eighty six and I spent a lot of time working with Solaris and on X eighty six and I thought, wonder if the emulator could run that. Yeah, it does. Just it’s fine. You know, I have an emulator that can run Solaris now. And Solaris was a pain to run, you know, and I mean it had to write an S three graphics card and emulate that in order to get it to work. But you know that’s you know the emulator graphics card and the bias for it. Yeah. Not a problem.
Just give me a few hours. I find increasingly it’s hard to think of things that are difficult enough that it can’t do. The performance stuff is the one thing that it’s still really struggling with in the emulator. But beyond that, most of the rest of the stuff is tractable. I think there is still a thing where it’s not able to step back and see the big picture in quite the way I would like it to. so that’s where I think a lot of value that Ed and I provide comes from, but you know, I mean at the point where it does that then maybe I do get to go and sit on the beach ’cause, you know, not needed anymore.
Anne (53:48) Yeah.
Well we just have to hope that there’s a nice Iain Banks Culture end to this story rather than a more negative one.
Martin Davidson (54:00) I don’t know. I yeah, I mean humans are pretty good at surviving at most things that have been thrown at us. So you kinda have to be optimistic and it’s very easy to look for all the things that could go wrong. you know, and we’re quite good at creating situations as well that could make things go wrong for us. I mean we’ve got global warming and nuclear missiles and now we’ve created AI as well. You know, so you know quite imaginative at ways that we might destroy ourselves with, but they seem to hang on. So
Anne (54:30) Yeah, we shouldn’t talk too much about this ’cause I remember from our last podcast episode, as soon as we started talking about dystopias, the recording just went off.
Martin Davidson (54:41) yeah. Well let's talk about happy things. I mean one of the things I’ve realized with it is you know, these little websites as well is I’ve learned so much in the past six months. I didn’t really know about dynamic recompilation and how all that worked. I didn’t know about the intricacies of the various different timers and things. I hadn’t really thought very much about software development life cycles for a while. There’s a whole lot of stuff that my son got me to create as well, which I’ve gone and read, which is you know, like the Permian extinction and the end of the dinosaurs and it turns out that he was interested in all these creatures that sort of half destroy themselves. So there’s a toad that when it’s under threat will fire blood from its eyes and other animals that they punch their skeletons and their ribs through their body in order to attack predators and things. there’s all other stuff, you know, this weird extra world. and yeah the BBC had this article about ocean currents and things and I thought that’d be interesting to learn about. So I created a website with ocean currents. And it’s far better than BBC article ’cause you know it shows all the ocean currents and you can drop bottles and things in and you can see how they follow and Lego bricks and you know this kind of things. And
It’s just wow, there’s so many toys to play with.
Anne (56:13) Yeah. You are massively, massively ahead of anyone else that I’ve come across in terms of actually actively using I mean, you started a long time ago. in AI years, you started you know, a century ago. You started two and a half years ago, three years ago.
Martin Davidson (56:33) Yeah, I think from very early on I was trying to get it to create code.
Anne (56:39) Mm.
Martin Davidson (56:40) and I mean it was kind of lucky because I’d been a manager for a long time and then I sort of realized I didn’t like being a manager. and I like the technical stuff and I’d always thought, you know, once I got to a certain age I would go back to being an IC and do the technical stuff. And there was a big sort of you know, there’s this whole sort of imposter syndrome thing of thinking like, well, the world’s moved on, you know.
Is an old C programmer, do they really have a place in this modern world? and you know, so I learned a bit of Rust but then AI sort of appeared at almost exactly the right point, you know, and I was able to throw myself into that. And it’s been good at the point where I was reinventing myself, the world was reinventing itself as well. you know, and I and I didn’t have any baggage, you know, I wasn’t tied to a particular project. I was, you know,
whatever a principal engineer and my job was to go and learn about new things and so it was all I was very lucky that everything just lined up. and also I think I’m really interested in it as well. I mean I like building and making things and you know you can build and make things like you never have before. Before we went on holiday I got Fable to go and make me a MIDI file. and they weren’t bad. And then I thought, well, the obvious next thing is for it to make a MIDI synthesizer. So it made a MIDI synthesizer, you know, and I now have a pretty half decent MIDI synthesizer. And I understand a lot more about MIDI and a lot more about LA synthesis and a whole lot of stuff that I was sort of tangentially aware of but now I you know and it’s just
It’s bonkers, you know? I used to dream of having a sound canvas when I first started playing computer games because that was how you got really good sound, but they were like, I don’t know, five hundred quid, a thousand quid or something back then, so that’s probably two thousand quid in today’s money, something like that. Nobody could afford that. And now for a handful of tokens you can build one.
Hmm, this is kind of cool. So yeah, you can get the models to express their creativeness and yeah. I think the sort of the limit is your imagination, isn’t it? And and I don’t I think a lot of people don’t quite know how to get started sometimes.
But if you’re willing to experiment and you’re happy to be wrong and you’re happy to fail. I think that’s the thing, isn’t it? I’m quite used to failing. Lots of things don’t
Anne (59:41) Yeah.
Martin Davidson (59:42) work. but then that’s okay.
Anne (59:47) So on that note that happy note, I’m gonna say
Martin Davidson (59:51) that’s all you’re getting.
Anne (59:52) our context windows are full. and with that’s been a really, really good episode. Thank you very much indeed for being on again.
Anne Currie (59:59) So so yes, so to all our viewers will have noticed that I’ve just jumped from one location to another because I’m having a bit of difficulty with my Wi-Fi connection. but that yeah feels like it reached a good point to end the chat because it was a very positive note. and so can I say thank you very much again, Martin, for being on the podcast because that was absolutely delightful. I really enjoyed it and I will
I say the good thing and the bad thing about these podcasts is I listen to them ten times ’cause I have to edit them. but the good thing is that then I listen to them ten times. So I pick up a lot more than I would have done otherwise just during the conversation. It’s an interesting new form of conversation.
Martin Davidson (1:00:48) Yes. Yeah. Well I’m I feel for you having to listen to me ten times over
Anne Currie (1:00:54) I always enjoy your ones, so it’ll be good. It’s easy actually, I enjoy all of them. so it’s always quite a delight to listen to all the conversations ten times. but thank you very, very much indeed for being back on. And hopefully I’ll be able to entice you back to be on future episodes because it’s always fun to chat about your crazy progress you’re making.
Martin Davidson (1:01:18) well thank you for having me. It’s lovely to have a chat. It’s lovely just to chat to you, Anne. I mean, whether it’s on a podcast or with a coffee and a and a bit of millionaire shortbread.
Anne Currie (1:01:29) Yeah. So I’m not up in Edinburgh as often we well, Jon and I are not up in Edinburgh as much as I would like. So possibly the next time that we will be up there again, we’ll meet up for a cup of tea and a millionaire short bread.
Martin Davidson (1:01:44) Yeah, sounds good. Well thanks for having me.
Anne Currie (1:01:48) Thank you very much.
Martin Davidson (1:01:50) Take care.
Anne (1:01:53) So thank you for listening to the last episode of Asynchronous Unreliable. But now I have some work for you. We’re very keen to answer questions that are submitted by listeners, take suggestions for who we should contact to come on the show next. we’re you we’re always very happy to do that. And you don’t have to record a snippet yourself, you can just send in a message. so you can contact me. If you’re on YouTube, you can add things to the comments. If you if you’re not
On YouTube, if you’re listening to this on a podcast, you can reach out to me on LinkedIn. I’m Anne Currie Very happy to connect to people. Just just reach out and connect, and I’ll and I’ll connect to you and you can send me your questions, your ideas, your thoughts. And of course, your thoughts on what W E T, wet programming, the opposite of dry programming. What should that TLA stand for? Now at the moment, the one that’s in the in the lead is Gemini’s suggestion, which is
Write everything twice, which I quite like. We do quite like that one. So but you know, if you if you’re a human and you have an idea, please do let us know. So thank you very much indeed and we are very, very keen to hear from you. Thank you.