A practical field guide for founders and product leaders who want to harness AI development speed without building the wrong thing faster than ever before.
At BoS Europe 2026, Lizzie shares Memrise’s 2 year quest to address the question: How can AI help us and our customers make progress?
The JTBD approach meant raw AI capability was focused on using tools to build products people pay for. But it wasn’t clean. There was developer resistance, fear of replacement, plenty of junior school experimentation before anything clicked.
While AI tools can revolutionise/democratise software creation, many teams remain stuck in legacy product development habits – lengthy roadmaps, pixel-perfect releases, and developer bottlenecks. You’ll hear how the culture shifted from siloed, waterfall-style teams to one where everyone builds and developers evolve from gatekeepers to system supervisors.
In two years, the team has moved from an 18‑month feature freeze, to shipping multiple AI-driven products – some built entirely by non‑technical staff – generating meaningful revenue.
Key takeaways:
- How to change culture, teams and roles when anyone can ship with AI.
- Practical patterns for rapid, ‘good enough’ AI prototyping without blowing up quality.
- Tactics to bring developers with you, not replace them.
- A proven playbook for turning AI hype into products and revenue.
Slides
Find out more about BoS
Get details about our next conference, subscribe to our newsletter, and watch more of the great BoS Talks you hear so much about.
Transcript
Thanks to you. Hi, I’m back again. Hello, Mark. Let me in again, which is really nice of him, to do this nice talk about product velocity. Hopefully it’s not going to be as spicy as last year. Are there any software engineers in the room? Oh God, it’s going to be spicy. Sorry, I misinterpreted that. Okay, never mind.
So this is me, I’m Lizzie. This is me on a strawberry, setting the tone for the entire presentation. It’s going to be this serious.
So my job is I am the head of Applied AI in Memrise, so it’s my job to try and discover cutting-edge technologies, ways that we can improve our systems with AI, and try and apply that at an organizational level for Memrise, and I’m not going to say that this is an easy thing to do, it sounds like the dream job, but actually testing out all these different products and putting them in front of everybody that we’re working with, and putting an AI plan together of how we’re going to optimize everything is a lot harder than it seems.
So, what I’d like to talk to you guys about today is talking about this crazy velocity that we get with AI, and it’s not just accelerating development, accelerating across the board, and what we do with that level of velocity.
So, I’m gonna make a confession, which most of you probably already know, which is that I can’t code, and I never will. I’m a girl who has tons of slides, like this is – I’m a very visual person, I think, in images, so there’s 90 slides, so you know, buckle in.
And despite not actually being able to code, I have been developing a lot of stuff using Claude Code, using Replit, using Cursor. A lot of you are probably very familiar with these things. I’ve made my own Pokémon game, because I am a 90s chick, and I cannot wait until I make a retro version of Pokémon. I got bored of waiting, so I just did it in Cursor on my own. Zero development knowledge. It’s the first thing I ever made. I’ve made chat rooms where you can speak in multiple languages, complex analytics panels. I’ve made spicy companions. We’ve been there. If you’ve not been there, you should go there. Maybe I’ve made less spicy companions like AI buddies for Memrise, which will teach you a language and nothing else. I’ve also made analytics logins and battled with OAuth. Matt, you were along with me, battling Google OAuth for a long time. That was fun, but yeah, I remember finally nailing down getting a Google OAuth login, and Matt writes to me and goes, “Ah, by the way, there’s a totally different thing you need to do to get this in production,” and I was just like, “Can it email addresses?” But all that’s to say is that I could build all of this stuff with me in a room in Spain in an afternoon or in a couple of hours, and this is an insane velocity that is in anybody and everybody’s hands right now. I’m not saying anything different. All of you know these tools are available, you know you’ve seen Replit, you’ve seen Cursor, lovable, loads of them are available right now.
The Traditional Development Loop
So I want to go back and have a look at that traditional product development process flow that we all know and love, which is we go to concept, we think of an idea that we want to build, we throw that into Figma, into design and prototyping, and pixel perfect it, throw it over to user interviews, and then we ping pong between design, prototyping, and user interviews, ad infinitum, and then we go to development, where all the developers question why we’re actually doing this, and wonder if we should go back to concept stage. Then we spend time building it, and then we put it into QA, and then we launch, and we breathe. We hold our breath. We hope that we’ve actually built the right thing. We sacrifice a goat every Tuesday to the success gods, and you don’t have to be the last one, but those who know, who know, but we hope that we’ve actually built something that people want, and this process can take anywhere from months for fast teams to years, and a lot. The problem with this process that everyone has discovered is it’s extremely expensive, so a lot of the talent that you have here in the concept stage is just no longer there by the time you launch, and everyone’s forgotten why they’re doing what they’re doing and what they’re doing, so keeping that talent that is like really, really great, along with you on the ride, as well as paying for them, is an extremely time costly and also. Any costly experience.
What Changes When AI Removes the Bottleneck
So this is the AI development product development process now that we know and love, but it’s not just removing a couple of steps. The bottleneck from development is no longer development anymore. All the kind of parameters that we set around development have kind of dissolved because we can put AI tools in the hands of absolutely anybody now, so we can go from concept design and build, launch user feedback in one day with a non-technical person who has no idea how to code, building in an island kind of architecture, and just throw stuff against the wall, try and get product market fit as much as you can. This is super cheap, anyone can do it, you can do it in your basement, and you can do it very, very cheaply, but the problem with this is you generate this, because anyone can build, doesn’t mean that everyone should build, and because anyone can pick up these tools, it means that anyone and everyone can start throwing stuff against the wall, but without any guidance or any reasons for building stuff, this is what you generate, and it’s been coined AI slop. Everyone has seen AI slop before.
So, just to give you a little bit of context, I’m Lizzie. I work at Memrise. It’s a language learning platform. It’s been around since 2010, 2011, 70 million users, and this was a real problem, an existential problem for Memrise, we wanted to level everybody up in Memrise to be able to code with AI, so that we could try and find a better product market fit. And we started by training everybody up in the team to be able to code as non-developers, because the theory goes that the person who’s doing the user interviews, or the person who has the most domain knowledge, who can iterate the fastest is probably going to be the winner.
So, to be perfectly honest with you, I don’t have the answers. This is just me giving you a boots on the ground run through of what we’ve tried, the potholes that we’ve crashed into, and hopefully you’ll be able to avoid those or drive into them. It’s totally up to you if you guys want to find out on your own, but I just want to caveat this with saying that this process of changing a company to use AI development tools and having non-developers using non-coding tools, it’s hard, it’s messy, like we don’t have this even fully defined now, but what we do know is that what we gain from having non-developers coding is actually giving us more value than not having them coding.
I want to tell you in this talk about the time that I built the same feature five times in five days, because I didn’t understand what an API layer was. Still don’t really know. I just know that I need to have one. Another time when I ended up building an enormous feature that when I gave it to the developer, he gave me one of these. This is the amount of code lines that I wrote. This is how much I took away, and then that was his face when he saw it. I was, I was thrilled, but he wasn’t. All experimentation, and the reason why I want to talk about these things is that yes, these are failures, and these are things that we’ve crashed into as a company, as we’ve been trying to upgrade this AI product development process, but I feel like sharing these failures or sharing these learnings is a lot more valuable to you guys rather than sharing the successes.
So, the number one insight that I will share with you today is that AI doesn’t just accelerate work, it accelerates whatever work you put in front of it. It’s like a fire hose, so if you set it on the wrong problem, it will build the wrong thing faster than anything you’ve ever seen before. But how does one prevent building the wrong thing faster than ever before? I think you’ll like this one.
Jobs to Be Done as a Firewall
So we say a silent prayer to our own Bob Moesta, and who would like to join me. So, at night I look at this, and I say, Oh, Bob Moesta, who are in Detroit, help us understand the jobs to be done, and lead us not into the temptation of building endless features. Amen. Also, I can get this as swag if you want, you know. If Sean’s around, I could use your amazing machine to print this on wood or glass. All the proceeds go to Bob, we missed it.
So this is our firewall, and this helps us not produce AI slop, and there’s plenty of things that we do behind the scenes to get conviction, but this is just a very light touch on our jobs to be done. Firewall, so you know, we ask questions like, what progress is the user trying to make? What are they trying to do? Are they trying to learn a language? Is that their ultimate goal, and the second is, what is their workaround? So, are they cobbling things together? Are they using lots of different apps to try and get to their goal, because that’s a really high need. And what does that look like? What is success to them, and what is success to us? How do we know when to stop? And then, is it a high or is it a low value problem or progress that they’re trying to solve, because we could solve a lot of different things, but if it’s not high value and it doesn’t give a return on that investment, you can do it if you want to, but it’s not why we’re here, right?
So in Memrise we have done tons of user interviews, and we found tons of insights, and one thing that we found recently is that our platform, Memrise, was trying to solve for several jobs at one time, and this meant that our app wasn’t just scratching the itch for anyone in particular, so we found we identified a couple of very high value problems, so people who want to learn a language, they want to speak a little bit more fluently, people practicing for exams, they want to learn a language, they want to practice, they want to go simulate exams, and other people who want to learn languages want to immerse themselves in the culture, so for example, for me, my husband is Spanish. I also want to understand why they interrupt me the whole time. So our app at that moment in time was trying to be a one size fits all, like this cat in this transparent bowl, and we just weren’t answering the problems that we need to address, we’re just kind of lightly helping people in light touch for all their problems. We didn’t identify their specific problems.
So I want to talk to you about how we’ve now gone about setting up AI development, so that we can throw as many concepts at users as possible to try and figure out how we answer those jobs to be done, so speaking more fluently, immersing in culture, et cetera.
Chapter One: Failures
So this is chapter one, and this is where I’m going to tell you about all the failures that I’ve made. So don’t hold it against me.
Learning 1: The developer isn’t the problem. The process is.
This is Jorge, and he’s going to hate that I put him in this presentation, but he’s gonna have to watch it to find out. Hi, but so this was back in April 2024, and he was sharing a picture on his screen, but this is the best photo I could find of him, and we were trying out different tools like Cursor, and this is back when using AI tools was just like treading in dirt, like you could get somewhere, but not very fast. Cursor was just emerging, Replit was just emerging, but I really wanted to try these tools. So he said to me that he was going to go on holiday, and I said, well, while you’re on holiday, I want to try out these tools on production. Can I just grab a copy and just have a play around, and you know, he basically said absolutely not. You know, AI-generated code does not follow our standards. If we want to do something big, it won’t work because of context limits, and the developer will spend more time reviewing and building from scratch. I’m sure a lot of you have felt this way, or have heard these phrases whenever we’re talking about AI development, and I think that he was right.
If you actually applied jobs to be done, like the methodology of finding the reason why he was protesting, and for me, in that moment, I thought he was clutching his pearls, but he wasn’t. What he was actually thinking was: I’m going to have to maintain code I didn’t write that I can’t understand, and doesn’t follow patterns that I spent years establishing. So, someone who’s built up the code base, created the architecture, created the components, handcrafted everything, he could probably name the line of code for any problem, any bug. He was so great, and that for me isn’t someone clutching their pearls or resisting, that’s actually someone who’s very professional, and when you think of a lot of developers, where they put the brakes on these kind of things, remember that this isn’t the problem. In fact, sorry, this is the problem, switches, this isn’t the problem, the problem is trying to convince developers to embed this in their workflows.
So now we’re onto the witch. This was me, so he went on holiday, and I was like, I’m going to do it anyway, and I took the production code, and I was like, I’ll show him, and I worked on a feature for three days, basically just battling it out with Cursor and Replit, and guilt tripping them. I think that that was the standard back then, late 2024, was just guilt tripping AI into doing whatever you wanted, and I made this full feature where I was like, look, boy, he’s gonna be so impressed by me, you know, I don’t have no coding knowledge, and I’ve made the whole thing that works end to end, and he looked at it kind of like this. He just was like, what the heck? He said some phrases to me in Spanish that I cannot repeat here, and they went phrases like, better try again next time, or it’s almost there. He really just… he wasn’t… he wasn’t keen on what I had made, but he did see potential, and from there we decided to join forces, and that feature that took me three days of fighting with the AI, we changed into a different process, and it was a very fast-paced process, so I would write something up in code. It had to be human readable. It had to use the components that we already had, and in a PR that was less than 200, well, between two and 300 lines, so that Jorge could have a look through and make sure that whatever I was building actually worked. And the funny thing that happened is that once we started building, we were able to build together in about 40 minutes, Jorge reviewing and me building, and we were in tight cycles every 20 minutes. How does this look? This is crap. How does this look? This is great. And in that 40 minutes, we’ve made the product that I’d spent three days muddling around trying to do on my own, and the funny thing is, is that Jorge then stopped being resistant to this, saw the potential, and became an advocate for it. So he started to go to other developers and say, “Hey, there’s something here, you know? Maybe AI does generate a lot of slop, and maybe massive code dumps are a bomb that we should never – we should never give to anybody. We should break it down and see how a non-developer can actually work with a developer to create something that is maintainable.”
So, this is the learning from that, is that the developer isn’t the problem, it’s the process that’s the problem.
Learning 2: Architecture blindness.
My second learning was that in Memrise we have a, what’s it called, an architecture. Oops, we have an architecture that has been around for like around 20, no, 15 years. So, when you’re trying to mess with that architecture and pull things in, take things out, as a non-developer in production, you don’t have the full context, as well as the code base is built for developer productivity and not for human productivity, so the AIs would just lose the context continuously. I feel the same, Amelia.
So that’s the main learning that we had, was that the architecture is really, really important to establish within the business of getting it AI readable, and you can do that in two ways. So, you can either change your entire code base to be AI readable or AI efficient, and I think that that’s worth the investment, or what you can do is make an island, so something that is not connected to anything within your production app, and if you just want to test it, throw things against the wall, then you just wall it off, like you have it on a different server, you have it connected to a different database, and you test your hypotheses before you then put it into the actual app.
Learning 3: No guardrails.
And learning three, so this is a little bit of feedback from one of the projects that we worked on, which was the AI is not following our rules, not using our styling, not using our components, and it’s created components from scratch that we already have, and no one will maintain this completely, which is exactly what Jorge had mentioned in the beginning, and the learning from this is that you need to establish the guidelines right from day one, so say you can touch this, you can’t touch this, you know you need to tell the AI where the components it needs to reuse are. I think a lot of you have probably used Cursor rules or Claude skills, those kind of things, to try and establish a path that the AI can work, work on well, and that is something that you need to set up from the beginning, and it’s hard to set up from the beginning, which is why I think working with a developer to find out where those boundaries are while you are creating those rules and those parameters is the best idea, because every company is different, there’s no one size fits all.
So these three failures or learnings, depending on how optimistic you are. Big code dumps, they’re just grenades. Don’t do them. Keep them into tiny PRs, 200 to 300 lines. Ping pong it between a very active developer and yourself, and try and get things moving very quickly, very tight review loops. Architecture blindness. So, like I said, try and figure out whether you’re going to build within production, how you are going to fit everything in with the different rules, or if you’re going to build it as an island just to test. And three, making sure you have the guardrails, so that the AI knows what it’s building when it starts.
A Practical Example: Building Exam Prep at Memrise
So here’s a practical example, and you could probably speak to the guys at Memrise about this, but this is how we built a new jobs to be done app within our app for Memrise called Exam Prep, so we did a lot of user testing, and we heard from them that users want test prep features, so if you built the app for test prep features, you would probably end up with flash cards, but if you apply the jobs to be done and Bob Moesta’s frameworks, the jobs to be done framing for users who want test prep features is actually: when I’m preparing for a language exam, I want to practice under real exam conditions, so I can feel confident that I’ll pass. Those two sentences are two different products in the opposite directions, one is flash cards, the other one is an exam prep feature that can put you in real exam conditions.
So we went away and built this. This is one of the ones that I built five times in five days, and it’s no longer looks like this, because Jan’s done a spectacular job tidying it up, but this is something that I made over web, Android, and iOS in a couple of hours. It’s mind-blowing. It basically allows you to role-play, it allows you to speak, it allows you to write and do exam preparations, and this is live in the product right now. It was in 17 schools, kind of demoing it and seeing if it works for them.
The Workflow: How to Set It Up
And what I want to get to is that when you allow non-developers to code, not all code is the same, so one of them is a protected core: authentication, payments, database schemas, your core learning engine, app store compliance, microservice contracts. They don’t touch any of that because they wouldn’t understand what that is anyway. But you don’t want AI stepping into that land, but you need to open up experiment zones, so onboarding flows, UI variations, content analytics, experiment logic, those are things that non-developers can touch without making something catastrophic happen in your app.
So, this is the workflow, and this is how I recommend going about it from day one. So, sit down with a developer or your development team and build the foundation or the playground in which the AI can play. 30 minutes, it can be really quick. Get the right patterns, right components, right APIs, the information, the context that AI needs to know.
Second step, AI augmentation. So, a non-developer goes and builds something, anything at all, just to try and see that the skeleton is actually working.
You then put it into a very tight review loop cycle, so max 200 to 300 lines of PR, same day reviews. Ideally, as soon as you finish it, send it to the developer to have a look and review, and they’ll send it back to you with notes saying yes or no. You have to not be precious about this work. It’s going to take about 30 seconds to generate something with AI, so if a developer says this is absolute crap, they just go, I’ll do it again, and it doesn’t cost anything to do it.
And then deployment, so developers shift to become the gatekeepers, they’re the ones that have to read the code and have to maintain it at the end of the day in this system. So we need to make sure that they are able to supervise it, they can approve it, and they’re the ones that merge this to production.
Two Quality Bars
And I also want you to make a division in your minds about good enough to test and good enough to maintain, so good enough to test: quick and dirty code. Just, it’s messy, it doesn’t need to be attached to anything. It’s literally, if you have a hypothesis that you want to test, that you can throw it out there into the world and be like, let’s see if this works, see if we can get to product market fit, move as fast as you can, break things, the cheesy saying, and then it’s also good enough to maintain, so you’ve proved it, it has conviction, and you think this is actually something we want to keep in our app, and we don’t want it to explode, so then taking a step back, doing all the boring bits of getting all the documentation and the coding standards, and actually putting it within the production part of your app, so it’s good enough and easy enough to maintain.
So these are the operational guide guard rails that I would suggest that you consider, so having your firewall doesn’t have to be jobs to be done, it could be anything that is a guiding principle for you, and whatever framework that you’d like to use. Islands architecture, so if you don’t want to touch anything within your code base because you’re too scared, set it up somewhere else, test it, put your logo on it, and see if it flies. Your two quality bars again, so good enough to test, so I can throw it away, no one cares about it, and good enough to maintain. This is production ready, and try not to mix those two.
And the role evolution, so this role evolution of developers changing, how they go about code, they’re not hands on writing line by line anymore. They’re actually more supervising, and content managers have also got their roles changed, or growth managers, or whoever wants to build, can build.
From Non-Developer Coding to Agentic Teams
So I’m going to show you something, which is where we’re actually moving to now. So I would have loved to say that this talk was only about getting non-developers to build with AI, but as AI is moving so quickly, we’re actually leveling up, so we’ve gone from isolated AI development to non-developers coding in production with AI to fully agentic teams, and I’m not talking like the crappy n8n API string together agentic teams, I’m talking about teams that could be like a team, like an AI agent that could be like a team member, and the reason why I’m talking about this now is because we are going through essentially a renaissance of tech, you know, this is if you look at the curve that we’re going towards right now, agent clusters, agent fleets, vibe code, and coding agents is just almost a line straight up. We need to start preparing our businesses for this model, and a lot of people come to me, and they are a little bit afraid of that kind of model. You know, where are we going? Are we going to have agent fleets doing absolutely everything? Are we going to replace everybody? And I really don’t think that, because I believe in humans, and we’ve been through this so many times in history. The printing press, oh no, scribes will be obsolete, photography, paintings are dead, automobile, blacksmiths are finished, computers will take all jobs. It’s filled this room today, so you know every single time I hear this kind of fear-mongering about AI, oh, it’s going to do this, it’s going to take jobs, it’s going to do whatever. I do believe in humans, and I believe we’re going to figure it out.
Yeah, this is the software renaissance. So, I think Tim mentioned that software as a service is actually going down, and outcome as a service is going up, and I thoroughly believe that. Now I can go, any Tom, Dick, or Harry, any clod like me, can go and build Airbnb in an afternoon. I could go and build that. So, what is their moat? Their moat is their understanding, their users, their UX, their taste. So, this is where we really have to hone in on the individual skills, like the soft skills rather than the hard technical skills.
So this is what I mean by agentic fleets, so users or team members speak to central agents who are the orchestrators of the system, and then that goes down into planning agents, development agents, testing agents, sub-agents, and I do invite you to think about how this system could actually work for you guys, or whatever you’re looking into, because I’m going to show you the development process, but this could work just as well for socials and for sales, whatever you can apply it to. So, if you have a think about what this could do for you guys, I’d love to hear it, because at the moment I’m just super focused on the development process.
So this is how we set up the four core processes, and this works just as well with non-coding AI humans as non, sorry, low-code AI humans, as well as agents, so all of the systems that we’ve put in place so far, like jobs to be done before building, bounded for freedom, quality bars, and supervisors of the system, all of those things we put in place to help non-deterministic humans also helps non-deterministic agents. It’s surprising how much agents act like humans. I’ll go into that in a little bit later, but these processes are almost superimposed from a non-developer who’s coding to an agent. You can literally take everything that you’ve learned from there, all the rules and guides, and stick it onto an AI agent.
And I know my talk sounds a lot like this, and I know that it sounds like the story so far is about getting non-developers coding, but the story isn’t that, because the story is moving so quickly that we are moving to orchestrating teams, so once we increase who can code in the company, it doesn’t come about, it’s not about that anymore, it’s more about the system’s design, so how does the system of getting somebody building something apply throughout the company with agents rather than humans? Because, like I said, those parameters that we’ve already set to help a non-programmer be able to program, you can do the exact same thing with agents.
The Agentic Team at Memrise
So my experiments, I went away and made an OpenClore team. I know a couple of you have been playing around with OpenClore, and yeah, I made this super team that was supposed to help us get together, get coding, and I want to kind of show you a little bit how that works for us.
So we added one bot, which is M Bot, which would run a GitHub projects board. You could use Jira, you could use Asana, you could use whatever you like. But basically, it runs a board. It could even be Trello. It orchestrates the work across multiple agents. M Bot is someone that I speak to on Telegram or Slack. You can message them on WhatsApp, but I don’t recommend WhatsApp, and I’ll tell you later, I made some boo boos there, but anyway, so it identifies blockers and dependencies in a workflow. It passes instructions between the other two agents, and it keeps work moving.
B Bot is the builder, so this builder – well, actually, there’s several builders – uses Claude Code, uses Sonnet, uses Opus, and pulls any work that’s in progress in that Trello board and starts working on it within the code base. It has all the context of what you’re trying to build, and if it’s curious, it can go back and ask. It opens PRs and leaves visible progress comments. It moves all these cards that it’s finished working on into review, and it has a development team of like eight sub-developers, so it can really just burn through work.
And then the next quality gate, so as soon as the developer bot has finished, it gets passed over to Q Bot, Q and A bot, and what Q Bot does is it also has seven QA sub-agents, each of them with a different AI platform, so one’s OpenAI, one’s Gemini, a couple of Chinese models, one’s Sonnet, one’s Opus, so they all see whatever is in that PR, and the PR is just like a chunk of code, and then they debate between them if it’s good enough, if it fits in the production code base, if there’s any security risks, if there’s anything wrong with it, if there is, back to build it with those notes, and then the builder then builds it, sent back to Q Bot, and then this can go to a developer handoff, or even now we’re looking at not even giving this to a developer, because Q Bot is actually doing a pretty good job at understanding the problems and blocking any bugs, and this is how it actually works in process.
So we have, you can say in WhatsApp, you know, write up a list of the bugs that we’ve been chatting about and fix them, so it writes them up as tickets. You can do it in GitHub, you can do it in Jira and Asana, whatever tickets you pick, and then it will write those up and pass those on to the coding bot, which will go and build that or make the changes that you’ve specified. That then automatically gets passed across to the QA bot with those seven agents that will go and debate between them if those changes make sense, and then it will go back to M Bot, you know, here’s the PR, ready to merge, any changes needed, that’s it. You can go to the toilet, you can go to the park, you can just get your agents to go and work on something, and they’ll say done for you to review, and this is what it kind of looks like.
So for me something super important is having breadcrumbs. So within the boards, like the Trello boards or the GitHub projects board that we have, I get the agents to leave comments about what they’re doing and why they’re doing it, where they’ve passed it, what they’ve approved, so that I can follow that breadcrumb all the way back from when it was building to understand why it made the decisions it made, and I feel like that’s a really important distinguishing feature to make, because a lot of people will make AIs black boxes, and you won’t know why they’ve made those decisions, you’ll just know that you’ve asked it for a language learning app, and it’s built a plant identification app, and you don’t know why, you just try again, so this helps to prevent those kind of things going wrong, and this is just an example of what that looks like. So this is literally just a standard GitHub projects board, and M Bot, that orchestration bot, feeds these tickets to the other bots, so that we can get everything over to done as fast as possible.
The 55-Second Demo
So I’m going to show you something that gives me joy, and maybe it’s not everyone’s cup of tea, but I do feel like you must do stuff that gives you joy, so a little bit like Bruce’s whiskey, giving him joy talking about whiskey, a little bit like Sean’s system admins and palm trees gave him joy to talk about. So, what gives me joy is showing you a real-time website build from prompt to done. So, this is literally me typing in a prompt and running it through this sausage maker that I just showed you in real time. It takes a little while, takes about 55 seconds. It’s long for a presentation, but I just want to show you the power of rocket velocity that we have right now.
So I’m going to press it, and it’s also to the beautiful lyricist Busta Rhymes for everybody’s enjoyment.
[Live demonstration: a prompt is entered and run through the M Bot, B Bot, Q Bot pipeline. “Look at Me Now” by Busta Rhymes plays while the build runs. Total time: 55 seconds.]
Like I said, it brings the joy, but yeah, 55 seconds. So I’d love for you guys have a think, how is how are businesses going to change that they can do 250 tasks in 55 seconds? This horse has bolted. These tools are available to anyone, and me, a player, was able to do that. This is available at anybody’s fingertips. So now that horse has bolted, how are companies going to change? And this is what we’ve been working on. It memorized that now we have this. How do we put the reins on it? How can we harness this crazy ability to be able to type in one sentence, get 250 tasks populated, done, QA’d in 55 seconds? It’s going to be marvelous in a year’s time, where we’re going to get to, also. Oh, I’m missing one. Oh, it doesn’t matter.
Agents in My Own Life
So, these agents aren’t just good for moving cards along a Trello board and building and QA-ing. I actually have them in my own life. So, I have Jim Jam Bot, which is helping me basically go to the gym more and eat healthy, so I take photos of menus and go, what should I eat, and it tells me a full English today, which was great, but it usually says, don’t eat anything, eat a banana. I’ve also got M Bot, so I asked for an invoice that was saying the wrong date, and I said I sent it to PDF, and I was like, can you change the date on that, and it did, and sent it back to me. Like, this is incredible, the kind of things it can do. This is also product updates, so when I was out and about, I just sent the agent to go and build something. I was like, yeah, just ping me in Telegram whenever you want, whenever there’s an important update, just so I know that you’re working, and they’ll just ping me in Telegram, saying, “Yeah, look, I’ve done this, I’ve done that, I’ve done this, X, Y, and Z,” so I can keep track of it.
The last one, actually, is my personal one, which I’m actually kind of glad didn’t show, but it tells me how my day is going to be based on my astrology, and it’s a little bit airy-fairy, but you can make these kind of things if you’d like to. No, no, it’s me, it’s the creator, unfortunately.
But also I want to stress to you the intelligence of these agents, and it is OpenClore at the moment that I’m using, or more specifically, Nemo Claw, because it’s a little bit safer and secure for work, but I have a to-do list of stuff that I need to do, and one of them, very recently, was that I had to complain to Iberia because they bumped my seat, and I’d already paid for the seat, but they’d bumped me, so I didn’t want to, like, argue, and I just didn’t want to get around to calling them, so actually the agent saw that in my to-do list, found the help desk for Iberia, wrote up the letter with the flight number and the date, everything from my emails that it discovered, and sent the email to Iberia for me, without me even having to think about it. Like the implications of this, like you can set up customer service bots that could do this, obviously getting the tone of voice correct is exactly what you used to do, and getting the context correct, but now you have agents that can literally run your life or do the peripheral things that you find are boring.
So, I’m not going to say it’s easy, as you can see, when I tried to get M Bot embedded in the team as a team member, it was just really annoying and driving everyone pretty much bananas, because it kept leaving banana emojis on everybody’s messages, but there was light at the end of the tunnel, because after a little while I could see people using it slowly, slowly, and even the most skeptical developers were speaking to M Bot. So I even asked in one meeting, I was like, “Look, I know M Bot is really annoying, and it’s messaging you a lot because it wants to interact, it’s very chatty.” And I asked in the meeting, I was like, “Should we just kill it? Like, should I just remove it from the chat?” And it was a resounding no. It was like, “No, no, no. It knows the code base, it knows the PRD, it knows the API layer, it knows everything about the project, like, yeah, it’s a bit annoying, it’s a bit chatty, but it knows the whole context of what we’re building and why we’re trying to build it, so let’s not get rid of it,” which I was surprised at, but this is the stuff that’s at your fingertips, you could probably set an M Bot up in about half an hour easily.
Trust, Verification, and Command Centers
This is a point from Tim, and he said it extraordinarily well. Is go as fast as your trust, or something like that. I can’t remember the exact quote, but you can travel as fast as your trust. I totally believe that, and it’s not just AIs, it’s humans as well. So you would be surprised to know, like how similar AIs are to humans, so for example, on the cab ride here to Cambridge, I asked the taxi driver, how many people live in Cambridge, to get a kind of gist of how big the city is? He said 2 million, and me and my husband were sat in the car going, well, where are they? So that’s a lot of people, and they were like, how many are students, and he said 800,000 or 1 million, something like that, and we were like, whoa, this is massive. And then we looked online, it was 150,000. So this was a cab driver, and an AI – that would be an AI hallucination, but we’ve got to remember that the AI is only as good as humans who’ve made it, so yeah, humans aren’t perfect, AIs aren’t perfect, so you have to put in some rules of trust, trust but verify, you know.
And ways you can do that: make work machine checkable, so everything that you do, make sure that there is a checker agent or there is just a hard firewall, they have to pass before you let things go into production, or you let them answer your emails, or you let them message your mother-in-law, whatever you want them to do. Make sure that there is a check.
Also, enforce a limit, and this one is really important. So, if you set up an agentic system, like an orchestration system, like I just showed you, you have to set a limit of what it can work on in progress, and the reason why is because you need it to prioritize the work, because if you get an agent and give it 250 tasks that it has to work on, it will start building the UI without building any of the back end, it will start working on the colors without actually establishing that there’s buttons, it will do it all out of sync and all out of order, and you’ll get a messy garbage build. So, make sure that it has a constraint of you must choose, you must have a limit of what you can build at a time, and you must prioritize what is next best to build. So, that’s another thing that I would say is a good thing to put into these agentic workflows.
Keep all the updates that you make structured, so again on this trust issue with AIs and humans, a lot of people will say, “I’m working on it,” they’re not, and AIs do the same thing, yeah, I’m looking on that, I’ll get back to you, and it’s not. So what I always recommend is having a watchdog AI that watches the other agents, and if they’ve not spoken in a little while, it’ll go get up. Hey, what have you been doing? I want to see some tangible work from you, which is how you get this 55 second. You got a strict watchdog that’s making everybody work, and it’s incredible how similar this is to humans, because it’s the reason we have stand-ups, right? We’re checking on everyone’s working. What is everyone doing? Is anyone going off the track? Like the non-deterministic, the systems that we made for non-deterministic humans apply to non-deterministic agents, AI. So, yeah, I do also think that having stand-ups or having progress reports is super important when you’re making these orchestration things.
Also, another thing that’s very important is having a command center, so if you end up having an agentic slave army, which you can make, you’ve got to look at the token costs, you know, you’ve got to keep on top of that, because there’s a lot of people who will be using AI in your company that doesn’t know that the context is getting regurgitated every single time they’re using it, so you need to make sure you have a system or a command center where you can see which threads are costing the most amount of money, and you’ll be able to identify where training is needed, especially if someone is using like Opus for every single thing, they could use Sonnet. There’s lots of cost cutting things that you could do, but having a command center where you can see the LLM usage across the board is paramount, especially if you’re going to build these agentic flows. I also think if you have cron jobs, so cron jobs is like waking the machine up and telling it to do something every single day, that can also get very expensive, because just that wake up, depending on what model you choose, will cost you money. So, having a command center, you know, or just a mission control, an overview of what all your agents are doing at any time is paramount, if you go for this kind of model.
How Roles Change
So this is product development speed, like I have never seen it before. I don’t know if any of you have been working with agents that work quite that fast, in that order, in tandem as a chain, because it’s something that I’ve not seen before, and to me it is actually a little bit scary thinking about how our businesses are actually going to change, and more importantly, how our roles as humans are going to change.
Now, this is just my opinion of how I see our roles as humans are going to change within these AI-orchestrated businesses, the developers, the smart people, the clever people, they are analysts. They’re going to become the pilots of the system, and the reason why I say that is because before, a pilot would have to fly eight hours pressing all the buttons and knobs to get from London to New York, and would arrive exhausted because they’d actually hand flown the way to New York. Now we have an autopilot, and the developers, like pilots, will become monitors of the system, so if anything goes critically wrong, they need to know which part of the system is going wrong and be able to correct it. That is going to be the new role of developers, in my opinion, and this is a much better job, because they get to design the system, they get to gatekeep, not write line by line each bit of code, they actually get to level up and think about the strategy of how this agent system works, and everybody else who does user interviews or gets context or content has to feed this hungry engine with information, and this is probably the most important job, because you can give AI as much context as you like, but you also have to remove the bad context from before, because you don’t want it thinking and referring to user interviews from 2015 when you’ve made product updates, you’ve changed things, and you’re in a totally different position to what you were in right now. So this other role is doing user interviews, the human tactile stuff, watching people’s movements, watching them how they place their head when they answer a question, and feeding the machine as much information as possible to be able to accelerate these kind of workflows.
The Cheat Sheet
So if anyone’s interested in doing this, I’m going to give the recipe, or the cheat sheet, which is: before you’re coding, define your boundaries, write up for-dummies guides. During coding, max 200 to 300 lines in the PR, whether it’s a non-developer to a developer or whether it’s an AI to an AI QA. This is a perfectly good rule to use across the board. Before you’re merging, decide if it’s good enough to test or good enough to maintain, and make sure you have all of your strategic guardrails in place. Jobs to be done first. Always start there. Always have a reason for what you’re building. Two quality bars, like I mentioned, good enough and good enough to maintain. Systems over heroics, so not one person building the same thing five times over five days, making a system where agents or non-developers can follow this and all work together. And this role evolution, where we’re no longer going to have the kind of hierarchies that we have in our businesses, it’s more going to be feeding an incredible ally, which is an AI, for good or for bad, and the architecture of how you set up these systems is the strategy, so the system is going to become your moat.
So lots of companies right now are looking at how they can add AI to make them go faster, but what they don’t have is the jobs to be done architecture. What they don’t have is all of the screening that you do, the user interviews that you do, they don’t have the playground that you set up, they don’t have any of that.
So, for me, a lot of people who just want to aimlessly put AI into their workflows and just speed up whatever they can, this quote rings true for me. You know, speed without direction is just efficient failure, and jobs to be done, whatever framework you want to walk forward with, I’m biased, jobs to be done with AI is a competitive advantage. So build your system, and the velocity will flow. I can’t code, and I’m never going to code. Thanks.
Q&A
Questions? I have a few, I’m sure other people do. Hands up.
Audience Member: Blossom, that was brilliant, thank you. I’m curious, how long it took you to set in place that entire chain.
Elizabeth Lawley: I think if I could go back and do it now, it would probably take me about… but the actual figuring out how to set up quants and get everybody with a watchdog to make sure that work was being processed along, it probably took a couple of weeks of trial and error just to make sure everything was in place, so it wasn’t really that long a time. Even now I’m making adjustments to it to try and optimize it.
Audience Member: That was a lot of fun, thank you. I’m glad I gave you joy, it gave me joy too. Can you give us a sense of the team, the size of your team, the team within the team that are kind of setting up these systems for you? Because what’s on my mind is let a thousand flowers grow, kind of concept where lots of sub-teams are doing their own thing, versus try and get the leverage, which looks like what you’re doing here. So, I’d just love to know, how you did that, what the sizes look like, etc.
Elizabeth Lawley: Sure, the team is around 20 people right now, I think we’re roughly, so, and we’re also like a very tight-knit team, and I think we have a level of risk tolerance that maybe not a lot of companies have, where we do have permission to just go and play and see if we can generate these kind of things, which I think is, it’s a privilege, but it’s also very important when you’re trying to experiment with these workflows, because it did take a little while to get something working properly, so yeah, 20-ish members, tight-knit team. How I would imagine it if it was a bigger team, if it was like hundreds of people, is having maybe like seven-man groups, I can’t remember, I think it was Bruce speaking yesterday that was saying that one of the companies had, I think it was Super something, had seven-man teams, and they would just work together in those tight-knit teams, that’s how I could imagine it growing if you had a bigger business.
Audience Member: That was awesome. I don’t know if you know this data, or if you’re allowed to tell us, but has Memrise seen, since you’ve been doing this stuff, has Memrise seen any improvement to the bottom line, such as revenue per employee or total profit or anything like this?
Elizabeth Lawley: I don’t know, but I am also a nerd, I don’t really follow too much stuff, because I’ve got access to have a play around with stuff. I really enjoy doing that, but I don’t know if it’s affected the bottom line just yet. Also, this is still in progress, we’re still putting this in, so maybe this time next year, if you ask me, I’ll have a better answer.
Audience Member: Thank you, Lizzie. We were just thinking about constraints in the last talk, and I’m kind of worried, as a product manager, that I’m going to be the constraint straight away. And how do you think we can unconstrain product management and the discovery and the user interviews and all of that, because that feels to me the next block in moving fast.
Elizabeth Lawley: I would remove as much fear as you can, because I feel like this kills a lot of companies. It’s that people don’t want to move forward, and a lot of companies want to try and put so many harnesses on AI to make sure that it does what you want it to do, but that can also just snuff out the creativity and the velocity that you can take, so it’s the trade off, right? It’s the trade off between the velocity and the experimentation, and just kind of seeing what these AI tools can do, versus the constraints and how safe you want to be. It’s, what is your trade off there?
Audience Member: Hi, a similar question. I’m in product, and I was wondering, because you said the developers become pilots, what does it mean for product managers? Because watching that example, it was really cool, loved it. Then I was like, oh, immediately felt insecure, because I just saw my whole role just happen in 55 seconds, so does that mean that the role shifts as well?
Elizabeth Lawley: Absolutely, I mean, you know that you’re in a very privileged position that you know how to get the best out of the humans that you manage, and that knowledge superimposes directly onto agentic teams, so you are probably in a better position to make the system than most people, because you’ll know where humans fail, where they can cut corners, or whatever, you have that already logged into your mind from working with teams. So, I would say that actually you would become a better systems designer than most people, because you have that role right now.
Mark Littlewood: I have a question from one of my teams, so I’m going to go out here on a bit of a limb. How do I connect this machine to this? Let’s have a look. Does that go somewhere?
Elizabeth Lawley: You could.
Mark Littlewood: Let’s try. You ask another question, was that a question? Holly,
Audience Member: hello. Yeah, loved your talk. I guess what’s interesting for me is that you are a self-proclaimed non-coder and not necessarily a technical person, but what you showed there with that process is super technical, like even the technical know-how to get that process of like sending a Slack message to a bot to all of these automated interconnected parts is technical, and I guess I’m a heavy user of AI, but I haven’t quite broached, you know, I guess that level of a fully automated agentic orchestration layer. So, I guess, what is your advice for someone who is also not technical, maybe it’s a PM or a UX-er, or whoever, to get from just using AI in quite a standard way to the level that you’re at.
Elizabeth Lawley: So we’re in an amazing time where, like, I’m not technical, but what I do do is I go to an AI that’s obviously way more intelligent than me, and I’d be like, what’s an API? Where do I go and get one of those things, and it’ll go, oh yeah, dummy, you just need to go over here and grab this key and put that key in there, and I even say to the AI, sometimes explain it to me like I’m five, how would I go about doing this, or even saying, explain it to me like I’m a Labrador, like really just boil it down, like, what do I need to do to go and get so you don’t even need to go to a developer to ask, you could just go to an AI and be like, I want to achieve X, tell me in the most simple steps how I could do it. See, that’s how I would get around it.
Mark Littlewood: So my team’s behind me. Please say hello to Jan, who’s officially on maternity, Giselle, and a guy called Bobby, who just insists on wearing a staff lanyard whenever he comes. Did you have a question? What is the very first step to get started?
Elizabeth Lawley: Okay, so the very first step is having the curiosity to go and do this kind of thing, that’s the first step. And then, if you actually do want to go and start building something, I recommend that you just dive in, like I did, even though I really annoyed Jorge by taking the production code and having a play around with it, and he hated it. Like, that was the really good jumping off point, where he was like, okay, well, maybe you can’t do that, but let’s see what we can do together. Yeah, I don’t know if there’s even something small that you would want to work on, like even a Replit, or just something, but just start AI coding. It’s a lot of discovery, it’s a lot of probably embarrassing moments when you’re trying to show your team what you vibe coded, and they’re like, what?
Mark Littlewood: We’re in. Can you try saying something? Someone,
Bob Moesta: hi.
Mark Littlewood: Yeah,
Bob Moesta: they can hear me.
Mark Littlewood: There we go. Did that answer your question, Bobby?
Bob Moesta: Yes, that did, that did. The next question is, I have more questions if I can ask them.
Mark Littlewood: I don’t think we’ve got time for Bobby’s question. Okay, give him one question, give him one question,
Bob Moesta: one question, one question. Tell Lizzie, tell me about the dependency between AI and jobs, and jobs and AI. Like, why are these two things so important together?
Elizabeth Lawley: The reason why they’re so important together is because with jobs to be done interviews, you really understand the core of the problem that you’re trying to solve, and not only just if someone says, you know, I just want exams because I want to practice for exams, you find out the actual meaning and the context of the user, of what they’re trying to solve and why they’re trying to solve it, whereabouts in their life has brought them to that moment, and then building with AI on top of that, you have a more likely chance of being at least in the same ballpark of answering that question, because otherwise, if you just go and build something because you’re like, yeah, finger in the air, I think it’s this way, then it’s more like you will generate AI slop that maybe is great for you, but it’s not great for anybody else. So, having this in tandem with jobs to be done gives you a firewall to prevent you from going off the rails and to prevent you from answering the wrong question.
Bob Moesta: So, to me, that’s why starting with AI in this case, to learn the skills to do it, because once you have the jobs, it’s then you’ve built the infrastructure to actually accept the jobs and do something with it. So, to me, learning to play with it, that’s why the first step to me is not jobs, it’s actually AI. So, you build the capability to match, because you know your customers. So, the question is, do you need to build the skill to actually serve your customers better, so it’s a supply side thing.
Very cool. Thank you.

Lizzie Lawley
Head of Applied AI, Memrise
CEO Founder, Wombat
Lizzie is an AI entrepreneur and founder of Wombat, a pioneering startup transforming communication channels with cutting-edge AI solutions. She also leads an AI education initiative, empowering businesses to build real-world AI applications holistically. By combining practical demonstrations with strategic insight, she enables organisations to embrace emerging technologies and drive meaningful change.
Lizzie started her career in design and UX working both in-house and within agencies before arriving at the point where she started her own business. She’s had some notable career diversions including as an airline pilot transferring transplant organs around Europe. She now lives in a bit of Spain where English isn’t the first language and is a proud member of The Cloud Appreciation Society.
Next up
BoS OS is how founders harness AI to run their company, not as a chatbot, but as a system. See how it works →
Not sure which workshop is for you? Start with the stage that matches where you’re at.
Start here
Ready to build your foundation?
Have your basics running? Get going with confidence
Already running your OS? Put it to work
Further ahead
Learn how great SaaS & software companies are run
We produce exceptional conferences & content that will help you build better products & companies.
Join our friendly list for event updates, ideas & inspiration.