I Needed a Lambda That Could Sleep


A few weeks ago I was designing a feature to fix an annoying problem in our deployment workflow.

Today, if a developer wants to validate a change before merging it, that work runs inside a shared deployment pipeline. While it's running, it occupies one of the pipeline stages.

That's fine when one person is working on the service. It's less great when several engineers are trying to make changes at the same time and end up waiting on each other.

And this isn't something we're building for one team. We're potentially solving this for tens of thousands of developers, so the scale is very different from fixing it for one service.

So we wanted to move that work into a developer's own isolated environment instead.

There was one problem.

The long-term version is supposed to be event driven. The deployment finishes, another system tells us about it, and our workflow continues.

That event doesn't exist yet.

So for now, we have to poll.

Start the work, wait a bit, check whether it's done, and repeat until we can move on.

This is the kind of problem where you can pretty quickly end up adding schedules, multiple Lambdas, Step Functions, EventBridge, or some combination of all of them.

Except we'd already solved almost exactly the same problem before.

With one Lambda.


We tried this once already

The first time we used Durable Lambda, it was still pretty new.

We needed to call another system to start a long-running operation, periodically check its status, and then send an event back into our workflow once it either completed or failed.

Durable Lambda looked dramatically simpler than the alternatives, so we decided to try it.

It was also a good excuse to actually use something new. AWS releases new stuff all the time, but you don't always have a real problem where it makes sense to try it.

This time we did.

And it worked.

We haven't really had any problems with that Lambda since.

So when I ran into another polling problem for this new feature, I didn't spend much time debating the architecture.

We've done this already. Use another Durable Lambda.


I really just needed it to sleep

What I need from this workflow is pretty boring.

It does some work, waits, wakes up and checks the status, then either waits again or moves to the next step.

The useful part is that "wait" doesn't mean keeping a Lambda running and paying for it to sit there doing nothing.

A durable wait suspends the execution. For an on-demand Lambda, AWS doesn't charge compute while it's suspended. There are still charges for durable operations and the checkpoint data AWS stores, but you're not paying Lambda compute for all of that idle time.

The other thing I like is that the workflow stays in Java.

You get a durable context and define the steps and waits in the code. AWS checkpoints the completed steps underneath, so when the execution wakes up again it doesn't have to redo work it already finished.

If you already understand the feature, you can look at the main handler and get a pretty good map of the workflow from the sequence of steps.

I've reviewed code using Step Functions before. Maybe I'm remembering it unfairly, but looking through the state machine definitions never felt as easy to me as reading normal Java and seeing what happens next.

Maybe an AI agent doesn't care.

I do.


I'm still the one on call

AI can generate either approach just fine, which has changed what I care about when I'm making these decisions.

I'm much less worried about which architecture takes less code to write. The agent is writing the code anyway.

For this feature, I had AI implement one step at a time. Each change was small enough that I could review it, merge it, and move on to the next one.

What matters more to me is what I'm left supporting afterward.

If this breaks months from now, I want to open one place and see the sequence of operations. Then I can look at the logs, figure out which steps executed, which didn't, and start looking around the place where things stopped.

I still go on-call. I still need to understand how the application works.

Making something easier for an agent to generate doesn't help much if I make it harder for myself to debug later.


We've already gotten the new flow through an end-to-end demo. The feature isn't completely finished yet, and the demo exposed some UX work around making it clearer what is happening while the workflow runs.

But the orchestration itself hasn't been the complicated part.

It's basically one Lambda with some extra durability configuration. The steps stay in the code, it can sleep without sitting there burning Lambda compute, and I don't need extra infrastructure just to wake something up every few minutes and ask whether the previous thing finished.

I'm not saying Durable Lambda replaces every workflow tool. I haven't used Step Functions enough myself to make that argument.

But for this kind of polling problem, Durable Lambda is becoming my default.

If one durable function can express the workflow clearly, I'm going to start there before inventing more infrastructure.

Cheers!

Evgeny Urubkov (@codevev)

600 1st Ave, Ste 330 PMB 92768, Seattle, WA 98104-2246
​Unsubscribe · Preferences​

codevev

Tools and workflows that survive real work. One short Friday email about better software work: tools I keep, workflow experiments, and AI habits I am testing in public.

Read more from codevev

Monday morning, I logged into my work computer, did the stupid auth I have to do on both my Mac and remote desktop, and resumed a few Codex sessions after my cloud desktop restarted over the weekend. I’d been using Astra for the last 3-4 weeks without much trouble. But that morning Codex barely responded. Even a fresh session with just hello sat there for about 45 minutes before I killed it. So I blamed an outage My first assumption was that Bedrock was having a bad day, since our internal AI...

It’s been just over a year since I had ACL surgery that put me out for about 3 months. I still remember all the pain associated with it, and I am still at only about 80% of where I thought I’d be. I also remember that before I left on medical leave, I had started using AI a lot more for all the coding work. I’d still open VS Code and use Cline because it felt cool seeing how AI made changes in your IDE. It was working great on some smaller tasks and much worse on anything a little bit...

After looking at the fourth PR for our service from someone we’d never even talked to and seeing the code comments that looked like Christmas trees, I thought that we have to create some rules so we don’t have to even bother reviewing the AI-generated slop nobody bothered reviewing before sending it to us. I posted a question in our team channel, and everyone agreed with my proposition of drafting an “away-team” guide. Ironically, I did use AI for that. However, I mostly used it for two...