That's fine when one person is working on the service. It's less great when several engineers are trying to make changes at the same time and end up waiting on each other.
And this isn't something we're building for one team. We're potentially solving this for tens of thousands of developers, so the scale is very different from fixing it for one service.
So we wanted to move that work into a developer's own isolated environment instead.
There was one problem.
The long-term version is supposed to be event driven. The deployment finishes, another system tells us about it, and our workflow continues.
That event doesn't exist yet.
So for now, we have to poll.
Start the work, wait a bit, check whether it's done, and repeat until we can move on.
This is the kind of problem where you can pretty quickly end up adding schedules, multiple Lambdas, Step Functions, EventBridge, or some combination of all of them.
Except we'd already solved almost exactly the same problem before.
With one Lambda.
We tried this once already
The first time we used Durable Lambda, it was still pretty new.
We needed to call another system to start a long-running operation, periodically check its status, and then send an event back into our workflow once it either completed or failed.
Durable Lambda looked dramatically simpler than the alternatives, so we decided to try it.
It was also a good excuse to actually use something new. AWS releases new stuff all the time, but you don't always have a real problem where it makes sense to try it.
This time we did.
And it worked.
We haven't really had any problems with that Lambda since.
So when I ran into another polling problem for this new feature, I didn't spend much time debating the architecture.
We've done this already. Use another Durable Lambda.
I really just needed it to sleep
What I need from this workflow is pretty boring.
It does some work, waits, wakes up and checks the status, then either waits again or moves to the next step.
The useful part is that "wait" doesn't mean keeping a Lambda running and paying for it to sit there doing nothing.
A durable wait suspends the execution. For an on-demand Lambda, AWS doesn't charge compute while it's suspended. There are still charges for durable operations and the checkpoint data AWS stores, but you're not paying Lambda compute for all of that idle time.
The other thing I like is that the workflow stays in Java.
You get a durable context and define the steps and waits in the code. AWS checkpoints the completed steps underneath, so when the execution wakes up again it doesn't have to redo work it already finished.
If you already understand the feature, you can look at the main handler and get a pretty good map of the workflow from the sequence of steps.
I've reviewed code using Step Functions before. Maybe I'm remembering it unfairly, but looking through the state machine definitions never felt as easy to me as reading normal Java and seeing what happens next.
Maybe an AI agent doesn't care.
I do.
I'm still the one on call
AI can generate either approach just fine, which has changed what I care about when I'm making these decisions.
I'm much less worried about which architecture takes less code to write. The agent is writing the code anyway.
For this feature, I had AI implement one step at a time. Each change was small enough that I could review it, merge it, and move on to the next one.
What matters more to me is what I'm left supporting afterward.
If this breaks months from now, I want to open one place and see the sequence of operations. Then I can look at the logs, figure out which steps executed, which didn't, and start looking around the place where things stopped.
I still go on-call. I still need to understand how the application works.
Making something easier for an agent to generate doesn't help much if I make it harder for myself to debug later.
We've already gotten the new flow through an end-to-end demo. The feature isn't completely finished yet, and the demo exposed some UX work around making it clearer what is happening while the workflow runs.
But the orchestration itself hasn't been the complicated part.
It's basically one Lambda with some extra durability configuration. The steps stay in the code, it can sleep without sitting there burning Lambda compute, and I don't need extra infrastructure just to wake something up every few minutes and ask whether the previous thing finished.
I'm not saying Durable Lambda replaces every workflow tool. I haven't used Step Functions enough myself to make that argument.
But for this kind of polling problem, Durable Lambda is becoming my default.
If one durable function can express the workflow clearly, I'm going to start there before inventing more infrastructure.
Cheers!
Evgeny Urubkov (@codevev)