I Thought My Codex Sessions Got Corrupted


Monday morning, I logged into my work computer, did the stupid auth I have to do on both my Mac and remote desktop, and resumed a few Codex sessions after my cloud desktop restarted over the weekend.

I’d been using Astra for the last 3-4 weeks without much trouble. But that morning Codex barely responded. Even a fresh session with just hello sat there for about 45 minutes before I killed it.


So I blamed an outage

My first assumption was that Bedrock was having a bad day, since our internal AI usage is all routed through it. I asked around in Slack whether anyone else was seeing the same thing. A couple of people reported no issues.

Then I asked which models they were using, and a couple people said Sol.

Except when I switched the model in my existing session to Sol, it immediately gave me an error about not being able to decrypt something.

That sounded a lot worse than a slow model. My cloud desktop had just restarted, and now Codex was telling me it couldn’t decrypt an existing session. For a minute, I thought the sessions I’d resumed had gotten corrupted. I usually keep a session around until I’m done with whatever I’m working on, so I really didn’t want to lose them.

So I gave up on Sol. I tried Terra instead. Terra actually worked.

At this point I had two different problems. Astra was taking forever to respond, while Sol wouldn’t even load some of my existing sessions.

I still assumed the first problem was some temporary service issue. Astra worked again the next day, so whatever caused the slowdown seemed to have resolved itself.


Then I noticed the pattern

Sessions that had started on Astra could resume on Astra. They could also resume on Terra. But trying to resume the same sessions with Sol failed immediately.

Fresh sessions with Sol were fine.

That made it pretty clear that whatever was happening with Astra’s speed wasn’t causing this particular problem. Something about the existing session mattered.

So I did what any reasonable engineer does after spending enough time being confused and asked Claude to figure it out. I could have asked Codex but at that time I didn’t know if I could trust it anymore.

Apparently, the answer was AWS Regions.

It always is, isn’t it?

For the in-region Bedrock endpoints I was using, Astra runs in us-west-2 (Oregon). Terra is also available in us-west-2. Sol runs in us-east-2 (Ohio).

Bedrock can also keep conversation state in the Region that served it. So when I started a session with Astra and later tried to resume it with Sol, I was effectively trying to continue encrypted conversation state in another AWS Region.

AWS did not appreciate that.

While looking into this later, I even found a Codex issue filed the same day by someone getting essentially the exact error:

Encrypted content cannot be used in a different region from the one that created it.

So at least I wasn’t the only one confused by this.


This feels very AWS

If you’ve used AWS for long enough, you’ve probably done this before.

You log into the console to check logs or just one of your services. But there’s nothing there. You refresh, search again, maybe panic that someone deleted production, and then after nothing making sense on that screen you finally notice that the region picker says us-east-2 when everything you care about is in us-west-2.

Phew.

Apparently AI conversations can now give you a version of the same experience.

I’d been treating a Codex session as just a Codex session. If I wanted to switch models, I assumed I could change the model and keep going.

Turns out there’s more infrastructure hiding underneath that abstraction than I realized. The model you choose can affect where the request runs, and once encrypted session state gets involved, changing models can mean crossing a Region boundary the existing conversation can’t cross.

At least now I know that if an old Codex session suddenly refuses to resume after switching models, starting over probably isn’t the first thing I should do. I should switch back to the model that created the session, or another model running in the same Region. If I really want to use a model in another Region, then I’ll need a new session.

Basically the same debugging step AWS has trained us to do for years.

Check the region.

Cheers!

Evgeny Urubkov (@codevev)

600 1st Ave, Ste 330 PMB 92768, Seattle, WA 98104-2246
​Unsubscribe · Preferences​

codevev

Tools and workflows that survive real work. One short Friday email about better software work: tools I keep, workflow experiments, and AI habits I am testing in public.

Read more from codevev

A few weeks ago I was designing a feature to fix an annoying problem in our deployment workflow. Today, if a developer wants to validate a change before merging it, that work runs inside a shared deployment pipeline. While it's running, it occupies one of the pipeline stages. That's fine when one person is working on the service. It's less great when several engineers are trying to make changes at the same time and end up waiting on each other. And this isn't something we're building for one...

It’s been just over a year since I had ACL surgery that put me out for about 3 months. I still remember all the pain associated with it, and I am still at only about 80% of where I thought I’d be. I also remember that before I left on medical leave, I had started using AI a lot more for all the coding work. I’d still open VS Code and use Cline because it felt cool seeing how AI made changes in your IDE. It was working great on some smaller tasks and much worse on anything a little bit...

After looking at the fourth PR for our service from someone we’d never even talked to and seeing the code comments that looked like Christmas trees, I thought that we have to create some rules so we don’t have to even bother reviewing the AI-generated slop nobody bothered reviewing before sending it to us. I posted a question in our team channel, and everyone agreed with my proposition of drafting an “away-team” guide. Ironically, I did use AI for that. However, I mostly used it for two...