This simple Codex change kept me from spending another $100 a month

I was on the brink of having to spend $100 more per month for AI because I was running out of tokens faster than they were replenishing.Little did I know a single change in how I used AI would make my tokens (and money) go further.I thought a long-running thread was the best way to get consistent results It's the best way to use tokens, at least Close For basically the last year and a half that I've been using AI heavily in various ways, I've always kept long-running threads to keep consistent output.

I would pin a thread that I liked the output of, rename it to something that went along with the content, and then I would just stick to using it.In some scenarios, this is fine.For instance, within the ChatGPT app, there are very generous usage limits, if they ever even kick in at all.

So, a long-running thread there is fine.Gemini is quite the same.Claude and Codex, however, apply usage limits across the board.

While I achieving my goal of getting a consistent output using a long-running thread, when I started doing this in Codex, I found something out: all that context took up a of tokens.You see, whenever you send a message in a thread, it's not just sending only that message.The AI doesn't actually know what the context of the conversation is, it only knows the exact prompt that is being sent.

So, to offer context, when you click "send" on a message, it sends the message with the context of that thread to the AI.This allows the AI to not just respond to the exact thing that you asked, like "Do you think this will work?" But, instead, it offers the context, so the AI knows what "this" is.Imagine having a conversation with someone where there was zero actual memory, and everything had to be restated—that's AI.

Related Stop wasting Claude's quota on long conversations—using projects can fix it Use your quota more efficiently.Posts 3 By  Adam Davidson The problem is, context takes up a of tokens when it starts to grow.If you have a thread that's been used for months, that's a lot of context to pack into a prompt to send back.

Lots of context equals high token use, and the more tokens you use, the faster you hit the limits.I found this out the hard way.I knew there had to be a better way, as I kept running up against my usage limits.

That's when a friend of mine told me to start making custom skills, which I had originally avoided.When I told him why I had avoided skills, he gave me the single best piece of advice about AI I have ever received—let the AI build the skills.I didn't realize how simple it was to build a custom skill I thought you had to hand-code it, but Codex can actually do it for you So, the reason I never used skills before is because I thought every skill had to be hand-written.

I don't know I thought that, but I did.So, I only used skills other people wrote.The advice my friend gave me, though, changed the game entirely.

His specific advice was to go into the long-running threads that I had going, and tell Codex to create a skill based on what we had done.Or, if I started a new thread and got an output I liked, tell it to turn into skill.So, I started doing that.

I started having Codex turn my long-running threads into skills, and it was a complete game changer for my AI usage.Using custom skills gives me more consistent results with less context usage I use far fewer tokens this way Now, I create skills for everything.If it's something I plan to do more than once, I make a skill for it.

I do this because I actually get way better and more consistent output using skills compared to when I just had long-running threads.This is because long-running threads are typically compacted by the AI platform when they get too big.You can only send a message with so many characters in it to an AI, so longer conversations get summarized and compacted before being sent.

That means a conversation can start to wander if the compaction and summarization isn't perfectly accurate.With a skill, however, the exact same set of instructions is sent every single time in a fresh thread.Using a skill means that you're only sending the absolutely necessary information to the AI in the most compact way possible.

So, using skills not only gives a more consistent output, but also uses way fewer tokens.I saw my token use drop by a good 30 to 40%, maybe more, when I started using skills routinely.This actively saved me money, as I was contemplating upgrading from the $100 ChatGPT Pro plan to the $200 ChatGPT Pro plan.

Moving to skills kept me from having to upgrade, and I'm so glad that my friend told me about it.Now, I'm telling you.AI usage is going to continue to morph and shift over the years When I first started using AI just 18 months ago, skills weren't nearly as big as they are now.

In fact, not all platforms supported custom skills.However, in mid 2026, skills are all the rage.I have no idea where AI will be 18 months from , but I do know it's going to continue to morph, adapt, and change.

My usage of AI will also have to morph, adapt, and change if I want to keep up with the times.Right now, I'm glad I'm using skills.Six months from now, I might not even be thinking about skills.

AI is a fast-paced wild west right now, and I'm all for it.

Read More
Related Posts