Blog / AI ROI

Token Maxxing: Don't Let Your AI Tokens Expire Unused

By Saurav Sharma||Updated September 22, 2026|9 min read

You pay a flat monthly fee for Codex, ChatGPT, Claude Code, Claude, Gemini or Cursor, and the tokens do not roll over. When a reset lands, whatever you did not use is gone, even though you already paid for it. I hit this on my own Codex plan: the weekly reset was two days away and I still had most of my limit left. So I collected the prompts I run in that situation and put them in a small open repo. This post holds every one of them in full, so you can copy them here without a GitHub account.

The repo: github.com/ravsau/token-maxxing-prompts (MIT licensed, works with Codex CLI, the ChatGPT app, Claude Code and any coding agent). If you have banked limit resets and the opposite problem, the videos on redeeming them are on the CloudYeti channel: Codex CLI and the ChatGPT app.

The one rule

Only spend leftover tokens on work that leaves a durable artifact: something committed, documented, decided, or reusable after the reset. A good use of spare tokens takes information you already have, synthesises it, and ends with an output you keep or a decision you make. A bad use produces more words, more brainstorming, and documents nobody opens.

"Tokenmaxxing" gets used two ways now. One is a status game where people burn tokens to prove they are serious about AI, leaderboards included. That is not this. Here it only means: you already paid for capacity that is about to expire, so point it at something you will still have next week.

Start here: the meta prompt

The seven prompts further down each do one job. This one asks the model to look at your actual world and tell you which jobs are worth the tokens. It works best in the tool that already holds your context: ChatGPT web with your chat history, Claude with memory, or a local agent sitting in your project folder. In a fresh chat with no history you get generic output. Run it there, then hand the ideas to your coding agent and build them before the reset.

I have a lot of AI tokens left before my usage resets, and I want to spend them on
high-leverage work instead of wasting them.

Based on what you know about me, my recent work, projects, notes, files, logs and
goals, give me 10-15 substantial tasks that would genuinely benefit from a lot of
tokens and a large context window.

Think along these lines:
  - mine old or unfinished projects for valuable ideas or assets
  - analyse past work or content and find patterns, opportunities, things worth reviving
  - evaluate whether an existing workflow or "factory" actually produces good results,
    before I optimise it
  - audit my recurring systems, agents, automations and processes; find what is broken,
    duplicated, or produces output nobody uses
  - synthesise scattered notes and logs into reusable knowledge, decisions, or
    source-of-truth documents
  - find recurring mistakes, bottlenecks, and things I keep relearning
  - stress-test an important project or plan
  - turn repeated manual work into a system
  - find contradictions or neglected opportunities across everything I am working on
  - take something partly built and do the heavy work needed to get it near finished

Do not give me generic productivity tasks or random brainstorming. Favour work where
you can read a lot, compare a lot, synthesise deeply, and leave me with something
reusable or actionable.

For each idea, say briefly what you would actually do and what output I would get.

Prioritise ideas that turn things I already have into more leverage.

If the short version returns shallow ideas, use the longer one below. It asks for the answer in a structure you can act on directly.

You have a large context window and a generous token budget. Work out the
highest-leverage way to spend it on what I am currently doing.

First, inspect whatever context you can reach: recent work, notes, files,
repositories, task lists, project folders, logs, documents, conversations,
unfinished drafts, research, plans, goals and recurring problems. Do not assume I
use any particular system. Find the equivalent of each thing in whatever I do use.

Do not give me generic productivity advice or small tasks I could do myself. Find
work that suits a model that can read many sources, compare at scale, classify,
critique, reconcile, and produce large finished outputs.

Look especially for chances to:
  - synthesise scattered information into one reusable artifact
  - mine old work for assets worth reviving, combining, repackaging or finishing
  - find recurring mistakes, bottlenecks and failure modes, and write drills against them
  - convert repeated manual work into a system
  - compare or classify dozens of items against consistent criteria
  - turn raw information into decision-ready material, not summaries
  - build a source of truth: knowledge map, evidence library, glossary, canonical brief
  - stress-test a plan: one case for it, one case against it, then reconcile
  - build practice material from my real weaknesses
  - finish substantial work that is already part done
  - audit my existing systems for duplication, stale projects and unused output
  - find contradictions where two plans need the same time, money or attention
  - build evaluation mechanisms: tests, rubrics, scorecards, feedback loops
  - compress thinking I keep repeating into a reference
  - move work from "AI researches" to "human only approves"

For each task tell me: what it is, why it is high leverage, what information it
should inspect, the final deliverable, what you do, the small part I still do, and
why a large context window matters for it.

Group the tasks as QUICK WINS, DEEP WORK, SYSTEMS and EXPERIMENTS.

Start by naming what looks most important in my current work, and where unfinished
work is piling up. Then give me the 10-20 best tasks, ranked by the downstream work
they unlock, not by how interesting they are. Say plainly if some of my apparent
projects are a poor use of tokens. Prefer "turn what exists into leverage" over
"make another new thing".

Finish with: the top 3 tasks to start with, one underrated task, one place I am
wasting AI capacity, one large task I can hand off almost entirely to an agent, and
one meta-level change that would make all my future AI work more effective.

A line worth adding to your CLAUDE.md or AGENTS.md: when the token budget is abundant, spend it compressing accumulated information into reusable systems, evidence, tests or decisions, not on generating more ideas.

See when your tokens actually expire

Half the reason this capacity gets wasted is that you cannot easily tell when it is about to. In Codex CLI, the /status command shows the 5-hour and weekly windows with their reset times. In the ChatGPT app, Settings then Usage and Billing shows the same, plus any banked limit resets. The community site whenreset.dev tracks the free resets OpenAI hands out to everyone. There is also an unofficial endpoint people pass around, which you can ask Codex itself to query:

curl -sS 'https://chatgpt.com/backend-api/wham/rate-limit-reset-credits'

It is unofficial and needs your authenticated session, so treat it as a rough gauge. The point is to know whether you have hours or days, because those are two different jobs.

If the answer is "work on this codebase": the seven prompts

The meta prompt usually points at a repo. These seven are the ready-made versions of the jobs it tends to name. The pattern across all of them is the same: read-heavy, artifact-producing, and safe to run while you do something else.

#PromptWhat you get back
1Parallel read-only repo auditA ranked bug and risk report from specialist reviewers
2Fix one bounded bug categoryA clean PR with a definition of done and a verification run
3Test-debt passMissing tests written, flaky tests isolated
4Code archaeologyA module map plus evidence-backed cleanup candidates
5Documentation passADRs, runbooks and user flows built from the real code
6Mine sessions into a reusable skillThe smallest reusable skill, CLI or subagent
7Repo maintenance sweepTriaged issues and PRs plus the next high-value actions

1. Parallel read-only repo audit

The single best use of a big token budget. Many read-only reviewers, each with its own context window and each looking for one class of problem, find far more than one agent reading the whole repo. Cap it at about six reviewers, and ask the main thread to verify each finding before reporting so the plausible-but-wrong noise gets dropped.

Audit this repository using parallel read-only subagents. Spawn one specialist
reviewer per area, each in its own context window, and do NOT let any of them
modify files:

  - Security (secrets, injection, auth gaps, unsafe deserialization)
  - Correctness / business logic (wrong behavior, edge cases, off-by-one)
  - Regressions (recent changes that break existing behavior)
  - Architecture (coupling, boundary violations, dead layers)
  - Performance (N+1s, needless allocation, blocking calls)
  - Test coverage (untested critical paths)

Each reviewer reports findings to the main thread. The main thread verifies each
finding against the actual code (drop anything it can't reproduce), dedupes, and
returns ONE ranked report: severity, file:line, why it's a problem, and the
smallest fix. Do not change any code. This pass is read-only.

2. Fix one bounded category of bugs

Do not say "fix bugs". Pick one category, give it a definition of done and a command that proves it worked. Good categories are one lint rule, one exception type, or one deprecated API call. The verification command is the whole point: no "done" claim without a passing run.

Fix one bounded category of bugs in this repo: <e.g. all unhandled promise
rejections / all missing null checks on API responses / all incorrect error
status codes>.

Rules:
  - Only touch that category. Do not refactor unrelated code.
  - Definition of done: every instance of this category is fixed or explicitly
    listed as "left alone, because <reason>".
  - After the changes, run <verification command, e.g. `npm test`, `pytest -q`,
    the linter> and paste the output. If it doesn't pass, keep going.
  - Produce a summary: what you changed, file by file, and what you deliberately
    skipped.

Open the work as a single focused commit/PR I can review in one sitting.

3. Test-debt pass

Testing is the ideal token sink: high value, low risk, and it produces artifacts you keep. "Rank by blast radius" stops the model from testing trivial getters to pad the numbers. The flaky-test isolation alone is often worth the whole run.

Do a test-debt pass on this repo. In order:

  1. Map coverage: which critical paths have no tests? Rank by blast radius
     (what breaks in production if this function is wrong).
  2. Write missing tests for the top untested critical paths. Real assertions,
     not smoke tests. Cover the edge cases, not just the happy path.
  3. Find flaky tests: run the suite a few times, flag any test that passes and
     fails non-deterministically, and diagnose why (timing, shared state, order
     dependence).
  4. Leave a prioritized report: what you added, what's still uncovered and why
     it matters, and which flaky tests need a human decision.

Run the suite at the end and paste the result.

4. Code archaeology

Point it at the parts of the repo nobody understands anymore. It has the patience to read every old PR and neglected module, and you do not. This one is explicitly read-only, because the value is the map, not premature deletion.

Do code archaeology on this repo. Investigate the neglected and least-understood
parts: modules with the oldest last-touched dates, code with no tests and no
recent commits, and closed/merged PRs that changed core behavior.

Produce:
  1. A module map: what each major module does, its dependencies, and its role.
  2. Evidence-backed cleanup candidates: dead code, duplicated logic, abandoned
     experiments, config nobody reads. For each, cite the file and the evidence
     (no callers, superseded by X, last touched <date>).
  3. Anything surprising or risky you found along the way.

Do not delete or change anything. This is a mapping pass. I decide what to cut.

5. Documentation pass

Docs written from the actual code, not from a template. This is the artifact that keeps paying off long after the reset. "Cite the files" and "write the question down" keep it honest instead of fluent but wrong. ADRs reverse-engineered from existing code come out surprisingly well, because the decisions are already there, just undocumented.

Write documentation for this repo, derived from the real code, not boilerplate.

Produce, where the code justifies it:
  - ADRs (architecture decision records) for the non-obvious choices already made,
    reverse-engineered from the code and commit history.
  - A runbook: how to deploy, roll back, and handle the top 3 failure modes.
  - The main user flows, described step by step from entry point to result.
  - UAT scenarios: concrete acceptance checks a human can run before a release.

Cite the files each doc is based on. If something is ambiguous, write the question
down instead of guessing.

6. Mine your history into a reusable skill

Turn the work you have already done into a tool you can reuse. This compounds, because every future session gets cheaper. Ask for the smallest version, otherwise you get an over-engineered framework.

Look back over my recent sessions, saved memories, and the corrections I've made
to you in this project. Find the workflow I repeat most often: the thing I keep
re-explaining or re-typing.

Then build the SMALLEST reusable version of it: a skill, a subagent definition, a
CLI script, or a saved prompt, whatever fits. Requirements:
  - It captures the pattern, not one specific instance.
  - It's documented well enough that I (or another agent) can use it cold.
  - It's the minimum that works. No framework, no config sprawl.

Show me the workflow you identified and why, then the artifact.

7. Repo maintenance sweep

Triage the backlog of issues and PRs you have been ignoring. "Do not close anything automatically" is the guardrail: it proposes, you review. Finding the already-fixed issues alone clears a lot of noise.

Do a maintenance sweep of this repo's issues and PRs (use parallel workers if the
backlog is large).

For issues and open PRs:
  - Identify duplicates and group them.
  - Flag anything that looks already-fixed in the current code (cite the code).
  - Flag stale items that need a human decision.
  - Rank the still-valid items by value-to-effort.

Output a single triage report with a recommended next 5 actions. Do NOT close,
merge, or comment on anything automatically. Just propose. I make the calls.

Two bonus jobs

A big enough budget is one of the few times a legacy refactor is realistic. Plan first, because a rolling 5-hour cap can stop a big refactor half way and you want a resumable plan, not a half-migrated repo.

This is a large legacy codebase (<N> lines). Propose a future-proof structure.
First produce a plan: target module boundaries, the migration order, and the
risks, before touching anything. Then, only if I approve the plan, execute it in
small verifiable steps, running the test suite after each step. Stop and report if
anything fails.

And a workspace cleanup, read-only. It proposes, you delete.

Scan this project (or my dev workspace) for stale and idle files: build artifacts,
old logs, orphaned branches' leftovers, dependency caches, large files nobody
references. Produce a list with sizes and a reason each is safe to remove. Do not
delete anything. I'll confirm the list.

If you are not in a rush: unattended runs

A reset that is days away is a different job from one that is hours away. With time to spare, the pattern the community keeps landing on is to queue work, walk away, and review a pile of pull requests in the morning. Three shapes show up again and again.

Overnight batch jobs

Many small, independent, verifiable jobs across many repos, with nothing auto-merged. One reported night produced 60 open PRs across 21 repos: 12 repos scaffolded, 817 pages given metadata and JSON-LD, a database migration in five sequential PRs, and error monitoring added to four apps. Jobs that suit this shape: add structured data to every page of a content site, add logging to every service that lacks it, write the missing README for every package in a monorepo, backfill type annotations file by file, generate a changelog from real commit history, or run the same audit across every repo you own and write one report each.

Large refactors, phased

The recipe people converge on: write a PLAN.md first, isolate the work in a git worktree, allowlist only the tools the job needs, run the test suite between phases, and commit after each green phase. The reported outcome is roughly 70% clean, 20% partial and 10% failed, so a resumable plan matters more than a clever prompt (playbook, what breaks). Refactors that survive: extract shared helpers and migrate every caller, deduplicate utilities that drifted into four copies, split one oversized module along a boundary you name in the plan, or replace a deprecated library call across the codebase.

Spec-queue runners

Instead of one prompt, you hand the agent a queue of small specs and it works down the list: implement, test, review, open a PR, next. By morning you get working branches and one page saying what was built and what needs you (example runner). Queues worth filling: ten bug reports each written as a failing test, fifteen small features each one screen or one endpoint, or every open issue labelled good-first-issue in your own repo.

Rules people learn the hard way

  • Write the plan before the run. A rolling cap can stop a job mid-way, and you want to resume, not restart.
  • Use a worktree or a branch. Overnight changes should never land in your working tree.
  • Allowlist tools. A job scoped to read, edit and test cannot start deleting things.
  • Open PRs, do not merge. The review is the point.
  • Run the tests between phases and stop on failure, rather than at the end.

Watch the caps while you spend

Big parallel jobs burn fast. Most plans have a rolling 5-hour cap that can stop a job mid-run, plus a weekly cap on top. So write the plan first, then run. To stretch what you have: run subagents on a smaller model for the grunt work, turn off fast mode for background jobs you are not waiting on, set a lower reasoning level for mechanical passes, and cap subagents at about six because more mostly duplicates.

Why this is a cost habit, not a hack

If you pay for a plan, your real cost per unit of work drops every time you use more of what you already bought. Letting a reset expire empty does the opposite. Ten minutes of copy-paste before a reset turns dead capacity into a backlog you would otherwise pay an engineer's time to clear.

If you ran one of these and got a job worth stealing, send it to the repo as a PR or an issue with the tool you ran it in and what it gave you back. Real examples help more than theory.

Work with me

I help teams get real return out of their AI and cloud spend, including setting up the agent workflows that turn idle capacity into shipped work. If that is worth a conversation, book a call at cloudyeti.io/meet.

Book a call