My Long-Running Codex Goal: Four Days and Still Working
On October 9, 2026, my long-running Codex goal showed 4 days, 14 hours, 8 minutes and 22 seconds. I took a screenshot while Codex was still working on an SEO product I have not launched yet. The goal remained open, with more tasks to complete.
I had asked it to complete the whole plan for the unreleased SEO product I am building. By that point, I had also asked it to use more agents, lower the reasoning effort, pause for updates, resume after a restart, and spend less time repeating tests.
This is my account of a long-running Codex goal: the work it produced, the work that slowed it down, and the decisions I still had to make. It is an unfinished build, not a benchmark or a launch announcement.

The timer above belongs to the goal. It does not establish four days of uninterrupted model computation. My session includes explicit pauses and resumptions, an app update, a computer restart, tool execution and waiting.
What I asked Codex to build
The conversation began on October 3 with a smaller question: do users have their own login credentials, can they sign in with Google, and can they generate personal MCP keys?
Then I expanded the requirements. Each user should see their own projects. Users should be able to invite people into a workspace. I wanted settings for personal preferences and crawler concurrency. The platform would eventually be commercial, with a free initial launch.
I supplied a much larger SEO product catalog to guide the build. The resulting plan covered 15 product areas, 42 app references and 19 public free tools, plus a website traffic ranking feature. It included technical SEO, keyword research, rankings, backlinks, AI visibility, content, analytics, local SEO and later enterprise features.
I also wanted the product to order articles from the content engine behind my own website.
So the instruction to complete the whole plan referred to a substantial product roadmap. I helped create that scope. A four-day timer on a small bug fix would tell a different story.
Which model and settings I used
The main session log identifies the model as GPT-6.1 Sol, with the identifier gpt-6.1-sol. Its recorded turns include ultra, medium and a small number of high reasoning settings. That identifies the main session; it does not establish the model behind every reviewer or external tool.
On October 5, I explicitly switched the effort to medium to save tokens. I then turned off turbo for the same reason. Those were my intentions, not measured savings: I do not have an audited cost comparison for the two configurations.
I also requested more parallel agents. The session used delegated work and Claude peer reviews for parts of the implementation. That let separate tasks move forward, but it also created work to reconcile their outputs and verify the combined result.
My earlier GPT-6.1 Sol comparison→ covers published model data. This build is a different kind of evidence: one real project with changing scope and settings, rather than a controlled comparison between models.
How long can a Codex goal keep working?
In this case, the interface displayed more than four days for the same ongoing goal. That is the observation I can support. It is not a maximum runtime guarantee, and I cannot calculate active inference time from the screenshot.
OpenAI describes Goals in Codex as objectives that persist across turns. Codex can continue toward an outcome, while the user can pause or resume it. Completion, interruptions, budgets and blockers affect whether work continues.
My experience fits that persistent workflow. I could return to the same objective after pauses and steer the next work. Keeping the target available was useful. It did not make the target smaller or ensure that each new turn moved the product closer to launch.
What had it delivered by October 9?
The delivery log on October 9 recorded 30 partially completed requirements, 39 not assessed, and zero fully accepted requirements out of a 69-point tracking list.
That count needs context. The list measures broad product acceptance. Zero fully accepted requirements does not mean zero working code. Thirty partial requirements also does not mean the product is 43 percent complete.
The log records a disposable local smoke test of a supported account baseline: create the first account, log in with a password, read the session and its private owner workspace, then log out and reject the old session. That is a concrete user flow with a bounded test result. It is not proof that the latest complete application is ready for customers.
Other progress included local password-reset source integration, backlink review drafts, saved Search Console report work, a customer-token GA4 report transport, and article-to-LinkedIn draft work. Several of those pieces still had disabled activation, incomplete integration or deferred verification.
The dated checkpoints explicitly reported no commit, push or deployment. Google login and full customer readiness remained unverified. I had a growing local implementation with useful evidence, rather than a released platform.
What slowed down my long-running Codex goal?
I expanded the goal into a product roadmap
Adding Google login sounds like one feature. Adding private customers changes who may access projects, reports, background jobs and integrations. I wanted that boundary to hold across the application.
Then I added keyword research, backlinks, content generation and a much larger catalog. Some of the elapsed time reflects necessary work on a broad target. My original goal made it easy to keep opening the next unfinished area.
The verification process became too repetitive
I asked for source-backed decisions, narrow changes, reviews and explicit evidence. Those instructions helped prevent vague claims that something was finished.
But the session accumulated repeated source checks, review preparation and test-fixture work. My judgment is that the balance drifted too far toward proving individual pieces before completing the next usable flow.
On October 9, I told it to stop spending so much time on tests, work toward completion and leave a larger test for later. That did not remove the need to verify account isolation. It changed the sequencing: focused checks during implementation, followed by broader validation of the assembled flow.
Some failures belonged to the test setup
One password-reset database attempt failed because a synthetic account record omitted a required display-name field. Repairing that fixture was necessary to run the test, but it was not a new product feature.
A later long-running database attempt ended when the workstation rebooted. The delivery log did not claim that attempt passed. Subsequent diagnosis found repeated computation of earlier verification contracts and buffered progress output, which made the running process harder to assess.
These details matter because waiting is ambiguous. A live process may be working, recomputing the same prerequisite, or producing output that I cannot yet see. The timer alone cannot tell me which.
More agents added coordination
Parallel work helped with separable tasks. The main session still needed to inspect results, resolve dependencies and incorporate changes into one application. I cannot attribute a speedup to adding agents because I did not run the same project with and without them.
The practical question became whether another agent could finish an independent slice or would create another handoff for the main session.
What I still do as the human
I choose the product direction and decide which features matter next. I check whether the reported progress describes working behavior, local source code, or an unverified proposal. I pause the work when I need to update Codex or restart the computer, then ask it to continue from its saved state.
I also challenge the pace. During this build, I:
The plan and delivery log give me something to inspect beyond chat messages. They also need discipline: a dated checkpoint is useful only if it says what changed and what remains unproven.
The experience continues the trade-off I described in I Thought AI Would Give Me More Free Time→. I can attempt a larger build, but I still spend time deciding what deserves to be built and checking the result.
The advantages and disadvantages so far
In my experience with this build, the biggest advantage is continuity. I can keep a substantial objective open and resume it after interruptions. Codex has produced local implementations, investigated failures and maintained detailed records that help me review the work.
It can also handle several kinds of work within the same project: database changes, API behavior, frontend flows, integration transports and documentation. That makes a broad build possible for me to direct.
The disadvantage is that activity can look like progress. Many successful checks can coexist with an unfinished product. High reasoning effort and more agents introduce budget and coordination choices without guaranteeing faster delivery.
Long runs also make scope discipline harder. A partially completed roadmap gives the agent many defensible next actions. I need to decide which one delivers the next useful result.
I would use a persistent goal again, but I would give each implementation stage a smaller acceptance target. For example: create an account, log in, see its private workspace, and reject access from a second account. Keep the larger roadmap as context, finish that flow, then move on.
What I will measure next
The SEO product is still in development and has not launched as of October 9. The next useful measure is a complete user flow against the current assembled application, with its remaining blockers listed explicitly.
After that, I want evidence for Google login, customer-bound integrations, personal MCP keys and a safe release. Codex continues to work through tasks; this article records the October 9 snapshot. I will judge the build by those outcomes rather than how long the goal stays open.
Sources and the limits of this account
The timer comes from my screenshot above. The model identifier, changes to settings and my interventions come from the main session history. The scope and progress counts come from the project plan and its dated delivery log. Those project records are private working documents; I have not published the logs or customer data.
OpenAI's Using Goals in Codex explains the persistent-objective workflow. It supports the description of Goals, not the delivery claims about my project.
The 69-point count is a broad acceptance snapshot, not a measure of hours remaining. The screenshot does not prove uninterrupted inference, the settings changes do not prove cost savings, and local checks do not establish production readiness. This build is still underway.
*This article was drafted with AI assistance from my screenshot, session history and project delivery records. It describes one ongoing build. It does not establish a general Codex runtime limit, a model ranking, an audited cost or production readiness.*



