This post is part of my TaskHub series. In the previous post, I described how I move from a system specification to a verified Oracle implementation with Codex. Here I focus on how I structure one development task before handing it to Codex.
I will use Task 50 — Tasks / Subtasks API as a running example. It covers creating, updating, and soft-deleting tasks and subtasks, including hierarchy rules, ownership checks, and task status behavior.
Implement task management. leaves too much open to interpretation. I want a bounded unit of work that defines the change, protects earlier decisions, and makes completion testable.
The structure below is what I currently use. It developed during the TaskHub lab. I do not present it as a universal or optimal format.
I usually divide each task into eight sections. Their size varies with the work involved.

The objective is one clear statement of the result I want. It provides direction without repeating the full specification.
Implement task and subtask operations in the existing application API.
The task builds on earlier work. I identify the relevant components and decisions so Codex does not redesign something that is already in place.
The task depends on established parts of TaskHub:
The scope lists the operations and behavior I want to change now. Related ideas do not automatically become part of the task.
The relevant public operations are:
CREATE_TASK — create a task or subtask.UPDATE_TASK — update properties, state, or parent relationship.DELETE_TASK — soft-delete a task and its subtree.Constraints define rules that the implementation must satisfy. A function that appears to work is not acceptable if it breaks one of these rules.
1..3.DONE must set COMPLETED_AT; leaving DONE must clear it.Acceptance Criteria state what must be true before I can consider the implementation successful. They are the target conditions, not the test output.
Tests are part of the task, not something I invent after the code has been written. I specify both allowed behavior and the failures the API must prevent.
This section protects decisions already made. It is different from behavioral constraints: it defines changes the agent must not make while solving the current problem.
SYS or SYSTEM.Some of these rules are specific to my lab. They are not universal Oracle development rules.
When Codex reports that a task is complete, I do not want to rely only on a statement that everything passed. I want to see what was actually checked and what the results were.
This is what I mean by Evidence: inspectable results that support the claim that a task was completed successfully.
The distinction matters:
Evidence makes the result reviewable. It also helps me investigate failures or verify a previous checkpoint later. A PASS statement, without supporting results, is not enough.
For Task 50, the evidence I would expect includes:
USER_ERRORS queries.CREATE_TASK, UPDATE_TASK, and DELETE_TASK.A small example shows the relationship. The diagram describes an illustrative verification path, not a captured test log.

For me, a task is not done when the source files exist. Three conditions must come together: the change is implemented, the required checks pass, and the evidence is available for review.

Codex can execute checks and collect results, but I decide whether the task is accepted. A failed check or unclear result means more work, not automatic closure.
This is a checkpoint for the defined task, not proof that the entire application is defect-free. The next post will focus on how I make Codex demonstrate that its Oracle changes worked.