AI, usefully · 4 min read ·

How an AI agent keeps track of its work

Once an AI agent can use tools, a less obvious question becomes important: how does it know what has already happened?

A request might require several steps. Retrieve data, check its quality, calculate a result, prepare a summary, then wait for someone to review it.

What happens if the third step fails? Or the review arrives tomorrow? Should the agent begin again, continue from where it stopped, or check whether the earlier results are still valid?

This is where state enters agent design. State is the application’s record of the task as it progresses: what was requested, what has completed, what evidence is available and what remains unresolved.

It is one of the foundations of a reliable agent.

A conversation is not a sufficient progress record

Consider an assistant preparing a weekly product report.

It needs to retrieve activation data, check for missing dates, calculate changes and draft a summary. A human reviews the draft before it is distributed.

You could put everything into one long conversation. But a message saying “the data looks complete” is difficult to use as an operational control. Which dates were checked? Against which dataset? Did the check actually pass, or did the model merely describe a plan?

A structured state record makes those distinctions explicit:

{
  "report_period": "2026-09-21/2026-09-27",
  "data_status": "retrieved",
  "quality_status": "failed",
  "missing_dates": ["2026-09-24"],
  "draft_status": "not_started",
  "review_status": "not_requested"
}

The application can use that record to decide which steps are eligible to run.

Here, drafting a definitive weekly comparison should wait. The missing date needs investigation first.

Separate plans from completed work

An agent may propose:

Retrieve the dataset, validate completeness, then calculate activation.

That is a plan. Each step still needs a recorded outcome.

Useful statuses might include not started, running, succeeded, failed, and waiting for review. The exact labels matter less than their meaning.

“Running” should not count as “succeeded.” A failed tool call should not disappear because the model continues writing. Human approval should have its own record, tied to the version that was reviewed.

For our report, the application should mark the retrieval successful only after receiving and checking the tool result. The model should not be able to declare a missing result complete through prose alone.

This also helps explain a failure. You can inspect the stage where progress stopped rather than reconstructing it from a confident final paragraph.

Save progress at useful boundaries

A checkpoint is a saved snapshot of the task’s state at a particular stage.

For the report, useful boundaries could be after retrieval, after validation and after the draft is prepared. If the process stops, a checkpoint provides a place from which to recover.

LangGraph’s documentation describes checkpointers that persist the state of a particular thread, supporting continuity, human review and recovery. It distinguishes those checkpoints from stores used for information shared across interactions.

Saving state does not train the model. It gives the application information it can load into a later step.

It also matters where that information is saved. A record held only in the process’s memory may disappear when the process restarts. Durable storage is needed when progress must survive that restart.

Resuming requires judgment

Suppose the report is ready for review on Monday, but approval arrives on Wednesday.

Can it simply continue?

Perhaps. But the underlying data may have been corrected. A metric definition may have changed. The approved draft may no longer match the latest calculation.

The state should retain enough information to check that relationship:

RecordWhy it matters
Reporting periodPreserves the intended scope
Dataset version or retrieval timeIdentifies the evidence used
Validation resultShows which checks passed
Draft versionIdentifies the reviewed content
Approval referenceConnects permission to that version

A checkpoint preserves progress; it does not establish that every saved fact remains current.

Before continuing, the application may need to recheck data freshness or request another review if the draft has materially changed.

A checkpoint does not prevent duplicate actions

There is another complication.

Imagine the report was distributed successfully, but the application stopped before saving “distribution complete.” On restart, the state still says distribution is pending.

Repeating the step could send it twice.

This is why actions outside the application need their own protection. An idempotent operation can be retried without creating an additional effect—for example, by attaching a unique report-and-version identifier that the receiving service recognizes.

When a service does not offer that protection, recovery may require checking whether the action already happened or involving a person. Blind retries are a poor substitute for knowing the outcome.

Try designing the state before the agent

Choose one recurring task and ask an AI assistant:

Propose a state record for this workflow. Separate requested work, completed work, evidence, failures and human approval. For every step, specify what must be true before it runs and what record proves completion. Identify where restarting could repeat an external action. Do not assume a plan has been executed.

Then inspect the design yourself.

Pause the imagined workflow after each step. Ask what would happen if the process restarted there, or if an earlier result changed.

A good state record should make those answers clear. You should be able to tell what happened, what remains uncertain and which next step is justified—without relying on the model to remember correctly.