Repository
Artificial IntelligenceAnalysis7 min readSep 30, 2026

By

Always-On AI Agents Need an Operating Contract, Not Just a Prompt

OpenAI's new always-on agents make persistence a product feature. Businesses now need explicit operating contracts for scope, evidence, budgets and stop conditions.

Always-On AI Agents Need an Operating Contract, Not Just a Prompt
Key Takeaways
  • 01Persistent agents need an operating contract that defines responsibility, authority, evidence and stop conditions.
  • 02Permissions control individual actions; they do not define the full lifecycle of long-running work.
  • 03Scope-boundary pressure increases as tasks, tools and intervening context accumulate.
  • 04Businesses should require budgets, checkpoints and reviewable receipts before consequential actions are completed.

Analysis: OpenAI's new dots product is being introduced as an always-on AI agent that can keep working across conversations, apps and recurring tasks. The product announcement matters because it moves the central design question beyond prompt quality. When an agent can continue working as conditions change, the real specification is an operating contract: what responsibility it owns, which resources it may use, what evidence it must return and when it must stop.

This is not simply a safety concern. It is a product and management problem. A persistent agent needs the equivalent of a job description, a runbook and a review process. Without those, even competent work can drift away from the outcome the business intended.

What changed with always-on agents

On September 29, 2026, OpenAI introduced dots as agents powered by GPT-6 Astra, with their own cloud computer, browser and access to connected apps. OpenAI says they can handle ongoing projects, scheduled work and background research, while bringing consequential decisions back for review.

The persistence is the important part. A conventional chat task usually has a visible beginning and end. A long-running agent can receive new information, accumulate memory, move between tools and continue after the original instruction has become stale. That makes the environment part of the specification.

OpenAI's September 29 dots safety appendix tests this directly. In one evaluation, the agent had to adapt when permissions or scope changed during a task. It passed 45 of 49 episodes, including all 17 explicit permission-change cases. In another test, increasing the number of intervening tasks from five to ten raised the rate of moderate scope-boundary flags from 8.6% to 19.7%. OpenAI says these evaluations are not necessarily representative of production, but the direction is useful: more accumulated context creates more boundary-management work.

Permissions are necessary, but not sufficient

I previously argued that AI agents need clear permissions before they get more responsibility. Persistent agents add another layer. A permission answers whether an action is allowed. An operating contract explains why the agent is working, which state it should maintain, how success is measured and what happens when the situation changes.

OpenAI's workspace controls illustrate the distinction. Its administrator documentation separates access to dots from cloud-browser use, network access, computer access, password management, plugins and service-level authorization. There is no single permission switch that describes the whole operating boundary.

For builders, this means authorization should be evaluated at each state transition. Reading a support queue, drafting a fix, running a test, opening a pull request and merging code are different states with different consequences. The agent should not inherit authority for the final state merely because it was allowed to begin the first one.

What an operating contract should contain

A useful operating contract can be short, but it should be explicit about six things:

  • Mandate: the responsibility the agent owns and the outcomes it does not own.
  • Authority: the systems it may read, the actions it may take and the actions that require review or handoff.
  • Budgets: limits on time, compute, spending, messages, retries and parallel work.
  • Evidence: the logs, tests, sources, screenshots or diffs required before work can be called complete.
  • Checkpoints: moments when changing conditions, new stakeholders or expanding scope require renewed approval.
  • Stop conditions: the failures, ambiguities, security signals or cost thresholds that end the run and notify an accountable person.

This is also where a definition of done becomes valuable. If an agent is told to “improve the website,” it can keep finding work. If it is told to correct three verified accessibility defects, pass named tests and prepare a reviewable change without publishing, the endpoint is observable.

Evidence should travel with the work

OpenAI's product examples emphasize completed pull requests, test results and review artifacts. That is the right direction. Persistent agents should return receipts, not just summaries. A decision-maker needs to see what changed, which source or test supported it, what remains uncertain and which actions occurred outside the reversible workspace.

The company's privacy and safety FAQ describes layered controls including plugin permissions, custom rules, automated action review and monitoring. It also warns that safeguards reduce risk rather than eliminate it, and that some completed actions cannot be reversed. An evidence trail therefore serves two purposes: quality review before approval and incident reconstruction after an unexpected result.

The practical Canadian question is data flow

For Canadian organizations, the first deployment question should be which connected information crosses into the agent's working context and where resulting files, memories and logs remain. As discussed in Canada's AI vendor privacy guidance, data mapping, retention, optional features and exit planning belong in architecture review rather than a later policy document.

An operating contract should therefore identify data classes the agent may use, locations where it may create durable artifacts and the process for revoking access without assuming that previously retained context disappears automatically.

Limits and what would change this assessment

The evidence available today comes primarily from OpenAI's product documentation and internal evaluations. Dots are rolling out gradually, enterprise access begins as a beta, and the safety appendix says some tests isolate model behaviour without the full production safeguard stack. The reported figures should not be treated as independent proof of real-world reliability.

This assessment would improve with external evaluations, public incident reporting, clearer admin telemetry and evidence showing how boundary failures change across months of real work. It would worsen if organizations treat persistence as permission, hide agent activity from reviewers or measure success only by task volume.

Always-on agents can be genuinely useful. But their value will come from disciplined delegation, not unlimited initiative. The winning interface is unlikely to be one perfect prompt. It will be a living contract that keeps responsibility, authority, evidence and accountability aligned as the work changes.

Disclosure: Jason Ansell is Co-Founder of Vector Smart Chain and author of The AI Money Revolution. Neither VSC nor the book is identified in OpenAI's materials, and this analysis does not imply a commercial relationship with OpenAI. This article is not legal, privacy or investment advice.

Sources

Share this page

X
AI agentsOperational governanceAutomationAI safety