Most agent experiments begin the same way.
You write a prompt. Test it. Refine it. Connect a tool. Test again. The output looks promising and someone asks: “can we let this run automatically?”
That is exactly where the prompt stops being enough.
A prompt can tell an agent how to reason or respond. It does not, by itself, settle questions such as:
- who owns the agent;
- which data it may access;
- which fields it may change;
- which actions require approval;
- how much it may spend;
- when it should stop and escalate;
- how quality is measured;
- how to shut it down without breaking the process.
Those are no longer prompt questions. They are operating design questions.
RevOpsHubs Thesis
A production agent needs an operating contract: an explicit definition of the work, context, autonomy, risk, economics and accountability under which it may operate.
What is an Agent Operating Contract?
The Agent Operating Contract is a RevOpsHubs framework. It is not a legal contract, and it should not become a 20-page technical specification.
It is an operational one-pager that accompanies an agent as it moves from experiment to pilot or production.
Its job is to turn something vague — “we have an agent helping Sales” — into a governable unit of work.
RevOps, the process owner and the technology administrator should all be able to read it. If only the agent builder can understand it, it is still too technical.
The 10 decisions in the contract
We propose ten blocks. The order is deliberate: you define the work before you spend time perfecting the prompt.
Objective — what outcome should it produce?
Not “help Sales.” Use something observable: prepare a meeting brief, classify an inbound request, flag risk in an opportunity, update specific data.
Owner — who is accountable for performance?
A named person should define acceptable quality, review failures and approve material changes. The technical builder may not be the business owner.
Context — what information may it use?
Define systems and sources: CRM, notes, email, documents, knowledge base, web, ERP. Available context is not the same as permitted context.
Tools — which capabilities does it have?
Search, create task, update field, send email, call API, use MCP. Tools are operational power and should be listed explicitly.
Permissions — where can it read and write?
Objects, fields, systems and scopes. “HubSpot access” is too broad. Reading Deals and changing Deal Stage are materially different capabilities.
Approval — where must it ask?
Define autonomous actions, approval-gated actions and prohibited actions. OpenAI supports human-review pauses for sensitive tool calls; the design principle matters regardless of platform.
Budget — how much capacity may it consume?
Runs, credits, tokens, compute time or external calls. As pricing becomes usage-based, budget becomes part of agent design rather than only procurement.
Quality — how will you know it works?
Define outcome and quality metrics, not just run count. Accuracy, acceptance without correction, rework, time saved, conversion or cost per outcome.
Escalation — when should it stop and hand off?
Missing data, conflicting rules, low confidence, commercial exceptions, retry limits or risky actions. A reliable agent also knows when not to continue.
Rollback — how do you go back?
How do you disable the agent, revoke connections, remove write access, restore data or return to the previous version? The exit path belongs in the design.
So does the prompt matter less?
No. It matters more because it no longer has to carry responsibilities that belong elsewhere.
Modern agent platforms already separate these layers. In OpenAI, an agent definition can include instructions, tools, guardrails, MCP servers, handoffs and structured outputs. Approval mechanisms can pause a run before a sensitive action. HubSpot lets teams configure MCP tools, simulate runs and set monthly limits.
The prompt is part of the contract. It is not the contract.
Prompt engineering tells the agent how to work. Operating design decides what work may exist and under which limits.
Example: an inbound qualification agent in HubSpot
Consider a B2B company using HubSpot that receives contact requests through its website. The goal is to reduce manual research and give Sales better context before the first response.
A lightweight operating contract could look like this:
Notice what is not here: two hundred lines about personality, tone or writing style.
Those can matter. But what makes this agent a reliable part of the Revenue System is the operating boundary around it.
Not every agent needs the same governance depth
A one-page contract should not become universal bureaucracy. Governance depth should match operational surface.
Research and recommendation
Uses context and produces suggestions without writing to systems or contacting customers. Lightweight contract, logs and ownership still apply.
Changes internal work
Updates fields, creates tasks, notes or tickets. Requires explicit permissions, simulation, auditability and rollback.
Acts on customers or money
Sends, publishes, deletes, commits pricing, changes critical states or takes hard-to-reverse actions. Approval and guardrails should be much stronger.
This connects directly to the Shadow Agent Risk Model: the larger the context, tool, write and cost surface, the more explicit the operating contract should be.
The contract also defines how to test
A common mistake is to test only whether the agent produces a good answer.
In production, the whole system needs testing:
- Did it choose the right tool?
- Did it refuse work outside scope?
- Did it request approval when required?
- Did it escalate when context was insufficient?
- Did it stay inside run and budget limits?
- Did it improve the process or merely create more review work?
OpenAI recommends evaluating agent workflows through traces, tool calls, guardrails and handoffs, rather than judging only the final response. HubSpot can simulate agent runs without changing CRM data or consuming credits, which gives teams a practical pre-production testing loop.
The operating contract provides the acceptance criteria for those tests.
Budget becomes an agent property
This is still new territory for many RevOps teams.
Historically, the marginal cost of another workflow execution was often invisible after the software subscription was purchased. Agents bring a more direct relationship between work performed and consumption.
In HubSpot, the new EMEA pricing model already makes credits an explicit unit of capacity. HubSpot recommends simulating runs to estimate cost and setting monthly limits.
That makes “budget” a behavioral constraint, not just an administrative line item.
An agent allowed to run 500 times a day is operationally different from one capped at 20 priority cases.
What to do on Monday morning
Pick one agent already in testing, or one somebody wants to move into production.
- Write the objective as one observable outcome.
- Name one business owner.
- List context sources and remove those that are not necessary.
- List every tool and mark read/write.
- Create three columns: autonomous, approval required, prohibited.
- Set a pilot usage limit.
- Choose two metrics: one quality metric and one outcome metric.
- Write the escalation conditions.
- Test 20–50 representative cases, including edge cases.
- Document how to disable or roll back.
If you cannot fill these sections, you do not have a prompt problem yet. You have an operating design problem.
Agents need freedom. But bounded freedom.
An agent that asks permission for everything is not especially useful. An agent that can do everything is not especially governable.
The interesting design space is in the middle.
Give the agent enough autonomy to remove human work, while making context, permissions, risk, economics and accountability explicit.
That is what the Agent Operating Contract is for.
Do not move an agent into production because the prompt is good. Move it into production when the system around it is clear.