An AI agent uses permitted information and tools to take steps towards a goal. That makes tool access, action boundaries and recovery part of the design. A broad instruction to support a team is not a manageable first specification.
Begin with one task that can be started, assessed and stopped. Add capabilities only after their outcomes can be checked and maintained.
Write a bounded task specification
A useful first task could read enquiries from one approved source, check five mandatory fields, summarise the request and prepare a draft task for a staff member. That is an illustrative specification, not evidence of a delivered customer system.
Name the trigger, source, checks, expected output and human next step. Explicitly exclude promises about price, customer-record changes, outgoing messages or unrelated dossiers when those actions are outside scope.
Choose sufficient task volume to evaluate behaviour, with manageable consequences of failure. Preparing proposals often provides a more controlled first experiment than immediate irreversible execution.
Treat each tool as a permission
For every search, mailbox, CRM or calendar connection, record why it is needed, which records may be read or changed, the technical identity, limits, failure responses and how access can be revoked.
Reading availability does not require deleting appointments. Writing a draft does not require sending it. Restrict the search space to approved sources and define precedence when records disagree. Use only the sensitive data needed for the task.
Separate preparation from execution
| Action category | Initial control to define |
|---|---|
| Reading and preparing | Approved sources, data scope and expected output |
| Reversible internal changes | Bounded quantities, review and a correction path |
| External or sensitive actions | Explicit authority, suitable approval and recovery limits |
For approval, show the intended action, recipient or affected record, reason and relevant evidence. Define how long approval remains valid and what happens if context changes. Use application-enforced limits for amounts, quantities, recipients or time windows where relevant. A prompt instruction alone is not an access-control boundary.
Build a reusable evaluation set
Use representative, protected workflow cases and define the expected source, classification, mandatory information, permitted action and stopping point. Wording can vary without being wrong; evaluate what matters to the business.
Include missing or contradictory information, outdated sources, duplicates, invalid tool output, prohibited requests, misleading instructions inside documents, expired access and partially completed actions. Reuse a stable core set between versions so changes can be compared meaningfully.
Design failures and recovery before release
Distinguish unavailable tools, invalid records, unclear instructions and an unknown business situation. Each needs a different response. A failed external action must not be retried blindly if it might already have completed.
Give someone operational authority
Define safe retries, duplicate detection, handoff, controlled restart and the evidence a reviewer needs. Diagnostic records should be bounded and exclude secrets or unnecessary personal content. Staff need visibility into completed tasks, stop points, approvals, failing tools and repeated corrections.
Give an owner authority to revoke access, disable the workflow and review changes. An agent without operational ownership is an experiment rather than a managed process.
Compare buying and building against the same controls
A standard tool may fit a standard workflow. Custom work becomes more relevant when your roles, data and approvals are distinctive. Compare source and evaluation control, data handling, access, export, supplier changes and operating responsibilities alongside features.
Record task, trigger, sources, tools, prohibited actions, approvals, success, stop conditions, recovery and owner in one specification. Explore custom software development, the personal AI assistant experiment and our working method. For customer-facing interactions, also test the whole chatbot journey.
