Knowledge base

Build an AI agent around one task, permissions and recovery

Define an AI agent through one task, permitted tools, action limits, evaluation cases and human ownership. Test failures and approvals before adding autonomy.

An AI agent uses permitted information and tools to take steps towards a goal. That makes tool access, action boundaries and recovery part of the design. A broad instruction to support a team is not a manageable first specification.

Begin with one task that can be started, assessed and stopped. Add capabilities only after their outcomes can be checked and maintained.

Write a bounded task specification

A useful first task could read enquiries from one approved source, check five mandatory fields, summarise the request and prepare a draft task for a staff member. That is an illustrative specification, not evidence of a delivered customer system.

Name the trigger, source, checks, expected output and human next step. Explicitly exclude promises about price, customer-record changes, outgoing messages or unrelated dossiers when those actions are outside scope.

Choose sufficient task volume to evaluate behaviour, with manageable consequences of failure. Preparing proposals often provides a more controlled first experiment than immediate irreversible execution.

Treat each tool as a permission

For every search, mailbox, CRM or calendar connection, record why it is needed, which records may be read or changed, the technical identity, limits, failure responses and how access can be revoked.

Reading availability does not require deleting appointments. Writing a draft does not require sending it. Restrict the search space to approved sources and define precedence when records disagree. Use only the sensitive data needed for the task.

Separate preparation from execution

Separate preparation from execution
Action categoryInitial control to define
Reading and preparingApproved sources, data scope and expected output
Reversible internal changesBounded quantities, review and a correction path
External or sensitive actionsExplicit authority, suitable approval and recovery limits

For approval, show the intended action, recipient or affected record, reason and relevant evidence. Define how long approval remains valid and what happens if context changes. Use application-enforced limits for amounts, quantities, recipients or time windows where relevant. A prompt instruction alone is not an access-control boundary.

Build a reusable evaluation set

Use representative, protected workflow cases and define the expected source, classification, mandatory information, permitted action and stopping point. Wording can vary without being wrong; evaluate what matters to the business.

Include missing or contradictory information, outdated sources, duplicates, invalid tool output, prohibited requests, misleading instructions inside documents, expired access and partially completed actions. Reuse a stable core set between versions so changes can be compared meaningfully.

Design failures and recovery before release

Distinguish unavailable tools, invalid records, unclear instructions and an unknown business situation. Each needs a different response. A failed external action must not be retried blindly if it might already have completed.

Give someone operational authority

Define safe retries, duplicate detection, handoff, controlled restart and the evidence a reviewer needs. Diagnostic records should be bounded and exclude secrets or unnecessary personal content. Staff need visibility into completed tasks, stop points, approvals, failing tools and repeated corrections.

Give an owner authority to revoke access, disable the workflow and review changes. An agent without operational ownership is an experiment rather than a managed process.

Compare buying and building against the same controls

A standard tool may fit a standard workflow. Custom work becomes more relevant when your roles, data and approvals are distinctive. Compare source and evaluation control, data handling, access, export, supplier changes and operating responsibilities alongside features.

Record task, trigger, sources, tools, prohibited actions, approvals, success, stop conditions, recovery and owner in one specification. Explore custom software development, the personal AI assistant experiment and our working method. For customer-facing interactions, also test the whole chatbot journey.