AI agent orchestration connects task assignment, shared state, tool use and decision-making across agents. Its real test is not whether several agents can exchange messages. It is whether the application can explain who owns a task, which actions are authorised, what counts as completion and what happens when a worker stops responding.
Start with a single workflow and make those decisions explicit. Add specialist agents only where the work genuinely needs separate capabilities or boundaries. The blueprint below joins coordination and reliability in one design: a component map, a pattern comparison, a state contract and a failure worksheet. It is a design method, not a claim that an untested multi-agent application is production-ready.
Decide whether multiple agents are justified
A deterministic workflow is the better starting point when the steps and routing rules are already known. A direct model call can handle a bounded language task; a single agent can choose tools dynamically without also coordinating a team. Microsoft’s architecture guidance recommends the least complex approach that meets the requirement.
Use multiple agents when separate expertise, tool sets or security boundaries solve an identifiable limitation of that simpler design. Record what each additional agent contributes and what it must not access. Requiring human approval does not, by itself, require another agent: approval can be a deterministic gate in a single-agent application.
Keep the original objective visible as you divide the work. A specialist’s successful response is only an intermediate result. The application still needs to decide whether the combined work answers the request and whether any proposed external action is allowed. Otherwise, a larger agent roster merely creates more places for ownership to become unclear.
Map components and policy boundaries together
Separate planning from permission, and execution from acceptance. The orchestration research describes distinct responsibilities for planning and policy, execution and control, state and knowledge, and quality and operations. Treat these as responsibilities to implement, not a requirement to deploy a separate service or language-model agent for every box.
Use the following component map as a design worksheet. Assign a concrete application component and an accountable operator to each responsibility before choosing a framework. The owner column specifies the proposed authority boundary; it does not describe features automatically supplied by a hosting provider.
| Component | What crosses the boundary | Owner and decision |
|---|---|---|
| Task intake | Request, authenticated identity, permitted scope and acceptance criteria | Application entry point admits or rejects the task. |
| Orchestrator | Task identifier, dependencies, selected worker and allowed work | Workflow controller owns scheduling and authoritative task progress. |
| Queue or dispatcher | Work assignment and acknowledgement, linked to that task identifier | Dispatch layer moves work; an acknowledgement is not proof of a completed business action. |
| Specialist worker | Restricted context in; structured result or explicit failure out | Worker owns its assigned computation, not unrestricted shared-state updates. |
| State store | Versioned task status, results, checkpoints and approval references | Workflow controller authorises state transitions; storage preserves the record. |
| Tool and policy boundary | Proposed operation, target, arguments and caller identity | Policy enforcement authorises access; the tool adapter records the external action and its outcome. |
| Human approval gate | Exact proposed change, relevant evidence and approval decision | Authorised reviewer approves or rejects that specific action. |
| Final validator | Required results, acceptance rules and outstanding failures | Application validation decides completion, failure or escalation. |
| Observability | Correlated lifecycle events, errors, usage and policy decisions | Operators can reconstruct the run without relying on a worker’s narrative. |
These responsibilities connect in a loop: intake establishes the task; the orchestrator dispatches permitted work; workers return results; validation accepts or rejects them; the controller records the next state. A tool action crosses a separate authorisation boundary. An approval request pauses that action rather than merely adding a comment to the transcript.
Operational state and knowledge also need different treatment. Task progress, an approval decision and a completed action belong in the workflow record. Retrieved documents and working context supply information to agents. Do not make a retrieved passage the authority for changing a task’s permissions or declaring it complete.
Compare the coordination patterns
Choose the dependency structure before choosing how many workers to start. These three patterns answer different coordination needs; they are not successive maturity levels that every application must adopt.
| Pattern | When the pattern fits | State and dependency ownership | Failure containment and completion rule |
|---|---|---|---|
| Sequential coordination | A later stage needs the accepted result of an earlier stage. | The workflow controller owns the ordered stages and passes a defined handoff result. | Stop dependent stages after a failed handoff; complete only when every required stage passes. |
| Concurrent coordination | Work can proceed independently and its results can be collected afterwards. | Each worker has a bounded assignment; the collector owns aggregation and shared-result updates. | Bound active workers, retain per-branch failures and apply the stated collection rule before completion. |
| Coordinator-led delegation | The required specialist or next task depends on intermediate findings. | The coordinator maintains the task ledger and delegates within an approved roster and scope. | Limit delegation and revision loops; the final validator, not a specialist’s confidence, determines completion or escalation. |
Sequential coordination makes dependencies explicit but does not remove failure propagation: accepting a poor intermediate result can corrupt every later stage. Concurrent coordination is appropriate only when the work really is independent. Giving several workers permission to rewrite the same record is not independence.
Coordinator-led delegation changes who chooses the next assignment. It does not remove the need for an enforceable permission boundary. Keep the coordinator’s ability to propose a plan separate from the application’s authority to approve a tool operation. A specialist should not gain broader credentials merely because the coordinator asks it to do more.
Combine patterns only at a named boundary. A sequential stage may dispatch independent analyses concurrently, then wait for its collector to return an accepted result. Document that join explicitly so the outer workflow cannot mistake an incomplete branch for a finished stage.
Give state and handoffs an explicit contract
For each task, record its identity, current status, owning component, required dependencies, accepted input version, result location and stopping conditions. Link attempts to the same logical task rather than treating every retry as unrelated work. These fields form the application’s design contract; exact database schemas and framework APIs depend on its implementation.
A handoff should say what work is requested, which context is relevant, which tools are permitted and what result format is accepted. The receiving worker returns a structured result or failure linked to the assignment. It should not silently change the original goal, widen its access or overwrite an accepted result because it received newer conversational instructions.
Choose who may commit shared state. Workers can submit candidate results while the controller or a designated collector performs the accepted transition. Record the input version used by a result so that a late response can be recognised as belonging to an older task state. This makes conflict handling a visible application decision rather than whichever write happens last.
Completion needs its own rule. Separate “the worker replied”, “the output passed validation” and “the authorised external action completed”. A run can finish generating text while an action is still awaiting approval. Report that state accurately instead of presenting all three events as success.
Contain retries, timeouts and duplicate actions
Retry a transient failure only within an explicit attempt and time limit. A permission denial or invalid request calls for correction or escalation, not repeated execution. Keep retry ownership at a layer that understands the task; nested retries can multiply work and delay failure reporting.
A timeout is not proof that an external action failed. The destination may have completed it before its response was lost. Restarting the worker can repeat the action. Use the destination’s idempotency mechanism where applicable, or reconcile the recorded operation with its destination before deciding to repeat it. An unresolved outcome should pause for investigation.
Apply this failure worksheet to the actual operations in your design. It prescribes responses to review and exercise in your own environment; it does not report incidents or measurements from a deployed system.
| Failure condition | Required evidence | Bounded response |
|---|---|---|
| Duplicate task arrives | Logical task identity, prior attempt and recorded action outcome | Reuse the recorded result or reconcile an unresolved action; do not automatically issue the external change again. |
| Worker times out | Assignment, last accepted checkpoint, elapsed deadline and outstanding action | Stop further dispatch for that attempt; distinguish incomplete computation from an unknown external outcome before retrying. |
| Workers return conflicting outputs | Input versions, separate results and the acceptance rule | Preserve both results, reject an unsupported merge and request a bounded re-evaluation or human decision. |
Recovery is not always a rollback. Some actions need a business-specific compensating operation; others cannot meaningfully be undone. Record those distinctions before granting write access. Compensation also needs progress tracking and can fail, so preserve the original action record and escalate when recovery remains unresolved.
Bound concurrency, usage and waiting time
Set limits at the workflow level as well as the worker level. A worker’s local iteration cap does not constrain a coordinator that keeps creating fresh workers. Define admission limits, maximum active assignments, delegation depth or allowed transitions, and an overall run deadline. Make reaching a limit produce an explicit stopped or escalated state.
Budget controls belong beside those execution controls. Track model usage, tool calls and retries against the application’s configured allowances. Decide which component can authorise further work and require a deliberate decision to increase the allowance. A worker must not reset the parent workflow’s budget by opening another task.
Set those limits from the application’s actual usage and approved operating budget. Keep recorded consumption attached to the parent task so an operator can see what exhausted its allowance. The useful design output is an enforceable boundary with an owner and a stop action, rather than an agent count treated as a proxy for cost.
Validate completion and make escalation actionable
Validate required output structure and the task’s substantive acceptance conditions before accepting a result. Keep quality checks distinct from permission checks: an accurate recommendation may still request an unauthorised action. Equally, an authorised tool call can return an unusable result that must not be passed to the next stage.
An escalation should carry the task identifier, the blocked decision, the relevant result or error, previous attempts and the exact action requiring approval. Preserve the paused state. Approval must apply to that action and context, not become a permanent permission for unrelated future work.
Use correlated logs to follow the entire path from intake to completion or failure. Record dispatch, result acceptance, policy decisions, approval, external action outcomes and budget exhaustion. Restrict sensitive payloads and credentials in those records. Operators need enough evidence to diagnose the boundary that failed, not an indiscriminate copy of every secret or document the agent encountered.