Agent Concepts 05 โ Multi-Agent Orchestration
Published:
Q1: What are the main multi-agent orchestration patterns?
Difficulty: Medium ยท Frequency: โ โ โ โ โ
Key points:
- Sequential pipeline: A โ B โ C, each specialized (e.g., spec compiler โ navigator โ comparator). Simple, debuggable, each stage independently testable.
- Supervisor / router: a central agent decomposes the task and delegates to workers, then synthesizes. Good for heterogeneous subtasks.
- Hierarchical teams: supervisors with sub-supervisors; scales to complex organizations of agents.
- Handoff networks: agents pass control to each other dynamically (like a relay) โ flexible, harder to trace.
- Debate / ensemble: multiple agents propose, then critique/vote โ quality via redundancy, at 3ร cost.
- Blackboard: shared state all agents read/write โ powerful, coordination-heavy.
- Communication substrate matters as much as topology: direct messages vs shared state vs schematized handoffs.
Follow-ups:
- โCode review vs open-ended research โ which pattern and why?โ โ Code review: pipeline (lint โ test โ review โ summarize โ verifiable stages). Research: supervisor + workers (parallel exploration, synthesis at the end).
Q2: What goes wrong in multi-agent systems that doesnโt in single-agent ones?
Difficulty: Medium-Hard ยท Frequency: โ โ โ โ
Key points โ failure modes unique to multi-agent:
- Error amplification across hops: a 95%-accurate stage ร 4 stages โ 81% end-to-end. Every handoff multiplies error.
- Context duplication cost: each agent re-reads shared context โ token cost scales with agent count.
- Responsibility diffusion: โsomeone else will verifyโ โ nobody does. (The fix: explicit verification ownership per stage.)
- Schema drift: upstream changes its output format, downstream silently misparses.
- Deadlock / livelock: agents waiting on each other, or ping-ponging a task back and forth.
- Coordination overhead exceeding the parallelism benefit.
Follow-ups:
- โHow do you debug a 5-agent pipeline failure?โ โ Per-layer tracing with structured inputs/outputs logged; find the first layer whose output diverged from its contract (not where the symptom appeared). This is why typed handoffs matter.
Q3: Schema-constrained handoffs โ what are they and why do they matter?
Difficulty: Medium-Hard ยท Frequency: โ โ โ โ
Key points:
- Agents pass typed, validated payloads (JSON schemas, protobufs) between each other โ not free text.
- Why: contracts between layers enable independent testing (โI can test the comparator without running the navigatorโ), versioning, and precise error attribution.
- Example: a UI-verification pipeline where the spec compiler emits
{screen_id, elements[], assertions[]}, the navigator emits{action_trace[], screenshots[]}, the comparator emits{verdict, diffs[]}โ each stageโs output is machine-checkable. - The deeper principle: deterministic interfaces between non-deterministic components. You canโt make LLMs deterministic, but you can make their contracts strict.
Follow-ups:
- โWhat breaks when agents pass raw text?โ โ Parsing failures, silently dropped fields, format drift over time, and untestable stages. It works in demos and rots in production.
Q4: What are self-improving agent loops? Sketch a concrete design.
Difficulty: Hard ยท Frequency: โ โ โ โ (Staff-level differentiator)
Key points:
- The loop: agent acts โ evaluator scores โ optimizer improves the agent (prompts, skills, few-shot examples โ or weights via methods like STaR). DSPy and TextGrad are the canonical frameworks: optimize the program, not just the prompt.
- Concrete design (real pattern): a 3-agent migration system โ
- Converter performs code conversion and emits key metrics (success rate, error classes);
- Brain consumes those metrics to open PRs improving the converterโs prompts, skills, and MCPs, and refines its own queries;
- Feeder surfaces additional signals (new failure modes, repo changes) to trigger further learning.
- Requirements: a trustworthy eval signal (garbage signal โ garbage improvement), versioning of every prompt/skill change, and rollback when the loop regresses.
Follow-ups:
- โWhat prevents reward hacking?โ โ Held-out eval sets the optimizer never sees, human spot-checks on a sample, and tracking multiple metrics (not a single number the loop can game).
Q5: When is multi-agent worse than a single agent?
Difficulty: Medium ยท Frequency: โ โ โ โ (judgment โ interviewers score the restraint)
Key points:
- Multi-agent adds latency (sequential hops), cost (context duplicated per agent), and complexity (coordination, contracts, debugging).
- Single agent wins for: narrow tasks, tasks where one context suffices, early prototypes.
- Multi-agent wins when: roles need different tools (a navigator needs a browser; a comparator needs vision), work is parallelizable, or subtasks are independently verifiable.
- The Staff answer: โStart single. Split when you can name the bottleneck the split removes โ specialization, parallelism, or verifiability. Never split for aesthetics.โ
Follow-ups:
- โYou inherited a 6-agent system for a 2-step task โ what do you do?โ โ Measure per-stage value-add; collapse stages that donโt earn their coordination cost; keep the split only where stages have distinct tools or verifiable contracts.
Share on
Twitter Facebook LinkedInโ Buy me a coffee! ๐
If you found this article helpful, consider buying me a coffee to support my work! ๐
