Agent Concepts 01 โ Agent Foundations
Published:
Q1: What is an AI agent, precisely? How does it differ from a chatbot, a workflow, and an โautonomous systemโ?
Difficulty: Easy-Medium ยท Frequency: โ โ โ โ โ (opener in almost every agent round)
Key points an interviewer wants to hear:
- An agent = an LLM (or model) in a perceptionโaction loop: it observes state, reasons, takes actions via tools, observes results, and repeats โ pursuing a goal across multiple steps.
- A chatbot responds to a single input with a single output; no tools, no persistent goal, no environment feedback.
- A workflow is a fixed DAG of steps (possibly with LLM calls inside); the path is predetermined, even if individual steps are non-deterministic.
- An agent chooses its own path dynamically at runtime based on observations. Thatโs the defining property: runtime control-flow decisions.
- Autonomy spectrum: copilot (human drives, AI suggests) โ supervised agent (AI acts, human approves checkpoints) โ autonomous agent with guardrails (AI acts within bounded envelopes, escalates on exceptions).
Follow-ups:
- โIs a 2-step ReAct loop an agent?โ โ Technically yes by the definition, but the interesting question is whether the dynamic path selection buys you anything over a workflow. Name the threshold: branching factor ร uncertainty.
- โWhat breaks when a team calls everything an agent?โ โ You lose the ability to reason about reliability: workflows get SLAs, agents get evals. Different engineering disciplines.
Q2: Explain the ReAct pattern. Why did interleaving reasoning and acting beat reasoning-only or acting-only approaches?
Difficulty: Medium ยท Frequency: โ โ โ โ โ
Key points:
- ReAct (Yao et al., 2022): Thought โ Action โ Observation loop. The model emits a reasoning trace, takes a tool action, reads the observation, and continues.
- Why it works: reasoning traces ground tool selection (the model commits to a rationale before acting); observations correct hallucinations mid-trajectory instead of at the end; it synergizes chain-of-thought with tool use.
- Empirically beat CoT-only and act-only baselines on HotpotQA/FEVER-style multi-hop tasks at the time.
- Limitations: error compounding over long trajectories, context bloat from verbose traces, no real long-horizon planning (itโs greedy/reactive), can loop forever without external bounds.
Follow-ups:
- โWhen does ReAct fail in production?โ โ Sparse or misleading tool observations, tasks needing upfront planning (multi-day migrations), cost blowup from long loops.
- โHow do you bound the loop?โ โ Max iterations, token/time budgets, stop conditions (goal check, no-progress detection), graceful fallback to human.
Q3: Compare ReAct vs Plan-and-Execute vs Reflexion. When would you pick each?
Difficulty: Medium-Hard ยท Frequency: โ โ โ โ
Key points:
- ReAct: interleaved, reactive, greedy. Strengths: adapts to surprises, simple to implement. Weaknesses: myopic, expensive on long tasks.
- Plan-and-Execute: generates an upfront plan, executes step by step, replans when steps fail. Strengths: better for long-horizon tasks with known structure, plan is inspectable. Weaknesses: brittle when the environment is highly dynamic; replanning is costly.
- Reflexion (Shinn et al., 2023): adds verbal reinforcement โ after a failed trial, the agent writes a self-critique into episodic memory and retries. No weight updates; learning happens in language. Strengths: improves across trials on the same task class. Weaknesses: needs a reliable success signal; critique quality bounds improvement.
- Selection heuristic: short/uncertain task โ ReAct; long/structured task โ Plan-and-Execute; repeated task class with clear success criteria โ Reflexion-style memory.
Follow-ups:
- โ30-minute coding task vs 5-minute QA task โ which pattern and why?โ โ Coding: plan-and-execute (or hierarchical) for structure + ReAct inside steps; QA: ReAct, cheap and reactive.
- โWhatโs the failure mode unique to Plan-and-Execute?โ โ Plan commitment: the agent follows a stale plan instead of reacting to new observations (mitigate with per-step verification).
Q4: What does โagenticโ behavior mean, and how do you measure it?
Difficulty: Medium ยท Frequency: โ โ โ
Key points:
- Observable properties: tool-use rate (actions per task), multi-step completion without human intervention, error recovery (does it fix its own mistakes?), autonomy under ambiguity (asks clarifying questions vs guessing).
- Benchmarks (know what each measures): SWE-bench (real GitHub issues โ code patches), ฯ-bench (tool-use consistency in customer-service-like tasks), WebArena (web navigation), GAIA (general assistant, multi-step reasoning), METRโs autonomy evals (long-horizon).
- Trajectory-level vs outcome-level: outcome = did the task succeed; trajectory = were the tool calls correct, efficient, safe. You need both โ a lucky success teaches nothing.
Follow-ups:
- โA model aces MMLU but fails agentic tasks โ why?โ โ Static QA tests knowledge recall; agentic tasks test sequential decision-making under uncertainty, tool grounding, and error recovery. Different capabilities, weakly correlated.
Q5: When should you NOT build an agent?
Difficulty: Medium ยท Frequency: โ โ โ โ (judgment question โ interviewers love this)
Key points:
- Agents are the most expensive, slowest, least reliable way to solve a problem. Default to simpler: rules โ workflow โ agent escalation ladder.
- Build an agent when: the task has high variance (inputs differ a lot), the branching is unknowable upfront, and the cost of a wrong step is manageable (or guarded).
- Donโt build one when: the path is deterministic, latency/cost budgets are tight, or errors are catastrophic and unguardable.
- Strong answer structure: โI ask three questions โ (1) can I enumerate the paths? (2) whatโs the cost of being wrong? (3) does the environment change mid-task?โ
Follow-ups:
- โGive an example where you replaced an agent with a workflow.โ โ Have one ready (e.g., a classification agent replaced by a small classifier + rules once the label space stabilized โ 10ร cheaper, more reliable).
Q6: What are the main failure modes of LLM agents in production?
Difficulty: Medium-Hard ยท Frequency: โ โ โ โ
Key points (name 5โ6 with a mitigation each):
- Compounding errors โ early small mistake snowballs โ mitigate with per-step verification, checkpoints.
- Tool misuse / hallucinated arguments โ schema validation, constrained outputs, dry-run modes.
- Infinite / unproductive loops โ iteration caps, no-progress detection, budgets.
- Context overflow / distraction โ compaction, structured state, retrieval instead of dumping.
- Prompt injection (especially via tool outputs) โ treat tool output as untrusted data, instruction hierarchy, output validation.
- Silent wrong answers (the worst) โ evals, verification steps, human sampling. A loud failure is a gift; a quiet one is a liability.
Follow-ups:
- โWhich is hardest to detect and why?โ โ Silent wrong answers โ everything looks green. This is why evals and verification agents exist.
- โYour agent worked in dev but fails in prod โ top 3 hypotheses?โ โ Data distribution shift (real tool outputs messier), latency/timeout differences, missing guardrails that dev never triggered.
Share on
Twitter Facebook LinkedInโ Buy me a coffee! ๐
If you found this article helpful, consider buying me a coffee to support my work! ๐
