Agent Concepts 06 โ Evals, Reliability & Cost
Published:
Outcome vs trajectory evals, LLM-as-judge calibration, SWE-bench / tau-bench / WebArena, reliability patterns, guardrail layers, and cutting $/task.
Published:
Outcome vs trajectory evals, LLM-as-judge calibration, SWE-bench / tau-bench / WebArena, reliability patterns, guardrail layers, and cutting $/task.
Published:
Supervisor vs pipeline vs handoffs, failure modes unique to multi-agent, typed handoffs, self-improving loops โ and when to stay single-agent.
Published:
End-to-end RAG, chunking tradeoffs, why hybrid search + rerank, retrieval metrics, and when agentic retrieval beats single-shot RAG.
Published:
Why context is the working memory, compaction strategies, episodic/semantic/procedural memory, and mitigating โlost in the middleโ.
Published:
Function calling mechanics, tool design principles, MCP vs bespoke integrations, error-handling contracts, and schema-constrained decoding.
Published:
Agent vs chatbot vs workflow, ReAct / Plan-and-Execute / Reflexion compared, the autonomy spectrum, and the โwhen not to build an agentโ decision framework.