
Author
Niyaz Puzhikkunnath
Niyaz Puzhikkunnath writes about evaluating and improving production AI agents at Papaya.
- AI agent evaluation
- Production agents
- Agent optimization
Articles and posts by Niyaz Puzhikkunnath
Agent evaluation
LLM-as-a-Judge for AI Agents: How to Build a Judge You Can Trust
LLM judges agree with people on preferences and struggle with correctness. How to build an agent judge from narrow, workflow-specific questions and calibrate it before you trust it.
Comparisons
Best AI Agent Evaluation Platforms for Startups (2026)
Free tiers, first paid steps, and startup programs for the leading agent evaluation platforms, and why most early teams need a few dozen reviewed runs before any platform.
Agent evaluation
How to Evaluate Multi-Step AI Agents Without Labeled Data
Agent runs are expensive to label and labels go stale. Six signals that work without them, how they mislead, and how to spend a small labeling budget where it counts.
Context engineering
Context Bloat: Too Much Context Makes Agents Worse
Agent prompts grow every turn: old messages, raw tool results, documents fetched just in case. This buildup makes agents slower, costlier, and less accurate. Here is how to trim it without losing information the agent still needs.
Agent evaluation
How to Evaluate AI Agents in Production: A Practical Framework
In one public benchmark we analyzed, 22% of runs confidently told the customer the job was done without ever calling the real API. A guide to evaluating agents beyond the pass rate.