Articles
On agent optimization.
Deeper writing on what makes production LLM agents reliable — context economy, tool quality, verification design, and the patterns behind continuous improvement.
Comparisons
Langfuse Alternatives (2026): 7 Tools Compared
Why teams leave Langfuse, what each alternative is best at, verified pricing and licensing as of October 2026, and when Langfuse is still the right call.
Macy Mody· October 1, 2026Comparisons
LangSmith Alternatives (2026): 7 Tools Compared
Per-seat pricing, retention, self-hosting, and framework fit: why teams leave LangSmith, what Engine does that you would give up, and the best alternative for each situation.
Macy Mody· October 1, 2026Comparisons
Arize Phoenix Alternatives (2026): 7 Tools Compared
Phoenix and AX are different products. The Elastic License, the Dynatrace acquisition, and the best alternative for each reason to move, verified as of October 2026.
Faiz Vadakkumpadath· October 1, 2026Comparisons
Braintrust Alternatives (2026): 7 Tools Compared
The $249 jump, data-plus-scores metering, and enterprise-only self-hosting: why teams leave Braintrust, the best alternative for each reason, and when Braintrust is still right.
Macy Mody· October 1, 2026Agent evaluation
LLM-as-a-Judge for AI Agents: How to Build a Judge You Can Trust
LLM judges agree with people on preferences and struggle with correctness. How to build an agent judge from narrow, workflow-specific questions and calibrate it before you trust it.
Niyaz Puzhikkunnath· September 29, 2026Agent optimization
How to Turn Agent Traces into Improvements
Most teams collect traces and never act on them. The agent improvement loop that turns traces into ranked, shipped, and verified fixes, and what those fixes usually look like.
Macy Mody· September 25, 2026Agent reliability
AI Agent Failure Patterns: A Field Guide from Production Traces
False completion, retry loops, ignored tool errors, context bloat, skipped verification. The ten patterns behind most production agent failures, as they actually appear in traces.
Macy Mody· September 23, 2026Agent reliability
How to Detect AI Agent Failures Automatically from Traces
Most agent failures never throw an error. A practical procedure for finding them in production traces, and an honest look at what automated detection can and cannot do.
Faiz Vadakkumpadath· September 17, 2026Comparisons
Best AI Agent Evaluation Platforms for Startups (2026)
Free tiers, first paid steps, and startup programs for the leading agent evaluation platforms, and why most early teams need a few dozen reviewed runs before any platform.
Niyaz Puzhikkunnath· September 14, 2026Agent evaluation
How to Evaluate Multi-Step AI Agents Without Labeled Data
Agent runs are expensive to label and labels go stale. Six signals that work without them, how they mislead, and how to spend a small labeling budget where it counts.
Niyaz Puzhikkunnath· September 11, 2026Agent reliability
How to Debug AI Agents in Production
Agent failures rarely leave a stack trace. How to read a trace backward to the first wrong decision, tell context bugs from tool bugs, and prove a fix works on real traffic.
Faiz Vadakkumpadath· September 8, 2026Comparisons
Best AI Agent Observability Tools in 2026
Ten agent observability tools compared on what matters in 2026: agent-aware tracing, automatic failure detection, evaluation, pricing units, and the best picks for startups.
Faiz Vadakkumpadath· September 3, 2026Context engineering
Context Bloat: Too Much Context Makes Agents Worse
Agent prompts grow every turn: old messages, raw tool results, documents fetched just in case. This buildup makes agents slower, costlier, and less accurate. Here is how to trim it without losing information the agent still needs.
Niyaz Puzhikkunnath· July 24, 2026AI observability
From Traces to Decisions: AI Observability Data Pipeline
Instrumentation makes agent behavior visible; data engineering makes it understandable at scale.
Faiz Vadakkumpadath· July 15, 2026Agent evaluation
How to Evaluate AI Agents in Production: A Practical Framework
In one public benchmark we analyzed, 22% of runs confidently told the customer the job was done without ever calling the real API. A guide to evaluating agents beyond the pass rate.
Niyaz Puzhikkunnath· July 14, 2026Article
Agent Optimization vs. Observability: Why Watching Isn't Fixing
Your dashboard saw the bad run and did nothing. The gap isn't between having data and not — it's between seeing what happened and knowing what to change.
Macy Mody· June 17, 2026