LangSmith Studio: Agent Evaluation and Monitoring

Ship LLM agents blind and you'll lose money. LangSmith Studio gives you visibility into what your agent actually does, catches problems before users do.

When you ship LLM agents to production, you quickly learn that testing in your local environment tells you almost nothing about what actually happens. I've dealt with this myself — you can have a perfectly working agent in development, then it goes into production and suddenly you're losing money on failed runs, latency spikes, or agents taking five steps when they should take one.

LangSmith Studio changes this. It gives you visibility into what your agent is actually doing, lets you evaluate its performance with real datasets, and catches problems before your users do. This is not theory — this is what we're running in production right now.

Why LangSmith Matters for Agent Evaluation

Let me be direct: without proper evaluation and monitoring, you're flying blind. Your agent could be:

  • Calling tools unnecessarily (burning through API costs)
  • Getting stuck in loops on certain inputs
  • Degrading in quality without you noticing
  • Making expensive mistakes that only show up after 10k calls

LangSmith Studio solves this by giving you:

  1. Complete tracing — Every tool call, LLM invocation, and decision is logged with full context
  2. Cost visibility — Not just LLM costs, but retrieval, tools, external APIs — the full picture
  3. Evaluation frameworks — Compare agent versions against datasets with consistent metrics
  4. Production monitoring — Alerts, regression detection, and performance tracking
  5. LangSmith Engine — AI-powered analysis that suggests what's breaking and why

The key insight: evaluation isn't just for testing. It's how you maintain quality over time as your agent evolves.

Setting Up LangSmith Studio

First, install the dependencies:

pip install langsmith langchain langchain-openai langgraph

Get your API key from langsmith.com and set it:

export LANGSMITH_API_KEY="your_key_here"
export LANGSMITH_PROJECT="your-project-name"

Building a Traced Agent

Here's a realistic agent setup that we actually use:

from langchain_openai import ChatOpenAI
from langchain.tools import tool
from langgraph.graph import StateGraph, END
from langgraph.prebuilt import ToolNode
from langsmith import trace, get_tracer
from typing import TypedDict, Annotated
import anthropic_sdk

class AgentState(TypedDict):
    messages: list
    tool_results: list
    cost: float
    step_count: int

@tool
def search_documentation(query: str) -> str:
    """Search internal documentation for relevant information."""
    # Simulate search with proper instrumentation
    results = f"Found 3 docs matching '{query}'"
    return results

@tool
def fetch_user_data(user_id: str) -> dict:
    """Fetch user context for personalization."""
    # In production, this hits your database
    return {
        "user_id": user_id,
        "subscription": "pro",
        "usage_this_month": 450
    }

@tool
def create_summary(content: str, max_length: int = 200) -> str:
    """Summarize content to specified length."""
    return f"Summary of {len(content)} chars to {max_length} max"

# Initialize LLM with tracing enabled
llm = ChatOpenAI(model="gpt-4", temperature=0.7)
tools = [search_documentation, fetch_user_data, create_summary]

def agent_node(state: AgentState):
    """Main agent reasoning step."""
    tracer = get_tracer()
    
    # Each step is traced automatically
    with tracer.trace("agent_reasoning") as run:
        response = llm.invoke(
            state["messages"],
            tools=tools,
        )
        
        run.add_tags(["step_count:" + str(state["step_count"])])
        return {
            "messages": state["messages"] + [response],
            "step_count": state["step_count"] + 1,
            "cost": state.get("cost", 0) + 0.002  # Track cumulative cost
        }

def should_continue(state: AgentState):
    """Decide if we need more steps or we're done."""
    if state["step_count"] > 10:
        return END
    
    last_message = state["messages"][-1]
    if hasattr(last_message, "tool_calls") and last_message.tool_calls:
        return "tools"
    return END

# Build the graph
workflow = StateGraph(AgentState)
workflow.add_node("agent", agent_node)
workflow.add_node("tools", ToolNode(tools))

workflow.add_edge("tools", "agent")
workflow.add_conditional_edges(
    "agent",
    should_continue,
    {"tools": "tools", END: END}
)
workflow.set_entry_point("agent")

agent = workflow.compile()

# Run with automatic tracing
result = agent.invoke({
    "messages": ["What's our top customer's usage?"],
    "tool_results": [],
    "cost": 0,
    "step_count": 0
})

print(f"Result: {result['messages'][-1]}")
print(f"Total steps: {result['step_count']}")
print(f"Estimated cost: ${result['cost']:.4f}")

Notice what's happening here:

  • Every agent step is automatically traced
  • Tool calls are captured with full context
  • We're tracking cost as we go
  • The tracer adds metadata (tags) for filtering later

In LangSmith Studio, you'll see each run as a tree. Click into any node and you see the exact prompt, model response, and latency.

Custom Evaluation Metrics

Tracing gives you visibility. Evaluation gives you benchmarks. Here's how we set up production-grade evaluation:

from langsmith import evaluate
from langsmith.evaluation import LangChainStringEvaluator
from langsmith.schemas import Example, Run

# Define your evaluation dataset
eval_dataset = [
    {
        "input": "What's our top customer's usage?",
        "expected_output": "Pro customer with 450 units used this month",
        "category": "data_retrieval"
    },
    {
        "input": "Summarize our Q3 performance",
        "expected_output": "Summary should mention revenue and key metrics",
        "category": "summary"
    },
    {
        "input": "Create a plan for scaling to 10x users",
        "expected_output": "Should consider infrastructure, costs, and strategy",
        "category": "planning"
    }
]

# Define custom evaluation logic
def evaluate_cost_efficiency(run: Run, example: Example) -> dict:
    """Check if the agent solved the task without excessive steps."""
    cost = run.outputs.get("cost", 0)
    steps = run.outputs.get("step_count", 0)
    
    # Cost efficiency: should solve in 2-4 steps, cost < $0.01
    is_efficient = steps <= 4 and cost < 0.01
    
    return {
        "key": "cost_efficiency",
        "score": 1.0 if is_efficient else 0.0,
        "comment": f"Used {steps} steps, cost ${cost:.4f}"
    }

def evaluate_accuracy(run: Run, example: Example) -> dict:
    """Check if output matches expected result."""
    output = run.outputs.get("messages", [])[-1]
    expected = example.outputs.get("expected_output", "")
    
    # Simple substring match (in production, use semantic similarity)
    accuracy = 1.0 if expected.lower() in str(output).lower() else 0.0
    
    return {
        "key": "accuracy",
        "score": accuracy,
        "comment": f"Output relevance: {'matched' if accuracy else 'missed'}"
    }

# Run evaluation suite
experiment_results = evaluate(
    agent.invoke,
    data=eval_dataset,
    evaluators=[evaluate_cost_efficiency, evaluate_accuracy],
    experiment_prefix="agent_v2",
)

print(f"Cost efficiency: {experiment_results['cost_efficiency'].get('pass_rate', 0):.1%}")
print(f"Accuracy: {experiment_results['accuracy'].get('pass_rate', 0):.1%}")

Here's what this does in practice:

  • Runs your agent against 3 test cases
  • Measures both cost efficiency and accuracy
  • Compares results across agent versions ("agent_v2", "agent_v3", etc.)
  • Catches regressions immediately

When you run this, LangSmith stores the results. In the UI, you can compare agent_v1 vs agent_v2 and see exactly which test cases got better or worse.

Production Monitoring and Alerting

Evaluation is offline testing. For production, you need continuous monitoring:

from langsmith.client import Client
from langsmith.utils import get_run_url
import time
from datetime import datetime, timedelta

client = Client()

def monitor_agent_health(project_name: str, check_interval: int = 300):
    """
    Monitor agent performance in production.
    Runs every check_interval seconds (default 5 minutes).
    """
    while True:
        try:
            # Get runs from last 5 minutes
            time_window = datetime.now() - timedelta(minutes=5)
            runs = client.list_runs(
                project_name=project_name,
                filter=f"create_time > {time_window.isoformat()}"
            )
            
            runs_list = list(runs)
            if not runs_list:
                print(f"[{datetime.now()}] No runs to analyze")
                time.sleep(check_interval)
                continue
            
            # Calculate metrics
            total_runs = len(runs_list)
            failed_runs = sum(1 for r in runs_list if r.error)
            avg_latency = sum(r.end_time - r.start_time 
                            for r in runs_list if r.end_time) / len(runs_list)
            
            failed_rate = failed_runs / total_runs if total_runs > 0 else 0
            
            print(f"\n[{datetime.now()}] Agent Health Report")
            print(f"  Runs: {total_runs}")
            print(f"  Failures: {failed_runs} ({failed_rate:.1%})")
            print(f"  Avg latency: {avg_latency:.2f}s")
            
            # Alert conditions
            if failed_rate > 0.05:  # > 5% failure rate
                print(f"  ⚠️  ALERT: High failure rate ({failed_rate:.1%})")
                # Send to monitoring system
                alert_slack(f"Agent failure rate: {failed_rate:.1%}")
            
            if avg_latency > 10:  # More than 10 seconds
                print(f"  ⚠️  ALERT: High latency ({avg_latency:.2f}s)")
                alert_slack(f"Agent latency spike: {avg_latency:.2f}s")
            
            # Cost tracking
            total_cost = calculate_total_cost(runs_list)
            print(f"  Cost (5min window): ${total_cost:.2f}")
            
        except Exception as e:
            print(f"Error in monitoring: {e}")
        
        time.sleep(check_interval)

def calculate_total_cost(runs: list) -> float:
    """Sum up costs from all runs."""
    return sum(
        run.outputs.get("cost", 0) 
        for run in runs 
        if run.outputs
    )

def alert_slack(message: str):
    """Send alert to Slack (implement your integration)."""
    # This is where you'd integrate with Slack, PagerDuty, etc.
    print(f"  📢 Slack: {message}")

# Start monitoring
if __name__ == "__main__":
    monitor_agent_health("production-agent", check_interval=300)

This runs in a background process (or on a schedule) and continuously checks your agent's health. When failure rate spikes or latency increases, you get alerted before customers complain.

Using LangSmith Engine for Auto-Optimization

Here's where it gets interesting. LangSmith Engine (new in 2026) analyzes your traces and suggests improvements:

from langsmith import Client
from langsmith.client import LangSmithApi

client = Client()

def analyze_failing_runs(project_name: str, limit: int = 10):
    """
    Let LangSmith Engine analyze why runs are failing.
    This uses AI to suggest fixes.
    """
    # Get recent failed runs
    failed_runs = client.list_runs(
        project_name=project_name,
        filter="error != null",
        limit=limit
    )
    
    print("LangSmith Engine Analysis:")
    print("=" * 50)
    
    for run in failed_runs:
        print(f"\nRun: {run.id}")
        print(f"Error: {run.error}")
        
        # LangSmith Engine analyzes this and suggests fixes
        # In the UI, you'll see AI-generated suggestions
        # Here we're just tracking it
        
        # Tag this run for manual review
        client.update_run(
            run.id,
            tags=list(set(run.tags or []) | {"needs_review", "engine_analysis"})
        )

def suggest_prompt_improvements(project_name: str):
    """
    Let LangSmith Engine suggest prompt improvements
    based on failing runs and low evaluation scores.
    """
    # In LangSmith UI:
    # 1. Go to Prompts tab
    # 2. Click "Get AI Suggestions"
    # 3. Engine analyzes your runs and suggests prompt changes
    # 4. Test suggestions against your evaluation dataset
    # 5. Deploy improved version
    
    print("Prompt improvement workflow:")
    print("1. Analyze failing runs in LangSmith Studio")
    print("2. Engine suggests prompt refinements")
    print("3. Test against evaluation dataset")
    print("4. Deploy to production if metrics improve")

# Usage
analyze_failing_runs("production-agent", limit=20)

In the LangSmith UI, when you click "Engine Analysis", it runs AI over your trace data and tells you things like:

  • "This run failed because the agent didn't find the right tool"
  • "Latency spiked because of a retry loop in tool X"
  • "Cost is high because the summary tool is being called twice"

Then you can iterate on the prompt or agent logic directly in LangSmith Playground and test it against your eval dataset.

Real-World Patterns We Use

Pattern 1: Version Comparison

# v1: Simple agent
result_v1 = agent_v1.invoke(inputs)

# v2: Enhanced with retrieval
result_v2 = agent_v2.invoke(inputs)

# LangSmith automatically tags these differently
# In the UI, go to: Experiments → Compare versions
# See side-by-side: cost, latency, accuracy for each

Pattern 2: Production Gradual Rollout

import random

def route_to_agent_version(request):
    """Route requests to agent versions for A/B testing."""
    
    if random.random() < 0.1:  # 10% to new version
        agent = agent_v2  # New, faster agent
        tags = ["version:v2", "canary"]
    else:
        agent = agent_v1  # Proven version
        tags = ["version:v1", "stable"]
    
    result = agent.invoke(request)
    
    # LangSmith tags this run so you can track v1 vs v2 metrics
    return result

Then in LangSmith, filter by version:v2 and version:v1 tags to compare real production metrics.

Pattern 3: Cost Attribution by Tool

def track_cost_per_tool(run: Run):
    """Break down costs by tool for optimization."""
    tool_costs = {}
    
    for trace in run.child_runs:
        tool_name = trace.name
        # Estimate based on tokens or API calls
        cost = estimate_cost(trace)
        tool_costs[tool_name] = tool_costs.get(tool_name, 0) + cost
    
    print("Cost breakdown:")
    for tool, cost in sorted(tool_costs.items(), key=lambda x: x[1], reverse=True):
        print(f"  {tool}: ${cost:.4f}")
    
    # Find expensive tools to optimize
    return tool_costs

This shows you exactly which tools are burning money. In production, we found that one retrieval call was being made unnecessarily, and fixing it saved 30% on costs.

Best Practices for Production

  1. Always tag your runs — Use tags like version:vX, environment:prod, user_tier:enterprise. This makes filtering and analysis trivial.
  2. Include metadata in the state — Track user_id, session_id, feature flags. When something breaks, you can correlate it with who was affected.
  3. Monitor cost continuously — Not just total cost, but cost per request. A 10% spike in average cost per request means something is wrong.
  4. Evaluation dataset is sacred — Your eval dataset should grow as you find edge cases. Every bug should become a test case.
  5. Use production data for improvement — Take your most interesting failing runs and add them to your eval dataset. This prevents regressions.
  6. Compare before deploying — Before pushing a new agent version, run it against your full eval dataset. If you lose on accuracy or cost, don't deploy.

Wrapping It Up

LangSmith Studio shifts you from "hope my agent works" to "I know exactly what my agent is doing and can prove it's better." The tracing is automatic, the evaluation framework catches regressions, and the production monitoring means you sleep better at night.

Start with basic tracing and evaluation. Once you have that working, add production monitoring. Then iterate on improvements using the Engine's suggestions and your eval dataset.

The agents that make money are the ones where you can answer: "What does this run cost? Why did it fail? Is the new version actually better?" LangSmith Studio gives you the tools to answer all three.