Promptfoo is an open-source CLI for testing, evaluating, and comparing LLM prompts. This guide covers writing test cases in YAML, running evaluations, comparing models, and catching prompt regressions in CI.
Promptfoo automates prompt evaluation: you define prompts, test cases, and assertions in a YAML config, then run `promptfoo eval` to test every combination and report pass/fail. This makes prompt changes as testable as code changes.
Promptfoo supports: testing a single prompt against multiple inputs, comparing the same prompt across multiple models, asserting on specific outputs (exact match, regex, JSON schema, LLM-as-judge), and running evaluations in CI pipelines.
Install: `npm install -g promptfoo` or `npx promptfoo init`. Create a `promptfooconfig.yaml` in your project. Basic structure: `prompts` (your prompt templates), `providers` (which models to test), and `tests` (input variables and assertions).
Minimal config: `prompts: ['Classify the sentiment of: {{message}}. Respond with POSITIVE, NEGATIVE, or NEUTRAL.'], providers: ['openai:gpt-4o-mini', 'anthropic:claude-haiku-20240307'], tests: [{vars: {message: 'I love this product!'}, assert: [{type: 'contains', value: 'POSITIVE'}]}]`. Run with `promptfoo eval` to test both models against your assertion.
Assertion types: `contains` (output includes substring), `regex` (output matches pattern), `json-schema` (valid JSON matching schema), `llm-rubric` (LLM evaluates against a criterion), `python` (custom Python function), and `similar` (semantic similarity to expected output).
Use `llm-rubric` for subjective quality: `{type: 'llm-rubric', value: 'The response is helpful, accurate, and mentions specific next steps'}`. Promptfoo uses a separate evaluator model (GPT-4o by default) to grade the response against the rubric. This lets you evaluate qualities that are hard to assert programmatically.
Run `promptfoo eval --view` to open the results in a browser UI. The table shows each test case as a row and each model as a column, with pass/fail indicators and the actual output for each cell. This makes it easy to spot cases where one model excels and another fails.
For cost-quality tradeoffs, add `--metrics cost,latency` to include token cost and response time in the comparison. You may find that a $0.002/call model achieves the same pass rate as a $0.02/call model on your specific task — a 10× cost reduction with no quality loss.
Add Promptfoo to your CI pipeline to catch prompt regressions before they reach production. In GitHub Actions: `- name: Test prompts; run: npx promptfoo eval --ci`. The `--ci` flag exits with a non-zero code if any assertion fails, failing the pipeline.
Store your eval results as CI artifacts: `npx promptfoo eval --output results.json`. Compare results between the current branch and main to surface regressions. Set a minimum pass rate threshold: `--pass-rate-threshold 0.95` fails the pipeline if fewer than 95% of assertions pass.
LangSmith is LangChain's observability platform for logging, tracing, and evaluating LLM applications. This guide covers setup, automatic tracing, custom traces, and using the dashboard to debug production issues.
Read guideLLM API costs can grow unexpectedly as usage scales. This guide covers cost attribution, per-request tracking, anomaly detection, and practical techniques that reduce costs by 50–80% without sacrificing quality.
Read guideOpenTelemetry (OTel) is the vendor-neutral standard for distributed tracing. The GenAI semantic conventions extend it to LLM calls. This guide covers setting up OTel tracing for LLM applications, exporting to Jaeger or Grafana, and the GenAI conventions.
Read guide