pricing

From first score to
prompt in production.

Evaluate for free. Gate regressions in CI and serve production prompts when you ship.

personal use
Free
Free

See your first prompt score in 60 seconds. No credit card.

✓ current plan
3 web evals/monthi
Eval API — lint 10/mo + BYOKi
Issues, warnings and strengthsi
Library — up to 5 promptsi
Free daily trainingi
web evaluation · prompts up to 8,000 characters
personal use
Basic
$9/month

Fix prompts before they hurt users in production. 30×/month.

cancel anytime · no lock-in

Improved prompti
30 credits/monthi
Production iteratori
Playground (BYOK) + versioned libraryi
Full daily training + tracksi
Eval API — lint 30/mo + BYOKi
web evaluation · prompts up to 12,000 characters
best value
◆ in production
Pro
$19/month
≈ $0.63/day · cancel anytime

Ship with a safety net: serving + CI gate.

cancel anytime · no lock-in

Everything in Basici
Slug serving (no redeploy)i
CI regression gate + GitHub Actioni
Full eval API — lint 75/moi
Unlimited web usagei
Batch A/B testingi
web evaluation · prompts up to 35,000 characters
◆ in production
Team
Custom

Govern prompts as a team — roles, approval, audit.

Contact us →
Everything in Proi
Workspaces + rolesi
Production approval workflowi
Audit logi
Eval API — lint 250/mo + BYOKi
Library export (JSON/CSV)i
Priority support (24h)i
web evaluation · prompts up to 60,000 characters

Output shape specification works because you're constraining the whole output distribution at once. Scripted edge-case responses are high-durability: pre-built templates survive ambiguity better than abstract rules.

— CodeMaitre · Reddit · came in skeptical, came out convinced

2,114 prompts evaluated · 77,601 tokens saved

Your prompts are processed and discarded — never used to train AI models.

What does PromptEval do beyond a chat (ChatGPT/Claude)?

A chat gives conversational, memoryless suggestions. Here you get a reproducible score (8 sub-criteria at temperature 0 against an anchored rubric — same prompt, same number) and, on top of it, what a chat doesn't have: versioning with diffs, a CI regression gate, and serving the production prompt by slug.

Is the score reliable? How is it calculated?

8 sub-criteria across 4 dimensions (clarity, specificity, structure, robustness), each at temperature 0 against an explicit rubric: below 60 = serious gaps, above 85 = genuinely robust. It adjusts ±8 for technical factors like instruction positioning (U-shaped attention weights the start and end more) and system/user separation. Structured analysis, not an opinion.

How does the Free plan work?

3 web evaluations per month, auto-renewed, no credit card. Includes score, 4 dimensions, issues and warnings, library up to 5 prompts, and 10 API lint calls/month (unlimited BYOK).

Are my prompts private?

Yes. They're sent to Claude for evaluation and discarded after processing — never used to train models. Everything is stored with Row Level Security: only your account can access it.

Can I switch plans or cancel anytime?

Yes. Switch plans or cancel anytime from the customer portal (Stripe), no lock-in and no fees. Your access continues until the end of the period you already paid for.

Can I use it in my CI/CD?

Yes, on Pro+. A REST API (POST /api/v1/eval) plus an official GitHub Action that fails the PR if the score drops, instructions contradict, or it regresses vs production. Lint mode is open on every plan; full needs Pro/Team or BYOK.

Do I need a redeploy to change a production prompt?

No (Pro+). Give the prompt a slug and serve the production version via GET /api/v1/prompts/{slug}. Change it in the library and it takes effect in ~60s, no deploy.

What's the difference between web credits and API quota?

They are separate meters. Web credits cover the evaluator, iterator and playground on the site (Free 3, Basic 30; Pro and Team unlimited). The API quota is only for HTTP calls (lint 10/30/75/250 per month). And BYOK is unlimited on both — it runs on your key.

secure payments via Stripe · prices in USD · cancel anytime · questions? get in touch