AI / Agents LLM Archived

Cite as: Real Problem AI problem “Why does evaluating something as simple as 'clarity' still require running a thousand prompts by hand?”. Opportunity score 6.2 out of 10 (severity 7, AI feasibility 6, market signal 6, competition gap 6). Category AI / Agents. Trend LLM. Source signal: Ask HN: How are people doing AI evals these days?, Hacker News item 47319587, 17 Mar 2026.. Canonical URL: https://www.realproblem.ai/archive/why-is-measuring-a-thousand-prompts-still-the-only-way-to-eval-clarity.

Why does evaluating something as simple as 'clarity' still require running a thousand prompts by hand?

AI agent builders doing evals face outputs that aren't binary right-or-wrong, forcing expensive, high-volume manual testing just to get statistically meaningful signal on subjective quality dimensions.

Who has it: Indie AI builders and small teams trying to eval their own agents/prompts.

Evidence

Builders describe subjective outputs as hard to measure because they are not simply right or wrong, so testing each idea takes a large volume of prompts and costly human review.

Our summary of the public post linked below, not a quote. Nobody submitted it to Real Problem AI.

Ask HN: How are people doing AI evals these days?, Hacker News item 47319587, 17 Mar 2026.

Why it is archived

Trimmed to 100-cap (lowest opportunity_score)

Scoring breakdown

6.2/ 10
Problem Severity7
Feasibility today6
Market Signal6
Competition Gap6

Existing players

  • Promptfoo / Braintrust / DeepEval · Handle structured evals well, subjective/dimensional quality still needs heavy human review.
  • LLM-as-judge · Cheaper than humans but introduces its own noise and bias.

What they are missing

An eval workflow purpose-built for subjective, multi-dimensional quality (tone, clarity, helpfulness) that combines cheap LLM-judge triage with targeted, minimal human review instead of all-or-nothing manual scoring.

Stack hint

01LLM-as-judge scoring pipeline
02Active-learning sample selection for human review
03Multi-dimensional rubric scoring
04Statistical confidence reporting

#ATB24 · Canonical URL: https://www.realproblem.ai/archive/why-is-measuring-a-thousand-prompts-still-the-only-way-to-eval-clarity