PREVIEW · static UI mock · screen 26 / 32 · existing-trace flow step 3 · all screens
‹ Back to inferred workflow

Generated eval set

We extracted 32 representative cases from your 142 traces. Review, deselect any you don't want, and we'll run baseline.

1 Connect source
2 Inferred workflow
3 Generated eval set
Total cases
32
All selected by default
Golden path
14
Modal happy-path traces
Boundary cases
12
Edge inputs · uncommon shapes
Known failures
6
Escalated or follow-up <24h
32 selected · estimated baseline cost $1.18 · ~4 min
Case
Type
Source trace
Expected
Single SKU, in-stock, standard payment terms
From trc_84e1a2 · clustered with 38 similar traces
Golden path
trc_84e1a2
order accepted
#ec-001
Multi-line order, mixed availability
From trc_84e1c7 · clustered with 22 similar traces
Golden path
trc_84e1c7
order accepted, 1 backorder
#ec-002
Net-30 customer requesting net-15
From trc_84e2f1 · payment-term mismatch path
Boundary
trc_84e2f1
escalate to human
#ec-003
Customer record not found, name-only match
From trc_84e387 · unresolved customer cluster
Boundary
trc_84e387
escalate to human
#ec-004
Bulk order, tiered pricing applied
From trc_84e44b · 14 similar traces
Golden path
trc_84e44b
order accepted, tier-2 price
#ec-005
Order accepted but customer followed up <1h later asking to change
From trc_84e511 · failure signal · 6 similar
Known failure
trc_84e511
should have asked clarifying Q
#ec-006
Inventory lookup timed out, agent retried
From trc_84e5d3 · tool-failure path
Boundary
trc_84e5d3
order accepted on retry
#ec-007
Email order with ambiguous SKU description
From trc_84e601 · 8 similar traces
Known failure
trc_84e601
picked wrong SKU variant
#ec-008
Late-night order, after-hours pricing rule
From trc_84e6a8 · 4 similar traces
Boundary
trc_84e6a8
order accepted, after-hours
#ec-009
Customer requested expedite, no expedite SKU exists
From trc_84e7c2 · escalation path
Boundary
trc_84e7c2
escalate with note
#ec-010
Show 22 more cases →
Sample case detail · #ec-006
Input (extracted from trace)
channel: chat customer_id: cust_84211 message: | hey can I get 200 of the red model 4 things asap same terms as last time
Expected outcome (graded by)
grader: llm_judge_v2 must_satisfy: - asks_clarifying_question - confirms_quantity - confirms_sku_variant must_not: - finalizes_order_without_confirmation
Graders are auto-generated from the trace pattern. Edit this grader →
Next: baseline run
We'll run your current agent against all 32 cases and record the score. That becomes the floor — every future change has to clear it. Estimated cost $1.18, runtime ~4 min. After baseline, you'll land on the workflow dashboard.