PREVIEW · static UI mock · screen 19 / 32 · workflow create step 3 — baseline complete · all screens
Describe
Review generated
3
Run baseline
Baseline complete
The agent ran against all 12 generated eval cases. You now have a starting metric and the improvement loop is ready to run on your production traces.
Composite score
85.8%
baseline · day 0
Eval cases passed
11 / 12
1 partial · #11 promo bundle
Total run time
3m 14s
parallel · 4 workers
Cost
$0.42
24,820 tokens · Sonnet 4.6
Per-case results 12 cases · 1 partial · 0 failures
1
PNW winter boot 2025-W47 (replay)
passed
14.2s
$0.04
2
Promo bundle uplift Q4 2024 (replay)
passed
11.8s
$0.03
3
Promo bundle uplift Q4 2025 (replay)
partial
12.1s
$0.04
4
Steady-demand baseline (no anomaly)
passed
8.4s
$0.02
5
Inventory-pegged forecast (low stock)
passed
9.1s
$0.03
6-12
Cold snap, supplier delay, hurricane evacuation, EOL, returns spike, holiday shift, new SKU
7 passed
avg 12.4s
$0.26
Your starting point
85.8%
Baseline · day 0
baseline 85.8% · improvements will climb above this line today
As the loop runs on your production traces, failure clusters surface and proposals appear in your inbox. Each approved improvement is gated against this 12-case suite plus any new cases promoted from clusters. The lift chart climbs from here.
The improvement loop will start running automatically on your next 100 traces. You'll see the first failure clusters within a few hours.