Skip to main content
The Evaluator Analysis page aggregates all transcriptions processed with a selected evaluator and organizes the results into KPI summaries, trend charts, a per-criterion evolution chart, a per-criterion performance ranking, a weekly quality heatmap, participant rankings, and a filterable transcription detail table. A period filter scopes the whole dashboard to a date range.
Each analysis generation consumes one unit of your account’s analytics_evaluator quota. Loading a previously cached result for the same evaluator ID does not consume quota.

Generating an analysis

Select an evaluator from the searchable selector at the top of the page and click Generate Analysis. The selector shows each evaluator’s name, ID, and criteria count.
Navigating directly to a URL that includes an evaluator ID (e.g. from a bookmark or shared link) auto-loads the last cached result — no quota consumed.
The Generate Analysis button is disabled when no evaluator is selected or when your quota is exhausted. If you click it without a selection, a warning is shown: “You must select an evaluator to generate the analysis.”

Dashboard header

Once data is loaded, the dashboard header shows:

Period filter

The first card of the dashboard — shown once an analysis has been run — scopes everything below it to a date range. If the start date is after the end date, Apply is disabled and a message appears: “The start date can’t be after the end date.”
Both Apply and Clear issue a real API call and consume one unit of analytics quota, exactly like Generate Analysis.
Navigating to another page and back keeps both the data and the From/To inputs. A browser reload does not — the dashboard re-queries unfiltered and the inputs come back empty.
There is no “quarter” mode — a quarter is just a range. For Q2 2026, set From 2026-04-01 and To 2026-06-30.
The filter matches on each transcription’s recording date (which the recording date controls set at upload) — so calls uploaded late but dated correctly land in the right period.

KPI Grid

Four metric cards summarize the overall state of the evaluator’s dataset.
When Average Score or Pass Rate has no value, the card shows (an em dash) with its /100 or % suffix, rather than a misleading 0.

Critical alert banner

If any pending critical calls exist, a clickable red banner appears above the KPI grid:
“N pending critical calls — Pending review · Click for details”
Clicking it jumps directly to the Transcription Details table with the Pending Critical filter pre-applied.

Charts

Monthly Comparison

Two side-by-side period cards — This Month and Previous Month — each showing:
  • Calls — total count
  • Average Score — numeric
  • Pass Rate — percentage
A delta summary below the cards shows the score change (green upward arrow for improvement, red downward for decline) and the pass rate change between periods.

Duration Analysis

Three tiles grouping calls by duration using thresholds dynamically calculated from the dataset: Each tile shows the call count and average score for that tier.
Duration thresholds are dynamically calculated using the p25 and p75 percentiles of your dataset — they adapt to the actual distribution of your calls, not a fixed value.

Timeline Evolution

A dual-axis chart overlaying two series:
  • Left Y-axis — Average Score (line)
  • Right Y-axis — Call Volume (bars)
Hover over any point to see both metrics simultaneously.

Score Distribution

A bar chart grouping all transcriptions into 20-point score ranges: 0–20 · 20–40 · 40–60 · 60–80 · 80–100. Shows how scores are distributed across the full dataset.

Evolution by criterion

A multi-line chart — one line per criterion, plotted per period — showing whether each criterion is getting better or worse over time. Rendered only when the analysis contains per-period criteria data. The header shows the covered range (e.g. Jan 2026 - Jun 2026) and an “N periods” pill. In Native mode, the pills for Yes/No and Strict criteria are disabled (“Yes/No criteria are only shown in %”) — a 0–1 compliance rate and a 1–5 score cannot share a Y axis. If there are no scale criteria at all, the Native button itself is disabled and the chart opens in %. Hovering a data point shows 85% (4.25/5) in % mode, 4.25/5 in Native mode, and “No data” where a period has none. Empty states:
  • No criterion selected → “Select at least one criterion to see its evolution.”
  • No evolution data at all → “No criterion evolution yet”“There’s no per-period criteria data yet. Assign a period when uploading recordings to see their evolution.”
Per-criterion data (this chart and the section below) only accumulates for transcriptions evaluated after this feature was released — older calls are not backfilled. If all your calls predate it, the chart stays empty until new recordings are evaluated. Set the recording date at upload so each call lands in the right period.

Performance by criterion

A ranked list of every criterion with data in the analysed range — scrollable, with no top-5 cut-off. Two sort buttons in the section header (the choice also drives the PDF report):
  • To improve (default) — worst first, highest fail rate. Subtitle: “To improve first · average score and fail %”
  • Best — best first.
The sort is not persisted; it resets to To improve on reload. Each row shows: rank circle → criterion name → type badge (Scale / Yes/No / Strict) → average-score badge (scale criteria only) → progress bar. On the right, the fail rate large (one decimal), with “fails · of N evaluated” beneath.
The large number is the fail %, computed over the calls where the criterion was actually evaluated — the “of N evaluated” count — not over total calls. A criterion evaluated in 5 of your 25 calls with 2 fails shows 40%, not 8%.
Color tiers. The fail rate drives the rank circle, the progress bar, and the big number — a fuller bar is always a worse criterion: The average-score badge has its own scale (higher = greener), because a low fail rate and a good score are different things — a criterion can pass every call and still score mediocre: The average badge is deliberately hidden for Yes/No and Strict criteria: for those, the average is just the mirror of the fail rate already shown.
Criteria with no data in the range are hidden entirely — not shown as perfect, not shown at all. This typically means their only calls in range were evaluated before per-criterion tracking was released (no backfill). If your evaluator has 12 criteria and you only see 9, this is why.

Weekly Heatmap

A day-of-week × hour-of-day grid showing the average quality score for each time slot across all evaluated calls. Two auto-generated insights are shown above the grid: Hover over any cell to see the exact average score for that time slot.
The hour axis comes from each call’s recording time. Recordings uploaded with a date-only recording date land at midnight UTC and pile into that column — add the time at upload for an accurate heatmap.

Participants Ranking

A table of the top 8 participants evaluated under this evaluator, sorted by average score. Consistency thresholds:
A low standard deviation means predictable, stable performance. A high value means the participant’s scores vary significantly from call to call.
This section is only shown when participant data is present.

Transcription Details

A full paginated table of every transcription included in the analysis (10 records per page).

Group tabs

Filter the table by workflow group with a single click: Each tab shows a count badge.

Filters

Click Filters to open the filter panel:

Table columns

Click any row to open the full transcription detail in a new tab.
Group changes made from this table are applied immediately and reflected without requiring a full page re-fetch.

PDF Export

The Download PDF button in the header is available once analysis data is loaded. Report filename: Evaluator_Report_[name]_[date].pdf Report contents:
  • Evaluator name and evaluation date
  • Period range — only when a date range is applied, e.g. Period: 2026-01-01 → 2026-03-31 (a missing side prints an ellipsis; dates print raw as YYYY-MM-DD)
  • KPIs: Total Calls · Average Score · Pass Rate · Critical Fails
  • Critical fails warning (if applicable)
  • Monthly comparison — omitted when a date range is applied
  • Performance by criterion — color-coded bars
  • Evolution by criterion — table
  • Participants ranking table
  • Transcription detail table (ID · Date · Score · Duration · Participant · Critical · Group)
The period line reflects the range that was applied — what the report’s numbers actually cover — not whatever is typed in the From/To boxes. Typing dates and downloading without pressing Apply produces a report with no period line and full-history figures.
With a date range applied, Monthly Comparison disappears from the PDF: it compares the current calendar month against the previous one, which is meaningless — and misleading — inside an arbitrary range. Its absence is intentional.
Performance by criterion mirrors the on-screen section: every criterion with data (no cap — the report paginates), following the sort you selected on screen, with the same four fail-rate color tiers on the bars and % labels, and avg 3.4/5 shown for scale criteria only. Evolution by criterion is printed as a table, not a chart — one row per criterion, one column per month. It deliberately differs from the on-screen chart in three ways: it is always monthly (the Quarter toggle has no effect), period labels are raw keys (2026-01, not Jan 2026), and it includes every criterion with data regardless of which pills are selected. Cells show 3.4 for scale criteria, 78% for Yes/No and Strict, and where a month has no data.

Evaluators

Manage your evaluators

Transcriptions

Browse individual transcription results