Skip to main content
Evaluator Shift Planning: Throughput Targets, QA Sampling and Scheduling Recipes

Evaluator Shift Planning: Throughput Targets, QA Sampling and Scheduling Recipes

Staffing math, rota templates and calibration routines for distributed idea-evaluator pools

Most idea programs don't fall apart because ideas dry up. They fall apart because the evaluation layer gets swamped. Submissions pile up, a handful of evaluators quietly do 70% of the scoring, quality drifts, and by the time anyone notices there's a three-week backlog and a bunch of frustrated submitters who think their ideas vanished into a void.

If you run evaluation as a distributed pool — people scattered across business units, time zones, and part-time commitments — the scheduling problem gets genuinely hard. This piece is about the operational plumbing: how many evaluators you actually need, how to build rotas that survive real life, what throughput targets are realistic, and how to sample for quality without turning it into a bureaucracy.

Start With The Throughput Math, Not The Headcount

Almost everyone staffs by gut feel. Someone says "we have 400 ideas a quarter, let's get 10 volunteers," and nobody checks whether those numbers actually connect.

  1. Inflow — ideas arriving per week (use your busiest recent 8-week average, not the annual mean)
  2. Handling time — realistic minutes per idea for a first-pass evaluation
  3. Available evaluator capacity — actual hours people give you, not the hours they promised

The trap is handling time. Teams estimate 5 minutes per idea. In practice, a proper first-pass — reading the submission, checking for duplicates, scoring against a rubric, leaving a one-line note — runs closer to 8 to 14 minutes depending on submission quality. Messy intake pushes it toward the high end. A disciplined intake form lands you nearer the low end.

So the base formula: Weekly evaluator hours needed = (weekly inflow × avg handling minutes × review passes) ÷ 60 The "review passes" multiplier matters. A single reviewer per idea means 1. Two independent scores on anything that clears a first threshold puts your effective multiplier around 1.4–1.6 — not 2, because plenty of ideas get killed on the first pass and never reach a second reviewer.

Staffing Formulas by Program Size

Below is a rough sizing guide based on patterns that show up across programs at different scales. Treat the evaluator counts as active contributors per week, not total pool size — you always recruit more people than you schedule, because attrition and no-shows are constant.

Program sizeWeekly inflowHandling timeReview passesWeekly hours neededActive evaluators/week
Small15–30 ideas~10 min1.0~3–5 hrs2–3
Medium60–120 ideas~11 min1.3~14–29 hrs6–10
Large250–450 ideas~12 min1.5~75–135 hrs25–40

The jump from medium to large isn't linear because coordination overhead grows. At small scale, one person can eyeball the whole queue. At large scale you're managing shift handoffs, calibration drift, and duplicate detection across dozens of people who never talk to each other. Budget an extra 15–20% capacity at large scale purely for coordination work — stuff that produces zero scored ideas but keeps the whole thing from drifting apart.

On pool size versus active count: if you need 8 active evaluators in a given week, your standing pool should be roughly 12–16 people. People take vacation, get pulled into their day jobs, burn out. A pool sized exactly to demand collapses the first time two people go quiet in the same week.

Rota Templates That Survive Real Life

"Everyone evaluates whenever they have time" sounds flexible. In practice it produces feast-or-famine: everyone waits for someone else, the queue balloons, and then three conscientious people binge-clear it on a Friday afternoon.

You want committed windows with named ownership. Three templates hold up in distributed settings:

The Zoned Relay (good for global pools). Split the week into three ownership blocks aligned to time zones — APAC morning, EMEA midday, Americas afternoon. Each zone owns the queue during its block and clears whatever arrived overnight. Nobody feels responsible for the whole thing, and the queue never sits untouched for more than 8 hours. This borrows from the coordination logic in a good moderator and curator playbook for high-volume idea streams — the same handoff discipline that keeps moderation queues moving keeps evaluation queues moving.

The Two-Slot Weekly (good for medium programs). Each evaluator commits to two fixed 45-minute slots per week — say Tuesday and Thursday. Stagger slots so coverage spreads rather than clusters. Works well for part-timers who can't handle a daily obligation but can protect two calendar blocks.

The Anchor + Surge (good for lumpy inflow). A small anchor crew of 2–4 people handles baseline flow every week. A larger surge list activates only when the queue crosses a threshold — more than 40 unscored ideas older than 48 hours, for example. The anchor keeps things stable; the surge prevents backlog blowups after a big campaign or hackathon drops 200 submissions in two days.

Process diagram

Here's a simple visual of the rota handoffs and surge activation flow.

Schedule occasional overlapping windows so two evaluators can resolve ambiguous submissions in real time instead of leaving them stuck for days.

One subtle rota mistake worth calling out: scheduling people with no overlap at all. You want occasional overlapping windows so two evaluators can resolve a genuinely ambiguous submission in real time instead of leaving it stuck for days waiting on someone else.

Throughput KPIs Worth Tracking

Four numbers tell you almost everything about whether your evaluation layer is healthy.

  1. Queue age (p50 and p90). Median tells you the normal experience; the 90th percentile shows how bad the worst cases get. A p50 of 2 days with a p90 of 19 days means a subset of ideas is quietly rotting — which is usually where "my idea disappeared" complaints come from.
  2. Throughput per active evaluator. Ideas scored per evaluator per week. Watch the spread, not just the average. If your top person does 45 and your median does 6, you don't have a pool — you have one hero and some spectators.
  3. Time-to-first-touch. How long before any human acknowledges a submission. Submitters tolerate slow decisions far better than silence. Getting first-touch under 48 hours does more for program trust than shaving days off the final decision timeline.
  4. Rework rate. Share of evaluations that get overturned, escalated, or re-scored. If it's creeping upward, calibration is slipping.

The pattern to watch: throughput looks fine on average while p90 queue age quietly climbs. Averages hide the ideas that fall through cracks, and those cracks are exactly what erodes participation over time.

QA Sampling Without The Bureaucracy

You can't re-check every evaluation. You also can't check zero, or quality quietly rots. The answer is proportional sampling that scales with both volume and stakes.

  1. Sample by decision weight, not uniformly. Pull a heavier sample from ideas that got advanced (a false positive wastes real pilot budget) and from ideas killed with high submitter effort (a false negative burns goodwill). Ideas that were obvious junk or obvious winners barely need auditing.
  2. Set a base rate around 8–12% of scored ideas per week for medium programs. Small programs can afford to audit more; large programs sample less by percentage but more in absolute terms.
  3. Rotate the auditor. The same person auditing every week develops their own drift. Rotate among your most calibrated evaluators.
  4. Flag disagreement, don't punish it. When an audit disagrees with the original score, feed that into calibration — not into a scoreboard of who's "wrong."

A concrete recipe for a medium program scoring around 90 ideas a week: audit all advanced ideas (say 12), plus 10% of the killed pile (~7), plus a random 5 from the middle. That's roughly 24 audits weekly — about 90 minutes of one person's time. Cheap insurance against a pool that's silently drifting off-rubric.

Weekly Calibration Checklist

Calibration decays fastest and gets skipped first. Distributed evaluators drift apart precisely because they never see each other's reasoning. A short weekly ritual holds the line. If your rubric is shaky, fix that first — a scoring frame built to resist bias, like the one covered in this breakdown of bias-resistant evaluation rubrics, gives calibration something solid to work against. Run this every week, keep it under 30 minutes:

  1. [ ] Pull 3–4 ideas where two evaluators disagreed by two or more points
  2. [ ] Have the two scorers each explain their reasoning in two sentences
  3. [ ] Note whether the gap was a rubric interpretation issue or a genuine judgment call
  4. [ ] Re-score one "anchor" idea everyone has seen before — check for group drift over time
  5. [ ] Update one line of rubric guidance if the same ambiguity keeps recurring
  6. [ ] Share the week's rework rate and any newly retired duplicates
  7. [ ] Confirm next week's rota coverage and flag any known absences

The most useful outcome isn't agreement — it's a shared vocabulary for why people score differently. Pools that calibrate weekly tend to converge over a couple of months. Pools that skip it spread further apart, and nobody notices until the rework rate spikes.

A Real Scenario

A mid-market manufacturer ran an internal idea program across five plants. Inflow averaged around 80 ideas a week, spiking past 200 after quarterly safety campaigns. They had 14 volunteer evaluators and no real schedule — people scored "when they had a minute."

The symptoms were predictable. Median queue age sat around 9 days, but p90 was pushing three weeks. Two engineers were doing well over half the scoring. Submitter complaints about ideas going nowhere were climbing, and participation had started to sag.

They didn't add people. They restructured. An anchor crew of four took two fixed slots a week each, a surge list of ten activated whenever the 48-hour backlog crossed 40 ideas, and one person ran a 20-minute Friday calibration. QA sampling landed around 10%. Within about six weeks, time-to-first-touch dropped under two days, p90 queue age settled near 6 days, and the scoring spread tightened — the two overloaded engineers went from doing roughly 55% of volume down to about a third between them. Nothing dramatic happened to the number of ideas. The evaluation layer just stopped being the bottleneck.

When This Level Of Structure Makes Sense — And When It Doesn't

If you're running fewer than 15 ideas a week with two or three reliable people, skip most of this. Formal rotas and sampling plans are overhead you don't need at that scale — a shared spreadsheet and a weekly glance will do. Over-engineering a small program mostly just adds friction that scares off volunteers.

The structure earns its keep once you cross roughly 50 ideas a week or once your pool grows past 6–8 people who don't naturally coordinate. That's where informal goodwill stops scaling and queue age quietly rots your program's reputation.

Worth noting: teams whose real problem is intake quality or unclear rubrics shouldn't bother with this yet. If half your submissions are unscorable or your evaluators can't agree on what "good" looks like, no scheduling recipe will save you. Fix the upstream mess first, then come back and tune the shift plan. Throughput math only works when the thing flowing through it is actually worth evaluating.

Built for Innovators Designed specifically for dynamic idea workflows & collaboration
Save Time Streamline idea submission, review, and execution
Engage Teams Boost participation with transparent feedback and voting
Drive Results Turn ideas into measurable business impact faster