Running a telehealth program at scale, Part 3: clinician capacity planning and a realistic turnaround SLA
At 1,000 encounters a month, clinician capacity stops being a hiring question and becomes a systems question: how many clinician-hours, licensed where, routed how, with what promise to the patient. Part 3 of our series for programs at scale.
Capacity is arithmetic: encounters per month by type, times minutes per type, divided by working minutes, gives clinician-hours — then multiplied by a peak factor, because demand arrives in Monday spikes and after every marketing push. Coverage is a second constraint: each encounter needs a clinician licensed in that patient’s state, so the plan is a matrix of hours by state, not a headcount. Use one queue routed by license and urgency, with separate lanes for red flags and for states that require a live visit. Define turnaround as a percentile from submission to decision, promise the 90th not the median, and expect hours, not minutes. Structured intake and drafted renewal summaries cut minutes per case more than any hiring plan. Measure turnaround p50 and p90, needs-more-information rate, utilization, and coverage gaps by state.
Capacity is arithmetic
Start with the encounters, not the clinicians. Split next month’s expected volume by type — initial, renewal, labs-only, live visit — because each has a different review time. Multiply each by the minutes a clinician actually spends on it, including charting and messaging. Sum, and you have clinician-minutes of demand. Divide by the productive minutes a clinician gives you in a month (fewer than you think once breaks, messaging, and administrative time are counted) and you have base headcount.
Then multiply by a peak factor. Demand is not flat: intake arrives in Monday spikes, after every email send, and around paydays. A program that staffs to the average meets its turnaround promise half the month. A peak factor between 1.3 and 1.5 is where most asynchronous programs land, and the right number is the one that keeps your 90th-percentile turnaround inside the promise.
| Term | What it is | Where to get it |
|---|---|---|
| Encounters by type | Initial · renewal · labs · live | Last three months, trended |
| Minutes per type | Review + chart + message | Measure from your platform, not from estimates |
| Productive minutes per clinician-month | Scheduled hours minus overhead | Ask; then measure |
| Peak factor | Peak-day demand ÷ average-day demand | Your own intake timestamps |
| Coverage matrix | Hours needed by state by window | From patient-state mix |
Coverage is the second constraint
The arithmetic above assumes any clinician can take any case. They cannot. Every encounter must be reviewed by a clinician licensed in the patient’s state, and for nurse practitioners, supervised as that state requires. So capacity is not a number; it is a matrix — hours needed per state per coverage window — and the plan has to fill every cell.
This is where in-house staffing breaks at scale. A team of ten clinicians licensed in a combined fifteen states has zero capacity in the other thirty-five, and licensing existing staff into new states takes months per state. A program that markets nationally needs coverage nationally, at every hour it promises turnaround. Most programs at volume solve this with a clinician network licensed in all fifty states, provisioned as capacity rather than headcount; our 50-state compliance checklist covers the licensure and collaborating-physician rules underneath.
Queue design
One queue, routed. Multiple queues by category or by clinician create idle capacity in one lane while another backs up. A single queue with routing rules — license match first, then urgency, then category preference, then clinician availability — keeps every licensed minute working.
Three lanes sit alongside the main queue:
- Red flags. Protocol-defined intake answers escalate immediately and jump the queue, with a clinician paged if none is active. These are configured, not left to judgment.
- Live visits. States and treatments that require a synchronous encounter route to a scheduling lane rather than the async queue, so they do not sit unreviewed.
- Needs more information. Encounters a clinician sends back wait for the patient, not for a clinician; they re-enter the queue when the patient responds, with priority preserved.
The design principle: a case should never wait for a clinician who cannot legally act on it, and a clinician should never wait for work while a case sits in the wrong lane.
What turnaround SLA is realistic
Define turnaround precisely: time from intake submission to the clinician’s decision, measured as a distribution. Report the median and the 90th percentile, and promise the 90th. A median of two hours with a 90th percentile of eighteen means one patient in ten waits most of a day, and that patient is the one who writes the review.
For asynchronous programs with coverage windows, hours is the honest unit. A promise of minutes requires clinicians idle and waiting, which is capacity you pay for and rarely use. A promise of hours with a reliable 90th percentile is what patients actually value: they submitted at 9pm and had an answer by morning. Renewals can and should be faster than initials, because the record is known and the check-in is short.
Minutes per case is the lever you control
Every term in the capacity equation is a constraint except one: minutes per case. Hiring adds hours slowly and expensively. Cutting minutes per case adds capacity immediately.
- Structured intake. A complete, schema-validated case reads in a fraction of the time of free text. See designing telehealth intake.
- Protocol pre-screening. Ineligible cases never reach the queue as prescriptions; they route to a live visit or a decline with reasons attached.
- Drafted renewal summaries. A check-in summarized as a diff against the last visit, with the protocol’s proposed next step, turns a renewal into a confirmation. This is where agentic workflows earn capacity; see refilling medication with AI.
- Templated charting. The note structure follows the protocol, so clinicians complete rather than compose.
Programs that do these four often find their existing clinician hours cover twice the volume before anyone is hired.
Metrics that show capacity is working
- Turnaround p50 and p90, by encounter type and by state. Rising p90 with flat p50 means a coverage gap somewhere.
- Needs-more-information rate. Above a few percent, intake is the problem, not capacity.
- Utilization. Clinician-minutes worked over minutes scheduled. Very high utilization with rising p90 means you are at the ceiling; low utilization with rising p90 means routing is wrong.
- Coverage gaps by state and hour. Cells in the matrix with demand and no licensed clinician. Each is a patient waiting.
- Escalation rate. Red flags caught by intake and routed correctly. Higher is better.
Lithos provides capacity as capacity: a licensed, board-certified physician network with multi-state licensure across all 50 states, AI-prepared consultations — intake, assessment, and clinical summary complete before a physician opens the encounter — which keeps minutes per case low, and state-by-state oversight requirements handled for you. Programs bring their volume and states; the matrix is our problem. Next in the series: the refill lapse problem.
Frequently asked questions
How many clinicians does a telehealth program need?
Work it from encounters: monthly encounters by type, times minutes per type, gives clinician-minutes; divide by productive minutes per clinician-month and multiply by a peak factor of roughly 1.3 to 1.5. Then check the answer against state coverage, because ten clinicians licensed in the wrong states is zero capacity in the right ones.
What is a realistic turnaround time for asynchronous telehealth?
Hours, measured at the 90th percentile from intake submission to clinician decision. Well-run asynchronous programs complete most reviews within a few hours during coverage windows; promising minutes invites failure and promising the median hides the patients who wait longest.
How does state licensure affect clinician capacity?
Every encounter must be reviewed by a clinician licensed in the patient’s state, so capacity is a matrix of hours by state. A program selling in thirty states needs coverage in thirty states at every hour it promises turnaround, which is why most programs at scale use a network rather than a staff.
Do nurse practitioners need a collaborating physician for telehealth?
In many states, yes, with limits on how many NPs a physician may supervise and requirements for chart review. Capacity plans that lean on NPs must plan physician hours alongside them, state by state.
How does structured intake change clinician capacity?
Directly: minutes per case is the largest term in the capacity equation. A clinician reading a complete, schema-validated intake with a drafted summary decides in a fraction of the time it takes to reconstruct a history from free text, which is capacity you did not have to hire.
Get the Journal by email
Our best guides on building compliant telehealth programs — a couple a week, unsubscribe any time.
From first call to first patient, in weeks.
A 15-minute intro call, sandbox credentials the same day, go-live in 3–4 weeks — new launches and existing patient bases alike.