A staffed seat costs about $255,000 a year to run around the clock. How many cameras it covers is the operator to camera ratio — and it decides whether proactive monitoring is a business or a subsidy.

Researchers put 73 people — 42 of them full-time CCTV operators — in front of 90 minutes of processing-plant footage and asked them to catch four kinds of target behaviour. They caught about half. Only 12% of operators detected more than three-quarters of the targets.

That is the ceiling every proactive monitoring business is built on top of, and it sets up the only number in the model that matters: the operator to camera ratio. How many cameras one person can meaningfully cover determines your cost per camera, and cost per camera determines whether you have a product or an expensive favour.

Start with the seat, because the seat is what you're actually buying.

Why the Ratio Is the Only Number That Matters

A monitoring centre's cost is not cameras, storage or bandwidth. It's people watching screens, and the arithmetic of covering a screen continuously is unforgiving.

A week has 168 hours. A full-time employee works 40. So a single seat staffed around the clock needs 4.2 full-time employees before you account for leave, training, sick cover or breaks — that floor is arithmetic, not an estimate, and real rosters run higher.

Price that floor. US Bureau of Labor Statistics data for May 2025 puts the mean annual wage for security guards and surveillance officers at $42,490. Wages are only part of employment cost: the BLS Employer Costs for Employee Compensation release puts benefits at 30.1% of total compensation in private industry, so the fully loaded figure is roughly $60,800 per employee per year. Multiply by 4.2 and one continuously staffed seat costs about $255,000 a year — around $21,300 a month.

That monthly number is the numerator. The ratio is the divisor. Everything else in a proactive monitoring P&L is rounding.

Cameras per operator Cost per camera / month Reality check
10 $2,130 More than the camera, the install and the bandwidth combined
25 $851 Still above most full VSaaS subscriptions
50 $426 Viable only at a premium remote-guarding price point
100 $213 Where the model starts to breathe
250 $85 Competitive with mid-tier managed video
500 $43 Only reachable when the human never searches

Swap in your own wage and roster assumptions — the inputs are all public and the arithmetic is three lines. The shape won't change. Below roughly fifty cameras per operator, you are charging more per camera than the hardware costs, and the only customers who accept that are the ones with something genuinely expensive to lose. Above a couple of hundred, the model looks like software.

So the whole business hangs on a number the research says is stubbornly small.

What a Human Can Actually Cover

Two studies define the ceiling, and they appear to disagree until you read what each one measured.

The first is the 2015 Applied Ergonomics study above — Donald, Donald and Thatcher. Detection ran around 50%, false alarms were high (a mean of 15.76 in the first stage), and performance split sharply by background. Novices and generalists declined significantly after the first thirty minutes. Specialists — operators whose regular work matched the task — held for a full hour and then got better. Experience mattered a great deal.

The second says the opposite. In a PLOS ONE study, 66% of participants failed to notice an unexpected event on a single screen in good conditions, and 84 military personnel with up to nineteen years of real CCTV duty performed no better than undergraduates.

Both are right, and the reconciliation is the most useful thing in this article: experience helps you find what you are looking for. It does nothing for noticing what you weren't. Donald tested trained detection of defined target behaviours — a search task, and specialists win search tasks. PLOS tested an unexpected event nobody was briefed on — and there, expertise bought nothing at all.

Which means the ratio you can safely run depends entirely on which job you are asking the operator to do. If the system tells them what to look at, experience compounds and the ratio can climb. If they are scanning a wall hoping something catches their eye, you are paying for the mode where twenty years of experience measures the same as none.

At which point most operations reach for the wrong lever.

The Video Wall Is Not the Lever

There is real research on how to arrange a multiplex display. A 2021 study tested 27 scenes in a 9×3 grid and found that grouping them by semantic category and separating them with borders beat a random, borderless layout. The improvement was 3,416 milliseconds versus 3,550.

A hundred and thirty-four milliseconds. Against a 50% detection rate and a thirty-minute attention cliff, wall layout is a rounding error. It's worth doing — it's nearly free — but no arrangement of tiles turns fifty cameras per operator into two hundred. Neither does a bigger monitor, a curved wall, or better chairs.

The lever is changing the job.

What Actually Moves the Operator-to-Camera Ratio

The ratio moves when the operator stops searching and starts adjudicating. Instead of scanning a hundred tiles for something anomalous, they are handed a queue of candidate events and asked a closed question: is this real, and what do I do about it? That's the task type experience is good at, and it removes the failure mode expertise can't fix.

Arithmetically it's straightforward. If classification surfaces four events an operator-hour and each takes two minutes to adjudicate, one operator absorbs the alert volume of a very large camera estate — and camera count stops being the constraint. The constraint becomes alerts per hour, which is a function of your site, your scene, and your detector's precision.

Here is the part vendors skip. There is no independent public benchmark for video-analytics false-positive rates in operational conditions. NIST evaluates face and object detection under controlled protocols, and none of that translates into how many times a specific camera pointed at a specific fence line in specific weather will fire on a plastic bag. Every precision figure in a sales deck is self-reported, measured on data the vendor chose.

So the ratio is not a specification you can buy. It's a number you measure on your own site, and any proposal that quotes you a ratio before it has seen your cameras is quoting you a hope.

Which makes measuring it the actual first task.

How to Measure Your Own Ratio

Run a week of real footage through the detector in shadow mode — alerting to a log, not to an operator — and collect three numbers:

  • Alerts per camera per hour. Broken out by camera, because in almost every estate a handful of cameras generate most of the noise, and fixing those individually moves the ratio more than any platform change.
  • Adjudication time. How long an operator actually takes to open an alert, watch enough to decide, and either dismiss it or act. Time it; don't estimate it. This is where a slow live path silently taxes the ratio.
  • Precision on your site. Of the alerts raised, how many were genuinely worth a human's attention. Not the vendor's number — yours, on your fence line, in your weather.

Alerts per hour divided into an operator's available adjudication time gives you the real ratio. Then hold back a fraction of capacity, because alerts are not evenly distributed and the ratio that works at 3 a.m. will fail at shift change.

And one input to that calculation is infrastructure rather than staffing.

What the Ratio Costs You in Infrastructure

Raising the ratio moves the bottleneck rather than removing it. More cameras per seat means more simultaneous streams reaching one workstation, and the ceiling there isn't bandwidth — watching 500 cameras is a decode problem. A single machine that can carry fifty low-resolution tiles will not carry three hundred.

Adjudication time is the other infrastructure cost hiding in a staffing number. If opening an alert costs the operator three seconds of waiting for a clean frame, and they adjudicate two hundred alerts a shift, that's ten minutes of paid time spent watching a spinner — and it degrades the decision itself, since the delay lands exactly when they're deciding. The live path is a line item in the ratio, not a technical detail underneath it.

Which puts the sourcing question in economic terms rather than architectural ones.

Build, Buy, or Deploy

  • Build on open source. go2rtc or MediaMTX for ingest and low-latency delivery, with your own classification layer on top. You own the cost curve completely, which matters when cost per camera is the business. Right when monitoring is the product and you have the team.
  • Use a managed cloud relay. Fastest to operators watching feeds — but per-stream pricing scales with exactly the number you are trying to grow. A relay that's cheap at 50 cameras per operator is a structural problem at 500.

Whichever route, the cameras are rarely in one building. Scaling the ratio across a customer base means connecting many sites to one operations centre, which is a topology problem the per-seat arithmetic above quietly assumes you have already solved.

  • Deploy a commercial platform. A full media stack under a flat license, run on-premise or as a managed cloud deployment — which changes the shape of the curve, since flat licensing means raising the ratio improves margin instead of raising the bill. Samvyo is one such option: based on SFU architecture, sub-second live delivery for adjudication, with the media path, TURN and recording under your control. Where it doesn't fit: a small estate, or an operation that only reviews footage after the fact.

Three routes, one question — does your cost per camera fall as the ratio rises, or follow it up?

The Bottom Line

Proactive monitoring is an arithmetic problem wearing a technology costume. A round-the-clock seat costs roughly $255,000 a year, so the operator-to-camera ratio sets your cost per camera and therefore your entire business model. The research says a human scanning a wall detects around half of what matters and slides after thirty minutes, and no video wall layout fixes that — 134 milliseconds is what the best layout buys. What moves the ratio is changing the operator's job from searching to adjudicating, because experience helps enormously with the second and not at all with the first. Just don't buy the ratio from a slide. Measure it on your own cameras, because nobody else's fence line is yours.

What's Next?

For the response model this economics sits inside — and how to tell genuine proactive from a faster alert — Proactive vs Reactive Surveillance draws the line.

For what happens after an operator decides to act, and how much delay the response loop can absorb, see Talk-Down at Ten Seconds Is Theater.

And the live path that all of it runs on starts with getting RTSP cameras into a browser.

Frequently Asked Questions

What is a good operator to camera ratio for a monitoring centre?

There isn't a universal figure, because it depends on alerts per hour rather than camera count. What the economics say is that below roughly fifty cameras per operator you are charging more per camera than the hardware costs, and the model only works with AI classification filtering the feed so the operator adjudicates events instead of scanning tiles.

How many cameras can one person actually watch?

Far fewer than most staffing plans assume. In a study of 73 people including 42 full-time CCTV operators, detection of target behaviours ran around 50%, only 12% caught more than three-quarters, and novices and generalists declined significantly after the first thirty minutes. Scanning is the failure mode; a filtered alert queue is not.

Does operator experience improve detection?

For expected targets, substantially — specialists in the Applied Ergonomics study held performance for an hour and then improved. For unexpected events, not at all: a PLOS ONE study found personnel with up to nineteen years of CCTV duty performed no better than untrained undergraduates. Experience helps you find what you're looking for, not notice what you weren't.

How much does a 24/7 monitoring seat cost?

Around $255,000 a year in the US on public figures. A week is 168 hours and a full-time role is 40, so continuous coverage needs at least 4.2 employees; BLS puts the mean security-officer wage at $42,490 and benefits at 30.1% of total compensation, giving roughly $60,800 fully loaded per person. Real rosters run higher once leave and training are covered.

Does a better video wall layout improve detection?

Marginally. Research testing 27 scenes found that semantic grouping plus borders beat a random layout by about 134 milliseconds. Worth doing because it's nearly free, but it doesn't change the ratio — the detection ceiling is an attention problem, not a display problem.

What false-positive rate should I expect from video analytics?

Whatever your own site produces, which nobody can tell you in advance. There is no independent public benchmark for analytics false-positive rates under operational conditions, and vendor precision figures are self-reported on data they selected. Run the detector in shadow mode against a week of your own footage before committing to a ratio.

What infrastructure does a high operator to camera ratio require?

More concurrent streams per workstation, which makes decode rather than bandwidth the ceiling, and a live path fast enough that adjudication isn't spent waiting for frames. Samvyo is based on SFU architecture and delivers sub-second live view with flat licensing, so raising the ratio improves margin rather than increasing a per-stream bill.