How to design a surveillance video wall: layouts and tile counts, the decode budget nobody quotes you, substreams, alert-driven switching, and what a video wall actually costs beyond the panels.
The panels are the commodity. The decode machine and the software are the real constraint, and neither appears on a display quotation.
Every tile on the wall is a continuous video decode. Specify the driving machine for your maximum tile count with every tile active, because that is the normal state rather than the peak.
Match the stream profile to the tile size. Using camera substreams on small tiles is the single largest saving available in a wall's decode budget.
More tiles is usually worse. Around nine scenes is the practical limit of one person's attention, and larger grids serve rooms rather than individuals.
A static wall is a screensaver. Alert driven switching is what converts a display into an instrument, and cycling must cooperate with alerting rather than fight it.
If you search for video wall you will find LED panels. Samsung, LG, Planar, Barco, Sony, and a long tail of resellers, all answering the question of what to buy with a screen. The most common question people ask alongside it is what it costs, and the answers cover panels, mounts and installation.
That framing is why surveillance video wall projects disappoint. The panels are a solved, competitive, commodity purchase. The part that decides whether operators actually use the wall is the software driving it and the machine decoding the streams, and almost nobody quotes you on either.
This guide covers the part that is left out.
A video wall is a set of displays driven as one logical surface, showing many camera streams at once so that a room full of people can see the estate without each of them driving a workstation.
It is worth separating three things that get conflated. The display layer is the physical panels, their bezels, mounts and cabling. LCD panels with thin bezels tile cheaply and are the default for control rooms, while direct view LED has no bezels at all, looks better, and costs materially more. This layer is a commodity purchase and you should treat it as one.
The signal layer is whatever gets pixels onto those panels. A hardware video wall controller or processor takes inputs and distributes them across a display array. This is the traditional approach, and for mixed source walls that show a spreadsheet, a news feed and cameras together, it is still the right one.
The video layer is the software that decides which camera appears in which tile, at what resolution, and what happens when something occurs. For a surveillance wall this is the layer that matters, and it is the one the display quotation will not mention. If every source on your wall is a camera from one VMS, you frequently do not need a hardware controller at all: a machine running the VMS fullscreen across its outputs is simpler, cheaper and easier to change.
This is the most searched question about video walls, and the published answers are almost all panel prices. A more honest breakdown has five lines, and the last two are the ones that get missed.
Displays are the commodity line, scaling with panel count, bezel width and whether you choose LCD or direct view LED. Mounts and installation cover wall structure, alignment, cable management and service access, which is underestimated but well understood by any AV integrator. The signal path is either a hardware controller or the graphics outputs of the driving machine, and a pure camera wall usually needs the latter, which is far cheaper.
The decode machine is the line that sinks projects. Every tile on the wall is a video stream being decoded in real time, so thirty six tiles is thirty six simultaneous decodes, continuously, forever. That is a hardware specification, not an afterthought.
Bandwidth and network is the fifth line. Thirty six streams arriving at one machine is a sustained load on the segment feeding it, and if the wall shares a switch with the recording path you should size it deliberately.
There is no single number, and any figure quoted without knowing your camera count, resolution and tile count is a guess. But the useful insight is that the ratio is not what people expect: you can buy excellent panels and still have an unusable wall, because the money went to the layer that was easy to quote.
Here is the calculation that should happen before anything is purchased. Each tile decodes a video stream, and the cost of that decode scales with resolution, frame rate and codec. Thirty six 4MP streams at twenty five frames per second is a serious continuous workload for any machine, and it is not a workload that pauses.
Three things make it tractable. The first is using substreams for small tiles. Almost every IP camera publishes at least two streams: a full resolution main stream and a lower resolution substream. A tile occupying a ninth of a display does not need 4MP and cannot show it, so decoding the main stream into a small tile burns capacity producing detail the scaler throws away. Matching the stream profile to the tile size is the single largest saving available, often by a wide margin.
The second is hardware decode. Modern GPUs and integrated graphics decode H.264 and H.265 in dedicated silicon, and software decode on the CPU will run out long before the hardware decoder would. Confirm the machine's decoder supports the codec your cameras publish, because H.265 support is not universal on older hardware and a wall that falls back to software decode will look fine for ten tiles and collapse at thirty.
The third is reducing frame rate on the wall. A wall is for situational awareness, not forensic review, so wall tiles do not need the frame rate the recording path captures. Recording at full rate while displaying at a lower one is normal practice and costs nothing operationally.
A useful rule: decide your maximum tile count, assume every tile is active simultaneously because it will be, and specify the decode machine for that. Then check what happens when someone expands a tile to fullscreen, because that tile switches to the main stream.
Available layouts are typically square grids and picture in picture variants, where one or more large tiles carry the attention and smaller tiles surround them.
The instinct is to maximise tile count. Resist it. Beyond roughly sixteen tiles on a single operator's field of view, additional tiles do not add awareness, they dilute it. There is a well documented ceiling on how many moving scenes a person can meaningfully monitor, and it is far lower than the number of tiles a modern grid will render.
Match the layout to the job. Use 1x1 and 2x2 for active investigation, where detail matters and someone is looking closely. Use 3x3 as the practical default for a monitored wall, because nine scenes is near the limit of what one person tracks. Use 4x4 and larger for overview walls seen by a room rather than scanned by an individual, where the purpose is presence and gross change detection rather than reading detail. Use picture in picture where there is a genuine hierarchy: a primary camera under active attention with contextual feeds around it.
The honest position is that large grids are for coverage optics and for rooms with several people, not for one operator's attention. If a wall exists so that a single person watches thirty six cameras, it is not doing what it appears to do.
Everything above is table stakes. The difference between a wall that gets used and a wall that becomes expensive decoration is whether the wall responds to events.
A static grid is a screensaver. Operators habituate to it within days, and the camera that matters is almost never one of the ones currently displayed. A wall earns its cost when it changes in response to what the system detects.
Two mechanisms do most of the work. Cycling handles estates far larger than the tile count, rotating through camera sets on a timer so that a hundred cameras get seen across a 3x3 grid, which is a coverage tool. Alert driven switching brings the relevant camera forward when the system detects something, and this is the mechanism that converts a wall from a display into an instrument.
The two interact badly if you are not careful, and that interaction is where most implementations are weak. If the wall is cycling and an alert fires on a camera that is not currently shown, a naive implementation either ignores it or jumps to it and then keeps cycling away mid incident, which is worse than not switching at all.
Being specific about what is implemented, because this section is a product description.
The Visylix video wall is a dedicated browser view, separate from the normal live view page, that renders fullscreen with no sidebar or application chrome so that it can be projected on dedicated displays using Chromium in kiosk mode. It is worth being precise about the architecture: this is a web page, not a native desktop application. There is no native multi monitor client. A multi display wall runs one browser instance per display, each configured with its own layout and camera set. For a wall driven by a single machine with several outputs this is straightforward, and it is a genuine architectural constraint rather than a feature.
Layouts are square grids from 1x1 through 6x6, giving up to thirty six tiles per display, plus two picture in picture variants for five and eight streams where one tile leads. The wall is keyboard driven: arrow keys move focus between tiles, Enter expands the focused tile, Escape returns to the grid, and number keys one through six switch layout directly. Control room operators work faster on a keyboard than a mouse, and a wall being driven from across a room often has no usable pointer at all.
Auto cycle rotates through stream sets on a configurable interval, so that fleets much larger than the tile count are covered. Alert pinned cycling is the part worth explaining, because it is where the naive implementation fails. When an unacknowledged critical alert exists on a camera in the cycling pool, the cycle pins to the page containing that camera and stops advancing, so the operator keeps looking at the thing that matters rather than watching it rotate away.
The detail we are more pleased with is what happens on acknowledgement. When the operator acknowledges the alert, the cycle does not resume from wherever a background timer had silently advanced to. It resumes from the page the operator was just looking at, and advances from there. The operator's context is preserved instead of being discarded, so acknowledging an alert does not teleport the wall somewhere unrelated.
An alert rail surfaces alerts alongside the grid rather than only inside a tile, so an operator scanning the wall sees severity without having to notice which tile changed. Audio is muted per tile independently, because at sixteen tiles and above an unmuted wall is unusable and a single global mute is too blunt when one camera's audio matters.
Small tiles request the camera's substream rather than the main stream, which is the decode saving described earlier. This is gated on whether a substream URL is actually configured for that camera, and that gate matters more than it sounds. Without it, a small tile requests a substream profile the engine does not publish, gets a 404, and the player falls through its transport fallback chain before finally rendering an offline tile, for a camera whose main stream was online the whole time. If your wall shows black tiles for healthy cameras, an unconditional substream request is the first thing to check, whatever platform you are running.
Decide the tile count first, then buy the decode machine, not the other way round. The panel quotation will arrive first and it is not the constraint.
Configure substreams on every camera. This is usually a five minute change per camera in the camera's own web interface, and it is the difference between a wall that runs comfortably and one that runs hot. Do it before commissioning, not after the wall stutters.
Put the wall on its own network segment, or at least confirm the segment can carry the sustained aggregate, because a wall pulling thirty six streams alongside the recording path is a real load.
Decide what the wall is for, in one sentence. Situational awareness for the duty officer, visible reassurance in reception, and incident response for a monitored estate produce different layouts, different cycling behaviour and different tile counts. Walls specified without this end up as 6x6 grids that nobody reads.
Plan for the display being on continuously, since static overlays and unchanging tiles can cause image retention on some panel technologies, and cycling helps here as a side effect. And test with the room lit as it will actually be lit, because control rooms are often dim by design and a wall commissioned in a bright room at midday is not the wall the night shift sees.
Worth saying plainly, since we would be paid either way.
If nobody is in the room, the wall is decoration. Recorded video with good search and alerting serves an unstaffed site far better than a display no one is watching.
If the wall exists to demonstrate that surveillance exists, that is a legitimate goal, but specify it as such and buy far fewer tiles than you think, because nobody is reading them.
If your operators work from desks, an ordinary multi monitor workstation running live view is usually more effective than a shared wall. Walls serve rooms, workstations serve people. And if your camera estate is small enough that every camera fits on one screen at usable size, you do not need a wall, you need a good live view page.
A video wall earns its place in staffed control rooms where several people need shared situational awareness; on estates far larger than one screen, where cycling gives coverage; in environments where alert driven switching converts detection into an operator response; and in command centres where video sits alongside other operational displays.
It disappoints in unstaffed rooms, where recorded search and alerting serve better; on very large grids scanned by one person, which dilute attention rather than adding it; in deployments where substreams were never configured and the decode machine is saturated; on static walls with no event response, which operators habituate to within days; and on projects where the panel budget consumed the money that should have specified the decode machine.
There is no single figure, because it depends on panel count and technology, tile count, camera resolution and the machine driving it. The more useful warning is that display quotations cover panels, mounts and installation, and omit the two lines that decide whether the wall works: the decode machine sized for your maximum tile count, and the network capacity to deliver those streams to it.
Not usually, if every source is a camera from one VMS. A machine running the VMS fullscreen across its display outputs is simpler and cheaper. Hardware controllers earn their cost on mixed source walls that combine cameras with other inputs such as dashboards, broadcast feeds or workstation outputs.
Technically as many as your layout has tiles, up to thirty six on a 6x6 grid per display, and more across multiple displays. Usefully, far fewer. Beyond roughly sixteen tiles in one person's field of view, extra tiles dilute attention rather than adding awareness. Use cycling to cover a large fleet rather than trying to display all of it at once.
LCD panels with narrow bezels are the cost effective default and are what most control rooms use. Direct view LED removes bezel lines entirely and looks considerably better, at a significantly higher price, which is usually justified when the room is a showpiece as well as an operations space. For pure operational use, thin bezel LCD is normally the better value.
The most common cause is that small tiles are requesting a substream profile the camera does not publish. The request fails, the player works through its transport fallback options, and the tile ends up rendering as offline even though the camera's main stream is healthy. Either configure substreams on the camera, or use a platform that checks whether a substream exists before requesting it.
A substream is a second, lower resolution stream published by the same camera alongside its full resolution main stream. It matters because a small tile cannot display full resolution anyway, so decoding the main stream into it wastes capacity producing detail the scaler discards. Using substreams for small tiles is the largest single reduction in a wall's decode load.
Yes on a platform that supports it, and this is the feature that makes a wall worth its cost. What to check is how alerting interacts with cycling. If the wall rotates on a timer, confirm that an unacknowledged alert holds the wall on the relevant camera rather than cycling away mid incident, and ask what happens after the operator acknowledges it.
Visylix does not charge per camera or per stream on any paid plan, so the number of cameras on the wall does not change the price. The video wall is part of the platform rather than a separately priced product.
It drives multiple displays by running one browser instance per display in Chromium kiosk mode, each with its own layout and camera set. It is a browser based wall rather than a native desktop application, and there is no native multi monitor client. For a wall driven by one machine with several outputs this works well, and it is a real architectural constraint worth knowing before you design around it.