MCP (Model Context Protocol) is an open protocol that lets AI agents call the functions of external systems as "tools." Register an MCP server with an agent such as Claude Code, and the agent reads the list of tools that server exposes, picks the ones that fit the question, and calls them.
In NOX, each device is itself an MCP server. Once an administrator issues a dedicated AI agent key and registers it in Claude Code or Claude Desktop, a question like "Did any camera have a gap in recording last night?" leads the agent to check the camera list, look up the recorded ranges, and even generate a playback link for that moment in its answer. Not a single tool changes the state of the device.
This post is not a feature introduction but a development log. It records why the server ended up in its current shape, what we deliberately chose not to do, and what broke during reviews and testing on real devices. The content is based on the design documents, review, QA and test reports, and commit history in our repository.
Two moments in the video are worth noting. Asked whether any camera had a recording problem, Claude answered that no camera was flagged with a problem, yet it pointed out that the recordings of all three recording cameras start at the same moment and asked what had happened at that time. The final object search returned zero results, but instead of just answering "zero," Claude also looked up event statistics for the same period and identified the likely cause: events exist, but object information is not coming in. Most of the design principles described below were chosen to make answers like these possible.
Two False Starts
What we first had in mind was not an MCP server. In April 2026 we designed a feature that put a chat window inside the product UI, so users could talk to a language model inside the device. The documents went as far as review and security assessment, but never turned into code. Looking back, what we wanted to build was not a chat window but a channel through which agents can safely read the device's information. Chat windows are better built by Claude Code or Cursor, which users already use.
At the end of July we restarted with the direction changed to an MCP server. And then we lost the code for that work. The working folder holding the design documents had been excluded from git tracking, and the code had not yet been committed. What remained were a few empty directories, tables created in the development database, and the SDK version and download timestamp left in the Go module cache. A document that reconstructs the design by reading those traces backward is still in the repository. Its first line reads: "No MCP-related code has ever been committed to the repo."
Since then, all design documents are managed in git as well. That is also why this post could be written.
The third start came in early October. This time we narrowed the scope from the outset: no new authentication scheme, just a thin layer on top of the existing external API keys and permission system. The device-agent pairing, separate agent table, and agent-only access mode from the July design were not carried over.
Six Decisions Made Up Front
1. Start read-only
We did not include tools that pan cameras, turn recording on and off, or change settings. If an agent reads something wrong, the answer is merely wrong; if an agent writes something wrong, recording on site stops. We decided to first confirm that reading alone is useful, and to design writing separately as a later step.
2. No new key — add one more "purpose" to keys
NOX already has API keys for integrating with external systems. We added a purpose called AI Agent (MCP). One option was to give agent keys a new role "instead of" the existing read role. But then three existing checks that verify a key is still valid, along with the permission cleanup job, would treat agent keys as revoked. So we chose to keep the existing read role and "add" an agent role on top. No existing code had to be touched.
The list of allowed tools is signed into the key itself. That way we know what a key can see without querying the database on every request. The trade-off is that changing a key's tool scope requires reissuing it. If the scope needs to be narrowed urgently, revoking the key and issuing a new one takes effect immediately.
3. The MCP server never queries with its own privileges
When the MCP server asks an internal service for data, it does not use its own service credentials; it forwards the identity of the calling key as is. Whether it is the camera list or events, the service that owns the data re-evaluates permissions using the original key and the real client IP. The checks on the MCP server side exist to reject early and leave a record; in the end, the side that owns the data makes the call.
This principle forced us to change one path. We could have routed internal queries through the device's web server one more time, but requests coming in that way are treated as "local access by an administrator sitting in front of the device," and the client IP is overwritten with the device's own address. That would have neutralized keys restricted to allowed IPs. So requests go directly to internal services, and delegated requests are always marked as remote access. The authentication key that internal services use among themselves is never attached to a delegated request. The moment that key is attached, it becomes a path that exceeds the caller's privileges.
4. No sessions
The July design maintained sessions. In October we switched to a sessionless approach that redoes authentication and permission checks from scratch for every single request. The fact that a previous request passed does not vouch for the next one. For the SDK we chose the official Go SDK 1.8, the latest at the time. The MCP protocol itself was also moving toward a sessionless model. Responses are returned as a single JSON body, without streaming.
5. Tools cut from 14 to 10
The July design had 14 tools. The tools that separately showed the dashboard, disks, and system metrics were merged into two, "system status" and "storage status," and person search was deferred to a later stage. In their place we added a tool that finds scenes in already-indexed recordings using natural language.
The more tools there are, the more an agent has to choose from, and the more conversation space the tool descriptions take up. The remaining 10 fall into five groups: cameras, events and objects, recording and playback, system, and AI. Each tool description names the tool that should come before it. For example, the camera list tool says it is "the source of the camera_id passed to other tools." That single line lets the agent work out on its own the order in which to chain tools.
6. Never disguise failure as an empty result
Errors during tool execution are returned not as protocol errors but as tool results. The result carries an error flag along with a code such as RATE_LIMITED or PERMISSION_DENIED, so the model can read it and adjust its next action on its own. The rule written in the design document is short: "No successful empty results." If permissions cannot be verified or an internal service does not respond, that is an error, not zero results. People are suspicious of zero results; agents take zero at face value and build an answer on it.
Limits That Protect Recording and Live View
Agents ask far faster and far more than people do. And an NVR's first job is recording. If an agent's questions cause recording or live video to fall behind, it would be better not to have this feature at all.
| Scope | Default limit |
|---|---|
| MCP requests per key | 300 per minute |
| Tool calls per key | 60 per minute, burst of 20 |
| Natural-language scene searches per key | 6 per minute |
| Concurrent executions per key | 4 |
| Concurrent scene searches, device-wide | 2 |
| Concurrent internal queries, device-wide | 8 |
The limits can be lowered or raised in settings, but not turned off. The query period is also capped at 31 or 7 days depending on the tool, and the size of the results returned at once is capped as well.
What the Reviews Caught
Once the feature roughly worked, we ran separate code, authentication, and UI reviews. The first verdict of the tool-side review was "changes requested."
Tests passed, but raw queries were left in the audit log
We decided not to keep natural-language queries and license plate strings verbatim in audit logs, since queries can contain personal information. The audit records written by the MCP server were masked properly, and the test checking that passed. But the search service that the MCP server sent delegated requests to was writing the raw text into its own audit log. The test passed because it only looked at rows written by the MCP server. Now, when the requester is an API key, the search service also masks queries and plates in its records.
One heavy tool starved the rest
A natural-language scene search can take up to two minutes. Because these searches shared the device-wide pool of 8 internal query slots, a few concurrent scene searches caused even a lightweight camera list query to be rejected as "temporarily unavailable." Scene searches now use their own 2 dedicated slots.
Cameras outside the key's scope read as "zero"
In event statistics, cameras outside the key's permissions were silently dropped from the results. A person would wonder, "Why is this camera missing?" but an agent would answer, "That camera had zero events." Now the dropped cameras are reported separately in excluded_camera_ids.
Right shape, wrong values
The defects found in QA had something in common. The shape of the response JSON was correct, but the side receiving those values handled them differently. Because we had hand-crafted the fake responses used in tests, the tests were only checking that our assumptions and our code agreed with each other.
- We built an "oldest first" option for event search, but the event service always sorted newest first. The option was accepted and had no effect. At first we removed the option; the next day we wired sorting all the way through and brought it back.
- The recording range tool did not count the gap before the first recording block or after the last one. For a camera whose recording had stopped three hours earlier, it would answer "no gaps." Gaps are now calculated by subtracting the recorded ranges from the entire query range.
The Day We Connected Claude Code to a Real Device
The final test was to actually connect Claude Code to a development device and ask questions. We downloaded and specified the device's CA certificate, registered it with claude mcp add, and confirmed ✔ Connected in claude mcp list.
The first two questions went smoothly. Asked about camera status, it chained the camera list and health check tools and answered "all 9 online," matching the device dashboard. Event statistics for the past 24 hours also matched the dashboard numbers.
The third question was about missing recordings. The recording range tool said every currently recording camera had been empty for the full 24 hours, and the playback link tool said there was no video at that time. Claude did not simply relay these results as "All recordings are empty." It answered that "the information is contradictory."
The cause was not MCP but something underneath it. The query that groups the recording timeline was converting time zones twice, so in Seoul time every recording block was shifted 18 hours later. A defect that had stayed hidden because no shipped feature called this path together with the device time zone surfaced for the first time through the agent's calls. Details are written up separately below.
After the fix, asked the same question again, Claude called the tools in the order camera list → recording ranges → playback link and answered, "Both channels have only about the last 25 minutes recorded, coverage 1.7%." The development device had only 8GB of recording storage, so old footage was constantly being deleted; the answer matched exactly the actual recording distribution left in the database.
Next, the playback link returned "no video" even at times when video existed. Once again, Claude was the first to point out that "the coverage and the link result don't match." This too was a defect in the same path. After the fix, the answer "There is video at that time" came back with an accurate playback link.
Finally, we revoked the key. The first rejection came about 0.2 seconds after the revocation time was recorded, and claude mcp list showed ✘ Failed to connect … HTTP 401 TOKEN_INVALIDATED.
What stayed with us longest that day were the two "these don't match" moments. When tool results contradict each other, the agent notices. Put the other way, if a tool returns an empty result or a wrong value without any indication, the agent believes it as is. This test showed why the principle we set earlier, "never disguise failure as an empty result," is necessary.
The Time Zone Defect Found in Pre-Release Testing
First, the scope. Recorded video and recording records were not affected. What was wrong was a single line of calculation that groups a wide-range timeline for display. And because no shipped feature called this calculation with the device time zone, the defect never appeared on a customer's screen. We found and fixed it in real-device testing before releasing the MCP server. The fix commit landed before the MCP feature commit.
We still describe it in detail because what this defect causes you to miss ties directly to what we learned building the MCP server.
What the query did
When viewing the recording timeline over a wide range (a day or a week), the device does not return recording segments one by one; it groups them into blocks of a fixed length. The block boundaries must line up with the hour and midnight in the device's local time. For a device in Seoul, the "today 00:00 to 01:00" block must be based on Seoul time. This alignment was handled by a single line in a database (PostgreSQL) query.
to_timestamp(FLOOR(EXTRACT(EPOCH FROM seg.start_time AT TIME ZONE $4) / $5) * $5) AT TIME ZONE $4 AS block_start
$4 is the device time zone (e.g. Asia/Seoul) and $5 is the block length in seconds.
Following it step by step
Let's trace an actual case where the first recording segment started at 09:38 UTC (18:38 Seoul).
seg.start_time AT TIME ZONE 'Asia/Seoul'— converts 09:38 UTC to Seoul local time. The result is18:38with the time zone information stripped. So far, this is as intended.EXTRACT(EPOCH FROM …)— computes seconds elapsed since 1970. But a value without time zone information is treated as UTC. Seoul 18:38 becomes 18:38 UTC. Here the time is shifted by 9 hours once. Even so, dividing this value by the block length and truncating gives boundaries aligned to the local hour, so for alignment purposes it is still usable up to this step.to_timestamp(…)— turns the aligned seconds back into a timestamp. The result is 18:38 UTC with a time zone attached. A value already shifted by 9 hours is now set in stone as if it were the real time.… AT TIME ZONE 'Asia/Seoul'— converts that 18:38 UTC to Seoul local time again: 03:38 the next day, once more without time zone information. The time has been shifted by 9 hours a second time.- Finally, the Go database driver reads values without time zone information as UTC. The result is 03:38 UTC the next day, or 12:38 noon in Seoul.
A recording that actually started at 18:38 Seoul time appeared on the timeline as starting at 12:38 the next afternoon. In Seoul (UTC+9), it is shifted by twice the offset: 18 hours.
Why the test showed a 0% recording rate
Every block was shifted 18 hours later, that is, into the future. The recording range tool treats anything after the current time as "time that hasn't arrived yet" and cuts it off. That leaves no blocks within the past 24 hours. This is why cameras that were recording normally and storing video showed up in tool results as "0% recording rate, gap across the full 24 hours." What was wrong was the calculation, not the recording. The playback link tool looks at the same timeline, so it returned "no video at that time."
Why it had gone unnoticed
Three things overlapped.
- On devices whose time zone is UTC, the defect disappears. Shifted twice, the amount of shift is still zero. Calls that do not specify a time zone likewise come out correct.
- There was no test that actually executed the database query. Tests for this path used fake objects instead of a database, so the query statement itself had never run.
- No shipped feature called this query with the device time zone. The device web UI defined a function that called this aggregation API with a time zone, but it was not used on any actual screen, and no integration partner used this path through the external API. The MCP server always sends the device time zone with every internal query. The first caller to honestly follow this path all the way through was the agent.
How we fixed it
We inserted one AT TIME ZONE 'UTC' between steps 3 and 4.
(to_timestamp(FLOOR(EXTRACT(EPOCH FROM seg.start_time AT TIME ZONE $4) / $5) * $5) AT TIME ZONE 'UTC') AT TIME ZONE $4 AS block_start
After turning the aligned seconds back into a timestamp, it is immediately converted back into a "local time without time zone information," and that value is then interpreted as a time in the device time zone. This way the offset is never applied an extra time. In the same commit we added a regression test that connects to a real PostgreSQL instance and checks block times in the Seoul time zone. Since fake objects never execute this query, it was a defect that could only be caught by running against a real database.
Another one from the same period
The defect where the playback link returned "no video" even though video existed also came from code in the same path. It, too, never appeared on a shipped screen. The database stores track types as Video, Audio, and Metadata, but the code that converts the timeline into a response was comparing against lowercase video. As a result, the zoomed-in timeline always showed video as "absent." We made the comparison case-insensitive and added a test using the actual database values as is.
Both defects share the same root as the earlier QA defects: the tests were only looking at values we had crafted by hand.
Connecting an Agent to an Air-Gapped NVR
NVRs usually sit on an on-site network isolated from the internet. Camera video must not leave, and there must be no path from outside into the device. So how can Claude, running in the cloud, ask this device questions?
The answer is that the one actually calling the tools is not the Claude model but the user's PC.
Claude model (cloud) ⇄ Internet ⇄ User PC: Claude Desktop · Claude Code ⇄ On-site network ⇄ NOX device /agent/mcp
- The user asks a question in Claude Desktop or Claude Code on their PC.
- The Claude model in the cloud only returns a request like "call the camera list tool with these arguments." The model never connects to the device directly.
- Claude Desktop (or Claude Code) on the PC receives that request and calls the tool over HTTPS on the NOX device in the same on-site network.
- The PC receives the device's result and passes it back to the model, which builds its answer from that result.
That is why there is no need to expose the NVR to the internet. No port forwarding, no external relay server. Only one condition is required: the PC running Claude Desktop or Claude Code must reach both the internet (Claude) and the on-site network (the NVR). An operator PC in the control room is the typical place.
We do, however, make clear what leaves the site. Tool results — text such as camera names, status, event counts, recorded ranges, and playback links — are sent to the cloud model as part of the conversation. Video, thumbnails, camera access addresses, and credentials are never included in tool results in the first place. A playback link is only an address that opens the device's web UI, so it will not open outside the on-site network.
"Remote connectors," added by URL in the Claude app's settings, work differently. A remote connector connects to the MCP server from Anthropic's servers, so it cannot reach a device on an air-gapped network that is not reachable from the internet. For an isolated NVR, use Claude Desktop's local configuration or Claude Code, which connect directly from the PC.
How to Connect Claude Code and Claude Desktop
In the device web UI, go to Settings → API Keys, choose the purpose AI Agent (MCP), and issue a key; the result dialog shows both the endpoint address and the registration command.
claude mcp add --transport http nox-<key name> https://<device address>/agent/mcp \
--header "Authorization: Bearer <AI agent key>"
- The device certificate is signed by the device's internal CA. Download the CA certificate from the web UI, set it with
NODE_EXTRA_CA_CERTS=<file path>, and then launch Claude Code. - A trailing slash in the address results in a 404. It is
/agent/mcp, not/agent/mcp/. - After registering, check the connection status with
claude mcp list.
Claude Desktop registers MCP servers in a local configuration file (claude_desktop_config.json). Servers registered in this file run on the PC, so they can reach devices on the on-site network. To connect to a remote HTTP server with a header, route through mcp-remote, a relay tool that runs on the PC (Node.js required).
{
"mcpServers": {
"nox": {
"command": "npx",
"args": ["-y", "mcp-remote", "https://<device address>/agent/mcp",
"--header", "Authorization:Bearer ${NOX_KEY}"],
"env": {
"NOX_KEY": "<AI agent key>",
"NODE_EXTRA_CA_CERTS": "<CA certificate file path>"
}
}
}
}
- Keep the key only in
NOX_KEYunderenv, and reference it inargsas${NOX_KEY}. - After saving the configuration, restart Claude Desktop and the NOX tools will appear in the tool list. The video at the top of this post shows it connected and in use with Claude Desktop.
- Other agents such as Cursor work the same way: put the same address and header into their remote (HTTP) MCP server settings. In every case, the PC running the program must be able to reach the device.
- On the issuance screen, you can choose which tools each key is allowed to use. Tools you do not select do not appear in that key's tool list at all.
- An AI agent key can query every camera on the device. The issuance screen always displays this notice as well.
Once connected, you can ask questions like these.
- "Are any cameras offline right now, or has any stopped recording?"
- "Tell me the five cameras with the most events over the past week."
- "Check whether the front gate camera had any gaps in yesterday's recording, and if so, give me a playback link from just before it."
- "Is the free storage space and disk health okay?"
- "Find the scene where a white truck passed through the parking lot yesterday evening." (devices with a natural-language scene search license)
By the Numbers
| Item | Details |
|---|---|
| Tools | 10 read, 0 write |
| Checks before a request reaches a tool | 8 stages |
| MCP server tests | 25 test files, 202 test functions |
| Additional tests in related services | 13 files with agent or mcp in the name, 93 test functions |
| Audit records | One row per tool call, for both successes and failures |
Not Done Yet
- Write tools. There are no tools that change cameras, recording, or settings.
- Per-camera key scope. Narrowing a key so that only specific cameras are visible is not yet supported.
- Access with user accounts. Login tokens cannot be used; only AI agent keys are accepted.
- Sessions and streaming. There are no server notifications or tool list change notifications.
- Remote devices. Cameras on other devices connected via federation are not queried.
- NIS security mode. On devices with this mode enabled, the MCP endpoint does not exist at all.
Conclusion
The skeleton of the MCP server — the path that accepts the protocol, serves the tool list, and answers by asking internal services — did not take long thanks to the official SDK. Where we spent our time was in limiting what an agent can do inside the device, and in making sure the answers an agent receives do not differ from the facts.
The defects caught in review and those found on the real device were mostly of the same kind: a result that is empty or a value that is wrong, with no indication of it. A person looking at the screen senses something is off, but an agent trusts the value it received and calls the next tool. So we settled on one criterion: whether an MCP server is well built should be judged not by how many tools it has, but by how rarely an agent receives wrong values as fact.
And the first to put that criterion to the test was not our test suite, but Claude Code.
👉 See the real operating screens on the NOX product page — natural-language scene search and the system dashboard
For NOX-based AI agent integration and custom development inquiries: yiyol.com/contact
Related Posts
- Building NVR Failover — Goals of n:1 Redundancy, Failure Cases, and the Work to Leave Beta — Another development log written the same way
- What Is an Agentic AI NVR? The Decisive Difference from an AI NVR — A structure that extends from detection to judgment and action
- Agentic AI with WebHooks: A Complete Guide to Parking Lot Automation — Another way to connect agents with external systems