Research · HRI 2026
SCOPE: A Natural-Language PTZ Camera Agent That Runs Entirely at the Edge
An air-gapped agent for pan-tilt-zoom cameras, and a digital twin environment where a result in simulation is a result you can deploy on.
SCOPE on a Digital Twin: Offshore Oil Rig
SCOPE on a real camera
Two ways SCOPE works
A person or an agent asks
“We might have someone in the water. Search the ocean cameras and alert the rescue coordinator if you find anyone.”
- A person or another agent asks in plain words
- SCOPE picks the cameras and looks
- SCOPE answers, and alerts people if needed
A sensor flags a change
Detectors watch every feed. When something changes, SCOPE is told.
- A detector or sensor notices a change
- SCOPE decides whether it matters and looks closer
- SCOPE alerts the right person, or stays quiet
On demand
Always on
The problem
A simple question turns into a multimodal reasoning problem. The agent has to resolve which camera it is being asked about, drive that camera to a named location, and then keep panning and counting, turn by turn, until it has the answer or can tell the operator the request cannot be answered.
SCOPE does all of this with small open-source vision and language models, sized to run inside the compute envelope of the edge hardware we already ship to sites. No frame and no instruction leaves the site.
Designing the system is one problem. Understanding what works best, and why, is another. That second one splits in two, and each half is measured separately:
- Accuracy. Which planner, paired with which perception model, gets the answer right, and what kind of mistake does it make when it does not.
- Can it run at the edge? Whether that pairing fits the latency, memory and compute a site actually has.
Both of those rest on something more basic: the camera has to be usable by a small model at all. Its capabilities are exposed as tools, and the model has to hold up both ends of that exchange:
- Call the right tool. Choose the right action and pass sane arguments. Pan by 15 degrees, not 15 percent.
- Understand what comes back. Read the tool's response, work out what it says about the scene, and decide the next move from it. In few enough tokens that the loop stays fast.
The interface was as much of the result as the models were. Writing a system prompt that tells a small model how to act in a real environment, and designing tool responses terse enough for it to actually use, moved accuracy as much as swapping the models did. That is the half that gets overlooked, and it is the half a simulator actually lets you improve: a prompt change or a shorter tool response can be A/B tested against a fixed scene, which is not something you can do against a live camera.
What we built
Our camera agent is built on three parts: a language model that plans, a vision model that looks, and a set of tools we wrote to drive the camera. Several of those tools are small workflows in their own right, wrapping a vision-language ability behind a single call rather than exposing the model directly. When the agent answers wrong, any of the three can be the cause, and a single accuracy score does not say which. That is the multimodal alignment problem, and it is the reason the architecture keeps all three separately addressable.
The action space is nine tools, exposed as an OpenAI-compatible tool schema, the same tool-calling pattern most modern agent runtimes share.
| Tool | What it does | Type |
|---|---|---|
| Pan, Tilt, Zoom | Move the camera by a given amount, as a relative move | Camera control |
| Go to Preset | Jump to a named viewpoint, for example the highway | Camera control |
| Go Home | Return to the home viewpoint, or to where the session started | Camera control |
| Get Presets | List the named viewpoints this camera actually has | Camera control |
| Take Image | Capture the current frame for the vision model to read | Camera control |
| Count Objects | Count how many of a described object are in view | Vision |
| Answer a Question | Answer a natural-language question about what is in view | Vision |
| Track Object | Find a named object, then keep driving the camera to hold it in frame | Vision + control |
| Zoom to Object | Find a described object, then zoom the camera until it fills the frame | Vision + control |
Why this needed a new benchmark
No existing benchmark could tell us how well natural-language PTZ control would actually deploy. The public ones score a model in isolation against static images and fixed answers, so they never exercise the loop at all: nothing moves, there is no second look to get wrong, and no run depends on the one before it. Whatever number comes out carries no latency, no memory, and no account of how the failure happened.
Part of the motivation was exactly that question. Frontier and open models publish scores on those public suites and compete on them, and we wanted to know whether a better leaderboard number translates into a better agent on a real operational task. It does, up to a point, and where it stops translating is the more useful finding.
Testing in deployment cannot fill the gap either, because the real world will not hold still. Traffic moves, the light changes, and what is parked where is different by the afternoon. Run one planner today and another tomorrow and you have not compared them, you have compared two different scenes. Nothing can be repeated, so someone has to sit and watch each run and form an opinion, which makes the whole exercise manual, noisy and subjective.
In simulation the scene holds still, so two runs differ only by what you changed. That is what makes the comparison mean anything, and it is what the twin buys.
So we built one. 536 tasks across eight categories, on four public Blender scenes: urban streets with signage, occlusion, clutter and varying conditions. Each scene carries several fixed camera presets at different positions and angles, which is what makes viewpoint control meaningful rather than decorative, and it mirrors how a real PTZ install is set up.
Most tasks need the agent to move before it can answer, which is what makes this a benchmark for an agent rather than for a vision model:
- Counting. “Sweep the scene and count all cars.”
- OCR. “Zoom in on a storefront window and read the sign.”
- Multi-step command. “Are there any cars in the north and south viewpoints? If so, zoom in and return the plates.”
Some tasks record an absence, so the correct answer is sometimes that the object is not there. All eight categories, their task counts, and how each is scored are in A1 and A3.
The agent’s system prompt is identical in simulation and on the physical AXIS camera. Blender exposes a camera object with the same affordances a real PTZ has, so the same nine tools drive both and nothing about the agent changes between them.
Robotics has used simulation for control policies for years. Language-driven agent systems mostly had not when we started. That is what this buys: a repeatable path from a code change to a deployment decision.
Why this matters for Armada
Armada deploys agents and models on edge hardware, often where the network is poor or absent. SCOPE is built for that constraint. Planning, perception, and camera control all run at the deployment site. The models are open-source and small enough to fit our container options and run air-gapped.
How small depends on which pairing you deploy, and the spread is wide enough that it changes the hardware conversation.
| Model | Role | Precision | Weights | Free on an 80 GB card |
|---|---|---|---|---|
| Moondream2-4bit | Perception | INT4 | ~1.4 GB | ~79 GB |
| Moondream2 | Perception | FP16 | 2.45 GB | ~78 GB |
| Qwen3-4B | Planner | FP8 | ~4 GB | ~76 GB |
| Qwen2.5-VL-7B | Perception | FP16 | ~17 GB | ~63 GB |
| Moondream3 (9B total, 2B active) | Perception | FP16 | ~18 GB | ~62 GB |
| Qwen3-30B-A3B (30.5B total, 3.3B active) | Planner | FP8 | ~31 GB | ~49 GB |
| Qwen3-30B-A3B (30.5B total, 3.3B active) | Planner | BF16 | ~61 GB | ~19 GB |
The free column is what is left of the card once that model is resident, and it is the budget everything else has to live in: the other half of the pair, the KV cache for each camera stream, and whatever else the customer is running on the box. Compute is a separate constraint on top of memory, and often the binding one, because a card with spare VRAM can still be saturated by concurrent vision calls. Paired up, the agent runs from about 5 GB for the smallest working configuration to about 79 GB for the best-scoring one.
What we found
We tested 19 pairings of Qwen3 planners with Moondream and Qwen vision models. The best pairing completes 73.8% of the benchmark.
Small language models are already good enough to plan. A weak planner hallucinates and calls the wrong tools, and planner choice moves accuracy by up to 12 points while you are below that bar. Clear it and the planner stops being what limits the system: it routes tools correctly, holds a multi-step instruction, and gets out of the way. Nothing here needed a frontier model to do the thinking.
Perception was what held accuracy back, and it is what we had to address. Score every vision-attributed failure as correct and average accuracy rises from 66.2% to 82.0%, so roughly sixteen points were sitting behind the vision model alone. Control tasks are close to solved at 97 to 100%. Counting (53 to 77%) and text recognition (39 to 83%) are where the work remains. The same holds for time: a planner call takes about 1.4 seconds and a vision call takes several times longer.
Modular beat monolithic, and planner size saturates. A mixture-of-experts planner activating 3B parameters per token matches or beats a dense 32B model and runs about twice as fast. Scaling to 80B buys no clear gain. Quantization is a real trade-off, not a free one: FP8 costs about half a point of accuracy on average, and more on some pairings (the best one drops from 73.8% to 69.1%). It is worth taking when memory or latency is the constraint. Two specialised halves, each replaceable on its own evidence, outperformed spending the budget on one larger model.
What we extended from OPUS
SCOPE is a continuation of our own earlier work. OPUS (ICCAR 2025) showed that a small language model can control a PTZ camera, but it had to be fine-tuned to do it. SCOPE does not. The camera's skills are exposed as tools that an off-the-shelf open-weight model can call, so the capability arrives without a training run.
The rest of the gap is perception and evidence. SCOPE swaps the fixed object detector for open-vocabulary vision, and adds a simulator that makes a run repeatable. That is what lets us say which models, at what latency, and failing how. Those questions were open when OPUS was written. They are narrower now, and still not obvious.
What has changed since we published
Vision-language models now ship with tool calling out of the box, and a planner and a perception model can sit on the same box without heroics. The integration cost of keeping them separate is genuinely lower than it was when we ran these experiments, and our own systems keep moving as each benchmark run makes those shifts visible.
What has not changed is structural rather than temporary. Model abilities advance asynchronously. A vision model gets better at counting in one release; a language model gets better at holding a multi-step instruction in another. A modular system made of two components, each optimized for half the problem, can beat a single generalist on latency, memory and accuracy at the same time, precisely because neither half is carrying capability it does not need.
So the most capable planner and the most capable perception model are still not the same model. Keeping them separate is what lets each one be chosen, quantized, and evaluated on a benchmark.
Agents that drive 3D tools, and real-to-sim-to-real pipelines, existed before. What is new is that current frontier models, such as the Claude Opus 5 and GPT-6 generations, now build good scenes when they drive tools like Blender. For robotics, that makes the sim-to-real approach faster to iterate on, and more exciting to work on, than it has been. It also makes agentic real-world evaluations like SCOPE easier to build and extend.
Technical supplement: benchmark composition, best configuration by category, failure modes, every configuration
A1. Benchmark composition
| Category | Tasks | Category | Tasks |
|---|---|---|---|
| Counting | 95 | Single call | 72 |
| Descriptor | 89 | Multi-step command | 57 |
| Location / spatial | 53 | Multi-step reasoning | 54 |
| OCR identification | 54 | Comparative / relational | 62 |
Most tasks require an explicit viewpoint change before they can be answered. Tasks include recorded absences (for example, “no pedestrians in the parking lot”), so the correct answer is sometimes that an object is not present.
A2. Best configuration by category
| Measurement | Value |
|---|---|
| Overall task success | 73.8% |
| Single call | 98.6% |
| Multi-step command | 98.2% |
| Multi-step reasoning | 79.6% |
| OCR identification | 77.8% |
| Counting | 76.8% |
| Location / spatial | 66.0% |
| Comparative / relational | 56.5% |
| Descriptor | 37.1% |
Accuracy is the fraction of executions judged correct by the LM-as-Judge (gpt-oss-120B) reading the full trace. All eight categories are listed, including the weakest. Descriptor tasks are open-ended free-text descriptions, and other pairings score higher there (up to 59.6%), so the best overall configuration is not the best in every category.
A3. What limits accuracy, and how it fails
Scoring is done by a model-based judge (gpt-oss-120B) that reads the entire execution trace, not just the final answer. For every failure it assigns exactly one of eight causes. That attribution is the point: a bare score says a configuration is worse, an attributed score says whether to change the planner, the perception model, or the prompt.
Six measurements decide where the effort goes. The left column is the question in plain terms, the right column is why it changes a deployment decision.
| What we measured | Value | Why it matters |
|---|---|---|
| Swapping the planner, best against worst | 5.6 to 12.4 pts | The biggest single lever, and it saturates |
| Swapping the vision model, best against worst | ≤ 6.9 pts | A narrower range, but it sets the ceiling |
| Score if the vision model never erred | 66.2% → 82.0% | 15.8 points sit behind perception alone |
| Failures that are perception's fault | 46.7% | Nearly half of everything still wrong |
| Asking about the current view against the whole scene | 66.9% vs 58.3% | Broader questions are answered worse |
| Time cost of sweeping the whole scene | +22.4% | A panorama is a larger-resolution image |
Three things follow from that table.
- Clear the planner bar, then stop paying. Planner choice moves accuracy more than anything else, and then the returns die. Buying a bigger planner past that point buys latency, not accuracy.
- Perception is the ceiling. Vision moves the score less, but it owns most of what is still broken. The 15.8-point oracle gap is headroom a better perception model would unlock, not a result we achieved.
- Ask the narrowest question that answers the operator. Full-scene sweeps cost 8.6 points of accuracy and 22.4% more time than a current-view query.
The failure side is what makes those numbers actionable. When the judge marks an execution incorrect it assigns exactly one cause out of eight, so a score becomes a diagnosis instead of an anecdote.
| Error mode | Meaning |
|---|---|
| Vision, query | Incorrect or incomplete visual read (OCR, attributes) |
| Vision, counting | Miscounts objects, or misses occluded items |
| Reasoning | Wrong conclusion despite correct tool outputs |
| Hallucination | Claims facts not grounded in tool evidence |
| Lack of tool call | A required tool was never invoked |
| Tool arguments | Wrong arguments or targets (wrong preset, wrong object) |
| Scope | Wrong observation scope (current view vs full scene) |
| Tool routing | Wrong tool chosen, or a required tool omitted |
Reading the distribution against the taxonomy gives the decomposition the paper is built on:
- Weak planners fail at the interface, not at the task. Their errors are hallucinations and argument-level mistakes, such as issuing “car” where the instruction said “car to the right”.
- Those failures fall away as the planner gets stronger. Tool-routing and scope-selection errors are rare for capable planners, which is exactly where the diminishing returns from further SLM scaling show up.
- What is left is vision. Query and counting errors dominate every pairing, and for the strongest planners they are the largest share of everything still wrong.
- The system moves from planner-limited to perception-limited as planner capacity rises. That transition is the most useful thing the decomposition shows, because it says where the next engineering effort belongs.
- Perception-heavy categories are simply harder. Counting (52.6 to 76.8%) and OCR (38.9 to 83.3%) sit far below tool-centric ones such as Single Call (97.2 to 100%) and Multi-step Command (86.0 to 98.2%).
A4. Every configuration we ran
Each block is one perception model; the rows inside it are the planners paired with it. Accuracy in percent by category. Blocks are VLM families; within a block, planners run from 4B up to 80B-A3B with quantized variants after their full-precision counterparts.
| Planner (SLM) | Comp/Rel | Count | Descr | Loc/Spat | MS Cmd | MS Reas | OCR | Single | Average |
|---|---|---|---|---|---|---|---|---|---|
| Perception (VLM): Moondream2 | |||||||||
| Qwen3-4B | 27.4 | 68.4 | 38.2 | 52.8 | 89.5 | 64.8 | 38.9 | 97.2 | 59.7 |
| Qwen3-4B-FP8 | 27.4 | 69.5 | 40.4 | 56.6 | 89.5 | 61.1 | 40.7 | 97.2 | 60.3 |
| Qwen3-30B-A3B | 43.5 | 70.5 | 44.9 | 52.8 | 93.0 | 81.5 | 53.7 | 98.6 | 67.3 |
| Qwen3-30B-A3B-FP8 | 50.0 | 72.6 | 50.6 | 56.6 | 89.5 | 79.6 | 59.3 | 98.6 | 69.6 |
| Qwen3-32B | 41.9 | 70.5 | 46.1 | 64.2 | 96.5 | 72.2 | 53.7 | 98.6 | 68.0 |
| Qwen3-Next-80B-A3B | 46.8 | 76.8 | 48.3 | 58.5 | 96.5 | 83.3 | 55.6 | 98.6 | 70.6 |
| Perception (VLM): Moondream2-4bit | |||||||||
| Qwen3-4B | 32.3 | 62.1 | 47.2 | 50.9 | 91.2 | 66.7 | 48.1 | 100.0 | 62.3 |
| Qwen3-4B-FP8 | 27.4 | 62.1 | 48.3 | 52.8 | 89.5 | 66.7 | 46.3 | 97.2 | 61.3 |
| Qwen3-30B-A3B | 41.9 | 64.2 | 49.4 | 58.5 | 98.2 | 72.2 | 51.9 | 98.6 | 66.9 |
| Qwen3-30B-A3B-FP8 | 50.0 | 65.3 | 55.1 | 50.9 | 93.0 | 72.2 | 40.7 | 98.6 | 65.7 |
| Qwen3-32B | 37.1 | 60.0 | 59.6 | 64.2 | 87.7 | 68.5 | 51.9 | 98.6 | 65.9 |
| Qwen3-Next-80B-A3B | 36.5 | 65.3 | 52.8 | 58.5 | 98.2 | 75.9 | 48.1 | 98.6 | 66.8 |
| Perception (VLM): Moondream3 | |||||||||
| Qwen3-4B | 32.3 | 68.4 | 33.7 | 47.2 | 87.7 | 55.6 | 66.7 | 100.0 | 61.4 |
| Qwen3-4B-FP8 | 30.6 | 73.7 | 34.8 | 58.5 | 87.7 | 59.3 | 57.4 | 98.6 | 62.6 |
| Qwen3-30B-A3B | 56.5 | 76.8 | 37.1 | 66.0 | 98.2 | 79.6 | 77.8 | 98.6 | 73.8 |
| Qwen3-30B-A3B-FP8 | 48.4 | 74.7 | 38.2 | 52.8 | 86.0 | 75.9 | 77.8 | 98.6 | 69.1 |
| Qwen3-32B | 45.2 | 68.4 | 39.3 | 52.8 | 89.5 | 72.2 | 83.3 | 100.0 | 68.8 |
| Qwen3-Next-80B-A3B | 45.2 | 72.6 | 34.8 | 47.2 | 98.2 | 74.1 | 83.3 | 98.6 | 69.3 |
| Perception (VLM): Qwen2.5-VL-7B | |||||||||
| Qwen3-Next-80B-A3B | 33.9 | 52.6 | 52.8 | 66.0 | 93.0 | 72.2 | 75.9 | 100.0 | 68.3 |
Cite
Hindsbo, Ehsani, Mishra (2026). SCOPE: Real-Time Natural Language Camera Agent at the Edge. In Proceedings of the 21st ACM/IEEE International Conference on Human-Robot Interaction (HRI '26).
@inproceedings{hindsbo2026scope,
title = {SCOPE: Real-Time Natural Language Camera Agent at the Edge:
A Sim-to-Real Benchmark and Analysis of Open-Source Vision
and Language Agents for PTZ Camera Tasks},
author = {Hindsbo, Nikolaj and Ehsani, Sina and Mishra, Pragyana},
booktitle = {Proceedings of the 21st ACM/IEEE International Conference
on Human-Robot Interaction},
series = {HRI '26},
year = {2026},
month = mar,
numpages = {9},
isbn = {979-8-4007-2128-1},
publisher = {ACM},
address = {New York, NY, USA},
location = {Edinburgh, Scotland, UK},
doi = {10.1145/3757279.3785641}
} Figures from the paper are published under CC BY 4.0.