← All work

Research · HRI 2026

SCOPE: A Natural-Language PTZ Camera Agent That Runs Entirely at the Edge

An air-gapped agent for pan-tilt-zoom cameras, and a digital twin environment where a result in simulation is a result you can deploy on.

Nikolaj Hindsbo presenting SCOPE at HRI 2026 beside a slide about PTZ camera agents.
Nikolaj Hindsbo presents the work at HRI 2026, Edinburgh

SCOPE on a Digital Twin: Offshore Oil Rig

See how SCOPE handles on-demand requests and reacts to sensor alerts on a digital twin of an oil rig.

SCOPE on a real camera

Same agent and tools on a physical AXIS PTZ camera; only the planner changes. First Qwen3-30B-A3B (mixture-of-experts), then the dense Qwen3-32B. The first plans and moves noticeably faster.

Two ways SCOPE works

On demand

A person or an agent asks

“We might have someone in the water. Search the ocean cameras and alert the rescue coordinator if you find anyone.”

  1. A person or another agent asks in plain words
  2. SCOPE picks the cameras and looks
  3. SCOPE answers, and alerts people if needed
Always on

A sensor flags a change

Detectors watch every feed. When something changes, SCOPE is told.

  1. A detector or sensor notices a change
  2. SCOPE decides whether it matters and looks closer
  3. SCOPE alerts the right person, or stays quiet

On demand

Always on

The problem

A simple question turns into a multimodal reasoning problem. The agent has to resolve which camera it is being asked about, drive that camera to a named location, and then keep panning and counting, turn by turn, until it has the answer or can tell the operator the request cannot be answered.

Seven camera frames showing the agent moving to the highway viewpoint and panning right in 15-degree steps, counting traffic cones at each step until it reaches seven.
A multi-step instruction (go to the highway, then pan right in 15-degree steps until six cones are in view), executed on a physical AXIS PTZ camera. The agent recalls the highway viewpoint, then alternates pan and count until the stop condition is met: three cones, then five, then seven. Every step is a tool call, and the whole trace is what the judge scores.

SCOPE does all of this with small open-source vision and language models, sized to run inside the compute envelope of the edge hardware we already ship to sites. No frame and no instruction leaves the site.

Designing the system is one problem. Understanding what works best, and why, is another. That second one splits in two, and each half is measured separately:

Both of those rest on something more basic: the camera has to be usable by a small model at all. Its capabilities are exposed as tools, and the model has to hold up both ends of that exchange:

The interface was as much of the result as the models were. Writing a system prompt that tells a small model how to act in a real environment, and designing tool responses terse enough for it to actually use, moved accuracy as much as swapping the models did. That is the half that gets overlooked, and it is the half a simulator actually lets you improve: a prompt change or a shorter tool response can be A/B tested against a fixed scene, which is not something you can do against a live camera.

What we built

Our camera agent is built on three parts: a language model that plans, a vision model that looks, and a set of tools we wrote to drive the camera. Several of those tools are small workflows in their own right, wrapping a vision-language ability behind a single call rather than exposing the model directly. When the agent answers wrong, any of the three can be the cause, and a single accuracy score does not say which. That is the multimodal alignment problem, and it is the reason the architecture keeps all three separately addressable.

SCOPE agent architecture The user sends a query to the language model. The language model calls PTZ tools, which drive camera control and change the scene, and perception tools, which call the vision model. Tool results return to the language model, which returns a final answer to the user. User Small Language Model TOOLS PTZ tools Perception tools Camera control Vision model Scene sim or camera query scene update call frame tool results final answer
The planner treats perception as a tool, not as a token stream. Images never enter the planner's conversation, so latency stays predictable and either model can be swapped without touching the other. Because perception sits behind a tool boundary it does not have to be a VLM at all: a YOLO-class detector can answer the same call, which was an intentional part of the design. The same system design and tools drive the Blender simulation and the physical camera.

The action space is nine tools, exposed as an OpenAI-compatible tool schema, the same tool-calling pattern most modern agent runtimes share.

The nine affordances we expose to the agent. Some are workflows rather than single actions: zoom to object is a detect call followed by zoom logic, packaged as one tool. That decomposition is deliberate, because a small model asked to chain those steps itself loses accuracy fast.
ToolWhat it doesType
Pan, Tilt, ZoomMove the camera by a given amount, as a relative moveCamera control
Go to PresetJump to a named viewpoint, for example the highwayCamera control
Go HomeReturn to the home viewpoint, or to where the session startedCamera control
Get PresetsList the named viewpoints this camera actually hasCamera control
Take ImageCapture the current frame for the vision model to readCamera control
Count ObjectsCount how many of a described object are in viewVision
Answer a QuestionAnswer a natural-language question about what is in viewVision
Track ObjectFind a named object, then keep driving the camera to hold it in frameVision + control
Zoom to ObjectFind a described object, then zoom the camera until it fills the frameVision + control

Why this needed a new benchmark

No existing benchmark could tell us how well natural-language PTZ control would actually deploy. The public ones score a model in isolation against static images and fixed answers, so they never exercise the loop at all: nothing moves, there is no second look to get wrong, and no run depends on the one before it. Whatever number comes out carries no latency, no memory, and no account of how the failure happened.

Part of the motivation was exactly that question. Frontier and open models publish scores on those public suites and compete on them, and we wanted to know whether a better leaderboard number translates into a better agent on a real operational task. It does, up to a point, and where it stops translating is the more useful finding.

Testing in deployment cannot fill the gap either, because the real world will not hold still. Traffic moves, the light changes, and what is parked where is different by the afternoon. Run one planner today and another tomorrow and you have not compared them, you have compared two different scenes. Nothing can be repeated, so someone has to sit and watch each run and form an opinion, which makes the whole exercise manual, noisy and subjective.

In simulation the scene holds still, so two runs differ only by what you changed. That is what makes the comparison mean anything, and it is what the twin buys.

So we built one. 536 tasks across eight categories, on four public Blender scenes: urban streets with signage, occlusion, clutter and varying conditions. Each scene carries several fixed camera presets at different positions and angles, which is what makes viewpoint control meaningful rather than decorative, and it mirrors how a real PTZ install is set up.

Most tasks need the agent to move before it can answer, which is what makes this a benchmark for an agent rather than for a vision model:

Some tasks record an absence, so the correct answer is sometimes that the object is not there. All eight categories, their task counts, and how each is scored are in A1 and A3.

Five representative tasks in one scene released from the benchmark.

The agent’s system prompt is identical in simulation and on the physical AXIS camera. Blender exposes a camera object with the same affordances a real PTZ has, so the same nine tools drive both and nothing about the agent changes between them.

Robotics has used simulation for control policies for years. Language-driven agent systems mostly had not when we started. That is what this buys: a repeatable path from a code change to a deployment decision.

Why this matters for Armada

Armada deploys agents and models on edge hardware, often where the network is poor or absent. SCOPE is built for that constraint. Planning, perception, and camera control all run at the deployment site. The models are open-source and small enough to fit our container options and run air-gapped.

How small depends on which pairing you deploy, and the spread is wide enough that it changes the hardware conversation.

ModelRolePrecision WeightsFree on an 80 GB card
Moondream2-4bitPerceptionINT4~1.4 GB~79 GB
Moondream2PerceptionFP162.45 GB~78 GB
Qwen3-4BPlannerFP8~4 GB~76 GB
Qwen2.5-VL-7BPerceptionFP16~17 GB~63 GB
Moondream3 (9B total, 2B active)PerceptionFP16~18 GB~62 GB
Qwen3-30B-A3B (30.5B total, 3.3B active)PlannerFP8~31 GB~49 GB
Qwen3-30B-A3B (30.5B total, 3.3B active)PlannerBF16~61 GB~19 GB

The free column is what is left of the card once that model is resident, and it is the budget everything else has to live in: the other half of the pair, the KV cache for each camera stream, and whatever else the customer is running on the box. Compute is a separate constraint on top of memory, and often the binding one, because a card with spare VRAM can still be saturated by concurrent vision calls. Paired up, the agent runs from about 5 GB for the smallest working configuration to about 79 GB for the best-scoring one.

What we found

We tested 19 pairings of Qwen3 planners with Moondream and Qwen vision models. The best pairing completes 73.8% of the benchmark.

Small language models are already good enough to plan. A weak planner hallucinates and calls the wrong tools, and planner choice moves accuracy by up to 12 points while you are below that bar. Clear it and the planner stops being what limits the system: it routes tools correctly, holds a multi-step instruction, and gets out of the way. Nothing here needed a frontier model to do the thinking.

Perception was what held accuracy back, and it is what we had to address. Score every vision-attributed failure as correct and average accuracy rises from 66.2% to 82.0%, so roughly sixteen points were sitting behind the vision model alone. Control tasks are close to solved at 97 to 100%. Counting (53 to 77%) and text recognition (39 to 83%) are where the work remains. The same holds for time: a planner call takes about 1.4 seconds and a vision call takes several times longer.

Modular beat monolithic, and planner size saturates. A mixture-of-experts planner activating 3B parameters per token matches or beats a dense 32B model and runs about twice as fast. Scaling to 80B buys no clear gain. Quantization is a real trade-off, not a free one: FP8 costs about half a point of accuracy on average, and more on some pairings (the best one drops from 73.8% to 69.1%). It is worth taking when memory or latency is the constraint. Two specialised halves, each replaceable on its own evidence, outperformed spending the budget on one larger model.

What we extended from OPUS

SCOPE is a continuation of our own earlier work. OPUS (ICCAR 2025) showed that a small language model can control a PTZ camera, but it had to be fine-tuned to do it. SCOPE does not. The camera's skills are exposed as tools that an off-the-shelf open-weight model can call, so the capability arrives without a training run.

The rest of the gap is perception and evidence. SCOPE swaps the fixed object detector for open-vocabulary vision, and adds a simulator that makes a run repeatable. That is what lets us say which models, at what latency, and failing how. Those questions were open when OPUS was written. They are narrower now, and still not obvious.

What has changed since we published

Vision-language models now ship with tool calling out of the box, and a planner and a perception model can sit on the same box without heroics. The integration cost of keeping them separate is genuinely lower than it was when we ran these experiments, and our own systems keep moving as each benchmark run makes those shifts visible.

What has not changed is structural rather than temporary. Model abilities advance asynchronously. A vision model gets better at counting in one release; a language model gets better at holding a multi-step instruction in another. A modular system made of two components, each optimized for half the problem, can beat a single generalist on latency, memory and accuracy at the same time, precisely because neither half is carrying capability it does not need.

So the most capable planner and the most capable perception model are still not the same model. Keeping them separate is what lets each one be chosen, quantized, and evaluated on a benchmark.

Agents that drive 3D tools, and real-to-sim-to-real pipelines, existed before. What is new is that current frontier models, such as the Claude Opus 5 and GPT-6 generations, now build good scenes when they drive tools like Blender. For robotics, that makes the sim-to-real approach faster to iterate on, and more exciting to work on, than it has been. It also makes agentic real-world evaluations like SCOPE easier to build and extend.

Technical supplement: benchmark composition, best configuration by category, failure modes, every configuration

A1. Benchmark composition

536 tasks across eight categories.
CategoryTasksCategoryTasks
Counting95Single call72
Descriptor89Multi-step command57
Location / spatial53Multi-step reasoning54
OCR identification54Comparative / relational62

Most tasks require an explicit viewpoint change before they can be answered. Tasks include recorded absences (for example, “no pedestrians in the parking lot”), so the correct answer is sometimes that an object is not present.

A2. Best configuration by category

Qwen3-30B-A3B planner × Moondream3 perception.
MeasurementValue
Overall task success73.8%
Single call98.6%
Multi-step command98.2%
Multi-step reasoning79.6%
OCR identification77.8%
Counting76.8%
Location / spatial66.0%
Comparative / relational56.5%
Descriptor37.1%

Accuracy is the fraction of executions judged correct by the LM-as-Judge (gpt-oss-120B) reading the full trace. All eight categories are listed, including the weakest. Descriptor tasks are open-ended free-text descriptions, and other pairings score higher there (up to 59.6%), so the best overall configuration is not the best in every category.

A3. What limits accuracy, and how it fails

Scoring is done by a model-based judge (gpt-oss-120B) that reads the entire execution trace, not just the final answer. For every failure it assigns exactly one of eight causes. That attribution is the point: a bare score says a configuration is worse, an attributed score says whether to change the planner, the perception model, or the prompt.

Six measurements decide where the effort goes. The left column is the question in plain terms, the right column is why it changes a deployment decision.

What we measuredValueWhy it matters
Swapping the planner, best against worst5.6 to 12.4 ptsThe biggest single lever, and it saturates
Swapping the vision model, best against worst≤ 6.9 ptsA narrower range, but it sets the ceiling
Score if the vision model never erred66.2% → 82.0%15.8 points sit behind perception alone
Failures that are perception's fault46.7%Nearly half of everything still wrong
Asking about the current view against the whole scene66.9% vs 58.3%Broader questions are answered worse
Time cost of sweeping the whole scene+22.4%A panorama is a larger-resolution image

Three things follow from that table.

  1. Clear the planner bar, then stop paying. Planner choice moves accuracy more than anything else, and then the returns die. Buying a bigger planner past that point buys latency, not accuracy.
  2. Perception is the ceiling. Vision moves the score less, but it owns most of what is still broken. The 15.8-point oracle gap is headroom a better perception model would unlock, not a result we achieved.
  3. Ask the narrowest question that answers the operator. Full-scene sweeps cost 8.6 points of accuracy and 22.4% more time than a current-view query.

The failure side is what makes those numbers actionable. When the judge marks an execution incorrect it assigns exactly one cause out of eight, so a score becomes a diagnosis instead of an anecdote.

Error modeMeaning
Vision, queryIncorrect or incomplete visual read (OCR, attributes)
Vision, countingMiscounts objects, or misses occluded items
ReasoningWrong conclusion despite correct tool outputs
HallucinationClaims facts not grounded in tool evidence
Lack of tool callA required tool was never invoked
Tool argumentsWrong arguments or targets (wrong preset, wrong object)
ScopeWrong observation scope (current view vs full scene)
Tool routingWrong tool chosen, or a required tool omitted
Grouped bar chart of error counts by error mode across all planner and perception pairings. Vision query is the tallest group, followed by reasoning, then hallucination and lack of tool call. Tool routing and scope are near zero.
Error counts by mode across every pairing. Figure from the HRI '26 paper (arXiv:2606.02951).

Reading the distribution against the taxonomy gives the decomposition the paper is built on:

  • Weak planners fail at the interface, not at the task. Their errors are hallucinations and argument-level mistakes, such as issuing “car” where the instruction said “car to the right”.
  • Those failures fall away as the planner gets stronger. Tool-routing and scope-selection errors are rare for capable planners, which is exactly where the diminishing returns from further SLM scaling show up.
  • What is left is vision. Query and counting errors dominate every pairing, and for the strongest planners they are the largest share of everything still wrong.
  • The system moves from planner-limited to perception-limited as planner capacity rises. That transition is the most useful thing the decomposition shows, because it says where the next engineering effort belongs.
  • Perception-heavy categories are simply harder. Counting (52.6 to 76.8%) and OCR (38.9 to 83.3%) sit far below tool-centric ones such as Single Call (97.2 to 100%) and Multi-step Command (86.0 to 98.2%).

A4. Every configuration we ran

Each block is one perception model; the rows inside it are the planners paired with it. Accuracy in percent by category. Blocks are VLM families; within a block, planners run from 4B up to 80B-A3B with quantized variants after their full-precision counterparts.

Average is the unweighted mean across the eight categories. The highlighted row is the best overall.
Planner (SLM)Comp/RelCountDescr Loc/SpatMS CmdMS Reas OCRSingleAverage
Perception (VLM): Moondream2
Qwen3-4B27.468.438.252.889.564.838.997.259.7
Qwen3-4B-FP827.469.540.456.689.561.140.797.260.3
Qwen3-30B-A3B43.570.544.952.893.081.553.798.667.3
Qwen3-30B-A3B-FP850.072.650.656.689.579.659.398.669.6
Qwen3-32B41.970.546.164.296.572.253.798.668.0
Qwen3-Next-80B-A3B46.876.848.358.596.583.355.698.670.6
Perception (VLM): Moondream2-4bit
Qwen3-4B32.362.147.250.991.266.748.1100.062.3
Qwen3-4B-FP827.462.148.352.889.566.746.397.261.3
Qwen3-30B-A3B41.964.249.458.598.272.251.998.666.9
Qwen3-30B-A3B-FP850.065.355.150.993.072.240.798.665.7
Qwen3-32B37.160.059.664.287.768.551.998.665.9
Qwen3-Next-80B-A3B36.565.352.858.598.275.948.198.666.8
Perception (VLM): Moondream3
Qwen3-4B32.368.433.747.287.755.666.7100.061.4
Qwen3-4B-FP830.673.734.858.587.759.357.498.662.6
Qwen3-30B-A3B56.576.837.166.098.279.677.898.673.8
Qwen3-30B-A3B-FP848.474.738.252.886.075.977.898.669.1
Qwen3-32B45.268.439.352.889.572.283.3100.068.8
Qwen3-Next-80B-A3B45.272.634.847.298.274.183.398.669.3
Perception (VLM): Qwen2.5-VL-7B
Qwen3-Next-80B-A3B33.952.652.866.093.072.275.9100.068.3

Cite

Hindsbo, Ehsani, Mishra (2026). SCOPE: Real-Time Natural Language Camera Agent at the Edge. In Proceedings of the 21st ACM/IEEE International Conference on Human-Robot Interaction (HRI '26).

@inproceedings{hindsbo2026scope,
  title     = {SCOPE: Real-Time Natural Language Camera Agent at the Edge:
               A Sim-to-Real Benchmark and Analysis of Open-Source Vision
               and Language Agents for PTZ Camera Tasks},
  author    = {Hindsbo, Nikolaj and Ehsani, Sina and Mishra, Pragyana},
  booktitle = {Proceedings of the 21st ACM/IEEE International Conference
               on Human-Robot Interaction},
  series    = {HRI '26},
  year      = {2026},
  month     = mar,
  numpages  = {9},
  isbn      = {979-8-4007-2128-1},
  publisher = {ACM},
  address   = {New York, NY, USA},
  location  = {Edinburgh, Scotland, UK},
  doi       = {10.1145/3757279.3785641}
}

Figures from the paper are published under CC BY 4.0.