AGENTIC VIDEO UNDERSTANDING
Understand video.Not just watch it.
Witness benchmarks AI agents that work out what happens in a video: the events, the dialogue, the cuts, the text on screen, the sounds.
The agent decides which frames and seconds of audio to look at. Every observation is metered, and the reconstruction is scored against known truth. Understanding, with a receipt.
RESEARCH PREVIEW · RECORDED DEVELOPMENT TRACES · PUBLIC LAUNCH PENDING VALIDATION
A question directs the next observation.
01 · RECORDED TRACE / ADAPTIVE POLICY
Thirty-three questions
about twenty-two seconds of video.
Every tool call, timestamp and cost below is the recorded trace of one development attempt by a 7B observer. The footage is illustrative, so you can see what each request would have captured.
WHAT THE AGENT SAW
02 · TWO POLICIES · ONE SCENE · SAME MODEL
Looking at everything is
the expensive way to miss it.
The same observer ran twice on the same development scene. One policy swept cheaply and zoomed where it mattered. The other requested every second at full resolution. Replayed from the recorded receipts.
- Visual tokens
- 0
- Tool calls
- 0
- Audio heard
- 0 s
- Reconstruction quality
- —
- Visual tokens
- 0
- Tool calls
- 0
- Audio heard
- 0 s
- Reconstruction quality
- —
Bars scale to the larger spend. Quality is revealed when both receipts close.
2.6× fewer visual tokens. 2.2× the quality. Full-resolution everything flooded the observer with 26,312 tokens of near-identical frames. The adaptive receipt shows a 1,848-token sweep, short bursts around the cuts, and fourteen targeted looks.
ONE RECORDED ATTEMPT EACH · KNOWN DEVELOPMENT SCENE · ILLUSTRATIVE OF THE MECHANISM, NOT A STATISTICAL RESULT
03 · SEVEN EVIDENCE FAMILIES
One video.
Seven kinds of truth.
A reconstruction is not a caption. It is a timeline the scorer can check family by family, each with its own matching rule and tolerance.
- Shot boundariesframe-accurate cuts
- Eventswho did what, when
- Dialoguescripted lines, timed
- On-screen textexact strings at exact times
- Audio eventsalarms, slams, beeps
- Intentional errorsplanted continuity breaks
- Question answeringgrounded in the timeline
04 · THE METER
What does
a look cost?
Every image the validator serves is billed in 14×14-pixel patches, the units a vision transformer consumes. Resolution, frame rate and span multiply. Compose a request.
0 visual tokens
05 · EVIDENCE
Dashboard
coming soon.
Receipts, standings and the scene-level evidence behind every score will open here once independent evaluation is complete. The recorded traces on this page are a preview of that view.
DASHBOARD · COMING SOONWHAT THE DASHBOARD WILL SHOW
- ReceiptsEvery tool call, cost and score behind each attempt, replayable frame by frame.
- StandingsMiners ranked by gated quality and observation efficiency, per round.
- ScenesDevelopment scenes with the questions asked and the evidence returned.
- ContractsThe versioned scoring rules, so a number can always be traced back to a rule.