VideoLLMs Conditioned on Visual Scene Graphs
A Global-Local Approach
1QpiAI, Bangalore, India
Abstract
Current VideoLLMs fit long videos into limited token budgets by aggressively downsampling frames, which can erase small objects and subtle spatial relationships. We instead detect objects at native resolution and encode every crop with a CLIP embedding and normalized spatial coordinates. Masked graph attention learns relationships among objects within a frame, and permutation-invariant pooling produces compact per-frame scene-graph tokens.
A complementary global stream encodes one reference frame for scene context. The language model receives both streams and reasons over the temporal sequence. Across three static-camera benchmarks, the method matches strong finetuned baselines while using 35% as many tokens as Qwen3-VL on average, transfers across LLM backbones, and generalizes zero-shot to dynamic-camera reasoning.
Method Overview
- Object extraction. RT-DETR detects objects in each frame. CLIP ViT-L/14 encodes native-resolution crops, and normalized box coordinates preserve spatial layout.
- Graph processing. Masked multi-head attention lets valid objects exchange visual and spatial information without requiring explicit edge annotations.
- Permutation-invariant aggregation. Max pooling converts each variable-sized object set into compact scene tokens independent of detector ordering.
- Global-local fusion. Scene tokens are concatenated with global tokens from a single reference frame and projected into the LLM input space for temporal reasoning.
Results
The central result is accuracy parity at a fraction of the token and memory cost.
633average tokens
65.4%fewer than Qwen3-VL
94.9%fewer than Cosmos-Reason1
18.71 GBaverage peak VRAM
Static-camera benchmarks
BLEU-4, ROUGE-L, METEOR, and BERTScore-F1; higher is better. The proposed Qwen3-VL variant remains within ±0.3 points of its finetuned base.
| Model | BridgeDataV2 | HoloAssist | RoboVQA | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| B-4 | R-L | Met | F1 | B-4 | R-L | Met | F1 | B-4 | R-L | Met | F1 | |
| Cosmos-Reason1-7B | 10.35 | 28.90 | 28.29 | 86.17 | 14.37 | 31.62 | 31.82 | 86.91 | 26.90 | 44.69 | 45.25 | 89.72 |
| Cosmos-Reason2-8B | 19.16 | 38.20 | 35.75 | 88.87 | 21.96 | 39.85 | 39.41 | 88.93 | 30.58 | 50.15 | 46.00 | 90.65 |
| Qwen3-VL-8B | 10.49 | 28.83 | 28.49 | 86.21 | 14.16 | 31.43 | 31.85 | 86.84 | 26.73 | 44.73 | 45.22 | 89.70 |
| InternVL3.5-8B | 9.34 | 19.13 | 13.25 | 83.42 | 9.66 | 18.68 | 11.26 | 83.68 | 12.54 | 20.86 | 15.48 | 84.09 |
| Ours (Qwen3-VL-8B) | 10.35 | 28.80 | 28.25 | 86.17 | 14.11 | 31.48 | 31.91 | 86.84 | 26.88 | 44.72 | 45.06 | 89.71 |
| Ours (InternVL3.5-8B) | 8.19 | 27.34 | 25.29 | 84.67 | 9.70 | 28.43 | 27.57 | 85.68 | 12.59 | 31.03 | 30.36 | 85.83 |
End-to-end efficiency
100-video, long and cluttered RoboVQA stress-test subset.
| Model | In tok ↓ | Out tok | Time ↓ | VRAM ↓ |
|---|---|---|---|---|
| Qwen3-VL-8B | 20,624 | 125 | 9.05 s | 35.33 GB |
| InternVL3.5-8B | 23,200 | 352 | 12.97 s | 56.92 GB |
| Ours (Qwen3-VL-8B) | 1,128 | 68 | 2.74 s | 18.08 GB |
| Ours (InternVL3.5-8B) | 8,618 | 85 | 4.25 s | 19.53 GB |
Representation ablation
500 unseen videos from BridgeDataV2 and HoloAssist.
| Model | METEOR ↑ | Tokens ↓ |
|---|---|---|
| Global stream only | 29.55 | 412 |
| Image + Symbolic | 29.20 | 4,777 |
| Image + Scene Graph | 30.84 | 507 |
Qualitative Results
BibTeX
Coming soon.