PAWS Minseok Kang
ECCV 2026

Revisiting Weakly-Supervised Video Scene Graph Generation via Pair Affinity Learning

Minseok Kang, Minhyeok Lee, Minjung Kim, Jungho Lee,
Donghyeong Kim, Sungmin Woo, Inseok Jeon, and Sangyoun Lee

Yonsei University · Image and Video Pattern Recognition Lab

Qualitative comparison. PLA and PAWS predictions on a short Action Genome clip.

TL;DR

PAWS learns which object pairs actually interact.

We address the object-pair distribution shift in weakly-supervised VSGG by refining pseudo-labels with vision-language grounding and suppressing non-interactive pairs throughout relational reasoning.

Weakly-supervised video scene graph generation relies on off-the-shelf detectors that produce numerous interaction-irrelevant objects, creating a substantial object-pair distribution shift from fully-supervised settings. We address this shift at both the pseudo-label generation and relational reasoning stages. Relation-Aware Matching (RAM) uses vision-language grounding to refine ambiguous detection-triplet matching, while Pair Affinity Learning and Scoring (PALS) learns pair interactivity for prediction ranking and Pair Affinity Modulation (PAM) suppresses non-interactive pairs during spatial-temporal reasoning. Our framework consistently improves different baselines and backbones on Action Genome, achieving state-of-the-art weakly-supervised VSGG performance.

01

Overall framework

Conventional WS-VSGG trains only on class-matched pairs and discards unmatched pairs, even though both appear at inference. Our pipeline refines the matched partition and learns predicate classes together with pair affinity from matched and unmatched pairs.

Comparison between the conventional weakly-supervised VSGG pipeline and PAWS, which adds relation-aware matching and pair affinity learning and modulation.
PAWS turns discarded unmatched pairs into explicit supervision for pair interactivity.
02

Relation-Aware Matching

RAM converts each unlocalized triplet into a relation-aware text query and extracts cross-modal attention from a vision-language model. Reliable attention grounds the interaction-consistent instance; uncertain cases safely fall back to class-level matching.

Relation-Aware Matching uses a relation-aware text query, cross-attention, and reliability estimation to select grounded or class-level matching.
Cleaner matching reduces false positives caused by multiple same-class detections.
03

Pair Affinity Learning & Modulation

PALS learns whether a subject-object pair is interactive and uses that affinity in final triplet ranking. PAM carries the same signal inside spatial-temporal attention, preventing non-interactive pairs from diluting relational context.

Pair Affinity Learning predicts predicate class and pair affinity, and modulates spatial-temporal attention using affinity embeddings.
Pair affinity improves both output-level scoring and internal contextual reasoning.

Scene graph detection on Action Genome with the STTran backbone. Higher Recall@K is better.

88.3% of Fully supervised
With-Constraint R@10
94.3% of Fully supervised
No-Constraint R@10
Backbone Supervision Method With Constraint No Constraint
R@10R@20R@50 R@10R@20R@50
STTran [1] Fully Vanilla 25.2034.1037.00 24.6036.2048.80
Weakly PLA [2] 15.3921.4426.24 15.8322.8331.74
PLA + Ours 22.2426.4828.00 23.2030.2437.47
TRKT 15.1120.6326.02 15.7322.6431.23
TRKT + Ours 19.6324.2927.55 21.1327.9535.89

Compare scene graphs, frame by frame.

Select a clip and toggle any combination of ground truth and model predictions. Every panel shares the same decoded video frame, playback position, and seek control.

Predictions are produced on sparsely sampled frames and held until the next prediction timestamp.

0:00.00 0:10.00
Loading interactive demo…
@inproceedings{kang2026paws,
  title     = {Revisiting Weakly-Supervised Video Scene Graph Generation via Pair Affinity Learning},
  author    = {Kang, Minseok and Lee, Minhyeok and Kim, Minjung and Lee, Jungho and Kim, Donghyeong and Woo, Sungmin and Jeon, Inseok and Lee, Sangyoun},
  booktitle = {European Conference on Computer Vision},
  year      = {2026}
}
  1. Y. Cong, W. Liao, H. Ackermann, B. Rosenhahn, and M. Y. Yang, “Spatial-Temporal Transformer for Dynamic Scene Graph Generation,” ICCV, 2021.
  2. S. Chen, J. Xiao, and L. Chen, “Video Scene Graph Generation from Single-Frame Weak Supervision,” ICLR, 2023.