Scene Understanding · Video Temporal Grounding
Vision-Language Multimodal Learning

Minseok Kang

Integrated M.S./Ph.D. Student in Electrical and Electronic Engineering at Yonsei University, working on computer vision and machine/deep learning for video.

My research centers on scene understanding and vision-language multimodal learning, including video temporal grounding, video scene graph generation, and video large language models. I am interested in building models that reason about motion, long-range temporal structure, and language for real-world video understanding.

Minseok Kang profile photo
01

About

Researcher Profile

I am an Integrated M.S./Ph.D. student in Electrical and Electronic Engineering at Yonsei University, in the Image and Video Pattern Recognition Lab (MVP Lab), advised by Prof. Sangyoun Lee. I received my B.S. in Electrical and Electronic Engineering from Korea University.

My research lies in computer vision and machine/deep learning, with a focus on scene understanding, video temporal grounding, and video scene graph generation, as well as vision-language multimodal learning for video large language models.

I am interested in models that can understand dynamic visual content, reason about motion and long-range temporal structure, and connect vision with language for robust real-world video understanding.

Affiliation Yonsei University, MVP Lab
Advisor Prof. Sangyoun Lee
Program Integrated M.S./Ph.D. (2022 – 2027 expected)
Interests Scene Understanding, Video Temporal Grounding, Vision-Language Multimodal Learning
02

News

Recent Updates

Jun 2026

ECCV 2026 Accepted

First-author paper "PAWS" accepted to ECCV 2026.

May 2026

NeurIPS 2026 Submission

First-author paper "OTT-Vid" under review at NeurIPS 2026.

Feb 2026

CVPR Findings 2026 Accepted

Two co-authored papers accepted to CVPR Findings 2026.

Feb 2026

Pattern Recognition 2026 Accepted

Co-authored paper on zero-shot anomaly detection accepted to Pattern Recognition.

Oct 2025

3rd Place — Perception Test Challenge 2025

3rd place in Task 3 (Temporal Action & Sound Localization), hosted by DeepMind at the ICCV 2025 Workshop.

Sep 2025

NeurIPS 2025 Accepted

First-author paper "DualGround" accepted to NeurIPS 2025.

May 2025

ICIP 2025 Accepted

Co-authored paper "CMTM" accepted to ICIP 2025.

03

Publications

Selected Publications

SwiftVGGT: A Scalable Visual Geometry Grounded Transformer for Large-Scale Scenes

Jungho Lee, Minhyeok Lee, Sunghun Yang, Minseok Kang, Sangyoun Lee

CVPR Findings 2026 (Accepted)

Seen-to-Scene: Keep the Seen, Generate the Unseen for Video Outpainting

Inseok Jeon, Minhyeok Lee, Seunghoon Lee, Minseok Kang, Suhwan Cho, Sangyoun Lee

CVPR Findings 2026 (Accepted)

Generalizing CLIP Prompts for Zero-Shot Anomaly Detection

Donghyeong Kim, Chaewon Park, Suhwan Cho, Hyeonjeong Lim, Minseok Kang, Jungho Lee, Sangyoun Lee

Pattern Recognition 2026 (Accepted)

CMTM: Cross-Modal Token Modulation for Unsupervised Video Object Segmentation

Inseok Jeon, Suhwan Cho, Minhyeok Lee, Seunghoon Lee, Minseok Kang, Jungho Lee, Chaewon Park, Donghyeong Kim, Sangyoun Lee

ICIP 2025 (Accepted)

Multilevel Artificial Electronic Synaptic Device of Direct Grown Robust MoS2 Based Memristor Array for In-Memory Deep Neural Network

Muhammad Naqi*, Minseok Kang*, Na Liu*, Taehwan Kim, Seungho Baek, Arindam Bala, Changgyun Moon, Jongsun Park, Sunkook Kim  (* equal contribution)

npj 2D Materials and Applications, 2022 (Accepted)

04

Research

Research Interests

01

Scene Understanding

Building models that parse complex visual scenes, reason about objects and their relations, and capture structured semantics in video.

Topics: Video Scene Graph Generation, Structured Reasoning

02

Video Temporal Grounding

Localizing moments and actions in video given natural language, with a focus on phrase- and sentence-level temporal grounding.

Topics: Temporal Grounding, Moment Retrieval, Action Localization

03

Vision-Language Multimodal Learning

Connecting vision and language to improve understanding and reasoning over dynamic visual content.

Topics: Video Large Language Models, Cross-Modal Learning

04

Efficient Video Understanding

Designing token compression and efficient architectures that make large-scale video models practical for real-world deployment.

Topics: Token Compression, Optimal Transport, Efficient Modeling

05

Robust Perception

Developing recognition systems that remain reliable under low-visibility, adverse weather, and multi-sensor conditions.

Topics: Multi-Sensor Fusion, Robust Detection, Domain Generalization

06

Video Object Segmentation

Studying motion-guided and cross-modal methods for segmenting salient objects in video, including unsupervised settings.

Topics: Unsupervised VOS, Motion Cues, Cross-Modal Modulation

Contact

Let's Connect

I am always happy to discuss research and collaboration.
please feel free to get in touch.