Interactive Video Object Mask Annotation


  • Trung-Nghia Le National Institute of Informatics
  • Tam V. Nguyen University of Dayton
  • Quoc-Cuong Tran University of Science VNU-HCM
  • Lam Nguyen University of Science VNU-HCM
  • Trung-Hieu Hoang University of Science VNU-HCM
  • Minh-Quan Le University of Science VNU-HCM
  • Minh-Triet Tran University of Science VNU-HCM



Video Object Mask Annotation, Video Object Segmentation, Human-computer Interaction


In this paper, we introduce a practical system for interactive video object mask annotation, which can support multiple back-end methods. To demonstrate the generalization of our system, we introduce a novel approach for video object annotation. Our proposed system takes scribbles at a chosen key-frame from the end-users via a user-friendly interface and produces masks of corresponding objects at the key-frame via the Control-Point-based Scribbles-to-Mask (CPSM) module. The object masks at the key-frame are then propagated to other frames and refined through the Multi-Referenced Guided Segmentation (MRGS) module. Last but not least, the user can correct wrong segmentation at some frames, and the corrected mask is continuously propagated to other frames in the video via the MRGS to produce the object masks at all video frames.




How to Cite

Le, T.-N., Nguyen, T. V., Tran, Q.-C., Nguyen, L., Hoang, T.-H., Le, M.-Q., & Tran, M.-T. (2021). Interactive Video Object Mask Annotation. Proceedings of the AAAI Conference on Artificial Intelligence, 35(18), 16067-16070.