SwiftNet: Real-time Video Object Segmentation

Abstract

In this work we present SwiftNet for real-time semi-supervised video objectsegmentation (one-shot VOS), which reports 77.8% J&F and 70 FPS on DAVIS 2017validation dataset, leading all present solutions in overall accuracy and speedperformance. We achieve this by elaborately compressing spatiotemporalredundancy in matching-based VOS via Pixel-Adaptive Memory (PAM). Temporally,PAM adaptively triggers memory updates on frames where objects displaynoteworthy inter-frame variations. Spatially, PAM selectively performs memoryupdate and match on dynamic pixels while ignoring the static ones,significantly reducing redundant computations wasted on segmentation-irrelevantpixels. To promote efficient reference encoding, light-aggregation encoder isalso introduced in SwiftNet deploying reversed sub-pixel. We hope SwiftNetcould set a strong and efficient baseline for real-time VOS and facilitate itsapplication in mobile vision.

Quick Read (beta)

loading the full paper ...