Systems and methods for open vocabulary instance segmentation in unannotated images
Embodiments described herein provide an open-vocabulary instance segmentation framework that adopts a pre-trained vision-language model to develop a pipeline in detecting novel categories of instances.