Scaling Open-Vocabulary Object Detection

Open-vocabulary object detection has benefited greatly from pretrained vision-language models, but is still limited by the amount of available detection training data. While detection training data can be expanded by using Web image-text pairs as weak supervision, this has not been done at scales comparable to image-level pretraining. Here, we scale up detection data with self-training, which uses an existing detector to generate pseudo-box annotations on image-text pairs. Major challenges in scaling self-training are the choice of label space, pseudo-annotation filtering, and training efficiency. We present the OWLv2 model and OWL-ST self-training recipe, which address these challenges. OWLv2 surpasses the performance of previous state-of-the-art open-vocabulary detectors already at comparable training scales (~10M examples). However, with OWL-ST, we can scale to over 1B examples, yielding further large improvement: With an L/14 architecture, OWL-ST improves AP on LVIS rare classes, for which the model has seen no human box annotations, from 31.2% to 44.6% (43% relative improvement). OWL-ST unlocks Web-scale training for open-world localization, similar to what has been seen for image classification and language modelling.

Adam: A Method forStochastic OptimizationAdam: A Method for Stochastic OptimizationEvaluatingLarge-Vocabulary Object…Evaluating Large-Vocabulary Object Detectors: The Devil is in the DetailsOpen-vocabulary ObjectDetection via Vision an…Open-vocabulary Object Detection via Vision and Language Knowledge DistillationDetectingTwenty-Thousand Classes…Detecting Twenty-Thousand Classes Using Image-Level SupervisionRegionCLIP: Region-basedLanguage-Image…RegionCLIP: Region-based Language-Image PretrainingExploring Plain VisionTransformer Backbones…Exploring Plain Vision Transformer Backbones for Object DetectionGrounded Language-ImagePre-trainingGrounded Language-Image Pre-trainingDetCLIPv2: ScalableOpen-Vocabulary Object…DetCLIPv2: Scalable Open-Vocabulary Object Detection Pre-training via Word-Region AlignmentSegment AnythingSegment AnythingThree ways to improvefeature alignment for…Three ways to improve feature alignment for open vocabulary detectionSigmoid Loss forLanguage Image…Sigmoid Loss for Language Image Pre-TrainingCombined Scaling forZero-shot Transfer…Combined Scaling for Zero-shot Transfer LearningDST-Det: Simple DynamicSelf-Training for…DST-Det: Simple Dynamic Self-Training for Open-Vocabulary Object DetectionYOLO-World: Real-TimeOpen-Vocabulary Object…YOLO-World: Real-Time Open-Vocabulary Object DetectionTokenize Anything viaPromptingTokenize Anything via PromptingTowards Open VocabularyLearning: A SurveyTowards Open Vocabulary Learning: A SurveySegment and CaptionAnythingSegment and Caption AnythingFlexCap: GeneratingRich, Localized, and…FlexCap: Generating Rich, Localized, and Flexible Captions in ImagesDetCLIPv3: TowardsVersatile Generative…DetCLIPv3: Towards Versatile Generative Open-Vocabulary Object DetectionHALC: ObjectHallucination Reduction…HALC: Object Hallucination Reduction via Adaptive Focal-Contrast DecodingSILC: Improving VisionLanguage Pretraining…SILC: Improving Vision Language Pretraining with Self-DistillationRevisiting Few-ShotObject Detection with…Revisiting Few-Shot Object Detection with Vision-Language ModelsSAM 3: Segment Anythingwith ConceptsSAM 3: Segment Anything with ConceptsRoboflow100-VL: AMulti-Domain Object…Roboflow100-VL: A Multi-Domain Object Detection Benchmark for Vision-Language ModelsScaling Open-VocabularyObject DetectionScaling Open-Vocabulary Object DetectionEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.