Viet-Anh on Software Logo

What is: OSCAR?

SourceOscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks
Year2000
Data SourceCC BY-SA - https://paperswithcode.com

OSCAR is a new learning method that uses object tags detected in images as anchor points to ease the learning of image-text alignment. The model take a triple as input (word-tag-region) and pre-trained with two losses (masked token loss over words and tags, and a contrastive loss between tags and others). OSCAR represents an image-text pair into semantic space via dictionary lookup. Object tags are used as anchor points to align image regions with word embeddings of pre-trained language models. The model is then fine-tuned for understanding and generation tasks.