Zero-Shot Text-to-Image Generation

Text-to-image generation has traditionally focused on finding better modeling assumptions for training on a fixed dataset. These assumptions might involve complex architectures, auxiliary losses, or side information such as object part labels or segmentation masks supplied during training. We describe a simple approach for this task based on a transformer that autoregressively models the text and image tokens as a single stream of data. With sufficient data and scale, our approach is competitive with previous domain-specific models when evaluated in a zero-shot fashion.

ImageNet: A large-scalehierarchical image…ImageNet: A large-scale hierarchical image databaseAdam: A Method forStochastic OptimizationAdam: A Method for Stochastic OptimizationGenerative AdversarialText to Image SynthesisGenerative Adversarial Text to Image SynthesisThe ConcreteDistribution: A…The Concrete Distribution: A Continuous Relaxation of Discrete Random VariablesNeural DiscreteRepresentation LearningNeural Discrete Representation LearningAttention Is All YouNeedAttention Is All You NeedGANs Trained by a TwoTime-Scale Update Rule…GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash EquilibriumFixing Weight DecayRegularization in AdamFixing Weight Decay Regularization in AdamGenerative PretrainingFrom PixelsGenerative Pretraining From PixelsDF-GAN: Deep FusionGenerative Adversarial…DF-GAN: Deep Fusion Generative Adversarial Networks for Text-to-Image SynthesisGenerative AdversarialNetworksGenerative Adversarial NetworksLearning TransferableVisual Models From…Learning Transferable Visual Models From Natural Language SupervisionMultiscale VisionTransformersMultiscale Vision TransformersGODIVA: GeneratingOpen-DomaIn Videos from…GODIVA: Generating Open-DomaIn Videos from nAtural DescriptionsComputer-Aided Design asLanguageComputer-Aided Design as LanguageHigh-Resolution ImageSynthesis with Latent…High-Resolution Image Synthesis with Latent Diffusion ModelsCogView2: Faster andBetter Text-to-Image…CogView2: Faster and Better Text-to-Image Generation via Hierarchical TransformersReduce Information Lossin Transformers for…Reduce Information Loss in Transformers for Pluralistic Image InpaintingSimVLM: Simple VisualLanguage Model…SimVLM: Simple Visual Language Model Pretraining with Weak SupervisionMultimodal DialogueResponse GenerationMultimodal Dialogue Response GenerationGALIP: GenerativeAdversarial CLIPs for…GALIP: Generative Adversarial CLIPs for Text-to-Image SynthesisLayout-BridgingText-to-Image SynthesisLayout-Bridging Text-to-Image SynthesisGLIGEN: Open-SetGrounded Text-to-Image…GLIGEN: Open-Set Grounded Text-to-Image GenerationTraining DiffusionModels with…Training Diffusion Models with Reinforcement LearningZero-Shot Text-to-ImageGenerationZero-Shot Text-to-Image GenerationEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.