Generate images from a prompt, edit them, or read what is inside one.
Visual instruction tuning towards large language and vision models with GPT-4 level capabilities