1 - Two key aspects of CL contribute to its effectiveness2 - To make a class prediction, we must extract the image logits and evaluate which class corresponds to the maximum.3 - Next, we can import a version of the clip model and its associated data processor. Note: the processor handles tokenizing input text and image preparation.4 - The basic idea behind using CLIP for 0-shot image classification is to pass an image into the model along with a set of possible class labels. Then, a classification can be made by evaluating which text input is most similar to the input image.5 - We can then match the best image to the input text by extracting the text logits and evaluating the image corresponding to the maximum.6 - The code for these examples is freely available on the GitHub repository.7 - We see that (again) the model nailed this simple example. But let's try some trickier examples.8 - Next, we'll preprocess the image/text inputs and pass them into the model.9 - Another practical application of models like CLIP is multimodal RAG, which consists of the automated retrieval of multimodal context to an LLM. In the next article of this series, we will see how this works under the hood and review a concrete example.10 - Another application of CLIP is essentially the inverse of Use Case 1. Rather than identifying which text label matches an input image, we can evaluate which image (in a set) best matches a text input (i.e. query)—in other words, performing a search over images.11 - This has sparked efforts toward expanding LLM functionality to include multiple modalities.12 - GPT-4o - Input: text, images, and audio. Output: text.FLUX - Input: text. Output: images.Suno - Input: text. Output: audio.13 - The standard approach to aligning disparate embedding spaces is contrastive learning (CL). A key intuition of CL is to represent different views of the same information similarly [5].14 - While the model is less confident about this prediction with a 54.64% probability, it correctly implies that the image is not a meme.15 - [8] Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities