This project uses the BLIP (Bootstrapping Language-Image Pretraining) model to generate natural language captions from images.
The project can be configured to use either BLIP Base or BLIP Large models.
- Python
- PyTorch
- Transformers
- Jupyter Notebook
- Install dependencies
- Open the notebook
- Run all cells
- Upload an image
- Get the generated caption