Upload one or many images and get captions from a fine-tuned BLIP model.
Supported: JPG, PNG, WebP, BMP. Inference at 224×224 (center-crop).