Machine Learning · Spring 2026
A look at automatic image description using vision-language models: what works, what breaks, and what it means for web accessibility.
01 - Project Motivation & Introduction
A vast majority of images that can be found on the web contain no alt-text: the short written descriptions that screen readers use to convey image content to visually impaired users. The goal of this project is to explore whether a modern vision-language model can generate useful alt-text automatically, and where it falls short.
Alt-text is an HTML attribute that allows screen readers describe images to users who can't see them. Below is an example of this tag
Without it, images are completely silent: users hear nothing, or just a raw filename.
According to WebAIM's annual accessibility survey, missing or inadequate alt-text is one of the most common accessibility failures on the web. The scale of the problem is enormous and mostly unaddressed.
WCAG 2.1 (the Web Content Accessibility Guidelines) defines several principles for good alt-text:
Based from our list above the last point is where automated tools struggle most.
This project asks: can a modern vision-language model generate alt-text that's actually useful? And where does it fail, and for whom? Based on these goals this site will walk through the full project in four sections
02 - How the Model Works
For this project we will be using BLIP (known as Bootstrapping Language-Image Pre-training), it is a vision-language model that translates images into natural language descriptions.
Unlike traditional computer vision systems that classify objects, BLIP generates full sentences. This makes it well suited for alt-text, which requires not just recognition, but explanation.
BLIP uses two core components which are:
These components are trained together on large image-text datasets, allowing the model to learn how visual patterns correspond to language.
Conceptually, the model performs this pipeline:
Alt-text is not just labeling images, it needs the images to contain a descriptive meaning. Unlike classifiers that output single words, BLIP generates full sentences, making it better suited for accessibility tasks.
03 - Implementation
This project was implemented in Google Colab using the PyTorch library and a dataset from Hugging Face. The project pipeline includes dataset loading, preprocessing, model fine-tuning, and inference. To access this project click here: Colab Notebook
---For this project, the model was trained on an NVIDIA A100 GPU (provided on Google Colab Pro) to handle large batch sizes and accelerate training. Essentially, it is used to enhance speed.
CUDA available: True
GPU: NVIDIA A100-SXM4-80GB
---
The dataset that is used for project is Flickr30k with original alt-text, not generic captions. Each image contains multiple human-written descriptions.
raw = load_dataset("Mozilla/flickr30k-transformed-captions")
Each image contributes several training samples, significantly increasing dataset size.
---A custom PyTorch Dataset from Hugging Face was built to expand each image into multiple (image, caption) pairs.
for item in hf_dataset:
for cap in item["original_alt_text"]:
self.samples.append({"image": image, "caption": cap})
This is important: that the model learns from multiple descriptions of the same image, which improves robustness but also introduces variation.
---The images and text are then processed together using the BLIP processor.
encoding = processor(
images=image,
text=caption,
padding="max_length",
truncation=True,
return_tensors="pt"
)
Padding tokens are masked out during training to prevent them from affecting loss.
---The model was then fine-tuned using Hugging Face's Seq2SeqTrainer.
gradient_accumulation_steps = 8
fp16 = True
---
| Epoch | Train Loss | Validation Loss |
|---|---|---|
| 1 | 16.72 | 2.08 |
| 2 | 14.16 | 2.01 |
| 3 | 12.57 | 2.02 |
As we can see, the training loss appears to decrease steadily, while the validation loss stabilizes, thus indicating that the model learns general patterns but plateaus quickly.
---A reusable function was created to generate captions from:
out = model.generate(
**inputs,
max_new_tokens=50,
num_beams=4
)
The usage of beam search improves caption quality by exploring multiple possible outputs.
---The model captures the scene correctly but simplifies important details.
---This pipeline demonstrates that performance depends on more than the model: dataset structure, preprocessing, and training strategy all shape the output.
04 - Processing
Before reaching our model, the images are transformed into a fixed format. These preprocessing choices directly impact caption quality.
The model is highly sensitive to detail and context. When information is lost, captions become vague.
The model's output depends heavily on how images are prepared.
Lower detail reduces caption specificity.
Removing context leads to vague outputs.
05 - Evaluation
This reveals a key limitation: the model is optimized for describing images, not for communicating meaning to a user with specific needs.
Instead of just interacting with the final model, this section lets you understand how the system works and experiment with it directly.
In this short video, I walk through the deliverables of this project: motivation, model, implementation, and key findings of the project.
This notebook contains the full implementation: dataset loading, preprocessing, training, experiments, and evaluation.
Use this if you want to reproduce results or understand the full pipeline.
Want to try your own images or modify the model behavior? Use the sandbox notebook below.
This version removes most of the training code and focuses on inference, making it easier to test how the model responds to different images.