Teaching an open model to read web pages
Multimodal models can read text and look at images, but they struggle with web pages: finding the right button, reading text inside a box, or guessing what a click will do. For our 11-777 (Multimodal Machine Learning) project, we tried to make a small open model, LLaVA-1.5-7B, better at these web tasks, and measured it on the VisualWebBench benchmark. This was fall 2024, when open models were far weaker at this than they are today.
- Action prediction, before
- 10.0%
- Action prediction, after
- 77.9%
The problem
VisualWebBench is 1,500 webpage screenshots from 139 real websites, with seven tasks: captioning a page, answering questions about it, reading the main heading, reading text inside a marked box (Element OCR), picking the right element from a description, choosing what to click for an instruction, and predicting which page a click leads to (Action Prediction).
Out of the box, LLaVA-1.5-7B averaged 13.07 across the seven tasks. Looking at its mistakes, it tended to focus on big, bold text in the middle of the page instead of the box it was asked about, and it often ignored the requested answer format.
What we did
First we ran baselines on several models (LLaVA, TroL, Phantom, InternLM and BLIP-2), including text-only and image-only versions, to see what each modality contributes.
Version 1 was fine-tuning. We trained LLaVA on a random subset of MultiUI, a large dataset of web-UI instructions, using LoRA (a cheap way to fine-tune by training small add-on matrices instead of the whole model): 12.5k samples for 5 epochs, then 2 more epochs with another 60k.
Version 2, which I led as an equal contributor, changed nothing in the weights. It changed what we send the model: task-specific prompts with examples and clear output rules (for example, answer options listed as A, B, C), and bounding-box tricks, first cropping the image to the marked box, then appending the box below the full page so the model sees both.
Results
Fine-tuning alone moved the average from 13.07 to 14.62. Adding the Version 2 prompts and preprocessing took it to 27.38.
The big jumps were in Action Prediction (9.96 to 77.94) and Element OCR (6.32 to 54.82). 77.94 was the highest Action Prediction score in our comparison table, above the closed models, including Gemini 1.5 Pro (74.4) and GPT-4 Vision (67.6).
What I took away
Cropping to the box helped in clear-cut cases, but it also raised a fair question: were we turning a multimodal task into a plain OCR task? Appending the box below the full page kept both the local and the page-wide context, and that fit the spirit of the benchmark better.
The gains were uneven. WebQA and captioning stayed well behind the best models, and hand-written prompts do not scale, so the report suggests automating prompt search as next steps.